Public opinion text classification method and device, electronic equipment and computer program product
By using text analytical models, differentiated information screening and hash calculation methods in public opinion text classification, the problem of insufficient accuracy of public opinion text classification in the existing technology is solved, and more efficient and accurate public opinion text classification is achieved.
Patent Information
- Application Number
- CN202510215991.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The existing classification methods for public opinion texts are insufficient in terms of classification accuracy, and it is difficult to effectively deal with the complexity and diversity of public opinion texts.
The initial public opinion label is determined based on the text analysis model, the differentiated information between the initial label and the public opinion text is calculated, the candidate public opinion label is screened, and the target public opinion label is determined through hash calculation to achieve more accurate classification.
It improves the accuracy of the classification of public opinion texts, ensures that the classification results are more objective and reliable, and can more accurately reflect the real theme and category attributes of public opinion texts.
Smart Images

Figure CN120144769A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular, to a method, device, electronic device, and computer program product for classifying public opinion texts. Background Art
[0002] With the rapid development of the Internet, a huge amount of public opinion information is generated on the Internet every day. These public opinion texts cover all aspects of social life, such as product evaluations, discussions on social hot events, topics related to corporate images, etc. For various organizations, enterprises, and relevant regulatory departments, it is of great significance to classify these public opinion texts in a timely and accurate manner and grasp their key information. For example, financial institutions can grasp the loan approval situation based on the public opinion generated by a certain enterprise.
[0003] However, most traditional public opinion classification methods are based on a simple keyword matching mechanism. A fixed keyword dictionary is constructed in advance. When a specific keyword appears in a public opinion text, it is classified into the corresponding category. In addition, some classification methods based on a single machine learning model can, to a certain extent, learn the semantic features of the text. However, due to the complexity and diversity of different public opinion texts, the model often has difficulty achieving an ideal classification effect for all types of texts.
[0004] In summary, the existing public opinion text classification methods still have deficiencies in terms of classification accuracy. Summary of the Invention
[0005] This application provides a method, device, electronic device, and computer program product for classifying public opinion texts, which can improve the accuracy of public opinion text classification.
[0006] In a first aspect, this application provides a method for classifying public opinion texts, including:
[0007] Determining a plurality of initial public opinion labels of the current public opinion text based on a text parsing model;
[0008] Calculating the differential information between each of the initial public opinion labels and the current public opinion text, and determining at least two candidate public opinion labels according to each of the differential information;
[0009] Performing a hash calculation on at least two of the candidate public opinion labels and the current public opinion text respectively to obtain a target public opinion label, so as to classify the current public opinion text according to the target public opinion label.
[0010] In a second aspect, this application provides a device for classifying public opinion texts, and the device includes:
[0011] An initial label determination module, configured to determine a plurality of initial public opinion labels of the current public opinion text based on a text parsing model;
[0012] A difference information determination module, configured to calculate the difference information between each of the initial public opinion tags and the current public opinion text respectively, and determine at least two candidate public opinion tags according to each of the difference information;
[0013] A target tag determination module, configured to perform hash calculation on at least two of the candidate public opinion tags and the current public opinion text respectively to obtain a target public opinion tag, so as to classify the current public opinion text according to the target public opinion tag.
[0014] In a third aspect, the present application further provides an electronic device, where the electronic device includes:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the public opinion text classification method according to any embodiment of the present application.
[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the public opinion text classification method according to any embodiment of the present application when executed by a processor.
[0019] In a fifth aspect, the present application further provides a computer program product, including a computer program, where the computer program implements the public opinion text classification method according to any embodiment of the present application when executed by a processor.
[0020] The public opinion text classification solution provided by the embodiments of the present application first determines multiple initial public opinion tags of the current public opinion text based on a text parsing model, so as to initially summarize the characteristics of the public opinion text through the text semantic understanding ability of the text parsing model; then further screens and determines candidate public opinion tags by calculating the difference information between the initial public opinion tags and the public opinion text, so as to screen out a set of tags more closely related to the current text from the initial public opinion tags, laying a foundation for subsequent classification; finally, a hash calculation method is used to determine a more accurate target public opinion tag from the candidate public opinion tags, and the public opinion text is accurately classified based on the target public opinion tag, solving the problem of poor classification accuracy in the existing solution and achieving the beneficial effect of improving the classification accuracy of public opinion text.
[0021] It should be noted that the above computer instructions can be stored in whole or in part on a computer-readable storage medium. Among them, the computer-readable storage medium can be packaged together with the processor of the public opinion text classification device, or can be separately packaged from the processor of the public opinion text classification device. This application does not make any limitations in this regard.
[0022] For the descriptions of the second, third, fourth, and fifth aspects in this application, reference can be made to the detailed description of the first aspect; and for the beneficial effects of the descriptions of the second, third, fourth, and fifth aspects, reference can be made to the analysis of the beneficial effects of the first aspect, which will not be elaborated here.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understood through the following description.
[0024] It can be understood that before using the technical solutions disclosed in the embodiments of this application, the types, usage scopes, usage scenarios, etc. of the personal information involved in this application should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 is a flowchart of a public opinion text classification method provided by an embodiment of this application;
[0027] Figure 2 is another flowchart of a public opinion text classification method provided by an embodiment of this application;
[0028] Figure 3 is a structural schematic diagram of a public opinion text classification device provided by an embodiment of this application;
[0029] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] To enable those skilled in the art to better understand the solution of this application, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in this embodiment. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] The following further elaborates on this application in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described here are only used to explain this application, rather than limiting this application. Additionally, it should be noted that for the sake of description, only the parts related to this application rather than all the structures are shown in the drawings.
[0033] Figure 1 FIG. 10 is a schematic flowchart of a public opinion text classification method provided by an embodiment of this application. This embodiment is applicable to the situation of classifying the collected public opinion texts. This method can be executed by a public opinion text classification device, which can be implemented in the form of hardware and / or software and integrated in an electronic device that executes this method. Preferably, the electronic device in the embodiment of this application can be a server or a computer device, etc.
[0034] Refer to Figure 1 , the public opinion text classification method of this embodiment includes but is not limited to the following steps:
[0035] S110. Determine multiple initial public opinion tags of the current public opinion text based on a text parsing model.
[0036] The text parsing model used in this embodiment is a model obtained in advance based on deep learning algorithms. This text parsing model can perform in-depth analysis and understanding of the input text content to extract semantic information, structural information, etc. contained in the text; thereby extracting tags corresponding to key elements, text themes, sentiment tendencies, etc. from the text, and initial public opinion tags reflecting the characteristics of public opinion texts from different perspectives are mined. For example, multiple initial public opinion tags such as "management change", "rating adjustment", "earnings warning", and "service quality" are obtained through the analysis of public opinion texts in the financial field. By analyzing multiple initial public opinion tags, basic analysis support is provided for subsequent public opinion-related processing to achieve the purpose of accurately classifying the current public opinion text.
[0037] Optionally, before determining multiple initial public opinion tags of the current public opinion text based on the text parsing model, preprocessing of the current public opinion text is required, including but not limited to cleaning the public opinion text, removing noise information such as redundant punctuation and special characters. The preprocessed public opinion text is conducive to analysis; for languages that require word segmentation such as Chinese, appropriate word segmentation tools are used for word segmentation processing, and the word segmentation results are spliced into strings according to the format required by the model. The specific process of preprocessing the public opinion text is not limited here.
[0038] The preprocessed current public opinion text is encoded into the input format corresponding to the text parsing model and input into the model to obtain the output result. In this embodiment, the text parsing model can generate initial public opinion tags for the current public opinion text in various ways. For example, based on the keyword extraction method, analyze the words corresponding to the hidden state output by the model, and find the word combinations with higher weights and key to semantics to form initial public opinion tags; or based on the topic classification method, select initial public opinion tags from the preset tag library of the corresponding category according to the output classification probability. The specific way to obtain the initial public opinion tags is not limited here.
[0039] In this embodiment, through the powerful semantic understanding ability of the text parsing model, the deep semantic information of the public opinion text is fully mined, avoiding the limitations brought by relying only on surface keyword matching, generating multiple initial public opinion tags that summarize the text features from different perspectives, laying a foundation for more accurate subsequent classification, and improving the comprehensiveness and accuracy of the overall classification.
[0040] S120. Calculate the differential information between each initial public opinion tag and the current public opinion text respectively, and determine at least two candidate public opinion tags according to each differential information.
[0041] The candidate public opinion tags in this embodiment are some tags selected based on the differential information between the initial public opinion tags and the public opinion text. Compared with the initial public opinion tags, the candidate public opinion tags are more matched with the current public opinion text in terms of semantics, relevance, etc., and the number is at least two, serving as an intermediate transition set for further determining the target public opinion tags for classification through hash calculation.
[0042] The differential information is the relevant data index used to measure the degree of difference between the initial public opinion tags and the current public opinion text in terms of semantics and content association. The specific way to obtain the differential information between each initial public opinion tag and the current public opinion text can be achieved through methods such as word vector similarity or calculating the edit distance, etc. Based on this, candidate public opinion tags that are more in line with the current public opinion text are selected from multiple initial public opinion tags.
[0043] By quantifying the degree of difference between the initial public opinion tags and the public opinion text in this embodiment, tags that are more relevant to the text and have more fitting semantics can be selected as candidate public opinion tags, narrowing the scope of subsequent further screening, removing the initial public opinion tags that are not very relevant to the actual content of the text, improving the efficiency and accuracy of subsequent determination of the target public opinion tags, and making the classification more focused on the theme truly expressed by the text.
[0044] S130. Perform hash calculation on at least two candidate public opinion tags and the current public opinion text respectively to obtain the target public opinion tags, and classify the current public opinion text according to the target public opinion tags.
[0045] The target public opinion tags are the final tags determined by further analyzing the candidate public opinion tags. The number of target public opinion tags is not less than one, and they can accurately represent the key attributes such as the category and theme of the current public opinion text, so as to accurately classify the public opinion text according to the target public opinion tags, making it fall into the corresponding category range, facilitating subsequent public opinion analysis, statistics and other work.
[0046] Specifically, the method for obtaining the target public opinion label by performing hash calculation on at least two candidate public opinion labels can be as follows: perform hash calculation on at least two candidate public opinion labels respectively to obtain at least two first hash values; perform hash calculation on the current public opinion text to obtain a second hash value; calculate the correlation between each first hash value and the second hash value, and determine the candidate public opinion labels that meet the preset rules in the correlation as the target public opinion labels. When classifying the current public opinion text into the corresponding category according to the target public opinion label, for example, classify the text corresponding to the "product quality problem" label into the "product quality category", etc. In this embodiment, the candidate public opinion labels and the public opinion text are converted into hash values with a fixed length through hash calculation, and the target public opinion label that can best represent the public opinion text can be efficiently and accurately determined from the candidate public opinion labels by calculating the correlation, making the classification of public opinion text more objective and accurate. At the same time, the hash calculation itself has characteristics such as certain irreversibility and uniqueness, enhancing the reliability of the classification result and helping to carry out precise analysis, statistics, etc. for different categories of public opinion text in the future.
[0047] The public opinion text classification method provided in this embodiment first determines multiple initial public opinion labels of the current public opinion text based on the text parsing model to initially summarize the characteristics of the public opinion text through the ability of the text parsing model to understand text semantics; then further screens and determines the candidate public opinion labels by calculating the differential information between the initial public opinion labels and the public opinion text, so as to screen out a set of labels that are more closely related to the current text from the initial public opinion labels and lay a foundation for subsequent classification; finally, uses the method of hash calculation to determine more accurate target public opinion labels from the candidate public opinion labels, so as to accurately classify the public opinion text based on the target public opinion labels. This solves the problem of poor classification accuracy in the existing solutions and achieves the beneficial effect of improving the classification accuracy of public opinion text.
[0048] Figure 2 This is another process schematic diagram of the public opinion text classification method provided in the embodiments of the present application. The embodiments of the present application are optimized based on the above embodiments, specifically optimized as follows: This embodiment provides a detailed explanation of the implementation process of "determining multiple initial public opinion labels of the current public opinion text based on the text parsing model", the implementation process of "calculating the differential information between each initial public opinion label and the current public opinion text respectively, and determining at least two candidate public opinion labels according to each differential information", and the implementation process of "performing hash calculation on at least two candidate public opinion labels and the current public opinion text respectively to obtain the target public opinion label".
[0049] See Figure 2 This embodiment's public opinion text classification method includes but is not limited to the following steps:
[0050] S210. Perform word segmentation processing on the current public opinion text to obtain multiple feature word vectors.
[0051] Since public opinion texts are usually complete reports described in natural language, it is difficult to accurately grasp various semantic elements contained therein by directly processing the whole of them. Through word segmentation processing, the text can be split into independent and meaningful individual words. For example, for a public opinion text like "The camera function of this mobile phone is very good, but the battery life is a bit poor", after word segmentation, words such as "this", "mobile phone", "camera function", "very", "good", "but", "battery life", "a bit poor" are obtained, enabling subsequent analysis to be carried out on these more refined units, mining the information conveyed by each component of the text, avoiding the loss of detailed content when the text is regarded as a whole, and helping to more comprehensively and deeply understand the theme, viewpoints, etc. expressed in the public opinion text.
[0052] For a public opinion text, different words have different weights and roles in reflecting the core features of the text. In the embodiment, by performing word segmentation processing on the current public opinion text and obtaining multiple feature word vectors based on the text after word segmentation, the relevant attributes of the vectors can be used to more conveniently extract key features. For example, by analyzing the numerical magnitudes of each dimension of the word vectors, the combination relationships between the vectors, etc., some word features that are of great significance for text theme judgment and classification can be discovered. In public opinion analysis, if it is a public opinion text about product quality, the feature word vectors corresponding to words such as "quality", "defect", and "fault" may have prominent performances in certain dimensions, helping to focus on these key semantic elements, so as to more accurately classify, label, etc. the text in subsequent operations, and improving the accuracy of natural language processing tasks.
[0053] In this embodiment, by performing word segmentation processing on the public opinion text to disassemble the text into meaningful words one by one, the constituent elements of the text can be analyzed more meticulously, avoiding semantic details that may be lost when the text is processed as a whole. Then, the words are converted into word vectors, enabling the text to participate in subsequent calculations and model analysis in the form of mathematical vectors, facilitating the computer to understand and process the semantic information of the text, and laying a foundation for accurately mining the features of the public opinion text and subsequent determination of the initial public opinion labels.
[0054] S211. For the pre-constructed label database, determine the label encoding corresponding to each public opinion label in the label database.
[0055] The data label library in this embodiment is to collect and sort out various labels related to public opinion through manual or automated means to construct a label database. These labels can cover different fields, themes, natures, etc., such as "product quality problems", "after-sales service complaints", and "social hot events", etc. These labels can be stored in a suitable data structure.
[0056] Further, in order to facilitate subsequent processing and comparison in the model, it is necessary to encode each public opinion label. A simple way is to use label encoding, which assigns a unique numerical identifier to each label in a certain order. For example, for a label database containing three labels: "Product quality problem" is encoded as 1, "After-sales service complaint" is encoded as 2, "Social hot event" is encoded as 3, etc.; optionally, other more complex and semantic information-rich encoding methods can also be selected according to actual needs, such as using hash encoding, vector encoding, etc., as long as it can ensure that each label has a unique and distinguishable encoding representation. The specific method for determining the label encoding corresponding to each public opinion label is not limited here.
[0057] In this embodiment, by encoding the public opinion labels, the labels originally in text form can participate in subsequent operations such as matching with text feature word vectors in a standardized and computer-friendly form, simplifying the comparison and calculation process, improving the processing efficiency of the entire process, and also facilitating quick search and positioning in large-scale label data.
[0058] S212. Input multiple feature word vectors into the text parsing model so that the text parsing model determines the initial public opinion labels according to the label encodings corresponding to each feature word vector and each public opinion label respectively.
[0059] Organize and input the multiple feature word vectors obtained previously in the format required by the text parsing model. The model matches with the label encodings corresponding to each public opinion label through a certain calculation method. For example, a certain similarity related to the model output features and each label encoding can be calculated, and the public opinion label corresponding to the label encoding with the highest similarity is selected as the initial public opinion label.
[0060] The method for determining the initial public opinion labels provided in this embodiment utilizes the semantic understanding ability of the text parsing model and combines the matching mechanism between feature word vectors and label encodings, and can more accurately screen out the initial public opinion labels that best fit the current public opinion text from the pre-constructed label database. This model-based method can better capture complex information such as the semantics and context of the text, improve the accuracy and comprehensiveness of determining the initial public opinion labels, contribute to enhancing the quality and efficiency of public opinion analysis work, and more accurately grasp the key information and category situation conveyed by the public opinion text.
[0061] A preferred implementation. In this embodiment, the specific implementation of "determining the initial public opinion label according to the label codes respectively corresponding to each feature word vector and each public opinion label" in step S212 above can be: for each feature word vector, determine the similarity values between the feature word vector and each public opinion label; determine the public opinion labels that meet the preset screening rules from the multiple similarity values as the initial public opinion labels. The advantage of determining the initial public opinion labels based on the preset screening rules in this embodiment is that it fully considers the degree of association between the semantic features of the words in the text and different public opinion labels. Compared with the method of simply relying on keyword matching or manual subjective judgment, this method based on vector similarity can measure the fit between the text and the label more comprehensively and objectively, so as to screen out the initial public opinion labels that more conform to the actual expression content of the public opinion text.
[0062] In this embodiment, the way of determining the public opinion labels that meet the preset screening rules as the initial public opinion labels can be: preset a similarity threshold in advance (such as 0.2), and as long as the label probability output by the model is greater than the similarity threshold, determine this public opinion label as the initial public opinion label; another way can be: sort all labels according to the similarity from high to low, and then select the top N labels as the initial public opinion labels. For example, select the first 20 public opinion labels in the sequential order as the initial public opinion labels, etc., to provide a basis for further analyzing and processing the public opinion text. The specific implementation of the preset screening rules in this embodiment is not limited here.
[0063] S220. Calculate the edit distance between the current initial public opinion label and the body information in the public opinion text, and obtain the differential information according to the edit distance.
[0064] The edit distance is an index used to measure the degree of difference between two strings. It can reflect the minimum number of edit operations required to convert one string into the other. These edit operations include: insertion, deletion, and replacement, etc. For the obtained multiple initial public opinion labels, calculate the edit distance between the current initial public opinion label and the body information of the public opinion text through a function, and store the label and its corresponding edit distance in a dictionary, which is used as the representation form of the differential information.
[0065] S221. Determine the initial public opinion labels that exceed the preset distance threshold in the differential information as the candidate public opinion labels.
[0066] Traverse the differential information stored in the previously constructed dictionary, and judge whether the edit distance corresponding to each label exceeds the preset distance threshold. If it exceeds, determine this label as the candidate public opinion label, and these candidate public opinion labels can be stored in a new list; finally, the candidate public opinion labels stored in the new list are the screened candidate public opinion labels.
[0067] Among them, the above preset distance threshold can be determined according to experience, multiple tests or specific requirements of the business scenario, and is used to screen candidate public opinion tags subsequently. For example, the set preset distance threshold is 3, etc. The determination of the specific preset distance threshold is not limited here.
[0068] In this embodiment, by calculating the edit distance to measure the degree of difference between the initial public opinion tag and the text information of the public opinion text to determine the candidate public opinion tags, the entire screening process no longer depends on subjective judgments or relatively rough methods such as simple keyword matching. There is a clear numerical measurement basis for the difference between each tag and the text. By setting a reasonable preset distance threshold, candidate public opinion tags that meet the requirements can be screened out relatively objectively, improving the scientificity and accuracy of the tag screening link and laying a good foundation for subsequent more accurate public opinion text classification, analysis and other work.
[0069] S230. Perform hash calculations on at least two candidate public opinion tags respectively to obtain at least two first hash values.
[0070] In this embodiment, to facilitate finding the correspondence between each candidate public opinion tag and the corresponding first hash value, the candidate public opinion tag list can be traversed, and a suitable hash algorithm is used to perform hash calculations on each candidate public opinion tag, and the calculated hash values are stored in a dictionary. In the current dictionary, the key is the candidate public opinion tag, and the value is the first hash value corresponding to the candidate public opinion tag.
[0071] S231. Perform a hash calculation on the current public opinion text to obtain a second hash value.
[0072] When performing a hash calculation on the current public opinion text, the same method of storing in a structured dictionary is also used to obtain its corresponding second hash value.
[0073] S232. Calculate the correlation between each first hash value and the second hash value respectively, and determine the candidate public opinion tags that meet the preset rules in the correlation as the target public opinion tags, so as to classify the current public opinion text according to the target public opinion tags.
[0074] In this embodiment, the way to measure the correlation can be to measure the number of different characters at the corresponding positions of each first hash value and the second hash value respectively. For example, the correlation between them can be judged by comparing the Hamming distance between the first hash value and the second hash value. The specific method can be as follows: Obtain the first hash value corresponding to each candidate public opinion label, calculate the Hamming distance from the second hash value of the current public opinion text, and store each candidate public opinion label and its corresponding Hamming distance in the first dictionary structure. Then, according to the preset rules, select the candidate public opinion labels with the Hamming distance less than or equal to the threshold from the first dictionary structure as the target public opinion labels, and store these target public opinion labels in the second dictionary structure; finally, what is stored in the second dictionary structure is the determined target public opinion labels. For example, in the above example, the obtained target public opinion labels are "price fluctuation problem", "poor after-sales service", etc. Then, according to the semantics of the target public opinion labels, in accordance with the pre-set classification rules, such as "price fluctuation problem" corresponding to "price category", "poor after-sales service" corresponding to "service category", etc., classify the current public opinion text and divide it into the corresponding categories.
[0075] In this embodiment, the candidate public opinion labels and the public opinion text are converted into hash values of a fixed length through hash calculation, and then the target public opinion labels are determined based on the correlation between the hash values. This method avoids the ambiguity and subjectivity that may exist when directly comparing text semantics. It can more accurately screen out the public opinion labels that are closely related to the current public opinion text in essence, thereby improving the accuracy of classifying the public opinion text, making the classification results more reliable, and better reflecting the true theme and category attribution of the public opinion text.
[0076] In another preferred embodiment, before determining multiple initial public opinion labels of the current public opinion text based on the text parsing model, the solution provided in this embodiment further includes: obtaining the attribute information of the public opinion text; judging whether at least one classification dimension can be determined according to the attribute information; if at least one classification dimension can be determined according to the attribute information, then perform the operation of determining multiple initial public opinion labels of the current public opinion text based on the text parsing model; if at least one classification dimension cannot be determined according to the attribute information, then determine that the public opinion text is an invalid text.
[0077] The attribute information in this embodiment can be the release source, release time, and text language of the public opinion text, etc. Among them, the release source can be from the source of obtaining public opinion information, such as the release sources of specific websites such as enterprises, finance, and current affairs; the release time is the release time of the public opinion text; the text language is used to detect whether the public opinion text belongs to several pre-concerned languages (such as Chinese, English, etc.), which is convenient for separately processing public opinion texts in different languages. The specific content included in the specific attribute information is not limited here.
[0078] In this embodiment, for any public opinion text, if it cannot be attributed to any of the above-listed dimensions (source of publication, publication time, and text language), it is determined that the current public opinion text is an invalid text and filtered out.
[0079] In this embodiment, by pre-judging whether the attribute information of the public opinion text can determine the classification dimension, texts that are obviously unable to be effectively classified and lack key information are screened out in advance, avoiding subsequent complex and possibly fruitless text parsing operations on these invalid texts, saving computing resources and time, and improving the overall processing efficiency.
[0080] In another preferred embodiment, after obtaining the target public opinion label, the solution provided in this embodiment can also perform the following operations: obtain the public opinion propagation parameters of the current public opinion text, where the public opinion propagation parameters at least include text discussion duration, text discussion volume, text propagation volume, and propagation speed; draw a public opinion map based on the text discussion duration, text discussion volume, text propagation volume, and propagation speed.
[0081] The above text discussion duration can be calculated by taking the time when the relevant discussion first appeared as the starting point and the current time as the end point to calculate the time difference; the text discussion volume can be obtained by summarizing the data related to the discussion obtained from various data sources, such as the number of comments on social media platforms and the number of messages on news websites, etc., as the text discussion volume; the text propagation speed can be obtained by integrating and calculating the data obtained from different channels, comprehensively considering data indicators reflecting the propagation range such as reposts, shares, and views; the propagation speed can be roughly estimated by dividing the text propagation volume by the text discussion duration.
[0082] The public opinion map in this embodiment includes, but is not limited to, a bar chart that can intuitively compare the numerical sizes of the parameters of text discussion duration, discussion volume, propagation volume, and propagation speed; a line chart that continuously monitors the changes in its propagation parameters within different time periods and shows the dynamic changes of these parameters over time; and a heat map that combines geographical information (such as the public opinion text propagation parameters corresponding to different cities and regions) and shows the distribution heat of the parameters geographically. The specific types included in the public opinion map are not limited here.
[0083] In this embodiment, by presenting the abstract public opinion propagation parameters in a visual map form, it is convenient for public opinion analysts to clarify the overall development trend of the public opinion. For example, by comparing the sizes of each parameter through a bar chart, it can be quickly understood which aspect of the public opinion text is prominent during the propagation process; the line chart showing the change trend over time can be used to know whether the public opinion is in an upward, stable, or downward stage, facilitating timely grasping of the public opinion dynamics and making corresponding decisions.
[0084] The public opinion text classification method provided in this embodiment first utilizes the semantic understanding ability of the text parsing model and combines the matching mechanism of feature word vectors and label encodings to accurately screen out the initial public opinion labels that best match the current public opinion text from the pre-constructed label database. Then, by calculating the edit distance to measure the degree of difference between the initial public opinion labels and the body information of the public opinion text to determine the candidate public opinion labels, the entire screening process no longer relies on subjective judgments or relatively rough methods such as simple keyword matching, improving the scientificity and accuracy of the label screening link and laying a good foundation for subsequent more accurate public opinion text classification, analysis, etc. Finally, by performing hash calculations on the candidate public opinion labels and the public opinion text to convert them into hash values of a fixed length, and then determining the target public opinion labels based on the correlation between the hash values, it is possible to more accurately screen out the public opinion labels that are closely related to the current public opinion text in essence, thereby improving the accuracy of public opinion text classification, making the classification results more reliable, and better reflecting the true theme and category attribution of the public opinion text.
[0085] Figure 3 FIG. is a structural schematic diagram of a public opinion text classification device provided in an embodiment of the present application, and this device is applicable to execute the public opinion text classification method provided in the embodiment of the present application. As Figure 3 shown, this device may specifically include: an initial label determination module 310, a difference information determination module 320, and a target label determination module 330, where:
[0086] The initial label determination module 310 is configured to determine multiple initial public opinion labels of the current public opinion text based on the text parsing model;
[0087] The difference information determination module 320 is configured to calculate the differential information between each of the initial public opinion labels and the current public opinion text, and determine at least two candidate public opinion labels according to each of the differential information;
[0088] The target label determination module 330 is configured to perform hash calculations on at least two of the candidate public opinion labels and the current public opinion text respectively to obtain target public opinion labels, so as to classify the current public opinion text according to the target public opinion labels.
[0089] The public opinion text classification device provided in this embodiment first determines multiple initial public opinion tags for the current public opinion text based on a text parsing model, so as to preliminarily summarize the characteristics of the public opinion text through the text parsing model's ability to understand text semantics; then further screens and determines candidate public opinion tags by calculating the differential information between the initial public opinion tags and the public opinion text, so as to screen out a set of tags that are more closely related to the current text from the initial public opinion tags and lay a foundation for subsequent classification; finally, uses the method of hash calculation to determine more accurate target public opinion tags from the candidate public opinion tags, and classifies the public opinion text based on the target public opinion tags, which solves the problem of poor classification accuracy in the existing solutions and achieves the beneficial effect of improving the classification accuracy of public opinion texts.
[0090] In one embodiment, the initial tag determination module 310 is specifically configured to perform word segmentation processing on the current public opinion text to obtain multiple feature word vectors; for a pre-constructed tag database, determine the tag encoding corresponding to each public opinion tag in the tag database; input the multiple feature word vectors into the text parsing model, so that the text parsing model determines the initial public opinion tags according to the tag encoding corresponding to each feature word vector and each public opinion tag respectively.
[0091] In one embodiment, the initial tag determination module 310 is specifically further configured to, for each of the feature word vectors, determine the similarity value between the feature word vector and each public opinion tag; determine the public opinion tags that meet the preset screening rules from the multiple similarity values as the initial public opinion tags.
[0092] In one embodiment, the differential information determination module 320 is specifically configured to calculate the edit distance between the current initial public opinion tag and the body information in the public opinion text, and obtain the differential information according to the edit distance; determine the initial public opinion tags whose differential information exceeds the preset distance threshold as the candidate public opinion tags.
[0093] In one embodiment, the target tag determination module 330 is specifically configured to perform hash calculation on at least two of the candidate public opinion tags respectively to obtain at least two first hash values; perform hash calculation on the current public opinion text to obtain a second hash value; calculate the correlation between each of the first hash values and the second hash value, and determine the candidate public opinion tags that meet the preset rules in the correlation as the target public opinion tags.
[0094] In one embodiment, the device further includes an attribute information acquisition module and a classification dimension determination module, where:
[0095] The attribute information acquisition module is used to acquire the attribute information of the public opinion text;
[0096] A classification dimension determination module, configured to determine whether at least one classification dimension can be determined based on the attribute information; if at least one classification dimension can be determined based on the attribute information, then perform an operation of determining multiple initial public opinion tags of the current public opinion text based on a text parsing model; if at least one classification dimension cannot be determined based on the attribute information, then determine that the public opinion text is an invalid text.
[0097] In one embodiment, the apparatus further includes a propagation parameter acquisition module and a public opinion map drawing module, where:
[0098] The propagation parameter acquisition module is configured to acquire the public opinion propagation parameters of the current public opinion text, and the public opinion propagation parameters at least include the text discussion duration, the text discussion volume, the text propagation volume, and the propagation speed;
[0099] The public opinion map drawing module is configured to draw a public opinion map according to the text discussion duration, the text discussion volume, the text propagation volume, and the propagation speed.
[0100] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used as an example. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described functional modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0101] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards of the relevant regions.
[0102] An embodiment of the present application further provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the public opinion text classification method according to any embodiment of the present application.
[0103] An embodiment of the present application further provides a computer-readable medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the public opinion text classification method according to any embodiment of the present application when executed.
[0104] Next, refer to Figure 4 ,Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. It shows a schematic structural diagram of a computer system 500 of the electronic device suitable for implementing the embodiment of the present application. Figure 4 The shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiment of the present application.
[0105] As Figure 4 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0106] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required so that a computer program read from it can be installed into the storage section 508 as required.
[0107] Particularly, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0108] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, and optical cable, etc., or any suitable combination of the above.
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0110] The modules and / or units involved in the embodiments described in this application can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes an initial tag determination module, a difference information determination module, and a target tag determination module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases.
[0111] As another aspect, the present application also provides a computer-readable medium. This computer-readable medium can be included in the device described in the above embodiments; it can also exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by such a device, the device includes: determining a plurality of initial public opinion tags of the current public opinion text based on a text parsing model; calculating the differential information between each of the initial public opinion tags and the current public opinion text, and determining at least two candidate public opinion tags according to each differential information; performing hash calculation on at least two of the candidate public opinion tags and the current public opinion text respectively to obtain target public opinion tags, so as to classify the current public opinion text according to the target public opinion tags.
[0112] According to the technical solution of this embodiment, first, a plurality of initial public opinion tags of the current public opinion text are determined based on the text parsing model to preliminarily summarize the characteristics of the public opinion text through the text parsing model's ability to understand text semantics; then, the candidate public opinion tags are further screened and determined by calculating the differential information between the initial public opinion tags and the public opinion text, so as to screen out a set of tags that are more closely related to the current text from the initial public opinion tags, laying a foundation for subsequent classification; finally, the target public opinion tags are accurately determined from the candidate public opinion tags by means of hash calculation, and the public opinion text is classified based on the target public opinion tags, solving the problem of poor classification accuracy in the existing solutions and achieving the beneficial effect of improving the classification accuracy of public opinion texts.
[0113] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A method for classifying public opinion texts, characterized in that: include: Determine multiple initial public opinion labels of the current public opinion text based on the text parsing model; Calculate the difference information between each of the initial public opinion labels and the current public opinion text, and determine at least two candidate public opinion labels according to each of the difference information; Hash calculations are performed on at least two of the candidate public opinion labels and the current public opinion text respectively to obtain a target public opinion label, so as to classify the current public opinion text according to the target public opinion label.
2. The method for classifying public opinion text according to claim 1, characterized in that: The method of determining a plurality of initial public opinion labels of the current public opinion text based on the text parsing model includes: Performing word segmentation processing on the current public opinion text to obtain multiple feature word vectors; For a pre-built tag database, determine the tag code corresponding to each public opinion tag in the tag database; Input a plurality of the feature word vectors into the text parsing model, so that the text parsing model determines the initial public opinion label according to the label code corresponding to each of the feature word vectors and each of the public opinion labels.
3. The method for classifying public opinion text according to claim 2, characterized in that: The determining the initial public opinion label according to the label code corresponding to each of the feature word vectors and each of the public opinion labels respectively includes: For each of the feature word vectors, determining a similarity value between the feature word vector and each of the public opinion labels; A public opinion label that satisfies a preset screening rule is determined from the multiple similarity values as the initial public opinion label.
4. The method for classifying public opinion text according to claim 1, characterized in that: The step of calculating the difference information between each of the initial public opinion labels and the current public opinion text, and determining at least two candidate public opinion labels according to each of the difference information, includes: Calculating the edit distance between the current initial public opinion label and the main text information in the public opinion text, and obtaining the differential information according to the edit distance; The initial public opinion labels in the differentiated information that exceed a preset distance threshold are determined as the candidate public opinion labels.
5. The method for classifying public opinion text according to claim 1, characterized in that: The step of performing hash calculation on at least two candidate public opinion labels and the current public opinion text to obtain a target public opinion label includes: Performing hash calculations on at least two of the candidate public opinion labels respectively to obtain at least two first hash values; Performing hash calculation on the current public opinion text to obtain a second hash value; The correlation between each of the first hash values and the second hash value is calculated, and the candidate public opinion label that meets the preset rule in the correlation is determined as the target public opinion label.
6. The method for classifying public opinion text according to claim 1, characterized in that: Before determining multiple initial public opinion labels of the current public opinion text based on the text parsing model, it also includes: Obtaining attribute information of the public opinion text; Determining whether at least one classification dimension can be determined based on the attribute information; If at least one classification dimension can be determined according to the attribute information, an operation of determining a plurality of initial public opinion labels of the current public opinion text based on the text parsing model is performed; If at least one classification dimension cannot be determined based on the attribute information, the public opinion text is determined to be an invalid text.
7. The method for classifying public opinion text according to claim 1, characterized in that: The method further comprises: Obtaining public opinion propagation parameters of the current public opinion text, wherein the public opinion propagation parameters at least include text discussion duration, text discussion volume, text propagation volume, and propagation speed; A public opinion map is drawn according to the text discussion duration, the text discussion volume, the text dissemination volume and the dissemination speed.
8. A public opinion text classification device, characterized in that: include: An initial label determination module is used to determine multiple initial public opinion labels of the current public opinion text based on a text parsing model; A difference information determination module is used to calculate the difference information between each of the initial public opinion labels and the current public opinion text, and determine at least two candidate public opinion labels according to each of the difference information; The target label determination module is used to perform hash calculations on at least two of the candidate public opinion labels and the current public opinion text respectively to obtain a target public opinion label, so as to classify the current public opinion text according to the target public opinion label.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the public opinion text classification method described in any one of claims 1-7.
10. A computer program product, comprising a computer program, which, when executed by a processor, implements the public opinion text classification method according to any one of claims 1-7.
Citation Information
Cited By
Traffic public opinion risk label classification early warning method based on large model
CN121071617A