Method, device, equipment and storage medium for generating intelligent guidance strategy for network public opinion based on large model
Through the big model generation method, network public opinion data is collected and analyzed, text clustering and emotional recognition are carried out, keyword and event relationships are identified, and guidance strategies are corrected, which solves the problem of lack of targetedness in network public opinion guidance and achieves accurate network public opinion guidance.
Patent Information
- Application Number
- CN202510647710.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing technology lacks targetedness in online public opinion guidance, and fails to deeply analyze the basis behind each specific speech by netizens, resulting in poor guidance effect.
Through the big model generation method, the initial viewpoint text of the target event is collected, text clustering and sentiment analysis are performed, keyword and event relationships are identified, and guidance strategies are corrected to improve accuracy.
It has achieved accurate guidance on online public opinion, timely discover and correct wrong views, optimize guidance strategies, and improve guidance effect.
Smart Images

Figure CN120162438B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for generating an intelligent guidance strategy for online public opinion based on a large model. Background Art
[0002] In today's digital age, the internet has become a vital platform for people to express and exchange views, making the guidance of online public opinion crucial. However, existing technologies for guiding online public opinion tend to operate at a macro level, releasing general information or viewpoints in an attempt to steer public opinion. These technologies fail to deeply analyze the motivations, rationales, and misconceptions behind each individual netizen's comments, nor do they provide detailed, targeted supplements and corrections to these specific comments. Consequently, this lack of targeted supplementation leads to poor results in guiding online public opinion, and even worse, can lead to further escalation of online incidents. Summary of the Invention
[0003] The main purpose of the embodiments of the present invention is to provide a method, device, equipment and storage medium for generating an intelligent network public opinion guidance strategy based on a large model, aiming to solve the problem in related technologies of attempting to guide public opinion through general information or opinions without in-depth analysis of the basis behind each specific speech of netizens, which leads to poor results in guiding network public opinion.
[0004] In a first aspect, an embodiment of the present invention provides a method for generating an intelligent network public opinion guidance strategy based on a large model, comprising:
[0005] Determine the target event and collect the target user's initial opinion text on the target event;
[0006] Obtaining the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and performing text clustering according to the opinion text vector and the target keywords to obtain a text clustering result;
[0007] Performing sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster;
[0008] Obtaining relevant keywords of the first sub-category cluster from the target keywords, and determining a target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster;
[0009] Performing event relationship recognition on the target viewpoint text and the target event to obtain a relationship type between the target viewpoint text and the target event;
[0010] Determining an initial guidance strategy for the target event according to the target emotion type and the relationship type;
[0011] Obtaining the latest opinion text of the target event under the initial guidance strategy, and determining a guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result;
[0012] The initial guidance strategy is modified according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event.
[0013] In a second aspect, an embodiment of the present invention provides a device for generating an intelligent network public opinion guidance strategy based on a large model, comprising:
[0014] A data collection module is used to determine a target event and collect the target user's initial opinion text on the target event;
[0015] A text clustering module, configured to obtain, based on the large model, an opinion text vector and target keywords corresponding to the initial opinion text, and perform text clustering based on the opinion text vector and the target keywords to obtain a text clustering result;
[0016] A sentiment analysis module, configured to perform sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster;
[0017] a text recognition module, configured to obtain relevant keywords of the first sub-category cluster from the target keywords, and determine a target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster;
[0018] A relationship identification module, configured to perform event relationship identification on the target opinion text and the target event to obtain a relationship type between the target opinion text and the target event;
[0019] A strategy determination module, configured to determine an initial guidance strategy for the target event according to the target emotion type and the relationship type;
[0020] A strategy evaluation module, configured to obtain the latest opinion text of the target event under the initial guidance strategy, and determine a guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result;
[0021] A strategy modification module is used to modify the initial guidance strategy according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event.
[0022] In the third aspect, an embodiment of the present invention also provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of any one of the methods for generating intelligent guidance strategies for network public opinion based on a large model provided in the specification of the present invention.
[0023] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any step of the method for generating an intelligent guidance strategy for network public opinion based on a large model as provided in the specification of the present invention.
[0024] The embodiment of the present invention provides a method, device, equipment and storage medium for generating an intelligent guidance strategy for network public opinion based on a large model. The method includes: determining a target event, and collecting the target user's initial opinion text on the target event, and then obtaining the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and performing text clustering according to the opinion text vector and the target keyword to obtain a text clustering result, thereby performing sentiment analysis on each first subclass cluster in the text clustering result to obtain the target sentiment type of the first subclass cluster; obtaining relevant keywords of the first subclass cluster from the target keywords, and determining the target opinion text of the first subclass cluster according to the relevant keywords and the first subclass cluster, and then performing event relationship recognition between the target opinion text and the target event. By obtaining the relationship type between the target opinion text and the target event, it is possible to determine whether there is a causal relationship or correlation relationship between the target opinion text and the target event, thereby determining the erroneous or abnormal opinions among the target users based on the target emotion type and relationship type, and thus obtaining the initial guidance strategy for the target event; and then obtaining the latest opinion text of the target event under the initial guidance strategy, and determining the guidance effect representation value of the initial guidance strategy based on the latest opinion text and text clustering results, thereby timely evaluating the effectiveness of the initial guidance strategy through the guidance effect representation value, and then correcting the initial guidance strategy based on the guidance effect representation value, and obtaining the target guidance strategy corresponding to the target event, thereby continuously optimizing the guidance strategy and improving the accuracy and effectiveness of guidance. Therefore, this method can timely discover whether there is a correlation between the target user's target opinion text on the target event, and then timely determine the target user's erroneous opinion information based on the target emotion type and relationship type. Once the target user's erroneous opinion information is determined, the initial guidance strategy corresponding to the target event is determined, and then the latest opinion text of the target event under the initial guidance strategy is obtained. The guidance effect representation value of the initial guidance strategy is determined based on the latest opinion text and text clustering results, so that the effectiveness of the initial guidance strategy is timely evaluated through the guidance effect representation value, and then the initial guidance strategy is corrected based on the guidance effect representation value to obtain the target guidance strategy corresponding to the target event. In this way, the guidance strategy can be continuously optimized to effectively improve the accuracy and effectiveness of guidance. This method also solves the problem in related technologies that attempts to guide public opinion through general information or opinions, and does not deeply analyze the basis behind each specific speech of netizens, which leads to poor results in guiding online public opinion. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] Figure 1 A flowchart of a method for generating an intelligent network public opinion guidance strategy based on a large model provided by an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of the module structure of a device for generating an intelligent network public opinion guidance strategy based on a large model provided by an embodiment of the present invention;
[0028] Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0031] It should be understood that the terms used in this specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0032] Embodiments of the present invention provide a method, apparatus, device, and storage medium for generating a large-scale model-based intelligent strategy for guiding online public opinion. The large-scale model-based intelligent strategy for guiding online public opinion can be applied to a terminal device, such as a tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device. The terminal device can also be a server or a server cluster.
[0033] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0034] Please refer to Figure 1 , Figure 1 A flowchart of a method for generating an intelligent network public opinion guidance strategy based on a large model provided by an embodiment of the present invention.
[0035] like Figure 1 As shown, the method for generating an intelligent guidance strategy for online public opinion based on a large model includes steps S101 to S108.
[0036] Step S101: determine a target event and collect the target user's initial opinion text on the target event.
[0037] For example, a target event is determined based on trending social topics, and then initial opinions of target users on the target event are obtained through web crawling on major social media platforms. Target users are those who have expressed their opinions or views on major social media platforms. The initial opinions are the target users' evaluations or opinions on the target event.
[0038] Step S102: obtaining the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and performing text clustering according to the opinion text vector and the target keywords to obtain a text clustering result.
[0039] For example, a large model refers to a large language model. The large model has powerful semantic understanding and feature extraction capabilities and can process various types of text. During the actual collection process, the initial opinion text may contain a large number of irrelevant characters, symbols, and stop words. For example, HTML tags on web pages, unnecessary punctuation, modal particles, etc. These elements are of no substantial help in subsequent text analysis. Instead, they increase processing complexity and affect the accuracy of the analysis results. Therefore, the large language model uses specific rules and algorithms to remove this irrelevant content, making the text more concise and pure.
[0040] For example, after completing text cleaning, the large language model will use a keyword extraction algorithm to identify target keywords in the initial opinion text. After determining the target keywords, the large language model will further obtain the word representation vectors corresponding to these target keywords. The large language model then fuses the word representation vectors corresponding to the target keywords to obtain the opinion text vector corresponding to the initial opinion text. For example, different weights are assigned to each target keyword based on its importance in the initial opinion text, and then these vectors are combined according to the weights to obtain an opinion text vector that can represent the semantic characteristics of the entire initial opinion text.
[0041] Exemplarily, the first distance between any two initial opinion texts is calculated using cosine similarity based on the opinion text vector, and then the target keywords in any two initial opinion texts are compared, the number of identical keywords is counted or the intersection ratio of the keywords is calculated to obtain the second distance between the two initial opinion texts.
[0042] Exemplarily, the weights corresponding to the first distance and the second distance are determined based on expert experience or historical experience. For example, if more emphasis is placed on the overall semantics of the text, a higher weight can be assigned to the first distance; if the matching of the target keyword is considered more critical, a higher weight is assigned to the second distance. Thus, for any two initial opinion texts, their first distance is multiplied by the corresponding weight, and their second distance is multiplied by the corresponding weight, and then the two results are added together to obtain the fused target distance information.
[0043] Exemplarily, the initial opinion text is clustered according to a clustering algorithm combined with target distance information to obtain a text clustering result.
[0044] In some embodiments, obtaining target keywords corresponding to the initial opinion text according to the large model includes: obtaining relevant news information corresponding to the target event, and obtaining event keywords involved in the target event from the relevant news information; obtaining a first text vector corresponding to the event keyword according to the large model; performing text preprocessing on the initial opinion text to obtain the initial keywords corresponding to the initial opinion text and the first score corresponding to the initial keywords, and obtaining a second text vector corresponding to the initial keywords according to the large model; determining the associated keywords related to the target event in the initial keywords according to the first text vector and the second text vector; performing word co-occurrence analysis based on the associated keywords in combination with the initial opinion text to obtain the co-occurrence keywords corresponding to the associated keywords and the co-occurrence frequency between the associated keywords and the co-occurrence keywords; establishing a word co-occurrence graph corresponding to the initial opinion text based on the associated keywords and the co-occurrence keywords in combination with the co-occurrence frequency; The associated keywords are removed from the initial keywords to obtain candidate keywords corresponding to the initial opinion text, and the second score corresponding to the candidate keywords is obtained according to the first score corresponding to the initial keywords; the connection keywords corresponding to the candidate keywords are obtained from the word co-occurrence graph, and the associated edge information corresponding to the candidate keywords and the connection keywords is determined; the external keywords corresponding to the connection keywords are obtained from the word co-occurrence graph, and the connection density between the candidate keywords and the connection keywords is determined according to the external keywords and the connection keywords; the frequency information of the candidate keywords in the initial opinion text is obtained, and the target score corresponding to the candidate keywords is determined according to the frequency information, the second score and the connection density; the selected keywords are obtained from the candidate keywords according to the target score, and the target keywords corresponding to the initial opinion text are determined according to the selected keywords and the associated keywords; wherein the target score is obtained according to the following formula:
[0045] ;
[0046] in, represents the target score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the second score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the jth connection keyword corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the connection density between the i-th candidate keyword and the j-th connection keyword corresponding to the m-th initial opinion text, max represents obtaining the maximum value, lg represents a logarithmic function with a base of 10, num represents the number of texts corresponding to all the initial opinion texts, represents the frequency information corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text in the t-th initial opinion text, Indicates the number of occurrences of the i-th candidate keyword corresponding to the m-th initial opinion text in all the initial opinion texts.
[0047] Exemplarily, relevant news information corresponding to the target event is collected from official media or platforms, where the relevant news information is authoritative news information released by official media or official platforms, and then the collected relevant news information is processed to obtain event keywords of the core content of the target event.
[0048] Exemplarily, a large model such as GPT-3 is used to input the extracted event keywords into the selected large model, and the embedding function of the large model is used to convert each event keyword into a corresponding vector representation, that is, the first text vector.
[0049] For example, the initial opinion text is cleaned to remove special characters, punctuation, and stop words, and lexical analysis (e.g., word segmentation) is performed. Keyword extraction algorithms such as TextRank are then used to identify the initial keywords corresponding to the cleaned initial opinion text. A first score is calculated for each initial keyword, which reflects its importance within the initial opinion text. The extracted initial keywords are then input into the previously selected large model to obtain a second text vector corresponding to each initial keyword.
[0050] Exemplarily, the similarity between the event keyword and the initial keyword is calculated based on the cosine similarity in combination with the first text vector and the second text vector, so as to filter out the associated keywords related to the target event from the initial keywords based on the similarity.
[0051] For example, the initial opinion text is processed to count the co-occurrences of associated keywords with other words in the initial opinion text. Two words are considered co-occurring when they appear together within a specific window (e.g., a sentence or paragraph). In this way, other keywords that co-occur with each associated keyword are identified, i.e., co-occurring keywords. The number of times each associated keyword appears with the co-occurring keyword is counted to obtain the co-occurrence frequency between them.
[0052] For example, a word co-occurrence graph is constructed using related keywords and co-occurring keywords as nodes and the co-occurrence relationships between them as edges. Each node in the graph represents a keyword, and the edge weights can be set to the corresponding co-occurrence frequency. In this way, the word co-occurrence graph can intuitively display the associations between keywords.
[0053] Exemplarily, related keywords are removed from the initial keywords, and the remaining keywords are candidate keywords, and then a second score corresponding to each candidate keyword is obtained from the first score.
[0054] For example, in the word co-occurrence graph, the connected keywords corresponding to each candidate keyword are searched, that is, other keywords that are connected to the candidate keyword by edges, and then the external keywords corresponding to the connected keywords are obtained from the word co-occurrence graph, that is, the keywords connected to the connected keywords are obtained from the word co-occurrence graph, so as to count the number of external keywords and the number of connected keywords, and then the minimum value between the number of external keywords and the number of connected keywords is determined as the connection density between the candidate keyword and the connected keyword; the connection density is used to represent the core degree of the candidate keyword in the word co-occurrence graph.
[0055] For example, Indicates obtaining the jth connection keyword of the i-th candidate keyword corresponding to the m-th initial opinion text Then, calculate the connection keywords The corresponding external keywords, thus counting the connection keywords The corresponding number of external keywords, and then the connected keywords The minimum value between the number of corresponding external keywords and the number of connection keywords corresponding to the i-th candidate keyword is determined as , thereby obtaining the connection density corresponding to each connection keyword in the i-th candidate keyword corresponding to the m-th initial opinion text , and then connect the corresponding keywords in the i-th candidate keyword corresponding to the m-th initial opinion text The maximum value in is determined as .
[0056] For example, the frequency information of each candidate keyword appearing in the initial opinion text is counted, and the target score is determined by combining the frequency information, the second score, and the connection density according to the following formula:
[0057] ;
[0058] in, represents the target score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the second score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, Indicates the jth connection keyword corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the connection density between the i-th candidate keyword and the j-th connection keyword corresponding to the m-th initial opinion text, max represents obtaining the maximum value, lg represents the logarithmic function with a base of 10, and num represents the number of texts corresponding to all initial opinion texts. Indicates the frequency information of the i-th candidate keyword corresponding to the m-th initial opinion text in the t-th initial opinion text, It represents the number of occurrences of the i-th candidate keyword corresponding to the m-th initial opinion text in all initial opinion texts.
[0059] For example, candidate keywords are sorted according to the target scores, and a portion of candidate keywords with higher target scores are selected as selected keywords. The selected keywords are combined with the previously determined related keywords to obtain target keywords corresponding to the initial opinion text. These target keywords can better represent the association information between the initial opinion text and the target event.
[0060] In some embodiments, the text clustering based on the opinion text vector and the target keyword to obtain the text clustering result includes: obtaining a first keyword and a second keyword from the target keyword, and calculating the degree of connection between the first keyword and the second keyword to obtain a connection value between the first keyword and the second keyword; determining the number of clusters corresponding to the text clustering result and the initial cluster center corresponding to the text clustering result according to the connection value; calculating the initial probability that the opinion text vector belongs to the initial cluster center, and determining the data distribution information corresponding to the opinion text vector under the initial cluster center according to the initial probability; determining distribution difference information according to the initial probability and the data distribution information, and determining the associated cluster center corresponding to the opinion text vector under the initial cluster center according to the distribution difference information; determining the initial clustering result corresponding to the initial cluster center according to the opinion text vector and the associated cluster center; updating the initial cluster center according to the initial clustering result to obtain the latest cluster center; re-clustering the text according to the latest cluster center and the opinion text vector until the latest cluster center no longer changes, thereby obtaining the text clustering result.
[0061] Exemplarily, a first keyword and a second keyword are randomly selected from the target keyword, and then the first number of times the first keyword and the second keyword appear simultaneously in the initial opinion text is counted, and then the second number of times the first keyword appears in the initial opinion text and the third number of times the second keyword appears in the initial opinion text are obtained, so as to determine the adjustment factor, and then the first number and the adjustment factor are added to obtain a first value, and then the second number and the third number are added to obtain a second value, and then the first value and the second value are divided to obtain a third value, and then the logarithm of the third value with base 10 is taken to obtain a fourth value, and then the logarithm of the first value with base 10 is taken to obtain a fifth value, and then the fourth value is divided by the fifth value to obtain the connection value between the first keyword and the second keyword, and the connection value is used to characterize the degree of connection between the first keyword and the second keyword.
[0062] For example, a cohesion value is calculated between any two target keywords from all target keywords corresponding to all initial opinion texts. This cohesion value can reflect the closeness of the association between any two keywords. Next, a keyword is randomly selected from the target keywords as the current keyword. Based on the previously calculated cohesion value, the target keywords are divided into two categories: one category is keywords that are associated with the current keyword, and the other category is keywords that are not associated with the current keyword. Subsequently, a keyword is randomly selected from the keywords that are not associated with the current keyword and set as the new current keyword. Similarly, based on the cohesion value, keywords that are not associated with the new current keyword are found from the keywords that are not associated with the current keyword. This process is repeated continuously, continuously using the newly determined current keywords as the basis for association judgment and screening based on the cohesion value. Ultimately, the associated keywords are grouped together, thereby obtaining multiple clusters, each of which has a certain degree of association between the keywords. The number of clusters corresponding to the text clustering results is thus determined based on the number of clusters of associated keywords.
[0063] Exemplarily, the word representation value corresponding to each word in the associated keywords is obtained according to the large model, and then all the word representation values of the associated keywords are fused to obtain the initial cluster center corresponding to the cluster of the associated keywords.
[0064] For example, a distance metric (such as Euclidean distance or cosine similarity) is used to calculate the similarity between each opinion text vector and the initial cluster center. The higher the similarity, the greater the probability of belonging to the initial cluster center. For example, cosine similarity is used to calculate the similarity between the opinion text vector and the initial cluster center, and then the similarity is normalized to obtain the initial probability.
[0065] Exemplarily, when obtaining the distribution information that the mth opinion text vector belongs to the kth initial cluster center, the first probability corresponding to the mth opinion text vector belonging to the kth initial cluster center is obtained from the initial probability, and then the second probability corresponding to all opinion text vectors belonging to the kth initial cluster center is obtained, and the second probabilities are summed to obtain a first sum value, and then the first probability is squared and divided by the first sum value to obtain the target value, and then all the target values are normalized to obtain the data distribution information corresponding to the opinion text vector under the initial cluster center.
[0066] Exemplarily, the KL divergence is used to determine the distribution difference information using the initial probability and data distribution information, so that when the distribution difference information is minimized, the initial cluster center is determined as the associated cluster center corresponding to the opinion text vector. The associated cluster center is also the center to which the opinion text vector belongs among all the initial cluster centers.
[0067] For example, based on the opinion text vectors and the associated cluster centers, each opinion text vector is assigned to the cluster where the corresponding associated cluster center is located to obtain the initial clustering result. That is, if the associated cluster center of a certain opinion text vector is the i-th initial cluster center, then the vector is classified as the i-th cluster.
[0068] Exemplarily, after obtaining the initial clustering results, the mean vector of all opinion text vectors in each cluster is calculated and used as the new cluster center to obtain the latest cluster center. The latest cluster center and the opinion text vector are then used to re-cluster the text, and the steps of obtaining the initial clustering results are repeated until the latest cluster center no longer changes or the change is less than the set threshold, at which time the final text clustering result is obtained.
[0069] Specifically, by utilizing the degree of connection between keywords to determine the number of clusters and the initial clustering centers, the semantic associations between texts can be better captured, avoiding the local optimal problem that may be caused by the random selection of initial clustering centers in traditional clustering methods, thereby improving the accuracy of clustering. In addition, in the clustering process, the data distribution information and distribution difference information of the opinion text vectors under different clustering centers are taken into account, and the cluster to which each vector belongs can be adaptively adjusted, so that the clustering results are more in line with the actual distribution of the data.
[0070] In some embodiments, the updating of the initial cluster center according to the initial clustering result to obtain the latest cluster center includes: obtaining the first distance information between any two first sub-text vectors in each second sub-cluster in the initial clustering result, and determining the first average distance information corresponding to the second sub-cluster according to the first distance information; obtaining the second distance information between the second sub-text vector in the second sub-cluster and the remaining sub-text vectors in the second sub-cluster after excluding the second sub-text vector; determining the first similar data information corresponding to the second sub-text vector in the second sub-cluster according to the second distance information and the first average distance information; determining the similar internal distance corresponding to the second sub-text vector according to the first similar data information and the second sub-text vector; obtaining a third sub-cluster after excluding the second sub-cluster in the initial clustering result, and obtaining The third distance information between any two third sub-text vectors in the third sub-class cluster is obtained, and the second average distance information corresponding to the third sub-class cluster is determined based on the third distance information; the fourth distance information between the fourth sub-text vector in the third sub-class cluster and the remaining sub-text vectors in the third sub-class cluster after excluding the third sub-text vector is obtained; the second similar data information corresponding to the fourth sub-text vector in the third sub-class cluster is determined based on the fourth distance information and the second average distance information; the number of data corresponding to the second similar data information is obtained, and the associated cluster corresponding to the second sub-class cluster is determined based on the number of data; the weight center representation value corresponding to the second sub-text vector is determined based on the associated cluster, the similar internal distance and the first similar data information; the initial cluster center is updated based on the weight center representation value to obtain the latest cluster center.
[0071] For example, in each second sub-cluster of the initial clustering results, the distance between any two first sub-text vectors is calculated using Euclidean distance or cosine distance to obtain first distance information. All first distance information within the second sub-cluster is then summed and divided by the number of first distance information to obtain the first average distance information corresponding to the second sub-cluster. This value can reflect the overall dispersion of the text vectors within the sub-cluster.
[0072] Exemplarily, for any second sub-text vector in the second sub-class cluster, the distance between it and the remaining sub-text vectors in the sub-class cluster after excluding itself is calculated to obtain second distance information.
[0073] For example, the second distance information is compared with the first average distance information. If a second distance information is less than the first average distance, the corresponding sub-text vector is considered to be similar to the second sub-text vector. All similar sub-text vectors are counted to obtain the first similar data information corresponding to the second sub-text vector in the second sub-cluster.
[0074] For example, the average distance between the subtext vector and the second subtext vector included in the first similar data information is calculated to obtain the corresponding internal distance of the second subtext vector. The internal distance can measure the closeness of the second subtext vector within its subclass cluster.
[0075] For example, the second subclass cluster is excluded from the initial clustering result to obtain the third subclass cluster. In the third subclass cluster, the distance between any two third subtext vectors is calculated to obtain the third distance information. The third distance information is summed and divided by its number to obtain the second average distance information corresponding to the third subclass cluster. For a fourth subtext vector in the third subclass cluster, the distance between it and the remaining subtext vectors after excluding itself in the subclass cluster is calculated to obtain the fourth distance information. The fourth distance information is compared with the second average distance information, and the subtext vectors whose fourth distance is less than the second average distance are counted to obtain the second similar data information corresponding to the fourth subtext vector in the third subclass cluster. The number of data corresponding to the second similar data information is counted, and the third subclass cluster with the largest number of data is the associated cluster corresponding to the second subclass cluster. This indicates that the second subclass cluster and the associated cluster may be closer in data distribution.
[0076] Exemplarily, the maximum number of data of the second similar data information corresponding to the fourth sub-text vector is found in the associated cluster. Subsequently, the data corresponding to the fourth sub-text vector with the maximum number of data in the associated cluster is determined as the core data of the third sub-class cluster. Next, the distance information between the fourth sub-text vector and the second sub-text vector corresponding to the core data is calculated. At the same time, the quantity information corresponding to the first similar data information is obtained. Afterwards, this quantity information is multiplied by the distance information between the fourth sub-text vector and the second sub-text vector, and then multiplied by the inverse of the similar internal distance, and finally the weight center representation value corresponding to the second sub-text vector is obtained. Simply put, this weight center representation value not only takes into account the situation of the second sub-text vector itself, but also takes into account the cluster distance between the second sub-class cluster and the third sub-class cluster in which it is located, as well as the relevant information within the second sub-class cluster.
[0077] Exemplarily, the sub-text vector corresponding to the maximum value of the weight center representation value of the second sub-text vector is determined as the latest cluster center corresponding to the sub-class cluster, or all sub-text vectors of the sub-class cluster are weighted averaged according to the weight center representation value to obtain the latest cluster center corresponding to the sub-class cluster.
[0078] Specifically, the similarity of text vectors within sub-clusters and the correlation between sub-clusters are taken into account to avoid unstable clustering results caused by local fluctuations in the data. When updating the cluster center, multiple factors are integrated into the calculation to increase the reliability of the cluster center update.
[0079] Step S103: Perform sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster.
[0080] For example, a model such as a traditional model based on machine learning or a model based on deep learning such as a convolutional neural network or a recurrent neural network can be selected, and then a text dataset containing labels of different emotion types is prepared, and the model is trained to obtain a sentiment classification model.
[0081] For example, each sub-opinion text in the first sub-cluster is preprocessed to meet the input requirements of the sentiment classification model. The preprocessing steps may include text cleaning (removing special characters, stop words, etc.), word segmentation, word vectorization, and other operations. The preprocessed sub-opinion text is input into the trained sentiment classification model. The model predicts the sentiment type of each sub-opinion text based on its learned patterns and features and outputs the corresponding initial sentiment type. Common sentiment types include positive, negative, neutral, etc. Specific classification labels can be defined according to task requirements.
[0082] For example, the initial sentiment types of all sub-opinion texts in the first sub-cluster are counted, and the number of occurrences of each sentiment type is recorded. Based on the statistical results, the sentiment type with the highest number of occurrences is selected as the target sentiment type for the first sub-cluster. For example, if the number of sub-opinion texts with a negative sentiment type is the highest, then the target sentiment type for the first sub-cluster is determined to be negative.
[0083] Step S104: Obtain relevant keywords of the first sub-category cluster from the target keywords, and determine the target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster.
[0084] Exemplarily, the target keywords are all keywords corresponding to all the initial opinion texts, and then the relevant keywords contained in each first sub-category cluster are obtained from the target keywords.
[0085] Exemplarily, similarity calculation is performed on each sub-opinion text in the first sub-category cluster based on relevant keywords to obtain the target opinion text of the first sub-category cluster, that is, the target opinion text that best represents the opinion of the first sub-category cluster is obtained from the sub-opinion texts of the first sub-category cluster based on relevant keywords.
[0086] Step S105: performing event relationship identification on the target opinion text and the target event to obtain the relationship type between the target opinion text and the target event.
[0087] For example, a large amount of text data containing logical relationships and not containing logical relationships is obtained, and these data are divided into different types such as logically coherent and illogical to obtain training data, and then the training data is used to train the pre-trained language model to obtain a logical classification model.
[0088] For example, a series of logical terms, such as "because...so..." and "if...then...", are determined to connect the target event and the target opinion text. The target event and the target opinion text are then connected using the selected logical terms to generate multiple different connective sentences. For example, if the target event is "it rained today" and the target opinion text is "the streets are clean," a connective sentence such as "because it rained today, so the streets are clean" can be generated. This connective sentence is then input into a trained logical classification model. The model predicts the type of connective sentence based on its learned patterns and features, determining whether the connective sentence is logically coherent or illogical.
[0089] For example, the type of connection sentence determines the type of relationship between the target viewpoint text and the target event. If the model determines that the connection sentence is logically incoherent, then it can be concluded that there is no correlation between the target viewpoint text and the target event. If the connection sentence is logically coherent, then it indicates that there is a correlation between the target viewpoint text and the target event.
[0090] In some embodiments, the performing event relationship identification on the target opinion text and the target event to obtain the relationship type between the target opinion text and the target event includes: performing event identification on the target opinion text to obtain a first opinion event corresponding to the target opinion text and a first event type corresponding to the first opinion event; determining a second event type corresponding to the target event, and performing event alignment on the first opinion event and the target event according to the first event type and the second event type to obtain a second opinion event corresponding to the first event type and a target sub-event corresponding to the second event type; obtaining an event descriptor corresponding to the target sub-event, and obtaining associated news information corresponding to the target sub-event; obtaining, from the associated news information according to the event descriptor, relevant description text of the target sub-event in the associated news information and a text correlation between the event description text and the relevant description text; obtaining opinion description text corresponding to the second opinion event from the target opinion text, and calculating the text similarity between the opinion description text and the relevant description text; determining the event correlation between the second opinion event and the target sub-event according to the text correlation and the text similarity; and determining the relationship type between the second opinion event and the target sub-event according to the event correlation.
[0091] Exemplarily, the rule-based method identifies event-related information from the target opinion text according to predefined language rules and patterns, and then extracts specific event content, namely the first opinion event, from the target opinion text, and classifies it according to the subject of the event, etc., to determine the first event type, which is the event subject corresponding to the first opinion event.
[0092] Exemplarily, it is determined that the classification method of the corresponding second event type is consistent with the first event type, so that the first viewpoint event and the target event are matched and aligned according to the first event type and the second event type. The second viewpoint event related to the target event under the first event type and the corresponding target sub-event under the second event type are found. For example, if the first event type and the second event type are both "sports events", the specific description of the sports event in the target viewpoint text is used as the second viewpoint event, and the relevant event links in the target event are used as target sub-events. That is, the target viewpoint text and the target event are further refined according to the first event type and the second event type, so as to provide a good basis for the subsequent relationship type determination.
[0093] For example, keywords or phrases that can summarize the main features and content of the target sub-event are extracted as event description words, and news reports, information and other information related to the target sub-event are collected using search engines, news databases and other tools to form related news information.
[0094] For example, a search and match is performed within the related news information based on event description terms to identify relevant descriptive text associated with the target sub-event. A text similarity calculation method (such as cosine similarity or edit distance) is then used to calculate the similarity between the event description terms and the relevant descriptive text, which is used as the text relevance. The relevant descriptive text represents the event text associated with the target sub-event within the related news information. In other words, event information associated with the target sub-event is found within the related news information based on the target sub-event.
[0095] Exemplarily, the specific description content corresponding to the second opinion event is extracted from the target opinion text as the opinion description text, and then the similarity between the opinion description text and the related description text is calculated using a text similarity calculation method to obtain text similarity.
[0096] For example, the text relevance and text similarity are weighted and summed to obtain the event relevance between the second viewpoint event and the target sub-event. Based on the magnitude of the event relevance, different thresholds are pre-set to categorize the relationship between the second viewpoint event and the target sub-event as either relevance or non-relevance. For example, when the event relevance is greater than the threshold, it is determined to be relevance; when the event relevance is less than the threshold, it is determined to be non-relevance.
[0097] Exemplarily, the present application obtains relevant descriptive text having a relationship with the target sub-event based on the associated news information corresponding to the target sub-event. When the relevant descriptive text also has a high text similarity with the opinion description text, it indicates that there is an association between the second opinion event and the target sub-event. That is, the present application uses the relevant descriptive text associated with the target sub-event in the associated news information as a medium to determine the relationship type between the second opinion event and the target sub-event.
[0098] In some embodiments, the method of obtaining, from the associated news information according to the event descriptor, relevant descriptive text associated with the target sub-event in the associated news information and the text correlation between the event descriptor and the relevant descriptive text includes: performing sentence segmentation on the associated news information to obtain segmented sentence information corresponding to the associated news information; obtaining first text information corresponding to the target sub-event from the segmented sentence information according to the event descriptor; removing the first text information from the segmented sentence information to obtain remaining sentence information, and obtaining second text information from the remaining sentence information; performing keyword extraction on the second text information to obtain sentence keywords corresponding to the second text information; obtaining a first word that appears simultaneously with the event descriptor from the first text information, and obtaining a second word that appears simultaneously with the sentence keyword from the second text information; determining a first correlation value between the event descriptor and the sentence keyword based on the first word and the second word; performing dependency analysis on the first text information to obtain a first dependent word corresponding to the event descriptor, and performing dependency analysis on the second text information to obtain a first dependent word corresponding to the sentence keyword. a second dependent word; determining a second association value between the event description word and the sentence keyword based on the first dependent word and the second dependent word; obtaining a third word that appears in the first text information and the second text information at the same time as the first dependent word from the first text information and the second text information based on the first dependent word; obtaining a fourth word that appears in the first text information and the second text information at the same time as the second dependent word from the first text information and the second text information based on the second dependent word; determining a third association value between the event description word and the sentence keyword based on the third word and the fourth word; determining a first event association degree between the target sub-event and the second text information by fusing the second association value and the third association value; determining a second event association degree between the target sub-event and the second text information by fusing the first event association degree and the first association value; obtaining the relevant description text corresponding to the target sub-event when there is an event association in the associated news information from the remaining sentence information based on the second event association degree; determining a text association degree between the event description word and the relevant description text based on the second event association degree.
[0099] For example, the related news information is divided into independent segmented sentence information based on rules such as punctuation marks (such as periods, question marks, exclamation marks, etc.) and the semantic integrity of sentences.
[0100] For example, in the segmented sentence information, sentences containing event description words are searched by string matching, and these sentences are used as the first text information corresponding to the target sub-event. In other words, the text information describing the target sub-event is found from the associated news information.
[0101] For example, the first text information is removed from the segmented sentence information, and the remaining sentences constitute the remaining sentence information. Then, a sentence is randomly selected from the remaining sentence information to be determined as the second text information.
[0102] Exemplarily, a keyword extraction algorithm such as TextRank is used to process the second text information to extract keywords that can represent the main content of the text. These keywords are sentence keywords.
[0103] Exemplarily, the first word that appears simultaneously with the event description word is found from the first text information, and the second word that appears simultaneously with the sentence keyword is found from the second text information, and then the first word and the second word are subjected to intersection processing to obtain the number of first words corresponding to the first word and the second word when the first word and the second word are the same, and the number of first words corresponding to the first word and the number of second words corresponding to the second word are determined, and then the minimum value between the number of first words and the number of second words is obtained, and then the number of first words is divided by the minimum value between the number of first words and the number of second words to determine as the first association value between the event description word and the sentence keyword.
[0104] Exemplarily, a dependency analysis algorithm such as Stanford CoreNLP or HanLP is used to perform dependency analysis on the first text information to obtain a first dependent word corresponding to the event description word, and a dependency analysis algorithm such as StanfordCoreNLP or HanLP is used to perform dependency analysis on the second text information to obtain a second dependent word corresponding to the sentence keyword.
[0105] Exemplarily, each sub-descriptor in the event description word is obtained, and the sub-dependency word corresponding to the sub-descriptor is obtained from the first dependent word, and then the intersection processing is performed on the sub-dependency word and the second dependent word respectively to obtain the number of second words corresponding to the sub-dependency word and the second dependent word when they are the same, and the number of third words corresponding to the sub-dependency word and the number of fourth words corresponding to the second dependent word are obtained, and then the minimum value between the number of third words and the number of fourth words is obtained, and then the number of second words is divided by the minimum value between the number of third words and the number of fourth words to determine it as the association value between the sub-descriptor and the sentence keyword, and then the association values of all sub-descriptors are summed to obtain the second association value between the event description word and the sentence keyword.
[0106] Exemplarily, based on the first dependent word, a third word that appears simultaneously with the first dependent word is searched in the first text information and the second text information; similarly, based on the second dependent word, a fourth word that appears simultaneously with the second dependent word is searched, and then the third word and the fourth word are intersection-processed to obtain the number of third words corresponding to the third word and the fourth word when the third word and the fourth word are the same, and the number of fifth words corresponding to the third word and the number of sixth words corresponding to the fourth word are determined, and then the minimum value between the number of fifth words and the number of sixth words is obtained, and then the number of third words is divided by the minimum value between the number of fifth words and the number of sixth words to determine the third association value between the event description word and the sentence keyword.
[0107] Exemplarily, the second correlation value and the third correlation value are fused by weighted summation or other methods to obtain the first event correlation between the target sub-event and the second text information. Then, the first event correlation and the first correlation value can also be fused by weighted summation to obtain the second event correlation between the target sub-event and the second text information.
[0108] For example, a suitable threshold is set based on the second event relevance. Sentences with a second event relevance greater than the threshold are selected from the remaining sentence information. These sentences serve as the relevant descriptive text corresponding to the target sub-event when it is associated with an event in the associated news information. In other words, the relevant descriptive text in the associated news information that is associated with the target sub-event is selected, and the second event relevance corresponding to the selected relevant descriptive text is used as the textual relevance between the event description word and the relevant descriptive text.
[0109] Exemplarily, the relevant descriptive text related to the target sub-event is obtained based on the associated news information corresponding to the target sub-event, thereby obtaining the medium required for the target sub-event to determine the relationship type between the second viewpoint event and the target sub-event, thereby providing a good basis for subsequent relationship type determination.
[0110] Step S106: Determine an initial guidance strategy for the target event according to the target emotion type and the relationship type.
[0111] For example, when the target emotion type is negative, it is determined whether the relationship type is unrelated. If the relationship type is determined to be unrelated after analysis, it means that the target user did not express his views based on the actual situation of the target event itself when commenting on the target event. On the contrary, the target user launched a divergent comment around the core keywords involved in the target event. Such divergent comments are often divorced from the real situation of the target event and may be affected by some false information or one-sided understanding. In order to effectively guide the public to form a correct understanding of the target event, the initial guidance strategy is to use the authority and influence of official media to elaborate in detail on the real situation of the core keywords involved in the first sub-cluster in the target event. At the same time, serious warnings should be given to those related reports that may cause false guidance to the target users.
[0112] For example, when the target sentiment type is determined to be negative and the relationship type is related, each first sub-cluster is analyzed to uncover the reasons why the target user made negative comments. For example, a topic model (such as LDA) is used to extract themes from each sub-cluster. Combined with the specific content of the comments, the specific factors that led to the target user's negative emotions are identified. Based on the discovered reasons for the negative emotions, a corresponding initial guidance strategy is then developed. For example, if the negative cause is unclear information communication, leading to misunderstanding among the target user, the initial guidance strategy may include re-posting detailed and clear event content.
[0113] For example, when the target sentiment type is positive and the relationship type is unrelated, the initial guidance strategy is no guidance required because the target user's positive comments are unrelated to the target event itself and will not have a substantial impact on the target event. When the target sentiment type is positive and the relationship type is related, the reasons for the positive comments can be further analyzed, such as whether they are due to a highlight or a point worth promoting in the target event. Based on these reasons, a corresponding initial guidance strategy can be formulated, such as promoting the highlights of the target event through official channels to attract more user attention and participation.
[0114] Step S107 : obtaining the latest opinion text of the target event under the initial guidance strategy, and determining the guidance effect representation value of the initial guidance strategy according to the latest opinion text and the text clustering result.
[0115] Exemplarily, after executing the initial guidance strategy, the latest opinion text corresponding to the target event under the initial guidance strategy is re-collected, and the first emotion type is determined to be negative and the second emotion type is determined to be positive. Then, the fourth subclass cluster corresponding to the target emotion type being the first emotion type is obtained from the text clustering results, and the fifth subclass cluster corresponding to the target emotion type being the second emotion type is obtained from the text clustering results. Then, the corresponding first cluster center is obtained based on the fourth subclass cluster and the corresponding second cluster center is obtained based on the fifth subclass cluster, thereby calculating the first text distance between the latest opinion text and the first cluster center and the second text distance between the latest opinion text and the second cluster center based on the cosine distance.
[0116] For example, the first text distance and the second text distance are compared. If the first text distance is smaller than the second text distance, the latest opinion text is considered to be more closely related to the fourth sub-cluster, and the associated sub-cluster to which the latest opinion text belongs is the fourth sub-cluster. Conversely, if the second text distance is smaller than the first text distance, the associated sub-cluster is the fifth sub-cluster.
[0117] Exemplarily, a first number of the latest opinion texts belonging to the fourth sub-cluster and a second number belonging to the fifth sub-cluster are counted. For example, for each latest opinion text, a corresponding counter is updated based on the associated sub-cluster to which it belongs. For example, if the associated sub-cluster of a latest opinion text is the fourth sub-cluster, the first number is incremented by 1; if the associated sub-cluster is the fifth sub-cluster, the second number is incremented by 1.
[0118] Exemplarily, a first ratio of the first quantity to the corresponding quantity of all the latest opinion texts is obtained, and a second ratio of the second quantity to the corresponding quantity of all the latest opinion texts is obtained, so that the difference between the first ratio and the second ratio is determined as the guidance effect characterization value.
[0119] In some embodiments, the determining of the guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result includes: calculating the target text distance between the latest opinion text and each of the first subclass clusters in the text clustering result; determining the associated subclass cluster to which the latest opinion text belongs in the text clustering result based on the target text distance; determining the associated emotion type and target relationship corresponding to the latest opinion text based on the target emotion class type and the relationship type corresponding to the first subclass cluster in combination with the associated subclass cluster; counting the number of targets corresponding to the latest opinion text under the preset emotion type and preset relationship based on the associated emotion type and the target relationship; and determining the guidance effect representation value corresponding to the initial guidance strategy based on the target number.
[0120] Exemplarily, a distance calculation is performed between the latest opinion text and each first subclass cluster in the text clustering result to obtain the target text distance between the latest opinion text and the first subclass cluster, and the first subclass cluster corresponding to the minimum target text distance is determined as the associated subclass cluster to which the latest opinion text belongs, and then the associated emotion type and target relationship corresponding to the associated subclass cluster are determined according to the target emotion type and relationship type corresponding to the first subclass cluster, and then the associated emotion type and target relationship corresponding to the associated subclass cluster are determined as the associated emotion type and target relationship corresponding to the latest opinion text.
[0121] Exemplarily, the preset emotion type is positive and the preset relationship is related, then the target number corresponding to the latest opinion text when the associated emotion type is positive and the target relationship is related is counted, and the total number corresponding to the latest opinion text is obtained, thereby obtaining the ratio between the target number and the total number, and then determining the ratio as the guidance effect representation value corresponding to the initial guidance strategy.
[0122] Step S108: Modify the initial guidance strategy according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event.
[0123] Exemplarily, the preset value is determined based on expert experience and historical experience. For example, when dealing with the guidance of public opinion for large-scale commercial activities, a reasonable preset value range can be given based on the guidance results of previous similar activities, taking into account factors such as the scale of the activity, audience characteristics, and industry characteristics. When the calculated guidance effect representation value is greater than the preset value, this indicates that the current initial guidance strategy has achieved a relatively ideal effect. At this time, in order to further expand this positive impact, it is necessary to expand the publicity scope corresponding to the initial guidance strategy. Expanding the publicity scope can be done from multiple dimensions. In terms of communication channels, in addition to existing mainstream social media platforms, news media, etc., it can also be expanded to professional forums in specific fields, etc., and then the target guidance strategy corresponding to the target event can be obtained by adjusting the publicity scope in the initial guidance strategy.
[0124] For example, when the guidance effect representation value is less than or equal to a preset value, it indicates that the current initial guidance strategy may not have achieved the expected effect. First, the target association cluster to which each latest opinion text belongs is obtained, and then the number of the divided target association clusters is counted. The target association cluster with the largest statistical number is obtained, and then the target emotion type and relationship type of the target association cluster with the largest statistical number are determined as the current emotion type and current association relationship of the target user under the initial guidance strategy. Then, a corresponding guidance strategy is re-formulated based on the current emotion type and current association relationship, and the guidance strategy replaces the original initial guidance strategy to obtain the target guidance strategy corresponding to the target event, until the guidance effect representation value exceeds the preset value.
[0125] See also Figure 2 , Figure 2 The embodiment of the present application provides a large-scale model-based network public opinion intelligent guidance strategy generation device 200, which includes a data acquisition module 201, a text clustering module 202, a sentiment analysis module 203, a text recognition module 204, a relationship recognition module 205, a strategy determination module 206, a strategy evaluation module 207, and a strategy modification module 208, wherein the data acquisition module 201 is used to determine the target event and collect the target user's initial opinion text on the target event; the text clustering module 202 is used to obtain the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and perform text clustering based on the opinion text vector and the target keywords to obtain text clustering results; the sentiment analysis module 203 is used to perform sentiment analysis on each first subclass cluster in the text clustering result to obtain the target sentiment of the first subclass cluster feeling type; a text recognition module 204, used to obtain relevant keywords of the first sub-cluster from the target keywords, and determine the target opinion text of the first sub-cluster based on the relevant keywords and the first sub-cluster; a relationship recognition module 205, used to perform event relationship recognition on the target opinion text and the target event to obtain the relationship type between the target opinion text and the target event; a strategy determination module 206, used to determine the initial guidance strategy of the target event based on the target emotion type and the relationship type; a strategy evaluation module 207, used to obtain the latest opinion text of the target event under the initial guidance strategy, and determine the guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result; a strategy modification module 208, used to modify the initial guidance strategy according to the guidance effect representation value to obtain the target guidance strategy corresponding to the target event.
[0126] In some embodiments, the large model-based network public opinion intelligent guidance strategy generation device 200 can be applied to terminal devices.
[0127] It should be noted that, those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working process of the above-described large-scale model-based network public opinion intelligent guidance strategy generation device 200 can refer to the corresponding process in the aforementioned large-scale model-based network public opinion intelligent guidance strategy generation method embodiment, and will not be repeated here.
[0128] See also Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0129] like Figure 3As shown, the terminal device 300 includes a processor 301 and a memory 302 , and the processor 301 and the memory 302 are connected via a bus 303 , such as an I 2 C (Inter-Integrated Circuit) bus.
[0130] Specifically, processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. Processor 301 can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0131] Specifically, the memory 302 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.
[0132] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the embodiment of the present invention is applied. The specific server may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0133] Among them, the processor is used to run the computer program stored in the memory, and when executing the computer program, implement any one of the methods for generating an intelligent network public opinion guidance strategy based on a large model provided by the embodiments of the present invention.
[0134] In one embodiment, the processor is configured to run a computer program stored in the memory, and implement the following steps when executing the computer program:
[0135] Determine the target event and collect the target user's initial opinion text on the target event;
[0136] Obtaining the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and performing text clustering according to the opinion text vector and the target keywords to obtain a text clustering result;
[0137] Performing sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster;
[0138] Obtaining relevant keywords of the first sub-category cluster from the target keywords, and determining a target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster;
[0139] Performing event relationship recognition on the target viewpoint text and the target event to obtain a relationship type between the target viewpoint text and the target event;
[0140] Determining an initial guidance strategy for the target event according to the target emotion type and the relationship type;
[0141] Obtaining the latest opinion text of the target event under the initial guidance strategy, and determining a guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result;
[0142] The initial guidance strategy is modified according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event.
[0143] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working process of the terminal device described above can refer to the corresponding process in the aforementioned embodiment of the method for generating intelligent guidance strategies for network public opinion based on a large model, and will not be repeated here.
[0144] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any step of the method for generating an intelligent guidance strategy for network public opinion based on a large model as provided in the description of the embodiment of the present invention.
[0145] The storage medium may be an internal storage unit of the terminal device described in the aforementioned embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the terminal device.
[0146] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units in the methods, systems, and devices disclosed above may be implemented as software, firmware, hardware, or any combination thereof. In hardware embodiments, the division between functional modules / units described above does not necessarily correspond to the division between physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage media encompasses both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0147] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further limitations, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.
[0148] The serial numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope of protection of the claims.
Claims
1. A method for generating an intelligent guidance strategy for online public opinion based on a large model, characterized in that: The method comprises: Determine the target event and collect the target user's initial opinion text on the target event; Obtaining the opinion text vector and target keywords corresponding to the initial opinion text according to the large model, and performing text clustering according to the opinion text vector and the target keywords to obtain a text clustering result; Performing sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster; Obtaining relevant keywords of the first sub-category cluster from the target keywords, and determining a target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster; Performing event relationship recognition on the target viewpoint text and the target event to obtain a relationship type between the target viewpoint text and the target event; Determining an initial guidance strategy for the target event according to the target emotion type and the relationship type; Obtaining the latest opinion text of the target event under the initial guidance strategy, and determining a guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result; Modifying the initial guidance strategy according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event; The step of performing event relationship identification on the target viewpoint text and the target event to obtain the relationship type between the target viewpoint text and the target event includes: Performing event recognition on the target opinion text to obtain a first opinion event corresponding to the target opinion text and a first event type corresponding to the first opinion event; Determine a second event type corresponding to the target event, and perform event alignment on the first viewpoint event and the target event according to the first event type and the second event type to obtain a second viewpoint event corresponding to the first event type and a target sub-event corresponding to the second event type; Obtaining an event description word corresponding to the target sub-event and obtaining related news information corresponding to the target sub-event; Obtaining, from the associated news information, the relevant description text associated with the target sub-event in the associated news information and the text correlation between the event description word and the relevant description text according to the event description word; Obtaining the opinion description text corresponding to the second opinion event from the target opinion text, and calculating the text similarity between the opinion description text and the related description text; Determining an event relevance between the second opinion event and the target sub-event according to the text relevance and the text similarity; The relationship type between the second viewpoint event and the target sub-event is determined according to the event correlation degree.
2. The method according to claim 1, characterized in that Obtaining target keywords corresponding to the initial opinion text according to the large model includes: Obtaining relevant news information corresponding to the target event, and obtaining event keywords involved in the target event from the relevant news information; Obtaining a first text vector corresponding to the event keyword according to the large model; Performing text preprocessing on the initial opinion text to obtain initial keywords corresponding to the initial opinion text and first scores corresponding to the initial keywords, and obtaining second text vectors corresponding to the initial keywords based on the large model; Determining, based on the first text vector and the second text vector, associated keywords in the initial keywords that are related to the target event; Performing a word co-occurrence analysis based on the associated keywords in combination with the initial opinion text to obtain co-occurrence keywords corresponding to the associated keywords and co-occurrence frequencies between the associated keywords and the co-occurrence keywords; Establishing a word co-occurrence graph corresponding to the initial opinion text based on the associated keywords and the co-occurrence keywords in combination with the co-occurrence frequency; Eliminating the associated keywords from the initial keywords to obtain candidate keywords corresponding to the initial opinion text, and obtaining second scores corresponding to the candidate keywords based on the first scores corresponding to the initial keywords; Obtaining a connection keyword corresponding to the candidate keyword from the word co-occurrence graph, and determining associated edge information corresponding to the candidate keyword and the connection keyword; Obtaining an external keyword corresponding to the connection keyword from the word co-occurrence graph, and determining a connection density between the candidate keyword and the connection keyword based on the external keyword and the connection keyword; Obtaining frequency information of the candidate keyword in the initial opinion text, and determining a target score corresponding to the candidate keyword based on the frequency information, the second score, and the connection density; Obtaining a selected keyword from the candidate keywords according to the target score, and determining the target keyword corresponding to the initial opinion text according to the selected keyword and the associated keyword; The target score is obtained according to the following formula: ; in, represents the target score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the second score corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the jth connection keyword corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text, represents the connection density between the i-th candidate keyword and the j-th connection keyword corresponding to the m-th initial opinion text, max represents obtaining the maximum value, lg represents a logarithmic function with a base of 10, num represents the number of texts corresponding to all the initial opinion texts, represents the frequency information corresponding to the i-th candidate keyword corresponding to the m-th initial opinion text in the t-th initial opinion text, Indicates the number of occurrences of the i-th candidate keyword corresponding to the m-th initial opinion text in all the initial opinion texts.
3. The method according to claim 1, characterized in that The performing text clustering according to the opinion text vector and the target keyword to obtain a text clustering result includes: Obtaining a first keyword and a second keyword from the target keyword, and calculating a degree of cohesion between the first keyword and the second keyword to obtain a cohesion value between the first keyword and the second keyword; Determining the number of clusters corresponding to the text clustering result and the initial cluster center corresponding to the text clustering result according to the connection value; Calculating an initial probability that the opinion text vector belongs to the initial cluster center, and determining data distribution information corresponding to the opinion text vector under the initial cluster center according to the initial probability; Determining distribution difference information based on the initial probability and the data distribution information, and determining the associated cluster center corresponding to the opinion text vector under the initial cluster center based on the distribution difference information; Determining the initial clustering result corresponding to the initial clustering center according to the opinion text vector and the associated cluster center; updating the initial cluster center according to the initial clustering result to obtain the latest cluster center; The text clustering is performed again according to the latest cluster center and the opinion text vector until the latest cluster center no longer changes, thereby obtaining the text clustering result.
4. The method according to claim 3, characterized in that The updating of the initial cluster center according to the initial clustering result to obtain the latest cluster center includes: first distance information between any two first sub-text vectors in each second sub-class cluster in the initial clustering result, and determining first average distance information corresponding to the second sub-class cluster according to the first distance information; Obtaining second distance information between a second sub-text vector in the second sub-class cluster and remaining sub-text vectors in the second sub-class cluster after excluding the second sub-text vector; Determining first similar data information corresponding to the second sub-text vector in the second sub-class cluster according to the second distance information and the first average distance information; Determining a similar internal distance corresponding to the second sub-text vector according to the first similar data information and the second sub-text vector; Obtaining a third subclass cluster after excluding the second subclass cluster from the initial clustering result, obtaining third distance information between any two third subtext vectors in the third subclass cluster, and determining second average distance information corresponding to the third subclass cluster based on the third distance information; Obtaining fourth distance information between a fourth subtext vector in the third subclass cluster and remaining subtext vectors in the third subclass cluster after excluding the third subtext vector; Determining second similar data information corresponding to the fourth sub-text vector in the third sub-class cluster according to the fourth distance information and the second average distance information; Obtaining the amount of data corresponding to the second similar data information, and determining an associated cluster corresponding to the second sub-cluster according to the amount of data; Determining a weight center representation value corresponding to the second sub-text vector according to the associated cluster, the similar internal distance and the first similar data information; The initial cluster center is updated according to the weight center representation value to obtain the latest cluster center.
5. The method according to claim 1, wherein The step of obtaining, from the associated news information according to the event description word, relevant description text associated with the target sub-event in the associated news information and the text correlation between the event description word and the relevant description text includes: Segmenting the related news information to obtain segmented sentence information corresponding to the related news information; Obtaining first text information corresponding to the target sub-event from the segmented sentence information according to the event description word; removing the first text information from the segmented sentence information to obtain remaining sentence information, and obtaining second text information from the remaining sentence information; performing keyword extraction on the second text information to obtain sentence keywords corresponding to the second text information; Obtaining a first word that appears simultaneously with the event description word from the first text information, and obtaining a second word that appears simultaneously with the sentence keyword from the second text information; Determining a first association value between the event description word and the sentence keyword according to the first word and the second word; Performing a dependency analysis on the first text information to obtain a first dependent word corresponding to the event description word, and performing a dependency analysis on the second text information to obtain a second dependent word corresponding to the sentence keyword; determining a second association value between the event description word and the sentence keyword according to the first dependent word and the second dependent word; obtaining, from the first text information and the second text information, a third word that appears in both the first text information and the second text information at the same time as the first dependent word; obtaining, from the first text information and the second text information, a fourth word that appears in both the first text information and the second text information at the same time as the second dependent word, according to the second dependent word; determining a third association value between the event description word and the sentence keyword according to the third word and the fourth word; fusing the second association value and the third association value to determine a first event association degree between the target sub-event and the second text information; fusing the first event relevance and the first relevance value to determine a second event relevance between the target sub-event and the second text information; Obtaining, from the remaining sentence information according to the second event relevance, the relevant description text corresponding to the target sub-event when there is an event association in the related news information; The text association degree between the event description word and the related description text is determined according to the second event association degree.
6. The method according to claim 1, characterized in that The determining of the guidance effect representation value of the initial guidance strategy according to the latest opinion text and the text clustering result includes: Calculating a target text distance between the latest opinion text and each of the first subclass clusters in the text clustering result; Determining the associated subclass cluster to which the latest opinion text belongs in the text clustering result according to the target text distance; Determine the associated sentiment type and target relationship corresponding to the latest opinion text according to the target sentiment type and the relationship type corresponding to the first subclass cluster in combination with the associated subclass cluster; According to the associated emotion type and the target relationship, the number of targets corresponding to the latest opinion text under the preset emotion type and the preset relationship is counted; The guidance effect representation value corresponding to the initial guidance strategy is determined according to the target quantity.
7. A device for generating intelligent guidance strategies for online public opinion based on a large model, characterized in that: include: A data collection module is used to determine a target event and collect the target user's initial opinion text on the target event; A text clustering module, configured to obtain, based on the large model, an opinion text vector and target keywords corresponding to the initial opinion text, and perform text clustering based on the opinion text vector and the target keywords to obtain a text clustering result; A sentiment analysis module, configured to perform sentiment analysis on each first subclass cluster in the text clustering result to obtain a target sentiment type of the first subclass cluster; a text recognition module, configured to obtain relevant keywords of the first sub-category cluster from the target keywords, and determine a target opinion text of the first sub-category cluster based on the relevant keywords and the first sub-category cluster; A relationship identification module, configured to perform event relationship identification on the target opinion text and the target event to obtain a relationship type between the target opinion text and the target event; A strategy determination module, configured to determine an initial guidance strategy for the target event according to the target emotion type and the relationship type; A strategy evaluation module, configured to obtain the latest opinion text of the target event under the initial guidance strategy, and determine a guidance effect representation value of the initial guidance strategy based on the latest opinion text and the text clustering result; A strategy modification module, configured to modify the initial guidance strategy according to the guidance effect representation value to obtain a target guidance strategy corresponding to the target event; The step of performing event relationship identification on the target viewpoint text and the target event to obtain the relationship type between the target viewpoint text and the target event includes: Performing event recognition on the target opinion text to obtain a first opinion event corresponding to the target opinion text and a first event type corresponding to the first opinion event; Determine a second event type corresponding to the target event, and perform event alignment on the first viewpoint event and the target event according to the first event type and the second event type to obtain a second viewpoint event corresponding to the first event type and a target sub-event corresponding to the second event type; Obtaining an event description word corresponding to the target sub-event and obtaining related news information corresponding to the target sub-event; Obtaining, from the associated news information, the relevant description text associated with the target sub-event in the associated news information and the text correlation between the event description word and the relevant description text according to the event description word; Obtaining the opinion description text corresponding to the second opinion event from the target opinion text, and calculating the text similarity between the opinion description text and the related description text; Determining an event relevance between the second opinion event and the target sub-event according to the text relevance and the text similarity; The relationship type between the second viewpoint event and the target sub-event is determined according to the event correlation degree.
8. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the method for generating an intelligent network public opinion guidance strategy based on a large model as described in any one of claims 1 to 6 when executing the computer program.
9. A computer storage medium for computer storage, characterized in that: The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the method for generating an intelligent guidance strategy for network public opinion based on a large model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Network public opinion analysis method, device and storage medium
CN109145215A
Public opinion monitoring method with emergency event analysis and extraction function
CN115658996A