Public opinion detection and analysis system
By constructing a word meaning link tree and a semantic link tree, analyzing the dynamic semantic correlation between words, the problem of insufficient accuracy of complex contexts and obscure emotions recognition in the existing technology is solved, and higher emotion recognition accuracy and dynamic adaptability are achieved.
Patent Information
- Application Number
- CN202510189243.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Existing keyword recognition technologies lack the accuracy of recognition in texts with complex contexts or obscure emotions expressions, and lack semantic dynamics, making it difficult to capture the ambiguity and changes in emotional tendencies of words in different contexts.
By constructing a word meaning link tree and a semantic link tree, the dynamic semantic correlation between words is analyzed, and the co-occurrence frequency matrix and keyword detection module are used for matching analysis, dynamically adjusting the emotional weight of words, and improving the accuracy of emotion recognition.
It significantly improves the accuracy of emotion recognition in complex contexts, enhances the dynamic adaptability of emotion analysis, improves the robustness and stability of emotion recognition, and can more accurately judge the user's emotional state.
Smart Images

Figure CN120124635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of public opinion detection and analysis, and specifically to a public opinion detection and analysis system. Background Art
[0002] With the rise of applications such as social media, intelligent customer service, and user experience analysis, the application demand for keyword recognition technology in fields such as social public opinion monitoring, user behavior analysis, and mental health monitoring has been increasing continuously. Keyword recognition technology can help systems or services better understand user needs by analyzing the expressions of users in text, speech, or video and identifying their emotional states. However, the existing keyword recognition technology has the following limitations in practical applications: Insufficient keyword recognition accuracy: Existing keyword recognition usually relies on keyword dictionaries or datasets based on specific keyword tags. This method works well in corpora with specific keyword expressions, but often performs poorly on texts with complex contexts or implicit emotional expressions. For example, for sentences containing homophonic words, existing models are difficult to effectively identify, resulting in low recognition accuracy. Lack of semantic dynamics: Most keyword recognition systems are based on static lexical associations or feature engineering, lacking the ability to analyze the dynamic semantic associations between words. Especially when words have polysemy or emotional tendency changes in different contexts, existing models are difficult to accurately capture, affecting the accuracy of keyword judgment. Summary of the Invention
[0003] Based on the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a public opinion detection and analysis system to solve the above technical problems.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A public opinion detection and analysis system, comprising:
[0005] A semantic connection module for constructing a semantic connection tree of words, where the root node is any word, the child nodes are words related to the root node, and values are assigned to the child nodes according to the co-occurrence frequency, such that the sum of the values of all child nodes is one;
[0006] A semantic connection module for constructing a semantic connection tree for each child node with the child node of the semantic connection tree of words as the root node until a preset number of layers is reached, to obtain a semantic connection tree;
[0007] A keyword detection module for performing matching analysis on the semantic connection tree of the document to be analyzed and keywords to obtain a set of public opinion keywords of the document to be analyzed.
[0008] The present invention is further configured such that the semantic connection module includes: a data processing unit, a node construction unit, and a node assignment unit;
[0009] A data processing unit, configured to extract words from historical texts, count their occurrence frequencies, generate a word frequency table, calculate the co-occurrence frequencies between words, and generate a co-occurrence frequency matrix;
[0010] A node construction unit, configured to select the corresponding words in descending order according to the word frequency table as root nodes, extract the words connected to the root nodes from the co-occurrence matrix, and fill the child nodes in order of co-occurrence frequency;
[0011] A node value assignment unit, configured to assign values to each child node according to the co-occurrence frequency between the word corresponding to the child node and the root node, and the sum of all child nodes is one.
[0012] The present invention is further configured such that calculating the co-occurrence frequencies between words and generating a co-occurrence frequency matrix includes:
[0013] Create a word pair co-occurrence matrix C, where the rows and columns of the word pair co-occurrence matrix are all words that appear in the historical text;
[0014] Set the window size of the word pair, perform co-occurrence calculation, for each word w i Traverse, search for other words w within the window range j , and update the value at the corresponding position in the co-occurrence matrix C. The update logic is: C ′ [w i ,w j = C[w i ,w j + 1, where C ′ [w i ,w j is the updated number of times, C[w i ,w j is the number of other words w j within the window range, with an initial value of zero. Whenever another word w j appears within the window range, the value is incremented by one. After the traversal is completed, count the co-occurrence times of word w i and word w j ;
[0015] Calculate the co-occurrence frequency according to the co-occurrence times, and generate a co-occurrence frequency matrix between words. The calculation logic of the co-occurrence frequency is: P(w i ,w j ) is the co-occurrence frequency of word w i and word w j , and k is the index of all words that co-occur with word w i in the co-occurrence matrix.
[0016] The present invention is further configured to assign values to each child node according to the co-occurrence frequency of the word corresponding to the child node and the root node, and the sum of all child nodes is one. The calculation logic for assigning values to each child node is as follows: where v(w j ) is the assigned value of the word w j on the child node, and N is the number of child nodes.
[0017] The present invention is further configured that the construction logic of the semantic connection tree includes:
[0018] Extract the set of child nodes of the current node from the semantic connection tree of word meanings. Starting from the first child node, create a new semantic connection tree with the current child node as the root node, and repeat to complete the set of child nodes;
[0019] For each new semantic connection tree, assign values to each child node according to the co-occurrence frequency of the word corresponding to the child node and its parent node, and the sum of all child nodes is one;
[0020] Repeat until the preset number of layers is reached to obtain a complete semantic connection tree.
[0021] The present invention is further configured that in the keyword detection module, keywords are generated through pre-setting, and a semantic connection tree of keywords is generated according to the pre-set keywords;
[0022] When performing matching analysis on the semantic connection tree of the document to be analyzed and the keywords to obtain the set of public opinion keywords of the document to be analyzed, it includes:
[0023] Perform word segmentation processing on the document to be analyzed to obtain the word sequence of the document to be analyzed;
[0024] Map the first character of the segmented words in the word sequence to the root node of the semantic connection tree of the keywords, and map the remaining characters in the word sequence to the corresponding nodes in the semantic connection tree of the keywords;
[0025] Calculate the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keywords, and input all words with a matching degree greater than the preset matching degree threshold into the keyword candidate set;
[0026] Calculate the semantic influence factor of each segmented word in the keyword candidate set, and select the top M words with the highest semantic influence factor as the final set of public opinion keywords.
[0027] The present invention is further configured such that the logic of word segmentation processing is as follows: format the document to be analyzed, remove punctuation and paragraphs to obtain a set of words in the document to be analyzed, obtain the pinyin of each word, and match it with the pinyin of the root node of the semantic connection tree of the keyword. When a match is found, starting from the matched word, obtain the same number of words as the preset number of layers and set them as segmented words. Combine all the segmented words to obtain the word sequence of the document to be analyzed.
[0028] The present invention is further configured such that the calculation logic of the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keyword includes:
[0029] Obtain the number of repeated words between the segmented words in the word sequence and the semantic connection tree of the keyword, and calculate the repeated similarity;
[0030] Calculate the co-occurrence similarity according to the values assigned to the repeated words in the semantic connection tree;
[0031] Calculate the hierarchical similarity according to the levels of the repeated words in the semantic connection tree;
[0032] Calculate the matching degree between the segmented words and the semantic connection tree of the keyword according to the repeated similarity, co-occurrence similarity, and hierarchical similarity.
[0033] The present invention is further configured such that the calculation logic of the repeated similarity is as follows: Wherein, R s is the repeated similarity, N r is the number of repeated words between the segmented words and the semantic connection tree, and N t is the total number of words in the segmented words; the calculation logic of the co-occurrence similarity is as follows: Wherein, C s is the co-occurrence similarity, v(w r ) is the value assigned to the r-th repeated word between the segmented words and the semantic connection tree in the semantic connection tree of the keyword; the calculation logic of the hierarchical similarity is as follows: Wherein, L s is the hierarchical similarity, D r is the absolute difference in the levels of the r-th repeated word between the segmented words and the semantic connection tree; the calculation logic of the matching degree between the segmented words and the semantic connection tree of the keyword is: M d =α*R s +β*C s +γ*L s , wherein, M d is the matching degree between the segmented words and the semantic connection tree of the keyword, and α, β, and γ are weight coefficients of the similarity, all of which are greater than zero, and α + β + γ = 1.
[0034] The present invention is further configured such that the calculation logic of the semantic influence factor of each segmented word in the keyword candidate set is as follows: wherein, F si is the semantic influence factor of each segmented word in the keyword candidate set, and f(S(word) i ) is the frequency of occurrence of the l-th segmented word in the word sequence in the document to be analyzed, and L is the number of segmented words in the word sequence.
[0035] The present invention provides a public opinion detection and analysis system, including: a semantic connection module for constructing a semantic connection tree of words, wherein the root node is any word, the child nodes are words related to the root node, and values are assigned to the child nodes according to the co-occurrence frequency, so that the sum of the values of all child nodes is one; a semantic connection module for constructing a semantic connection tree of each child node with the child node of the semantic connection tree as the root node until a preset number of layers is reached to obtain a semantic connection tree; a keyword detection module for performing matching analysis on the document to be analyzed and the semantic connection tree of keywords to obtain a set of public opinion keywords of the document to be analyzed. The beneficial effects generated include:
[0036] 1. Improve the accuracy of emotion recognition: By constructing a word association tree, the dynamic semantic association between words can be analyzed, enabling the system to dynamically adjust the emotion weights of words according to the context. The dynamic analysis model based on semantic association significantly improves the emotion recognition accuracy in complex contexts. Especially when processing texts containing polysemous words and homophones, it can more accurately judge the user's emotion state than traditional static dictionaries or emotion tagging methods;
[0037] 2. Enhance the dynamic adaptability of emotion analysis: By constructing a co-occurrence frequency matrix and a keyword detection module, the keywords of the document can be effectively analyzed, enabling the system of the present invention to dynamically adapt to keyword changes in scenarios such as public opinion monitoring and long-term user interaction analysis;
[0038] 3. Improve the robustness and stability of emotion recognition. By introducing the calculation of multi-level similarity and matching degree, the system of the present invention shows stronger robustness and stability when processing texts with different contexts and different emotion expression methods. Whether it is the keyword recognition of a single text or the long-term keyword trend tracking, the present invention can provide consistent and reliable emotion analysis results to meet diverse emotion recognition needs.
[0039] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Brief Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0041] Figure 1 It is a schematic structural diagram of a public opinion detection and analysis system shown in an exemplary embodiment of the present invention. Detailed implementation manners
[0042] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.
[0043] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0044] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0045] A public opinion detection and analysis system, as Figure 1 shown, includes:
[0046] A semantic connection module, configured to construct a semantic connection tree of words, where the root node is any word, the child nodes are words related to the root node, and values are assigned to the child nodes according to the co-occurrence frequency, such that the sum of the values of all child nodes is one;
[0047] A semantic connection module, configured to construct a semantic connection tree for each child node with the child node of the semantic connection tree of words as the root node until a preset number of layers is reached, to obtain a semantic connection tree;
[0048] The keyword detection module is used to perform matching analysis on the semantic connection tree of the document to be analyzed and keywords, and obtain the public opinion keyword set of the document to be analyzed.
[0049] The present invention is further configured such that the semantic connection module includes: a data processing unit, a node construction unit, and a node assignment unit;
[0050] The data processing unit is used to extract characters from historical texts, count the occurrence frequencies, generate a character frequency table, calculate the co-occurrence frequencies between characters, and generate a co-occurrence frequency matrix; specifically, format the historical texts, including removing punctuation marks, spaces, and line breaks, ensuring that only valid Chinese characters are retained, convert the texts to a unified character encoding (such as UTF-8) to ensure the consistency of subsequent processing; traverse the processed text characters one by one, extract the Chinese characters therein, ignore other non-Chinese characters, and store the extracted Chinese characters in a list or set for frequency statistics; use a dictionary or hash table data structure to record the number of times each Chinese character appears. For each Chinese character, if the Chinese character already exists in the dictionary, its count value is incremented by one; if the Chinese character does not exist, add the Chinese character to the dictionary and initialize its count value to one; sort the character frequency table in descending order according to the occurrence frequencies.
[0051] The node construction unit is used to select the corresponding characters in descending order according to the character frequency table as the root nodes, extract the characters connected to the root nodes from the co-occurrence matrix, and fill in the sub-nodes in the order of co-occurrence frequencies;
[0052] The node assignment unit is used to assign values to each sub-node according to the co-occurrence frequency between the character corresponding to the sub-node and the root node, and the sum of all sub-nodes is one.
[0053] The present invention is further configured such that calculating the co-occurrence frequencies between characters and generating a co-occurrence frequency matrix includes:
[0054] Create a co-occurrence matrix C for word pairs, where the rows and columns of the co-occurrence matrix for word pairs are all the characters that appear in the historical texts;
[0055] Set the window size of the word pairs for co-occurrence calculation. For each character w i traverse to find other characters w j within the window range, and update the value at the corresponding position in the co-occurrence matrix C. The update logic is: C ′ [w i , w j = C[w i , w j + 1, where C ′ [w i , w j is the updated number of times, and C[w i , w jFor other words w within the window range j The number of which has an initial value of zero. Whenever other word w appears within the window range j , the value is incremented by one. After the traversal is completed, count the co-occurrence times of word w i and word w j ; Specifically, the window range is preferably 2 - 10. No specific value is limited here, but a window exceeding 10 will exponentially increase the computational complexity when generating the semantic connection tree;
[0056] Calculate the co-occurrence frequency based on the co-occurrence times, and generate a co-occurrence frequency matrix between words. The calculation logic of the co-occurrence frequency is: P(w i ,w j ) is the co-occurrence frequency of word w i and word w j , and k is the index of all words co-occurring with word w i in the co-occurrence matrix. Specifically, the above calculation logic generates a co-occurrence frequency matrix of word pairs by counting the co-occurrence times between words, quantifying the co-occurrence relationship between words. The co-occurrence frequency matrix reflects the probability of each word co-occurring with other words in the text;
[0057] The present invention is further configured to assign a value to each child node according to the co-occurrence frequency between the word corresponding to the child node and the root node, and the sum of all child nodes is one. The calculation logic for assigning a value to each child node is: wherein, v(w j ) is the assigned value of word w j on the child node, and N is the number of child nodes. Specifically, the above calculation logic assigns a weight (i.e., the assigned value) to each child node by calculating the co-occurrence frequency between the child node and the root node to reflect the importance of the child node in the association tree. When calculating, the assigned value of each child node is based on its co-occurrence frequency with the root node, ensuring that the sum of the assigned values of all child nodes is 1. It can quantify the influence or association degree of the child node relative to the root node.
[0058] The present invention is further configured that the construction logic of the semantic connection tree includes:
[0059] Extract the set of child nodes of the current node from the semantic connection tree. Starting from the first child node, create a new semantic connection tree with the current child node as the root node, and repeat to complete the set of child nodes;
[0060] For each new semantic connection tree, assign a value to each child node according to the co-occurrence frequency between the word corresponding to the child node and its parent node, and the sum of all child nodes is one;
[0061] Repeat until the preset number of layers is reached to obtain a complete semantic connection tree. Specifically, the semantic connection tree is constructed recursively to reflect the hierarchical relationship between words in a layer-by-layer expansion manner. Specifically, starting from the root node in the word meaning connection tree, extract the child nodes layer by layer and calculate the weights based on the co-occurrence frequency to form a semantic connection tree with a hierarchical structure. After reaching the preset number of layers, a complete semantic connection tree is obtained. The assigned value of each child node is based on its degree of association with its parent node, and it is ensured that the sum of the assigned values of each layer is 1.
[0062] The present invention is further configured such that in the keyword detection module, keywords are generated through presetting, and a semantic connection tree of the keywords is generated according to the preset keywords; specifically, in the keyword detection module, a semantic connection tree is generated through a preset keyword set. The semantic connection tree of the keywords is used as a reference structure when detecting public opinion, so as to identify words or phrases semantically related to the preset keywords in the document to be analyzed. The preset semantic connection tree of the keywords helps to improve the detection efficiency and accuracy, focuses the calculation in the semantic space related to the keywords, and avoids calculating irrelevant words for the entire text;
[0063] When performing a matching analysis on the document to be analyzed and the semantic connection tree of the keywords to obtain the set of public opinion keywords of the document to be analyzed, it includes:
[0064] Perform word segmentation on the document to be analyzed to obtain the word sequence of the document to be analyzed; the present invention is further configured such that the logic of word segmentation is: format the document to be analyzed, remove punctuation and paragraphs to obtain the text set of the document to be analyzed, obtain the pinyin for each character, and match it with the pinyin of the root node of the semantic connection tree of the keywords. When a match is found, starting from the matched character, obtain the same number of characters as the preset number of layers and set them as the segmented words. Combine all the segmented words to obtain the word sequence of the document to be analyzed.
[0065] Map the first character of the segmented words in the word sequence to the root node of the semantic connection tree of the keywords, and map the remaining characters in the word sequence to the corresponding nodes in the semantic connection tree of the keywords; specifically, by performing a character-by-character mapping of the word sequence in the document to be analyzed and the semantic connection tree of the keywords, determine the degree of association between each segmented word in the word sequence and the keywords. The mapping can help the system quickly determine whether a word is semantically related to the preset keywords, thereby improving the detection efficiency and accuracy of the keywords;
[0066] Calculate the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keywords, and input all the words with a matching degree greater than the preset matching degree threshold into the keyword candidate set; the present invention is further set that the calculation logic of the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keywords includes: obtaining the number of repeated characters between the segmented words in the word sequence and the semantic connection tree of the keywords, and calculating the repeated similarity; calculating the co-occurrence similarity according to the assigned value of the repeated characters in the semantic connection tree; calculating the hierarchical similarity according to the level of the repeated characters in the semantic connection tree; calculating the matching degree between the segmented words and the semantic connection tree of the keywords according to the repeated similarity, co-occurrence similarity and hierarchical similarity. The present invention is further set that the calculation logic of the repeated similarity is: wherein, R s is the repeated similarity, N r is the number of repeated characters between the segmented words and the semantic connection tree of the keywords, N t is the total number of words in the segmented words; the calculation logic of the co-occurrence similarity is: wherein, C s is the co-occurrence similarity, v(w r ) is the assigned value of the r-th repeated character between the segmented words and the semantic connection tree of the keywords in the semantic connection tree of the keywords; the calculation logic of the hierarchical similarity is: wherein, L s is the hierarchical similarity, D r is the absolute difference in the number of layers of the r-th repeated character between the segmented words and the semantic connection tree of the keywords; the calculation logic of the matching degree between the segmented words and the semantic connection tree of the keywords is: M d =α*R s +β*C s +γ*L s , wherein, M dis the matching degree between the segmented word and the semantic connection tree of the keyword. α, β, and γ are the weight coefficients of the similarity, all greater than zero, and α + β + γ = 1. Specifically, the above logical calculation calculates the matching degree between the segmented words in the document to be analyzed and the semantic connection tree of the keyword. By quantifying multiple dimensions of the matching degree, including repetition similarity, co-occurrence similarity, and hierarchical similarity, the relevance between the segmented words and the keyword in the text to be analyzed is judged. The specific steps include calculating the repetition similarity, co-occurrence similarity, hierarchical similarity, and comprehensively calculating the final matching degree based on these similarities; the repetition similarity measures the number of overlapping words between the segmented words in the word sequence to be analyzed and the semantic connection tree of the keyword, and is used to quantify the literal overlap degree between the segmented word and the semantic connection tree of the keyword. The more overlap, the stronger the possible semantic association between the two words; the co-occurrence similarity measures the sum of the co-occurrence frequencies of the segmented word and the child nodes of the keyword in the semantic connection tree, and is used to quantify the co-occurrence relationship between the segmented word and the keyword in the semantic connection tree. The higher the assigned value, the higher the importance of the node, and the larger the sum of the co-occurrence similarities, the higher the relevance between the two words in the semantic connection tree; the hierarchical similarity measures the proximity of the corresponding child nodes in the semantic connection trees of the segmented word and the keyword at the hierarchy level, and is used to evaluate the relative positions of the two words in the semantic connection tree. The closer the hierarchy, the stronger their possible semantic relationship. By weighted summation of these three similarity indicators, a quantified matching degree indicator can be obtained for the final determination of the semantic relevance between the segmented word and the keyword.
[0067] Calculate the semantic influence factor of each segmented word in the keyword candidate set, and select the top M words with the highest semantic influence factors as the final public opinion keyword set. The present invention is further set that the calculation logic of the semantic influence factor of each segmented word in the keyword candidate set is: wherein, F si is the semantic influence factor of each segmented word in the keyword candidate set, f(S(word) i ) is the frequency of the l-th segmented word in the word sequence appearing in the document to be analyzed, and L is the number of segmented words in the word sequence. Specifically, the above calculation logic evaluates the importance in the text by calculating the semantic influence factor of each segmented word in the keyword candidate set. The semantic influence factor comprehensively considers the frequency of occurrence and the matching degree of the word in the text to be analyzed, weights these two factors, and obtains a comprehensive indicator reflecting the importance of the word. According to the size of the semantic influence factor, the top M words with the strongest semantic relevance can be selected as the final public opinion keyword set; through the calculation of the semantic influence factor, the frequency of occurrence of the word can be combined with the semantic matching degree to more comprehensively evaluate the importance of the word in the text, ensuring that the selected keywords are highly relevant to the theme semantically.
[0068] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0069] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0070] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0071] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not indicate the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0072] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0073] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0074] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0075] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0076] In addition, the functional units in the various embodiments of this application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0077] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0078] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A public opinion detection and analysis system, characterized in that: include: The word sense association module is used to construct a word sense association tree of characters, where the root node is any character, the child nodes are characters related to the root node, and the child nodes are assigned values according to the co-occurrence frequency, so that the sum of the values of all child nodes is one; A semantic connection module, used to construct a word sense connection tree for each child node based on the child node of the word sense connection tree as a root node, until a preset number of layers is reached to obtain a semantic connection tree; The keyword detection module is used to match and analyze the semantic connection tree of the document to be analyzed and the keywords to obtain the public opinion keyword set of the document to be analyzed.
2. A public opinion detection and analysis system according to claim 1, characterized in that: The semantic connection module includes: a data processing unit, a node construction unit and a node assignment unit; A data processing unit, used to extract characters from historical texts and count their occurrence frequencies, generate a character frequency table, calculate the co-occurrence frequencies between characters, and generate a co-occurrence frequency matrix; A node construction unit, used to select the corresponding word as the root node according to the descending order of the word frequency table, extract the words connected to the root node from the co-occurrence matrix, and fill the child nodes according to the co-occurrence frequency sorting; The node assignment unit is used to assign a value to each child node according to the co-occurrence frequency of the word corresponding to the child node and the root node, and the sum of all child nodes is one.
3. A public opinion detection and analysis system according to claim 2, characterized in that: Calculate the co-occurrence frequency between words and generate a co-occurrence frequency matrix, including: Create a word pair co-occurrence matrix C, wherein the rows and columns of the word pair co-occurrence matrix are all words that appear in the historical text; Set the window size of the word pair and perform co-occurrence calculations. i Traverse and find other words w within the window range j , and update the value of the corresponding position in the co-occurrence matrix C. The update logic is: C ′ [w i ,w j ]=C[w i ,w j ]+1,C ′ [w i ,w j ] is the number of updates, C[w i ,w j ] is other characters within the window range j The initial value is zero. Whenever other characters w appear in the window range, j , the value is increased by one, and after the traversal is completed, the word w is counted i and word w j The number of co-occurrences of The co-occurrence frequency is calculated according to the co-occurrence times, and a co-occurrence frequency matrix between characters is generated. The calculation logic of the co-occurrence frequency is: P(w i ,w j ) is the word w i and word w j The co-occurrence frequency of word w is k, and k is the co-occurrence matrix. i Index of co-occurring words.
4. A public opinion detection and analysis system according to claim 3, characterized in that: A value is assigned to each child node according to the co-occurrence frequency of the word corresponding to the child node and the root node, and the sum of all child nodes is one. The calculation logic of the value assigned to each child node is: Among them, v(w j ) is the word w j The value assigned to the child node, N is the number of child nodes.
5. A public opinion detection and analysis system according to claim 1, characterized in that: The construction logic of the semantic connection tree includes: Extracting a set of child nodes of the current node from the word sense association tree, starting from the first child node, taking the current child node as the root node, creating a new semantic association tree for the current child node, and repeatedly completing the set of child nodes; For each new semantic association tree, assign a value to each child node according to the co-occurrence frequency of the word corresponding to the child node and its parent node, and the sum of all child nodes is one; Repeat until the preset number of layers is reached to obtain a complete semantic connection tree.
6. A public opinion detection and analysis system according to claim 1, characterized in that: In the keyword detection module, keywords are generated by presetting, and a semantic connection tree of keywords is generated based on the presetting keywords; When matching and analyzing the semantic connection tree of the document to be analyzed and the keyword, and obtaining the public opinion keyword set of the document to be analyzed, it includes: Perform word segmentation on the document to be analyzed to obtain the word sequence of the document to be analyzed; Mapping the first character of the segmented word in the word sequence to the root node of the semantic connection tree of the keyword, and mapping the remaining characters in the word sequence to the corresponding nodes in the semantic connection tree of the keyword; Calculate the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keyword, and input all words with matching degree greater than a preset matching degree threshold into the keyword candidate set; Calculate the semantic influence factor of each segmented word in the keyword candidate set, and select the top M words with the highest semantic influence factors as the final public opinion keyword set.
7. A public opinion detection and analysis system according to claim 6, characterized in that: The logic of word segmentation processing is: format the document to be analyzed, remove punctuation and paragraphs, obtain the text set of the document to be analyzed, obtain the pinyin of each word, and match it with the pinyin of the root node of the semantic connection tree of the keyword. When a match is found, start from the matched word, obtain the same number of words as the preset number of layers, set them as word segmentation words, and collect all the word segmentation words to obtain the word sequence of the document to be analyzed.
8. A public opinion detection and analysis system according to claim 6, characterized in that: The calculation logic of the matching degree between the segmented words in the word sequence of the document to be analyzed and the semantic connection tree of the keywords includes: Obtain the number of repeated characters in the semantic connection tree between the segmented words and the keywords in the word sequence, and calculate the repeated similarity; Calculate the co-occurrence similarity based on the values assigned to the repeated words in the semantic association tree; Calculate the hierarchical similarity based on the hierarchical level of the repeated words in the semantic connection tree; The matching degree of the semantic connection tree between the segmented words and the keywords is calculated based on repetition similarity, co-occurrence similarity and hierarchical similarity.
9. A public opinion detection and analysis system according to claim 8, characterized in that: The calculation logic of the repeated similarity is: Among them, R s is the repetition similarity, N r is the number of repeated words in the segmented words and the semantic connection tree, N t is the total number of words in the segmented words; the calculation logic of the co-occurrence similarity is: Among them, C s is the co-occurrence similarity, v(w r ) is the distribution value of the rth word in the semantic connection tree that is repeated in the segmented word and the semantic connection tree in the semantic connection tree of the keyword; the calculation logic of the hierarchical similarity is: Among them, L s is the level similarity, D r is the absolute difference between the number of layers of the rth word repeated in the segmentation word and the semantic connection tree; the calculation logic of the matching degree between the segmentation word and the semantic connection tree of the keyword is: d =α*R s +β*C s +γ*L s , where M d is the matching degree of the semantic connection tree between the segmented words and the keywords, α, β and γ are weight coefficients of similarity, all of which are greater than zero, and α+β+γ=1.
10. A public opinion detection and analysis system according to claim 9, characterized in that: The calculation logic of the semantic impact factor of each segmented word in the keyword candidate set is: Among them, F si is the semantic influence factor of each word in the keyword candidate set, f(S(word) i ) is the frequency of the lth segmentation word in the word sequence appearing in the document to be analyzed, and L is the number of segmentation words in the word sequence.
Citation Information
Patent Citations
Intent recognition method, device and equipment based on artificial intelligence and storage medium
CN113935333A
Electronic file archiving and classifying method based on semantic analysis
CN117273015A
Cited By
Data set quality evaluation method and device, equipment, medium and program product
CN120372325A