Text sentiment analysis method and system based on natural language processing
Through the text emotion analysis method and system based on natural language processing, the problem of difficulty in analyzing the fluctuations and changes of emotional fluctuations among characters in long texts in the prior art is solved, and the detailed analysis and classification of character emotions in the text is realized, which improves the accuracy and accuracy of the analysis.
Patent Information
- Application Number
- CN202411857855.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively analyze the fluctuations and changes in emotional fluctuations between characters in long texts, and it is impossible to conduct detailed sentiment analysis on smart devices, limiting the display and expansion of information.
The text sentiment analysis method and system based on natural language processing is adopted to identify and classify subjective and objective emotions in the text through steps such as data extraction, data preprocessing, subjective and objective emotions analysis, and then realize detailed analysis of character emotions.
It improves the accuracy and pertinence of text sentiment analysis, enhances the classification accuracy of text in public opinion products, and can more effectively analyze and display the role emotions in the text.
Smart Images

Figure CN120045970A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text intelligence technology, and particularly to a text sentiment analysis method and system based on natural language processing. Background Art
[0003] In domestic research, the analysis of natural language processing for long texts is gradually increasing. However, research on the emotional fluctuations between characters in event analysis and research on the four great classical novels is still relatively scarce. Moreover, when organizing and displaying the required information through intelligent devices, the information displayed on the intelligent device screen is limited. It can only analyze the subjective structure of the text through the sentiment analysis database, unable to analyze the emotions of different characters towards the searched character, and unable to provide users with sufficient extended thinking. Therefore, the present invention provides a text sentiment analysis method and system based on natural language processing. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of the low functionality of the existing technology text system in assisting in searching text information.
[0005] To achieve the above objective, the present invention adopts the following technical solutions: A text sentiment analysis method and system based on natural language processing, including: collecting text information extraction based on the data, and then subsequent operations can be coordinated with the intelligent device. The data extraction can generate data cleaning for data preprocessing, cleaning subjective words and objective words irrelevant to the character stories in the text documents in the data preprocessing, comparing unnecessary subjective feature sentences in the words in the data cleaning to distinguish the character features in the book. In response to the necessity of the irrelevant subjective words in the data cleaning, it is determined that the data preprocessing requires data cleaning of unnecessary subjective feature sentences, and it is determined that the data preprocessing sets data cleaning. The sentences of the data processing are stored with the retention of the second object sentiment before the necessary features and full stops;
[0006] Based on the data preprocessing and the data cleaning, a corresponding topic model analysis of the data preprocessing is generated. The topic model analysis is coordinated with the data of the data preprocessing to display a concise and intuitive first sentence sentiment analysis result.
[0007] As a preferred implementation, sentiment analysis is generated based on the topic model analysis. The personal information of the text characters in the topic model analysis is compared. The processing of the text sentences in the topic model analysis is relatively cumbersome. In response to the large amount of sentence information in the topic model analysis, it is determined that the topic model analysis requires concise sentiment analysis, and it is determined that the sentiment analysis makes the sentence analysis concise and intuitive.
[0008] As a preferred embodiment, after the step of determining the second sentiment analysis result corresponding to the to-be-analyzed statement according to the second sentiment statement that has not been cleaned by data cleaning, the method further includes: determining the second sentiment statement analysis result according to the screening result of the Python program based on the hash value corresponding to the statement;
[0009] Input the key-value pair of the second sentiment statement analysis result into the sentiment database;
[0010] After receiving the key-value of the second sentiment statement analysis result, the sentiment database updates the second sentiment statement result.
[0011] As a preferred embodiment, the step of determining the sentiment analysis result of the to-be-analyzed text according to the first sentiment statement analysis and the second sentiment statement analysis includes: matching the sentiment results of each statement in the first sentiment statement analysis to obtain the matching data of the first sentiment statement analysis result;
[0012] Determine the first sentiment analysis result of the to-be-analyzed text by using the sentiment analysis results of the statements in the sentiment database that match similar or identical statements.
[0013] As a preferred embodiment, count the corresponding text, obtain the positions of the full stops in the corresponding text, and determine all the statements included in the corresponding text and calculate the lengths of the statements according to the positions of the full stops in the corresponding text.
[0014] As a preferred embodiment, data preprocessing is used to determine the first sentiment statement analysis result according to each sentiment statement in the corresponding text and a preset sentiment database, where the preset relationship between the preset statements and the sentiments is stored in the sentiment database;
[0015] Topic model analysis is used to determine the to-be-analyzed statements in each statement according to the first sentiment analysis result;
[0016] Sentiment analysis is used to determine the sentiment analysis result of the corresponding text according to the first sentiment analysis result and the second sentiment analysis result;
[0017] The data preprocessing is also used for: performing discrimination processing on the corresponding text to determine each statement constituting the corresponding text; performing hash operations on the statements respectively to obtain hash values corresponding to the statements, inputting the hash values corresponding to the first sentiment statements into a preset sentiment database respectively, determining whether the hash values are included in the sentiment database, regarding the statements not including the hash values as second sentiment statements, determining the judgment results as the first sentiment analysis result and the second sentiment analysis result, and processing the second sentiment analysis result through a Jieba program to determine the second sentiment statement analysis result.
[0018] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0019] The present invention provides a text sentiment analysis method and system based on natural language processing. The method first obtains data extraction, and then through data preprocessing, performs steps such as word segmentation, stop word removal, and part-of-speech tagging for the description of specific characters in "Water Margin", and finally based on sentiment analysis and three types of character sentiment classification of positive, negative, and neutral, obtains the sentiment types corresponding to the sentiment analysis, can identify the subjective and objective sentiment types of the sentiment analysis text, enables each character in "Water Margin" to utilize the corresponding theme model to analyze the detailed sentiment types, effectively improves the accuracy and pertinence of text sentiment analysis, and thus improves the classification accuracy of the text in the public opinion products. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a flowchart of a text sentiment analysis method based on natural language processing provided by the present invention:
[0022] Figure 2 It is a flowchart of step S102 in the text sentiment analysis method based on natural language processing provided by the present invention:
[0023] Figure 3 It is a flowchart of step S103 in the text sentiment analysis method based on natural language processing provided by the present invention:
[0024] Figure 4 It is a flowchart of step S104 in the text sentiment analysis method based on natural language processing provided by the present invention.
[0025] Reference signs: S101, data extraction module; S102, data preprocessing module; S103, topic model analysis module; S104, sentiment analysis module. Detailed implementation manners
[0026] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] The present invention will be further described below with reference to the embodiments.
[0028] Embodiment: Refer to Figures 1 to 4 , the present invention provides a technical solution: a text sentiment analysis method and system based on natural language processing, including: collecting text information based on the data, and then subsequent operations can be cooperated with an intelligent device. The data extraction can generate data cleaning of data preprocessing, and subjective words and objective words irrelevant to the character stories in the text documents in the data preprocessing are cleaned. Unnecessary subjective feature sentences and phrases appearing in the words in the data cleaning are compared to distinguish the character features in the book. In response to the necessity of the irrelevant subjective words in the data cleaning, it is determined that the data preprocessing needs to clean unnecessary subjective feature sentences and phrases, and it is determined that the data preprocessing sets data cleaning. The sentences of the data processing are reserved for the second object sentiment before the necessary features and full stops.
[0029] Based on the data preprocessing and the data cleaning, a topic model analysis corresponding to the data preprocessing is generated. The topic model analysis cooperates with the data of the data preprocessing, and then displays a concise and intuitive first sentence sentiment analysis result.
[0030] As Figure 1 shown, sentiment analysis is generated based on the topic model analysis. The personal information of the text characters in the topic model analysis is compared. The text sentences processed by the topic model analysis are relatively cumbersome. In response to the large amount of sentence information in the topic model analysis, it is determined that the topic model analysis needs to be concise for sentiment analysis, and it is determined that the sentiment analysis makes the sentence analysis concise and intuitive.
[0031] In actual operation, the corresponding literary masterpieces, such as *Water Margin* among the four great classical novels, can be extracted from online resources, book scans or electronic versions. At the same time, if social network analysis is to be carried out, data on character relationships also need to be collected. These data can be obtained through Internet searches or text mining tools. Then, when searching for a character in *Water Margin*, the descriptions of the character's personal character and emotions are screened through the data preprocessing and topic model analysis modules, and then the first-person description results of the character are obtained through the sentiment analysis module;
[0032] In the data preprocessing module, first convert the text content into a hash value, and then match according to the hash value to be searched. Since some full stops or other symbols in *Water Margin* distinguish character descriptions, in order to facilitate the subsequent sentiment analysis module to obtain the first sentiment analysis conclusion after hash value matching, the hash values and non-hash values containing a part of the same content after the full stop can be input into the Python program first. The statement symbols for searching characters through others' descriptions are retained and then stored as the second sentiment analysis. The rest are implanted into data cleaning for cleaning, which can reduce the system extraction and analysis time, and at the same time, the second sentiment analysis can also be additionally analyzed, facilitating readers to expand their thinking and not being restricted by the analysis conclusion of the sentiment analysis module, thereby improving the practicality of the system.
[0033] Such as Figure 2 And Figure 3 As shown, after the step of determining the second sentiment analysis result corresponding to the statement to be analyzed according to the second sentiment statement that has not been cleaned by data cleaning, the method further includes: determining the second sentiment statement analysis result according to the screening result of the hash value corresponding to the statement through the Python program, and inputting the key-value pair of the second sentiment statement analysis result into the sentiment database. After receiving the key of the second sentiment statement analysis result, the sentiment database updates the second sentiment statement result. The step of determining the sentiment analysis result of the text to be analyzed according to the first sentiment statement analysis and the second sentiment statement analysis includes: matching the sentiment results of each statement in the first sentiment statement analysis to obtain the matching data of the first sentiment statement analysis result, and determining the sentiment analysis result of the statement in the sentiment database that matches similar or identical statements as the first sentiment analysis result of the text to be analyzed.
[0034] After the second emotion is screened out by the Python program, in order to further improve the description of the characters in the search text in the second person or the third person, the specific word segmentation (jieba program) is used. Word segmentation is an important task in natural language processing (NLP). Its main function is to segment continuous Chinese text into individual words according to certain rules. For example, for the sentence "I love natural language processing", Jieba can segment it into "I / love / natural language / processing" and perform secondary screening on the specific word segmentation after the corresponding hash value, so as to analyze a more objective description of the second emotion. For example, the sentence analysis result is obtained through the sentiment analysis module by hash value matching. The descriptions of other characters after the symbol will be described by specific analysis, so as to remove the ones without specific word segmentation, simplify and improve the accurate analysis of the second emotion sentence analysis result.
[0035] As Figure 4 As shown, count the corresponding text, obtain the positions of the full stops in the corresponding text, and determine all the sentences included in the corresponding text and calculate the lengths of the sentences according to the positions of the full stops in the corresponding text. Data preprocessing is used to determine the first emotion sentence analysis result according to each emotion sentence in the corresponding text and a preset emotion database, where the preset sentence and emotion correspondence relationship is stored in the emotion database.
[0036] Since the text to be analyzed contains multiple sentences and it cannot be guaranteed that the emotion analysis results expressed by each sentence are the same, it is necessary to count all the sentences. For example, a series of emotions such as "happy", "sad", "negative", "positive", etc. involved in the emotion analysis result can be processed by the (jieba program). The emotion word that appears the most times indicates that the emotion to be expressed in this text is the most, so it can be used as the final emotion analysis result of the text to be analyzed, while the ones with fewer occurrences can be used as a reference for readers for divergent thinking to avoid restricting the readers' divergent thinking.
[0037] Operation steps: First, extract from network resources, book scans or electronic versions. At the same time, if social network analysis is carried out, data on the relationships between people also needs to be collected. When searching for one of the characters in "Water Margin", the description of the character's personal character and emotion is screened by data preprocessing and the topic model analysis module, and then the first-person description result of the character is obtained through the emotion analysis module;
[0038] After that, the part of the hash value and the input without the hash value after the full stop are input into the Python program. The statement symbols for searching roles described by others are retained, and then stored as the second sentiment analysis. The rest of the implanted data is cleaned up, which can reduce the system extraction and analysis time. At the same time, the second sentiment analysis can also be additionally analyzed, which is convenient for readers to expand their thinking and will not be restricted by the analysis conclusion of the sentiment analysis module, improving the practicality of the system.
[0039] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text sentiment analysis method based on natural language processing, characterized in that: include: Based on the data extraction text information collection, subsequent operations can be carried out in conjunction with the operation of intelligent devices. The data extraction can generate data cleaning for data preprocessing, clean up subjective words and objective words that are irrelevant to the character story in the text document in the data preprocessing, compare unnecessary subjective characteristic words and sentences that appear in the words in the data cleaning to distinguish the characteristics of the characters in the book, and in response to the necessity of the data cleaning irrelevant subjective words, determine that the data preprocessing needs to clean up unnecessary subjective characteristic words and sentences, determine that the data preprocessing is set to clean up data, and the data processing sentences are stored in the second object emotion before the necessary features and the period; Based on the data preprocessing and the data cleaning, a topic model analysis corresponding to the data preprocessing is generated, and the topic model analysis is combined with the data of the data preprocessing to display a concise and intuitive first sentence sentiment analysis result.
2. According to the method for text sentiment analysis based on natural language processing described in claim 1, sentiment analysis is generated based on the topic model analysis, the text character personal information is analyzed by the topic model, and the text sentences processed by the topic model analysis are more complicated. In response to the fact that the topic model analysis has more sentence information, it is determined that the topic model analysis is concise and requires sentiment analysis, and it is determined that the sentiment analysis makes the sentence analysis concise and intuitive.
3. The text sentiment analysis method based on natural language processing according to claim 1, characterized in that: After the step of cleaning the uncleaned second emotional sentence according to the data and determining the second emotional analysis result corresponding to the sentence to be analyzed, the method further comprises: determining the second emotional sentence analysis result according to the result of screening the hash value corresponding to the sentence by the Python program; Inputting the second emotion statement analysis result key-value pair into the emotion database; After receiving the second emotion sentence analysis result key value, the emotion database updates the second emotion sentence result.
4. The text sentiment analysis method based on natural language processing according to claim 1, characterized in that: The step of determining the sentiment analysis result of the text to be analyzed according to the first sentiment sentence analysis and the second sentiment sentence analysis comprises: matching the sentiment results of each of the sentences in the first sentiment sentence analysis to obtain matching data of the first sentiment sentence analysis result; The sentiment analysis result of the sentence that is similar or identical to the sentence in the sentiment database is determined as the first sentiment analysis result of the text to be analyzed.
5. The text sentiment analysis method based on natural language processing according to claim 1, characterized in that: The corresponding texts are counted, and the positions of periods in the corresponding texts are obtained. According to the positions of the periods in the corresponding texts, all sentences included in the corresponding texts are determined and the lengths of the sentences are calculated.
6. The text sentiment analysis system based on natural language processing according to claim 1, characterized in that: Data preprocessing, for determining a first emotional sentence analysis result according to each emotional sentence in the corresponding text and a preset emotional database, wherein the emotional database stores a correspondence between preset sentences and emotions; Topic model analysis, used to determine the sentences to be analyzed in each of the sentences according to the first sentiment analysis result; Sentiment analysis, used to determine a sentiment analysis result of the corresponding text according to the first sentiment analysis result and the second sentiment analysis result; The data preprocessing is also used to: distinguish the corresponding text and determine the various sentences that constitute the corresponding text; perform hash operations on the sentences respectively to obtain the hash values corresponding to the various sentences, input the hash values corresponding to the first emotional sentences into a preset emotional database respectively, judge whether the emotional database contains the hash values, and regard the sentences that do not contain the hash values as the second emotional sentences, and determine the judgment results as the first emotional analysis results and the second emotional analysis results, and process the second emotional analysis results through the stuttering program to determine the second emotional sentence analysis results.