A method and device for text sentiment analysis

By determining the tendency strength values ​​of sentiment words in the corpus and extracting third sentiment words with similar syntactic structures, an extended dictionary is generated, which solves the problems of inaccurate results and insufficient coverage in Chinese text sentiment analysis and achieves higher accuracy and comprehensiveness.

CN115098636BActive Publication Date: 2025-09-09SHENZHEN TAIJI SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210749866.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-09-09
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

When performing sentiment analysis on Chinese text, existing technologies tend to ignore key feature information, resulting in inaccurate results, high dependence on sample data, and insufficient coverage.

Method used

By determining the emotional tendency intensity value according to the number of occurrences of emotional words in the corpus, extracting third emotional words with similar syntactic structure, generating an extended second emotional dictionary, and combining the emotional polarity probability to judge the emotional polarity of the test text.

Benefits of technology

It improves the accuracy and coverage of text sentiment analysis and ensures the accuracy and comprehensiveness of sentiment polarity analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115098636B_ABST
    Figure CN115098636B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of natural language processing technology, and provides a method and device for text sentiment analysis. The method includes: determining the sentiment tendency strength value of a first sentiment word based on the number of times the first sentiment word appears in a corpus; extracting a third sentiment word having a similar syntactic structure to the second sentiment word from the corpus based on the second sentiment word in the first sentiment dictionary; generating a second sentiment dictionary based on the sentiment tendency strength value and the third sentiment word, the second sentiment dictionary including the first sentiment word, the second sentiment word and the third sentiment word; analyzing the sentiment polarity of the test text based on the second sentiment dictionary. The present application can improve the coverage and accuracy of sentiment polarity analysis of the test text based on the sentiment dictionary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of natural language processing technology, and in particular relates to a method and device for text sentiment analysis. Background Art

[0002] With the development of internet technology, a vast amount of textual information, including user sentiment, has been aggregated across various online platforms, encompassing content such as criticisms of products or preferences. Capturing the emotions embedded in these texts in a timely manner would be crucial for both business operations and academic research, enabling applications such as product analysis and recommendation, and consumer spending forecasting. Therefore, textual sentiment analysis technology is currently an active research area with ample room for development.

[0003] Sentiment analysis is the process of analyzing, processing, and extracting emotionally charged, subjective text using natural language processing and text mining techniques. Currently, sentiment analysis techniques can be categorized into paragraph-level, sentence-level, and word-level sentiment analysis, depending on the level of sentiment analysis performed.

[0004] Among them, word-level text sentiment analysis is to identify the sentiment words included in the text and infer the sentiment of the entire text based on the emotional tendencies expressed by these sentiment words.

[0005] However, due to the high complexity of language expression, especially when analyzing Chinese text, where words are used in a very flexible and diverse manner, word-level sentiment analysis can easily produce biased results. Existing techniques for natural language processing may overlook key text features, resulting in inaccurate sentiment analysis results.

[0006] In addition, existing technologies usually require a large amount of sample data to support text sentiment analysis and have a high dependence on samples, which can lead to insufficient coverage when analyzing the text to be tested. Summary of the Invention

[0007] In view of this, the present application provides a method and apparatus for text sentiment analysis, which can improve the coverage and accuracy of text sentiment analysis based on sentiment dictionaries.

[0008] In a first aspect, an embodiment of the present application provides a method for text sentiment analysis, comprising:

[0009] Determining the sentiment tendency intensity value of the first sentiment word according to the number of times the first sentiment word appears in the corpus;

[0010] Extracting a third sentiment word from the corpus based on the second sentiment word, where the second sentiment word refers to a word in the first sentiment dictionary, and the third sentiment word has a similar syntactic structure to the second sentiment word;

[0011] generating a second emotional dictionary according to the emotional tendency strength value and the third emotional word, wherein the second emotional dictionary includes the first emotional word, the second emotional word, and the third emotional word;

[0012] The sentiment polarity of the test text is analyzed according to the second sentiment dictionary.

[0013] The embodiment of the present application determines the emotional tendency intensity value of the first emotional word by the number of occurrences of the first emotional word in the corpus, thereby improving the accuracy of determining the emotional intensity expressed by the emotional word. In addition, the embodiment of the present application extracts the third emotional word from the corpus and generates a second emotional dictionary based on the emotional tendency intensity value and the third emotional word, thereby expanding the coverage of the second emotional dictionary. Therefore, when analyzing the emotional polarity of the test text according to the second emotional dictionary, the embodiment of the present application can improve the accuracy and coverage of the text emotional polarity analysis results.

[0014] In a possible implementation, before analyzing the sentiment polarity of the test text according to the second sentiment dictionary, the method further includes:

[0015] Determining the sentiment polarity probability of the text to be tested;

[0016] The step of analyzing the sentiment polarity of the text to be tested according to the second sentiment dictionary includes:

[0017] The sentiment polarity of the text to be tested is analyzed according to the sentiment polarity probability and the second sentiment dictionary.

[0018] In a possible implementation, the first sentiment words include positive sentiment words and negative sentiment words, and the corpus includes positive sentiment corpus and negative sentiment corpus;

[0019] The step of determining the sentiment tendency strength value of the first sentiment word according to the number of times the first sentiment word appears in the corpus includes:

[0020] Determine the sentiment tendency intensity value of the positive sentiment word according to one or more of the number of occurrences of the positive sentiment word when expressing positive semantics in the positive sentiment corpus, the number of occurrences of the positive sentiment word when expressing positive semantics in the negative sentiment corpus, the number of occurrences of the positive sentiment word in the negative sentiment corpus, and the sum of the number of occurrences of all sentiment words expressing positive semantics in the positive sentiment corpus;

[0021] The emotional tendency intensity value of the negative emotional word is determined based on one or more of the number of times the negative emotional word appears when expressing negative semantics in the negative emotional corpus, the number of times the negative emotional word appears when expressing negative semantics in the positive emotional corpus, the number of times the negative emotional word appears in the positive emotional corpus, and the sum of the number of times all emotional words expressing negative semantics appear in the negative emotional corpus.

[0022] In a possible implementation, the sentiment tendency strength value of the positive sentiment word satisfies the following formula:

[0023]

[0024] Among them, t i is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of positive semantics in the positive sentiment corpus, pwords is the set of all positive sentiment words, is the sum of the occurrence times of all sentiment words expressing positive semantics in the positive sentiment corpus, t i The number of occurrences of positive semantics in the negative sentiment corpus, t i The number of occurrences in the negative sentiment corpus;

[0025] The sentiment tendency strength value of the negative sentiment word satisfies the following formula:

[0026]

[0027] Among them, t i is the negative sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of negative semantics in the negative sentiment corpus, nwords is the set of all negative sentiment words, is the sum of the occurrence times of all sentiment words expressing negative semantics in the negative sentiment corpus, t i The number of occurrences of negative semantics in the positive sentiment corpus, t i The number of occurrences in the positive sentiment corpus.

[0028] In a possible implementation, extracting a third sentiment word from the corpus according to the second sentiment word includes:

[0029] Performing syntactic analysis on the text in the corpus to obtain syntactic analysis results;

[0030] dividing the text into a set of short sentences;

[0031] Determine, according to the second sentiment word, a first short sentence in which the second sentiment word is located, wherein the first short sentence is a short sentence in the short sentence set;

[0032] Annotating the second sentiment word and the first short sentence to obtain a syntactic structure annotation result;

[0033] Determining a second short sentence and the third sentiment word according to the syntactic analysis result and the syntactic structure annotation result, wherein the second short sentence is a short sentence in which the third sentiment word is located in the short sentence set, and the second short sentence has a similar syntactic structure to the first short sentence, wherein the number of occurrences of the third sentiment word in the corpus is greater than a first threshold;

[0034] Determining the emotional tendency of the third emotional word according to the second emotional word, wherein the emotional tendency includes a positive emotional tendency and a negative emotional tendency;

[0035] According to the emotional tendency of the third emotional word, the emotional tendency strength value of the third emotional word is determined.

[0036] It should be understood that, based on the similarity in syntactic structure between the second and third sentiment words, the present embodiment acquires the third sentiment word from the corpus by annotating the second sentiment word and the first short sentence. The second sentiment word generated by the present embodiment includes the first, second, and third sentiment words. Therefore, the present embodiment effectively expands the coverage of the second sentiment word dictionary.

[0037] In a possible implementation, determining the emotional tendency of the third emotional word according to the second emotional word includes:

[0038] Determine a sentiment word graph according to the co-occurrence relationship between the second sentiment word and the third sentiment word in the same text, wherein the sentiment word graph includes a positive sentiment subgraph and a negative sentiment subgraph;

[0039] Determine a first separation cost and a second separation cost, wherein the first separation cost refers to the separation cost of the third emotion word and the positive emotion subgraph, and the second separation cost refers to the separation cost of the third emotion word and the negative emotion subgraph;

[0040] comparing the first separation cost and the second separation cost;

[0041] The sentiment tendency corresponding to the sentiment subgraph having the largest separation cost between the first separation cost and the second separation cost is determined as the sentiment tendency of the third sentiment word.

[0042] In a possible implementation, the first separation cost satisfies the following formula:

[0043]

[0044] Among them, SepCost is the first separation cost, s i is the emotional tendency strength value of the second emotional word in the positive emotional subgraph, G is the positive emotional subgraph, d i represents the co-occurrence times of the second sentiment word and the third sentiment word in the positive sentiment subgraph, and |G| is the number of the second sentiment words in the positive sentiment subgraph.

[0045] In a possible implementation, the second separation cost satisfies the following formula:

[0046]

[0047] Among them, SepCost is the second separation cost, s i is the emotional tendency strength value of the second emotional word in the negative emotional subgraph, G is the negative emotional subgraph, d i represents the number of co-occurrences of the second sentiment word and the third sentiment word in the negative sentiment subgraph, and |G| is the number of the second sentiment words in the negative sentiment subgraph.

[0048] Since the embodiment of the present application provides different calculation methods for the tendency strength values ​​of positive sentiment words and the sentiment tendency strength values ​​of negative sentiment words, the embodiment of the present application provides a calculation method for calculating the sentiment tendency strength value of the third sentiment word by proposing the first separation cost, thereby further improving the second sentiment dictionary.

[0049] In a possible implementation, the sentiment polarity probability includes a positive sentiment polarity probability, a neutral sentiment polarity probability, and a negative sentiment polarity probability. Before analyzing the sentiment polarity of the test text based on the sentiment polarity probability and the second sentiment dictionary, the method further includes:

[0050] Determining a fuzzy prediction degree according to the sentiment polarity probability, wherein the fuzzy prediction degree is used to represent the degree of dispersion between different sentiment polarity probabilities of the text to be tested;

[0051] Determining whether it is necessary to perform a first processing on the sentiment polarity of the text to be tested according to the fuzzy prediction degree;

[0052] The step of analyzing the sentiment polarity of the text to be tested according to the sentiment polarity probability and the second sentiment dictionary includes:

[0053] When the fuzzy prediction degree meets a preset condition, performing a first processing on the sentiment polarity of the text to be tested, wherein the first processing includes determining the sentiment polarity of the text to be tested according to the sentiment dictionary;

[0054] When the fuzzy prediction degree does not meet the preset conditions, the sentiment polarity corresponding to the largest sentiment polarity probability among the positive sentiment polarity probability, the neutral sentiment polarity probability and the negative sentiment polarity probability is determined as the sentiment polarity of the text to be tested.

[0055] It should be understood that the embodiment of the present application uses fuzzy prediction to determine whether the text to be tested needs to be subjected to the first processing, and the fuzzy prediction is used to represent the degree of discreteness between the probabilities of different sentiment polarities of the text to be tested. Therefore, when the fuzzy prediction meets the preset conditions, the sentiment polarity of the text to be tested is further analyzed through the first processing. The first processing is to determine the sentiment polarity of the text to be tested based on the sentiment dictionary, and this process is used to improve the accuracy of text sentiment polarity analysis.

[0056] In a possible implementation, determining the emotional tendency of the text to be tested according to the second emotional dictionary includes:

[0057] Determine a first sentiment value of a sentence in the text to be tested according to the second sentiment dictionary;

[0058] Determining a second sentiment value of the text to be tested based on the first sentiment value and a weight, wherein the weight refers to the weight of the sentence corresponding to the first sentiment value in the text to be tested;

[0059] Determine the sentiment polarity of the text to be tested according to the threshold range in which the second sentiment value is located, wherein:

[0060] When the second sentiment value is within the first threshold range, the sentiment polarity of the text to be tested is positive;

[0061] When the second sentiment value is within a second threshold range, the sentiment polarity of the text to be tested is neutral sentiment polarity;

[0062] When the second sentiment value is within the third threshold range, the sentiment polarity of the text to be tested is negative sentiment polarity.

[0063] In a possible implementation, the second emotion value satisfies the following formula:

[0064]

[0065] Wherein, f2 is the second sentiment value, n is the number of sentences in the text to be tested, α i is the weight of the i-th sentence in the test text, n i is the number of sentiment words in the i-th sentence in the text to be tested, n ij is the number of times negative words appear in the short sentence of the jth sentiment word in the i-th sentence in the test text, c ij is the maximum degree value of the degree adverbs in the short sentence containing the jth sentiment word in the i-th sentence of the test text, s ij is the emotional tendency strength value of the jth emotional word in the i-th sentence in the text to be tested.

[0066] In a second aspect, an embodiment of the present application provides a device for text sentiment analysis, for executing the method in the above-mentioned first aspect or any possible implementation of the first aspect. Specifically, the device includes a module (or unit) for executing the method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0067] In a third aspect, an embodiment of the present application provides a device for text sentiment analysis, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, for executing the method in any possible implementation of the first aspect described above.

[0068] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, it can implement the method in any possible implementation of the first aspect above.

[0069] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program runs on a text sentiment analysis device, it enables the text sentiment analysis device to execute the method in any possible implementation of the first aspect above.

[0070] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0071] The beneficial effects of the embodiments of the present application compared to the prior art are as follows: the present application proposes a method for calculating the emotional tendency intensity value, which is achieved by combining the number of occurrences of the first emotional word in the corpus, whether the corpus is a positive emotional corpus or a negative emotional corpus, and whether the first emotional word expresses positive semantics or negative semantics, so as to achieve a method for calculating the emotional tendency intensity value when representing the emotional intensity of the first emotional word with high accuracy. At the same time, the embodiment of the present application also proposes a method for extracting the third emotional word based on the second emotional word, extracting the third emotional word from the corpus based on the similarity of the syntactic structure, and improving the coverage when performing text emotional polarity analysis based on the second emotional dictionary. In addition, the present application also proposes a method for performing text emotional polarity analysis in combination with the second emotional dictionary, determining the fuzzy prediction degree through the text emotional polarity probability to determine whether the first processing is required for the text to be tested, and determining the emotional polarity of the text to be tested in combination with the second emotional dictionary, and improving the accuracy when performing text emotional polarity analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the present application;

[0073] Figure 2 This is a flowchart of a text sentiment analysis process provided by an embodiment of the present application;

[0074] Figure 3 This is a flow chart of extracting new sentiment words provided by an embodiment of the present application;

[0075] Figure 4 This is a schematic diagram of the result of marking the syntactic structure, old sentiment words, and short sentences containing the old sentiment words provided by the embodiment of the present application;

[0076] Figure 5 This is a flow chart of performing sentiment polarity analysis on a text based on a second sentiment dictionary provided in an embodiment of the present application;

[0077] Figure 6 This is a flowchart of a text sentiment analysis process provided by an embodiment of the present application;

[0078] Figure 7 is a schematic block diagram of an apparatus for performing text sentiment analysis provided in an embodiment of the present application;

[0079] Figure 8 This is a schematic diagram of the structure of the device for performing text sentiment analysis provided in an embodiment of the present application. DETAILED DESCRIPTION

[0080] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0081] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0082] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0083] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0084] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0085] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "comprise," "include," "have," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0086] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0087] The text sentiment analysis method provided in the embodiment of the present application is applicable to the emotions expressed in texts composed of language and characters, and is specifically used to analyze whether the text expresses negative, positive or neutral emotions, without limiting the source of the text and the language used.

[0088] In a possible application scenario, the text sentiment analysis method provided in the embodiment of the present application is suitable for analyzing text from a terminal device, and the terminal device is connected to a server via a network. Figure 1 This is a schematic diagram of the application scenario of the embodiment of the present application. Figure 1 This possible embodiment is described below.

[0089] like Figure 1 As shown, the application scenario may include terminal devices, networks and servers. Figure 1 Mobile phones and computers are used as examples of terminal devices.

[0090] The terminal device is used to provide the text to be tested, and the terminal includes but is not limited to mobile phones, tablet computers, desktop computers, laptop computers and handheld computer devices.

[0091] The network is used as a medium for providing a communication link between a terminal device and a server. The network may include various connection types, such as wired and / or wireless communication links, etc.

[0092] The server is used to perform sentiment polarity analysis on the text to be tested.

[0093] Specifically, the terminal device establishes a communication connection with the server through the network. The server obtains the text in the terminal device, performs sentiment analysis on the text data, and obtains the data of each specific object in the text. The server then analyzes the sentiment polarity of the text through the data in the text and displays the analysis results so that the terminal can obtain the corresponding display results through the network.

[0094] It should be understood that Figure 1 The scenario in the example is only an example of an embodiment of the present application, but the embodiment of the present application is not limited thereto. Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0095] Optionally, the text sentiment analysis method provided in the embodiments of the present application can be executed by a server. Accordingly, the text sentiment analysis device provided in the embodiments of the present application can generally be set in a server. It should also be understood that the server is only an exemplary execution entity, and the execution entity of the embodiments of the present application is not limited to this, and may also include other devices capable of executing the text sentiment analysis method, such as cloud servers and terminals.

[0096] The following combination Figure 2 A detailed description is given of the text sentiment analysis method proposed in the embodiments of the present application. Figure 2 This is a flowchart of the text sentiment analysis process of this application. Figure 2 As shown, the method includes the following steps:

[0097] Step 201: Determine the sentiment tendency strength value of the first sentiment word according to the number of times the first sentiment word appears in the corpus.

[0098] It should be understood that the first sentiment word introduced here is only used for differentiation. The first sentiment word is used to generally refer to the sentiment word targeted when determining the sentiment tendency strength value in the embodiment of the present application. Specifically, the first sentiment word includes any word that can adopt the sentiment tendency strength value calculation method provided by the embodiment of the present application. For ease of description, "sentiment word" may be used in the following text to express the first sentiment word for which the sentiment tendency strength value is calculated.

[0099] The corpus is used to provide reference text for calculating the sentiment intensity value of sentiment words, and the corpus resources included in the corpus have determined sentiment polarity. The corpus resources can be understood as text or a collection of texts composed of words (including sentiment words). The sentiment polarity includes positive sentiment polarity, negative sentiment polarity, or neutral sentiment polarity.

[0100] For unified explanation here, the terminology involving "positive" in the embodiments of the present application can be equivalent to or replaced by "positive", and the terminology involving "negative" can be equivalent to or replaced by "negative". For example, positive emotional polarity can be expressed as positive emotional polarity; positive emotional corpus can be expressed as positive emotional corpus, etc. For another example, negative emotional polarity can be expressed as negative emotional polarity; negative emotional corpus can be expressed as negative emotional corpus, etc. Similarly, the terminology involving "positive" in the embodiments of the present application can also be equivalent to or replaced by "positive", and the terminology involving "negative" can be equivalent to or replaced by "negative".

[0101] Specifically, the corpus includes positive emotion corpus and negative emotion corpus. Optionally, the corpus may also include neutral emotion corpus.

[0102] The embodiments of the present application do not limit the source of the corpus. Optionally, the corpus may be from an existing corpus. The following provides two ways to obtain a corpus.

[0103] In one possible implementation, a corpus that has been labeled on the network can be downloaded, that is, the corpus resources included in the corpus have been divided into positive sentiment corpus, negative sentiment corpus, and neutral sentiment corpus.

[0104] Here is a unified explanation. Labeling refers to the process of annotating the sentiment polarity of the corpus resources in the corpus. It can be understood that the corpus resources after labeling can be divided into corpora of different sentiment types, for example, positive sentiment corpus, negative sentiment corpus, and neutral sentiment corpus, etc.

[0105] In another possible implementation, crawl the corpus resources on the network, and then manually label them as positive sentiment corpus, negative sentiment corpus, and neutral sentiment corpus to obtain a corpus. Among them, crawling means collecting data on the network.

[0106] It should be understood that the corpus adopted in the embodiments of the present application may be a corpus obtained by any of the above two implementation methods, or a collection of the corpora obtained by the above two implementation methods, and no specific limitation is made thereto.

[0107] Exemplarily, the corpus resources on the network include but are not limited to various data such as Weibo comments, Douban movie reviews, hotel evaluations, user evaluations on takeaway platforms, restaurant evaluations, etc. After crawling these corpus resources, the above-mentioned manual labeling is performed on the corpus resources according to the sentiment polarity of the corpus resources, and the corpus required by the embodiments of the present application can be obtained.

[0108] The sentiment tendency intensity value is used to represent the intensity of the sentiment that a sentiment word can express. For example, the intensity of the sentiment expressed by the sentiment word "extremely gratifying" is obviously greater than the intensity of the sentiment expressed by the sentiment word "satisfactory" literally. Correspondingly, the sentiment tendency intensity value corresponding to the sentiment word "extremely gratifying" is greater than the sentiment tendency intensity value corresponding to the sentiment word "satisfactory".

[0109] In the process of calculating the sentiment tendency intensity value of a sentiment word in the embodiments of the present application, the sentiment tendency intensity value is calculated according to the number of times the sentiment word appears in the corpus.

[0110] Exemplarily, the sentiment tendency intensity value of a sentiment word can be calculated according to the number of times the sentiment word appears in corpora with different sentiment polarities when expressing different meanings of sentiment.

[0111] Optionally, when calculating the sentiment tendency strength value of a sentiment word, whether the sentiment word is used to express positive semantics or negative semantics, the number of occurrences of all positive sentiment words, and the number of occurrences of all negative sentiment words can also be taken into consideration.

[0112] When confirming the emotional tendency strength value of an emotional word, the embodiment of the present application adopts different calculation methods for the emotional tendency strength values ​​of positive emotional words and negative emotional words, where positive emotional words refer to emotional words that express positive emotions, and negative emotional words refer to emotional words that express negative emotions.

[0113] Optionally, as an implementation method, when determining the sentiment tendency strength value of a positive sentiment word, the factors affecting the sentiment tendency strength value of the positive sentiment word include one or more of the following: the number of occurrences of the positive sentiment word when expressing positive semantics in a positive sentiment corpus, the number of occurrences of the positive sentiment word when expressing positive semantics in a negative sentiment corpus, the number of occurrences of the positive sentiment word in a negative sentiment corpus, and the sum of the number of occurrences of all sentiment words expressing positive semantics in a positive sentiment corpus.

[0114] In one possible implementation, the positive sentiment word t i The emotional tendency strength value The calculation is done as follows:

[0115]

[0116] Among them, t i is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of positive semantics in positive sentiment corpus, pwords is the set of all positive sentiment words, is the sum of the occurrence times of all sentiment words expressing positive semantics in the positive sentiment corpus, t i The number of occurrences of positive semantics in negative sentiment corpus, t i The number of occurrences in negative sentiment corpus.

[0117] Optionally, as an implementation method, when determining the sentiment tendency strength value of a negative sentiment word, the factors affecting the sentiment tendency strength value of the negative sentiment word include one or more of the following: the number of occurrences of the negative sentiment word when expressing negative semantics in the negative sentiment corpus, the number of occurrences of the negative sentiment word when expressing negative semantics in the positive sentiment corpus, the number of occurrences of the negative sentiment word in the positive sentiment corpus, and the sum of the number of occurrences of all sentiment words expressing negative semantics in the negative sentiment corpus.

[0118] In one possible implementation, the negative sentiment word t i The emotional tendency strength value and is calculated as follows:

[0119]

[0120] Among them, t i is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of negative semantics in negative sentiment corpus, nwords is the set of all negative sentiment words, is the sum of the occurrence times of all sentiment words expressing negative semantics in the negative sentiment corpus, t i The number of occurrences of negative semantics in positive sentiment corpus, t i The number of occurrences in positive sentiment corpus.

[0121] It is worth noting that the expressions of "positive emotional corpus" and "negative emotional corpus" are used to correspond to "positive emotional words" and "negative emotional words" for easy understanding. Specifically, "positive emotional corpus" can be replaced by the "positive emotional corpus" in the corpus obtained in the embodiment of the present application in the foregoing text, and "negative emotional corpus" can be replaced by the "negative emotional corpus" in the corpus obtained in the embodiment of the present application in the foregoing text.

[0122] In addition, this application does not limit the numerical range of the emotional tendency strength value, or the corresponding relationship between the size of the emotional tendency strength value and the emotional strength.

[0123] For example, in the embodiment of the present application, the emotional tendency strength value of a positive emotional word is a positive number, and the emotional tendency strength value of a negative emotional word is a negative number, and the larger the absolute value of the emotional tendency strength value, the stronger the emotional intensity expressed by the corresponding emotional word.

[0124] It is understood that this is just an example description, and the numerical range of the emotional tendency strength value, the corresponding relationship between the size of the emotional tendency strength value and the emotional intensity is not limited to this. In fact, other corresponding methods may also be achieved by adjusting the formula.

[0125] Optionally, as an implementation method, after obtaining the sentiment tendency strength value, the sentiment tendency strength value may be normalized.

[0126] For example, the normalized emotional tendency strength value of positive emotional words ranges from [0, 1], and the normalized emotional tendency strength value of negative emotional words ranges from [-1, 0]. The following calculation method is used:

[0127]

[0128] Normalized sentiment intensity value of negative sentiment words The following calculation method is used:

[0129]

[0130] It should be noted that in the above-mentioned calculation formula for the intensity value of sentiment tendency, and the calculation formula for the normalized intensity value of sentiment tendency, some parameters contain superscripts and / or subscripts. The value of the superscript is used to characterize the sentiment polarity of the corpus; the value of the subscript is used to characterize the semantic type corresponding to the sentiment word, and the semantic type includes positive semantics and negative semantics. Exemplarily, the value of the superscript can be 0 or 1. If the superscript is 0, it indicates negative sentiment corpus, and if the superscript is 1, it indicates positive sentiment corpus. Exemplarily, the value of the subscript can be 0 or 1. For example, if the value of the subscript is 1, it indicates that the sentiment word has positive semantics, and if the value of the subscript is 0, it indicates that the sentiment word has negative semantics.

[0131] for Among these parameters, for example, The 0 in the superscript is used to represent negative sentiment corpus, so Refers to the word t i The number of occurrences in negative sentiment corpus; for example, The 1 in the superscript represents positive sentiment corpus, and the 1 in the subscript is used to represent positive semantics, so Refers to the word t i In the positive sentiment corpus, it indicates the number of occurrences of positive semantics. It should be understood that this is just a parameter and Taking this as an example, the superscripts and / or subscripts of other parameters can also be understood based on similar principles, which will not be elaborated here for the sake of brevity.

[0132] Positive semantics refers to sentiment words that express positive emotions, while negative semantics refers to sentiment words that express negative emotions. It is worth noting that due to the complexity of language expression, the meaning of sentiment words may change within a text. Therefore, positive sentiment words may be used to express positive or negative semantics, and negative sentiment words may also be used to express positive or negative semantics.

[0133] In order to accurately determine whether a sentiment word expresses positive or negative semantics, the embodiment of the present application uses the sentiment word as a basis, combined with factors in the text that can change the semantics, to jointly determine whether the sentiment word expresses positive or negative semantics. For example, the factor that can change the semantics is a negation word.

[0134] Optionally, when determining whether a sentiment word represents positive or negative semantics in a text, the number of times the negative word is used may be used as a criterion.

[0135] For example, the positive emotional word "like" is used to illustrate the positive and negative semantics involved in the embodiments of this application. "I like this book" does not include a negative word, "like" is used to express a positive meaning; "I don't like this book" includes the negative word "no", and "like" is used to express a negative meaning; "I don't dislike this book" includes the two negative words "no" and "no", and "like" is used to express a positive meaning, that is, double negation expresses affirmation.

[0136] Optionally, a window of length k can be set in the text, and then the number of times a negative word that has a co-occurrence relationship with a sentiment word appears in the window is determined, and the number of times is used to define whether the semantics expressed by the sentiment word is positive or negative. Here, a co-occurrence relationship refers to a relationship in which sentiment words and negative words co-occur within a limited range. For example, when the number of negative words co-occurring with positive sentiment words in the window is odd, the positive sentiment words represent negative semantics; when the number of negative words co-occurring with positive sentiment words in the window is even, the positive sentiment words represent positive semantics. When the number of negative words co-occurring with negative sentiment words in the window is odd, the negative sentiment words represent positive semantics; when the number of negative words co-occurring with negative sentiment words in the window is even, the negative sentiment words represent negative semantics.

[0137] Optionally, the window for determining the number of occurrences of negative words may be set at a position adjacent to the sentiment words that need to be judged.

[0138] In one possible implementation, a window of length k is defined before a sentiment word. Specifically, the number of negative words in the text of length k preceding the sentiment word is determined. K represents the size of the window. The value of K can be a specified value or a default value, which can be defined or preset. The value of k is a positive integer. If no specified value is set, the default value is used, for example, a default value of 3.

[0139] Optionally, when determining the negative words in the text, a negative word list can be constructed in the second sentiment dictionary generated in the embodiment of the present application, based on the negative words in the negative word list. The specific content included in the second sentiment dictionary will be described in detail later.

[0140] Step 202: extracting a third sentiment word from the corpus according to the second sentiment word.

[0141] The second sentiment word refers to the sentiment word in the first sentiment dictionary, or the old sentiment word. The third sentiment word refers to the sentiment word that is not in the first sentiment dictionary but in the corpus, so it can also be said that the third sentiment word is a new sentiment word.

[0142] It should be understood that the second and third sentiment words are only used for differentiation.

[0143] The purpose of step 202 is to extract new emotion words, i.e., third emotion words, from the corpus based on the emotion words already in the first emotion dictionary. Then, based on the second emotion words and the third emotion words, a second emotion dictionary in the embodiment of the present application can be constructed or generated to expand the coverage of the second emotion dictionary. The second emotion dictionary refers to the newly generated emotion dictionary.

[0144] For the convenience of description, the following text may use "existing sentiment words" or "old sentiment words" to express the second sentiment word, and use "new sentiment words" to express the third sentiment word.

[0145] Optionally, the new sentiment words have similar syntactic structures to the existing sentiment words in the sentiment dictionary.

[0146] For example, when a word and an old sentiment word constitute similar components in a sentence, and the sentence containing this word has a similar syntactic structure to the sentence containing the old sentiment word, it can be inferred that this word may also be a sentiment word, and then it can be considered to extract this word as a new sentiment word and include it in the second sentiment dictionary.

[0147] How to analyze the syntactic structure of old and new sentiment words and extract new sentiment words based on the syntactic structure will be discussed later. Figure 2 Give a detailed description of the method of extracting new sentiment words.

[0148] The above-mentioned first emotional dictionary is an emotional lexicon that can be directly obtained or is an existing emotional lexicon. Generally speaking, the first emotional dictionary can include positive emotional words and negative emotional words. In the embodiment of the present application, the first emotional dictionary can be used as the basis for generating the second emotional words, and then new emotional words (i.e., third emotional words) are extracted from the corpus to expand the second emotional dictionary. In other words, the generated second emotional dictionary not only includes the emotional words already in the first emotional dictionary, but also includes the newly extracted emotional words, which can greatly improve the coverage of the second emotional dictionary.

[0149] This application does not limit the source of the first emotion dictionary. The first emotion dictionary can be an existing emotion dictionary, such as the emotion dictionary launched by HowNet, or an emotion dictionary provided on the Internet.

[0150] Step 203: determining a second emotional dictionary according to the emotional tendency strength value and the third emotional word, where the second emotional dictionary includes the first emotional word, the second emotional word, and the third emotional word.

[0151] Step 203 may also be expressed in another way, that is, generating or constructing a second emotion dictionary based on the emotion tendency strength value and the third emotion word.

[0152] Through step 203, not only the second emotional word is used as the emotional word of the second emotional dictionary, but also the new emotional word extracted in step 202 is included in the second emotional dictionary. In this way, the second emotional dictionary not only includes the emotional words already in the first emotional dictionary, but also expands the new emotional words (i.e., the third emotional word), thereby improving the coverage of the dictionary. In addition, each emotional word in the second emotional dictionary may also have a corresponding emotional tendency intensity value. The specific calculation method of the emotional tendency intensity value can refer to the calculation method of the emotional tendency intensity value of the first emotional word in the previous article. For the sake of brevity, it will not be repeated here. By introducing the emotional tendency intensity value, preparation can be made for the subsequent analysis of the emotional polarity of the text to be tested.

[0153] Specifically, the second emotional words are the emotional words in the existing first emotional dictionary. The second emotional dictionary is based on the first emotional dictionary. Therefore, the second emotional words are the old emotional words in the second emotional dictionary. The third emotional words refer to the new emotional words extracted from the corpus, that is, the new emotional words in the second emotional dictionary.

[0154] It is worth noting that the embodiment of the present application can use the second sentiment dictionary for sentiment polarity analysis of the text. Based on the sentiment words in the second sentiment dictionary and the corresponding sentiment tendency strength values, the sentiment expressed in the text can be quantified into numerical values, thereby facilitating the subsequent analysis of the sentiment polarity of the text. Therefore, all sentiment words in the second sentiment dictionary have corresponding sentiment tendency strength values. The first sentiment word refers to the sentiment word for which the sentiment tendency strength value is calculated, that is, the first sentiment word can be the second sentiment word or the third sentiment word.

[0155] Optionally, the second sentiment dictionary includes a positive sentiment word list and a negative sentiment word list. The positive sentiment word list includes one or more positive sentiment words, and each positive sentiment word has a corresponding sentiment tendency strength value. The negative sentiment word list includes one or more negative sentiment words, and each negative sentiment word has a corresponding sentiment tendency strength value. That is, each sentiment word in the second sentiment dictionary corresponds to a specific sentiment tendency strength value.

[0156] Optionally, the second sentiment dictionary constructed in the embodiment of the present application may further include one or more items of a stop word vocabulary, a negation word vocabulary, or a degree adverb vocabulary.

[0157] Stop words are words that can be ignored during natural language processing. In other words, stop words are words that have no effect on analyzing the sentiment expressed in a text.

[0158] Negative words are used to negate the semantics expressed by words of parts of speech such as verbs or adjectives, and are usually used in conjunction with these parts of speech such as verbs or adjectives. When negative words are used together with emotional words, they are used to negate the emotions expressed by the emotional words, thereby changing the meaning of the sentence in which the emotional words are used. For example, a sentence uses a negative word while using a positive emotional word, and this sentence expresses a negative meaning. For example, in the sentence "I am not reconciled", the negative word "not" is used together with the positive emotional word "willing", and "not" negates the positive semantics expressed by "willing", so that this sentence expresses a negative semantics.

[0159] Adverbs of degree are adverbs that limit or modify the degree of an adjective or adverb, and are usually used before the modified adjective or adverb. For example, adverbs of degree include very, extremely, quite, etc. Adverbs of degree do not affect the semantics of the text, but rather describe the degree of semantics. When adverbs of degree are used in conjunction with emotional words, the adverbs of degree will increase or decrease the intensity of the emotion expressed by the corresponding emotional words. For example, the use of the adverb of degree "extremely" in the sentence "He is extremely sad" makes the negative emotion expressed in the sentence stronger than the negative emotion expressed by "sadness" alone.

[0160] It should be understood that this application does not limit the source of the stop word list, negation word list or degree adverb list, which can be derived from online resources or set by the user.

[0161] Step 204: Analyze the sentiment polarity of the text to be tested according to the second sentiment dictionary.

[0162] The embodiment of the present application analyzes the sentiment polarity of the text based on the second sentiment dictionary. Moreover, the sentiment words included in the second sentiment dictionary have sentiment tendency strength values, which can provide a more sufficient basis for analyzing the sentiment polarity of the text.

[0163] In the embodiment of the present application, by calculating the sentiment intensity value of the sentiment word, the accuracy of analyzing the sentiment intensity expressed by the sentiment word can be improved; and by extracting the third sentiment word (i.e., the new sentiment word) from the second sentiment word to construct a second sentiment dictionary, so as to expand the sentiment words included in the second sentiment dictionary. Therefore, when performing text sentiment polarity analysis based on the second sentiment dictionary, the accuracy and coverage of the text sentiment polarity analysis method can be effectively improved.

[0164] For text sentiment polarity analysis, the existing technology uses a trained model to analyze the sentiment polarity of the text. However, this method has a high dependence on the samples used during training and cannot take into account contextual information, resulting in limited accuracy. In addition, the text sentiment analysis method based on deep learning adopted by the existing technology requires a large amount of data to support it, and its accuracy is limited. Compared with the existing text sentiment analysis method based on deep learning, the sentiment tendency intensity value calculation method provided by the embodiment of the present application is based on the number of occurrences of sentiment words, which is more accurate and takes into account contextual information by counting the number of times sentiment words express positive and negative semantics. In addition, when performing text sentiment polarity analysis, the embodiment of the present application combines the sentiment words and corresponding sentiment tendency intensity values ​​in the second sentiment dictionary. Therefore, under the same data scale, the embodiment of the present application has a higher accuracy rate when analyzing text sentiment polarity than the existing technology. At the same time, by extracting the third sentiment word, the number of sentiment words in the second sentiment dictionary generated by the embodiment of the present application is higher than that of the dictionary commonly used in the existing technology, so it has a wider coverage when analyzing text sentiment polarity.

[0165] Furthermore, in the process of extracting the third sentiment word from the corpus according to the second sentiment word involved in step 202, the embodiment of the present application also provides a method for extracting new sentiment words, so as to expand the coverage of sentiment polarity analysis of the test text and achieve further improvement of the text sentiment analysis method. Figure 3 Schematic diagram of the process of extracting new sentiment words.

[0166] For ease of description and understanding, Figure 3 When explaining, "old emotional words" are used to represent "second emotional words", and "new emotional words" are used to represent "third emotional words".

[0167] Optionally, as an embodiment, step 202 may include: Figure 3 As shown in the steps. Figure 3 As shown, the details are as follows:

[0168] Step 301: Perform syntactic analysis on the text in the corpus to obtain syntactic analysis results.

[0169] Obtaining the syntactic analysis result refers to obtaining the part of speech corresponding to each word in the text.

[0170] Specifically, the object of syntactic analysis includes each text in the corpus.

[0171] Optionally, the text may be segmented, and different parts of speech of each word may be obtained and marked to perform syntactic analysis on the text.

[0172] The embodiment of the present application does not limit the tools used for word segmentation and part-of-speech tagging. In one possible implementation, the part-of-speech tagging tool posseg of the jieba library of Python is used to segment the text, obtain the part of speech corresponding to each word in the segmentation result, and tag each character with the part of speech.

[0173] For example, the B(I)-postag method is used for annotation. Postag represents the part of speech. For example, if a noun is "n," B represents the first character in the word, and I represents all characters in the word. For example, the word "traffic" is a noun, so the annotation result is ['B-n', 'I-n'].

[0174] Step 302: Divide the text into short sentence sets.

[0175] Optionally, the text is first divided into sentence sets, and then the sentences in the sentence sets are divided into short sentence sets based on the sentence sets. In the embodiment of the present application, a sentence refers to the smallest unit that expresses a complete meaning in the text, and a short sentence is a component of a sentence.

[0176] Alternatively, in a possible implementation, sentences and short sentences can be divided by punctuation marks. For example, the text can be divided into sentence sets based on “\n\r.。:;?”; and the sentences in the sentence set can be further divided into short sentence sets based on “,,,”.

[0177] For example, let's say the text to be tested is "I've stayed here three times in two months. Of course it's good, and the transportation is very convenient." This text contains a period, so it's a sentence. Furthermore, it contains two commas, which divide the sentence into the following three sub-sentences: "I've stayed here three times in two months," "Of course it's good," and "The transportation is very convenient."

[0178] Step 303: Locate the short sentences containing the old sentiment words according to the old sentiment words, and mark the old sentiment words and the short sentences containing the old sentiment words to obtain syntactic structure marking results.

[0179] Among them, the syntactic structure refers to the structure of the short sentence containing the old sentiment word. In the subsequent steps of extracting new sentiment words and the short sentences containing the new sentiment words involved in the embodiment of this application, it is necessary to infer the new sentiment words from the short sentences containing the new sentiment words, and the syntactic structure will provide a basis for this inference. The syntactic annotation results may specifically include: the annotation results for the old sentiment words, and the annotation results for the short sentences containing the old sentiment words.

[0180] In one possible implementation, the "BIO" method is used for labeling. Among them, B represents the first letter of the old sentiment word or the short sentence containing the old sentiment word, I represents the characters of the old sentiment word or the short sentence containing the old sentiment word except the first letter, and O is used to represent the characters excluding the old sentiment word or the short sentence containing the old sentiment word. Specifically, the first letter of the short sentence containing the old sentiment word is marked as "B-SENT", and the remaining characters in the short sentence containing the old sentiment word are marked as "I-SENT". The first letter of the old sentiment word in the short sentence containing the old sentiment word is marked as "B-SW", and the remaining characters in the old sentiment word are marked as "I-SW".

[0181] Step 304: Determine the short sentence containing the new sentiment word and the new sentiment word according to the syntactic analysis result and the syntactic structure annotation result.

[0182] In one possible implementation, a Bi-directional Long Short-Term Memory and Conditional Random Field (BiLSTM-CRF) classifier is used to analyze syntactic features and extract new sentiment words.

[0183] Among them, the BiLSTM-CRF classifier is a method for identifying named entities, and the named entity refers to the type of text sequence that the user wants to extract. In the embodiment of the present application, the named entity refers to the sentiment word and the short sentence containing the sentiment word. For example, the syntactic analysis results and the syntactic structure annotation results can be input into the BiLSTM-CRF classifier to extract and infer the syntactic structure features of the short sentence corresponding to the old sentiment word, and locate the new sentiment word and the short sentence containing the new sentiment word in the text based on the syntactic analysis results and the syntactic structure annotation results.

[0184] Step 304: Based on the syntactic analysis results and the syntactic structure annotation results, determine the short sentence where the new sentiment word is located and the new sentiment word.

[0185] Specifically, step 304 locates the short sentence containing the new emotional word with a similar syntactic structure in the text through the marked old emotional words and the short sentence containing the old emotional words, and then infers the new emotional word in the short sentence containing the new emotional word based on the syntactic structure of the short sentence containing the old emotional word, and extracts the new emotional word.

[0186] It should be understood that the BiLSTM-CRF classifier in the embodiment of the present application is only used as an example, and the present application does not limit the tools used to infer and extract new sentiment words based on the syntactic structure annotation results.

[0187] To improve the accuracy of extracting new sentiment words, we need to avoid the following situation: extracting words from incidental events and treating them as new sentiment words. Therefore, after locating new words in the text based on the annotation results, we can also consider the number of times the extracted new words appear in the corpus before confirming them as new sentiment words.

[0188] Optionally, determining the short sentence containing the new sentiment word and the new sentiment word according to the syntactic analysis result and the syntactic structure annotation result further includes:

[0189] It is determined whether the number of occurrences of the new word in the corpus is greater than a first threshold.

[0190] When the number of occurrences of the new word in the corpus is greater than a first threshold, the new word is confirmed as a new sentiment word.

[0191] The first threshold value may be a specified threshold value or a default threshold value, which may be defined or preset. If no specified threshold value is set, a default threshold value may be used for determination, for example, a default threshold value of 100.

[0192] In one implementation, taking the third sentiment word mentioned above as an example, during the extraction of the third sentiment word, the number of occurrences of the third sentiment word in the corpus may be determined. For example, the number of occurrences of the third sentiment word in the corpus is greater than a first threshold.

[0193] It is worth noting that the new sentiment words extracted in step 304 need to have their corresponding sentiment tendency strength values ​​calculated before being added to the second sentiment dictionary.

[0194] Alternatively, for new emotion words whose emotion tendency has not yet been judged, they can also be expressed as "candidate new emotion words". Introducing candidate new emotion words is to distinguish the different stages in which the method for extracting new emotion words provided by the embodiment of the present application resides. "Candidate new emotion words" can be understood as being only temporarily as candidate new emotion words. As for whether the candidate new emotion words can eventually be included in the second emotion dictionary as new emotion words, it also depends on the processing of subsequent steps. In the subsequent steps, the candidate new emotion words also need to be further processed, and the further processing includes the judgment of emotion tendency and the calculation of emotion tendency intensity value.

[0195] Step 305: Determine the sentiment tendency of the new sentiment word based on the old sentiment word.

[0196] After extracting new sentiment words, we need to calculate their sentiment intensity before adding them to the sentiment dictionary. The calculation of the sentiment intensity requires reference to the sentiment of the words. Therefore, we must first determine the sentiment of the new words. Sentiment intensity refers to whether the words are positive or negative.

[0197] Optionally, factors affecting the emotional tendency of new emotional words include: the number of co-occurrences of old emotional words that co-occur with the new emotional words in the same text, the emotional tendency strength value of the old emotional words that co-occur with the new emotional words, and the number of old emotional words that co-occur with the new emotional words in a set of positive or negative emotional words.

[0198] Optionally, an emotional word graph of the new emotional word and the first emotional word is established through the co-occurrence relationship in the same text, and the positive emotional word graph and the negative emotional word graph are respectively regarded as a subgraph. By cutting and separating the edges of the new emotional word, the separation cost of the new emotional word and the positive emotional word graph, as well as the separation cost SepCost of the new emotional word and the negative emotional subgraph are calculated. The emotional tendency corresponding to the subgraph with the largest separation cost is the emotional tendency of the new emotional word.

[0199] Optionally, all sentiment words included in the sentiment subgraph are old sentiment words. The positive sentiment subgraph refers to the set of positive sentiment words; the negative sentiment subgraph refers to the set of negative sentiment words. Furthermore, when a new sentiment word co-occurs with an old sentiment word, the new sentiment word is connected to the corresponding sentiment word in the subgraph. The degree of the edge connecting the new sentiment word and the old sentiment word is the number of times the new sentiment word and the old sentiment word co-occur, referred to as the co-occurrence count.

[0200] In order to facilitate distinction, the “first separation cost” is used to express the separation cost of the positive emotion subgraph, and the “second separation cost” is used to express the separation cost of the negative emotion subgraph.

[0201] Optionally, the first separation cost, that is, the separation cost SepCost of the positive sentiment subgraph, can be calculated using the following method:

[0202]

[0203] Among them, s i is the emotional tendency strength value of the emotional word in the positive emotional subgraph, and the emotional word and the new emotional word have a co-occurrence relationship in the same text, G is the positive emotional subgraph, d i represents the number of co-occurrences of sentiment words in the positive sentiment subgraph that have a co-occurrence relationship with the new sentiment word (i.e., the degree of the edge of the sentiment word connected to the new sentiment word in the sentiment subgraph), and |G| is the number of sentiment words in the positive sentiment subgraph that have a co-occurrence relationship with the new sentiment word.

[0204] The second separation cost, that is, the separation cost SepCost of the negative sentiment subgraph, can be calculated as follows:

[0205]

[0206] Among them, s iis the emotional tendency strength value of the emotional word in the negative emotional subgraph, and the emotional word and the new emotional word have a co-occurrence relationship in the same text, G is the negative emotional subgraph, d i represents the number of co-occurrences of sentiment words in the negative sentiment subgraph that have a co-occurrence relationship with the new sentiment word, and |G| is the number of sentiment words in the negative sentiment subgraph that have a co-occurrence relationship with the new sentiment word.

[0207] After obtaining the sentiment tendency of the new sentiment word, the sentiment tendency intensity value of the new sentiment word can be further calculated.

[0208] Step 306: Calculate the sentiment tendency strength value of the new sentiment word and add the new sentiment word to the second sentiment dictionary.

[0209] Optionally, the calculation method of the emotional tendency strength value of the new emotional word can adopt the calculation method proposed in step 201. Specifically, if the emotional tendency of the new emotional word is a positive emotional tendency, the emotional tendency strength value of the new emotional word can refer to the calculation method of the positive emotional tendency strength value in the previous text; if the emotional tendency of the new emotional word is a negative emotional tendency, the emotional tendency strength value of the new emotional word can refer to the calculation method of the negative emotional tendency strength value in the previous text. For the sake of brevity, the specific calculation process of the emotional tendency strength value will not be repeated here.

[0210] For the extraction method of this new sentiment word, in order to facilitate the understanding of the processing of a specific text from step 301 to step 303, combined with Figure 4 The examples in are described. Figure 4 This is an example diagram of the annotation results of an embodiment of the present application.

[0211] like Figure 4 As shown, taking the text "I have stayed here three times in two months, of course it is good, the transportation is also very convenient" as an example, combined with Figure 4 Provide an explanation of the specific processing results. Figure 4 The first column is the text; the second column is the syntactic analysis result of the text, that is, the part of speech corresponding to each character marked in step 301; the third column is the syntactic structure marking result, that is, the marking of the old sentiment words and the short sentences where the old sentiment words are located. Specifically, in step 302, the sentence is divided into three short sentences "I have lived here three times in two months", "Of course it is good" and "The transportation is also very convenient" according to the comma. In step 303, two corresponding short sentences are located according to the old sentiment words "good" and "convenient", namely "Of course it is good" and "The transportation is also very convenient" and the two old sentiment words and two short sentences are marked, namely Figure 4 The third column in , thus we get Figure 4 The annotation results in .

[0212] It should be understood that the marking rules of B(I)-postag and BIO adopted in the embodiments of the present application are only used as an example to mark the syntactic analysis results and syntactic structures. The embodiments of the present application do not specifically limit the selection of marking rules.

[0213] As a further improvement to the text sentiment analysis method, an embodiment of the present application further provides a method for analyzing text sentiment polarity by combining a second sentiment dictionary with text sentiment polarity probability.

[0214] Specifically, the embodiment of the present application combines the second sentiment dictionary with a machine learning method for analyzing text sentiment polarity and makes further improvements, wherein the machine learning algorithm is used to obtain the probability of text sentiment polarity.

[0215] It is understood that this application does not limit the choice of machine learning method, and the method can be any method that can derive the probability of text sentiment polarity. The sentiment polarity of a text refers to whether the emotion expressed in the text is positive, negative, or neutral. The sentiment polarity probability of a text includes the probability of positive sentiment polarity, the probability of neutral sentiment polarity, and the probability of negative sentiment polarity.

[0216] Optionally, the machine learning method for analyzing text sentiment polarity may be deep learning. Deep learning refers to a multi-layer neural network and corresponding training methods that takes numerical information as input and generates output results based on the training method.

[0217] In one possible implementation method, the present application uses a long short-term memory neural network (LSTM) to analyze the sentiment polarity of the text to obtain the sentiment polarity probability.

[0218] Specifically, before inputting the text into the neural network, the text data needs to be preprocessed to obtain vectorized data that the neural network can analyze.

[0219] In one possible implementation method, text data preprocessing and sentiment polarity analysis of text using a neural network may specifically include the following steps:

[0220] Step A: First, perform word segmentation and part-of-speech tagging on the text, then remove stop words and keep only words with specified parts of speech to improve the training efficiency of the neural network.

[0221] Optionally, the words of the designated part of speech include one or more of nouns, verbs, adjectives, etc.

[0222] The word segmentation process refers to dividing a complete text into independent words. This application does not limit the tools used for word segmentation.

[0223] Optionally, when removing stop words, a stop word list can be constructed in the second sentiment dictionary provided in the embodiment of the present application, and the removal is performed based on the stop word list. The present application does not limit the selection of stop words.

[0224] Step B: Vectorize the retained words, integrate semantic information and context information, and obtain text feature vectors.

[0225] Semantic information refers to information that helps us understand the meaning of a text. Contextual information refers to information before and after the words that need to be vectorized in the text, and this information may be related to the meaning of the word.

[0226] The embodiment of the present application does not limit the choice of the tool used for vectorizing words. In one possible implementation method, the words in the text are vectorized using a model for generating word vectors.

[0227] Optionally, the model used to generate word vectors can use the word2vec model for word vectorization. Word2vec is a small neural network model that generates word vectors that incorporate contextual and semantic information. When using the word2vec model for word vectorization, the semantic and contextual information refers to the set of words surrounding the target word, where the target word is the word to be vectorized.

[0228] Step C: Input the text feature vector into the neural network for training to obtain the probability f1 of each sentiment polarity of the text:

[0229] f1=(f 11 ,f 12 ,f 13 )

[0230] Among them, f 11 Indicates the probability that the text sentiment polarity is positive, f 12 Indicates the probability that the text sentiment polarity is neutral, f 13 Indicates the probability that the sentiment polarity of the text is negative.

[0231] It should be understood that steps A, B, and C are merely exemplary illustrations of the calculation of the sentiment polarity probability of a text, and the embodiments of the present application are not limited thereto. In fact, those skilled in the art can perform equivalent transformations based on the above steps A-C, and the steps obtained by the transformation still fall within the scope of protection of the embodiments of the present application.

[0232] Figure 5 The flowchart for sentiment polarity analysis of text based on the second sentiment dictionary includes:

[0233] Step 501: Determine the fuzzy prediction degree according to the sentiment polarity probability.

[0234] For example, as described in steps A to C above, the sentiment polarity probability is obtained through the LSTM model f1, and the sentiment polarity probability is expressed as f1=(f 11 ,f 12 ,f 13 ), where the meaning of each parameter can refer to the above description.

[0235] Before analyzing the text to be tested, the embodiment of the present application can calculate the fuzzy prediction degree based on f1. The fuzzy prediction degree is used to represent the degree of dispersion between the probabilities of different sentiment polarities of the text to be tested.

[0236] Optionally, when the emotion polarity probability includes a positive emotion polarity probability, a neutral emotion polarity probability and a negative emotion polarity probability, the fuzzy prediction degree refers to the degree of discreteness between the positive emotion polarity probability, the neutral emotion polarity probability and the negative emotion polarity probability.

[0237] In one possible implementation, the fuzzy predictability is measured by variance, which can be understood as the variance between the probability of positive sentiment polarity, the probability of neutral sentiment polarity, and the probability of negative sentiment polarity.

[0238] For example, the fuzzy prediction degree can be f 11 , f 12 , f 13 The variance of the three values, where variance is a property used to measure the degree of dispersion of a random variable or a set of data.

[0239] Step 502: Determine whether it is necessary to perform a first processing on the sentiment polarity of the text to be tested according to the fuzzy prediction degree.

[0240] Specifically, when determining whether to perform the first processing, when the fuzzy prediction degree meets the preset conditions, the emotional polarity of the text to be tested is determined according to the probabilities of each emotional polarity of the text; when the fuzzy prediction degree does not meet the preset conditions, the text needs to be processed first.

[0241] Optionally, the fuzzy prediction degree meeting the preset condition includes: the fuzzy prediction degree being less than a second threshold. When the fuzzy prediction degree is less than the second threshold, the first processing is performed on the test text, i.e., step 504 is executed below. When the fuzzy prediction degree is greater than or equal to the second threshold, step 503 is executed to determine the sentiment polarity of the test text.

[0242] The second threshold value may be a specified threshold value or a default threshold value, which may be defined or preset. If no specified threshold value is set, a default threshold value may be used for determination, for example, the default threshold value is 0.01.

[0243] Optionally, based on the sentiment polarity probability obtained in step 204, the sentiment polarity of the text can actually be judged. The sentiment polarity analysis of the text based on the second sentiment dictionary proposed in the embodiment of the present application is a further improvement to this method. Therefore, the determination of the sentiment polarity probability can also be considered as the first analysis of the sentiment polarity of the text, and the first processing can be considered as the second analysis of the sentiment polarity of the text.

[0244] Step 503: Determine the sentiment polarity P of the text to be tested according to the sentiment polarity probability.

[0245] For example, the positive emotion polarity probability, the neutral emotion polarity probability and the negative emotion polarity probability are compared, and the emotion polarity probability with the largest probability value is selected among these emotion polarity probabilities. Then, the emotion polarity corresponding to the emotion polarity probability with the largest probability value is used as the emotion polarity of the text to be tested.

[0246] Step 504: Calculate the second sentiment value of the text to be tested.

[0247] For ease of description, the second sentiment value may be expressed as "text sentiment value" in the following text. For example, the text sentiment value is expressed as f2.

[0248] The magnitude of the sentiment value of the text to be tested is used to represent the sentiment polarity tendency of the text, and this application does not limit the numerical range of the text sentiment value.

[0249] In one possible implementation, when the sentiment value of the text is positive, the sentiment polarity of the text tends to be positive; when the sentiment value of the text is negative, the sentiment polarity of the text tends to be negative; when the sentiment value of the text is close to 0, the sentiment polarity of the text tends to be neutral.

[0250] Optionally, factors affecting the calculation of text sentiment value include: the weight of sentences in the text, the number of sentiment words included in the sentences, the sentiment tendency strength value of the sentiment words, the number of occurrences of negative words in the text, and the degree value of degree adverbs in the text.

[0251] According to the degree expressed by the degree adverb, the degree adverb corresponds to different degree values. The embodiment of the present application does not limit the source of the degree value. Optionally, the sentiment value in the network resource is obtained and assigned.

[0252] In a possible implementation method, the six adverb word lists in the HowNet sentiment dictionary are assigned values: degree words such as "one hundred percent" and "double" are assigned a degree value of 1.8, degree words such as "excessive" and "too" are assigned a degree value of 1.6, degree words such as "very" and "extraordinarily" are assigned a degree value of 1.5, degree words such as "relatively" and "more" are assigned a degree value of 0.8, degree words such as "slightly" and "slightly" are assigned a degree value of 0.7, and degree words such as "some" and "relatively" are assigned a degree value of 0.5.

[0253] Optionally, the degree value can be obtained in the sentiment dictionary by constructing a degree adverb vocabulary in the sentiment dictionary provided in the embodiment of the present application and assigning values ​​to the degree adverbs in the vocabulary.

[0254] Optionally, the first sentiment value of the sentence in the text is calculated based on the sentiment dictionary, and then the sentiment value of the text is calculated by combining the sentiment value of the sentence and the weight of the sentence in the text.

[0255] For the sake of convenience, the following text may use "sentiment value of the sentence" to express the first sentiment value.

[0256] In one possible implementation, the weight of the sentence is obtained through the Python third-party library textrank4zh.

[0257] In one possible implementation method, the sentiment value of a text is calculated using the following method:

[0258]

[0259] Among them, f2 is the sentiment value of the text, n is the number of sentences in the text, α i is the weight of the i-th sentence in the text, n i is the number of sentiment words in the i-th sentence in the text, where the sentiment word is the second sentiment word or the third sentiment word, n ij is the number of times the negative word appears in the short sentence of the jth sentiment word in the i-th sentence in the text, c ij is the maximum degree value of the degree adverbs in the short sentence of the jth sentiment word in the i-th sentence in the text, s ij is the sentiment tendency strength value of the j-th sentiment word in the i-th sentence in the text.

[0260] It should be understood that the above formula calculates the sentiment value of the text by calculating the sentiment value of the sentence in the text. Specifically, the sentiment value of the sentence is calculated as follows: The text sentiment value is the weighted sum of the sentiment values ​​of the sentences.

[0261] Step 505: Determine the sentiment polarity P of the text to be tested according to the second sentiment value of the text to be tested.

[0262] Specifically, when the second emotion value is in the first threshold interval, the emotion polarity of the text to be tested is positive emotion polarity; when the second emotion value is in the second threshold interval, the emotion polarity of the text to be tested is neutral emotion polarity; when the second emotion value is in the third threshold interval, the emotion polarity of the text to be tested is negative emotion polarity.

[0263] Optionally, a threshold interval within which the text sentiment value falls is determined by setting a third threshold and a fourth threshold. In one possible implementation, when the text sentiment value is greater than or equal to the third threshold, the sentiment polarity of the text is positive; when the text sentiment value is greater than the fourth threshold and less than the third threshold, the sentiment polarity of the text is neutral; and when the text sentiment value is less than or equal to the fourth threshold, the sentiment polarity of the text is negative.

[0264] For example, the third threshold is set to 0.15 and the fourth threshold is set to -0.15. Then, by dividing the intervals by 0.15 and -0.15, three threshold intervals can be obtained. The first threshold interval is when the sentiment value is less than or equal to 0.15, the second threshold interval is when the sentiment value is greater than -0.15 and less than 0.15, and the third threshold interval is when the sentiment value is less than or equal to -0.15. Accordingly, the sentiment polarity P in step 503 can be determined according to the following formula:

[0265]

[0266] Under this setting condition, when the text sentiment value f2 is greater than or equal to 0.15, the sentiment polarity P of the text is positive sentiment polarity; when the text sentiment value f2 is greater than -0.15 and less than 0.15, the sentiment polarity P of the text is neutral sentiment polarity; when the text sentiment value f2 is less than or equal to -0.15, the sentiment polarity P of the text is negative sentiment polarity.

[0267] It should be understood that the above-mentioned threshold intervals divided by the third threshold and the fourth threshold, and the emotional polarity corresponding to each threshold interval are merely exemplary descriptions, and the embodiments of the present application are not limited thereto.

[0268] This application provides a comprehensive embodiment for the method of calculating the sentiment tendency intensity value, the method of extracting new sentiment words, and the method of analyzing the sentiment polarity of text in combination with the sentiment dictionary proposed in this application. It should be understood that this embodiment is only used as an example, and the embodiment corresponding to this application can be a combination of one or more of the above three methods.

[0269] Figure 6 The flowchart of the text sentiment polarity analysis provided in the embodiment of this application is as follows: Figure 6 This embodiment is described by way of example only. It should be understood that Figure 6 This is just a specific example and does not completely correspond to the possible implementation methods of the embodiments of this application. Figure 6 For the terms and explanations involved, as well as the specific method flow, please refer to the previous description.

[0270] Figure 6 610 includes the construction of the sentiment dictionary and the LSTM model of the embodiment of the present application. The construction of the sentiment dictionary includes: calculation of sentiment tendency intensity values ​​and extraction of new sentiment words.

[0271] For example, Figure 6 The sentiment dictionary shown in can be the second sentiment dictionary generated by the embodiment of the present application, including: a positive sentiment word list, a negative sentiment word list, a stop word word list, a degree adverb word list and a negation word list.

[0272] Furthermore, for the sentiment words in the positive sentiment word list and the negative sentiment word list, the sentiment tendency intensity value calculation method proposed in the embodiment of the present application is used to calculate the corresponding sentiment tendency intensity value. Specifically, the sentiment tendency intensity value is calculated by the number of occurrences of the positive sentiment words and the negative sentiment words in the corpus, wherein the corpus includes positive sentiment corpus, neutral sentiment corpus and negative sentiment corpus. Optionally, the calculation of the sentiment tendency intensity value also takes into account whether the sentiment words are used to express positive or negative meanings, and the negative words in the negative word list can be used to judge whether the sentiment words express positive or negative meanings.

[0273] Optionally, the text in the corpus provided in the embodiment of the present application can also be used as the test text to input into LSTM to perform text sentiment polarity analysis. For example, Figure 6 The texts in the neutral sentiment corpus, positive sentiment corpus, and negative sentiment corpus shown in the figure can be input into LSTM to perform text sentiment polarity analysis and obtain sentiment polarity probability.

[0274] In the embodiment of the present application, new sentiment words are extracted from the corpus, and the specific steps include:

[0275] Step 611: Marking the syntactic structure of the text in the corpus;

[0276] Step 612: Input the syntactic structure labeling result into the BiLSTM-CRF classifier;

[0277] Step 613: Obtain candidate new sentiment words through the BiLSTM-CRF classifier;

[0278] Step 614: performing sentiment tendency analysis on the candidate new sentiment words based on the old sentiment words;

[0279] Step 615: Calculate the sentiment intensity value of the candidate new sentiment word according to the sentiment orientation;

[0280] Step 616: The candidate new sentiment word is used as a new sentiment word, and the new sentiment word is added to the sentiment dictionary.

[0281] Figure 6 The portion other than 610 represents inputting a text to be tested and performing sentiment polarity analysis on the text to be tested based on the sentiment polarity probability and the second sentiment dictionary. The steps of performing sentiment polarity analysis on the text specifically include:

[0282] Step 621: Input the text to be tested into the LSTM model;

[0283] Step 622: Obtain the probability of each sentiment polarity of the text through the LSTM model;

[0284] Step 623: Calculate the fuzzy prediction degree based on the probability of each sentiment polarity of the text;

[0285] Step 624: Determine whether to perform a secondary analysis on the text to be tested based on whether the fuzzy prediction degree meets the preset conditions;

[0286] Step 625: When secondary analysis is required, the sentiment value of the text is calculated based on the sentiment dictionary;

[0287] Step 626: determining the sentiment polarity of the text according to the threshold range of the text sentiment value;

[0288] Step 627: deriving the sentiment polarity of the text;

[0289] Step 628: When secondary analysis is not required, the sentiment polarity of the text is obtained according to the probability of each sentiment polarity.

[0290] It is worth noting that the text sentiment analysis method provided in the embodiment of the present application is a word-based analysis method. For example, Figure 6 The text sentiment value is calculated based on the sentiment words in the second sentiment dictionary generated by the embodiment of the present application. Therefore, when executing step 625, the embodiment of the present application calculates the text sentiment value based on the sentiment words in the sentiment dictionary and the sentiment tendency intensity values ​​corresponding to the sentiment words. Optionally, the calculation of the text sentiment value may also consider degree adverbs and their degree values, which may be derived from a degree adverb vocabulary.

[0291] It should be understood that Figure 6 This is just an example of the method for analyzing text sentiment polarity, and the embodiments of the present application are not limited thereto. In fact, those skilled in the art can perform equivalent transformations based on the above steps, and the steps obtained by the transformation still fall within the scope of protection of the embodiments of the present application.

[0292] Figure 7FIG1 shows a schematic block diagram of an apparatus 700 for text sentiment analysis according to an embodiment of the present application. Optionally, the specific form of the apparatus 700 can be a general-purpose computer device or a chip in a general-purpose computer device, which is not limited in the embodiment of the present application. Figure 7 As shown, the apparatus 700 includes: a determination module 710 , an extraction module 720 , a generation module 730 , and an analysis module 740 .

[0293] In a possible implementation, the determination module 710 is configured to calculate the sentiment tendency strength value of the sentiment word.

[0294] The extraction module 720 is used to extract new sentiment words from the corpus.

[0295] The generating module 730 is configured to generate a second sentiment dictionary.

[0296] The analysis module 740 is used to analyze the sentiment intensity value of the text to be tested.

[0297] As a possible embodiment, the determination module 710 is configured to determine the sentiment tendency strength value of the first sentiment word according to the number of times the first sentiment word appears in the corpus.

[0298] In one possible implementation, the positive sentiment word t i The emotional tendency strength value The calculation is done as follows:

[0299]

[0300] Among them, t i is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of positive semantics in positive sentiment corpus, pwords is the set of all positive sentiment words, is the sum of the occurrence times of all sentiment words expressing positive semantics in the positive sentiment corpus, t i The number of occurrences of positive semantics in negative sentiment corpus, t i The number of occurrences in negative sentiment corpus.

[0301] Negative sentiment words i Sentiment intensity value and is calculated as follows:

[0302]

[0303] Among them, ti is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of negative semantics in negative sentiment corpus, nwords is the set of all negative sentiment words, is the sum of the occurrence times of all sentiment words expressing negative semantics in the negative sentiment corpus, t i The number of occurrences of negative semantics in positive sentiment corpus, t i The number of occurrences in positive sentiment corpus.

[0304] As a possible embodiment, the extraction module 720 is used to extract new sentiment words from the corpus based on the old sentiment words, specifically including:

[0305] Performing syntactic analysis on the text in the corpus to obtain syntactic analysis results;

[0306] dividing the text into a set of short sentences;

[0307] Locate the short sentences where the old emotional words are located according to the old emotional words, and annotate the old emotional words and the short sentences where the old emotional words are located to obtain the syntactic structure annotation results;

[0308] According to the syntactic analysis results and syntactic structure annotation results, determine the short sentences where the new sentiment words are located and the new sentiment words;

[0309] Determine the emotional tendency of new emotional words based on old emotional words;

[0310] Calculate the sentiment intensity value of the new sentiment word.

[0311] As a possible embodiment, the generating module 730 is configured to generate a second sentiment dictionary.

[0312] In a possible implementation, the second emotion dictionary includes the second emotion words in the first emotion dictionary and the third emotion words extracted by the extraction module 720. In addition, the emotion tendency strength values ​​of the emotion words in the second emotion dictionary are determined by the determination module 710.

[0313] As a possible embodiment, the analysis module 740 is used to analyze the sentiment value of the text to be tested, including:

[0314] First, the text is segmented and POS tagged, then stop words are removed and only words with specified POS are retained;

[0315] Perform vectorization on the retained words, integrate semantic information and context information, and obtain text feature vectors;

[0316] Input the text feature vector into the neural network for training to obtain the probability of each sentiment polarity of the text;

[0317] Analyze the sentiment polarity of text.

[0318] Optionally, as a possible embodiment, the analysis module 740 is further configured to analyze the sentiment polarity of the text by combining the sentiment dictionary and the text sentiment polarity probability, specifically including:

[0319] Determine the fuzzy prediction degree according to the sentiment polarity probability;

[0320] Determining whether it is necessary to perform a first processing on the sentiment polarity of the text to be tested according to the fuzzy prediction degree;

[0321] When the text to be tested needs to be first processed, the sentiment value of the text to be tested is calculated, and the sentiment polarity of the text to be tested is determined according to the sentiment value of the text to be tested;

[0322] When the first processing is not required for the text to be tested, the sentiment polarity of the text to be tested is determined according to the probabilities of the sentiment polarities of the text to be tested.

[0323] Figure 8 This is a schematic diagram of the structure of the apparatus 800 for text sentiment analysis provided in one embodiment of the present application. Figure 8 As shown, the device 800 includes: at least one processor 80 ( Figure 8 Only one is shown in the figure) a processor, a memory 81, and a computer program 82 stored in the memory 81 and executable on the at least one processor 80, wherein the processor 80 executes the computer program 82 to implement any of the above-mentioned method embodiments for text sentiment analysis (e.g. Figure 2 The method in Figure 3 The method in Figure 5 Methods in or Figure 6 ) in the method.

[0324] The apparatus 800 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The apparatus for text sentiment analysis may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will appreciate that Figure 8 This is merely an example of the apparatus 800 for text sentiment analysis and does not constitute a limitation on the apparatus 800 . The apparatus 800 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the apparatus 800 may also include input and output devices, network access devices, etc.

[0325] The processor 80 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0326] In some embodiments, the memory 81 may be an internal storage unit of the device 800, such as a hard disk or memory of the device 800 for text sentiment analysis. In other embodiments, the memory 81 may also be an external storage device of the device 800, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the device 800. Furthermore, the memory 81 may also include both an internal storage unit of the device 800 and an external storage device. The memory 81 is used to store an operating system, an application, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 81 may also be used to temporarily store data that has been output or is to be output.

[0327] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0328] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0329] An embodiment of the present application also provides a device for text sentiment analysis, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the computer program.

[0330] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0331] An embodiment of the present application provides a computer program product. When the computer program product is run on a device for text sentiment analysis, the device for text sentiment analysis can implement the steps in the above-mentioned various method embodiments when executed.

[0332] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process of the above-mentioned method embodiment by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, mobile hard drive, magnetic disk, or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.

[0333] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0334] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0335] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0336] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0337] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for text sentiment analysis, characterized in that: include: Determining the sentiment tendency intensity value of the first sentiment word according to the number of times the first sentiment word appears in the corpus; Extracting a third sentiment word from the corpus based on the second sentiment word, where the second sentiment word refers to a word in the first sentiment dictionary, and the sentence containing the third sentiment word has a similar syntactic structure to the sentence containing the second sentiment word; generating a second emotional dictionary according to the emotional tendency strength value and the third emotional word, wherein the second emotional dictionary includes the first emotional word, the second emotional word, and the third emotional word; Analyzing the sentiment polarity of the text to be tested according to the second sentiment dictionary; Wherein, the first emotional words include positive emotional words and negative emotional words, and the corpus includes positive emotional corpus and negative emotional corpus; The step of determining the sentiment tendency strength value of the first sentiment word according to the number of times the first sentiment word appears in the corpus includes: Determine the sentiment tendency intensity value of the positive sentiment word according to one or more of the number of occurrences of the positive sentiment word when expressing positive semantics in the positive sentiment corpus, the number of occurrences of the positive sentiment word when expressing positive semantics in the negative sentiment corpus, the number of occurrences of the positive sentiment word in the negative sentiment corpus, and the sum of the number of occurrences of all sentiment words expressing positive semantics in the positive sentiment corpus; The emotional tendency intensity value of the negative emotional word is determined based on one or more of the number of times the negative emotional word appears when expressing negative semantics in the negative emotional corpus, the number of times the negative emotional word appears when expressing negative semantics in the positive emotional corpus, the number of times the negative emotional word appears in the positive emotional corpus, and the sum of the number of times all emotional words expressing negative semantics appear in the negative emotional corpus.

2. The method according to claim 1, characterized in that Before analyzing the sentiment polarity of the test text according to the second sentiment dictionary, the method further includes: Determining the sentiment polarity probability of the text to be tested; The step of analyzing the sentiment polarity of the text to be tested according to the second sentiment dictionary includes: The sentiment polarity of the text to be tested is analyzed according to the sentiment polarity probability and the second sentiment dictionary.

3. The method according to claim 1, characterized in that The emotional tendency strength value of the positive emotional word satisfies the following formula: Among them, t i is the positive sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of positive semantics in the positive sentiment corpus, pwords is the set of all positive sentiment words, is the sum of the number of occurrences of all positive sentiment words expressing positive semantics in the positive sentiment corpus, t i The number of occurrences of positive semantics in the negative sentiment corpus, t i The number of occurrences in the negative sentiment corpus; The sentiment tendency strength value of the negative sentiment word satisfies the following formula: Among them, t i is the negative sentiment word, t i The emotional tendency intensity value, t i The number of occurrences of negative semantics in the negative sentiment corpus, nwords is the set of all negative sentiment words, is the sum of the occurrence times of all negative sentiment words expressing negative semantics in the negative sentiment corpus, t i The number of occurrences of negative semantics in the positive sentiment corpus, t i The number of occurrences in the positive sentiment corpus.

4. The method according to any one of claims 1 to 3, characterized in that The extracting the third sentiment word from the corpus according to the second sentiment word comprises: Performing syntactic analysis on the text in the corpus to obtain syntactic analysis results; dividing the text into a set of short sentences; Determine, according to the second sentiment word, a first short sentence in which the second sentiment word is located, wherein the first short sentence is a short sentence in the short sentence set; Annotating the second sentiment word and the first short sentence to obtain a syntactic structure annotation result; Determining a second short sentence and the third sentiment word according to the syntactic analysis result and the syntactic structure annotation result, wherein the second short sentence is a short sentence in which the third sentiment word is located in the short sentence set, and the second short sentence has a similar syntactic structure to the first short sentence, wherein the number of occurrences of the third sentiment word in the corpus is greater than a first threshold; Determining the emotional tendency of the third emotional word according to the second emotional word, wherein the emotional tendency includes a positive emotional tendency and a negative emotional tendency; According to the emotional tendency of the third emotional word, the emotional tendency intensity value of the third emotional word is determined.

5. The method according to claim 4, characterized in that Determining the emotional tendency of the third emotional word according to the second emotional word includes: Determining a sentiment word graph according to a co-occurrence relationship between the second sentiment word and the third sentiment word in the same text, wherein the sentiment word graph includes a positive sentiment subgraph and a negative sentiment subgraph; Determine a first separation cost and a second separation cost, wherein the first separation cost refers to the separation cost of the third emotion word and the positive emotion subgraph, and the second separation cost refers to the separation cost of the third emotion word and the negative emotion subgraph, and the first separation cost or the second separation cost satisfies the following formula: Among them, when SepCost is the first separation cost, s i is the emotional tendency strength value of the second emotional word in the positive emotional subgraph, G is the positive emotional subgraph, d i represents the co-occurrence times of the second sentiment word and the third sentiment word in the positive sentiment subgraph, |G| is the number of the second sentiment words in the positive sentiment subgraph; when SepCost is the second separation cost, s i is the emotional tendency strength value of the second emotional word in the negative emotional subgraph, G is the negative emotional subgraph, d i represents the number of co-occurrences of the second sentiment word and the third sentiment word in the negative sentiment subgraph, and |G| is the number of the second sentiment words in the negative sentiment subgraph; comparing the first separation cost and the second separation cost; The emotional tendency corresponding to the emotional subgraph having the largest separation cost between the first separation cost and the second separation cost is determined as the emotional tendency of the third emotional word.

6. The method according to any one of claims 1 to 3, characterized in that The emotion polarity probability includes a positive emotion polarity probability, a neutral emotion polarity probability and a negative emotion polarity probability; Before analyzing the sentiment polarity of the text to be tested according to the sentiment polarity probability and the second sentiment dictionary, the method further includes: Determining a fuzzy prediction degree according to the sentiment polarity probability, wherein the fuzzy prediction degree is used to represent the degree of dispersion between different sentiment polarity probabilities of the text to be tested; Determining whether it is necessary to perform a first processing on the sentiment polarity of the text to be tested according to the fuzzy prediction degree; The step of analyzing the sentiment polarity of the text to be tested according to the sentiment polarity probability and the second sentiment dictionary includes: When the fuzzy prediction degree meets a preset condition, performing a first processing on the sentiment polarity of the text to be tested, wherein the first processing includes determining the sentiment polarity of the text to be tested according to the second sentiment dictionary; When the fuzzy prediction degree does not meet the preset conditions, the sentiment polarity corresponding to the largest sentiment polarity probability among the positive sentiment polarity probability, the neutral sentiment polarity probability and the negative sentiment polarity probability is determined as the sentiment polarity of the text to be tested.

7. The method according to claim 6, characterized in that Determining the emotional tendency of the text to be tested according to the second emotional dictionary includes: Determine a first sentiment value of a sentence in the text to be tested according to the second sentiment dictionary; Determining a second sentiment value of the text to be tested based on the first sentiment value and a weight, wherein the weight refers to the weight of the sentence corresponding to the first sentiment value in the text to be tested; Determine the sentiment polarity of the text to be tested according to the threshold range in which the second sentiment value is located, wherein: When the second sentiment value is within the first threshold range, the sentiment polarity of the text to be tested is positive; When the second sentiment value is within a second threshold range, the sentiment polarity of the text to be tested is neutral sentiment polarity; When the second sentiment value is within the third threshold range, the sentiment polarity of the text to be tested is negative sentiment polarity.

8. The method according to claim 7, characterized in that The second emotion value satisfies the following formula: Wherein, f2 is the second sentiment value, n is the number of sentences in the text to be tested, α i is the weight of the i-th sentence in the test text, n i is the number of sentiment words in the i-th sentence in the text to be tested, n ij is the number of times negative words appear in the short sentence of the jth sentiment word in the i-th sentence in the test text, c ij is the maximum degree value of the degree adverbs in the short sentence containing the jth sentiment word in the i-th sentence of the test text, s ij is the emotional tendency strength value of the jth emotional word in the i-th sentence in the text to be tested.

9. A text sentiment analysis device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. A chip, characterized in that: The method comprises a processor, wherein when the processor executes instructions, the processor performs the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sentiment value based web text sentiment analysis method

    CN104008091A

  • Sentiment analysis method and device of text information

    CN106547924A

  • A method and device for analyzing emotional polarity of network public opinion

    CN109446404A

  • Online shopping comment new sentiment word extraction method based on syntactic analysis

    CN112926318A