A text sentiment analysis method and device

By decomposing the text into subtext identifiers and processing using text sentiment analysis model, the problem that text sentiment analysis in the prior art cannot accurately analyze the overall semantics of text in the text is solved, and the accuracy and efficiency of sentiment analysis are improved.

CN115081456BActive Publication Date: 2025-05-30GUANGZHOU YOUMI INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210645924.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-05-30
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

Existing text sentiment analysis methods cannot accurately analyze the overall semantics of the text, resulting in low accuracy of sentiment analysis.

Method used

By decomposing the text to be analyzed into subtext identifiers, and using a pre-trained text sentiment analysis model, processing based on the text vector and position vector of each subtext identifier, multiple sentiment analysis results and probability information are output, and the target sentiment analysis results that meet the pre-set filter conditions are finally determined.

Benefits of technology

It improves the accuracy and efficiency of text sentiment analysis, can analyze the overall semantics of the text more accurately, and reduces the problems of single emotion analysis methods and unreasonable focus due to simple splicing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081456B_ABST
    Figure CN115081456B_ABST
Patent Text Reader

Abstract

The present invention discloses a text sentiment analysis method and device, including: determining text information corresponding to the text to be analyzed; inputting the text information into a text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and probability information corresponding to each sentiment analysis result; determining, from all the sentiment analysis results, a target sentiment analysis result whose corresponding probability information meets a preset screening condition as the sentiment analysis result corresponding to the text to be analyzed; wherein, the text sentiment analysis model determines all the sentiment analysis results based on the text vector corresponding to each sub-text identifier and the position vector corresponding to each sub-text identifier. It can be seen that implementing the present invention can use the text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed according to the text vector and position vector of the text to be analyzed, which is beneficial to accurately analyzing the overall semantics of the text, thereby improving the accuracy of text sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic analysis, and in particular, to a method and device for text sentiment analysis. Background Art

[0002] In real life, there are a large number of comments on the Internet participated by users, such as comments on people and events. In order to obtain the public opinion on a certain thing, text sentiment analysis processing can be performed on a large number of comments on the thing on the Internet. The current text sentiment analysis methods mainly use convolutional neural networks or recurrent neural networks to disassemble the text and extract the keywords therein according to a series of pre-established sentiment dictionaries and rules, calculate the sentiment value corresponding to the text according to the keywords, and finally confirm the sentiment tendency of the text according to the sentiment words. However, it is found in practice that the existing text sentiment analysis methods only regard the text to be analyzed as a set of words, resulting in the inability to accurately analyze the overall semantics of the text, and further resulting in low accuracy of text sentiment analysis. It can be seen that how to improve the accuracy of text sentiment analysis is particularly important. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for text sentiment analysis, which can be conducive to accurately analyzing the overall semantics of the text, thereby improving the accuracy of text sentiment analysis.

[0004] To solve the above technical problem, a first aspect of the present invention discloses a method for text sentiment analysis, the method comprising:

[0005] Determine the text information corresponding to the text to be analyzed, the text information including at least one sub-text identifier and the text content corresponding to each sub-text identifier;

[0006] Input the text information into a pre-trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output a plurality of sentiment analysis results and the probability information corresponding to each sentiment analysis result;

[0007] According to the probability information corresponding to each sentiment analysis result, determine, from all the sentiment analysis results, a target sentiment analysis result whose corresponding probability information meets a pre-set screening condition as the sentiment analysis result corresponding to the text to be analyzed;

[0008] wherein, the text sentiment analysis model determines all the sentiment analysis results based on the text vector corresponding to each sub-text identifier and the position vector corresponding to each sub-text identifier.

[0009] As an optional implementation manner, in the first aspect of the present invention, the text sentiment analysis model processes the text information, including:

[0010] Encoding operations are performed on the text information by the encoding structure corresponding to the text sentiment analysis model to obtain the encoding results corresponding to each of the sub - text identifiers; the encoding result corresponding to each sub - text identifier includes the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier;

[0011] Vector conversion processing is performed on the encoding result corresponding to each sub - text identifier by the vector processing structure of the text sentiment analysis model to obtain the target matrix corresponding to each sub - text identifier;

[0012] Average pooling operations are performed on the target matrix corresponding to the sub - text identifier by the average pooling structure of the text sentiment analysis model to obtain the average pooling result corresponding to each sub - text identifier;

[0013] Concatenation operations are performed on the average pooling results corresponding to all the sub - text identifiers by the concatenation structure of the text sentiment analysis model to obtain a concatenation result;

[0014] The concatenation result is processed by the fully - connected layer of the text sentiment analysis model.

[0015] As an optional implementation manner, in the first aspect of the present invention, the vector processing structure of the text sentiment analysis model performs vector conversion processing on the encoding result corresponding to each sub - text identifier to obtain the target matrix corresponding to each sub - text identifier, including:

[0016] When the text sentiment analysis model includes only one vector processing structure, the vector processing structure performs vector conversion processing on the encoding result corresponding to each text identifier to obtain the target matrix corresponding to each sub - text identifier;

[0017] When the text sentiment analysis model includes multiple text processing structures, for each sub - text identifier, the vector processing structure matching the sub - text identifier performs vector conversion processing on the encoding result corresponding to the sub - text identifier to obtain the target matrix corresponding to the sub - text identifier.

[0018] As an optional implementation manner, in the first aspect of the present invention, the determination of the text information corresponding to the text to be analyzed includes:

[0019] According to at least one pre - set sub - text identifier, text elements corresponding to each sub - text identifier in the text to be analyzed are extracted as the text content corresponding to the sub - text identifier;

[0020] And, before inputting the text information into a pre-trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result, the method further includes:

[0021] Performing a preprocessing operation on the text information;

[0022] Wherein, performing the preprocessing operation on the text information includes:

[0023] For the text content corresponding to each sub-text identifier, detecting whether there is a text element to be removed that meets a preset clearing condition in the text content, and when the detection result is yes, removing the text element to be removed in the text content.

[0024] As an optional implementation manner, in the first aspect of the present invention, detecting whether there is a text element to be removed that meets a preset clearing condition in the text content corresponding to each sub-text identifier includes:

[0025] For the text content corresponding to each sub-text identifier, detecting whether there is a first text element whose element type matches a preset first element type in the text content, and when it is detected that there is the first text element in the text content, determining that there is a text element to be removed that meets a preset clearing condition in the text content, wherein the text element to be removed includes the first text element; and / or,

[0026] For the text content corresponding to each sub-text identifier, detecting whether there are adjacent second text elements with the same element type and the element type matching a preset second element type in the text content, and when it is detected that there are the second text elements in the text content, determining that there is a text element to be removed that meets a preset clearing condition in the text content, wherein the text element to be removed includes some of the second text elements;

[0027] And, before removing the text element to be removed in the text content, the method further includes:

[0028] Judging whether the influence degree of the text element to be removed in the text content on the sentiment of the text content is greater than a preset degree, and when the judgment result is no, triggering the operation of removing the text element to be removed in the text content.

[0029] As an optional implementation manner, in the first aspect of the present invention, before determining the text information corresponding to the text to be analyzed, the method further includes:

[0030] When the original text for which the sentiment is to be analyzed includes at least two text structures, determine the sentiment information for each of the text structures, where the sentiment information for each of the text structures includes the sentiment subject corresponding to the text structure and / or the sentiment object corresponding to the text structure;

[0031] Determine at least one text to be analyzed from the original text according to the sentiment information for each of the text structures;

[0032] Among them, the determining at least one text to be analyzed from the original text according to the sentiment information for each of the text structures includes:

[0033] When the sentiment information for all of the text structures matches, determine the original text as the text to be analyzed;

[0034] When the sentiment information for all of the text structures does not match, classify all of the text structures according to the sentiment information for each of the text structures to obtain at least two texts to be analyzed, where if a text to be analyzed includes at least two of the text structures, the sentiment information for all of the text structures in the text to be analyzed matches.

[0035] As an optional implementation manner, in the first aspect of the present invention, the determining the sentiment information for each of the text structures includes:

[0036] Determine at least one sentiment entity corresponding to each of the text structures, and determine the confirmation keywords in each of the text structures that meet the preset confirmation conditions;

[0037] For each of the text structures, according to the position order of each of the sentiment entities in the text structure and the position order of the confirmation keywords in the text structure, determine the entity attribute of each of the sentiment entities in the text structure as the sentiment information for the text structure;

[0038] Among them, the determining the confirmation keywords in each of the text structures that meet the preset confirmation conditions includes:

[0039] For each of the text structures, determine whether there is a first keyword for introducing a sentiment object in the text structure. When the determination result is yes, determine the first keyword as the confirmation keyword in the text structure that meets the preset confirmation conditions. When the determination result is no, determine the second keyword for describing the sentiment content in the text structure as the confirmation keyword in the text structure that meets the preset confirmation conditions.

[0040] A second aspect of the present invention discloses a text sentiment analysis device, where the device includes:

[0041] A determination module, configured to determine text information corresponding to the text to be analyzed, where the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier;

[0042] An emotion analysis module, configured to input the text information into a pre - trained text emotion analysis model, so as to trigger the text emotion analysis model to process the text information and output multiple emotion analysis results and the probability information corresponding to each emotion analysis result;

[0043] The determination module is further configured to determine, according to the probability information corresponding to each emotion analysis result, a target emotion analysis result whose corresponding probability information meets a preset screening condition from all the emotion analysis results as the emotion analysis result corresponding to the text to be analyzed;

[0044] Wherein, the text emotion analysis model determines all the emotion analysis results based on the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier.

[0045] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the text emotion analysis model processes the text information includes:

[0046] The encoding structure corresponding to the text emotion analysis model performs an encoding operation on the text information to obtain an encoding result corresponding to each sub - text identifier; the encoding result corresponding to each sub - text identifier includes the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier;

[0047] The vector processing structure of the text emotion analysis model performs a vector conversion process on the encoding result corresponding to each sub - text identifier to obtain a target matrix corresponding to each sub - text identifier;

[0048] The average pooling structure of the text emotion analysis model performs an average pooling operation on the target matrix corresponding to the sub - text identifier to obtain an average pooling result corresponding to each sub - text identifier;

[0049] The splicing structure of the text emotion analysis model performs a splicing operation on the average pooling results corresponding to all the sub - text identifiers to obtain a splicing result;

[0050] The fully - connected layer of the text emotion analysis model processes the splicing result.

[0051] As an alternative implementation, in the second aspect of the present invention, the specific manner in which the vector processing structure of the text sentiment analysis model performs vector conversion processing on the encoding result corresponding to each sub-text identifier to obtain the target matrix corresponding to each sub-text identifier includes:

[0052] When the text sentiment analysis model includes only one vector processing structure, the vector processing structure performs vector conversion processing on the encoding result corresponding to each text identifier to obtain the target matrix corresponding to each sub-text identifier;

[0053] When the text sentiment analysis model includes multiple text processing structures, for each sub-text identifier, the vector processing structure matching the sub-text identifier performs vector conversion processing on the encoding result corresponding to the sub-text identifier to obtain the target matrix corresponding to the sub-text identifier.

[0054] As an alternative implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the text information corresponding to the text to be analyzed includes:

[0055] According to at least one pre-set sub-text identifier, text elements corresponding to each sub-text identifier in the text to be analyzed are extracted as the text content corresponding to the sub-text identifier;

[0056] In addition, the device further includes:

[0057] A preprocessing module, configured to perform a preprocessing operation on the text information before the sentiment analysis module inputs the text information into a pre-trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result;

[0058] Wherein, the specific manner in which the preprocessing module performs a preprocessing operation on the text information includes:

[0059] For the text content corresponding to each sub-text identifier, it is detected whether there are text elements to be removed in the text content that meet the preset clearing conditions, and when the detection result is yes, the text elements to be removed in the text content are removed.

[0060] As an alternative implementation, in the second aspect of the present invention, the specific manner in which the preprocessing module detects whether there are text elements to be removed in the text content corresponding to each sub-text identifier that meet the preset clearing conditions includes:

[0061] For each piece of text content corresponding to the sub - text identifier, detect whether there is a first text element in the text content whose element type matches a preset first element type. When it is detected that there is the first text element in the text content, determine that there is a text element to be removed in the text content that meets the preset removal condition, where the text element to be removed includes the first text element; and / or,

[0062] For each piece of text content corresponding to the sub - text identifier, detect whether there are adjacent second text elements with the same element type and the element type matching a preset second element type in the text content. When it is detected that there are the second text elements in the text content, determine that there is a text element to be removed in the text content that meets the preset removal condition, where the text element to be removed includes some of the second text elements;

[0063] Moreover, the pre - processing module is further configured to, before performing the operation of removing the text element to be removed in the text content, determine whether the influence degree of the text element to be removed in the text content on the sentiment of the text content is greater than a preset degree. When the judgment result is negative, trigger the execution of the operation of removing the text element to be removed in the text content.

[0064] As an optional implementation manner, in the second aspect of the present invention, the determination module is further configured to, before performing the operation of determining the text information corresponding to the text to be analyzed, when the original text of the sentiment to be analyzed includes at least two text structures, determine the sentiment information of each text structure, where the sentiment information of each text structure includes the sentiment subject corresponding to the text structure and / or the sentiment object corresponding to the text structure; and determine at least one text to be analyzed from the original text according to the sentiment information of each text structure;

[0065] Among them, the specific manner in which the determination module determines at least one text to be analyzed from the original text according to the sentiment information of each text structure includes:

[0066] When the sentiment information of all the text structures matches, determine the original text as the text to be analyzed;

[0067] When the sentiment information of all the text structures does not match, classify all the text structures according to the sentiment information of each text structure to obtain at least two texts to be analyzed, where if a text to be analyzed includes at least two of the text structures, the sentiment information of all the text structures in the text to be analyzed matches.

[0068] As an alternative implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the sentiment information of each text structure includes:

[0069] Determine at least one sentiment entity corresponding to each text structure, and determine the confirmation keywords in each text structure that meet the preset confirmation conditions;

[0070] For each text structure, according to the position order of each sentiment entity in the text structure and the position order of the confirmation keywords in the text structure, determine the entity attributes of each sentiment entity in the text structure as the sentiment information of the text structure;

[0071] Among them, the specific manner in which the determination module determines the confirmation keywords in each text structure that meet the preset confirmation conditions includes:

[0072] For each text structure, determine whether there is a first keyword for introducing a sentiment object in the text structure. When the determination result is yes, determine the first keyword as the confirmation keyword in the text structure that meets the preset confirmation conditions. When the determination result is no, determine the second keyword for describing the sentiment content in the text structure as the confirmation keyword in the text structure that meets the preset confirmation conditions.

[0073] The third aspect of the present invention discloses another text sentiment analysis device, and the device includes:

[0074] A memory storing executable program code;

[0075] A processor coupled to the memory;

[0076] The processor calls the executable program code stored in the memory and executes the text sentiment analysis method disclosed in the first aspect of the present invention.

[0077] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute the text sentiment analysis method disclosed in the first aspect of the present invention when called.

[0078] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0079] In an embodiment of the present invention, text information corresponding to the text to be analyzed is determined, where the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier; the text information is input into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result; according to the probability information corresponding to each sentiment analysis result, a target sentiment analysis result whose corresponding probability information meets a pre - set screening condition is determined from all the sentiment analysis results as the sentiment analysis result corresponding to the text to be analyzed; wherein, the text sentiment analysis model determines all the sentiment analysis results based on the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier. It can be seen that implementing the present invention can use the text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed according to the text vector and the position vector of the text to be analyzed, improve the relevance between different words in the text sentiment analysis process, thereby facilitating the accurate analysis of the overall semantics of the text, improving the accuracy and efficiency of text sentiment analysis. Moreover, by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation where the sentiment analysis method is relatively single and the focus of sentiment analysis is unreasonable due to using the simply concatenated text as the same input source to input into the text sentiment analysis model, further improving the accuracy of text sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0081] Figure 1 is a schematic flowchart of a text sentiment analysis method disclosed in an embodiment of the present invention;

[0082] Figure 2 is a schematic flowchart of another text sentiment analysis method disclosed in an embodiment of the present invention;

[0083] Figure 3 is a schematic structural diagram of a text sentiment analysis device disclosed in an embodiment of the present invention;

[0084] Figure 4 is a schematic structural diagram of another text sentiment analysis device disclosed in an embodiment of the present invention;

[0085] Figure 5 is a schematic structural diagram of yet another text sentiment analysis device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0087] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0088] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0089] The present invention discloses a text sentiment analysis method and device, which can use a text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed according to the text vector and position vector of the text to be analyzed, improve the relevance between different words in the process of text sentiment analysis, thereby facilitating the accurate analysis of the overall semantics of the text, improving the accuracy and efficiency of text sentiment analysis, and, by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation that the sentiment analysis method is relatively single and the focus of sentiment analysis is unreasonable due to using the simply concatenated text as the same input source to input into the text sentiment analysis model, and further improve the accuracy of text sentiment analysis. The following will be described in detail respectively.

[0090] Embodiment 1

[0091] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a text sentiment analysis method disclosed in an embodiment of the present invention. Among them, Figure 1The described text sentiment analysis method can be applied to the analysis process of the sentiment tendency of any text, such as movie reviews, which is not limited in the embodiments of the present invention. For example, Figure 1 As shown, the text sentiment analysis method may include the following operations:

[0092] 101. Determine the text information corresponding to the text to be analyzed.

[0093] In the embodiments of the present invention, the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier. Optionally, the sub - text identifier may include one or more of the text title, the specific text content, the text type, the text source, the text release time, etc. Preferably, the sub - text identifier includes the text title and the specific text content. Among them, when the text to be analyzed is used to train the text sentiment analysis model to be trained, the sub - text identifier may also include the text sentiment. This can improve the diversity of the text information. And by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation that the sentiment analysis method is relatively single and the focus of sentiment analysis is unreasonable due to taking the simply spliced text as the same input source and inputting it into the text sentiment analysis model, which further helps to improve the accuracy of text sentiment analysis.

[0094] As an optional implementation manner, determining the text information corresponding to the text to be analyzed may include:

[0095] According to at least one pre - set sub - text identifier, extract the text elements corresponding to each sub - text identifier in the text to be analyzed as the text content corresponding to the sub - text identifier.

[0096] For example, the text information may be "‘title (text title)’: movie review, ‘content (specific text content)’: This movie is really good...". When the text to be analyzed is used to train the text sentiment analysis model to be trained, the text information may be "‘title’: movie review, ‘content’: This movie is really good..., ‘label (sentiment label)’: positive".

[0097] It can be seen that implementing this optional implementation manner can extract the text elements in the text to be analyzed according to multiple sub - text identifiers, which is beneficial to accurately extract the text features in the text to be analyzed, improve the accuracy and reliability of determining the text information, and further help to improve the accuracy and efficiency of text sentiment analysis.

[0098] It should be noted that when determining the text information of the text to be analyzed, the text content corresponding to each sub - text identifier can also be determined by manual annotation.

[0099] In an embodiment of the present invention, optionally, the text length of the text to be analyzed needs to be less than or equal to the text length threshold corresponding to the text sentiment analysis model. This can reduce the occurrence of inaccurate text sentiment analysis caused by the inability of the text sentiment analysis model to fully analyze the text.

[0100] 102. Input the text information into a pre-trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result.

[0101] In an embodiment of the present invention, the text sentiment analysis model determines all sentiment analysis results based on the text vector corresponding to each sub-text identifier and the position vector corresponding to each sub-text identifier. Among them, the text vector corresponding to each sub-text identifier is used to represent the semantic information corresponding to each text element in the text content corresponding to the sub-text identifier, and the position vector corresponding to each sub-text identifier is used to represent the position order corresponding to each text element in the text content corresponding to the sub-text identifier. The text elements may include words and / or single characters. In this way, by determining the sentiment analysis results of the text to be analyzed based on the text vector and the position vector, the relevance between different words in the text sentiment analysis process can be improved, which is conducive to accurately analyzing the overall semantics of the text and improving the accuracy of text sentiment analysis.

[0102] Preferably, the text sentiment analysis model is a BERT model. This is beneficial to accurately capturing the global semantic features of the text while capturing the local features of the text, improving the relevance between different text structures in the text sentiment analysis process, and thus improving the accuracy and reliability of text sentiment analysis.

[0103] As an alternative implementation, the processing of the text information by the text sentiment analysis model may include:

[0104] Performing an encoding operation on the text information by the encoding structure corresponding to the text sentiment analysis model to obtain the encoding result corresponding to each sub-text identifier; the encoding result corresponding to each sub-text identifier includes the text vector corresponding to each sub-text identifier and the position vector corresponding to each sub-text identifier;

[0105] Performing a vector conversion process on the encoding result corresponding to each sub-text identifier by the vector processing structure of the text sentiment analysis model to obtain the target matrix corresponding to each sub-text identifier;

[0106] Performing an average pooling operation on the target matrix corresponding to the sub-text identifier by the average pooling structure of the text sentiment analysis model to obtain the average pooling result corresponding to each sub-text identifier;

[0107] The concatenation operation is performed on the average pooling results corresponding to all sub - text identifiers by the concatenation structure of the text sentiment analysis model to obtain a concatenation result;

[0108] The fully - connected layer of the text sentiment analysis model processes the concatenation result.

[0109] It can be seen that implementing this optional implementation method can perform vector transformation, average pooling, and concatenation operations on the text vector and position vector of the text, which is beneficial to improving the relevance between different text structures in the text sentiment analysis process, and thus is beneficial to accurately analyzing the global semantic information of the text. Moreover, performing a fully - connected process on the concatenation result can classify the sentiment corresponding to different text structures, which is beneficial to improving the accuracy and reliability of text sentiment analysis.

[0110] In this optional implementation method, optionally, the vector processing structure includes multiple processing units arranged in sequence. Among them, the first processing unit uses the encoding result corresponding to each sub - text identifier as input, and any other processing unit except the first processing unit uses the output of the previous processing unit corresponding to this processing unit as input, which can gradually improve the globality of the extracted semantic information.

[0111] Preferably, the vector processing structure is a transformer structure, and the processing unit is an encoder structure in the transformer structure, which is beneficial to improving the relevance between different text structures in the text through the attention mechanism of the transformer structure, so as to accurately analyze the global semantic information of the text.

[0112] Optionally, the target matrix corresponding to each sub - text identifier may include the global semantic analysis result corresponding to this sub - text identifier determined based on the text vector corresponding to this sub - text identifier and the position vector corresponding to this sub - text identifier. The global semantic analysis result corresponding to each sub - text identifier may include the global semantic vector corresponding to this sub - text identifier or the global semantic matrix corresponding to this sub - text identifier. Further optionally, the target matrix corresponding to each sub - text identifier may also include the own parameters of the text sentiment analysis model, such as the amount of data that can be processed at one time when the text sentiment analysis model analyzes the text, the number of nodes in the fully - connected layer of the text sentiment analysis model, etc. This can improve the diversity of the information contained in the target matrix.

[0113] In this optional implementation method, optionally, the encoding operation is performed on the text information by the encoding structure corresponding to the text sentiment analysis model to obtain the encoding result corresponding to each sub - text identifier, which may include:

[0114] The encoding structure corresponding to the text sentiment analysis model determines the mapping encoding corresponding to each text element in the text content corresponding to each sub - text identifier according to the vocabulary corresponding to the encoding structure, and generates the text vector corresponding to the sub - text identifier according to the mapping encoding corresponding to each text element in the text content corresponding to each sub - text identifier;

[0115] The encoding structure generates the position vector corresponding to the sub - text identifier according to the position order of each text element in the text content corresponding to each sub - text identifier.

[0116] It should be noted that the generation steps of the text vector and the position vector have no sequence relationship.

[0117] It can be seen that implementing this optional implementation can improve the efficiency and accuracy of the encoding of the text sentiment analysis model.

[0118] In this optional implementation, further optionally, the method may further include:

[0119] Judge whether the text length of the text content corresponding to each sub - text identifier is less than the preset text length. When the judgment result is yes, fill a target number of preset vector elements at the preset positions of the text vector corresponding to the sub - text identifier and the position vector corresponding to the sub - text identifier, where the target number matches the difference in the text length corresponding to the sub - text identifier, and the difference in the text length corresponding to the sub - text identifier includes the difference in length between the text length of the text content corresponding to the sub - text identifier and the preset text length.

[0120] It can be seen that implementing this optional implementation can improve the matching degree between the vector lengths of the text vector and the position vector and the model parameters of the text sentiment analysis model, and when the text sentiment analysis model is a BERT model, it can improve the reliability of the attention mechanism of the BERT model.

[0121] In this optional implementation, optionally, the vector processing structure of the text sentiment analysis model performs vector conversion processing on the encoding result corresponding to each sub - text identifier to obtain the target matrix corresponding to each sub - text identifier, which may include:

[0122] When the text sentiment analysis model includes only one vector processing structure, such as the Siamese BERT model, the vector processing structure performs vector conversion processing on the encoding result corresponding to each text identifier to obtain the target matrix corresponding to each sub - text identifier;

[0123] When the text sentiment analysis model includes multiple text processing structures, such as the Dual BERT model, for each sub - text identifier, the vector processing structure matched with the sub - text identifier performs vector conversion processing on the encoding result corresponding to the sub - text identifier to obtain the target matrix corresponding to the sub - text identifier.

[0124] In this optional embodiment, optionally, different sub - text identifiers correspond to different vector processing structures.

[0125] It can be seen that implementing this optional embodiment can also analyze text sentiment using different types of text sentiment analysis models, improving the diversity and flexibility of the text sentiment analysis method. Moreover, when different text processing structures are used to process the encoding results corresponding to different sub - text identifiers, the matching degree between the text type and the structural parameters of the text processing structure can be improved, reducing the situation where some encoding results do not match the text processing structure and thus the encoding results are inaccurately processed when using the same text processing structure to process the encoding results corresponding to multiple sub - text identifiers. This is conducive to improving the accuracy and reliability of text sentiment analysis.

[0126] 103. According to the probability information corresponding to each sentiment analysis result, determine, from all sentiment analysis results, a target sentiment analysis result whose corresponding probability information meets a preset screening condition as the sentiment analysis result corresponding to the text to be analyzed.

[0127] In the embodiment of the present invention, optionally, the probability information corresponding to each sentiment analysis result may be the probability value corresponding to this sentiment analysis result. This probability value may be a normalized probability value or a non - normalized probability value, which is not limited in the embodiment of the present invention.

[0128] As an optional implementation manner, according to the probability information corresponding to each sentiment analysis result, determining, from all sentiment analysis results, a target sentiment analysis result whose corresponding probability information meets a preset screening condition as the sentiment analysis result corresponding to the text to be analyzed may include:

[0129] According to the probability information corresponding to each of the sentiment analysis results, determine, as the target sentiment analysis result that meets the preset screening condition and is used as the sentiment analysis result corresponding to the text to be analyzed, the sentiment analysis result whose probability value represented by the corresponding probability information among all the sentiment analysis results is the maximum probability value.

[0130] It can be seen that implementing this optional implementation manner can use the sentiment analysis result with the maximum probability value as the sentiment analysis result of the text to be analyzed, improving the accuracy of text sentiment analysis.

[0131] In an optional embodiment, before determining the text information corresponding to the text to be analyzed, the method may further include:

[0132] When the original text of the sentiment to be analyzed includes at least two text structures, determine the sentiment information of each text structure, where the sentiment information of each text structure includes the sentiment subject corresponding to this text structure and / or the sentiment object corresponding to this text structure;

[0133] Determine at least one text to be analyzed from the original text according to the sentiment information of each text structure.

[0134] In this optional embodiment, optionally, the text structure may include text paragraphs or text statements.

[0135] It can be seen that implementing this optional embodiment can determine the text to be analyzed from the original text according to the sentiment subject and sentiment object of different text structures, which is beneficial to improving the matching degree between the text to be analyzed and the actual requirements, and is beneficial to reducing the unnecessary text structures in the text to be analyzed, and improving the efficiency and accuracy of text sentiment analysis.

[0136] In this optional embodiment, as an optional implementation manner, determining at least one text to be analyzed from the original text according to the sentiment information of each text structure may include:

[0137] When the sentiment information of all text structures matches, determine the original text as the text to be analyzed;

[0138] When the sentiment information of all text structures does not match, classify all text structures according to the sentiment information of each text structure to obtain at least two texts to be analyzed, where if a text to be analyzed includes at least two text structures, the sentiment information of all text structures in this text to be analyzed matches.

[0139] It can be seen that implementing this optional implementation manner can uniformly process all text structures when the sentiment information of all text structures matches, improving the efficiency of text sentiment analysis, and when the sentiment information of all text structures does not match, classify the text structures according to the sentiment subject and sentiment object and process them separately, which is beneficial to improving the singularity of the sentiment information of the text to be analyzed in the input model, reducing the difficulty and complexity of text sentiment analysis, improving the efficiency and accuracy of text sentiment analysis, and is beneficial to reducing the situation where the sentiment data of some text structures is lost or submerged by the sentiment data of other text structures due to focusing on the global semantic information during text sentiment analysis, thereby improving the comprehensiveness and accuracy of text sentiment analysis.

[0140] In this optional implementation manner, optionally, when the sentiment information of all text structures does not match, the method may further include:

[0141] After determining the sentiment analysis result corresponding to each text to be analyzed, determine the sentiment analysis result corresponding to the original text according to the sentiment analysis result corresponding to each text to be analyzed.

[0142] It can be seen that implementing this optional implementation manner can further improve the comprehensiveness of text sentiment analysis.

[0143] In this optional embodiment, as another optional implementation manner, determining the sentiment information of each text structure may include:

[0144] Determining at least one sentiment entity corresponding to each text structure, and determining the confirmation keywords in each text structure that meet the preset confirmation conditions;

[0145] For each text structure, according to the position order of each sentiment entity in the text structure and the position order of the confirmation keywords in the text structure, determine the entity attributes of each sentiment entity in the text structure as the sentiment information of the text structure.

[0146] It can be seen that implementing this optional implementation manner can determine the entity attributes of sentiment entities according to the position order of sentiment entities and confirmation keywords in the text structure, obtain the sentiment information of the text structure, and improve the accuracy and reliability of determining the sentiment information of the text structure.

[0147] In this optional implementation manner, optionally, determining the confirmation keywords in each text structure that meet the preset confirmation conditions may include:

[0148] For each text structure, determine whether there is a first keyword for introducing a sentiment object in the text structure. When the determination result is yes, determine the first keyword as the confirmation keyword in the text structure that meets the preset confirmation conditions. When the determination result is no, determine the second keyword for describing the sentiment content in the text structure as the confirmation keyword in the text structure that meets the preset confirmation conditions.

[0149] For example, in the text structure A "I have a special preference for bread", the sentiment entities include "I" and "bread", and there is a first keyword "for" for introducing the sentiment object "bread" in this text structure. Since "I" is before "for" and "bread" is after "for", "I" is the sentiment subject and "bread" is the sentiment object; in the text structure B "She likes him", the sentiment entities include "She" and "him", and it can be determined that the second keyword "likes" in this text structure is the confirmation keyword. Since "She" is before "likes" and "him" is after "likes", "She" is the sentiment subject and "him" is the sentiment object; in the text structure C "Dogs are worthy of the love of all mankind", the sentiment entities include "Dogs" and "all mankind", and it can be determined that the second keyword "love" in this text structure is the confirmation keyword. Although both "Dogs" and "all mankind" are before "love", "all mankind" is closer to "love". Therefore, it can be determined that "all mankind" is the sentiment subject and "Dogs" is the sentiment object.

[0150] It can be seen that implementing this optional implementation manner can also determine the confirmation keyword of the text structure according to the type of keyword included in the text structure, that is, adopt different emotional information determination methods for different types of text structures, improving the diversity, flexibility, and accuracy of the emotional information determination method.

[0151] It can be seen that implementing the present invention can use a text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed according to the text vector and position vector of the text to be analyzed, improving the correlation between different words in the text sentiment analysis process, thereby facilitating the accurate analysis of the overall semantics of the text and improving the accuracy and efficiency of the text sentiment analysis. Moreover, by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation where the sentiment analysis method is relatively single and the sentiment analysis focus is unreasonable due to using the simply concatenated text as the same input source to input into the text sentiment analysis model, further improving the accuracy of the text sentiment analysis.

[0152] Embodiment 2

[0153] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another text sentiment analysis method disclosed in the embodiments of the present invention. Among them, Figure 2 the described text sentiment analysis method can be applied to the analysis process of the sentiment tendency of any text, such as movie reviews, and the embodiments of the present invention do not make limitations. As Figure 2 shown, the text sentiment analysis method may include the following operations:

[0154] 201. Determine the text information corresponding to the text to be analyzed, where the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier.

[0155] 202. Perform a pre - processing operation on the text information.

[0156] As an optional implementation manner, performing a pre - processing operation on the text information may include:

[0157] For the text content corresponding to each sub - text identifier, detect whether there is a text element to be removed in the text content that meets the preset removal condition. When the detection result is yes, remove the text element to be removed in the text content.

[0158] It can be seen that implementing this optional implementation manner is beneficial to reducing the unnecessary text information in the text content and is beneficial to improving the matching degree between the text content and the requirements of the text sentiment analysis model, improving the efficiency and accuracy of the text sentiment analysis.

[0159] In this optional implementation manner, optionally, for the text content corresponding to each sub-text identifier, detecting whether there is a text element to be removed in the text content that meets a preset removal condition may include:

[0160] For the text content corresponding to each sub-text identifier, detecting whether there is a first text element in the text content whose element type matches a preset first element type. When it is detected that there is a first text element in the text content, it is determined that there is a text element to be removed in the text content that meets the preset removal condition, where the text element to be removed includes the first text element; and / or,

[0161] For the text content corresponding to each sub-text identifier, detecting whether there are second text elements that are adjacent, have the same element type, and whose element type matches a preset second element type in the text content. When it is detected that there are second text elements in the text content, it is determined that there is a text element to be removed in the text content that meets the preset removal condition, where the text element to be removed includes some of the second text elements.

[0162] In this optional implementation manner, further optionally, the first element type may include an emoji type, and the second element type may include a punctuation type.

[0163] It can be seen that implementing this optional implementation manner can also remove text elements with incorrect formats or repeated and adjacent ones in the text content, increase the possibility of the text sentiment analysis model successfully analyzing the text sentiment, be beneficial to reducing the space occupied by non-essential text information in the text content, improve the efficiency of text sentiment analysis, and can also reduce the situation where the necessary text information is lost due to the excessive text content caused by the excessive space occupied by non-essential text information, improving the accuracy and reliability of text sentiment analysis.

[0164] In this optional implementation manner, further optionally, before removing the text element to be removed in the text content, the method may further include:

[0165] Judging whether the influence degree of the text element to be removed in the text content on the sentiment of the text content is greater than a preset degree. When the judgment result is no, triggering the operation of removing the text element to be removed in the text content as described above.

[0166] It can be seen that implementing this optional implementation manner can remove the text element only when the influence degree of the text element to be removed on the sentiment of the text content is relatively low, reducing the situation where the text semantics cannot be accurately analyzed due to the removal of some text elements, and improving the accuracy of text sentiment analysis.

[0167] In this optional embodiment, further optionally, for the text content corresponding to each sub - text identifier, when it is determined that the influence degree of the text elements to be removed in the text content on the sentiment of the text content is greater than a preset degree, the method further includes:

[0168] Determine whether the element form of the text elements to be removed in the sub - text content matches the target element form corresponding to the text sentiment analysis model. When the determination result is negative, convert the element form of the text elements to be removed into the target element form.

[0169] For example, if the target element form is in character form, punctuation marks (such as "??") can retain their original element form, while emoticons need to be converted into the corresponding emoticon characters. For example, a smiling face emoticon can be converted into "happy".

[0170] It can be seen that implementing this optional embodiment can convert the text elements of the text to be removed with a greater influence on text sentiment into the element form required by the text sentiment analysis model, thereby increasing the possibility of the text sentiment model successfully analyzing text sentiment and the accuracy and comprehensiveness of text sentiment analysis.

[0171] 203. Input the text information into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result.

[0172] 204. According to the probability information corresponding to each sentiment analysis result, determine, from all the sentiment analysis results, the target sentiment analysis result whose corresponding probability information meets the preset screening conditions as the sentiment analysis result corresponding to the text to be analyzed.

[0173] It should be noted that for other descriptions of steps 201, 203, and 204, please refer to the detailed descriptions of steps 101 - 103 in Embodiment 1 of the present invention, and the embodiments of the present invention will not be elaborated herein.

[0174] It can be seen that implementing the embodiments of the present invention can use the text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed according to the text vector and the position vector of the text to be analyzed, improve the relevance between different words in the text sentiment analysis process, thereby facilitating the accurate analysis of the overall semantics of the text, improving the accuracy and efficiency of text sentiment analysis. Moreover, inputting the text information into the text sentiment analysis model after preprocessing is conducive to improving the matching degree between the text content of the text information and the text requirements of the text sentiment analysis model. Additionally, by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation where the sentiment analysis method is relatively single and the focus of sentiment analysis is unreasonable due to using the simply concatenated text as the same input source to input into the text sentiment analysis model, further improving the accuracy of text sentiment analysis.

[0175] Embodiment Three

[0176] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a text sentiment analysis device disclosed in the embodiments of the present invention. Among them, Figure 3 the described text sentiment analysis device can be applied to the analysis process of the sentiment tendency of any text, such as movie reviews, which is not limited in the embodiments of the present invention. As Figure 3 shown, the text sentiment analysis device may include:

[0177] A determination module 301, configured to determine the text information corresponding to the text to be analyzed, where the text information includes at least one sub-text identifier and the text content corresponding to each sub-text identifier;

[0178] A sentiment analysis module 302, configured to input the text information into a pre-trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result;

[0179] The determination module 301 is further configured to determine, according to the probability information corresponding to each sentiment analysis result, a target sentiment analysis result whose corresponding probability information meets a preset screening condition from all the sentiment analysis results as the sentiment analysis result corresponding to the text to be analyzed;

[0180] Among them, the text sentiment analysis model determines all sentiment analysis results based on the text vector corresponding to each sub-text identifier and the position vector corresponding to each sub-text identifier.

[0181] It can be seen that implementing Figure 3The described device can use a text sentiment analysis model to determine the sentiment analysis result of the text to be analyzed based on the text vector and position vector of the text to be analyzed, improve the correlation between different words in the text sentiment analysis process, thereby facilitating the accurate analysis of the overall semantics of the text, and improving the accuracy and efficiency of text sentiment analysis. Moreover, by processing the text into multiple input sources and then inputting them into the text sentiment analysis model for analysis, it is beneficial to reduce the situation where the sentiment analysis method is relatively single and the focus of sentiment analysis is unreasonable due to using the simply concatenated text as the same input source to input into the text sentiment analysis model, and further improve the accuracy of text sentiment analysis.

[0182] In an alternative embodiment, as Figure 3 shown, the specific manner in which the text sentiment analysis model processes text information may include:

[0183] Performing an encoding operation on the text information by the encoding structure corresponding to the text sentiment analysis model to obtain the encoding result corresponding to each sub - text identifier; the encoding result corresponding to each sub - text identifier includes the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier;

[0184] Performing a vector conversion process on the encoding result corresponding to each sub - text identifier by the vector processing structure of the text sentiment analysis model to obtain the target matrix corresponding to each sub - text identifier;

[0185] Performing an average pooling operation on the target matrix corresponding to the sub - text identifier by the average pooling structure of the text sentiment analysis model to obtain the average pooling result corresponding to each sub - text identifier;

[0186] Performing a concatenation operation on the average pooling results corresponding to all sub - text identifiers by the concatenation structure of the text sentiment analysis model to obtain a concatenation result;

[0187] Processing the concatenation result by the fully - connected layer of the text sentiment analysis model.

[0188] It can be seen that implementing Figure 3 the described device can also perform vector transformation, average pooling, and concatenation operations on the text vector and position vector of the text, which is beneficial to improving the correlation between different text structures in the text sentiment analysis process, thereby facilitating the accurate analysis of the global semantic information of the text. Moreover, performing a fully - connected process on the concatenation result can classify the sentiment corresponding to different text structures, thereby being beneficial to improving the accuracy and reliability of text sentiment analysis.

[0189] In another alternative embodiment, as Figure 3As shown, the specific manner in which the vector processing structure of the text sentiment analysis model performs vector conversion processing on the encoding results corresponding to each sub - text identifier to obtain the target matrix corresponding to each sub - text identifier may include:

[0190] When the text sentiment analysis model includes only one vector processing structure, the vector processing structure performs vector conversion processing on the encoding results corresponding to each text identifier to obtain the target matrix corresponding to each sub - text identifier;

[0191] When the text sentiment analysis model includes multiple text processing structures, for each sub - text identifier, the vector processing structure matching the sub - text identifier performs vector conversion processing on the encoding result corresponding to the sub - text identifier to obtain the target matrix corresponding to the sub - text identifier.

[0192] It can be seen that implementing Figure 3 the described device can also analyze text sentiment using different types of text sentiment analysis models, improving the diversity and flexibility of the text sentiment analysis method. Moreover, when different text processing structures are used to process the encoding results corresponding to different sub - text identifiers, the matching degree between the text type and the structural parameters of the text processing structure can be improved, reducing the situation where some encoding results do not match the text processing structure and thus leading to inaccurate processing of the encoding results when using the same text processing structure to process the encoding results corresponding to multiple sub - text identifiers. This is conducive to improving the accuracy and reliability of text sentiment analysis.

[0193] In yet another alternative embodiment, as Figure 4 shown, the specific manner in which the determination module 301 determines the text information corresponding to the text to be analyzed may include:

[0194] According to at least one pre - set sub - text identifier, extract the text elements corresponding to each sub - text identifier in the text to be analyzed as the text content corresponding to the sub - text identifier;

[0195] And, the device may further include:

[0196] A pre - processing module 303, configured to perform pre - processing operations on the text information before the sentiment analysis module 302 inputs the text information into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result;

[0197] Among them, the specific manner in which the pre - processing module 303 performs pre - processing operations on the text information includes:

[0198] For the text content corresponding to each sub - text identifier, detect whether there is a text element to be removed in the text content that meets the preset clearing conditions. When the detection result is yes, remove the text element to be removed from the text content.

[0199] It can be seen that implementing Figure 4 the described device can extract text elements in the text to be analyzed according to multiple sub - text identifiers, which is beneficial to accurately extract text features in the text to be analyzed, improve the accuracy and reliability of determining text information, and further improve the accuracy and efficiency of text sentiment analysis. In addition, inputting the text information into the text sentiment analysis model only after pre - processing the text information is beneficial to improving the matching degree between the text content of the text information and the text requirements of the text sentiment analysis model, and removing text elements to be removed that meet the preset clearing conditions during pre - processing is beneficial to reducing unnecessary text information in the text content and improving the matching degree between the text content and the text requirements of the text sentiment analysis model, further improving the efficiency and accuracy of text sentiment analysis.

[0200] In another alternative embodiment, as Figure 4 shown, the specific ways for the pre - processing module 303 to detect whether there is a text element to be removed in the text content that meets the preset clearing conditions for the text content corresponding to each sub - text identifier may include:

[0201] For the text content corresponding to each sub - text identifier, detect whether there is a first text element whose element type matches the preset first element type. When it is detected that there is a first text element in the text content, determine that there is a text element to be removed in the text content that meets the preset clearing conditions, where the text element to be removed includes the first text element; and / or,

[0202] For the text content corresponding to each sub - text identifier, detect whether there are adjacent second text elements with the same element type and whose element type matches the preset second element type. When it is detected that there are second text elements in the text content, determine that there is a text element to be removed in the text content that meets the preset clearing conditions, where the text element to be removed includes some of the second text elements;

[0203] In addition, the pre - processing module 303 is further configured to, before performing the operation of removing the text element to be removed from the text content, determine whether the influence degree of the text element to be removed on the sentiment of the text content is greater than a preset degree. When the judgment result is no, trigger the execution of the operation of removing the text element to be removed from the text content.

[0204] It can be seen that implementing Figure 4The described device can also remove text elements with incorrect formats or adjacent duplicates in the text content, increasing the likelihood of the text sentiment analysis model successfully analyzing the text sentiment, facilitating the reduction of the space occupied by non-essential text information in the text content, improving the efficiency of text sentiment analysis, and reducing the occurrence of situations where the necessary text information is lost due to excessive text content caused by the overoccupation of space by non-essential text information, thereby improving the accuracy and reliability of text sentiment analysis. Additionally, removing the text element only when the degree of influence of the text element to be removed on the sentiment of the text content is low can reduce the occurrence of situations where the text semantics cannot be accurately analyzed due to the removal of some text elements, further improving the accuracy of text sentiment analysis.

[0205] In yet another alternative embodiment, as Figure 4 shown, the determination module 301 is further configured to, before performing the operation of determining the text information corresponding to the text to be analyzed described above, when the original text whose sentiment is to be analyzed includes at least two text structures, determine the sentiment information of each text structure, where the sentiment information of each text structure includes the sentiment subject corresponding to the text structure and / or the sentiment object corresponding to the text structure; and determine at least one text to be analyzed from the original text according to the sentiment information of each text structure.

[0206] Among them, the specific manner in which the determination module 301 determines at least one text to be analyzed from the original text according to the sentiment information of each text structure may include:

[0207] When the sentiment information of all text structures matches, determine the original text as the text to be analyzed;

[0208] When the sentiment information of all text structures does not match, classify all text structures according to the sentiment information of each text structure to obtain at least two texts to be analyzed, where if a text to be analyzed includes at least two text structures, the sentiment information of all text structures in the text to be analyzed matches.

[0209] It can be seen that implementing Figure 4 the described device can also uniformly process all text structures when the sentiment information of all text structures matches, improving the efficiency of text sentiment analysis, and when the sentiment information of all text structures does not match, classify the text structures according to the sentiment subject and sentiment object and process them separately, thereby facilitating the improvement of the singularity of the sentiment information of the text to be analyzed in the input model, reducing the difficulty and complexity of text sentiment analysis, improving the efficiency and accuracy of text sentiment analysis, and facilitating the reduction of the occurrence of situations where the sentiment data of some text structures is lost or submerged by the sentiment data of other text structures due to excessive focus on global semantic information during the text sentiment analysis process, thereby improving the comprehensiveness and accuracy of text sentiment analysis.

[0210] In yet another alternative embodiment, as Figure 4 shown, the specific manner in which the determination module 301 determines the sentiment information of each text structure may include:

[0211] Determining at least one sentiment entity corresponding to each text structure, and determining confirmation keywords in each text structure that meet a preset confirmation condition;

[0212] For each text structure, according to the position order of each sentiment entity in the text structure and the position order of the confirmation keywords in the text structure, determine the entity attribute of each sentiment entity in the text structure as the sentiment information of the text structure;

[0213] Among them, the specific manner in which the determination module 301 determines the confirmation keywords in each text structure that meet the preset confirmation condition includes:

[0214] For each text structure, determine whether there is a first keyword for introducing a sentiment object in the text structure. When the determination result is yes, determine the first keyword as the confirmation keyword in the text structure that meets the preset confirmation condition. When the determination result is no, determine the second keyword for describing the sentiment content in the text structure as the confirmation keyword in the text structure that meets the preset confirmation condition.

[0215] It can be seen that implementing Figure 4 the described device can also determine the entity attribute of the sentiment entity according to the position order of the sentiment entity and the confirmation keyword in the text structure, obtain the sentiment information of the text structure, improve the accuracy and reliability of determining the sentiment information of the text structure, and, by determining the confirmation keyword of the text structure according to the type of keyword included in the text structure, that is, adopting different sentiment information determination methods for different types of text structures, improve the diversity, flexibility, and accuracy of the sentiment information determination method.

[0216] Embodiment 4

[0217] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another text sentiment analysis device disclosed in the embodiments of the present invention. As Figure 5 shown, the text sentiment analysis device may include:

[0218] A memory 401 storing executable program code;

[0219] A processor 402 coupled to the memory 401;

[0220] The processor 402 calls the executable program code stored in the memory 401 and executes the steps in the text sentiment analysis method described in Embodiment 1 or Embodiment 2 of the present invention.

[0221] Embodiment 5

[0222] An embodiment of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the steps in the text sentiment analysis method described in Embodiment 1 or Embodiment 2 of the present invention when the computer instructions are called.

[0223] Embodiment 6

[0224] An embodiment of the present invention discloses a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the text sentiment analysis method described in Embodiment 1 or Embodiment 2.

[0225] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0226] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0227] Finally, it should be noted that: What is disclosed in an embodiment of a text sentiment analysis method and device of the present invention is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention and is not intended to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text sentiment analysis method, characterized in that, the method includes: determining text information corresponding to the text to be analyzed, where the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier; inputting the text information into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result; determining, according to the probability information corresponding to each sentiment analysis result, a target sentiment analysis result whose corresponding probability information meets a pre - set screening condition from all the sentiment analysis results as the sentiment analysis result corresponding to the text to be analyzed; wherein, the text sentiment analysis model determines all the sentiment analysis results based on the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier; and, before determining the text information corresponding to the text to be analyzed, the method further includes: when the original text to be analyzed for sentiment includes at least two text structures, determining the sentiment information of each text structure, where the sentiment information of each text structure includes the sentiment subject and / or the sentiment object corresponding to the text structure, and the text structure includes a text paragraph or a text sentence; determining at least one text to be analyzed from the original text according to the sentiment information of each text structure; and, the determining at least one text to be analyzed from the original text according to the sentiment information of each text structure includes: when the sentiment information of all the text structures matches, determining the original text as the text to be analyzed; when the sentiment information of all the text structures does not match, classifying all the text structures according to the sentiment information of each text structure to obtain at least two texts to be analyzed, where if a text to be analyzed includes at least two text structures, the sentiment information of all the text structures in the text to be analyzed matches.

2. The text sentiment analysis method according to claim 1, characterized in that, the processing of the text information by the text sentiment analysis model includes: performing an encoding operation on the text information by the encoding structure corresponding to the text sentiment analysis model to obtain an encoding result corresponding to each sub - text identifier; the encoding result corresponding to each sub - text identifier includes the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier; performing a vector conversion process on the encoding result corresponding to each sub - text identifier by the vector processing structure of the text sentiment analysis model to obtain a target matrix corresponding to each sub - text identifier; performing an average pooling operation on the target matrix corresponding to the sub - text identifier by the average pooling structure of the text sentiment analysis model to obtain an average pooling result corresponding to each sub - text identifier; performing a splicing operation on the average pooling results corresponding to all the sub - text identifiers by the splicing structure of the text sentiment analysis model to obtain a splicing result; The fully connected layer of the text sentiment analysis model processes the splicing result.

3. The text sentiment analysis method according to claim 2, wherein, the vector processing structure of the text sentiment analysis model performs vector conversion processing on the encoding result corresponding to each sub - text identifier to obtain a target matrix corresponding to each sub - text identifier, including: when the text sentiment analysis model only includes one vector processing structure, the vector processing structure performs vector conversion processing on the encoding result corresponding to each text identifier to obtain a target matrix corresponding to each sub - text identifier; when the text sentiment analysis model includes multiple text processing structures, for each sub - text identifier, the vector processing structure matched with the sub - text identifier performs vector conversion processing on the encoding result corresponding to the sub - text identifier to obtain a target matrix corresponding to the sub - text identifier.

4. The text sentiment analysis method according to any one of claims 1 - 3, wherein, the determining the text information corresponding to the text to be analyzed includes: extracting text elements corresponding to each sub - text identifier in the text to be analyzed according to at least one pre - set sub - text identifier as the text content corresponding to the sub - text identifier; and, before inputting the text information into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and probability information corresponding to each sentiment analysis result, the method further includes: performing a pre - processing operation on the text information; wherein, the performing a pre - processing operation on the text information includes: for the text content corresponding to each sub - text identifier, detecting whether there is a text element to be removed that meets a preset clearing condition in the text content, and when the detection result is yes, removing the text element to be removed in the text content.

5. The text sentiment analysis method according to claim 4, wherein, the detecting whether there is a text element to be removed that meets a preset clearing condition in the text content corresponding to each sub - text identifier includes: for the text content corresponding to each sub - text identifier, detecting whether there is a first text element whose element type matches a pre - set first element type in the text content, and when it is detected that there is the first text element in the text content, determining that there is a text element to be removed that meets a preset clearing condition in the text content, wherein the text element to be removed includes the first text element; and / or, for the text content corresponding to each sub - text identifier, detecting whether there are second text elements that are adjacent, have the same element type, and the element type matches a pre - set second element type in the text content, and when it is detected that there are the second text elements in the text content, determining that there is a text element to be removed that meets a preset clearing condition in the text content, wherein the text element to be removed includes some of the second text elements; Also, before removing the text elements to be removed from the text content, the method further includes: Determining whether the degree of influence of the text elements to be removed in the text content on the sentiment of the text content is greater than a preset degree. When the determination result is negative, triggering the operation of removing the text elements to be removed from the text content.

6. The text sentiment analysis method according to claim 5, characterized in that the determining the sentiment information of each text structure includes: determining at least one sentiment entity corresponding to each text structure, and determining confirmation keywords in each text structure that meet preset confirmation conditions; for each text structure, according to the position order of each sentiment entity in the text structure and the position order of the confirmation keywords in the text structure, determining the entity attribute of each sentiment entity in the text structure as the sentiment information of the text structure; wherein, the determining the confirmation keywords in each text structure that meet preset confirmation conditions includes: for each text structure, determining whether there is a first keyword for introducing a sentiment object in the text structure. When the determination result is positive, determining the first keyword as the confirmation keyword in the text structure that meets preset confirmation conditions. When the determination result is negative, determining a second keyword for describing sentiment content in the text structure as the confirmation keyword in the text structure that meets preset confirmation conditions.

7. A text sentiment analysis device, characterized in that the device includes: a determination module, configured to determine text information corresponding to the text to be analyzed, where the text information includes at least one sub - text identifier and the text content corresponding to each sub - text identifier; a sentiment analysis module, configured to input the text information into a pre - trained text sentiment analysis model to trigger the text sentiment analysis model to process the text information and output multiple sentiment analysis results and the probability information corresponding to each sentiment analysis result; the determination module is further configured to determine, according to the probability information corresponding to each sentiment analysis result, a target sentiment analysis result whose corresponding probability information meets a preset screening condition from all the sentiment analysis results as the sentiment analysis result corresponding to the text to be analyzed; wherein, the text sentiment analysis model determines all the sentiment analysis results based on the text vector corresponding to each sub - text identifier and the position vector corresponding to each sub - text identifier; the determination module is further configured to, before performing the operation of determining the text information corresponding to the text to be analyzed, when the original text of the sentiment to be analyzed includes at least two text structures, determine the sentiment information of each text structure, where the sentiment information of each text structure includes the sentiment subject and / or the sentiment object corresponding to the text structure, and the text structure includes a text paragraph or a text sentence; and determine at least one text to be analyzed from the original text according to the sentiment information of each text structure. Moreover, the specific manner in which the determining module determines at least one text to be analyzed from the original text according to the sentiment information of each text structure includes: When the sentiment information of all the text structures matches, the original text is determined as the text to be analyzed; When the sentiment information of all the text structures does not match, all the text structures are classified according to the sentiment information of each text structure to obtain at least two texts to be analyzed, wherein if a text to be analyzed includes at least two of the text structures, the sentiment information of all the text structures in the text to be analyzed matches.

8. A text sentiment analysis device Characterized in that The device includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the text sentiment analysis method according to any one of claims 1-6.

9. A computer storage medium Characterized in that The computer storage medium stores computer instructions which, when called, are used to execute the text sentiment analysis method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Text analysis method, device, computer device, and storage medium

    CN109271627A

  • Information analysis method and device, electronic equipment and computer readable storage medium

    CN113609390A