Data preprocessing method and device for text emotional tendency analysis

By performing sentence segmentation and dictionary analysis of emotional tendency on the text, the situation where words in the text are opposite to the actual expression of semantics or unclear emotions and evaluations are eliminated, and the problem of semantic distortion in the text's emotional tendency analysis is solved, and the accuracy of the analysis is improved.

CN119990126APending Publication Date: 2025-05-13GRG BANKING IT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205193.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the analysis of text emotional tendency, the semantic diversity and flexibility of texts have caused natural language processing to distort the true semantics of text, and the analysis results are inaccurate.

Method used

A data preprocessing method is proposed to clear the situation where words in the text are opposite to the actual expression semantics or the emotions and evaluation are unclear through statement segmentation, emotional tendency dictionary analysis and rewriting or removal operations.

Benefits of technology

It improves the accuracy of text emotional tendency analysis and avoids the emotional tendency that distorts the truth of the text during the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990126A_ABST
    Figure CN119990126A_ABST
Patent Text Reader

Abstract

The invention discloses a data preprocessing method and device for text emotional tendency analysis, and belongs to the technical field of computers. The method comprises the following steps: performing statement segmentation processing on target text data to obtain a first target statement; for each first target statement, based on a pre-stored emotional tendency dictionary library, analyzing whether the first target statement contains a phrase of a first target structure or a phrase of a second target structure; performing rewriting operation on the first target statement containing the phrase of the first target structure, and / or performing removal operation on the first target statement containing the phrase of the second target structure, the condition that the use words in the first target statement are opposite to the actual expression semantics of the first target statement or the emotion and evaluation expressed by the first target statement are indefinite is eliminated. According to the data preprocessing method for text emotional tendency analysis, the accuracy of text emotional tendency analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a data preprocessing method and device for text sentiment analysis. Background Art

[0002] In scenarios such as social media, e-commerce, or digital government, in order to obtain the public's comment tendencies on social phenomena, products sold on e-commerce platforms, or government policies, natural language processing is usually used to conduct sentiment analysis on social media public opinion, e-commerce product reviews, or suggestions for digital government policies.

[0003] Generally, before conducting sentiment analysis based on natural language processing, the acquired text data is preprocessed by cleaning, word segmentation, part-of-speech tagging, etc. However, due to the complexity of text components and the randomness of writing, the semantics of the text are highly diverse and flexible. Some interrogative sentences, rhetorical questions, double negative words, or words based on specific punctuation marks and sentiment words will cause natural language processing to distort the true semantics of the text, resulting in inaccurate results of the text sentiment analysis. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a data preprocessing method and device for text sentiment analysis to improve the accuracy of text sentiment analysis.

[0005] In a first aspect, the present application provides a data preprocessing method for text sentiment analysis, the method comprising:

[0006] Perform sentence segmentation processing on the target text data to obtain a first target sentence;

[0007] For each first target sentence, based on a pre-stored sentiment tendency dictionary, analyzing whether the first target sentence contains a phrase of the first target structure or a phrase of the second target structure;

[0008] A rewriting operation is performed on the first target sentence containing phrases of the first target structure, and / or a removal operation is performed on the first target sentence containing phrases of the second target structure to eliminate the situation where the words in the first target sentence are contrary to the actual semantics expressed by the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear.

[0009] According to the data preprocessing method for text sentiment tendency analysis of the present application, the target text data is subjected to sentence segmentation processing to obtain the first target sentence, and then each first target sentence is subjected to a rewriting operation for clearing the words in the first target sentence that are contrary to the actual semantics of the first target sentence or a removal operation for the expressed sentiment and evaluation that are unclear, so that before the sentiment tendency analysis is performed, the words in the first target sentence that are contrary to the actual semantics of the first target sentence or the semantics that are unclear are cleared, so as to avoid distorting the sentiment tendency actually expressed by the target sentence during the sentiment tendency analysis, thereby improving the accuracy of the text sentiment tendency analysis.

[0010] According to an embodiment of the present application, the wording of the phrases of the first target structure is opposite to the actual semantics expressed by the first target sentence; the sentiment and evaluation expressed by the phrases of the second target structure are unclear.

[0011] According to one embodiment of the present application, performing a rewriting operation on a first target sentence containing a phrase of a first target structure includes:

[0012] In the case where the first target sentence contains phrases of the first target structure, the phrases of the first target structure are rewritten to obtain a second target sentence, so that the second target sentence and the first target sentence actually express the same semantics, and the second target sentence does not contain the phrases of the first target structure.

[0013] According to one embodiment of the present application, the first target structure includes at least one of the following:

[0014] The structure of quotation marks enclosing sentiment words or evaluation words;

[0015] The structure of double negation of sentiment words or evaluation words;

[0016] A structure that contains an adverb of degree, a single negative word, and an emotional word, or an adverb of degree, a single negative word, and an evaluation word.

[0017] According to one embodiment of the present application, the second target structure includes: a structure including a single negative word and a sentiment word, or a single negative word and an evaluation word.

[0018] According to an embodiment of the present application, performing sentence segmentation processing on target text data to obtain a first target sentence includes:

[0019] Obtain target text data in a target text format;

[0020] Based on the first punctuation mark, the target text data is segmented to obtain a plurality of first sentences; wherein the first punctuation mark is a punctuation mark for pausing at the end of a sentence;

[0021] Clear the target sentence pattern in the first sentence to obtain a second sentence; wherein the target sentence pattern includes at least one of a question and a rhetorical question;

[0022] Based on the second punctuation mark, each second sentence is segmented to obtain the first target sentence; wherein the second punctuation mark is a punctuation mark indicating a pause within a sentence.

[0023] In a second aspect, the present application provides a data preprocessing device for text sentiment analysis, the device comprising:

[0024] A first processing module is used to perform sentence segmentation processing on the target text data to obtain a first target sentence;

[0025] An analysis module, for analyzing, for each first target sentence, whether the first target sentence contains a phrase of the first target structure or a phrase of the second target structure based on a pre-stored sentiment tendency dictionary library;

[0026] The second processing module is used to rewrite the first target sentence containing phrases of the first target structure, and / or remove the first target sentence containing phrases of the second target structure to eliminate the situation where the words in the first target sentence are opposite to the actual semantics expressed by the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear.

[0027] According to the data preprocessing device for text sentiment tendency analysis of the present application, sentence segmentation processing is performed on the target text data to obtain the first target sentence, and then a rewriting operation for clearing the words in the first target sentence that are opposite to the actual semantics of the first target sentence or a removal operation for clearing the emotions and evaluations that are unclear is performed on each first target sentence. Before the sentiment tendency analysis is performed, the words in the first target sentence that are opposite to the actual semantics of the first target sentence or the emotions and evaluations that are unclear are cleared, so as to avoid distorting the emotional tendency actually expressed by the target sentence during the sentiment tendency analysis, thereby improving the accuracy of the text sentiment tendency analysis.

[0028] In a third aspect, the present application provides a text sentiment tendency analysis method based on the data preprocessing method for text sentiment tendency analysis according to the first aspect, characterized in that it includes:

[0029] Based on the data preprocessing method for text sentiment analysis, the target sentence is obtained;

[0030] Conduct text sentiment analysis on the target sentence.

[0031] According to the text sentiment tendency analysis method of the present application, a target sentence is obtained by a data preprocessing method for text sentiment tendency analysis; a text sentiment tendency analysis is performed on the target sentence processed based on the data preprocessing method for text sentiment tendency analysis, so as to avoid distorting the emotional tendency truly expressed by the target sentence in the process of performing text sentiment tendency analysis on the target sentence, thereby improving the accuracy of the text sentiment tendency analysis.

[0032] In a fourth aspect, the present application provides a text sentiment tendency analysis device based on the data preprocessing method for text sentiment tendency analysis according to the first aspect, the device comprising:

[0033] An acquisition module, for acquiring a target sentence based on a data preprocessing method for text sentiment tendency analysis;

[0034] The target analysis module is used to perform text sentiment analysis on the target sentence.

[0035] According to the text sentiment tendency analysis device of the present application, a target sentence is obtained by a data preprocessing method for text sentiment tendency analysis; a text sentiment tendency analysis is performed on the target sentence processed based on the data preprocessing method for text sentiment tendency analysis, so as to avoid distorting the emotional tendency truly expressed by the target sentence in the process of performing text sentiment tendency analysis on the target sentence, thereby improving the accuracy of the text sentiment tendency analysis.

[0036] In a fifth aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data preprocessing method for text sentiment analysis as described in the first aspect above is implemented.

[0037] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0039] Figure 1 It is a flow chart of a data preprocessing method for text sentiment analysis provided by an embodiment of the present application;

[0040] Figure 2 It is one of the flow diagrams of the text sentiment tendency analysis method based on the data preprocessing method for text sentiment tendency analysis provided in the embodiment of the present application;

[0041] Figure 3 This is a second flow chart of a text sentiment tendency analysis method based on a data preprocessing method for text sentiment tendency analysis provided in an embodiment of the present application;

[0042] Figure 4 It is a structural schematic diagram of a data preprocessing device for text sentiment tendency analysis provided by an embodiment of the present application;

[0043] Figure 5 It is a structural schematic diagram of a text sentiment tendency analysis device based on a data preprocessing method for text sentiment tendency analysis provided in an embodiment of the present application;

[0044] Figure 6 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0046] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0047] In combination with the accompanying drawings, the data preprocessing method for text sentiment tendency analysis, the data preprocessing device for text sentiment tendency analysis, the text sentiment tendency analysis method, the text sentiment tendency analysis device and the electronic device provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0048] Among them, the data preprocessing method for text sentiment tendency analysis can be applied to the terminal, and can be specifically executed by hardware or software in the terminal.

[0049] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0050] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse and a joystick.

[0051] The data preprocessing method for text sentiment tendency analysis provided in the embodiment of the present application can be an electronic device or a functional module or functional entity in the electronic device that can implement the data preprocessing method for text sentiment tendency analysis. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, wearable devices, etc. The data preprocessing method for text sentiment tendency analysis provided in the embodiment of the present application is described below using the electronic device as an example of the execution subject.

[0052] like Figure 1 As shown, the data preprocessing method for text sentiment analysis includes: step 110, step 120 and step 130.

[0053] Step 110: Segment the target text data to obtain a first target sentence;

[0054] In actual implementation, the target text data can be text data on the Internet that can reflect user emotions and is collected based on big data analysis technology. The target text data includes but is not limited to buyers' comments on products on e-commerce websites, consumers' comments on food on catering websites, comments on social media, and readers' comments on news on portal websites.

[0055] In some embodiments, the target text data can be divided based on punctuation marks in the target text data, and the text data of sentences in the target text data that interfere with the real semantic recognition can be cleared to obtain the first target data, and then the first target data can be sentence segmented to obtain the first target sentence.

[0056] Step 120: for each first target sentence, based on a pre-stored sentiment tendency dictionary, analyze whether the first target sentence contains a phrase of the first target structure or a phrase of the second target structure.

[0057] In actual implementation, the sentiment tendency dictionary library may include a large number of sentiment words and the sentiment tendencies corresponding to the sentiment words. The sentiment tendency dictionary library may include positive sentiment words, negative sentiment words, positive evaluation words, negative evaluation words, degree adverbs (degree level words), double negative words, one-way negative words and advocacy words.

[0058] In some embodiments, the wording of the phrases of the first target structure is opposite to the actual semantics expressed by the first target sentence, and the sentiment and evaluation expressed by the phrases of the second target structure are unclear.

[0059] Step 130: rewrite the first target sentence containing phrases of the first target structure, and / or remove the first target sentence containing phrases of the second target structure to eliminate the situation where the words in the first target sentence are opposite to the actual semantics expressed by the first target sentence or the emotion and evaluation expressed by the first target sentence are unclear.

[0060] In some embodiments, when the first target sentence includes a phrase of the first target structure, the first target sentence may be rewritten. When the first target sentence includes a phrase of the second target structure, the first target sentence may be removed.

[0061] In some embodiments, after the rewriting operation is performed on the first target sentence, a text sentiment analysis may be performed on the target sentence obtained based on the rewriting operation.

[0062] According to the data preprocessing method for text sentiment tendency analysis in an embodiment of the present application, sentence segmentation processing is performed on the target text data to obtain the first target sentence, and then a rewriting operation for clearing the words in the first target sentence that are opposite to the actual semantics of the first target sentence or a removal operation for clearing the emotions and evaluations that are unclear is performed on each first target sentence. Before the sentiment tendency analysis is performed, the words in the first target sentence that are opposite to the actual semantics of the first target sentence or the emotions and evaluations that are unclear are cleared through target preprocessing, so as to avoid distorting the emotional tendency actually expressed by the target sentence in the process of sentiment tendency analysis, thereby improving the accuracy of text sentiment tendency analysis.

[0063] In some embodiments, the words used in the phrases of the first target structure are opposite to the actual semantics expressed by the first target sentence; the sentiment and evaluation expressed by the phrases of the second target structure are unclear.

[0064] In some embodiments, when a first target sentence contains phrases of the first target structure, a rewriting operation can be performed on the phrases of the first target structure to obtain a second target sentence, so that the second target sentence actually expresses the same semantics as the first target sentence, and the second target sentence does not contain phrases of the first target structure.

[0065] In actual execution, it is possible to search whether the first target sentence includes a phrase with a structure in which sentiment words or evaluation words are enclosed in quotation marks, or to determine whether the first target sentence includes a phrase with a structure in which sentiment words or evaluation words are double negated, or to determine whether the first target sentence includes a phrase with a structure containing an adverb of degree, a single negative word and a sentiment word, or an adverb of degree, a single negative word and an evaluation word. If the first target sentence includes a phrase with the above first target structure, the first target sentence is rewritten.

[0066] In actual execution, it is possible to search whether the first target sentence includes one-way negative words and sentiment words or the first target sentence includes one-way negative words and evaluation words. If the first target sentence includes a phrase of the above second target structure, the first target sentence is removed.

[0067] According to the data preprocessing method for text sentiment tendency analysis in an embodiment of the present application, sentence segmentation processing is performed on the target text data to obtain a first target sentence, and then a rewrite operation and a removal operation are performed on the first target sentence to eliminate words in the first target sentence that are opposite to the actual semantics of the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear. Before performing the sentiment tendency analysis, the rewrite operation and the removal operation are used to eliminate the situations in which the words in the first target sentence are opposite to the actual semantics of the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear, so as to avoid distorting the emotional tendency actually expressed by the target sentence in the process of performing the sentiment tendency analysis on the target sentence, thereby improving the accuracy of the text sentiment tendency analysis.

[0068] In some embodiments, the first target structure includes at least one of the following: a structure in which sentiment words or evaluation words are enclosed in quotation marks; a structure in which sentiment words or evaluation words are double negated; a structure containing adverbs of degree, single negations and sentiment words, or adverbs of degree, single negations and evaluation words.

[0069] In some embodiments, the sentiment tendency dictionary library may include double negative words and sentiment words or evaluation words; the first target sentence may be rewritten when the first target sentence includes double negative words and sentiment words or the first target sentence includes double negative words and evaluation words, or the first target sentence includes sentiment words or evaluation words enclosed in quotation marks.

[0070] In actual implementation, double negative words are usually composed of two "not" or a negative adverb plus the word "not", for example, "have to..., can't not..., won't not..., not not..., dare not not..., no not..., nothing but..., no not..., not no..., impossible not..., must... must", etc. The third punctuation mark can be a punctuation mark used for quotation, such as "", ''.

[0071] In some embodiments, when the first target sentence includes double negative words and sentiment words or evaluation words, the first target sentence can be rewritten to obtain the second target sentence. In some embodiments, when the first target sentence includes double negative words and sentiment words or the first target sentence includes double negative words and evaluation words, the double negative words in the first target sentence can be removed. For example, the first target sentence is "She can't help but like it", and the double negative word "can't help" in the first target sentence is removed to obtain the second target sentence "She likes it".

[0072] In some embodiments, when the first target sentence includes sentiment words or evaluation words wrapped in quotation marks, the first target sentence can be rewritten to obtain a second target sentence. In some embodiments, when the first target sentence includes sentiment words or evaluation words wrapped in quotation marks, the quotation marks can be removed, and the sentiment words or evaluation words can be replaced with sentiment words or evaluation words that are semantically opposite to the sentiment words or evaluation words. For example, if the first target sentence is "This is a big 'surprise'", remove "''" and replace "surprise" with the antonym "shock" to obtain the second target sentence "This is a big shock".

[0073] In some embodiments, the sentiment tendency dictionary library includes adverbs of degree and one-way negative words; when the first target sentence includes adverbs of degree, one-way negative words and sentiment words or the first target sentence includes adverbs of degree, one-way negative words and evaluation words, the first target sentence can be rewritten to obtain the second target sentence.

[0074] In actual implementation, adverbs of degree can include words such as "incredible, extreme, absolute, extremely, how much, and special". One-way negative words include "no, no, other, not, no, not, not, not, rest, lose, exempt, lack, prohibit, taboo, caution, prevent, no, absolutely, and difficult to" and other negative words.

[0075] In actual execution, the first target sentence can be retrieved, and when the first target sentence includes degree adverbs, one-way negative words and sentiment words or the first target sentence includes degree adverbs, one-way negative words and evaluation words, the first target sentence can be rewritten to obtain the first target sentence.

[0076] In some embodiments, when the first target sentence includes a degree adverb, a one-way negative word and a sentiment word or the first target sentence includes a degree adverb, a one-way negative word and an evaluation word, the degree adverb and the one-way negative word in the first target sentence can be removed, and the sentiment word or the evaluation word can be replaced with a sentiment word or an evaluation word that is semantically opposite to the sentiment word or the evaluation word. For example, the first target sentence "I absolutely don't like you" includes the degree adverb "absolutely", the one-way negative word "no" and the sentiment word or the evaluation word "like", remove the degree adverb "absolutely" and the one-way negative word "no", and replace the sentiment word or the evaluation word "like" with "hate", and obtain the third target sentence "I hate you".

[0077] According to the data preprocessing method for text sentiment tendency analysis in the embodiment of the present application, a first target sentence is obtained by performing sentence segmentation processing for sentence segmentation and interference removal on the target text data, and then, when the first target sentence includes a first target structural phrase, a rewrite operation is performed on the first target sentence, and then a removal operation is performed to remove the words in the first target sentence that are contrary to the actual semantics expressed by the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear, so as to avoid distorting the emotional tendency actually expressed by the target sentence in the process of performing sentiment tendency analysis on the target sentence, thereby improving the accuracy of text sentiment tendency analysis.

[0078] In some embodiments, the second target structure includes: a structure including a single negative word and a sentiment word, or a single negative word and an evaluation word.

[0079] In some embodiments, when the first target sentence includes a one-way negative word and a sentiment word or the first target sentence includes a one-way negative word and an evaluation word, a removal operation may be performed on the first target sentence.

[0080] In some embodiments, when the first target sentence includes a one-way negative word and a sentiment word or the first target sentence includes a one-way negative word and an evaluation word, a removal operation may be performed on the first target sentence. For example, the first target sentence "I don't hate you" includes a one-way negative word "no" and a sentiment word or an evaluation word "hate", and a removal operation is performed on the first target sentence.

[0081] According to the data preprocessing method for text sentiment analysis in an embodiment of the present application, a first target sentence is obtained by performing sentence segmentation processing for sentence segmentation and interference removal on the target text data, and a rewrite operation is performed on the first target sentence, and / or when the first target sentence includes a one-way negative word and a sentiment word or the first target sentence includes a one-way negative word and an evaluation word, a removal operation is performed on the first target sentence, so that before the sentiment analysis is performed on the target sentence, the rewrite operation and the removal operation are used to remove the words in the first target sentence that are opposite to the actual semantics of the first target sentence or have unclear semantics, so as to avoid distorting the sentiment actually expressed by the target sentence in the process of performing sentiment analysis on the target sentence, thereby improving the accuracy of the text sentiment analysis.

[0082] In some embodiments, target text data in a target text format is obtained; based on a first punctuation mark, the target text data is segmented to obtain a plurality of first sentences; wherein the first punctuation mark is a punctuation mark for a pause at the end of a sentence; sentences of a target sentence pattern in the first sentence are cleared to obtain a second sentence; wherein the target sentence pattern includes at least one of a question and an interrogative sentence; based on a second punctuation mark, each second sentence is segmented to obtain a first target sentence; wherein the second punctuation mark is a punctuation mark for indicating a pause within a sentence.

[0083] In actual execution, the first punctuation mark is the punctuation mark indicating the complete end of a sentence, and the second punctuation mark is the punctuation mark before the complete end of a sentence.

[0084] In actual execution, the target text format may be word, pdf, txt, excel, txt, etc. In some embodiments, the tika tool class of Apache may be used to obtain text content data in files in target text formats such as word, pdf, txt, excel, etc., and store the text content data in a txt format file named with "original file name + suffix".

[0085] In some embodiments, after obtaining the target text data in the target text format, the target text data can be segmented based on the target technology corresponding to the target text format, and the paragraph data obtained by the segmentation process can be stored in the database, and the paragraph data includes the paragraph ID, paragraph text and paragraph name. In some embodiments, for the target text data in the target text format of Word, the POI technology can be used to segment the target text data to obtain the paragraph data. For the target text data in the target text format of PDF, the PDFBox technology can be used to segment the target text data to obtain the paragraph data.

[0086] In actual execution, the first punctuation mark is a punctuation mark that indicates the complete end of a sentence, such as punctuation marks such as ".", "!", "?", etc. The target text data may be divided into the first sentence based on the first punctuation mark.

[0087] In some embodiments, after acquiring the paragraph data, the paragraph data may be segmented based on the first punctuation mark to obtain a first sentence, which may be a type of sentence data, and the sentence data may be stored in a database. The sentence data may include a sentence ID, a paragraph ID, and a sentence text.

[0088] In actual execution, the target sentence pattern may be a question sentence or a rhetorical question sentence. In some embodiments, after obtaining the first sentence, the first sentence may be traversed to retrieve the sentences in the first sentence that are question sentences or rhetorical questions, and the sentences in the first sentence that are question sentences or rhetorical questions may be deleted to obtain the second sentence.

[0089] In actual execution, the second punctuation mark is the punctuation mark before the complete end of the sentence, such as punctuation marks such as ",",",":,",";," etc. The second sentence data can be divided into the first target sentence based on the second punctuation mark.

[0090] In some embodiments, after obtaining the second sentence, the second sentence can be fine-grained sentence-processed based on the second punctuation mark to obtain a first target sentence, the first target sentence can be a fine-grained sentence data, and the first target sentence is stored in a database. The first target sentence can include a sentence ID, a fine-grained sentence ID, and a fine-grained sentence text.

[0091] According to the data preprocessing method for text sentiment tendency analysis of the embodiment of the present application, the target text data in the acquired target text format is divided into a first sentence based on a first punctuation mark, and the sentences of the target sentence pattern in the first sentence are cleared to obtain a second sentence. Based on a second punctuation mark, the second sentence data is divided into a first target sentence to remove text data in the target text data that has no obvious tendency in the emotion and evaluation dimensions, so as to converge the amount of text data and improve the efficiency of text sentiment tendency analysis.

[0092] In some embodiments, Figure 2 As shown, the present application also provides a text sentiment tendency analysis method based on the data preprocessing method for text sentiment tendency analysis, and the text sentiment tendency analysis method includes: step 210 and step 220.

[0093] Step 210: Obtain a target sentence based on a data preprocessing method for text sentiment analysis.

[0094] In actual execution, the target text data can be subjected to sentence segmentation processing for sentence segmentation and interference removal to obtain a first target sentence, and the first target sentence can be subjected to target preprocessing to remove situations in which the words in the first target sentence are contrary to the actual semantics of the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear, to obtain the target sentence.

[0095] Step 220: Perform text sentiment analysis on the target sentence.

[0096] In actual implementation, the sentiment tendency of the target sentence can be obtained based on natural language processing. For example, based on natural language processing, the sentiment tendency of the target sentence can be obtained as positive or negative.

[0097] According to the text sentiment tendency analysis method of the embodiment of the present application, according to the text sentiment tendency analysis method of the present application, a target sentence is obtained through a data preprocessing method for text sentiment tendency analysis; a text sentiment tendency analysis is performed on the target sentence processed based on the data preprocessing method for text sentiment tendency analysis, so as to avoid distorting the emotional tendency truly expressed by the target sentence in the process of performing text sentiment tendency analysis on the target sentence, thereby improving the accuracy of the text sentiment tendency analysis.

[0098] In order to better understand the text sentiment tendency analysis method based on the data preprocessing method for text sentiment tendency analysis provided in the embodiment of the present application, it is further explained below. It should be understood that the discussion below is only exemplary.

[0099] This application provides a text sentiment analysis method based on a data preprocessing method for text sentiment analysis. The specific steps can be as follows: Figure 3 As shown:

[0100] Step 310: Obtain target text data in a target text format.

[0101] In actual execution, the target text format may be word, pdf, txt, excel, txt, etc. In some embodiments, the tika tool class of Apache may be used to obtain text content data in files in target text formats such as word, pdf, txt, excel, etc., and store the text content data in a txt format file named with "original file name + suffix".

[0102] In some embodiments, after obtaining the target text data in the target text format, the target text data can be segmented based on the target technology corresponding to the target text format, and the paragraph data obtained by the segmentation process can be stored in the database, and the paragraph data includes the paragraph ID, paragraph text and paragraph name. In some embodiments, for the target text data in the target text format of Word, the POI technology can be used to segment the target text data to obtain the paragraph data. For the target text data in the target text format of PDF, the PDFBox technology can be used to segment the target text data to obtain the paragraph data.

[0103] Step 312: segment the target text data to obtain a plurality of first sentences, wherein the first punctuation mark is a punctuation mark for pausing at the end of a sentence.

[0104] In actual execution, the first punctuation mark is a punctuation mark that indicates the complete end of a sentence, such as punctuation marks such as ".", "!", "?", etc. The target text data may be divided into the first sentence based on the first punctuation mark.

[0105] In some embodiments, after acquiring the paragraph data, the paragraph data may be segmented based on the first punctuation mark to obtain a first sentence, which may be a type of sentence data, and the sentence data may be stored in a database. The sentence data may include a sentence ID, a paragraph ID, and a sentence text.

[0106] Step 314: clear the target sentence pattern in the first sentence and obtain the second sentence, wherein the target sentence pattern is a question or rhetorical question.

[0107] In actual execution, the target sentence pattern may be a question sentence or a rhetorical question sentence. In some embodiments, after obtaining the first sentence, the first sentence may be traversed to retrieve the sentences in the first sentence that are question sentences or rhetorical questions, and the sentences in the first sentence that are question sentences or rhetorical questions may be deleted to obtain the second sentence.

[0108] Step 316: Based on the second punctuation mark, each second sentence is segmented to obtain the first target sentence; wherein the second punctuation mark is a punctuation mark indicating a pause within a sentence.

[0109] In actual execution, the second punctuation mark is the punctuation mark before the complete end of the sentence, such as punctuation marks such as ",",",":,",";," etc. The second sentence data can be divided into the first target sentence based on the second punctuation mark.

[0110] In some embodiments, after obtaining the second sentence, the second sentence can be fine-grained sentence-processed based on the second punctuation mark to obtain a first target sentence, the first target sentence can be a fine-grained sentence data, and the first target sentence is stored in a database. The first target sentence can include a sentence ID, a fine-grained sentence ID, and a fine-grained sentence text.

[0111] In some embodiments, after obtaining the first target sentence, the sentences in the first target sentence that do not include sentiment and evaluation tendency words can be cleared based on a preset sentiment tendency dictionary library, where the sentiment and evaluation tendency words are words in the sentiment tendency dictionary library.

[0112] In actual implementation, the preset emotional tendency dictionary library can be shown in the following table:

[0113] Table 1

[0114]

[0115] Step 318: When the first target sentence includes double negative words and sentiment words, or the first target sentence includes double negative words and evaluation words, or the first target sentence includes sentiment words or evaluation words enclosed in quotation marks, perform a first rewrite operation on the first target sentence to obtain a second target sentence.

[0116] In actual implementation, a double negative word may be a negative word expressing a double negative meaning, usually composed of two "not"s or a negative adverb plus a "not", for example, "have to..., cannot not..., will not not..., is not not..., dare not not..., have to..., nothing but..., no not..., not no..., impossible not..., must... must not", etc. The third punctuation mark may be a punctuation mark used for quotation, such as "", ''.

[0117] In some embodiments, when the first target sentence includes double negative words and sentiment words or the first target sentence includes double negative words and evaluation words, the first target sentence can be rewritten to obtain the second target sentence. In some embodiments, when the first target sentence includes double negative words and sentiment words or evaluation words, the double negative words in the first target sentence can be removed. For example, the first target sentence is "She can't help but like it", and the double negative word "can't help" in the first target sentence is removed to obtain the second target sentence "She likes it".

[0118] In some embodiments, when the first target sentence includes sentiment words or evaluation words wrapped in quotation marks, the first target sentence can be subjected to a first rewriting operation to obtain a second target sentence. In some embodiments, when the first target sentence includes sentiment words or evaluation words wrapped in quotation marks, the quotation marks can be removed, and the sentiment words or evaluation words can be replaced with sentiment words or evaluation words that are semantically opposite to the sentiment words or evaluation words. For example, if the first target sentence is "This is a big 'surprise'", the quotation marks "''" are removed, and "surprise" is replaced with the antonym "shock" to obtain the second target sentence "This is a big shock".

[0119] Step 320: When the second target sentence includes an adverb of degree, a one-way negative word and an emotional word or the second target sentence includes an adverb of degree, a one-way negative word and an evaluation word, a second rewriting operation is performed on the second target sentence to obtain a third target sentence; wherein the one-way negative word is a negative word that expresses one-way negative semantics.

[0120] In actual implementation, degree adverbs can include words such as "incredible, extreme, absolute, extremely, how much, special", etc. One-way negative words can be negative words that express one-way negative semantics, such as "no, no, other, not, no, not, not, not, rest, loss, exemption, lack, prohibition, taboo, caution, prevention, no, absolutely not, difficult to", etc.

[0121] In actual execution, the second target sentence can be retrieved. When the second target sentence includes adverbs of degree, one-way negative words and sentiment words or the second target sentence includes adverbs of degree, one-way negative words and evaluation words, a second rewriting operation is performed on the second target sentence to obtain the third target sentence.

[0122] In some embodiments, when the second target sentence includes a degree adverb, a one-way negative word and a sentiment word or the second target sentence includes a degree adverb, a one-way negative word and an evaluation word, the degree adverb and the one-way negative word in the second target sentence can be removed, and the sentiment word or the evaluation word can be replaced with a sentiment word or an evaluation word that is semantically opposite to the sentiment word or the evaluation word. For example, the second target sentence "I absolutely don't like you" includes the degree adverb "absolutely", the one-way negative word "no" and the sentiment word or the evaluation word "like", remove the degree adverb "absolutely" and the one-way negative word "no", and replace the sentiment word or the evaluation word "like" with "hate", and obtain the third target sentence "I hate you".

[0123] Step 322: When the third target sentence includes a one-way negative word and a sentiment word or the third target sentence includes a one-way negative word and an evaluation word, a removal operation is performed on the third target sentence.

[0124] In some embodiments, when the third target sentence includes a one-way negative word and a sentiment word or the third target sentence includes a one-way negative word and an evaluation word, a removal operation may be performed on the third target sentence. For example, the third target sentence "I don't hate you" includes a one-way negative word "no" and a sentiment word or an evaluation word "hate", and a removal operation is performed on the third target sentence.

[0125] In some embodiments, based on the target logic algorithm, the first rewrite operation, the second rewrite operation and the removal operation may be sequentially performed on the first target statement. The target logic algorithm content may be as follows:

[0126]

[0127]

[0128]

[0129] In some embodiments, after obtaining the target sentence, the target sentence can be reintegrated based on the fine-grained sentence ID and sentence ID corresponding to the target sentence, and the reintegrated target sentence is stored in a txt file named "original file name + suffix + 'processed'", and then the target sentence is pre-processed by performing word segmentation, removing stop words, removing repeated words and other pre-processing operations.

[0130] The embodiment of the present application also provides a data preprocessing device for text sentiment tendency analysis.

[0131] like Figure 4 As shown, the data preprocessing device 400 for text sentiment analysis includes: a first processing module 410, an analysis module 420 and a second processing module 430.

[0132] The first processing module 410 is used to perform sentence segmentation processing on the target text data to obtain a first target sentence;

[0133] An analysis module 420 is used to analyze, for each first target sentence, whether the first target sentence contains a phrase of the first target structure or a phrase of the second target structure based on a pre-stored sentiment tendency dictionary library;

[0134] The second processing module 430 is used to rewrite the first target sentence containing phrases of the first target structure, and / or remove the first target sentence containing phrases of the second target structure to eliminate the situation where the words in the first target sentence are opposite to the actual semantics expressed by the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear.

[0135] According to the data preprocessing device for text sentiment tendency analysis of the present application, sentence segmentation processing is performed on the target text data to obtain the first target sentence, and then a rewriting operation for clearing the words in the first target sentence that are opposite to the actual semantics of the first target sentence or a removal operation for clearing the emotions and evaluations that are unclear is performed on each first target sentence. Before the sentiment tendency analysis is performed, the words in the first target sentence that are opposite to the actual semantics of the first target sentence or the emotions and evaluations that are unclear are cleared, so as to avoid distorting the emotional tendency actually expressed by the target sentence during the sentiment tendency analysis, thereby improving the accuracy of the text sentiment tendency analysis.

[0136] In some embodiments, the wording of the phrases of the first target structure is opposite to the actual semantics of the first target sentence; the sentiment and evaluation expressed by the phrases of the second target structure are unclear.

[0137] In some embodiments, the second processing module 430 includes:

[0138] A rewriting unit is used to rewrite the phrases of the first target structure when the first target sentence contains the phrases of the first target structure, and obtain a second target sentence, so that the second target sentence has the same actual expression semantics as the first target sentence, and the second target sentence does not contain the phrases of the first target structure.

[0139] In some embodiments, the first target structure includes at least one of the following:

[0140] The structure of quotation marks enclosing sentiment words or evaluation words;

[0141] The structure of double negation of sentiment words or evaluation words;

[0142] A structure that contains an adverb of degree, a single negative word, and an emotional word, or an adverb of degree, a single negative word, and an evaluation word.

[0143] In some embodiments, the second target structure includes: a structure including a single negative word and a sentiment word, or a single negative word and an evaluation word.

[0144] In some embodiments, the first processing module 410 includes:

[0145] An acquisition unit, used for acquiring target text data in a target text format;

[0146] A first segmentation unit is used to segment the target text data based on the first punctuation mark to obtain a plurality of first sentences; wherein the first punctuation mark is a punctuation mark for pausing at the end of a sentence;

[0147] A clearing unit, used to clear the target sentence pattern in the first sentence and obtain the second sentence; wherein the target sentence pattern includes at least one of a question and a rhetorical question;

[0148] The second segmentation unit is used to segment each second sentence based on the second punctuation mark to obtain the first target sentence; wherein the second punctuation mark is a punctuation mark indicating a pause within a sentence.

[0149] The embodiment of the present application also provides a text sentiment tendency analysis device based on the data preprocessing method for text sentiment tendency analysis.

[0150] like Figure 5 As shown, the text sentiment tendency analysis device 500 includes: an acquisition module 510 and a target analysis module 520.

[0151] An acquisition module 510 is used to acquire a target sentence based on a data preprocessing method for text sentiment analysis;

[0152] The target analysis module 520 is used to perform text sentiment analysis on the target sentence.

[0153] According to the text sentiment tendency analysis device of the embodiment of the present application, a target sentence is obtained by a data preprocessing method for text sentiment tendency analysis; a text sentiment tendency analysis is performed on the target sentence processed based on the data preprocessing method for text sentiment tendency analysis, so as to avoid distorting the emotional tendency truly expressed by the target sentence in the process of performing text sentiment tendency analysis on the target sentence, thereby improving the accuracy of the text sentiment tendency analysis.

[0154] The data preprocessing device and text sentiment tendency analysis device for text sentiment tendency analysis in the embodiment of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0155] The data preprocessing device and text sentiment tendency analysis device for text sentiment tendency analysis in the embodiment of the present application may be a device having an operating system. The operating system may be a Microsoft (Windows) operating system, an Android (Android) operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0156] The data preprocessing device 400 for text sentiment analysis provided in the embodiment of the present application can achieve Figures 1 to 3 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0157] In some embodiments, Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, each process of the above-mentioned data preprocessing method embodiment for text sentiment tendency analysis is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0158] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0159] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned data preprocessing method embodiment for text sentiment tendency analysis are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0160] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0161] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned data preprocessing method for text sentiment tendency analysis.

[0162] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0163] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned data preprocessing method embodiment for text sentiment tendency analysis, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0164] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0165] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0166] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0167] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0168] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0169] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A data preprocessing method for text sentiment analysis, characterized in that: include: Perform sentence segmentation processing on the target text data to obtain a first target sentence; For each of the first target sentences, based on a pre-stored sentiment tendency dictionary, analyzing whether the first target sentence contains a phrase of a first target structure or a phrase of a second target structure; A rewriting operation is performed on the first target sentence containing phrases of the first target structure, and / or a removal operation is performed on the first target sentence containing phrases of the second target structure, so as to eliminate the situation in which the words in the first target sentence are contrary to the actual semantics expressed by the first target sentence or the emotion and evaluation expressed by the first target sentence are unclear.

2. The data preprocessing method for text sentiment analysis according to claim 1, characterized in that: The words used in the phrases of the first target structure are opposite to the actual semantics expressed by the first target sentence; the emotions and evaluations expressed by the phrases of the second target structure are unclear.

3. The data preprocessing method for text sentiment analysis according to claim 1, characterized in that: The rewriting operation on the first target sentence containing the phrase of the first target structure includes: In the case where the first target sentence contains phrases of the first target structure, a rewriting operation is performed on the phrases of the first target structure to obtain a second target sentence, so that the second target sentence and the first target sentence actually express the same semantics, and the second target sentence does not contain the phrases of the first target structure.

4. The data preprocessing method for text sentiment analysis according to claim 2, characterized in that: The first target structure includes at least one of the following: The structure of quotation marks enclosing sentiment words or evaluation words; The structure of double negation of sentiment words or evaluation words; A structure that contains an adverb of degree, a single negative word, and an emotional word, or an adverb of degree, a single negative word, and an evaluation word.

5. The data preprocessing method for text sentiment analysis according to claim 2, characterized in that: The second target structure includes: a structure including a single negative word and a sentiment word, or a structure including a single negative word and an evaluation word.

6. The data preprocessing method for text sentiment analysis according to any one of claims 1 to 5, characterized in that: The step of performing sentence segmentation processing on the target text data to obtain a first target sentence includes: Obtain target text data in a target text format; Based on the first punctuation mark, the target text data is segmented to obtain a plurality of first sentences; wherein the first punctuation mark is a punctuation mark for pausing at the end of a sentence; Clear the target sentence pattern in the first sentence to obtain a second sentence; wherein the target sentence pattern includes at least one of a question and a rhetorical question; Based on the second punctuation mark, each of the second sentences is segmented to obtain the first target sentence; wherein the second punctuation mark is a punctuation mark indicating a pause within a sentence.

7. A data preprocessing device for text sentiment analysis, characterized in that: include: A first processing module is used to perform sentence segmentation processing on the target text data to obtain a first target sentence; An analysis module, configured to analyze, for each of the first target sentences, whether the first target sentence contains a phrase of a first target structure or a phrase of a second target structure based on a pre-stored sentiment tendency dictionary; The second processing module is used to rewrite the first target sentence containing phrases of the first target structure, and / or remove the first target sentence containing phrases of the second target structure to eliminate the situation where the words in the first target sentence are contrary to the actual semantics of the first target sentence or the emotions and evaluations expressed by the first target sentence are unclear.

8. A text sentiment tendency analysis method based on the data preprocessing method for text sentiment tendency analysis as described in any one of claims 1 to 6, characterized in that: include: Based on the data preprocessing method for text sentiment analysis, the target sentence is obtained; A text sentiment analysis is performed on the target sentence.

9. A text sentiment tendency analysis device based on the data preprocessing method for text sentiment tendency analysis according to any one of claims 1 to 6, characterized in that: include: An acquisition module, for acquiring a target sentence based on a data preprocessing method for text sentiment tendency analysis; The target analysis module is used to perform text sentiment analysis on the target sentence.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the data preprocessing method for text sentiment analysis as described in any one of claims 1-6 is implemented.