Method and apparatus for comment sentiment analysis
By preprocessing and similarity comparison of comment texts, combined with a sentiment classification model, the target attribute labels and sentiment polarity are identified and determined. This solves the problems of insufficient efficiency and accuracy in existing methods, and achieves efficient identification and improved accuracy in fine-grained sentiment analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-06-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing sentiment analysis methods for comments are inefficient and inaccurate, especially when dealing with fine-grained sentiment analysis, where they struggle to effectively identify implicit attribute words and handle large-scale attributes.
By preprocessing the comment text, using a preset attribute representation matrix for similarity comparison, candidate attribute representations are determined, and the target attribute label and sentiment polarity, including negative and non-negative sentiments, are determined through a sentiment classification model.
It improves the efficiency and accuracy of sentiment analysis for comments, can identify implicit attribute words, reduce cascading errors, and supports flexible expansion of attributes.
Smart Images

Figure CN115129873B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fine-grained sentiment analysis technology, and in particular to a comment sentiment analysis method and related apparatus. Background Technology
[0002] Currently, an increasing number of users are sharing their experiences with certain products on social media. Effective analysis and integration of this public opinion data are crucial for the stable operation of products. Unlike overall sentiment analysis, attribute-based sentiment analysis offers finer granularity. By analyzing comments to identify the attributes mentioned and their positive or negative evaluations, it helps understand user groups' preferences for various product attributes, allowing for targeted improvements. Existing methods mostly utilize machine learning or deep learning for analysis, but their efficiency and accuracy need improvement. Summary of the Invention
[0003] In view of this, this application provides a method and related apparatus for comment sentiment analysis, which can improve the efficiency and accuracy of comment sentiment analysis.
[0004] In a first aspect, embodiments of this application provide a comment sentiment analysis method, the method comprising:
[0005] The comment text undergoes a first preprocessing step to obtain the original comment representation, which is used to indicate the user's evaluation of the target product;
[0006] The original comment representation is compared with the preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations.
[0007] A second preprocessing is performed on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation;
[0008] The target comment representation is input into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
[0009] Secondly, embodiments of this application provide a comment sentiment analysis device, the device comprising:
[0010] The first preprocessing unit is used to perform a first preprocessing on the comment text to obtain the original comment representation, wherein the comment text is used to indicate the user's evaluation of the target product;
[0011] The comparison unit is used to compare the original comment representation with the preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations.
[0012] The second preprocessing unit is used to perform a second preprocessing on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation.
[0013] The sentiment analysis unit is used to input the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
[0014] Thirdly, embodiments of this application provide an electronic device, including a processor, a communication module, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any method of the first aspect of this application.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in any method of the first aspect of this application.
[0016] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of this application. The computer program product may be a software installation package.
[0017] As can be seen, the above-described comment sentiment analysis method and related apparatus firstly preprocess the comment text to obtain an original comment representation, which indicates the user's evaluation of the target product. Then, the original comment representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. Next, the comment text and the candidate attribute texts corresponding one-to-one with the at least one candidate attribute representation are subjected to a second preprocessing to obtain the target comment representation. Finally, the target comment representation is input into a sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, where the target sentiment polarity includes negative and non-negative. This method can improve the efficiency and accuracy of sentiment analysis of comments. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A system architecture diagram of a comment sentiment analysis method provided in this application embodiment;
[0020] Figure 2 A flowchart illustrating a comment sentiment analysis method provided in an embodiment of this application;
[0021] Figure 3 A flowchart illustrating another comment sentiment analysis method provided in this application embodiment;
[0022] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0023] Figure 5 A functional unit block diagram of a comment sentiment analysis device provided in this application embodiment;
[0024] Figure 6 A block diagram of the functional units of another comment sentiment analysis device provided in an embodiment of this application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0026] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0027] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.
[0028] In this application, the term "connection" refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices. This application does not impose any limitations on this.
[0029] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0030] The background technology and related terms of this application are explained below.
[0031] Background technology related:
[0032] Fine-grained sentiment analysis (ABSA) traditionally divides the task into two subtasks: attribute identification, which mines the attributes involved in the sentence; and attribute sentiment analysis, which analyzes each attribute to identify the sentiment polarity they express. Machine learning or deep learning methods are commonly used, but these traditional methods have high system complexity, a high risk of cascading errors throughout the process, and some attributes may be unmentioned or implicitly mentioned, making it impossible to directly extract attribute words.
[0033] One approach involves converting each attribute into a reading comprehension question, such as "How would you rate the screen display effect?" This is then converted into question-and-answer pairs, and named entity recognition models, such as MRC models, are used to identify the positive or negative sentiment of the answer in the question-and-answer pair. However, this method is less efficient and requires multiple calculations to obtain the result when there are many attributes.
[0034] Based on the above-mentioned technical problems, this application provides a comment sentiment analysis method and related apparatus, which can identify implicit attribute words, reduce cascading errors, process a large number of attributes, support flexible attribute expansion, and has high efficiency and accuracy.
[0035] Let's combine the following... Figure 1 The system architecture of a comment sentiment analysis method according to an embodiment of this application is described below. Figure 1 The present application provides a system architecture for a comment sentiment analysis method. The system architecture 100 includes a server 110, which is equipped with a matching model 111 and a sentiment classification model 112. The server 110 receives comment texts for target products from various channels, then uses the matching model 111 to determine the candidate attributes most likely related to the comment text, and finally uses the sentiment classification model 112 to determine the candidate attributes that are truly related to the comment text as the target attributes, and determines the sentiment polarity of the target attributes.
[0036] In one possible embodiment, the server 110 may also be equipped with a validity classification model 113, which is used to perform validity analysis on the received comment text and filter out invalid comment text.
[0037] Among them, the matching model 111, sentiment classification model 112, and validity classification model 113 can all be trained using training data, which is labeled training text. Labeling can be performed from at least the following dimensions:
[0038] Comment validity: Whether the entire comment text is a valid comment.
[0039] Evaluation attributes: words or phrases in the text that relate to the attributes.
[0040] Evaluation words: Evaluation words or phrases for attributes.
[0041] Attribute pairs: Pair evaluation attributes with evaluation terms for labeling; if an evaluation attribute does not appear, it can be labeled as null.
[0042] Attribute category tags: Based on entity pairs and the tag system, predefined multi-level tag mapping is performed.
[0043] Sentiment polarity: Label each attribute pair with negative and non-negative sentiment polarity.
[0044] It is understandable that the target product has preset attribute category labels. These preset attribute category labels, as well as the evaluation attributes and evaluation words under each attribute category label, can form a preset attribute representation matrix, which is built into the matching model 111 mentioned above.
[0045] For example, a complete annotation example is as follows:
[0046]
[0047] in, <asp-1> and< / asp-1> The "run" attribute between the annotations is an evaluation attribute. <term-1> and< / term-1> The word "fast" between the annotations is an evaluation term. <term-2> and< / term-2> The word "expensive" between the annotations is an evaluation term, and its corresponding evaluation attribute is empty. 1 indicates that the training text is a valid comment, 0 indicates that the sentiment polarity is positive, and -1 indicates that the sentiment polarity is negative.
[0048] It is evident that training data labeled with this subdivision dimension can make the outputs of matching model 111, sentiment classification model 112, and validity classification model 113 more accurate, thereby improving the efficiency and accuracy of sentiment analysis of comments.
[0049] The following is combined Figure 2 This application describes a comment sentiment analysis method as described in its embodiments. Figure 2 A flowchart illustrating a sentiment analysis method for comments provided in this application embodiment specifically includes the following steps:
[0050] Step 201: Perform the first preprocessing on the comment text to obtain the original comment representation.
[0051] In this embodiment, the meaning of "representation" is the term "representation" in the field of deep learning. The original comment representation can be understood as the semantic representation of the comment text, which will not be elaborated here.
[0052] The comment text is used to indicate the user's evaluation of the target product, which can be electronic products or other products that users can comment on on the online platform.
[0053] The comment text can be encoded to obtain a comment text vector, and then the comment text vector can be averaged to obtain the original comment representation.
[0054] In one possible embodiment, the comment text vector can be input into a BERT-like encoder, and the output can be averaged and then aggregated to obtain the original comment representation. The BERT-like encoder can be a BERT model (Bidirectional Encoder Representations from Transformer), etc., which will not be elaborated here.
[0055] It is evident that by performing the first preprocessing on the comment text to obtain the original comment representation, a reliable reference can be provided for subsequent similarity comparisons.
[0056] Step 202: Compare the original comment representation with the preset attribute representation matrix to determine at least one candidate attribute representation.
[0057] The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product, and the at least one candidate attribute representation is a subset of the multiple preset attribute representations.
[0058] To facilitate understanding, the composition of the preset attribute representation matrix is explained below. First, the preset description text of the multi-level classification labels corresponding to the target product is obtained. Then, at least one evaluation object text and at least one evaluation word text corresponding to the evaluation object text are obtained from the training data under each of the multi-level classification labels. The at least one evaluation object text is determined by the evaluation object labels in the training data, and the at least one evaluation word text corresponding to the evaluation object text is determined by the evaluation word labels in the training data. Finally, the evaluation object text, the evaluation word text, and the preset description text under each multi-level classification label are concatenated and then encoded and average pooled to obtain the preset attribute representation matrix.
[0059] It is evident that by selecting descriptive information from the training data and concatenating it into the preset descriptive text, the representation of attributes can be made more complete, thereby improving the accuracy of similarity comparison between the original comment text and the preset attribute representation matrix.
[0060] Specifically, the original comment representation can be compared with each of the preset attribute representations in the preset attribute representation matrix to determine a first similarity sequence, which is sorted from high to low. Then, at least one preset attribute representation is selected as the at least one candidate attribute representation based on the first similarity sequence.
[0061] In one possible embodiment, a matching model can be constructed, which may adopt a dual-tower structure. The matching model includes a comment encoder. After the comment text vector is input into the comment encoder, the original comment representation is obtained. Then, the original comment representation is compared with the preset attribute representation matrix stored offline in the matching model to obtain the similarity distance between the original comment representation and each preset attribute representation. Then, the similarity distances are sorted in descending order to obtain the first similarity sequence. Finally, the top-ranked preset attribute representations are selected as candidate attribute representations.
[0062] The matching model can be trained using a contrastive training approach, where each training set consists of one positive attribute and N randomly selected negative attributes. This minimizes the similarity distance with positive attributes and maximizes the similarity distance with negative attributes. Furthermore, to ensure efficiency, the training loss function can be Margin Ranking Loss.
[0063]
[0064] Where N is the number of negative example attributes sampled, S + To determine the similarity to the attributes of positive examples, The similarity to the attributes of the i-th negative example is not discussed further here.
[0065] It should be noted that the number of candidate attribute representations can be set according to the needs. The more the number is selected, the more accurate the recall rate will be, but the subsequent computational load will be greater. For example, in the mobile phone field, there are generally 200 predefined attributes. At this time, selecting the top 10 candidate attribute features can ensure a high similarity while keeping the computational load of subsequent sentiment analysis within an ideal range.
[0066] It is evident that by comparing the original comment representation with the preset attribute representation matrix to determine at least one candidate attribute representation, the approximate range of attributes related to the comment text can be selected in advance, and then targeted sentiment analysis can be performed, which can improve the efficiency of sentiment analysis of comments.
[0067] Step 203: Perform a second preprocessing on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation.
[0068] Specifically, the comment text and the candidate attribute text corresponding to the at least one candidate attribute representation can be concatenated and encoded to obtain the target comment feature.
[0069] In one possible embodiment, all candidate attribute representations can first be converted into text format, i.e., candidate attribute text, and then both the comment text and the candidate attribute text can be converted into vector form and concatenated to obtain the target comment representation.
[0070] Step 204: Input the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text.
[0071] The target emotional polarity includes negative and non-negative. The above emotional classification model can be a classification model that performs a three-classification task, with classification labels of {not mentioned, non-negative, negative}.
[0072] Specifically, the target comment representation can be input into the aforementioned sentiment classification model to classify each candidate attribute text. Candidate attribute representations not mentioned in the comment text are labeled as "not mentioned," while candidate attribute representations mentioned in the comment text are identified as target attribute representations. The target attribute label and target sentiment polarity corresponding to the target attribute representation are then determined. For example, when the comment text is "Battery charges quickly but easily gets hot," one target attribute label is ultimately determined as "Battery|Charging speed," with a corresponding target sentiment polarity of "non-negative," and the other target attribute label is "Battery|Getting hot," with a corresponding target sentiment polarity of "negative." This will not be elaborated further here.
[0073] It should be noted that the sentiment classification model can be trained in the following way: First, determine the second similarity sequence between the training text representation of the training data and each of the preset attribute representations in the preset attribute representation matrix. The second similarity sequence is sorted from low to high. Then, determine at least one negative example attribute representation based on the second similarity sequence, and add the negative example attribute text corresponding one-to-one with the at least one negative example attribute representation to the training data. Finally, train the preset classification model with the training data containing the negative example attribute text to obtain the sentiment classification model.
[0074] In one possible embodiment, since the training data only contains the sentiment polarity labels of preset attributes, it is also necessary to dynamically generate unmentioned attributes for each training sample. This can be achieved by randomly sampling from the preset attribute set that is not mentioned in the training text to form attribute negative sample samples, which are then added to the training data to improve the output accuracy of the sentiment classification model.
[0075] In one possible embodiment, attribute matching calculation can also be performed on each training sample. That is, the second similarity sequence between each training sample and each preset attribute representation in the preset attribute representation matrix is determined, the top few preset attribute representations with the worst similarity are selected, and the preset attributes corresponding to the selected top few preset attribute representations with the worst similarity are determined as attribute negative sample samples. The attribute negative sample samples are added to the training data to improve the output accuracy of the sentiment classification model. At the same time, since the similarity is calculated, the problem that when the preset attribute set is large, most of the selected attributes are not related to the training sample, thereby reducing the classification difficulty and resulting in poor training effect can be solved.
[0076] In one possible embodiment, the sentiment classification model can be a BERT model. The specific training process can employ a two-stage classification method, where the training comment text and the training attribute description are used as inputs to two BERT segments, segment 1 and segment 2, respectively. Cross-attention is then used for sufficient interactive training. The training loss function typically uses cross-entropy loss, Focal loss, or CrossEntropy loss with weight parameters, etc., without specific limitations.
[0077] It is evident that by inputting the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, the efficiency and accuracy of sentiment analysis of comment text can be greatly improved.
[0078] The following is combined Figure 3 Another comment sentiment analysis method provided in the embodiments of this application will be described. Figure 3 A flowchart illustrating another comment sentiment analysis method provided in this application embodiment is shown, specifically including the following steps:
[0079] Step 301: Input the comment text into the validity classification model to determine the validity of the comment text.
[0080] The validity classification model is a model trained using training data, which includes validity labels. When the comment text is invalid, the comment text is deleted; when the comment text is valid, step 302 is executed.
[0081] Specifically, the validity classification model can be obtained by training a pre-defined binary classification model using training data. Since the training data is labeled with the validity dimension, it can be used to train the pre-defined binary classification model, which will not be elaborated here.
[0082] It is evident that inputting the comment text into the validity classification model to determine its validity can eliminate a large number of noisy comments, save computational resources, and improve the processing efficiency of a single comment text.
[0083] Step 302: Perform the first preprocessing on the comment text to obtain the original comment representation.
[0084] Step 303: Compare the original comment representation with the preset attribute representation matrix to determine at least one candidate attribute representation.
[0085] Step 304: Perform a second preprocessing on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation.
[0086] Step 305: Input the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text.
[0087] As can be seen, the above-described sentiment analysis method firstly preprocesses the comment text to obtain the original comment representation, which indicates the user's evaluation of the target product. Then, the original comment representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. Next, a second preprocessing is performed on the comment text and the candidate attribute texts corresponding one-to-one with the at least one candidate attribute representation to obtain the target comment representation. Finally, the target comment representation is input into a sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, where the target sentiment polarity includes negative and non-negative. This method can improve the efficiency and accuracy of sentiment analysis of comments.
[0088] For steps not detailed above, please refer to Figure 2 The descriptions of all or part of the methods are omitted here.
[0089] The following is combined Figure 4 An electronic device according to an embodiment of this application will be described. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 4As shown, the electronic device 400 includes a processor 401, a communication interface 402, and a memory 403, which are interconnected. The electronic device 400 may also include a bus 404, through which the processor 401, communication interface 402, and memory 403 are interconnected. The bus 404 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 404 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 4 The bus is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The memory 403 is used to store a computer program, which includes program instructions. The processor is configured to call the program instructions and execute the above-mentioned... Figure 2 , Figure 4 All or part of the methods described herein.
[0090] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the electronic device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0091] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0092] When dividing functional modules according to their respective functions, the following is combined with... Figure 5 This application provides a detailed description of a comment sentiment analysis device as described in its embodiments. Figure 5A functional unit block diagram of a comment sentiment analysis device 500 provided in this application embodiment, the comment sentiment analysis device 500 includes:
[0093] The first preprocessing unit 510 is used to perform a first preprocessing on the comment text to obtain the original comment representation, wherein the comment text is used to indicate the user's evaluation of the target product;
[0094] The comparison unit 520 is used to compare the original comment representation with the preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations.
[0095] The second preprocessing unit 530 is used to perform a second preprocessing on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation.
[0096] The sentiment analysis unit 540 is used to input the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
[0097] As can be seen, the above-described comment sentiment analysis method and related apparatus firstly preprocess the comment text to obtain an original comment representation, which indicates the user's evaluation of the target product. Then, the original comment representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. Next, the comment text and the candidate attribute texts corresponding one-to-one with the at least one candidate attribute representation are subjected to a second preprocessing to obtain the target comment representation. Finally, the target comment representation is input into a sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, where the target sentiment polarity includes negative and non-negative. This method can improve the efficiency and accuracy of sentiment analysis of comments.
[0098] When using integrated units, the following is combined with Figure 6Another comment sentiment analysis device 600 in the embodiments of this application will be described in detail. The comment sentiment analysis device 600 includes a processing unit 601 and a communication unit 602. The processing unit 601 is used to perform any step as described in the above method embodiments, and when performing data transmission such as sending, the communication unit 602 can be selectively invoked to complete the corresponding operation.
[0099] The comment sentiment analysis device 600 may further include a storage unit 603 for storing program code and data. The processing unit 601 may be a processor, the communication unit 602 may be a wireless communication module, and the storage unit 603 may be a memory.
[0100] The processing unit 601 is specifically used for:
[0101] The comment text undergoes a first preprocessing step to obtain the original comment representation, which is used to indicate the user's evaluation of the target product;
[0102] The original comment representation is compared with the preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations.
[0103] A second preprocessing is performed on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation;
[0104] The target comment representation is input into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
[0105] As can be seen, the above-described comment sentiment analysis method and related apparatus firstly preprocess the comment text to obtain an original comment representation, which indicates the user's evaluation of the target product. Then, the original comment representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. Next, the comment text and the candidate attribute texts corresponding one-to-one with the at least one candidate attribute representation are subjected to a second preprocessing to obtain the target comment representation. Finally, the target comment representation is input into a sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, where the target sentiment polarity includes negative and non-negative. This method can improve the efficiency and accuracy of sentiment analysis of comments.
[0106] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments.
[0107] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.
[0108] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0109] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0110] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0111] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0112] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0113] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0114] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0115] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for sentiment analysis of comments, characterized in that, The method includes: The comment text undergoes a first preprocessing step to obtain the original comment representation, which is used to indicate the user's evaluation of the target product; The original review representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. The preset attribute representation matrix is determined according to the following steps: obtaining the preset description text of the multi-level classification labels; obtaining at least one evaluation object text and at least one evaluation word text corresponding to the evaluation object text under each multi-level classification label in the training data, wherein the at least one evaluation object text is determined by the evaluation object label in the training data, and the at least one evaluation word text corresponding to the evaluation object text is determined by the evaluation word label in the training data; concatenating the evaluation object text, the evaluation word text, and the preset description text under each multi-level classification label, and performing encoding and average pooling processing to obtain the preset attribute representation matrix. A second preprocessing is performed on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation, wherein the comment text and the candidate attribute text that corresponds to the at least one candidate attribute representation are concatenated and encoded to obtain the target comment feature; The target comment representation is input into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
2. The method according to claim 1, characterized in that, The first preprocessing of the comment text to obtain the original comment representation includes: The comment text is encoded to obtain a comment text vector; The original comment representation is obtained by performing average pooling on the comment text vector.
3. The method according to claim 1, characterized in that, The original comment representation is compared with a preset attribute representation matrix to determine at least one candidate attribute representation, including: The original comment representation is compared with each of the preset attribute representations in the preset attribute representation matrix to determine a first similarity sequence, which is sorted from high to low. At least one preset attribute representation is selected as the at least one candidate attribute representation based on the first similarity sequence.
4. The method according to claim 1, characterized in that, The method further includes: Determine a second similarity sequence between the training text representation of the training data and each of the preset attribute representations in the preset attribute representation matrix, wherein the second similarity sequence is sorted from low to high; At least one negative example attribute representation is determined based on the second similarity sequence, and the negative example attribute text corresponding one-to-one with the at least one negative example attribute representation is added to the training data; The sentiment classification model is obtained by training the preset classification model with the addition of the negative example attribute text.
5. The method according to claim 1, characterized in that, Before performing the first preprocessing on the comment text to obtain the original comment representation, the method further includes: The comment text is input into a validity classification model to determine the validity of the comment text. The validity classification model is a model trained with training data, which includes validity labels. If the comment text is invalid, delete the comment text. When the comment text is valid, the step of performing the first preprocessing on the comment text to obtain the comment representation is performed.
6. A sentiment analysis device for comments, characterized in that, The device includes: The first preprocessing unit is used to perform a first preprocessing on the comment text to obtain the original comment representation, wherein the comment text is used to indicate the user's evaluation of the target product; A comparison unit is used to perform a similarity comparison between the original review representation and a preset attribute representation matrix to determine at least one candidate attribute representation. The preset attribute representation matrix includes multiple preset attribute representations, which are representations corresponding to the multi-level classification labels of the target product. The at least one candidate attribute representation is a subset of the multiple preset attribute representations. The preset attribute representation matrix is determined according to the following steps: obtaining the preset description text of the multi-level classification labels; obtaining at least one evaluation object text and at least one evaluation word text corresponding to the evaluation object text under each multi-level classification label in the training data, wherein the at least one evaluation object text is determined by the evaluation object label in the training data, and the at least one evaluation word text corresponding to the evaluation object text is determined by the evaluation word label in the training data; concatenating the evaluation object text, the evaluation word text, and the preset description text under each multi-level classification label, and performing encoding and average pooling processing to obtain the preset attribute representation matrix. The second preprocessing unit is used to perform a second preprocessing on the comment text and the candidate attribute text that corresponds one-to-one with the at least one candidate attribute representation to obtain the target comment representation, wherein the comment text and the candidate attribute text that corresponds to the at least one candidate attribute representation are concatenated and encoded to obtain the target comment feature. The sentiment analysis unit is used to input the target comment representation into the sentiment classification model to determine the target attribute label and target sentiment polarity corresponding to the comment text, wherein the target sentiment polarity includes negative and non-negative.
7. An electronic device, characterized in that, Includes a processor, the processor being configured to execute instructions for the steps of the method as described in any one of claims 1-5.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer program, the computer program including program instructions, which, when executed by the baseband chip, cause the baseband chip to perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Hotel comment text attribute description extraction method
CN110750646A
Sales prediction method based on product comment viewpoint mining
CN111242679A