A cross-domain fine-grained sentiment analysis method, device and storage medium
By integrating grammar knowledge and common sense knowledge in cross-domain fine-grained sentiment analysis, using BERT encoder and graph convolution network GCN, a cross-domain fine-grained sentiment analysis model was constructed, which solved the problem of domain differences and improved the prediction effect of domain adaptability and sentiment analysis.
Patent Information
- Application Number
- CN202210660427.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-13
AI Technical Summary
The prior art faces domain differences in cross-domain fine-grained sentiment analysis, resulting in migration errors and it is difficult to effectively use unlabeled data for domain adaptation.
The pre-trained grammar knowledge feature vector representation module and the pre-trained common sense knowledge relation vector representation module are used to construct a cross-domain fine-grained sentiment analysis model. Through the integration of grammatical knowledge and common sense knowledge, the data differences in the source and target fields can be narrowed to achieve domain adaptation.
The prediction effect of model migration from source domain to target domain is improved and sentiment analysis is predicted, migration errors are reduced, and adaptability to label-free data in the target domain is enhanced.
Smart Images

Figure CN115221272B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a cross-domain fine-grained sentiment analysis method, device and storage medium. Background Art
[0002] Fine-grained sentiment analysis aims to extract aspect terms and identify the sentiment polarity of each extracted aspect term. Different from the coarse-grained analysis based on the entire sentence or text, fine-grained analysis specifically analyzes the attributes and functions in the sentence to determine whether the sentiment expressed is positive, negative, or neutral. However, a large number of supervised models now rely heavily on large-scale labeled data. However, due to the multiplicity of human emotions, the labeling process of labeled data for fine-grained sentiment analysis is very time-consuming and labor-intensive.
[0003] To alleviate the reliance on large-scale labeled data, cross-domain fine-grained sentiment analysis, which aims to transfer knowledge from rich labeled source domains to unlabeled target domains, becomes a promising direction. The main challenge of cross-domain sentiment analysis is the domain difference, since aspect terms in different domains are almost disjoint.
[0004] Faced with this challenge, existing research on sentiment analysis domain adaptation mainly focuses on coarse-grained domain adaptation. Since fine-grained is based on the word level, the challenges faced will be more arduous. There are not many explorations of fine-grained sentiment analysis across domains, but currently syntactic information is regarded as a bridge for domain adaptation and has achieved remarkable performance. By using domain-shared grammatical features to construct auxiliary tasks or adversarial training to obtain domain-invariant features, a domain-shared feature space is constructed to achieve domain adaptation. And with the emergence of pre-trained language models, traditional neural network models using pre-trained models have also achieved good results. Current methods focus on extracting domain-shared syntactic information to bridge the gap between domains. However, the transferable syntactic knowledge is complex and diverse, which leads to the problem of migration errors in domain adaptation. Summary of the invention
[0005] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a cross-domain fine-grained sentiment analysis method, device and storage medium.
[0006] The technical solution adopted by the present invention is:
[0007] A cross-domain fine-grained sentiment analysis method,
[0008] The following steps are involved:
[0009] Constructing a fine-grained sentiment analysis model for the target domain, the fine-grained sentiment analysis model includes a pre-trained grammatical knowledge feature vector representation module, a pre-trained common sense knowledge relationship vector representation module, and a classifier; the pre-trained grammatical knowledge feature vector representation module includes a BERT encoder, and the pre-trained common sense knowledge relationship vector representation module includes a graph convolutional network GCN;
[0010] Inputting the unlabeled texts from the source domain and the target domain into the BERT encoder to obtain a grammatical knowledge feature vector representation of each word in the text;
[0011] Input the text of the source domain or the target domain into the pre-trained common sense knowledge relationship vector representation module; wherein the words of the specified part of speech of each sentence in each domain are input into the ConceptNet common sense knowledge base, and the path and relationship to the domain concept are obtained, so as to construct a domain common sense graph; input the domain common sense graph into the graph convolutional network GCN to obtain the pre-trained common sense knowledge feature vector representation of the word by predicting the relationship between nodes;
[0012] Taking sentences as units, we obtain the paths and relationships between the specified parts of speech and domain concepts in the sentences through the domain common sense graph, construct a sub-graph, input the sub-graph into the graph convolutional network GCN, obtain the vector representation of the sub-graph, and map it to the same distribution space as the BERT encoder through the feature space conversion layer to obtain the common sense knowledge feature vector representation of the word; concatenate the grammatical knowledge feature vector representation and the common sense knowledge feature vector representation of words as the word feature representation;
[0013] Taking each word in the sentence as a unit, input the word feature representation into the training classifier, train the classifier, and obtain the optimal model parameters;
[0014] The unlabeled data of the target domain is input into the trained fine-grained sentiment analysis model, and the classification task is performed on the finally concatenated word feature representation vector to output a predicted label, thereby completing the recognition of aspect words in the target domain and their sentiment polarity.
[0015] Furthermore, the steps of pre-training the BERT encoder are also included:
[0016] The Pos-tag and dependency relation in the Spacy library are used to change the two sub-supervision tasks of the BERT encoder, thereby improving the sensitivity of the BERT encoder to grammatical common sense.
[0017] Furthermore, the extracted grammatical knowledge feature vector is expressed as:
[0018] h i = transformer(e i)
[0019]
[0020]
[0021] Among them, ei is the continuous word embedding vector corresponding to the index word, h i Represents mapping ei to the corresponding pre-trained word embedding vector through multiple layers of transfermer. They correspond to the two self-supervised tasks of the BERT encoder respectively; Wp and bp are weight matrices; and are the representations of the head tag and sub-tag of the corresponding index word in the dependency tree, Predict the dependencies between indexed words; [;], [-] and ⊙ represent concatenation, subtraction and multiplication operations respectively; W d is the weight matrix for relation classification.
[0022] Furthermore, the loss function of the grammatical knowledge feature vector representation implemented by the BERT encoder is as follows:
[0023]
[0024] Among them, D U represents the unlabeled samples; I1(i) represents whether the index word i is [MASK] in the self-supervised task, if yes, it is 1, otherwise it is 0. represents the real POS-tag label of the index word; similarly, I2(ij) also predicts the dependency relationship only for the index words that have connections in the dependency tree; λ is a hyperparameter that weighs the weight of the loss function; T represents the word set of unlabeled samples; represents the grammatical relationship label between index words i and j.
[0025] Furthermore, the graph convolutional network GCN is trained in the following way:
[0026] The sub-graph of each unlabeled sample in the source domain and the target domain is used to train the graph autoencoder of the graph convolutional network GCN. The common sense relational structural features are captured by convolving the features of adjacent nodes. The extracted feature vector is expressed as:
[0027]
[0028] Where N(t) represents the neighborhood nodes of word t; represents the hidden features of node g in layer l; W l and b lis a weight matrix; using the representation of nodes in the sub-graph, we perform self-supervised relationship classification tasks and optimize the graph autoencoder;
[0029] In order to mine the representation of each node in the graph, a random sampling method is adopted, and two nodes are given to perform pre-training on the classification task of judging the node relationship.
[0030] Furthermore, during the training process, the cross entropy loss function is used as follows:
[0031]
[0032] Among them, [o i ;o j ] represents the concatenation of the representations of the i-th node and the j-th node; p(y r |[o i ;o j ] represents the predicted relationship y given two nodes r The probability of; N and M are the number of random draws.
[0033] Furthermore, the word feature is expressed as:
[0034] v i =[t c ;t s ]
[0035] p i =softmax(Wv i +b)
[0036] Among them, t s represents the grammatical knowledge feature vector representation obtained by the BERT encoder, t c It represents the feature vector representation with domain common sense relationship after passing through the graph convolution network GCN and then mapped to the same spatial dimension as the BERT encoder through spatial mapping. [;] represents vector concatenation; the word feature representation vector also has node-level common sense knowledge feature representation t c and word-level grammatical knowledge feature representation t s , thereby improving the fine-grained sentiment analysis task in the target domain; finally, the probability of each index word belonging to all the final labels is predicted through the full connection and softmax layer, and W and b are both weight matrices.
[0037] Furthermore, the step of inputting the word feature representation into a training classifier and training the classifier to obtain optimal model parameters includes:
[0038] Taking each word in the sentence as a unit, the word feature representation is input into the training classifier, and the Adam optimizer is used to train the model parameters to obtain the optimal model parameters and achieve domain adaptation at the word level;
[0039] Among them, the cross entropy loss function is used to optimize the model parameters. The specific loss function is:
[0040]
[0041] Where T is the number of words in the sentence, p(y|v i ) is the probability of predicting the label y for a given instance vector.
[0042] Another technical solution adopted by the present invention is:
[0043] A cross-domain fine-grained sentiment analysis device, comprising:
[0044] at least one processor;
[0045] at least one memory for storing at least one program;
[0046] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0047] Another technical solution adopted by the present invention is:
[0048] A computer-readable storage medium stores a program executable by a processor, wherein the program executable by the processor is used to execute the method described above when executed by the processor.
[0049] The beneficial effect of the present invention is that the present invention utilizes grammatical knowledge relations and common sense knowledge relations to narrow the difference between source domain and target domain data, thereby improving the prediction effect of aspect recognition and sentiment analysis of the model migration from the source domain to the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 is a flow chart of a cross-domain fine-grained sentiment analysis method based on knowledge fusion in an embodiment of the present invention;
[0052] Figure 2It is a schematic diagram of the model structure of cross-domain fine-grained sentiment analysis based on knowledge fusion in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0054] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0055] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0056] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0057] like Figure 1 As shown, this embodiment provides a cross-domain fine-grained sentiment analysis method based on knowledge fusion, the method constructs a cross-domain fine-grained sentiment analysis model, the model structure is as follows Figure 2 As shown, it includes pre-training the language model BERT grammatical knowledge module, building a graph pre-training common sense knowledge relationship module, vector splicing as a word feature representation and training a classifier, and the method includes the steps of:
[0058] Step 1: Input the text of the source domain or the target domain into the BERT grammatical knowledge module of the pre-trained language model of the cross-domain fine-grained sentiment analysis model, and map each word in the text into a feature vector with a grammatical structure.
[0059] The word embedding vector obtained by mapping each word in the text is expressed as:
[0060] h i = transformer(e i )
[0061]
[0062]
[0063] Among them, e i is the continuous word embedding vector corresponding to the index word, h i It represents mapping ei to the corresponding pre-trained word embedding vector through multi-layer transfermer. The remaining two formulas correspond to the two self-supervised tasks of BERT. After randomly replacing 20% of the words and their pos-tag labels with [MASK], the softmax layer outputs the pos-tag label type. The probability of. Wp and bp are weight matrices; and are the representations of the head tag and sub-tag of the corresponding index word in the dependency tree, Predict the dependencies between indexed words. [;], [-], and ⊙ represent concatenation, subtraction, and multiplication operations, respectively. d is the weight matrix of the relation classification, with the number of dependencies. When the index words are directly connected in the dependency tree, we predict the dependency between them through the softmax layer.
[0064] In order to obtain domain-shared grammatical features and reduce domain differences, it is necessary to obtain the optimal parameters of the pre-trained language model BERT. The loss function of the language knowledge feature vector implemented by BERT is as follows:
[0065]
[0066] Among them, D U represents the unlabeled sample (Domain unlabeled), and we use the cross entropy function to optimize it. I1(i) represents whether the index word i is [MASK] in the self-supervised task, if yes, it is 1, otherwise it is 0. represents the real POS-tag label of the index word. Similarly, I2(ij) also predicts the dependency relationship only for the index words that have connections in the dependency tree. λ is a hyperparameter that weighs the weight of the loss function and is used to control the contribution of the two self-supervised tasks.
[0067] Step 2: In the process of pre-training common sense knowledge relations, take sentences as units, find all paths with less than n hops between the specified part-of-speech words in the sentence, such as nouns, verbs, and adjectives, and the domain concept words (such as Restaurant nodes in the Restaurant domain) in the ConceptNet common sense relation library, and construct small graphs. These small graphs can together form the sub-graph of the sentence unit. The sub-graph of each unlabeled comment in the source domain and the target domain is used to train the graph autoencoder of the graph convolutional network (GCN), so as to capture the common sense relation structure features by convolving the features of adjacent nodes. Therefore, the extracted feature vector is expressed as:
[0068]
[0069] Where N(t) represents the neighborhood nodes of word t; represents the hidden features of node g in layer l; W l and b l is the weight matrix. Using the representation of nodes in the subgraph, we perform self-supervised relation classification tasks and optimize the graph autoencoder.
[0070] In order to mine the representation of each concept (i.e., node) in the graph, a random extraction method is used to pre-train two nodes for the classification task of determining the node relationship. In order to obtain the best model parameters, the cross entropy loss function is used as follows:
[0071]
[0072] Among them, [o i ;o j ] represents the concatenation of the representations of the i-th node and the j-th node; p(y r |[o i ;o j ] represents the predicted relationship y given two nodes r The probability of; N and M are the number of random draws.
[0073] Step 3: Taking sentences as units, obtain the sub-graph constructed by the paths and relationships between the specified parts of speech and domain concepts in the sentence through the domain common sense graph, input it into GCN to obtain the vector representation of the sub-graph and map it to the same distribution space as BERT through the feature space conversion layer to obtain the common sense knowledge feature vector representation of the word; the BERT grammatical knowledge relationship feature vector representation and the domain shared common sense knowledge relationship vector representation of the word and the domain concept are connected as the word feature representation.
[0074] That is, each sentence is encoded by a BERT encoder and a GCN encoder respectively, so that the word vector representation has word-level grammatical knowledge feature representation and node-level common sense knowledge feature. The final vector of each word in the specific text is specifically represented as:
[0075] v i =[t c ;t s ]
[0076] p i =softmax(Wv i +b)
[0077] Among them, t s Represents the grammatical knowledge vector representation obtained after the BERT module, t c It represents the feature vector representation with domain common sense relationship after passing through the GCN module and then being mapped to the same spatial dimension as BERT through spatial mapping. [.;.] represents vector concatenation. This makes the final word feature representation vector also have node-level common sense knowledge feature representation t c and word-level grammatical knowledge feature representation t s , thereby improving the fine-grained sentiment analysis task in the target domain. Finally, the probability of each indexed word belonging to all the final labels is predicted through the full connection and softmax layer. W and b are both weight matrices.
[0078] Step 4: In the process of inputting word feature representation into the training classifier and using Adam optimizer to train model parameters, the cross entropy loss function is used to optimize the model parameters. The specific loss function is:
[0079]
[0080] Where n is the number of training samples, p(y|v i ) is the probability of predicting the label y for a given instance vector representation. y∈{B-POS,I-POS,E-POS,S-POS,B-NEG,I-NEG,E-NEG,S-NEG,B-NEU,I-NEU,E-NEU,S-NEU,O}. Where {B,I,E,S,O} represent the boundary of the aspect, the beginning, the middle, the end of the multi-word aspect, the single aspect, and not the aspect. T is the number of words in the sentence.
[0081] Step 5. After obtaining the final cross-domain fine-grained sentiment analysis model, when the unlabeled data of the target domain is input, the grammatical knowledge vector representation and the common sense knowledge vector representation are generated through two modules respectively. Finally, the two vectors are concatenated as the word feature representation vector for classification task output prediction label, and the recognition of aspect words and their sentiment polarity in the target domain is completed.
[0082] As can be seen from the above, this embodiment is a cross-domain fine-grained sentiment analysis method based on knowledge fusion, which constructs aspect word recognition and a sentiment analysis model based on aspect words. The method comprises: inputting text from a source domain or a target domain, inputting unlabeled samples into a BERT pre-trained language model to obtain a grammatical knowledge vector representation of each word; and treating words of specified parts of speech, i.e., nouns, adjectives, and verbs in a sentence as domain concept seeds, using a ConceptNet common sense knowledge base to capture the distance and relationship between the words and the domain concept seeds to construct a domain common sense graph, and pre-training a graph autoencoder of a graph convolutional network (GCN), thereby capturing common sense relationship structural features by convolving features of adjacent nodes and mapping them to the same word-level dimensional vector space as BERT to obtain a feature vector representation of common sense knowledge; concatenating a grammatical knowledge vector representation and a common sense knowledge vector representation as the final feature representation of a word; inputting them into a training classifier and optimizing model parameters using an Adam optimizer; inputting target domain data, generating a grammatical knowledge vector representation and a common sense relationship knowledge vector representation through two modules respectively, and finally concatenating the two vectors as a word feature representation vector, and inputting them into a classifier, thereby predicting the sentiment label of a given sample, and realizing the recognition of aspect words and their sentiment polarity in the target domain. This embodiment reduces the domain differences between different fields in the same distribution space by combining grammatical knowledge and common sense relationship knowledge, has strong adaptability to target fields with fewer resources, and improves the prediction effect of aspect extraction and sentiment analysis in the target field.
[0083] This embodiment also provides a cross-domain fine-grained sentiment analysis device, including:
[0084] at least one processor;
[0085] at least one memory for storing at least one program;
[0086] When the at least one program is executed by the at least one processor, the at least one processor implements Figure 1 The method shown.
[0087] A cross-domain fine-grained sentiment analysis device of this embodiment can execute a cross-domain fine-grained sentiment analysis method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0088] The present application also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes Figure 1The method shown.
[0089] This embodiment also provides a storage medium, which stores instructions or programs that can execute a cross-domain fine-grained sentiment analysis method provided by an embodiment of the method of the present invention. When the instructions or programs are run, any combination of implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method can be obtained.
[0090] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0091] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0092] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0093] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0094] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0095] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0096] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0097] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0098] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A cross-domain fine-grained sentiment analysis method, characterized in that: The following steps are involved: Constructing a fine-grained sentiment analysis model for the target domain, the fine-grained sentiment analysis model includes a pre-trained grammatical knowledge feature vector representation module, a pre-trained common sense knowledge relationship vector representation module, and a classifier; the pre-trained grammatical knowledge feature vector representation module includes a BERT encoder, and the pre-trained common sense knowledge relationship vector representation module includes a graph convolutional network GCN; Inputting the unlabeled texts from the source domain and the target domain into the BERT encoder to obtain a grammatical knowledge feature vector representation of each word in the text; Inputting text from a source domain or a target domain into the pre-trained common sense knowledge relationship vector representation module; The words of the specified part of speech in each sentence in each field are input into the ConceptNet common sense knowledge base, and the path and relationship between each word of the specified part of speech and the field concept are obtained to construct a field common sense graph; the field common sense graph is input into the graph convolutional network GCN to predict the relationship between nodes, thereby obtaining the pre-trained common sense knowledge feature vector representation of the word; Taking sentences as units, we obtain the paths and relationships between the specified parts of speech and domain concepts in the sentences through the domain common sense graph, construct a sub-graph, input the sub-graph into the graph convolutional network GCN, obtain the vector representation of the sub-graph, and map it to the same distribution space as the BERT encoder through the feature space conversion layer to obtain the common sense knowledge feature vector representation of the word; The grammatical knowledge feature vector representation and the common sense knowledge feature vector representation of the words are concatenated as the word feature representation; Taking each word in the sentence as a unit, input the word feature representation into the training classifier, train the classifier, and obtain the optimal model parameters; Input the unlabeled data of the target domain into the trained fine-grained sentiment analysis model, perform classification tasks on the finally concatenated word feature representation vectors to output predicted labels, and complete the recognition of aspect words in the target domain and the sentiment polarity of aspect words; Also included are the steps for pre-training the BERT encoder: Use the Pos-tag and dependency relation in the Spacy library to change the two sub-supervision tasks of the BERT encoder, thereby improving the sensitivity of the BERT encoder to grammatical common sense; The extracted grammatical knowledge feature vector is expressed as: h i =transformer(e i ) Among them, ei is the continuous word embedding vector corresponding to the index word, h i Represents mapping ei to the corresponding pre-trained word embedding vector through multiple layers of transfermer. They correspond to the two self-supervised tasks of the BERT encoder respectively; Wp and bp are weight matrices; and are the representations of the head tag and sub-tag of the corresponding index word in the dependency tree, Predict the dependencies between indexed words; [;], [-] and ⊙ represent concatenation, subtraction and multiplication operations respectively; W d is the weight matrix for relation classification.
2. A cross-domain fine-grained sentiment analysis method according to claim 1, characterized in that: The loss function of the grammatical knowledge feature vector representation implemented by the BERT encoder is as follows: Among them, D U represents the unlabeled samples; I1(i) represents whether the index word i is [MASK] in the self-supervised task, if yes, it is 1, otherwise it is 0. represents the real POS-tag label of the index word; similarly, I2(ij) also predicts the dependency relationship only for the index words that have connections in the dependency tree; λ is a hyperparameter that weighs the weight of the loss function; T represents the word set of unlabeled samples; represents the grammatical relationship label between index words i and j.
3. A cross-domain fine-grained sentiment analysis method according to claim 1, characterized in that: The graph convolutional network GCN is trained in the following way: The sub-graph of each unlabeled sample in the source domain and the target domain is used to train the graph autoencoder of the graph convolutional network GCN. The common sense relational structural features are captured by convolving the features of adjacent nodes. The extracted feature vector is expressed as: Where N(t) represents the neighborhood nodes of word t; Represents the hidden features of node g in layer l; W l and b l is the weight matrix; Using the representation of nodes in the subgraph, we perform self-supervised relation classification tasks and optimize graph autoencoders. In order to mine the representation of each node in the graph, a random sampling method is adopted, and two nodes are given to perform pre-training on the classification task of judging the node relationship.
4. A cross-domain fine-grained sentiment analysis method according to claim 3, characterized in that: During the training process, the cross entropy loss function is used as follows: Among them, [o i ;o j ] represents the concatenation of the representations of the i-th node and the j-th node; p(y r |[o i ;o j ] represents the predicted relationship y given two nodes r The probability of; N and M are the number of random draws.
5. The cross-domain fine-grained sentiment analysis method according to claim 1, characterized in that: The word feature is expressed as: v i =[t c ;t s ] p i =softmax(Wv i +b) Among them, t s represents the grammatical knowledge feature vector representation obtained by the BERT encoder, t c It represents the feature vector representation with domain common sense relationship after passing through the graph convolution network GCN and then mapped to the same spatial dimension as the BERT encoder through spatial mapping. [;] represents vector concatenation; the word feature representation vector also has node-level common sense knowledge feature representation t c and word-level grammatical knowledge feature representation t s , thereby improving the fine-grained sentiment analysis task in the target domain; finally, the probability of each index word belonging to all the final labels is predicted through the full connection and softmax layer, and W and b are both weight matrices.
6. A cross-domain fine-grained sentiment analysis method according to claim 1, characterized in that: The inputting the word feature representation into the training classifier, training the classifier, and obtaining the best model parameters includes: Taking each word in the sentence as a unit, the word feature representation is input into the training classifier, and the Adam optimizer is used to train the model parameters to obtain the optimal model parameters and achieve domain adaptation at the word level; Among them, the cross entropy loss function is used to optimize the model parameters. The specific loss function is: Where T is the number of words in the sentence, p(y|v i ) is the probability of predicting the label y for a given instance vector.
7. A cross-domain fine-grained sentiment analysis device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 6.
8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 6 when executed by the processor.
Citation Information
Patent Citations
Aspect-level cross-domain emotion analysis method based on CNN
CN112163091A
System And Method For Fuzzy Concept Mapping, Voting Ontology Crowd Sourcing, And Technology Prediction
US20140075004A1