Feature enhancement method for multi-modal aspect level sentiment analysis
By using the BERT model and semantic alignment module to obtain aspect word features in the multimodal aspect-level sentiment analysis model, and combining the syntactic knowledge enhancement module to obtain syntactic relationships, the difficulties in the existing model in acquiring aspect-level word features and image visual features are solved, and richer feature representations and more accurate emotional discrimination are achieved.
Patent Information
- Application Number
- CN202510197333.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing multimodal aspect-level sentiment analysis model has difficulties in acquiring aspect-word feature information, especially ignoring the syntactic information interaction and the alignment of image visual features with aspect-word semantics, resulting in insufficient feature representation.
The BERT model is used to encode the aspect word, and the image data is converted into adjective noun pairs through the semantic alignment module, and the semantic similarity is fused to obtain supplementary information of the aspect word. At the same time, a syntactic knowledge enhancement module is designed, a syntactic relationship dependency diagram with the root node as the aspect word is constructed, and a syntactic relationship is obtained to enhance the text feature representation.
Through the combination of the semantic alignment module and the syntactic knowledge enhancement module, the local area information and syntactic relationships related to the aspect word can be more effectively obtained, enrich the expression of aspect word characteristics, and improve the accuracy of emotional discrimination.
Smart Images

Figure CN120031044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and sentiment analysis, and in particular to a feature enhancement method for multimodal aspect-level sentiment analysis. Background Art
[0002] For early pure text speeches, researchers focused on aspect-based sentiment analysis (ABSA) tasks, such as Figure 1 As shown in the figure, the current ABSA model has changed the previous practice of predicting the overall sentiment trend of the text, and has made many achievements in judging the sentiment polarity of specific aspect words in the text. In recent years, as the type of comment data on the Internet has changed from pure text to graphic mode, researchers have turned their attention to multimodal aspect-based sentiment analysis (MABSA). Figure 2 As shown in Figure 2, MABSA can analyze the sentiment polarity of a specific aspect word from both the text and image perspectives. The challenge of the existing MABSA model is how to obtain feature information focused on aspect words.
[0003] Most MABSA models only use LSTM or BERT to extract semantic relations in sentences when representing text modalities, but ignore the syntactic information interaction between words, so that the extracted text features cannot more comprehensively represent the information contained in the text, resulting in ignoring key information related to aspect words when modeling the relationship between text and aspect words and pictures.
[0004] Compared with the ABSA task, another challenge of the MABSA task is how to process the image visual features so that they can be consistent with the semantics of the aspect words. For the image visual feature representation extracted by the pre-trained model, it is necessary to model its interaction with the aspect words and text. The traditional way is to use the attention mechanism to obtain the image-text modality pointing to a specific aspect word. For example, the attention mechanism is used to fuse the aspect words and the image-text modality separately, and the aspect word information is embedded in the final feature representation of the image-text modality. However, the above work only roughly combines the aspect words and the image, without considering the noise in the area where the image is not related to the aspect words. Subsequently, a distance-based correlation algorithm is used to obtain the correlation between the aspect words and the local area of the image using the obtained local semantic similarity, and the reliability of the local semantic similarity is obtained using the global semantic similarity, so as to measure the contribution level of the local area of the image to the final sentiment discrimination and obtain the image content related to the aspect words. After that, a new learning architecture is proposed to jointly perform coarse-grained classification and fine-grained alignment of the correlation between the local area of the image and the aspect words, and assign weight coefficients to the local area of the image according to the given aspect words, so as to filter out irrelevant image noise. However, the heterogeneity of the image-text modality itself leads to a semantic gap between the two. In this regard, one is to use non-autoregressive generated text to generate the title of the image, use the title to match the aspect words, and extract the facial features of the face as emotional clues to enhance the emotional orientation of the title. The second is to improve the impact caused by the structural differences of the image and text modalities themselves by building an external knowledge base to mine the local areas with the highest matching degree with the aspect words and the corresponding emotional information. However, the semantic alignment of the image and specific aspect words does not pay attention to the importance of the image to the aspect word vector, and it is difficult to accurately match the local area of the image with the highest correlation with the aspect word. At the same time, the emotional information in the picture is not better integrated into the final feature representation, making the emotional content contained in the feature vector used for emotion prediction not rich enough. Summary of the invention
[0005] In order to overcome the semantic alignment problem between images and specific aspect words and the technical defects of lack of syntactic dependency in text feature representation, the present invention provides a feature enhancement method for multimodal aspect-level sentiment analysis.
[0006] The present invention provides a feature enhancement method for multimodal aspect-level sentiment analysis, comprising the following steps:
[0007] S1, aspect word encoding;
[0008] For aspect words, the BERT model is used to perform initial encoding on the aspect words, obtain the feature information of the aspect words, and obtain the aspect word feature vector H a; Through the semantic alignment module, the image data is converted into multiple sets of adjective noun pairs. The nouns contain information related to the aspect words, and the adjectives point to the emotional information contained in the corresponding nouns. The semantic alignment module determines which noun is the best choice for matching the auxiliary image with the aspect words, and uses the trilinear similarity algorithm to calculate the semantic similarity between the aspect words and the nouns, and finds the noun with the greatest similarity. As supplementary information for the aspect word representation, the nouns calculated by the semantic alignment module are and aspect word feature vector H a Fusion is performed to obtain the final aspect word representation H' a ;
[0009] S2, text modality coding;
[0010] For text S, the BERT model is used to initially encode the text, obtain the feature information of the text, and obtain the text feature vector Get the long-distance dependencies of the text and capture the real text information. n is the number of words in the text S. S The syntactic relationship H between aspect words and all words in text S is obtained through the syntactic knowledge enhancement module. r , so as to notice all the words related to the aspect words; secondly, H r With H S The improved text representation H' is obtained by splicing and fusion S , the improved text representation H' S With the final aspect word H' a The multi-head cross attention mechanism is input together to obtain the text feature representation H guided by the aspect words. S-a ;
[0011] S3, image visual encoding;
[0012] For the original image, resize and normalize it to standardize the image format, and then use the ResNet-152 network to perform image feature H V Extraction; then the final aspect word is represented by H' a With image feature H V The multi-head cross attention mechanism is used for fusion to obtain the image feature representation H guided by aspect words V-a ;
[0013] S4, multimodal fusion;
[0014] The text features guided by aspect words are represented as H S-a Image feature representation H guided by aspect words V-a After fusion and splicing, the image-text modality feature vector H containing aspect word information is obtained cAt the same time, the adjectives in the adjective-noun pair are added to H through the sentiment enhancement module c In this way, the final sentiment tendency is more directed to the aspect words rather than the entire sentence, and H c ', and then fully integrate all the internal sentiment information about aspect words through the self-attention mechanism, and finally obtain the vector representation H for sentiment discrimination c ”.
[0015] Compared with the prior art, the technical solution provided by the present invention has the following technical effects:
[0016] (1) To address the semantic alignment problem between images and specific aspect words, the present invention uses text to represent the local area information of the image related to the aspect words, and then constructs a semantic alignment module to narrow the semantic gap between images and texts, enrich the feature representation of aspect words, and more effectively guide the feature representation of image and text modalities to focus on the vicinity of aspect words. At the same time, the emotion enhancement module is used to give the final feature representation more emotional information pointing to aspect words.
[0017] (2) To address the problem of lack of syntactic dependencies in text feature representation, the present invention designs a syntactic knowledge enhancement module. By constructing a syntactic dependency graph with aspect words as the root node, the syntactic knowledge focused on aspect words is acquired, thereby solving the problem of lack of syntactic information in the process of text modal feature extraction and enhancing the learning of text modal feature information representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 It is the framework diagram of the ABSA model in the background technology;
[0021] Figure 2 It is the framework diagram of the MABSA model in the background technology;
[0022] Figure 3 A schematic diagram of a flow chart of a feature enhancement method for multimodal aspect-level sentiment analysis described in an embodiment of the present invention;
[0023] Figure 4 It is an overall framework diagram of a feature enhancement method for multimodal aspect-level sentiment analysis described in an embodiment of the present invention;
[0024] Figure 5 A schematic diagram of a semantic alignment module of a feature enhancement method for multimodal aspect-level sentiment analysis described in an embodiment of the present invention;
[0025] Figure 6 A schematic diagram of a syntactic knowledge enhancement module of a feature enhancement method for multimodal aspect-level sentiment analysis described in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.
[0028] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0029] In one embodiment, Figure 1 As shown, a feature enhancement method for multimodal aspect-level sentiment analysis is disclosed, comprising the following steps:
[0030] S1, aspect word encoding;
[0031] For aspect words, the BERT model is used to perform initial encoding on the aspect words, obtain the feature information of the aspect words, and obtain the aspect word feature vector H a ; The data in the dataset appear in the form of image-text pairs, so it is necessary to consider the contribution of images to the representation of aspect words. The image data can be converted into text form through the semantic alignment module (Semantic Alignment Module, SAM), which can not only extract the content in the image that is highly related to aspect words, but also narrow the semantic structure difference between image vision and aspect words; Through the semantic alignment module, the image data is converted into 5 groups of adjective-noun pairs, the nouns contain information related to aspect words, and the adjectives point to the emotional information contained in the corresponding nouns. The semantic alignment module determines which noun is the best choice for assisting the image to match the aspect words, and uses the trilinear similarity algorithm to calculate the semantic similarity between aspect words and nouns, and finds the noun with the greatest similarity. As supplementary information for the aspect word representation, the nouns calculated by the semantic alignment module are and aspect word feature vector H aFusion is performed to obtain the final aspect word representation H' a ; The BERT model introduces a masked language model to perform unsupervised pre-training on the input data to fully understand the contextual meaning of relevant information on the left and right sides of the text;
[0032] S2, text modality coding;
[0033] For text S, the BERT model is used to initially encode the text, obtain the feature information of the text, and obtain the text feature vector Get the long-distance dependencies of the text and capture the real text information. n is the number of words in the text S. S The syntactic relationship H between the aspect words and all the words in the text S is obtained through the Syntactic Knowledge Enhancement Module (SKEM) r , so as to notice all the words related to the aspect words; secondly, H r With H S The improved text representation H' is obtained by splicing and fusion S , the improved text representation H' S With the final aspect word H' a The multi-head cross attention mechanism is used to obtain information focused on aspect words in the text modality, and the text feature representation H guided by aspect words with deep semantic and syntactic knowledge is obtained. S-a ;
[0034] S3, image visual encoding;
[0035] For the original image, resize and normalize it to standardize the image format, and then use the ResNet-152 network to perform image feature H V Extraction; then the final aspect word is represented by H' a With image feature H V The multi-head cross attention mechanism is used for fusion to obtain the image feature representation H guided by aspect words V-a ;
[0036] S4, multimodal fusion;
[0037] The text modality uses syntactic knowledge to concentrate the sentiment information of the given aspect words in the text modality, while the sentiment information of the image modality is hidden in the extracted adjectives. In the aspect word encoding stage, nouns are used as references for the final representation of aspect words to enhance the representation of aspect words. The corresponding adjectives are fused into the final feature representation through the sentiment enhancement module in the multimodal fusion stage to enrich the sentiment information in the features.
[0038] The text features guided by aspect words are represented as HS-a Image feature representation H guided by aspect words V-a After fusion and splicing, the image-text modality feature vector H containing aspect word information is obtained c At the same time, the adjectives in the adjective-noun pair are added to H through the emotional enhancement module (EEM). c In this way, the final sentiment tendency is more directed to the aspect words rather than the entire sentence, and H c ', and then fully integrate all the internal sentiment information about aspect words through the self-attention mechanism, and finally obtain the vector representation H for sentiment discrimination c ”.
[0039] Based on the above embodiment, in a preferred embodiment, in step S1, the processing process of the semantic alignment module is:
[0040]
[0041] In formulas (1) and (2), Contains the i-th noun to represent the result, and b a are all learnable weight parameters, α m Yes H a and The highest similarity score obtained by the trilinear similarity algorithm, is α m The corresponding noun, H' N It is the weighted noun representation, which represents the matching of image information with aspect words;
[0042] In order to enrich the information content of aspect words, the semantic alignment module calculates the noun and aspect word feature vector H a After fusion, formula (3) is:
[0043] H' a =H a +ω n H' N (3)
[0044] In formula (3), ω n is an adjustable parameter used to dynamically change the final aspect word representation H' a This allows the aspect word features to not only contain text features, but also better reflect the important information represented by the image.
[0045] Based on the above embodiments, in a preferred embodiment, in step S2, the syntactic knowledge enhancement module discovers the syntactic structure of a sentence through dependency parsing. Dependency parsing is a task of generating a dependency graph to represent the grammatical structure. The relationship between words in a sentence is represented by directed edges and labels. Assume that there are n nodes in the dependency graph, each node represents a word in the sentence, and the neighborhood node set of node i is represented by N i Indicates that r ij Represents the directed edge between node i and node j, which contains the syntactic relationship between words; the traditional graph attention network uses grammatical knowledge to construct the root nodes, so that the grammatical information related to aspect words may be lost. At the same time, the way of encoding the entire dependency tree is very easy to introduce noise irrelevant to the task. Based on the basic syntactic dependency graph, the syntactic knowledge enhancement module reconstructs the root node and constructs a virtual syntactic relationship to obtain the syntactic relationship between all words and aspect words in the text context, so as to pay attention to all words related to aspect words and enhance the quality of text feature representation;
[0046] Firstly, the Biaffine parse tree is used to learn the syntactic relationship between words in a sentence, and the root node of the obtained syntactic dependency tree is placed on the aspect word to obtain the improved dependency graph. Secondly, the dependency relationship between nodes is used to iteratively update the representation of each node.
[0047] Taking a single dependency graph attention hidden layer as an example, we first need to use multi-head attention to aggregate all neighborhood node representations. The process is shown in formulas (4) and (5):
[0048]
[0049] In formulas (4) and (5), h α is the weight coefficient between node i and node j. Since the node encoder used is BERT, the value of the number of attention heads K is 12; Represents the vector from x 1 to x k The splicing, is the normalized attention coefficient of the kth attention head, W k is the input transformation matrix, att(·) is a dot product operation;
[0050] Secondly, since the relationship graph attention network adds the importance calculation of the edge to the node, the syntactic relationship information is integrated into the feature representation of the node, which can better realize the aggregation and update of the node. ij Weigh the weight of each directed edge in the improved dependency graph to node i, and aggregate the relationship feature representations obtained from all neighboring nodes by summing and concatenating them. Take a single dependency graph attention layer as an example:
[0051]
[0052] In formulas (6), (7), and (8), μ is the number of relation heads. In this implementation, 5 relation heads are used. It is the attention weight assigned to the relationship between nodes after calculation using linear transformation and Relu activation function. μ is the transformation matrix for vector dimension conversion, h β The hidden layer finally aggregates the features with syntactic relations; the obtained h α and h β Directly spliced together, the generated h r contains both node and directed edge syntactic information:
[0053] h r =h α ||h β (9)
[0054] In formula (8), || is a concatenation operation. Considering the complexity of the syntactic structure, the output results of two graph attention layers are accumulated as the final output H of the syntactic knowledge enhancement module. r .
[0055] Based on the above embodiment, in a preferred embodiment, in step S2, in order to make the final text feature representation contain deep semantic information in terms of both semantics and syntax, H r With H S The improved text representation H' is obtained by concatenation S :
[0056] H' S =H r ||H S (10)
[0057] The self-attention mechanism encodes the input vector by itself. It obtains three different representations of an input vector using different matrix coefficients: query vector Q, key vector K, and value vector V, and then obtains the influence of each dimension within the input vector:
[0058]
[0059] In formula (9), d k is the vector dimension of Q, and f(·) calculates the weight result of a single attention layer;
[0060] The multi-head cross attention mechanism focuses on the importance distribution of the relationship between different modalities or different sequences. By calculating the similarity between the Q vector of one modality A and the K and V vectors of another modality B, it dynamically adjusts the attention of each modality to other modalities, thereby achieving comprehensive utilization of information.
[0061] The multi-head cross attention mechanism first calculates the value of a single cross attention layer to obtain the correlation coefficient between the various parts of the input, and then uses multiple attention heads to focus on the influence of different dimensions between the input vectors to enrich the semantic features:
[0062]
[0063] In formula (10), are the query and key transformation matrices of the mth head, respectively. The output of C-att(·) is the operation result of a single-layer cross attention mechanism;
[0064] Then H' S With H' a In the common input multi-head cross attention mechanism, the text feature representation H guided by aspect words S-a The formula is:
[0065] H S-a =Mul(H' a ,H' S ) (13).
[0066] On the basis of the above embodiments, in a preferred embodiment, in step S3, ResNet uses residual modules and residual connections to construct a deep residual network, which solves the gradient vanishing problem that occurs in the process of training the deep network while retaining the original features, and improves the robustness of the model learning deep neural network; the original image is divided into multiple regions using the ResNet-152 network, and the hidden features of the image are mined through multiple layers of convolution, and the output of the last layer of convolution is used as the final representation vector of the image; each region is represented by a 2048-dimensional vector Represents, j is the label value of the selected area; a linear function is used to project the visual features into the same space of the text features:
[0067]
[0068] In formula (12), W V ∈R d×2048 is a learnable parameter, d is the dimension of the text feature vector;
[0069] H' a With H V Using multi-head cross attention mechanism for fusion, H' aAs the query value input attention mechanism, capture the image feature representation H guided by aspect words V-a , reduce the impact of irrelevant noise on image feature representation:
[0070] H V-a =Mul(H' a ,H V ) (15).
[0071] 6. A feature enhancement method for multimodal aspect-level sentiment analysis according to claim 1, characterized in that in step S4, after obtaining the semantic vector of the adjective using BERT, the one-to-one correspondence between the adjective and the noun is used to find the noun Corresponding adjective And use it as the emotional information of the target area of the image to assist the final emotional prediction:
[0072]
[0073] In formula (14), ANP(·) refers to the operation of finding adjectives using nouns;
[0074] H c , H' c and H c The calculation formulas are:
[0075] H c =H S-a ||H V-a (17)
[0076]
[0077] H' c '=f(H' c ,H' c ,H' c ) (19)
[0078] In formulas (17), (18), and (19), ω Adj It is an adjustable parameter.
[0079] 7. A feature enhancement method for multimodal aspect-level sentiment analysis according to claim 5, characterized in that in step S4, in order to learn the nouns with the highest relevance, a loss function L is constructed using mean square error m , adjust the parameters by minimizing the mean square error to improve the final prediction effect:
[0080]
[0081] In formula (20), M is the total number of samples;
[0082] The true label y i Input into Softmax and calculate the final sentiment score p i , using the cross entropy loss function L p Calculate the error value of the predicted label; the final loss function L s Combined with L m With L p :
[0083] L p =-∑y i logp i (twenty one)
[0084] L s =(1-ω)L p +ωL m (twenty two)
[0085] In formulas (21) and (22), ω is an adjustable parameter.
[0086] The above is only a specific implementation of the present invention, which enables those skilled in the art to understand or implement the present invention. Although detailed descriptions are given with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A feature enhancement method for multimodal aspect-level sentiment analysis, characterized in that: The following steps are involved: S1, aspect word encoding; For aspect words, the BERT model is used to perform initial encoding on the aspect words, obtain the feature information of the aspect words, and obtain the aspect word feature vector H a ; Through the semantic alignment module, the image data is converted into multiple sets of adjective noun pairs. The nouns contain information related to the aspect words, and the adjectives point to the emotional information contained in the corresponding nouns. The semantic alignment module determines which noun is the best choice for matching the auxiliary image with the aspect words, and uses the trilinear similarity algorithm to calculate the semantic similarity between the aspect words and the nouns, and finds the noun with the greatest similarity. As supplementary information for the aspect word representation, the nouns calculated by the semantic alignment module are and aspect word feature vector H a Fusion is performed to obtain the final aspect word representation H' a ; S2, text modality coding; For text S, the BERT model is used to initially encode the text, obtain the feature information of the text, and obtain the text feature vector Get the long-distance dependencies of the text and capture the real text information. n is the number of words in the text S. S The syntactic relationship H between aspect words and all words in text S is obtained through the syntactic knowledge enhancement module. r , so as to notice all the words related to the aspect words; secondly, H r With H S The improved text representation H' is obtained by splicing and fusion S , the improved text representation H' S With the final aspect word H' a The multi-head cross attention mechanism is input together to obtain the text feature representation H guided by the aspect words. S-a ; S3, image visual encoding; For the original image, resize and normalize it to standardize the image format, and then use the ResNet-152 network to perform image feature H V Extraction; then the final aspect word is represented by H' a With image feature H V The multi-head cross attention mechanism is used for fusion to obtain the image feature representation H guided by aspect words V-a ; S4, multimodal fusion; The text features guided by aspect words are represented as H S-a Image feature representation H guided by aspect words V-a After fusion and splicing, the image-text modality feature vector H containing aspect word information is obtained c At the same time, the adjectives in the adjective-noun pair are added to H through the sentiment enhancement module c In this way, the final sentiment tendency is more directed to the aspect words rather than the entire sentence, and H c ', and then fully integrate all the internal sentiment information about aspect words through the self-attention mechanism, and finally obtain the vector representation H for sentiment discrimination c ”.
2. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 1, characterized in that: In step S1, the processing process of the semantic alignment module is: In formulas (1) and (2), Contains the i-th noun to represent the result, and b a are all learnable weight parameters, α m Yes H a and The highest similarity score obtained by the trilinear similarity algorithm, is α m The corresponding noun, H' N It is the weighted noun representation, which represents the matching of image information with aspect words; The semantic alignment module calculates the nouns using formula (3). and aspect word feature vector H a After fusion, formula (3) is: H' a =H a +ω n H' N (3) In formula (3), ω n is an adjustable parameter used to dynamically change the final aspect word representation H' a .
3. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 2, characterized in that: In step S2, the syntactic knowledge enhancement module discovers the syntactic structure of a sentence through dependency parsing. Dependency parsing is a task of generating a dependency graph to represent the grammatical structure. The relationship between words in a sentence is represented by directed edges and labels. Assume that there are n nodes in the dependency graph, each node represents a word in the sentence, and the set of neighboring nodes of node i is N i Indicates that r ij represents the directed edge between node i and node j, which contains the syntactic relationship between words; Firstly, the Biaffine parse tree is used to learn the syntactic relationship between words in a sentence, and the root node of the obtained syntactic dependency tree is placed on the aspect word to obtain the improved dependency graph. Secondly, the dependency relationship between nodes is used to iteratively update the representation of each node. Taking a single dependency graph attention hidden layer as an example, we first need to use multi-head attention to aggregate all neighborhood node representations. The process is shown in formulas (4) and (5): In formulas (4) and (5), h α is the weight coefficient between node i and node j. Since the node encoder used is BERT, the value of the number of attention heads K is 12; Represents a vector from x1 to x k The splicing, is the normalized attention coefficient of the kth attention head, W k is the input transformation matrix, att(·) is a dot product operation; Secondly, the syntactic relationship information is integrated into the feature representation of the node, using r ij Weigh the weight of each directed edge in the improved dependency graph to node i, and aggregate the relationship feature representations obtained from all neighboring nodes by summing and concatenating them. Take a single dependency graph attention layer as an example: In formulas (6), (7), and (8), μ is the number of relation heads. It is the attention weight assigned to the relationship between nodes after calculation using linear transformation and Relu activation function. μ is the transformation matrix for vector dimension conversion, h β The hidden layer finally aggregates the features with syntactic relations; the obtained h α and h β Directly spliced together, the generated h r contains both node and directed edge syntactic information: h r =h α ||h β (9) In formula (8), || is a concatenation operation. Considering the complexity of the syntactic structure, the output results of multiple graph attention layers are accumulated as the final output H of the syntactic knowledge enhancement module. r .
4. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 3, characterized in that: In step S2, H r With H S The improved text representation H' is obtained by concatenation S : H' S =H r ||H S (10) The self-attention mechanism encodes the input vector by itself. It obtains three different representations of an input vector using different matrix coefficients: query vector Q, key vector K, and value vector V, and then obtains the influence of each dimension within the input vector: In formula (9), d k is the vector dimension of Q, and f(·) calculates the weight result of a single attention layer; The multi-head cross attention mechanism focuses on the importance distribution of the relationship between different modalities or different sequences. By calculating the similarity between the Q vector of one modality A and the K and V vectors of another modality B, it dynamically adjusts the attention of each modality to other modalities, thereby achieving comprehensive utilization of information. The multi-head cross attention mechanism first calculates the value of a single cross attention layer to obtain the correlation coefficient between the various parts of the input, and then uses multiple attention heads to focus on the influence of different dimensions between the input vectors to enrich the semantic features: In formula (10), are the query and key transformation matrices of the mth head, respectively. The output of C-att(·) is the operation result of a single-layer cross attention mechanism; Then H' S With H' a In the common input multi-head cross attention mechanism, the text feature representation H guided by aspect words S-a The formula is: H S-a =Mul(H' a ,H' S ) (13)。 5. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 1, characterized in that: In step S3, the original image is segmented into multiple regions using the ResNet-152 network, and the hidden features of the image are mined through multiple layers of convolution. The output of the last layer of convolution is used as the final representation vector of the image; each region is represented by a 2048-dimensional vector Represents, j is the label value of the selected area; a linear function is used to project the visual features into the same space of the text features: In formula (12), W V ∈R d×2048 is a learnable parameter, d is the dimension of the text feature vector; H' a With H V Using multi-head cross attention mechanism for fusion, H' a As the query value input attention mechanism, capture the image feature representation H guided by aspect words V-a : H V-a =Mul(H' a ,H V ) (15)。 6. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 1, characterized in that: In step S4, after obtaining the semantic vector of the adjective using BERT, the one-to-one correspondence between adjectives and nouns is used to find the noun Corresponding adjective And use it as the emotional information of the target area of the image to assist the final emotional prediction: In formula (14), ANP(·) refers to the operation of finding adjectives using nouns; H c , H' c and H c The calculation formulas are: H c =H S-a ||H V-a (17) H' c '=f(H' c ,H' c ,H' c ) (19) In formulas (17), (18), and (19), ω Adj It is an adjustable parameter.
7. The feature enhancement method for multimodal aspect-level sentiment analysis according to claim 5, characterized in that: In step S4, in order to learn the most relevant nouns, the loss function L is constructed using the mean square error. m , adjust the parameters by minimizing the mean square error to improve the final prediction effect: In formula (20), M is the total number of samples; The true label y i Input into Softmax and calculate the final sentiment score p i , using the cross entropy loss function L p Calculate the error value of the predicted label; the final loss function L s Combined with L m With L p : L p =-∑y i logp i (21) L s =(1-ω)L p +ωL m (22) In formulas (21) and (22), ω is an adjustable parameter.
Citation Information
Patent Citations
Text retrieval matching method and system
CN114428850A
Aspect-level sentiment analysis method fusing multi-modal data
CN114936623A
Emotion analysis method and device, electronic equipment and storage medium
CN116541520A
Visual language navigation method based on double semantic graphs and modal alignment
CN117889864A
Emotion analysis method based on multi-modal resources
CN118070809A
Cited By
Video multi-modal content management system and method based on multi-modal features
CN120599105A
Text and image bimodal fusion-based risk identification system
CN121960712A