A Cross-modal Sentiment Analysis Method and Device Incorporating a Knowledge Graph
By using pre-trained models in multimodal emotion classification to extract multimodal knowledge graphs of images and texts, and combining visual matrix to supplement structural information, the problem of insufficient fusion of images and text features in the prior art is solved, and more efficient multimodal emotion classification is achieved, reducing the need for data and labeling.
Patent Information
- Application Number
- CN202310909452.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-07-24
AI Technical Summary
The existing multimodal emotion classification methods have shortcomings in the fusion and interaction of image and text features, resulting in poor emotional classification effects, and the pre-training method requires a large amount of data and labeling, which is costly.
The pre-trained model is used to extract the image subtitles in text form and the first multimodal knowledge graph at the image level, and extend it into the second multimodal knowledge graph that transmits the image and text connection through entity node alignment, and supplement structural information with the visual matrix, and finally use the sentiment analysis model of the encoder, decoder and the fully connected layer for sentiment analysis.
Through the fusion and serialization of multimodal knowledge graphs, the connection between text and images is enhanced, the accuracy of emotion classification is improved, and the most advanced multimodal emotion classification capabilities are achieved, while reducing the need for data and annotation.
Smart Images

Figure CN116861367B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cross-modal sentiment classification, and specifically relates to a cross-modal sentiment analysis method and device integrating a knowledge graph. Background Art
[0002] In multi-modal sentiment classification, the purpose of multi-modal sentiment learning is to imitate humans in analyzing and understanding complex multi-modal information. For given image and text information, sentiment classification is performed on the sentiment subject, which is beneficial for application in scenarios such as multi-modal tweet sentiment detection in social networks.
[0003] In the technical solutions disclosed in the literatures "Xu, N., Mao, W., Chen, G.: Multi-interactive memory network for aspect-based multimodal sentiment analysis. In: AAAI. pp. 371-378. AAAI Press (2019)" and "Ju, X., Zhang, D., Xiao, R., Li, J., Li, S., Zhang, M., Zhou, G.: Joint multi-modal aspect-sentiment analysis with auxiliary cross-modal relation detection. In: EMNLP(1). pp. 4395-4405. Association for Computational Linguistics (2021)", in both cases, multi-modal sentiment classification obtains a feature representation of an image and text respectively through their respective encoders, and then the two are simply concatenated or mapped to obtain an overall representation of multi-modal information for sentiment classification. Inevitably, these methods will overfit, and it is not conducive to fusing and interacting the features of the image and text, resulting in poor performance of sentiment classification itself.
[0004] In recent years, some pre-training and attention-based methods have emerged, such as the technical solution methods disclosed in the literature Ling, Y., Yu, J., Xia, R.: Vision-language pre-training for multimodal aspect-based sentiment analysis. In: ACL(1). pp. 2149-2159. Association for Computational Linguistics(2022), Wang, J., Liu, Z., Sheng, V.S., Song, Y., Qiu, C.: Saliencybert: Recurrent attention network for target-oriented multimodal sentiment classification. In: PRCV(3). Lecture Notes in Computer Science, vol. 13021, pp. 3-15. Springer(2021). Such methods can greatly promote the interaction between the text and image modalities and can improve the performance of the model to a certain extent. However, the pre-training method itself requires a large amount of data and a large amount of manual annotation, which is quite difficult and costly. Summary of the Invention
[0005] In view of the above, the object of the present invention is to provide a cross-modal sentiment analysis method and device integrating a knowledge graph, which can achieve better performance in the multi-modal sentiment classification task on the basis of further capturing the fine-grained information of the image and constructing the connection among the image, text, and target subject.
[0006] To achieve the above object of the invention, a cross-modal sentiment analysis method integrating a knowledge graph provided by an embodiment includes the following steps:
[0007] Use a pre-trained model to extract image captions in text form and a first multi-modal knowledge graph at the image level from the image;
[0008] Extract the target subject that needs sentiment analysis from the description text and construct a mask template for the target subject;
[0009] Convert the description text into a knowledge graph at the text level through the pre-trained model;
[0010] Expand the first multi-modal knowledge graph and the knowledge graph into a second multi-modal knowledge graph that transmits the connection between the image and the text by means of entity node alignment;
[0011] Flatten the triples in the second multi-modal knowledge graph to obtain serialized triples, and use a visualization matrix to supplement the structural information of the serialized triples;
[0012] Perform sentiment analysis using a sentiment analysis model including an encoder, a decoder, and a fully connected layer. Specifically: the encoder combines the visualization matrix to encode the serialized triples, the descriptive text, and the image captions to obtain output features, the decoder decodes the masked position representations based on the output features of the encoder and the mask template, and the fully connected layer performs connection mapping on the masked position representations to obtain the main body sentiment representation.
[0013] Preferably, the pre-trained model uses an image-to-text model. After extracting global features from the image, the global features are converted into text-form image captions;
[0014] The pre-trained model uses a picture scene graph extraction model to convert the picture into a first multi-modal knowledge graph with structured information. Among them, the first multi-modal knowledge graph contains two forms of triples. One is the triple relationship between entities and entities, and the other is the triple relationship between sub-images and entities. The sub-image is an image block segmented from the input image.
[0015] Preferably, the pre-trained model uses a text scene graph extraction model to convert the descriptive text into a knowledge graph with structured information, where the knowledge graph establishes the relationships between key entities in the text.
[0016] Preferably, the first multi-modal knowledge graph and the knowledge graph are extended into a second multi-modal knowledge graph that transmits the connection between images and texts by means of entity node alignment, including:
[0017] When the text semantic similarity between two nodes in the first multi-modal knowledge graph and the knowledge graph is greater than the set threshold, the two nodes are merged into one new node, thereby obtaining a multi-modal knowledge graph that transmits the connection between images and texts, where the set threshold is preferably 0.9.
[0018] Preferably, the flattening of the triples in the second multi-modal knowledge graph to obtain serialized triples includes:
[0019] The head entity, relationship, and tail entity of each triple are concatenated in a comma-separated manner to form a serialized triple in the form of a text sequence, and special flag bits are used between different triples. <ts>"Separation. For the sub-images in the second multi-modal knowledge graph, use a special flag bit " " to represent.
[0020] Preferably, the visualization matrix realizes the supplementation of structured information by defining the correlation between elements in the serialized triples;
[0021] The definition rule is as follows: In the serialized triples input to the encoder, the elements in the same triple are visible to each other, the shared entities in each triple are visible to each other, and the remaining triples are invisible; the descriptive text, image captions, and other special markers in the input encoder should be visible to each other.
[0022] Preferably, the encoder combines the visualization matrix to encode the serialized triples, descriptive text, and image captions to obtain output features, including:
[0023] Convert each part of the serialized triples, descriptive text, and image captions into embedding vectors, and then input the embedding vectors into the encoder, where the embedding vectors include feature embeddings, position embeddings, and type embeddings;
[0024] In the encoder, for the text marked by the text flag bit, it is encoded by the text encoder, and for the sub-images marked by the special flag " " it is encoded by the image encoder.
[0025] Preferably, the fully connected layer performs a connection mapping on the masked position representation to obtain the main body sentiment representation, which is expressed by the formula:
[0026] p(y||H [m] ) = softmax(θ Linear Dropout(H [m] ))
[0027] where, H [m] represents the masked position " <mask>The eigenvector representation of “, θ Linear Dropout(·) represents a learnable non - linear layer, and softmax(·) represents an activation function. p(y∣∣H [m] ) represents the probability distribution of predicting the sentiment classification label y based on H [m]
[0028] To achieve the above - mentioned invention purpose, an cross - modal sentiment analysis device integrating a knowledge graph provided by an embodiment includes a first multi - modal knowledge graph construction module, a mask template construction module, a knowledge graph construction module, a second multi - modal knowledge graph construction module, a structure information supplement module, and a sentiment analysis module;
[0029] The first multi - modal knowledge graph construction module is used to extract image captions in text form and the first multi - modal knowledge graph at the image level from images by using a pre - trained model;
[0030] The mask template construction module is used to extract the target subject that needs sentiment analysis from the description text and construct a mask template for the target subject;
[0031] The knowledge graph construction module is used to convert the description text into a knowledge graph at the text level by using a pre - trained model;
[0032] The second multi - modal knowledge graph construction module is used to expand the first multi - modal knowledge graph and the knowledge graph into a second multi - modal knowledge graph that transmits the connection between images and texts by means of entity node alignment;
[0033] The structure information supplement module is used to serialize and flatten the triples in the second multi - modal knowledge graph to obtain serialized triples, and supplement the structure information of the serialized triples by using a visualization matrix;
[0034] The sentiment analysis module is used to perform sentiment analysis by using a sentiment analysis model including an encoder, a decoder, and a fully - connected layer. Specifically, the encoder combines the visualization matrix to encode the serialized triples, the description text, and the image captions to obtain output features, the decoder decodes the mask position representation according to the output features of the encoder and the mask template, and the fully - connected layer performs connection mapping on the mask position representation to obtain the subject sentiment representation.
[0035] To achieve the above - mentioned invention purpose, an cross - modal sentiment analysis device integrating a knowledge graph provided by an embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above - mentioned cross - modal sentiment analysis method integrating a knowledge graph is implemented.
[0036] Compared with the prior art, the beneficial effects of the present invention at least include:
[0037] The second multi-modal knowledge graph constructed based on text and pictures performs cross-modal information fusion, enabling the obtained visual features and text features to take into account each other and enhancing the connection between text and images. With the assistance of the serialization and visualization matrix of the multi-modal knowledge graph, the serialized triples, descriptive texts, and image captions form a closer relationship with the target subject. In this way, when performing downstream multi-modal sentiment classification, the accuracy of sentiment classification can be improved, achieving the most advanced multi-modal sentiment classification ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 is a flowchart of the cross-modal sentiment analysis method for the fusion knowledge graph provided by the embodiment;
[0040] Figure 2 is a schematic structural diagram of the sentiment analysis model provided by the embodiment;
[0041] Figure 3 is an example diagram of the visualization matrix provided by the embodiment;
[0042] Figure 4 is a schematic structural diagram of the cross-modal sentiment analysis device for the fusion knowledge graph provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further details the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0044] At present, social networks are becoming increasingly developed, and people always post their own dynamics on social networks. The dynamics include text and picture information. However, it is a difficult and cumbersome process for network auditors to judge the emotional tendency of people through one-by-one multi-modal data, and the misjudgment rate is very high, which is likely to cause network security problems. Therefore, multi-modal emotion learning is needed to perform emotion classification to judge the emotional tendency of people who post dynamics on social networks. However, most of the existing multi-modal emotion learning methods cannot obtain too fine-grained information from pictures, and the connection between pictures and texts is relatively weak, resulting in low accuracy of multi-modal emotion classification. For this reason, the embodiments of the present invention provide a cross-modal emotion analysis method and device integrating a knowledge graph, which can perform fast and accurate emotion classification on multi-modal information at a relatively low cost.
[0045] As Figure 1 and Figure 2 shown, the cross-modal emotion analysis method integrating a knowledge graph provided by the embodiment includes the following steps:
[0046] Step 1, use a pre-trained model to extract image captions in text form and a first multi-modal knowledge graph at the image level from the image.
[0047] In the embodiment, the pre-trained model is used to obtain the fine-grained feature representations of two parts from the picture. Among them, one part of the fine-grained feature representation is the image caption in text form, and the other part of the fine-grained feature representation is the first multi-modal knowledge graph at the image level. Specifically, the pre-trained model adopts a picture-to-text model. After extracting the global features from the image, the global features are converted into the image caption in text form. Among them, the picture-to-text model can be the NIC model disclosed in the literature [Vinyals O, Toshev A, Bengio S, et al. Show and tell: A neural image caption generator [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2015: 3156-3164.], or it can also be the DETR model disclosed in the literature [Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. arXiv:2005.12872 [cs] (May 2020). arXiv:2005.12872 [cs]].
[0048] Specifically, the pre-trained model adopts a picture scene graph extraction model to convert the picture into the first multi-modal knowledge graph with structured information. Among them, the picture scene graph extraction model can be the TDE model disclosed in the literature [Tang, K., Niu, Y., Huang, J., Shi, J., Zhang, H.: Unbiased scene graph generation from biased training. In: CVPR. pp. 3713-3722. Computer Vision Foundation / IEEE (2020)], or it can also be the VCTree model disclosed in the literature [K. Tang, H. Zhang, B. Wu, W. Luo, and W. Liu. Learning to compose dynamic tree structures for visual contexts. In CVPR, 2019.].
[0049] The first multi-modal knowledge graph contains two forms of triples. One is the triple relationship between entities, for example, (woman, wearing, jacket), (woman, next to, table); the other is the triple relationship between a sub-image and an entity, where the sub-image is an image patch segmented from the input image, for example, (woman, in the image, sub-image ), the head entity is a specific entity, and the tail entity is the feature representation of the image.
[0050] Step 2: Extract the target subject that needs sentiment analysis from the descriptive text and construct a mask template for the target subject.
[0051] In the embodiment, the target subject that needs sentiment analysis is extracted from the descriptive text and a mask template is constructed for the target subject, such as "the target entity is <mask>The mask template of “.”
[0052] Step 3: Convert the description text into a knowledge graph at the text level through a pre-trained model.
[0053] In the embodiment, the description text is converted into a knowledge graph at the text level through another pre-trained model. Specifically, the pre-trained model adopts a text scene graph extraction model to convert the description text into a knowledge graph with structured information, where the knowledge graph establishes the relationships between key entities in the text. Among them, the text scene graph extraction model can be the UniVSE model disclosed in the literature Wu H, Mao J, Zhang Y, et al. Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 6609-6618, or it can also be the VSE++ model disclosed in the literature F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler. VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In Proceedings of British Machine Vision Conference (BMVC), 2018. 2, 3, 5, 6.
[0054] Step 4: Expand the first multimodal knowledge graph and the knowledge graph into a second multimodal knowledge graph that transmits the connection between images and texts through entity node alignment.
[0055] In the embodiment, through entity node alignment, the first multimodal knowledge graph and the knowledge graph are expanded into a second multimodal knowledge graph that contains rich information and transmits the connection between images and texts. Specifically, the fusion process is as follows: When the text semantic similarity between two nodes in the first multimodal knowledge graph and the knowledge graph is greater than the set threshold, the two nodes are merged into one new node, thereby obtaining a multimodal knowledge graph that transmits the connection between images and texts. Among them, the set threshold is preferably 0.9.
[0056] Step 5: Serialize and flatten the triples in the second multimodal knowledge graph to obtain serialized triples, and use a visualization matrix to supplement the structural information of the serialized triples.
[0057] In the embodiment, during the process of serializing and flattening the triples in the second multi-modal knowledge graph, the head entity, relationship, and tail entity of each triple are concatenated with commas to form a serialized triple in the form of a text sequence, and special flag bits are used between different triples. <ts>"Separation. For the sub-images in the second multi-modal knowledge graph, use a special flag bit " " to indicate. During subsequent encoding, use a frozen image encoder to extract features, thus transforming a multi-modal knowledge graph with structured information into a flattened sequence.
[0058] In the embodiment, as Figure 3 shown, a visualization matrix is also established to represent the correlation between elements in the knowledge graph sequence, that is, by defining the correlation between elements in the serialized triple, the supplementation of structured information is realized. Specifically, the definition rules of the correlation are: (1) In the serialized triple input to the encoder, the elements in the same triple are visible to each other, the shared entities in each triple are visible to each other, and the remaining triples are invisible, which reduces the noise from irrelevant information and ensures that the implicit information between entities is modeled; (2) The descriptive text, image captions, and other special markers in the input encoder should be visible to each other so that the text information can interact with the triple information extracted from the multi-modal knowledge graph.
[0059] Step 6, using
[0060] In the embodiment, the specific process of using the sentiment analysis model for sentiment analysis is as follows:
[0061] (a) The encoder combines the visualization matrix to encode the serialized triple, descriptive text, and image caption to obtain output features. Specifically, each part of the serialized triple, descriptive text, and image caption is converted into an embedding vector, and the embedding vector is then input into the encoder. The embedding vector includes feature embedding, position embedding, and type embedding. Among them, for the text marked by the text flag bit, it is encoded by the text encoder, and for the sub-image marked by the special flag " " it is encoded by the image encoder. The visualization matrix encoding contains the correlation between information such as triples, descriptive text, and image captions. Therefore, the visualization matrix makes the word embedding of a word only come from the relevant context guided by the matrix, and the irrelevant words do not affect each other.
[0062] (b) The decoder decodes the mask position representation according to the output features of the encoder and the mask template. Specifically, the output features of the encoder and the mask template are input into the decoder at the same time, and the decoder obtains the embedding vector representation of the mask position through decoding prediction.
[0063] (d) The fully connected layer performs connection mapping on the mask position representation to obtain the main body sentiment representation. Specifically, the following formula is used for the main body sentiment classification prediction:
[0064] p(y||H [m] ) = softmax(θ Linear Dropout(H [m] ))
[0065] where H [m] represents the masked position <mask>The eigenvector representation of “, θ Linear Dropout(·) represents a learnable non - linear layer, and softmax(·) represents an activation function. p(y∣∣H [m] ) represents the probability distribution of predicting the sentiment classification label y based on H [m]
[0066] It should also be noted that the sentiment analysis model needs to be trained before being applied. The specific training process is as follows: The hyperparameters of the sentiment analysis model and the hyperparameters of the loss function are trained on the training dataset. The trained sentiment analysis model performs multi - modal sentiment classification prediction on unseen data samples.
[0067] The cross - modal sentiment analysis method integrating the knowledge graph provided in the above - mentioned embodiment can well integrate the information of pictures and texts in the form of a knowledge graph, better capture the fine - grained information of images, and further strengthen the connection between images, texts, and subjects. It has good practical value in multi - modal sentiment classification tasks.
[0068] Based on the same inventive concept, as Figure 4 shown, the embodiment also provides a cross - modal sentiment analysis device integrating a knowledge graph, including a first multi - modal knowledge graph construction module, a mask template construction module, a knowledge graph construction module, a second multi - modal knowledge graph construction module, a structure information supplement module, and a sentiment analysis module;
[0069] Among them, the first multi - modal knowledge graph construction module is used to extract image captions in text form and the first multi - modal knowledge graph at the image level from images using a pre - trained model; the mask template construction module is used to extract the target subject that needs sentiment analysis from the descriptive text and construct a mask template for the target subject; the knowledge graph construction module is used to convert the descriptive text into a knowledge graph at the text level through a pre - trained model; the second multi - modal knowledge graph construction module is used to expand the first multi - modal knowledge graph and the knowledge graph into a second multi - modal knowledge graph that transmits the connection between images and texts by aligning entity nodes; the structure information supplement module is used to serialize and flatten the triples in the second multi - modal knowledge graph to obtain serialized triples, and supplement the structure information of the serialized triples using a visualization matrix; the sentiment analysis module is used to perform sentiment analysis using a sentiment analysis model including an encoder, a decoder, and a fully - connected layer.
[0070] It should be noted that when the cross-modal sentiment analysis device integrating a knowledge graph performs cross-modal sentiment analysis of the integrated knowledge graph, the above-described examples of the functional modules should be used for illustration. The above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the cross-modal sentiment analysis device integrating a knowledge graph provided in the above embodiments and the embodiments of the cross-modal sentiment analysis method of the integrated knowledge graph belong to the same concept. For the specific implementation process, please refer to the embodiments of the cross-modal sentiment analysis method of the integrated knowledge graph, which will not be elaborated here.
[0071] Based on the same inventive concept, the embodiment also provides a cross-modal sentiment analysis device integrating a knowledge graph, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above cross-modal sentiment analysis method of the integrated knowledge graph is implemented, including the following steps:
[0072] Step 1, use a pre-trained model to extract image captions in text form and a first multi-modal knowledge graph at the image level from the image;
[0073] Step 2, extract the target subject that needs sentiment analysis from the descriptive text and construct a mask template for the target subject;
[0074] Step 3, convert the descriptive text into a knowledge graph at the text level through a pre-trained model;
[0075] Step 4, expand the first multi-modal knowledge graph and the knowledge graph into a second multi-modal knowledge graph that transmits the connection between the image and the text through entity node alignment;
[0076] Step 5, serialize and flatten the triples in the second multi-modal knowledge graph to obtain serialized triples, and use a visualization matrix to supplement the structural information of the serialized triples;
[0077] Step 6, perform sentiment analysis using a sentiment analysis model including an encoder, a decoder, and a fully connected layer.
[0078] In practical applications, the memory can be a volatile memory proximal, such as RAM, or a non-volatile memory, such as ROM, FLASH, floppy disk, mechanical hard disk, etc., or a remote storage cloud. The processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), that is, the steps of the cross-modal sentiment analysis method of the integrated knowledge graph can be implemented through these processors.
[0079] The specific embodiments described above have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.< / mask> < / ts> < / mask> < / mask> < / ts>
Claims
1. A cross-modal sentiment analysis method integrating a knowledge graph, characterized in that Including the following steps: Using a pre-trained model to extract image captions in text form and the first multi-modal knowledge graph at the image level from the image; Extracting the target subject that needs sentiment analysis from the description text and constructing a mask template for the target subject; Converting the description text into a knowledge graph at the text level through a pre-trained model; Expanding the first multi-modal knowledge graph and the knowledge graph into a second multi-modal knowledge graph that transmits the connection between images and texts through entity node alignment, including: when the text semantic similarity between two nodes in the first multi-modal knowledge graph and the knowledge graph is greater than the set threshold, the two nodes are merged into one new node, so as to obtain a multi-modal knowledge graph that transmits the connection between images and texts; Serializing and flattening the triples in the second multimodal knowledge graph to obtain serialized triples, including: concatenating the head entity, relation, and tail entity of each triple with commas to form a serialized triple in the form of a text sequence, and using a special flag " <ts>"Separate. For the sub-images in the second multi-modal knowledge graph, use the special flag bit " " to indicate;< / ts> Using a visualization matrix to supplement the structural information of the serialized triples, where the visualization matrix realizes the supplement of structural information by defining the correlation between the elements in the serialized triples; the definition rule is: in the serialized triples input to the encoder, the elements in the same triple are visible to each other, the shared entities in each triple are visible to each other, and the rest of the triples are invisible; the description text, image captions, and other special tokens in the encoder should be visible to each other; Performing sentiment analysis using a sentiment analysis model including an encoder, a decoder, and a fully connected layer. Specifically: the encoder combines the visualization matrix to encode the serialized triples, the description text, and the image captions to obtain output features, the decoder decodes the mask position representation based on the output features of the encoder and the mask template, and the fully connected layer performs connection mapping on the mask position representation to obtain the subject sentiment representation, which is expressed by the formula: p(y∣∣H [m] ) = softmax(θ Linear Dropout(H [m] )) Among them, H [m] represents the mask position <mask>The eigenvector representation of “, θ Linear Dropout(·) represents a learnable non - linear layer, and softmax(·) represents an activation function. p(y∣∣H [m] ) represents the probability distribution of predicting the sentiment classification label y based on H [m] < / mask> 2. The cross-modal sentiment analysis method integrating a knowledge graph according to claim 1, characterized in that The pre-trained model uses an image-to-text model to convert the global features extracted from the image into image captions in text form after extracting the global features from the image; The pre-trained model uses a picture scene graph extraction model to convert the picture into a first multi-modal knowledge graph with structural information. Among them, the first multi-modal knowledge graph contains two forms of triples, one is the triple relationship between entities and entities, and the other is the triple relationship between sub-images and entities. The sub-image is an image block segmented from the input image; 3. The cross-modal sentiment analysis method integrating a knowledge graph according to claim 1, wherein The pre-trained model uses a text scene graph extraction model to convert the description text into a knowledge graph with structural information, where the knowledge graph establishes the relationship between key entities in the text; 4. The cross-modal sentiment analysis method integrating a knowledge graph according to claim 1, wherein The set threshold is 0.9; 5. The cross-modal sentiment analysis method integrating a knowledge graph according to claim 1, characterized in that The encoder combines the visualization matrix to encode the serialized triples, the description text, and the image captions to obtain output features, including: Converting each part of the serialized triples, the description text, and the image captions into embedding vectors and inputting the embedding vectors into the encoder, where the embedding vectors include feature embeddings, position embeddings, and type embeddings; In the encoder, the text marked by the text flag bit is encoded by the text encoder, and the sub-image marked by the special flag " " is encoded by the image encoder.
6. A cross-modal sentiment analysis device integrating a knowledge graph, characterized in that, Including a first multi-modal knowledge graph construction module, a mask template construction module, a knowledge graph construction module, a second multi-modal knowledge graph construction module, a structural information supplement module, and a sentiment analysis module; The first multi-modal knowledge graph construction module is used to extract image captions in text form and the first multi-modal knowledge graph at the image level from the image using a pre-trained model; The mask template construction module is used to extract the target subject that needs sentiment analysis from the description text and construct a mask template for the target subject; The knowledge graph construction module is used to convert the description text into a knowledge graph at the text level through a pre-trained model; The second multi-modal knowledge graph construction module is used to expand the first multi-modal knowledge graph and the knowledge graph into a second multi-modal knowledge graph that transmits the connection between images and texts by aligning entity nodes, including when the text semantic similarity between two nodes in the first multi-modal knowledge graph and the knowledge graph is greater than a set threshold, the two nodes are merged into one new node, so as to obtain a multi-modal knowledge graph that transmits the connection between images and texts; The structure information supplement module is used to serialize and flatten the triples in the second multi-modal knowledge graph to obtain serialized triples, including: forming a serialized triple in the form of a text sequence by connecting the head entity, relation, and tail entity of each triple with commas, and using a special flag " <ts>"Separation. For the sub-images in the second multi-modal knowledge graph, use the special flag bit " " to indicate;< / ts> The structure information supplement module is further used to supplement the structure information of the serialized triples by using a visualization matrix, where the visualization matrix realizes the supplement of the structured information by defining the correlation between the elements in the serialized triples; the definition rule is: in the serialized triples input to the encoder, the elements in the same triple are visible to each other, the shared entities in each triple are visible to each other, and the remaining triples are invisible; the description text, image captions, and other special markers in the input encoder should be visible to each other; The sentiment analysis module is used to perform sentiment analysis by using a sentiment analysis model including an encoder, a decoder, and a fully connected layer. Specifically: the encoder combines the visualization matrix to encode the serialized triples, the description text, and the image captions to obtain output features, the decoder decodes the mask position representation according to the output features of the encoder and the mask template, and the fully connected layer performs connection mapping on the mask position representation to obtain the subject sentiment representation, which is expressed by the formula: p(y||H [m] ) = softmax(θ Linear Dropout(H [m] )) Among them, H [m] represents the mask position" <mask>The eigenvector representation of “, θ Linear Dropout(·) represents a learnable non-linear layer, and softmax(·) represents an activation function. p(y∣∣H [m] ) represents based on H [m] The probability distribution of predicting the sentiment classification label y.< / mask> 7. A cross-modal sentiment analysis device integrating a knowledge graph, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the cross-modal sentiment analysis method of the fusion knowledge graph according to any one of claims 1-6.
Citation Information
Patent Citations
Method for constructing and displaying multi-modal emotion knowledge graph
CN112417172A
Text-image enhanced multi-modal knowledge graph embedding method
CN115099409A