A fine-grained sentiment analysis method, system, terminal and medium based on deep learning
Through a fine-grained sentiment analysis method based on deep learning, using hierarchical feature extraction and cross-modal feature fusion, the problem of ignoring the deep semantic connection between pictures and text in the existing technology is solved, and higher sentiment analysis accuracy is achieved.
Patent Information
- Application Number
- CN202411864173.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing multimodal sentiment analysis method based on machine learning ignores the deep semantic connection between pictures and text, and cannot make full use of the complementary information between pictures and text, resulting in low accuracy of sentiment analysis.
A fine-grained sentiment analysis method based on deep learning is adopted, and through hierarchical feature extraction and cross-modal feature fusion, multi-label classification is used to achieve fine-grained sentiment analysis.
By making full use of complementary information between pictures and text, the classification accuracy of sentiment analysis is improved, and the sentiment analysis task is accurate from binary classification to specific emotional expression.
Smart Images

Figure CN119322985B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sentiment analysis, and specifically relates to a fine-grained sentiment analysis method, system, terminal and medium based on deep learning. Background Art
[0002] Currently, information is widely disseminated in the form of a combination of pictures and texts, and the content of the combination of pictures and texts contains rich emotional information. Traditional sentiment analysis methods mainly rely on text data to identify emotional tendencies through keyword matching, sentiment dictionaries, or machine learning models. However, although the existing multimodal sentiment analysis methods based on machine learning have improved the generalization and accuracy of the analysis, most of them use simple feature splicing or weighted averaging, ignoring the deep semantic connection between pictures and texts, and failing to fully utilize the complementary information between pictures and texts, resulting in low accuracy of sentiment analysis. Summary of the invention
[0003] To solve the above problems, the present invention provides a fine-grained sentiment analysis method, system, terminal and medium based on deep learning, which performs hierarchical feature extraction on semantic information, and performs cross-modal feature fusion on the hierarchical semantic information. Based on the multi-label cross entropy loss function, the fused cross-modal features are subjected to multi-label classification to realize fine-grained sentiment analysis, make full use of the complementary information between pictures and texts, and improve classification accuracy.
[0004] In a first aspect, the technical solution of the present invention provides a fine-grained sentiment analysis method based on deep learning, comprising the following steps:
[0005] Step 1, a step of training a fine-grained sentiment analysis model using image-text pairs with sentiment labels;
[0006] Input the text file and the sentiment label file into the initial text semantic information extraction model to extract different levels of initial text semantic information;
[0007] Input the image file into the initial image semantic information extraction model to extract initial image semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial image semantic information is the same;
[0008] Input the initial text semantic information and initial image semantic information of the same layer into a cross-modal contextual attention network to obtain cross-modal fusion semantic information;
[0009] The cross-modal fusion semantic information of each layer is input into the discriminator, and multi-label classification is performed based on the multi-label cross entropy loss function;
[0010] Step 2, using the trained fine-grained sentiment analysis model to perform fine-grained sentiment analysis on the image-text pair to be analyzed;
[0011] Get the image-text pair to be analyzed;
[0012] The image-text pair to be analyzed is input into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0013] In an optional implementation, the text file and the sentiment tag file are input into the initial text semantic information extraction model to extract different levels of initial text semantic information, specifically including:
[0014] Break the text content in a text file into a sequence of words , Indicates the number of words;
[0015] Detection Whether it exceeds the preset maximum value;
[0016] If so, the excess part is cut off, otherwise, the insufficient part is replaced by a specific symbol;
[0017] The text files and sentiment label files after the text content is decomposed are input into the Bert model for encoding, and the initial text semantic information at different levels is extracted. The initial text semantic information is in vector form and is expressed as ,in correspond Encoded features.
[0018] In an optional implementation, the image file is input into the initial image semantic information extraction model to extract initial image semantic information at different levels, specifically including:
[0019] Convert pictures to standard size images;
[0020] Split the image of standard size into several image blocks;
[0021] Flatten all image patches along the RGB channels into feature vectors of a certain size;
[0022] Use the ViT model to perform a multi-head attention mechanism on the feature vectors of all image blocks in the entire image to obtain feature vector sets at different levels , that is, different levels of initial image semantic information.
[0023] In an optional embodiment, the Bert model and the ViT model are both 12-layer network structures, and the Bert model and the ViT model are divided into three parts in groups of four layers. The first part of the Bert model extracts the first shallow initial text semantic information, the second part extracts the second shallow initial text semantic information, and the third part extracts the deep initial text semantic information. The first part of the ViT model extracts the first shallow initial image semantic information, the second part extracts the second shallow initial image semantic information, and the third part extracts the deep initial image semantic information.
[0024] In an optional implementation, the initial text semantic information and the initial image semantic information of the same layer are input into a cross-modal contextual attention network to obtain fused semantic information, specifically including:
[0025] Pre-constructing a cross-modal contextual attention network, the constructed cross-modal contextual attention network includes a first Transformer unit and a second Transformer unit, and the two Transformer units are connected to each other;
[0026] Input the initial text semantic information into the first Transformer unit to obtain the intermediate text semantic information, and use the intermediate text semantic information as a supplement to the second Transformer unit to obtain the context text semantic information;
[0027] Input the initial image semantic information into the second Transformer unit to obtain the intermediate image semantic information, and use the intermediate image semantic information as a supplement to the first Transformer unit to obtain the context image semantic information;
[0028] The contextual text semantic information and the contextual image semantic information are combined to obtain cross-modal fusion semantic information.
[0029] In an optional implementation, the initial text semantic information is input into the first Transformer unit to obtain the intermediate text semantic information, and the intermediate text semantic information is used as a supplement to the second Transformer unit to obtain the context text semantic information, specifically including:
[0030] Step 1: The initial text semantic information is Calculate the attention weight matrix within the text modality , expressed by the following formula:
[0031] ;
[0032] In the formula, and Represents the cross-modal attention network of the first Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result;
[0033] Step 2: and The vectors are multiplied, and the initial text semantic information is added and normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula:
[0034] ;
[0035] Step 3: Transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the intermediate text semantic information, is expressed by the following formula:
[0036] ;
[0037] Step 4: The initial image semantic information and the intermediate text semantic information are Calculate the attention weight matrix within the image modality , expressed by the following formula:
[0038] ;
[0039] In the formula, and Represents the cross-modal attention network of the second Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result;
[0040] Step 5: and The vectors are multiplied, and the initial image semantic information is added and then normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula:
[0041] ;
[0042] Step 6: Transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the contextual text semantic information, is expressed by the following formula:
[0043] .
[0044] In an optional implementation, the context text semantic information and the context image semantic information are combined to obtain cross-modal fusion semantic information, specifically including:
[0045] Configuration control parameters ;
[0046] The cross-modal fusion semantic information is calculated by the following formula ,
[0047] ;
[0048] In the formula, is the context text semantic information, It is the context image semantic information.
[0049] In a second aspect, the technical solution of the present invention provides a fine-grained sentiment analysis system based on deep learning, comprising:
[0050] A model training module for training a fine-grained sentiment analysis model using image-text pairs with sentiment labels;
[0051] A sentiment analysis module, which is used to perform fine-grained sentiment analysis on the image-text pairs to be analyzed using the trained fine-grained sentiment analysis model;
[0052] The model training module includes:
[0053] An initial text language information extraction unit is used to input a text file and a sentiment tag file into an initial text semantic information extraction model to extract initial text semantic information at different levels;
[0054] An initial picture semantic information extraction unit is used to input the picture file into the initial picture semantic information extraction model to extract initial picture semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial picture semantic information is the same;
[0055] A cross-modal fusion semantic information acquisition unit, used for inputting the initial text semantic information and the initial image semantic information of the same layer into a cross-modal context attention network to obtain cross-modal fusion semantic information;
[0056] The multi-label classification unit is used to input the cross-modal fusion semantic information of each layer into the discriminator and perform multi-label classification based on the multi-label cross entropy loss function;
[0057] The sentiment analysis module includes:
[0058] A data acquisition unit to be processed, used to acquire a picture-text pair to be analyzed;
[0059] The sentiment analysis unit is used to input the image-text pair to be analyzed into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0060] In a third aspect, the technical solution of the present invention provides a terminal, including:
[0061] A memory for storing a fine-grained sentiment analysis program based on deep learning;
[0062] A processor is used to implement the steps of the fine-grained sentiment analysis method based on deep learning as described in any of the above items when executing the fine-grained sentiment analysis program based on deep learning.
[0063] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a fine-grained sentiment analysis program based on deep learning is stored. When the fine-grained sentiment analysis program based on deep learning is executed by a processor, the steps of the fine-grained sentiment analysis method based on deep learning as described in any of the above items are implemented.
[0064] The present invention provides a deep learning-based fine-grained sentiment analysis method, system, terminal and medium, which have the following beneficial effects compared with the prior art: first, hierarchical feature extraction of semantic information is performed based on a neural network model, and cross-modal feature fusion of the hierarchical semantic information is performed, and then multi-label classification is performed on the fused cross-modal features based on a multi-label cross entropy loss function to achieve fine-grained sentiment analysis, and the sentiment analysis task is refined from binary classification to specific sentiment expression, and the contextual attention mechanism and hierarchical model are used to fully utilize the complementary information between images and texts to improve classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0066] Figure 1 It is a flowchart of a fine-grained sentiment analysis method based on deep learning provided by an embodiment of the present invention.
[0067] Figure 2 It is a schematic diagram of the structure of the fine-grained sentiment analysis model.
[0068] Figure 3A It is a schematic diagram of the cross-modal semantic information fusion architecture.
[0069] Figure 3B It is a schematic diagram of the contextual attention network fusion in the cross-modal semantic information fusion architecture.
[0070] Figure 4 It is a schematic block diagram of the structure of a fine-grained sentiment analysis system based on deep learning provided by an embodiment of the present invention.
[0071] Figure 5 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0072] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0074] The key terms appearing in the present invention are explained below.
[0075] Bert model: A pre-trained language model used to process text data, breaking text into sequences of words or characters and encoding them to obtain text semantic information.
[0076] <pad>: Fill symbol, used to fill or truncate text when the text length is less than or exceeds the set length.
[0077] Vision Transform model: ViT model is used to process image data. It preprocesses the image, splits it into image blocks, flattens it into feature vectors, and then uses a multi-head self-attention mechanism to obtain a vector set to obtain image semantic information.
[0078] Multi-head self-attention: An extended mechanism that uses multiple different self-attention modules (i.e., multiple heads) in parallel to allow the model to dynamically focus on other elements in the sequence from multiple different representation subspaces when processing each element of a sequence, thereby more comprehensively capturing the information of the input sequence and further improving the ability to capture long-distance dependencies and different feature dimensions.
[0079] Q matrix (query matrix), K matrix (key matrix), V matrix (value matrix): are the three key matrices in the self-attention mechanism.
[0080] Cross-modal information fusion: A method that aims to integrate data information from different modalities to achieve more comprehensive and accurate analysis and understanding. By fusing data from multiple modalities such as images and text, we can fully utilize the advantages of each modality, make up for the limitations of a single modality, and enable the model to better handle the relationship between cross-modal data.
[0081] In view of the problem that the prior art ignores the contribution of different levels of semantics of text and image to the analysis results, resulting in inaccurate sentiment analysis results, an embodiment of the present invention provides a fine-grained sentiment analysis method based on deep learning. The method uses a hierarchical pre-training model to improve the utilization of semantic information of each modality, and then captures cross-modal information between modalities by combining a contextual attention neural network model based on the Transformer model. Finally, a multi-head discriminator is used to perform specific downstream tasks of sentiment analysis, thereby obtaining a model that can accurately classify the sentiment of multi-modal information at a fine granularity, thereby avoiding the shortcomings of traditional methods in utilizing semantic information and capturing cross-modal information, and effectively improving the accuracy of sentiment analysis.
[0082] Figure 1 : is a flow chart of a fine-grained sentiment analysis method based on deep learning provided by an embodiment of the present invention. Figure 1 The execution subject may be a fine-grained sentiment analysis system based on deep learning. The fine-grained sentiment analysis method based on deep learning provided in the embodiment of the present invention is executed by a computer device, and accordingly, the fine-grained sentiment analysis system based on deep learning runs in the computer device. According to different requirements, the order of the steps in the flowchart can be changed, and some can be omitted.
[0083] like Figure 1 As shown, the method includes the following steps.
[0084] S1, the step of training a fine-grained sentiment analysis model using image-text pairs with sentiment labels.
[0085] S1.1, input the text file and the sentiment label file into the initial text semantic information extraction model to extract the initial text semantic information at different levels.
[0086] S1.2, inputting the image file into the initial image semantic information extraction model to extract initial image semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial image semantic information is the same.
[0087] S1.3, the initial text semantic information and the initial image semantic information of the same layer are input into a cross-modal contextual attention network to obtain cross-modal fusion semantic information.
[0088] S1.4, the cross-modal fusion semantic information of each layer is input into the discriminator, and multi-label classification is performed based on the multi-label cross entropy loss function.
[0089] S2, the step of using the trained fine-grained sentiment analysis model to perform fine-grained sentiment analysis on the image-text pairs to be analyzed.
[0090] S2.1, obtain the image-text pair to be analyzed.
[0091] S2.2, input the image-text pair to be analyzed into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0092] The deep learning-based fine-grained sentiment analysis method of this embodiment includes two stages: training and execution. In the training stage, the fine-grained sentiment analysis model is trained. In the execution stage, the trained fine-grained sentiment analysis model is used to analyze the data to be analyzed to obtain the sentiment classification in the data to be analyzed.
[0093] In this embodiment, during the training phase, the model receives image-text pairs that have been marked with specific emotional tags as input. First, the image-text pairs are split to generate different levels of image semantic information and text semantic information for each image and text. Then, after performing cross-modal information fusion operations on these semantic information, they are sent to the discriminator for cross-modal analysis to obtain the emotional information represented by the image-text pair, which may involve fine-grained emotions such as happiness, sadness, anger, and irony.
[0094] Figure 2 This is a schematic diagram of the fine-grained sentiment analysis model structure. Figure 2 The structure of the training of the fine-grained sentiment analysis model is described in detail. The training process includes the following steps.
[0095] Step 1: Input the text file and the sentiment label file into the initial text semantic information extraction model to extract initial text semantic information at different levels, and input the image file into the initial image semantic information extraction model to extract initial image semantic information at different levels.
[0096] In an optional implementation, the initial text semantic information extraction model is the Bert model, the initial image semantic information extraction model is the ViT model, the text file and the txt file representing the specific emotion are input into the Bert model, and the image file is input into the VisionTransform model. Since both the Bert model and the VisionTransform model have 12-layer network structures, the two models are divided into three parts in groups of four layers, and the information contained in the hidden layers of the first two parts, which are groups of four and eight layers, is used as shallow semantic information, and the semantic information output by the overall model in the last part is used as deep semantic information, which is sent to three different cross-modal context attention networks.
[0097] In the sentiment analysis model, it is necessary to convert text information and image information into vector form and input them into the neural network, and output the semantic features of the text and image. However, since the pre-trained model is a black box for users, users only need to pay attention to the results of the model analysis without paying attention to the process of acquiring semantic features. This results in the semantic information parsed by the model not being fully utilized, and the progressive relationship of semantic information from shallow to deep is not reflected in the subsequent cross-modal fusion module, affecting the performance and accuracy of the overall model. Therefore, this embodiment uses a strategy for hierarchical extraction of semantic information, including the following steps.
[0098] Step 1.1: Decompose the text content in the text file into a sequence of words , Indicates the number of words.
[0099] Step 1.2, Detection Whether it exceeds the preset maximum value.
[0100] Step 1.3, if yes, cut off the excess part, otherwise, replace the insufficient part with a specific symbol.
[0101] Step 1.4: Input the text file and sentiment label file after the text content is decomposed into the Bert model for encoding, and extract the initial text semantic information at different levels. The initial text semantic information is in vector form and is expressed as ,in correspond Encoded features.
[0102] For example, The maximum value is set to 100, the excess will be truncated, and the insufficient part will be used as a symbol <pad>replace.
[0103] Step 1.5, convert the image to a standard size.
[0104] Step 1.6, divide the image of standard size into several image blocks.
[0105] Step 1.7, flatten all image patches along the RGB channels into feature vectors of a certain size.
[0106] Step 1.8: Use the ViT model to perform a multi-head attention mechanism on the feature vectors of all image blocks of the entire image to obtain feature vector sets at different levels. , that is, different levels of initial image semantic information.
[0107] For example, for a given image content, the image is first preprocessed to convert it to a standard size of 3×224×22, then it is split into 16×16 image blocks, and the entire image is flattened along the RGB channel into a 16×16×3 feature vector. After a multi-head self-attention mechanism is performed on all image blocks of the entire image, a vector set is obtained. .
[0108] Step 2: Input the initial text semantic information and initial image semantic information of the same layer into a cross-modal contextual attention network to obtain cross-modal fusion semantic information.
[0109] In this embodiment, after obtaining the feature vectors of text and picture respectively in step 1, the cross-modal semantic information at each level is fused. The picture-text pair is preprocessed and combined with the specific emotion label obtained by manual annotation to form the initial vector representation of the picture-text pair at the input end of the model, and the semantic information of different levels represented by the picture-text pair is input into a contextual attention neural network model using Transformer for training and evaluation. The model uses two Transformers to capture the cross-modal semantic information of text with picture as the context and picture with text as the context respectively.
[0110] Figure 3A It is a schematic diagram of the cross-modal semantic information fusion architecture. Figure 3B This is a schematic diagram of the context attention network fusion in the cross-modal semantic information fusion architecture. The cross-modal semantic fusion specifically includes the following steps.
[0111] Step 2.1, pre-construct a cross-modal contextual attention network, the constructed cross-modal contextual attention network includes a first Transformer unit and a second Transformer unit, and the two Transformer units are connected to each other.
[0112] In this embodiment, the first Transformer unit is Figure 3B The context attention network A on the left side of the middle, the second Transformer unit is Figure 3B Middle right context attention network B.
[0113] Step 2.2, input the initial text semantic information into the first Transformer unit to obtain the intermediate text semantic information, and use the intermediate text semantic information as a supplement to the second Transformer unit to obtain the context text semantic information.
[0114] Step 2.2.1: Transform the initial text semantic information through Calculate the attention weight matrix within the text modality , expressed by the following formula:
[0115] ;
[0116] In the formula, and Represents the cross-modal attention network of the first Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result.
[0117] Step 2.2.2, and The vectors are multiplied, and the initial text semantic information is added and normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula:
[0118] .
[0119] Step 2.2.3, transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the intermediate text semantic information, is expressed by the following formula:
[0120] .
[0121] The above steps 2.2.1 to 2.2.3 convert the text feature vector , image feature vector Input the context Transformer units on the left and right respectively to obtain the unimodal semantic information of the text.
[0122] Step 2.2.4: The initial image semantic information and the intermediate text semantic information are Calculate the attention weight matrix within the image modality , expressed by the following formula:
[0123] ;
[0124] In the formula, and Represents the cross-modal attention network of the second Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result.
[0125] Step 2.2.5, and The vectors are multiplied, and the initial image semantic information is added and normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula:
[0126] .
[0127] Step 2.2.6, transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the contextual text semantic information, is expressed by the following formula:
[0128] .
[0129] After obtaining text semantic information in the above steps 2.2.4 to 2.2.6, it is input into the Transformer unit on the right, and the image-related information is extracted from the learned text semantic information to supplement the image semantic information. It should be noted that when calculating the cross-modal attention matrix of text and image, the dot product operation should be performed on the image information and the text semantic information to output the cross-modal attention matrix. The subsequent operation process is the same as the operation process of text semantic information, and the final result is the context semantic information.
[0130] Step 2.3, input the initial image semantic information into the second Transformer unit to obtain the intermediate image semantic information, and use the intermediate image semantic information as a supplement to the first Transformer unit to obtain the context image semantic information.
[0131] The above step 2.2 analyzes the text as a single-modal input and the image as context information to output contextual text semantic information (i.e., text feature vector ). This step 2.3 uses the image as the single-modal input and the text as the context information, and outputs the context image semantic information through the same steps. . These two vectors are then concatenated to represent multimodal fusion semantic information.
[0132] In step 2.4, the context text semantic information and the context image semantic information are combined to obtain cross-modal fusion semantic information.
[0133] Step 2.4.1, configure control parameters .
[0134] Step 2.4.2, calculate the cross-modal fusion semantic information by the following formula ,
[0135] ;
[0136] In the formula, is the context text semantic information, It is the context image semantic information.
[0137] This embodiment takes into account that the degree of dependence of different levels of single-modal semantic information on another modality as context for feature fusion may be different, and uses a learnable semantic representation parameter Control the representation of contextual semantic information and finally obtain multimodal fusion semantic information.
[0138] Step 3: Input the cross-modal fusion semantic information of each layer into the discriminator and perform multi-label classification based on the multi-label cross entropy loss function.
[0139] In this embodiment, the fine-grained sentiment analysis task is modeled as a multi-label classification task in the model and is implemented by using a multi-label cross entropy loss function. Specifically, the obtained multimodal fusion semantic information is input into a discriminator composed of multiple linear layers for cross-modal analysis. The discriminator includes and Two activation functions. The multi-label cross entropy loss function is then used to calculate the classification result loss of the multi-label sentiment analysis.
[0140]
[0141]
[0142] in, , and , They represent the parameters of the first and second linear layers respectively, express Activation function, , Represent the sample index and label index respectively, That means The sample The true value of each label (0 or 1).
[0143] The above describes in detail an embodiment of a fine-grained sentiment analysis method based on deep learning. Based on the fine-grained sentiment analysis method based on deep learning described in the above embodiment, an embodiment of the present invention also provides a fine-grained sentiment analysis system based on deep learning corresponding to the method.
[0144] Figure 4 It is a schematic block diagram of the structure of a fine-grained sentiment analysis system based on deep learning provided by an embodiment of the present invention. In this embodiment, the fine-grained sentiment analysis system based on deep learning 400 can be divided into multiple functional modules according to the functions it performs. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory.
[0145] The model training module 410 is used to train a fine-grained sentiment analysis model using image-text pairs with sentiment labels.
[0146] The sentiment analysis module 420 is used to perform fine-grained sentiment analysis on the image-text pair to be analyzed using the trained fine-grained sentiment analysis model.
[0147] The model training module 410 includes:
[0148] The initial text language information extraction unit 410_1 is used to input the text file and the sentiment tag file into the initial text semantic information extraction model to extract different levels of initial text semantic information.
[0149] The initial picture semantic information extraction unit 410_2 is used to input the picture file into the initial picture semantic information extraction model to extract initial picture semantic information of different levels; the number of levels of the extracted initial text semantic information and the initial picture semantic information is the same.
[0150] The cross-modal fusion semantic information acquisition unit 410_3 is used to input the initial text semantic information and the initial image semantic information of the same layer into a cross-modal context attention network to obtain cross-modal fusion semantic information.
[0151] The multi-label classification unit 410_4 is used to input the cross-modal fusion semantic information of each layer into the discriminator and perform multi-label classification based on the multi-label cross entropy loss function.
[0152] The sentiment analysis module 420 includes:
[0153] The to-be-processed data acquisition unit 420_1 is used to acquire a to-be-analyzed picture-text pair.
[0154] The sentiment analysis unit 420_2 is used to input the image-text pair to be analyzed into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0155] The deep learning-based fine-grained sentiment analysis system of this embodiment is used to implement the aforementioned deep learning-based fine-grained sentiment analysis method. Therefore, the specific implementation method of the system can be seen in the embodiment section of the deep learning-based fine-grained sentiment analysis method in the previous text. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part, which will not be introduced in detail here.
[0156] In addition, since the deep learning-based fine-grained sentiment analysis system of this embodiment is used to implement the aforementioned deep learning-based fine-grained sentiment analysis method, its function corresponds to that of the above-mentioned method and will not be repeated here.
[0157] Figure 5 A schematic diagram of the structure of a terminal 500 provided in an embodiment of the present invention includes: a processor 510, a memory 520 and a communication unit 530. The processor 510 is used to implement the following steps when implementing the fine-grained sentiment analysis program based on deep learning stored in the memory 520:
[0158] Step 1, a step of training a fine-grained sentiment analysis model using image-text pairs with sentiment labels;
[0159] Input the text file and the sentiment label file into the initial text semantic information extraction model to extract different levels of initial text semantic information;
[0160] Input the image file into the initial image semantic information extraction model to extract initial image semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial image semantic information is the same;
[0161] Input the initial text semantic information and initial image semantic information of the same layer into a cross-modal contextual attention network to obtain cross-modal fusion semantic information;
[0162] The cross-modal fusion semantic information of each layer is input into the discriminator, and multi-label classification is performed based on the multi-label cross entropy loss function;
[0163] Step 2, using the trained fine-grained sentiment analysis model to perform fine-grained sentiment analysis on the image-text pair to be analyzed;
[0164] Get the image-text pair to be analyzed;
[0165] The image-text pair to be analyzed is input into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0166] The present invention also provides a computer storage medium, wherein the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0167] The computer storage medium stores a deep learning-based fine-grained sentiment analysis program, and when the deep learning-based fine-grained sentiment analysis program is executed by a processor, the following steps are implemented:
[0168] Step 1, a step of training a fine-grained sentiment analysis model using image-text pairs with sentiment labels;
[0169] Input the text file and the sentiment label file into the initial text semantic information extraction model to extract different levels of initial text semantic information;
[0170] Input the image file into the initial image semantic information extraction model to extract initial image semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial image semantic information is the same;
[0171] Input the initial text semantic information and initial image semantic information of the same layer into a cross-modal contextual attention network to obtain cross-modal fusion semantic information;
[0172] The cross-modal fusion semantic information of each layer is input into the discriminator, and multi-label classification is performed based on the multi-label cross entropy loss function;
[0173] Step 2, using the trained fine-grained sentiment analysis model to perform fine-grained sentiment analysis on the image-text pair to be analyzed;
[0174] Get the image-text pair to be analyzed;
[0175] The image-text pair to be analyzed is input into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result.
[0176] The above disclosure is only a preferred embodiment of the present invention, but the present invention is not limited thereto. Any non-creative changes that can be thought of by a person skilled in the art, as well as several improvements and modifications made without departing from the principle of the present invention, should fall within the protection scope of the present invention.< / pad> < / pad>
Claims
1. A fine-grained sentiment analysis method based on deep learning, characterized in that: The following steps are involved: Step 1, a step of training a fine-grained sentiment analysis model using image-text pairs with sentiment labels; (1.1) Inputting the text file and the sentiment label file into the initial text semantic information extraction model to extract different levels of initial text semantic information; (1.2) Inputting the image file into the initial image semantic information extraction model to extract initial image semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial image semantic information is the same; (1.3) Input the initial text semantic information and initial image semantic information of the same layer into a cross-modal contextual attention network to obtain cross-modal fusion semantic information; (1.3.1) Pre-constructing a cross-modal contextual attention network, the constructed cross-modal contextual attention network includes a first Transformer unit and a second Transformer unit, and the two Transformer units are connected to each other; (1.3.2) Input the initial text semantic information into the first Transformer unit to obtain the intermediate text semantic information, and use the intermediate text semantic information as a supplement to the second Transformer unit to obtain the context text semantic information; (1.3.3) Input the initial image semantic information into the second Transformer unit to obtain the intermediate image semantic information, and use the intermediate image semantic information as a supplement to the first Transformer unit to obtain the context image semantic information; (1.3.4) Combine the contextual text semantic information and the contextual image semantic information to obtain cross-modal fusion semantic information; (1.4) Input the cross-modal fusion semantic information of each layer into the discriminator, and perform multi-label classification based on the multi-label cross entropy loss function; Step 2, using the trained fine-grained sentiment analysis model to perform fine-grained sentiment analysis on the image-text pair to be analyzed; (2.1) Obtain the image-text pair to be analyzed; (2.2) Input the image-text pair to be analyzed into the trained fine-grained sentiment analysis model to obtain the sentiment analysis results.
2. The fine-grained sentiment analysis method based on deep learning according to claim 1, characterized in that: Input the text file and sentiment label file into the initial text semantic information extraction model to extract different levels of initial text semantic information, including: Break the text content in a text file into a sequence of words , Indicates the number of words; Detection Whether it exceeds the preset maximum value; If so, the excess part is cut off, otherwise, the insufficient part is replaced by a specific symbol; The text files and sentiment label files after the text content is decomposed are input into the Bert model for encoding, and the initial text semantic information at different levels is extracted. The initial text semantic information is in vector form and is expressed as ,in correspond Encoded features.
3. The fine-grained sentiment analysis method based on deep learning according to claim 2 is characterized in that: Input the image file into the initial image semantic information extraction model to extract different levels of initial image semantic information, including: Convert pictures to standard size images; Split the image of standard size into several image blocks; Flatten all image patches along the RGB channels into feature vectors of a certain size; Use the ViT model to perform a multi-head attention mechanism on the feature vectors of all image blocks in the entire image to obtain feature vector sets at different levels , that is, different levels of initial image semantic information.
4. The fine-grained sentiment analysis method based on deep learning according to claim 3 is characterized in that: Both the Bert model and the ViT model have a 12-layer network structure. The Bert model and the ViT model are divided into three parts with four layers as a group. The first part of the Bert model extracts the first shallow initial text semantic information, the second part extracts the second shallow initial text semantic information, and the third part extracts the deep initial text semantic information. The first part of the ViT model extracts the first shallow initial image semantic information, the second part extracts the second shallow initial image semantic information, and the third part extracts the deep initial image semantic information.
5. The fine-grained sentiment analysis method based on deep learning according to claim 4 is characterized in that: Inputting the initial text semantic information into the first Transformer unit to obtain the intermediate text semantic information, and using the intermediate text semantic information as a supplement to the second Transformer unit to obtain the context text semantic information, specifically including: Step 1: The initial text semantic information is Calculate the attention weight matrix within the text modality , expressed by the following formula: ; In the formula, and Represents the cross-modal attention network of the first Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result; Step 2: and The vectors are multiplied, and the initial text semantic information is added and normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula: ; Step 3: Transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the intermediate text semantic information, is expressed by the following formula: ; Step 4: The initial image semantic information and the intermediate text semantic information are Calculate the attention weight matrix within the image modality , expressed by the following formula: ; In the formula, and Represents the cross-modal attention network of the second Transformer unit , The fully connected layer corresponding to the matrix, Represents the dimension of the vector, used to normalize the dot product result; Step 5: and The vectors are multiplied, and the initial image semantic information is added and then normalized to obtain the feature vector after the attention mechanism. , expressed by the following formula: ; Step 6: Transform the feature vector After a two-layer fully connected neural network After nonlinear transformation and layer normalization, the final text feature vector is obtained. , that is, the contextual text semantic information, is expressed by the following formula: 。 6. The fine-grained sentiment analysis method based on deep learning according to claim 5 is characterized in that: The context text semantic information and the context image semantic information are combined to obtain cross-modal fusion semantic information, including: Configuration control parameters ; The cross-modal fusion semantic information is calculated by the following formula , ; In the formula, is the context text semantic information, It is the context image semantic information.
7. A fine-grained sentiment analysis system based on deep learning, characterized in that: include: A model training module for training a fine-grained sentiment analysis model using image-text pairs with sentiment labels; A sentiment analysis module, which is used to perform fine-grained sentiment analysis on the image-text pairs to be analyzed using the trained fine-grained sentiment analysis model; The model training module includes: An initial text language information extraction unit is used to input a text file and a sentiment tag file into an initial text semantic information extraction model to extract initial text semantic information at different levels; An initial picture semantic information extraction unit is used to input the picture file into the initial picture semantic information extraction model to extract initial picture semantic information of different levels; the number of layers of the extracted initial text semantic information and the initial picture semantic information is the same; A cross-modal fusion semantic information acquisition unit, used for inputting the initial text semantic information and the initial image semantic information of the same layer into a cross-modal context attention network to obtain cross-modal fusion semantic information; The multi-label classification unit is used to input the cross-modal fusion semantic information of each layer into the discriminator and perform multi-label classification based on the multi-label cross entropy loss function; The sentiment analysis module includes: A data acquisition unit to be processed, used to acquire a picture-text pair to be analyzed; A sentiment analysis unit, used to input the image-text pair to be analyzed into the trained fine-grained sentiment analysis model to obtain the sentiment analysis result; Among them, the cross-modal fusion semantic information acquisition unit is specifically used for: Pre-constructing a cross-modal contextual attention network, the constructed cross-modal contextual attention network includes a first Transformer unit and a second Transformer unit, and the two Transformer units are connected to each other; Input the initial text semantic information into the first Transformer unit to obtain the intermediate text semantic information, and use the intermediate text semantic information as a supplement to the second Transformer unit to obtain the context text semantic information; Input the initial image semantic information into the second Transformer unit to obtain the intermediate image semantic information, and use the intermediate image semantic information as a supplement to the first Transformer unit to obtain the context image semantic information; The contextual text semantic information and the contextual image semantic information are combined to obtain cross-modal fusion semantic information.
8. A terminal, characterized in that: include: A memory for storing a fine-grained sentiment analysis program based on deep learning; A processor, used to implement the steps of the fine-grained sentiment analysis method based on deep learning as described in any one of claims 1 to 6 when executing the fine-grained sentiment analysis program based on deep learning.
9. A computer-readable storage medium, characterized in that: The readable storage medium stores a deep learning-based fine-grained sentiment analysis program, which, when executed by a processor, implements the steps of the deep learning-based fine-grained sentiment analysis method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-modal fine-grained sentiment analysis method based on momentum contrast learning
CN117435732A
Cross-modal positive and negative semantic classification method based on text emotion and image content perception
CN118690259A