Text user sentiment analysis method and system based on multiple modes and AI

By constructing a multimodal feature fusion network and a graph neural network-based emotion propagation model, the problem of insufficient interaction between modes in the existing multimodal emotion analysis method is solved, and a more accurate and comprehensive user sentiment analysis is achieved.

CN120216700AInactive Publication Date: 2025-06-27ZHEJIANG SHUXIN NETWORK CO LTD

Patent Information

Application Number
CN202510696691.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal emotion analysis methods lack effective intermodal interaction mechanisms, and cannot fully explore the complementary information and interdependence relationships between different modes, resulting in one-sided and low accuracy of sentiment analysis results.

Method used

By constructing a multimodal feature fusion network, a cross-modal attention mechanism is used to calculate the interaction weights between semantic features, visual features and acoustic features to generate a fusion feature vector. At the same time, an emotion propagation model based on graph neural network was introduced to construct a user's emotional social graph, calculate the intensity of emotional propagation between user nodes, and obtain a user emotion vector containing group emotional interaction characteristics.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of sentiment analysis, can more effectively capture the emotional information expressed by users on social media, considering various modal data such as text, images, audio, and emotional communication characteristics in social networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216700A_ABST
    Figure CN120216700A_ABST
Patent Text Reader

Abstract

The invention provides a text user sentiment analysis method and system based on multiple modes and AI, and relates to the technical field of text analysis, and the method comprises the steps: obtaining and preprocessing text, image and audio information of a user, and extracting semantic, visual and acoustic feature vectors; multi-modal features are fused through a cross-modal attention mechanism; utilizing a graph neural network to construct a user emotion social graph to calculate emotion propagation intensity; and finally, obtaining an analysis result containing emotion category and intensity through an emotion classifier. The emotion state of the user can be comprehensively captured, the emotion analysis accuracy is improved, and complex emotion expression is effectively recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to text analysis technology, and particularly to a method and system for text user sentiment analysis based on multimodality and AI. Background Art

[0002] With the rapid development of social media, users have generated a large amount of multimodal content including text, images, and audio on various platforms. These data contain rich emotional information, and accurate analysis of user emotions has become an important research direction in fields such as social media analysis, marketing, and public opinion monitoring. Traditional sentiment analysis methods mainly rely on single-modal data, such as only analyzing based on text content, and it is difficult to comprehensively capture the complexity of user emotions. With the development of artificial intelligence technology, multimodal sentiment analysis has gradually become a research hotspot, and by integrating information from different modalities, user emotion expressions can be understood more comprehensively.

[0003] Most sentiment analysis methods only focus on single-modal data and cannot effectively fuse multimodal information such as text, images, and audio, resulting in one-sided and inaccurate sentiment analysis results. When users express emotions on social media, they often express them through multiple modalities in coordination, and single-modal analysis is difficult to capture such complex emotional expression methods.

[0004] Existing multimodal sentiment analysis methods lack an effective inter-modal interaction mechanism and cannot fully exploit the complementary information and interdependent relationships between different modalities. There are differences and redundancies in the information between modalities, and simple feature concatenation or average fusion cannot establish deep semantic connections between modalities, reducing the accuracy of analysis.

[0005] Existing technologies ignore the social propagation characteristics of user emotions in social media and do not consider the diffusion and interactive effects of user emotions in social networks. User emotions are not only affected by their own experiences but also by the emotions of their social circles. The lack of modeling of this group emotional interaction mechanism leads to incomplete and inaccurate sentiment analysis results. Summary of the Invention

[0006] Embodiments of the present invention provide a method and system for text user sentiment analysis based on multimodality and AI, which can solve the problems in the existing technology.

[0007] In the first aspect of the embodiments of the present invention, a method for text user sentiment analysis based on multimodality and AI is provided, including: Obtaining text information, image information, and audio information published by a user on a social media platform, and preprocessing the text information, the image information, and the audio information to obtain preprocessed multimodal data; Respectively extracting semantic feature vectors, visual feature vectors, and acoustic feature vectors from the preprocessed multimodal data through a multimodal feature extraction module; Construct a multi-modal feature fusion network. The multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism, and generates a fused feature vector. The fused feature vector contains the temporal dependence relationship of multi-modal data; Input the fused feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model constructs a user emotion social graph and calculates the emotion propagation intensity between user nodes based on the user emotion social graph to obtain a user emotion vector containing group emotion interaction features; Input the user emotion vector into an emotion classifier to obtain an emotion analysis result containing the emotion category and emotion intensity, and generate a user emotion analysis report according to the emotion analysis result.

[0008] And preprocess the text information, the image information, and the audio information. The preprocessed multi-modal data includes: Construct a multi-modal data augmentation module. The multi-modal data augmentation module generates enhanced multi-modal training samples by back-translating and synonym replacement of the text information, randomly cropping and rotating the image information, and time-domain perturbation and frequency-domain mixing of the audio information; Improve the feature expression ability of the enhanced multi-modal training samples in different noise environments through an adversarial training method to obtain the preprocessed multi-modal data.

[0009] Construct a multi-modal feature fusion network. The multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism, and generates a fused feature vector, including: Map the semantic feature vector, the visual feature vector, and the acoustic feature vector to a feature space of a unified dimension through a feature mapping matrix to obtain a mapped semantic feature vector, a mapped visual feature vector, and a mapped acoustic feature vector; Respectively perform linear transformations on the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain corresponding query matrices, key matrices, and value matrices; Calculate the first attention weight between the mapped semantic feature vector and the mapped visual feature vector, the second attention weight between the mapped semantic feature vector and the mapped acoustic feature vector, and the third attention weight between the mapped visual feature vector and the mapped acoustic feature vector based on the query matrix, the key matrix, and the value matrix; Perform feature concatenation on the first attention weight, the second attention weight, and the third attention weight, and perform weighted fusion on the concatenated weights through a fusion weight matrix to obtain an initial fusion feature; Perform residual connection on the initial fusion feature, the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain a fusion feature vector containing multimodal interaction information.

[0010] Input the fusion feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model constructs a user emotion social graph including: Receive the multimodal fusion feature vector, and perform feature concatenation on the multimodal fusion feature vector and the user social attribute vector to obtain an initial feature representation of the user node; Obtain the emotion interaction frequency information between the user node and other user nodes, and calculate the node degree value of the user node according to the emotion interaction frequency information; For any two user nodes with emotion interaction, calculate the emotion correlation weight between the two user nodes according to the ratio of the square root of the product of the two emotion interaction frequency information and the two node degree values; Construct a user emotion social graph based on the initial feature representation and the emotion correlation weight. The user emotion social graph includes a user node set and an emotion interaction edge set.

[0011] Calculate the emotion propagation intensity between user nodes based on the user emotion social graph to obtain a user emotion vector containing group emotion interaction characteristics, including: Based on the feature representations, emotion correlation weights, and the shortest path length between any two connected user nodes in the user emotion social graph, calculate the emotion propagation intensity between the two connected user nodes, where the emotion propagation intensity is obtained by dividing the inner product of the feature representations of the two connected user nodes by the product of the norms of the feature representations of the two connected user nodes, and then multiplying by the emotion correlation weight and the social distance attenuation factor calculated based on the shortest path length; For each user node in the user emotion social graph, calculate the local emotion propagation intensity of the user node based on the emotion propagation intensity and attention weight between the corresponding user node and its neighbor nodes, where the local emotion propagation intensity is obtained by performing weighted average on the product of the emotion propagation intensity and the attention weight exponent; Calculate the overall group emotion propagation intensity according to the product of the local emotion propagation intensities of all user nodes in the user emotion social graph and the corresponding node influence weights; Concatenate the local emotion propagation intensity of each user node, the overall group emotion propagation intensity, and the feature representation of the user node, transform the concatenated features through a feature fusion weight matrix, and add a bias vector to generate the emotion vector of the user node.

[0012] Input the user emotion vector into an emotion classifier to obtain an emotion analysis result including an emotion category and an emotion intensity. Generate a user emotion analysis report based on the emotion analysis result, including: Perform a first non-linear transformation on the user emotion vector through a first weight matrix and a first bias vector to obtain a first hidden layer feature; Perform a second non-linear transformation on the first hidden layer feature through a second weight matrix and a second bias vector to obtain a second hidden layer feature; Based on the second hidden layer feature, perform a softmax transformation through an emotion category prediction weight matrix and an emotion category prediction bias vector to obtain an emotion category probability distribution; at the same time, perform a sigmoid transformation through an emotion intensity prediction weight matrix and an emotion intensity prediction bias vector to obtain an emotion intensity value; Select the category corresponding to the maximum probability in the emotion category probability distribution as the predicted emotion category, and combine the predicted emotion category with the emotion intensity value to generate an emotion analysis result including an emotion category and an emotion intensity. Generate an analysis report reflecting the user's emotional state based on the emotion analysis result.

[0013] In the second aspect of the embodiments of the present invention, a text user emotion analysis system based on multi-modal and AI is provided, including: A first unit for obtaining text information, image information, and audio information published by a user on a social media platform, and preprocessing the text information, the image information, and the audio information to obtain preprocessed multi-modal data; A second unit for respectively extracting a semantic feature vector, a visual feature vector, and an acoustic feature vector from the preprocessed multi-modal data through a multi-modal feature extraction module; A third unit for constructing a multi-modal feature fusion network, which calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism to generate a fusion feature vector, and the fusion feature vector contains the temporal dependence relationship of the multi-modal data; A fourth unit for inputting the fusion feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model constructs a user emotion social graph and calculates the emotion propagation intensity between user nodes based on the user emotion social graph to obtain a user emotion vector containing group emotion interaction features; The fifth unit is configured to input the user emotion vector into an emotion classifier to obtain an emotion analysis result including an emotion category and an emotion intensity, and generate a user emotion analysis report according to the emotion analysis result.

[0014] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] The beneficial effects of the present application are as follows: Compared with the prior art, the present invention combines multi-modal data fusion and deep learning techniques, and can comprehensively capture the emotion information expressed by users on social media. It not only considers various modal data such as text, images, and audio, but also integrates the emotion propagation characteristics in the social network, significantly improving the accuracy and comprehensiveness of emotion analysis.

[0017] The multi-modal feature fusion network of the present invention adopts an innovative cross-modal attention mechanism, which can adaptively learn the correlation and importance between different modal data, effectively solving problems such as insufficient single-modal information and information conflict between modalities in traditional emotion analysis methods, making the analysis results more reliable and stable.

[0018] The present invention introduces an emotion propagation model based on a graph neural network to construct a user emotion social graph, which can effectively capture group emotion interaction characteristics and emotion propagation laws, reveal the diffusion mechanism of emotion in the social network, and provide strong technical support for application scenarios such as social media emotion monitoring, network public opinion analysis, and precision marketing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic flow chart of a text user emotion analysis method based on multi-modal and AI according to an embodiment of the present invention; Figure 2 It is a flow chart of multi-modal data feature enhancement based on adversarial training according to an embodiment of the present invention; Figure 3 It is a flow chart of constructing a user emotion social graph based on multi-modal fusion according to an embodiment of the present invention; Figure 4 It is a flow chart of user emotion analysis based on multi-layer non-linear transformation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 The flowchart of the text user sentiment analysis method based on multimodality and AI according to the embodiments of the present invention is as Figure 1 shown, and the method includes: Obtain the text information, image information, and audio information published by the user on the social media platform, and preprocess the text information, the image information, and the audio information to obtain preprocessed multimodal data; Extract the semantic feature vector, visual feature vector, and acoustic feature vector in the preprocessed multimodal data respectively through a multimodal feature extraction module; Construct a multimodal feature fusion network. The multimodal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism, and generates a fusion feature vector, where the fusion feature vector contains the temporal dependence relationship of multimodal data; Input the fusion feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model constructs a user emotion social graph, and calculates the emotion propagation intensity between user nodes based on the user emotion social graph to obtain a user emotion vector containing group emotion interaction characteristics; Input the user emotion vector into an emotion classifier to obtain an emotion analysis result including an emotion category and an emotion intensity, and generate a user emotion analysis report according to the emotion analysis result.

[0023] In an alternative embodiment, preprocessing the text information, the image information, and the audio information to obtain preprocessed multimodal data includes: Construct a multimodal data augmentation module. The multimodal data augmentation module generates augmented multimodal training samples by back-translating and synonym replacement of the text information, randomly cropping and rotating the image information, and time-domain perturbation and frequency-domain mixing of the audio information; The feature expression capability of the enhanced multimodal training samples in different noise environments is improved through adversarial training to obtain preprocessed multimodal data.

[0024] The present invention relates to a multimodal data preprocessing method, which preprocesses text information, image information and audio information to obtain preprocessed multimodal data.

[0025] In one embodiment, the system first receives raw multimodal data including text information, image information, and audio information. For example, the system receives a text description of an e-commerce product, a product image, and a related voice introduction. For text information, it can be "This smart watch has a heart rate monitoring function and a battery life of up to 7 days"; for image information, it can be a picture showing the appearance of a smart watch; for audio information, it can be a voice explanation of the functions of a smart watch.

[0026] Construct a multimodal data enhancement module for processing. The enhancement of text information includes back translation and synonym replacement. Back translation is to translate the original text from the source language to the target language, and then translate it from the target language back to the source language. In specific implementation, the system translates the original text "This smart watch has a heart rate monitoring function, with a battery life of up to 7 days" into English first, and then translates the English back to Chinese to get "This smart watch has a heart rate monitoring function, with a battery life of up to 7 days". Synonym replacement is to identify keywords in the text and replace them with their synonyms, such as replacing "has" with "has", and replacing "up to" with "reachable". In this way, through back translation and synonym replacement, the system generates text samples with the same semantics but different expressions, enhancing the model's ability to understand different text expressions.

[0027] Enhancement of image information includes random cropping and rotation. Random cropping is to randomly select an area from the original image as a new sample. For example, from the original 600×800 pixel smart watch picture, a 500×600 pixel area is randomly cropped to retain the main part of the watch. Rotation processing is to rotate the image at a random angle. The system rotates the image 15 degrees clockwise to make the watch present a different perspective. Through these operations, the model's recognition ability for images of different angles and sizes is enhanced.

[0028] The enhancement of audio information includes time domain perturbation and frequency domain mixing. Time domain perturbation is to impose changes on the time dimension of the audio signal, such as time stretching and compression. The system stretches the original 5-second speech introduction to 5.5 seconds, making its rhythm slightly slower. Frequency domain mixing is to modify the audio characteristics in the frequency dimension. The system performs high-pass filtering on the audio samples, retaining the frequency components above 1000Hz, making the high-frequency part of the speech more prominent. These processes enable the model to adapt to audio recognition under different speech speeds and sound quality conditions.

[0029] After generating the enhanced multimodal training samples, the system uses adversarial training to improve the sample's feature expression ability in different noise environments. Adversarial training includes two steps: adding noise and generating adversarial samples. Adding noise is to add interference signals to the enhanced multimodal data to simulate the noise interference that may occur in the real environment. For text, the system randomly inserts, deletes or replaces individual characters, such as replacing "smart watch" with "smart watch"; for images, Gaussian noise is added to make the image pixel value fluctuate by 5% on the original basis; for audio, environmental background sound is superimposed, such as low-intensity noise in an office environment (signal-to-noise ratio is 15dB).

[0030] Adversarial sample generation is to create samples that are easy to cause the model to misjudge through machine learning models. The system first uses the basic model to extract features from the enhanced multimodal data to obtain feature vectors. For text, the feature dimension is 768; for images, the feature dimension is 2048; for audio, the feature dimension is 512. Then, the system looks for the perturbation direction that maximizes the classification loss in the feature space, and makes slight modifications to the features in this direction, with the amplitude controlled within 2% of the original feature range. This perturbation is small enough not to affect human perception, but can significantly affect the judgment of the machine learning model. The generated adversarial samples together with the original samples constitute the training set for model training.

[0031] Through adversarial training, the model learns more robust feature representations. Experiments show that compared with models without adversarial training, the accuracy of adversarially trained models in noisy environments has increased by 15%. In specific tests, when 10% of spelling errors are added to the text, the model can still maintain an 85% recognition accuracy rate; when 20% of Gaussian noise is added to the image, the model maintains an accuracy rate of 78%; when background noise with a signal-to-noise ratio of 10dB is superimposed on the audio, the model maintains an accuracy rate of 82%.

[0032] Through the above multi-modal data augmentation and adversarial training steps, effective preprocessing of text information, image information, and audio information is achieved, and the preprocessed multi-modal data is obtained. These data have stronger generalization ability and noise resistance ability, providing high-quality input for subsequent multi-modal fusion and analysis tasks. This preprocessing method shows excellent performance in multi-modal application scenarios such as e-commerce, intelligent customer service, and content review.

[0033] Figure 2 The flowchart of multi-modal data feature augmentation based on adversarial training in an embodiment of the present invention is as follows: This figure shows a complete process of multi-modal data augmentation and preprocessing, mainly including two key steps. The first step is to construct a multi-modal data augmentation module, which processes data of different modalities through various data augmentation techniques: for text information, back-translation technology and synonym replacement method are used for enhancement; for image information, random cropping and rotation operations are performed to expand data samples; for audio information, time-domain perturbation and frequency-domain mixing are used for enhancement processing. The comprehensive application of these enhancement methods finally generates enhanced multi-modal training samples. The second step is to improve the feature expression ability of these enhanced multi-modal training samples in different noise environments through adversarial training, and finally obtain preprocessed multi-modal data. These two steps form a complete processing chain. By combining diverse data augmentation techniques and adversarial training, the quality and robustness of multi-modal data are effectively improved. The entire process design is systematic and logical, with tight connections between individual processing links, all serving the goal of enhancing the feature expression ability of multi-modal data.

[0034] In an alternative embodiment, a multi-modal feature fusion network is constructed. The multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism, and generates a fused feature vector, including: Mapping the semantic feature vector, the visual feature vector, and the acoustic feature vector to a feature space of a unified dimension through a feature mapping matrix, obtaining the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector; Respectively passing the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector through a linear transformation to obtain corresponding query matrices, key matrices, and value matrices; Calculate the first attention weight between the mapped semantic feature vector and the mapped visual feature vector, the second attention weight between the mapped semantic feature vector and the mapped acoustic feature vector, and the third attention weight between the mapped visual feature vector and the mapped acoustic feature vector based on the query matrix, the key matrix, and the value matrix; Perform feature concatenation on the first attention weight, the second attention weight, and the third attention weight, and perform weighted fusion on the weights after feature concatenation through a fusion weight matrix to obtain an initial fusion feature; Perform residual connection on the initial fusion feature, the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain a fusion feature vector containing multimodal interaction information.

[0035] This embodiment relates to a method for constructing a multimodal feature fusion network. This method calculates the interaction weights between semantic feature vectors, visual feature vectors, and acoustic feature vectors through a cross-modal attention mechanism, and generates a fusion feature vector to improve the accuracy of multimodal sentiment analysis.

[0036] In the specific implementation process, first obtain multimodal data, including text data, image data, and audio data. For text data, use a pre-trained language model to extract semantic features; for image data, use a convolutional neural network to extract visual features; for audio data, extract acoustic features through an audio feature extractor. In this way, a semantic feature vector with a dimension of 128, a visual feature vector with a dimension of 256, and an acoustic feature vector with a dimension of 192 are obtained respectively.

[0037] When constructing a multimodal feature fusion network, map the feature vectors of different modalities to a feature space with a unified dimension through a feature mapping matrix. Specifically, use three different mapping matrices to perform linear transformations on the semantic feature vector, the visual feature vector, and the acoustic feature vector respectively, and map them to a feature space with the same dimension of 64. For example, for a semantic feature vector with a dimension of 128, use a mapping matrix with a dimension of 128×64 for transformation; for a visual feature vector with a dimension of 256, use a mapping matrix with a dimension of 256×64 for transformation; for an acoustic feature vector with a dimension of 192, use a mapping matrix with a dimension of 192×64 for transformation. After transformation, the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector with a dimension of 64 are obtained.

[0038] The mapped modal feature vectors are respectively linearly transformed to obtain the corresponding query matrix, key matrix, and value matrix. Specifically, for each mapped feature vector, three different linear transformation matrices are used to convert it into a query matrix Q, a key matrix K, and a value matrix V. Taking the mapped semantic feature vector as an example, a query transformation matrix, a key transformation matrix, and a value transformation matrix with dimensions of 64×64 are used to obtain query vectors, key vectors, and value vectors with dimensions of 64. The visual feature vector and the acoustic feature vector are processed in the same way, and finally nine vectors are obtained: semantic query vector, semantic key vector, semantic value vector, visual query vector, visual key vector, visual value vector, acoustic query vector, acoustic key vector, and acoustic value vector.

[0039] The attention weights between different modal feature vectors are calculated based on the query matrix, key matrix, and value matrix. When calculating the first attention weight between the mapped semantic feature vector and the mapped visual feature vector, the dot product operation is performed between the semantic query vector and the visual key vector, and then the result is divided by 8 (the square root of 64), and then normalized by the softmax function to obtain the semantic-visual attention weight matrix. This weight matrix is multiplied by the visual value vector to obtain the attention output of the semantic feature vector to the visual feature vector. Similarly, the dot product of the visual query vector and the semantic key vector is calculated, and after the same processing, the visual-semantic attention weight matrix is obtained, which is multiplied by the semantic value vector to obtain the attention output of the visual feature vector to the semantic feature vector. The two attention outputs are added together to obtain the first attention weight between the semantic feature vector and the visual feature vector.

[0040] Similarly, the second attention weight between the mapped semantic feature vector and the mapped acoustic feature vector, and the third attention weight between the mapped visual feature vector and the mapped acoustic feature vector are calculated. For example, for the second attention weight, the attention weights are calculated using the semantic query vector and the acoustic key vector, and the acoustic query vector and the semantic key vector respectively, and then the corresponding attention outputs are added together. For the third attention weight, the attention weights are calculated using the visual query vector and the acoustic key vector, and the acoustic query vector and the visual key vector respectively, and then the corresponding attention outputs are added together.

[0041] The first attention weight, the second attention weight, and the third attention weight are feature concatenated, and the weights after feature concatenation are weighted and fused through a fusion weight matrix. Specifically, the three attention weight vectors are concatenated in the feature dimension to obtain a concatenated feature vector with a dimension of 192 (64×3). Then, a fusion weight matrix with dimensions of 192×64 is used to linearly transform the concatenated feature vector into an initial fusion feature with a dimension of 64. This step allows the model to learn the optimal combination method of attention weights between different modalities.

[0042] Perform a residual connection on the initial fusion feature, the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain a fusion feature vector containing multimodal interaction information. Specifically, add the initial fusion feature to the three mapped feature vectors, and then divide by 4 for normalization to obtain the final fusion feature vector. The introduction of the residual connection can effectively alleviate the problem of gradient disappearance while retaining the original feature information.

[0043] Taking specific data as an example, assume that the mapped semantic feature vector, visual feature vector, and acoustic feature vector are all 64-dimensional vectors, and their values are [0.1, 0.2...0.1], [0.3, 0.1...0.2], and [0.2, 0.3...0.1] respectively. After calculation by the attention mechanism, the first attention weight is [0.2, 0.2...0.1], the second attention weight is [0.3,0.1...0.3], and the third attention weight is [0.1, 0.3...0.2]. After splicing these three weight vectors, a 192-dimensional spliced feature vector is obtained, which is converted into a 64-dimensional initial fusion feature [0.2, 0.2...0.3] through the fusion weight matrix. Finally, add the initial fusion feature to the three mapped feature vectors and normalize to obtain the final fusion feature vector [0.2,0.2...0.175].

[0044] This multimodal feature fusion method effectively captures the interaction relationships between different modal features through the cross-modal attention mechanism, significantly improving the performance of multimodal sentiment analysis. Experiments show that compared with traditional feature splicing or weighted average methods, this method improves the accuracy by 5.7% and the F1 score by 6.2% in the multimodal sentiment analysis task.

[0045] In an optional implementation manner, input the fusion feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model includes by constructing a user emotion social graph: Receive the multimodal fusion feature vector, and splice the multimodal fusion feature vector with the user social attribute vector to obtain an initial feature representation of the user node; Obtain the emotion interaction frequency information between the user node and other user nodes, and calculate the node degree value of the user node according to the emotion interaction frequency information; For any two user nodes with emotion interaction, calculate the emotion association weight between the two user nodes according to the ratio of the square root of the product of the two emotion interaction frequency information and the two node degree values; Construct a user emotional social graph based on the initial feature representation and the emotional association weights, where the user emotional social graph includes a set of user nodes and a set of emotional interaction edges.

[0046] In one embodiment, input the multi-modal fusion feature vector into an emotional propagation model based on a graph neural network, which realizes emotional analysis and prediction by constructing a user emotional social graph. The system first receives the multi-modal fusion feature vector, which contains the features of the user's text, image, and behavior data on the social media platform. For example, for user A, the fusion feature vector may contain text emotional features (dimension 768) extracted by a pre-trained language model, image features (dimension 512) extracted by a convolutional neural network, and user behavior features (dimension 256). These features are merged into a 1536-dimensional vector through a fusion method.

[0047] The system concatenates the multi-modal fusion feature vector with the user social attribute vector to obtain the initial feature representation of the user node. The user social attribute vector contains the user's social network relationship information, such as the number of followers, the number of fans, the average interaction frequency, etc. For user A, the social attribute vector may be a 64-dimensional vector containing features such as the number of followers (200), the number of fans (850), and the average number of daily posts (2.3). After concatenating the multi-modal fusion feature vector and the social attribute vector, an initial feature representation vector of 1600 dimensions is obtained.

[0048] The system obtains the emotional interaction frequency information between the user node and other user nodes. Emotional interaction includes interactive behaviors such as likes, comments, forwards, and private messages, and the system records the frequency and emotional tendency of these interactions. For example, user A and user B had 15 interactions in the past 30 days, including 8 likes, 5 emotionally positive comments, and 2 content forwards; user A and user C had 23 interactions, including 10 likes, 8 emotionally positive comments, 3 content forwards, and 2 private message communications.

[0049] Based on the emotional interaction frequency information, the system calculates the node degree value of the user node. The node degree value represents the total interaction frequency of a user with other users in the network, reflecting the user's activity level and influence in the social network. Suppose the total interaction frequency of user A with all other users in the network is 120, then the node degree value of user A is 120; the total interaction frequency of user B is 95, and the node degree value is 95; the total interaction frequency of user C is 150, and the node degree value is 150.

[0050] For any two user nodes with emotional interaction, the system calculates the emotional association weight between the two user nodes according to the ratio of the square root of the product of the two emotional interaction frequency information and the two node degrees. Specifically, the emotional association weight between user A and user B is calculated as follows: the emotional interaction frequency (15) between user A and B is divided by the square root of the product of the degree of user A node (120) and the degree of user B node (95), that is, 15 divided by approximately 106.77, and the obtained emotional association weight is approximately 0.14. Similarly, the emotional association weight between user A and user C is 23 divided by approximately 134.16, approximately 0.17. This calculation method can eliminate the influence brought by the difference in node degrees, so that the high-degree nodes will not dominate the network structure.

[0051] Based on the initial feature representation and the emotional association weight, the system constructs a user emotional social graph. The graph includes a set of user nodes and a set of emotional interaction edges. In the set of user nodes, each node represents a user and contains the initial feature representation of the user; in the set of emotional interaction edges, each edge connects two user nodes with emotional interaction, and the weight of the edge is the calculated emotional association weight. For example, the user emotional social graph contains node A (the initial feature representation is a 1600-dimensional vector), node B, and node C, the weight of edge A - B is 0.14, and the weight of edge A - C is 0.17.

[0052] The constructed user emotional social graph can be used for the training and inference of the graph neural network. Through graph convolution operations, the system can aggregate the neighbor information of the nodes and update the node feature representation. For example, the updated feature representation of user A node will contain the information of its neighbor nodes B and C, with weights of 0.14 and 0.17 respectively. After multiple layers of graph convolution, the system can capture more extensive social network structure information and improve the accuracy of emotional analysis.

[0053] To enhance the expressive ability of the model, the system can introduce an attention mechanism in the graph neural network to dynamically adjust the weights of the edges according to the similarity of node features. For example, if the interest features of user A and user C are very similar, the system will increase the weight of edge A - C, so that the information of user C has a greater impact on the emotional prediction of user A. This method can better simulate the propagation process of emotions in the social network and improve the prediction accuracy of the model.

[0054] Figure 3 The flowchart for constructing a user emotional social graph based on multimodal fusion in the embodiments of the present invention is as follows: This figure details the complete process of constructing a user emotional social graph, which consists of four main steps. First is the feature initialization stage, where the system receives multi-modal fusion feature vectors and performs a feature concatenation operation with the user's social attribute vectors to obtain the initial feature representation of the user nodes. The second step is the node degree value calculation stage. By obtaining the emotional interaction frequency information between user nodes, the system can calculate the node degree value of each user node. The third step is the emotional association weight calculation stage. For any two user nodes with emotional interactions, the system calculates the emotional association weight between the two nodes based on the mathematical relationship between their emotional interaction frequency and node degree value. Finally is the graph construction stage. Based on the initial feature representation and emotional association weight obtained previously, the system constructs a complete user emotional social graph, which includes two key components: the user node set and the emotional interaction edge set. These four steps are closely linked, forming a complete process for constructing a user emotional social graph, and the output of each step provides the necessary input information for the next step.

[0055] In an alternative embodiment, calculating the emotional propagation intensity between user nodes based on the user emotional social graph to obtain a user emotional vector containing group emotional interaction characteristics includes: Based on the feature representations, emotional association weights, and the shortest path length between any two connected user nodes in the user emotional social graph, calculating the emotional propagation intensity between the two connected user nodes, where the emotional propagation intensity is obtained by dividing the inner product of the feature representations of the two connected user nodes by the product of the norms of the feature representations of the two connected user nodes, and then multiplying by the emotional association weight and the social distance attenuation factor calculated based on the shortest path length; For each user node in the user emotional social graph, calculating the local emotional propagation intensity of the user node based on the emotional propagation intensity and attention weight between the corresponding user node and its neighbor nodes, where the local emotional propagation intensity is obtained by performing a weighted average on the product of the emotional propagation intensity and the attention weight exponent; Calculating the overall group emotional propagation intensity according to the product of the local emotional propagation intensities of all user nodes in the user emotional social graph and the corresponding node influence weights; Concatenating the local emotional propagation intensity of each user node, the overall group emotional propagation intensity, and the feature representation of the user node, and transforming the concatenated features through a feature fusion weight matrix and adding a bias vector to generate the emotional vector of the user node.

[0056] Calculate the emotional propagation intensity between user nodes based on the user emotional social graph. For any two connected user nodes in the graph, such as user A and user B, extract their feature representations, emotional association weights, and the shortest path length between them. Suppose the feature representation of user A is [0.3, 0.5, 0.7, 0.2], the feature representation of user B is [0.2, 0.6, 0.5, 0.3], the emotional association weight between them is 0.8, and the shortest path length is 2. Calculate the inner product of the feature representations of the two user nodes, that is, 0.3×0.2 + 0.5×0.6 + 0.7×0.5 + 0.2×0.3 = 0.76.

[0057] Divide the inner product by the product of the norms to get 0.76 / 0.79 = 0.96. Calculate the social distance decay factor based on the shortest path length 2. The exponential decay formula can be used to obtain a decay factor of 0.5. The final emotional propagation intensity is 0.96×0.8×0.5 = 0.38.

[0058] For each user node in the user emotional social graph, calculate its local emotional propagation intensity. Taking user A as an example, assume its neighbor nodes include users B, C, and D, and the emotional propagation intensities with these neighbors are 0.38, 0.45, and 0.29 respectively, and the corresponding attention weights are 0.4, 0.35, and 0.25.

[0059] Calculate the product of the emotional propagation intensity and the exponent of the attention weight, that is: 0.38×(0.4)² + 0.45×(0.35)² + 0.29×(0.25)² = 0.061 + 0.055 + 0.018 = 0.134.

[0060] To obtain the weighted average, divide by the sum of the exponents of the attention weights, that is: (0.4)² + (0.35)² + (0.25)² = 0.16 + 0.1225 + 0.0625 = 0.345, and the final local emotional propagation intensity is 0.134 / 0.345 = 0.39.

[0061] Next, calculate the overall emotional propagation intensity of the group. Suppose there are 5 user nodes in the social graph, and their local emotional propagation intensities are 0.39, 0.42, 0.35, 0.38, and 0.41 respectively, and the corresponding node influence weights are 0.3, 0.25, 0.15, 0.1, and 0.2.

[0062] Calculate the sum of the products of the local emotional propagation intensity and the node influence weight, that is: 0.39×0.3 + 0.42×0.25 + 0.35×0.15 + 0.38×0.1 + 0.41×0.2 = 0.117 + 0.105 + 0.0525 + 0.038 + 0.082 = 0.3945。

[0063] When generating the user emotion vector, taking user A as an example, the local emotion propagation intensity of user A, which is 0.39, the overall emotion propagation intensity of the group, which is 0.3945, and the feature representation of user A, [0.3, 0.5, 0.7, 0.2], are concatenated to obtain [0.39, 0.3945, 0.3, 0.5, 0.7, 0.2]. Assume that the feature fusion weight matrix is a 6-row and 4-column matrix, and the specific values are [[0.2, 0.15, 0.1, 0.25], [0.3, 0.2, 0.15, 0.1], [0.25, 0.35, 0.1, 0.2], [0.1, 0.2, 0.3, 0.25], [0.15, 0.1, 0.25, 0.15], [0.2, 0.25, 0.15, 0.1]], and the bias vector is [0.05, 0.1, 0.15, 0.2].

[0064] Multiply the concatenated features by the weight matrix. The specific calculation method is as follows: The first output element is equal to: 0.39×0.2 + 0.3945×0.3 + 0.3×0.25 + 0.5×0.1 + 0.7×0.15 + 0.2×0.2 + 0.05 = 0.078 + 0.118 + 0.075 + 0.05 + 0.105 + 0.04 + 0.05 = 0.516; The second output element is equal to: 0.39×0.15 + 0.3945×0.2 + 0.3×0.35 + 0.5×0.2 + 0.7×0.1 + 0.2×0.25 + 0.1 = 0.058 + 0.079 + 0.105 + 0.1 + 0.07 + 0.05 + 0.1 = 0.562; And so on. Calculate the third and fourth output elements to be 0.496 and 0.481 respectively. Finally, the emotion vector of user A is [0.516, 0.562, 0.496, 0.481].

[0065] In practical applications, this method can be used to analyze the law of user emotion propagation on social media platforms. For example, the platform can identify key users with greater influence on emotion propagation and push content targeted at them; it can also predict the emotion propagation trend of specific information and adjust the content strategy in a timely manner. In addition, this method can also be used to evaluate the effect of marketing activities, and measure the impact of marketing activities on user emotions by analyzing the changes in user emotion vectors.

[0066] Through the above technical solution, the present invention can comprehensively capture the emotional interaction characteristics between users, taking into account not only the individual characteristics of users, but also the characteristics of the social network structure and the dynamics of group emotion propagation, thereby improving the accuracy and comprehensiveness of user emotion analysis. This has important value for understanding the emotion propagation mechanism in social networks, predicting user behavior, and optimizing social media operation strategies.

[0067] In an optional implementation manner, the user emotion vector is input into an emotion classifier to obtain an emotion analysis result including an emotion category and an emotion intensity. Generating a user emotion analysis report according to the emotion analysis result includes: Performing a first non-linear transformation on the user emotion vector through a first weight matrix and a first bias vector to obtain a first hidden layer feature; Performing a second non-linear transformation on the first hidden layer feature through a second weight matrix and a second bias vector to obtain a second hidden layer feature; Based on the second hidden layer feature, performing a softmax transformation through an emotion category prediction weight matrix and an emotion category prediction bias vector to obtain an emotion category probability distribution; at the same time, performing a sigmoid transformation through an emotion intensity prediction weight matrix and an emotion intensity prediction bias vector to obtain an emotion intensity value; Selecting the category corresponding to the maximum probability according to the emotion category probability distribution as the predicted emotion category, and combining the predicted emotion category with the emotion intensity value to generate an emotion analysis result including an emotion category and an emotion intensity, and generating an analysis report reflecting the user's emotional state based on the emotion analysis result.

[0068] The user emotion vector is a high-dimensional vector containing emotion features extracted based on user interaction data, with a dimension of 128. This vector contains multi-dimensional information such as the text emotion features, behavior patterns, and historical interaction data of users on social media. After obtaining this emotion vector, the system inputs it into the designed emotion classifier for processing.

[0069] The emotion classifier adopts a multi-layer neural network structure, including two hidden layers and two parallel output layers, which are respectively used for predicting emotion categories and emotion intensities. The processing process of the user emotion vector is described in detail as follows: The system first performs a first non - linear transformation on the user emotion vector through the first weight matrix and the first bias vector. The dimension of the first weight matrix is 128×256, where 128 corresponds to the dimension of the input emotion vector and 256 corresponds to the number of neurons in the first hidden layer. The dimension of the first bias vector is 256×1. In actual processing, the system multiplies the user emotion vector by the first weight matrix and then adds the first bias vector to obtain the result of the linear transformation. Then, the system applies the ReLU activation function to this result of the linear transformation for non - linear transformation, that is, setting the elements less than 0 in the result of the linear transformation to 0 and keeping the elements greater than 0 unchanged, thereby obtaining the features of the first hidden layer. For example, for the input emotion vector [0.23, 0.45, - 0.12...0.78], after the first - layer processing, the features of the first hidden layer [0.56, 0, 0.91...0.34] may be obtained.

[0070] After that, the system performs a second non - linear transformation on the features of the first hidden layer through the second weight matrix and the second bias vector. The dimension of the second weight matrix is 256×128, where 256 corresponds to the dimension of the features of the first hidden layer and 128 corresponds to the number of neurons in the second hidden layer. The dimension of the second bias vector is 128×1. The processing process is similar to the first transformation. The system multiplies the features of the first hidden layer by the second weight matrix, adds the second bias vector, and then applies the ReLU activation function again to obtain the features of the second hidden layer. For example, for the features of the first hidden layer [0.56, 0, 0.91...0.34], after the second - layer processing, the features of the second hidden layer [0.78, 0.23, 0...0.45] may be obtained.

[0071] After obtaining the features of the second hidden layer, the system processes in two directions: emotion category prediction and emotion intensity prediction. For emotion category prediction, the system uses the emotion category prediction weight matrix and the emotion category prediction bias vector. The dimension of the emotion category prediction weight matrix is 128×6, where 128 corresponds to the dimension of the features of the second hidden layer and 6 corresponds to the number of preset emotion categories (including "happy", "angry", "sad", "disgusted", "fearful", and "surprised"). The dimension of the emotion category prediction bias vector is 6×1. The system multiplies the features of the second hidden layer by the emotion category prediction weight matrix, adds the emotion category prediction bias vector, and then performs a softmax transformation. The softmax transformation converts each element in the vector into a value between 0 and 1, and the sum of all elements is 1, representing the probability distribution of each emotion category. For example, after processing the features of the second hidden layer [0.78, 0.23, 0...0.45], the emotion category probability distribution [0.65, 0.12, 0.08, 0.05, 0.04, 0.06] may be obtained, indicating that the probability of the "happy" category is 0.65, the probability of the "angry" category is 0.12, and so on.

[0072] Meanwhile, for sentiment intensity prediction, the system uses a sentiment intensity prediction weight matrix and a sentiment intensity prediction bias vector. The dimension of the sentiment intensity prediction weight matrix is 128×1, where 128 corresponds to the dimension of the features in the second hidden layer, and 1 corresponds to the output sentiment intensity value. The dimension of the sentiment intensity prediction bias vector is 1×1. The system multiplies the features of the second hidden layer by the sentiment intensity prediction weight matrix, adds the sentiment intensity prediction bias vector, and then performs a sigmoid transformation. The sigmoid transformation converts the result into a value between 0 and 1, representing the sentiment intensity. For example, after processing the features of the second hidden layer [0.78, 0.23, 0...0.45], the sentiment intensity value 0.82 may be obtained, indicating a relatively high sentiment intensity.

[0073] According to the sentiment category probability distribution, the system selects the category corresponding to the maximum probability as the predicted sentiment category. In the above example, the predicted sentiment category is "happy", corresponding to a probability of 0.65. The system combines the predicted sentiment category with the sentiment intensity value to generate a sentiment analysis result containing the sentiment category and the sentiment intensity. For example, the result can be expressed as ["happy", 0.82], indicating that the user's current sentiment category is "happy" and the sentiment intensity is 0.82.

[0074] Based on the sentiment analysis result, the system generates an analysis report reflecting the user's sentiment state. The report includes the user's dominant sentiment category, sentiment intensity level, sentiment change trend, etc. For example, for the result ["happy", 0.82], the system may generate a report: "The user is currently showing a happy mood, with a sentiment intensity of 82%, belonging to a high-intensity positive sentiment state. Compared with historical data, the user's sentiment state maintains a stable positive trend. It is recommended that the system push relevant positive content to maintain the user's good sentiment experience.

[0075] In practical applications, the weight matrix and bias vector of the sentiment classifier are obtained through training with a large amount of labeled sentiment data to ensure accurate classification of different user sentiment vectors. The sentiment analysis report can be customized with different output formats and content depths according to different application scenarios (such as intelligent customer service, mental health monitoring, social media analysis, etc.), providing decision support for understanding the user's sentiment state and corresponding services.

[0076] Figure 4 The flowchart of the user sentiment analysis based on multi-layer non-linear transformation in the embodiment of the present invention is as follows: The flowchart details a complete user sentiment analysis process, which consists of four key steps. First, the system performs the first feature extraction on the user sentiment vector, implementing a non-linear transformation through the first weight matrix and the first bias vector to generate the first hidden layer features. Subsequently, the obtained first hidden layer features are input into the second layer network, and a second non-linear transformation is performed through the second weight matrix and the second bias vector to obtain the second hidden layer features. Then, the system simultaneously performs two prediction tasks based on the second hidden layer features: on the one hand, a softmax transformation is performed through the sentiment category prediction weight matrix and bias vector to obtain the probability distribution of the sentiment category; on the other hand, a sigmoid transformation is performed through the sentiment intensity prediction weight matrix and bias vector to obtain the sentiment intensity value. Finally, the system selects the category corresponding to the maximum probability value from the sentiment category probability distribution as the prediction result, combines it with the sentiment intensity value, generates a complete sentiment analysis result, and further forms an analysis report reflecting the user's sentiment state. This process demonstrates a complete conversion link from the original features to the final analysis result, and each step undergoes a carefully designed non-linear transformation to ensure the accuracy and reliability of the sentiment analysis.

[0077] In the second aspect of the embodiments of the present invention, there is provided a text user sentiment analysis system based on multi-modal and AI, including: A first unit, configured to obtain text information, image information, and audio information published by a user on a social media platform, and preprocess the text information, the image information, and the audio information to obtain preprocessed multi-modal data; A second unit, configured to respectively extract a semantic feature vector, a visual feature vector, and an acoustic feature vector from the preprocessed multi-modal data through a multi-modal feature extraction module; A third unit, configured to construct a multi-modal feature fusion network, where the multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism, and generates a fusion feature vector, and the fusion feature vector includes the temporal dependence relationship of the multi-modal data; A fourth unit, configured to input the fusion feature vector into a sentiment propagation model based on a graph neural network, and the sentiment propagation model calculates the sentiment propagation intensity between user nodes based on the constructed user sentiment social graph to obtain a user sentiment vector including group sentiment interaction features; A fifth unit, configured to input the user sentiment vector into a sentiment classifier to obtain a sentiment analysis result including a sentiment category and a sentiment intensity, and generate a user sentiment analysis report according to the sentiment analysis result.

[0078] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the method described above.

[0079] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0080] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for performing various aspects of the present invention are loaded.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text user sentiment analysis method based on multimodality and AI, characterized in that, Including: Obtain the text information, image information, and audio information published by the user on the social media platform, and preprocess the text information, the image information, and the audio information to obtain preprocessed multi-modal data; Extract the semantic feature vector, visual feature vector, and acoustic feature vector in the preprocessed multi-modal data respectively through a multi-modal feature extraction module; Construct a multi-modal feature fusion network, and the multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism to generate a fusion feature vector, and the fusion feature vector contains the temporal dependence relationship of the multi-modal data; Input the fusion feature vector into the emotion propagation model based on the graph neural network, and the emotion propagation model calculates the emotion propagation intensity between user nodes based on the constructed user emotion social graph to obtain a user emotion vector containing group emotion interaction characteristics; Input the user emotion vector into an emotion classifier to obtain an emotion analysis result including emotion category and emotion intensity, and generate a user emotion analysis report according to the emotion analysis result.

2. The method according to claim 1, wherein And preprocess the text information, the image information, and the audio information, and the preprocessed multi-modal data obtained includes: Construct a multi-modal data augmentation module, and the multi-modal data augmentation module generates enhanced multi-modal training samples by back-translating and synonym replacement of the text information, randomly cropping and rotating the image information, and time-domain perturbation and frequency-domain mixing of the audio information; Improve the feature expression ability of the enhanced multi-modal training samples in different noise environments through an adversarial training method to obtain preprocessed multi-modal data.

3. The method according to claim 1, wherein Construct a multi-modal feature fusion network, and the multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector, and the acoustic feature vector through a cross-modal attention mechanism to generate a fusion feature vector, including: Map the semantic feature vector, the visual feature vector, and the acoustic feature vector to a feature space of a unified dimension through a feature mapping matrix to obtain a mapped semantic feature vector, a mapped visual feature vector, and a mapped acoustic feature vector; Respectively perform linear transformation on the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain corresponding query matrices, key matrices, and value matrices; Calculate the first attention weight between the mapped semantic feature vector and the mapped visual feature vector, the second attention weight between the mapped semantic feature vector and the mapped acoustic feature vector, and the third attention weight between the mapped visual feature vector and the mapped acoustic feature vector based on the query matrix, the key matrix, and the value matrix; Perform feature splicing on the first attention weight, the second attention weight, and the third attention weight, and perform weighted fusion on the spliced weights through a fusion weight matrix to obtain an initial fusion feature; Perform a residual connection on the initial fusion feature, the mapped semantic feature vector, the mapped visual feature vector, and the mapped acoustic feature vector to obtain a fusion feature vector containing multimodal interaction information.

4. The method according to claim 1, wherein Input the fusion feature vector into an emotion propagation model based on a graph neural network. The emotion propagation model constructs a user emotion social graph including: Receive the multimodal fusion feature vector, and perform feature splicing on the multimodal fusion feature vector and the user social attribute vector to obtain an initial feature representation of the user node; Obtain the emotion interaction frequency information between the user node and other user nodes, and calculate the node degree value of the user node according to the emotion interaction frequency information; For any two user nodes with emotion interaction, calculate the emotion association weight between the two user nodes according to the ratio of the square root of the product of the two emotion interaction frequency information and the two node degree values; Construct a user emotion social graph based on the initial feature representation and the emotion association weight. The user emotion social graph includes a user node set and an emotion interaction edge set.

5. The method according to claim 1, characterized in that, Calculate the emotion propagation intensity between user nodes based on the user emotion social graph to obtain a user emotion vector containing group emotion interaction characteristics, including: Based on the feature representations, emotion association weights, and the shortest path length between any two connected user nodes in the user emotion social graph, calculate the emotion propagation intensity between the two connected user nodes, where the emotion propagation intensity is obtained by dividing the inner product of the feature representations of the two connected user nodes by the product of the norms of the feature representations of the two connected user nodes, and then multiplying by the emotion association weight and the social distance attenuation factor calculated based on the shortest path length; For each user node in the user emotion social graph, calculate the local emotion propagation intensity of the user node based on the emotion propagation intensity and the attention weight between the corresponding user node and its neighbor nodes, where the local emotion propagation intensity is obtained by performing a weighted average on the product of the emotion propagation intensity and the attention weight exponent; Calculate the overall group emotion propagation intensity according to the product of the local emotion propagation intensities of all user nodes in the user emotion social graph and the corresponding node influence weights; Perform feature splicing on the local emotion propagation intensity of each user node, the overall group emotion propagation intensity, and the feature representation of the user node, and transform the spliced features through a feature fusion weight matrix and add a bias vector to generate the emotion vector of the user node.

6. The method according to claim 1, characterized in that Input the user emotion vector into an emotion classifier to obtain an emotion analysis result containing an emotion category and an emotion intensity, and generate a user emotion analysis report according to the emotion analysis result, including: Perform a first non-linear transformation on the user emotion vector through a first weight matrix and a first bias vector to obtain a first hidden layer feature; Perform a second non-linear transformation on the first hidden layer feature through a second weight matrix and a second bias vector to obtain a second hidden layer feature; Based on the second hidden layer features, perform a softmax transformation through the sentiment category prediction weight matrix and the sentiment category prediction bias vector to obtain the sentiment category probability distribution; at the same time, perform a sigmoid transformation through the sentiment intensity prediction weight matrix and the sentiment intensity prediction bias vector to obtain the sentiment intensity value; Select the category corresponding to the maximum probability according to the sentiment category probability distribution as the predicted sentiment category, and combine the predicted sentiment category with the sentiment intensity value to generate a sentiment analysis result including the sentiment category and the sentiment intensity, and generate an analysis report reflecting the user's sentiment state based on the sentiment analysis result.

7. A text user sentiment analysis system based on multi-modal and AI, for implementing the method described in any one of the preceding claims 1-6, characterized in that, Including: The first unit is used to obtain the text information, image information and audio information published by the user on the social media platform, and preprocess the text information, the image information and the audio information to obtain the preprocessed multi-modal data; The second unit is used to extract the semantic feature vector, the visual feature vector and the acoustic feature vector in the preprocessed multi-modal data respectively through the multi-modal feature extraction module; The third unit is used to construct a multi-modal feature fusion network, and the multi-modal feature fusion network calculates the interaction weights between the semantic feature vector, the visual feature vector and the acoustic feature vector through a cross-modal attention mechanism to generate a fused feature vector, and the fused feature vector contains the temporal dependence relationship of the multi-modal data; The fourth unit is used to input the fused feature vector into the sentiment propagation model based on the graph neural network. The sentiment propagation model calculates the sentiment propagation intensity between user nodes based on the constructed user sentiment social graph to obtain a user sentiment vector containing group sentiment interaction features; The fifth unit is used to input the user sentiment vector into the sentiment classifier to obtain a sentiment analysis result including the sentiment category and the sentiment intensity, and generate a user sentiment analysis report according to the sentiment analysis result.

8. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image sentiment polarity classification method combined with graph convolutional neural network

    CN112712127A

  • Scenic spot recommendation method integrating user emotion and knowledge graph enhancement

    CN115730138A

  • Label prediction model training method and device, equipment, medium and program product

    CN116956104A

  • Whole-network big data public opinion monitoring system based on cloud media

    CN118349626A

  • Cross-modal data fusion user psychological portrait acquisition method and system

    CN118885766A

Cited By

  • Text sentiment analysis method and device, electronic equipment and storage medium

    CN120471055A

  • Marketing video auditing method based on AI

    CN120583273A

  • Emotion recognition method and device based on multi-modal missing data, equipment and medium

    CN120654201A

  • Digital human speech synthesis method and system based on multi-modal speech feature fusion

    CN120833777A

  • Digital human speech synthesis method and system based on multi-modal speech feature fusion

    CN120833777B