Social emotion analysis method and system based on multi-modal data
By using MobileNet and BiLSTM in the social emotion analysis system for real-time emotion classification and in-depth analysis combined with the multimodal joint coding model of the GPU cluster, the accuracy and consistency problems in multimodal data fusion and processing are solved, and efficient social emotion analysis is achieved.
Patent Information
- Application Number
- CN202510277827.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems that accuracy and consistency are difficult to guarantee in the fusion and processing of multimodal data, especially in real-time analysis and in-depth analysis modules.
Real-time emotion classification was performed using MobileNet combined with BiLSTM, and the multimodal joint coding model was run through GPU cluster for in-depth analysis to generate a fine-grained emotion map.
It realizes efficient real-time analysis of multimodal data, can quickly identify and display complex emotional information, and provides strong support for related decisions.
Smart Images

Figure CN120123508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social emotion analysis, and particularly to a social emotion analysis method and system based on multi-modal data. Background Art
[0002] In a social emotion analysis system based on multi-modal data, the data acquisition layer obtains data through multiple channels, including API access, web crawler scraping, and user-uploaded content. These data cover various formats such as text, pictures, short videos, and live audio. The preprocessing layer deploys edge computing nodes, which are responsible for data cleaning, format standardization, and preliminary feature extraction. The cleaning process removes invalid information such as advertising text, and the standardization process includes operations such as video frame segmentation. The preliminary feature extraction provides a basis for subsequent analysis.
[0003] The analysis engine layer is divided into a real-time analysis module and a deep analysis module. The real-time analysis module uses lightweight models, such as MobileNet combined with BiLSTM, to achieve second-level emotion classification, meeting the immediate needs of feedback information monitoring. The deep analysis module then calls the GPU cluster to run a multi-modal joint encoding model to generate a fine-grained emotion map, such as "angry - 80% | surprised - 15% | neutral - 5%", providing more accurate emotion analysis results.
[0004] The application layer provides a visualization dashboard, API interfaces, and an early warning system, supporting the display of emotion heat maps. The visualization dashboard intuitively presents the emotion analysis results, the API interfaces facilitate other systems to call the analysis data, and the early warning system issues emotion anomaly alarms in a timely manner. The emotion heat map reflects the emotion intensity through the depth of color, helping users quickly grasp the social emotion distribution.
[0005] In the entire system, the key to technical problems lies in the fusion and processing of multi-modal data. Different formats of data require a unified processing process to ensure the accuracy and consistency of the analysis results. The real-time analysis module needs to ensure speed without sacrificing the accuracy of emotion classification. The deep analysis module needs to make full use of the computing power of the GPU cluster to process complex multi-modal data joint encoding and generate a fine-grained emotion map. The solution of these technical problems directly affects the performance and user experience of the entire system. Summary of the Invention
[0006] The purpose of the present invention is to propose a social emotion analysis method and system based on multi-modal data to solve the problems existing in the above-mentioned prior art.
[0007] To achieve the above purpose, the present invention provides the following solutions:
[0008] A social emotion analysis method based on multi-modal data, comprising:
[0009] Obtain multimodal data; wherein, the multimodal data includes: text data, picture data, short video data, and live audio data;
[0010] Preprocess the multimodal data;
[0011] Perform preliminary feature extraction on the preprocessed multimodal data to generate a basic feature vector;
[0012] Perform second-level sentiment classification on the basic feature vector;
[0013] Perform in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map;
[0014] Generate a sentiment heat map according to the fine-grained sentiment map, wherein the sentiment heat map reflects the sentiment intensity through the depth of color and displays the sentiment heat map through a visualization dashboard.
[0015] Optionally, preprocessing the multimodal data includes:
[0016] Perform data cleaning on the multimodal data to remove invalid information;
[0017] Perform format standardization processing on the cleaned data.
[0018] Optionally, performing preliminary feature extraction on the preprocessed multimodal data includes:
[0019] Use MobileNet to obtain a basic feature vector from the preprocessed multimodal data; wherein, the basic feature vector includes: basic feature vector of text data: syntactic analysis results, semantic features of word frequency information; basic feature vector of picture data: image content features such as color distribution, texture features, shape features, and local region features; basic feature vector of short video data: feature vectors of each frame of the picture, feature vectors of the audio stream, and relevant information features such as video duration and background music; basic feature vector of live audio data: audio content features of pitch and tone.
[0020] Optionally, performing second-level sentiment classification on the basic feature vector includes:
[0021] Use BiLSTM to perform temporal modeling on the basic feature vector to capture the time-dependent relationship of the features;
[0022] Perform dimensionality reduction processing on the feature vector after temporal modeling to obtain a low-dimensional feature representation;
[0023] Based on the dimension-reduced feature vectors, a Softmax classifier is used for sentiment classification; if the confidence of the classification result is higher than a preset threshold, the sentiment category is determined; according to the determined sentiment category, a sentiment label dataset is generated; using the sentiment label dataset, a sentiment analysis model is constructed to achieve second-level sentiment classification.
[0024] Optionally, perform in-depth analysis on the second-level sentiment classification result to generate a fine-grained sentiment map, including:
[0025] Use a GPU cluster to load a pre-trained multi-modal joint encoding model to extract features from the input second-level sentiment classification result;
[0026] Fuse the extracted features through the multi-modal joint encoding model to generate a high-dimensional joint feature vector;
[0027] Deeply analyze the high-dimensional joint feature vector to capture the implicit fine-grained information;
[0028] According to the deep analysis result, use a clustering algorithm to classify the feature vectors; if the clustering result meets the preset category distribution condition, determine the sentiment category label; based on the determined sentiment category label, construct the node information of the sentiment map;
[0029] Model the node information of the sentiment map through a graph neural network to generate a fine-grained sentiment map.
[0030] Optionally, generate a sentiment heat map according to the fine-grained sentiment map, including:
[0031] Based on the fine-grained sentiment map, extract the intensity value corresponding to the sentiment value;
[0032] According to the intensity value distribution, use a mapping method to convert the intensity value into a color value;
[0033] Generate a sentiment heat map through the correspondence between the color value and the intensity value.
[0034] A social emotion analysis system based on multi-modal data, the system includes: a data acquisition module, a preprocessing module, a preliminary feature extraction module, a second-level sentiment classification module, a deep analysis module, and a sentiment intensity display module;
[0035] The data acquisition module is used to acquire multi-modal data; wherein, the multi-modal data includes: text data, picture data, short video data, and live audio data;
[0036] The preprocessing module is used to preprocess the multi-modal data;
[0037] The preliminary feature extraction module is used to perform preliminary feature extraction on the preprocessed multi-modal data to generate a basic feature vector;
[0038] The second-level emotion classification module is used to perform second-level emotion classification on the basic feature vectors;
[0039] The in-depth analysis module is used to perform in-depth analysis on the second-level emotion classification results to generate a fine-grained emotion map;
[0040] The emotion intensity display module is used to generate an emotion heat map according to the fine-grained emotion map. Among them, the emotion heat map reflects the emotion intensity through the depth of color, and the emotion heat map is displayed through a visualization dashboard.
[0041] Optionally, the preliminary feature extraction module performs preliminary feature extraction on the preprocessed multi-modal data, including:
[0042] Using MobileNet, obtain basic feature vectors from the preprocessed multi-modal data; among them, the basic feature vectors include: basic feature vectors of text data: syntactic analysis results, semantic features of word frequency information; basic feature vectors of picture data: image content features such as color distribution, texture features, shape features, and local region features; basic feature vectors of short video data: feature vectors of each frame of the picture, feature vectors of the audio stream, and relevant information features such as video duration and background music; basic feature vectors of live audio data: audio content features of pitch and tone.
[0043] Optionally, the second-level emotion classification module performs second-level emotion classification on the basic feature vectors, including:
[0044] Using BiLSTM to perform temporal modeling on the basic feature vectors to capture the time-dependent relationship of the features;
[0045] Perform dimensionality reduction processing on the feature vectors after temporal modeling to obtain low-dimensional feature representations;
[0046] Based on the feature vectors after dimensionality reduction, use a Softmax classifier for emotion classification; if the confidence of the classification result is higher than the preset threshold, determine the emotion category; according to the determined emotion category, generate an emotion label data set; use the emotion label data set to construct an emotion analysis model to achieve second-level emotion classification.
[0047] The beneficial effects of the present invention are:
[0048] The present invention discloses a social emotion analysis method and system based on multi-modal data. The method obtains multi-modal data, including text, pictures, short videos, and live audio, through API access, web crawler scraping, and user upload. After data cleaning and format standardization processing at the edge computing node, a lightweight model combining MobileNet and BiLSTM is used for real-time analysis to obtain second-level emotion classification results. Subsequently, a GPU cluster is called to run a multi-modal joint encoding model to deeply analyze the results and generate a fine-grained emotion map. The present invention also provides a visualization dashboard to intuitively display the analysis results through an emotion heat map, and is equipped with an API interface and an early warning system to support other systems to call and analyze data. The method of the present invention realizes efficient real-time analysis of multi-modal data, can quickly identify and display complex emotion information, and provides strong support for relevant decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flow chart of a social emotion analysis method based on multi-modal data according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0052] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0053] As Figure 1 shown, this embodiment proposes a social emotion analysis method based on multi-modal data, including:
[0054] Obtain multi-modal data; wherein, the multi-modal data includes: text data, picture data, short video data, and live audio data;
[0055] Preprocess the multi-modal data;
[0056] Perform preliminary feature extraction on the preprocessed multimodal data to generate basic feature vectors;
[0057] Perform second-level sentiment classification on the basic feature vectors;
[0058] Perform in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map;
[0059] Generate a sentiment heat map based on the fine-grained sentiment map. The sentiment heat map reflects the sentiment intensity through the depth of color and displays the sentiment heat map through a visualization dashboard.
[0060] Specifically, in this embodiment, first, multimodal data is obtained from API access, crawler scraping, and user-uploaded content, including text, pictures, short videos, and live audio. Secondly, data cleaning is performed to remove invalid information such as advertising text and retain valid data. Then, format standardization processing is performed on the cleaned data, such as frame splitting for videos and segmenting for audio to unify them into a processable format. Then, preliminary feature extraction is performed on the standardized data to generate basic feature vectors. And a lightweight model combining MobileNet and BiLSTM is used to perform real-time analysis on the basic feature vectors to obtain second-level sentiment classification results. Then, the GPU cluster is called to run the multimodal joint encoding model to perform in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map. Finally, a sentiment heat map is generated based on the fine-grained sentiment map, which reflects the sentiment intensity through the depth of color and is displayed through a visualization dashboard. And an API interface and an early warning system are provided to support other systems to call and analyze data. This method realizes the efficient and real-time analysis of multimodal data, can quickly identify and display complex sentiment information, and provides strong support for relevant decisions.
[0061] Furthermore, the preprocessing of multimodal data includes:
[0062] Perform data cleaning on the multimodal data to remove invalid information;
[0063] Perform format standardization processing on the cleaned data.
[0064] Specifically, in this embodiment, first, data cleaning is performed at the edge node to obtain the original text content from the data source. The pre-established advertising text recognition rules are used to filter advertising text and retain valid data. If the text content conforms to the characteristics of invalid information, the text is discarded; otherwise, it enters the next step of processing. According to the text characteristics of the valid data, keywords and topic distributions are extracted. The keywords are classified through machine learning algorithms to determine the text category. Combining the text category and the topic distribution, the structured content after data cleaning is generated. The structured content is stored in a preset database to complete the data cleaning process.
[0065] Secondly, adopt the preset video frame segmentation rule to segment the video data frame by frame to obtain the image sequence after frame segmentation. According to the audio segmentation rule, divide the audio data into time windows to generate the segmented audio clips. For the image sequence after frame segmentation, extract the image features of each frame to generate an image feature dataset. For the segmented audio clips, extract the audio features of each segment to generate an audio feature dataset. If the formats of the image feature dataset and the audio feature dataset do not match, then use the conversion processing rule to adjust them to a unified format. According to the image feature dataset and the audio feature dataset, generate the multi-modal feature fusion result. Store the multi-modal feature fusion result in the preset database to complete the data standardization processing process.
[0066] Furthermore, for the preprocessed multi-modal data, perform preliminary feature extraction including:
[0067] Adopt MobileNet to obtain the basic feature vectors from the preprocessed multi-modal data; among them, the basic feature vectors include:
[0068] Basic feature vectors of text data: semantic features including syntactic analysis results and word frequency information;
[0069] Basic feature vectors of picture data: image content features including color distribution, texture features, shape features and local region features;
[0070] Basic feature vectors of short video data: feature vectors of each frame of the picture, feature vectors of the audio stream, and relevant information features of the video duration and background music;
[0071] Basic feature vectors of live audio data: audio content features including tone and pitch.
[0072] Furthermore, perform second-level emotion classification on the basic feature vectors including:
[0073] Use BiLSTM to perform temporal modeling on the basic feature vectors to capture the temporal dependence relationship of the features;
[0074] Perform dimensionality reduction processing on the feature vectors after temporal modeling to obtain low-dimensional feature representations;
[0075] Based on the feature vectors after dimensionality reduction, use a Softmax classifier for emotion classification; if the confidence of the classification result is higher than the preset threshold, then determine the emotion category; according to the determined emotion category, generate an emotion label dataset; use the emotion label dataset to construct an emotion analysis model to achieve second-level emotion classification.
[0076] Specifically, in this embodiment, MobileNet is used to extract basic feature vectors from the original data to obtain high-dimensional feature representations. Combining with BiLSTM, temporal modeling is performed on the basic feature vectors to capture the temporal dependence relationships of the features. The feature vectors after temporal modeling are dimensionally reduced to obtain low-dimensional feature representations. Based on the dimensionally reduced feature vectors, a Softmax classifier is used for sentiment classification. If the confidence level of the classification result is higher than a preset threshold, the final sentiment category is determined. According to the determined sentiment category, a sentiment label dataset is generated. Using the sentiment label dataset, a sentiment analysis model is constructed to achieve second-level sentiment classification.
[0077] MobileNet feature extraction is a lightweight image recognition architecture that can extract key characterization information for visual data of different scales. For example, in the shopping scenario of face expression collection, multi-dimensional information such as the positions of facial feature points and expression changes can be extracted from face images to form basic feature vectors for subsequent sentiment analysis. Temporal modeling is the dynamic analysis of continuously changing information. Taking the restaurant dining scenario as an example, the continuous behaviors of customers after entering the store may include stages such as ordering, waiting, dining, and leaving the store, and corresponding emotional changes will occur in each stage. Through temporal modeling, the emotional change rules of customers in different stages can be captured. For example, too long waiting time may lead to an increase in negative emotions. Dimensionality reduction helps to extract the most representative feature combinations. In the analysis of shopping mall passenger flow, the original data may include multiple dimensions such as passenger flow, stay time, and purchase behavior. Through dimensionality reduction, the correlations between these features can be discovered. For example, the correlation between passenger flow and purchase behavior may be more significant than the stay time, so as to retain the most valuable feature combinations. Sentiment classification is a quantitative evaluation of the user experience. In the education scenario, the performance of students in class changes over time. By real-time analyzing information such as the expressions and postures of students, their understanding degree and engagement status of the course content can be judged. When it is recognized that the attention of a certain student continues to decrease and the confidence level exceeds the preset threshold, the student will be classified as an object that needs attention. The construction of the sentiment label dataset needs to consider the diversity of scenarios. Taking medical services as an example, patients will experience multiple links such as registration, consultation, and examination during the medical treatment process, and different degrees of emotional fluctuations may occur in each link. By collecting the sentiment labels of each link, a complete medical service experience evaluation system can be constructed. The real-time application of the sentiment analysis model requires the system to have the ability of rapid response. In the intelligent customer service scenario, the system needs to real-time analyze the text input or voice content of users to identify their emotional tendencies. When it is detected that the emotions of users change significantly, the system can adjust the service strategy in a timely manner. For example, transfer to a human customer service when the emotion is negative, and recommend relevant products when the emotion is positive, so as to provide a personalized service experience. This real-time sentiment analysis ability is of great significance to improving service quality and user satisfaction.
[0078] Furthermore, perform in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map, including:
[0079] Use a GPU cluster to load a pre-trained multi-modal joint encoding model to extract features from the input second-level sentiment classification results;
[0080] Fuse the extracted features through the multi-modal joint encoding model to generate a high-dimensional joint feature vector;
[0081] Perform in-depth analysis on the high-dimensional joint feature vector to capture implicit fine-grained information;
[0082] According to the in-depth analysis results, use a clustering algorithm to classify the feature vectors; if the clustering results meet the preset category distribution conditions, determine the sentiment category label; based on the determined sentiment category label, construct the node information of the sentiment map;
[0083] Model the node information of the sentiment map through a graph neural network to generate a fine-grained sentiment map.
[0084] Specifically, the multimodal joint encoding model in this embodiment requires a large amount of computing resources during the pre-training stage, which necessitates the use of a graphics processing unit cluster with powerful parallel computing capabilities. In a typical distributed training scenario, a cluster consisting of eight nodes can be used, with each node equipped with four high-performance graphics processors, and high-speed network interconnection is employed to achieve fast data transmission. For the input second-level sentiment classification results, the model first extracts multimodal features such as audio and text. For example, acoustic features such as fundamental frequency and intensity are extracted from speech, and semantic features such as word vectors are extracted from text. In the feature fusion stage, an attention mechanism is adopted to dynamically allocate weights to features of different modalities. For instance, in anger emotion recognition, the weight of speech features may be higher because the emotional expression in the voice is more direct. The fused high-dimensional joint feature vector usually contains hundreds of dimensions, and each dimension contains rich emotional information. In the in-depth analysis link, dimensionality reduction and visualization techniques can be used to display the distribution law of features. Taking a specific scenario as an example, after projecting the joint feature vector into a three-dimensional space, it can be clearly observed that samples of different emotion categories show a clustering phenomenon in the space, indicating that the model has successfully captured the discriminative features of emotions. The choice of clustering algorithm needs to consider the characteristics of data distribution. In the sentiment analysis scenario, common clustering methods include density clustering algorithms, which can effectively handle data distributions with irregular shapes. The category distribution condition can be set such that the clustering purity is not less than 85%, that is, at least 85% of the samples in each cluster belong to the same emotion category. The construction of node information in the sentiment map needs to consider the correlation between emotions. For example, in social media sentiment analysis, a comment may contain multiple emotions simultaneously, and there are transformation relationships between these emotions. Through graph neural network modeling, the dependency relationships between nodes can be learned to generate fine-grained sentiment feature representations. Specifically, each node can represent a fine-grained emotion, and the edges between nodes represent the correlation strength of emotions. In practical applications, this fine-grained sentiment map can more accurately depict the emotional state of users. For example, in a customer service scenario, the system can identify the emotional change process of users from dissatisfaction to satisfaction, including the initial complaining emotion, the hesitant emotion during the transition period, and the final approving emotion. This fine-grained emotional portrayal helps to understand the emotional demands of users and provides an important reference for improving service quality. Such a sentiment analysis system demonstrates strong robustness in practical applications and can adapt to the emotional recognition requirements in different scenarios.
[0085] Furthermore, generating a sentiment heat map based on the fine-grained sentiment map includes:
[0086] Extracting the intensity value corresponding to the sentiment value based on the fine-grained sentiment map;
[0087] According to the intensity value distribution, converting the intensity value into a color value using a mapping method;
[0088] Generate an emotional heat map based on the correspondence between color values and intensity values.
[0089] Based on the heat map data, construct the layout structure of the visualization dashboard. Optimize and adjust the dashboard using distribution values. Associate the optimized dashboard with the data source. Output the dashboard as a visual emotional heat map through a display method.
[0090] Specifically, to generate the heat map of the emotional spectrum in this embodiment, it is first necessary to extract the intensity values corresponding to the emotional values. For example, in the customer service dialogue scenario, the emotional judgment results of each conversation will carry intensity values. For example, for the angry emotion, it may be manifested as different degrees from slight dissatisfaction to extreme anger. These intensity values are usually represented numerically, such as using floating-point numbers from zero to one to represent emotional intensity. When mapping the emotional intensity to color values, a gradient color system can be used. For negative emotions, a cold color system is adopted, with light blue to dark blue indicating increasing intensity; for positive emotions, a warm color system is adopted, with light yellow to dark red indicating intensity changes. In the medical scenario, the degree of pain when a patient describes symptoms can be represented by different shades of red, helping doctors intuitively grasp the condition. When constructing the heat map, a reasonable layout structure needs to be designed. Taking the smart city scenario as an example, the urban area can be divided into grids, each grid representing a regional unit, and the emotional state of the people in that area is shown through the shade of color. Layout optimization can consider the factor of population density. The grids in densely populated areas can be appropriately densified to provide a more detailed display of emotional distribution. Associating the dashboard with the data source requires establishing a real-time update mechanism. In the analysis of mall customer flow, the real-time customer flow emotional data collected by sensing devices can dynamically update the color distribution on the dashboard. For example, the happy area shows warm orange-red, and the rest area shows peaceful light cyan, helping mall managers timely grasp the customer emotion distribution. When visualizing, a multi-level display method can be adopted. Taking the education scenario as an example, the classroom emotional heat map can be divided into the overall class level and the individual level. At the class level, the overall learning atmosphere is shown through color blocks, and clicking on a specific area can drill down to view the emotional change curve of individual students, which helps teachers accurately grasp the teaching effect. In social media analysis, the emotional heat map can be used to show the spread effect of a certain topic over time and space. The intensity of the emotional response triggered by the topic is intuitively shown through the change in color shade, helping operators identify the hot spots of feedback information and intervene and guide in a timely manner. This multi-dimensional visualization effect can help decision-makers quickly perceive the group emotional situation and achieve precise decision-making and management.
[0091] This embodiment also uses a preset API interface to obtain call requests from external systems. When obtaining requests, it parses the request parameters and extracts relevant data. For the extracted data, it calculates the sentiment value through a pre-established sentiment analysis model. According to the sentiment value result, it determines whether it exceeds the preset threshold range. If the sentiment value exceeds the threshold, it triggers the warning system to generate an alarm message. The warning system sends the alarm message to the specified recipient. After the alarm is issued, it records the alarm message and updates the system log.
[0092] This embodiment also provides a social emotion analysis system based on multimodal data. The system includes: a data acquisition module, a preprocessing module, a preliminary feature extraction module, a second-level sentiment classification module, a depth analysis module, and a sentiment intensity display module;
[0093] The data acquisition module is used to acquire multimodal data; among them, the multimodal data includes: text data, picture data, short video data, and live audio data;
[0094] The preprocessing module is used to preprocess the multimodal data;
[0095] The preliminary feature extraction module is used to perform preliminary feature extraction on the preprocessed multimodal data to generate a basic feature vector;
[0096] The second-level sentiment classification module is used to perform second-level sentiment classification on the basic feature vector;
[0097] The depth analysis module is used to perform in-depth analysis on the second-level sentiment classification result to generate a fine-grained sentiment map;
[0098] The sentiment intensity display module is used to generate a sentiment heat map according to the fine-grained sentiment map. Among them, the sentiment heat map reflects the sentiment intensity through the depth of color and displays the sentiment heat map through a visualization dashboard.
[0099] Furthermore, the preliminary feature extraction module's preliminary feature extraction of the preprocessed multimodal data includes:
[0100] Using MobileNet to obtain a basic feature vector from the preprocessed multimodal data; among them, the basic feature vector includes: the basic feature vector of text data: syntactic analysis results, semantic features of word frequency information; the basic feature vector of picture data: image content features such as color distribution, texture features, shape features, and local region features; the basic feature vector of short video data: the feature vector of each frame of the picture, the feature vector of the audio stream, and the relevant information features of the video duration and background music; the basic feature vector of live audio data: audio content features such as pitch and tone.
[0101] Furthermore, the second-level sentiment classification module's second-level sentiment classification of the basic feature vector includes:
[0102] Use BiLSTM to perform temporal modeling on the basic feature vectors to capture the temporal dependence relationship of the features;
[0103] Perform dimensionality reduction on the feature vectors after temporal modeling to obtain low-dimensional feature representations;
[0104] Based on the dimensionality-reduced feature vectors, use a Softmax classifier for sentiment classification; if the confidence of the classification result is higher than the preset threshold, determine the sentiment category; according to the determined sentiment category, generate a sentiment label dataset; use the sentiment label dataset to construct a sentiment analysis model to achieve second-level sentiment classification.
[0105] Furthermore, the in-depth analysis module conducts in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map, including:
[0106] Use a GPU cluster to load a pre-trained multi-modal joint encoding model to extract features from the input second-level sentiment classification results;
[0107] Fuse the extracted features through the multi-modal joint encoding model to generate high-dimensional joint feature vectors;
[0108] Conduct in-depth analysis on the high-dimensional joint feature vectors to capture the implicit fine-grained information;
[0109] According to the in-depth analysis results, use a clustering algorithm to classify the feature vectors; if the clustering results meet the preset category distribution conditions, determine the sentiment category label; based on the determined sentiment category label, construct the node information of the sentiment map;
[0110] Model the node information of the sentiment map through a graph neural network to generate a fine-grained sentiment map.
[0111] Furthermore, the sentiment intensity display module generates a sentiment heat map according to the fine-grained sentiment map, including:
[0112] Based on the fine-grained sentiment map, extract the intensity values corresponding to the sentiment values;
[0113] According to the intensity value distribution, use a mapping method to convert the intensity values into color values;
[0114] Generate a sentiment heat map through the correspondence between the color values and the intensity values.
[0115] According to the heat map data, construct the layout structure of the visualization dashboard. Optimize and adjust the dashboard using the distribution values. Associate the optimized dashboard with the data source. Output the dashboard as a visual sentiment heat map through the display method.
[0116] Use a preset API interface to obtain call requests from an external system. When obtaining the requests, parse the request parameters and extract relevant data. For the extracted data, calculate the sentiment value through a pre-established sentiment analysis model. According to the sentiment value result, determine whether it exceeds the preset threshold range. If the sentiment value exceeds the threshold, trigger the warning system to generate an alarm message. Send the alarm message to the specified recipient through the warning system. After the alarm is issued, record the alarm message and update the system log.
[0117] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A social sentiment analysis method based on multimodal data, characterized in that: include: Acquire multimodal data; wherein the multimodal data includes: text data, picture data, short video data and live audio data; Preprocessing the multimodal data; Perform preliminary feature extraction on the preprocessed multimodal data to generate basic feature vectors; Performing second-level sentiment classification on the basic feature vector; Conduct in-depth analysis on the second-level sentiment classification results to generate a fine-grained sentiment map; According to the fine-grained emotion map, an emotion heat map is generated, wherein the emotion heat map reflects the intensity of emotion through the depth of color, and the emotion heat map is displayed through a visual dashboard.
2. The social sentiment analysis method based on multimodal data according to claim 1 is characterized in that: Preprocessing the multimodal data includes: Performing data cleaning on the multimodal data to remove invalid information; The format of the cleaned data is standardized.
3. The social emotion analysis method based on multimodal data according to claim 1 is characterized in that: The preliminary feature extraction of preprocessed multimodal data includes: Using MobileNet, basic feature vectors are obtained from the preprocessed multimodal data; wherein the basic feature vectors include: Basic feature vector of text data: including syntactic analysis results and semantic features of word frequency information; Basic feature vector of image data: image content features including color distribution, texture features, shape features and local area features; Basic feature vectors of short video data: including feature vectors of each frame, feature vectors of audio streams, and related information features of video duration and background music; Basic feature vector of live audio data: including audio content features such as tone and pitch.
4. The social emotion analysis method based on multimodal data according to claim 1, characterized in that: Performing second-level sentiment classification on the basic feature vector includes: Use BiLSTM to perform time series modeling on basic feature vectors to capture the temporal dependency of features; Perform dimensionality reduction on the feature vector after time series modeling to obtain low-dimensional feature representation; Based on the feature vector after dimensionality reduction, the Softmax classifier is used for sentiment classification; if the confidence of the classification result is higher than the preset threshold, the sentiment category is determined; according to the determined sentiment category, a sentiment label data set is generated; using the sentiment label data set, a sentiment analysis model is constructed to achieve sentiment classification in seconds.
5. The social sentiment analysis method based on multimodal data according to claim 1, characterized in that: In-depth analysis of the second-level sentiment classification results to generate fine-grained sentiment maps includes: Use GPU cluster to load pre-trained multimodal joint encoding model to extract features from input second-level sentiment classification results; The extracted features are fused through a multimodal joint coding model to generate a high-dimensional joint feature vector; Deeply analyze high-dimensional joint feature vectors to capture implicit fine-grained information; According to the results of deep analysis, a clustering algorithm is used to classify the feature vectors; if the clustering results meet the preset category distribution conditions, the emotional category labels are determined; based on the determined emotional category labels, the node information of the emotional map is constructed; The node information of the sentiment graph is modeled through graph neural network to generate a fine-grained sentiment graph.
6. The social sentiment analysis method based on multimodal data according to claim 1, characterized in that: Based on the fine-grained sentiment map, generating a sentiment heat map includes: Based on the fine-grained sentiment map, extract the intensity value corresponding to the sentiment value; According to the intensity value distribution, a mapping method is used to convert the intensity value into a color value; The emotional heat map is generated through the correspondence between color value and intensity value.
7. A social emotion analysis system based on multimodal data, characterized in that: Used to implement the social emotion analysis method based on multimodal data as described in any one of claims 1-6, the system includes: a data acquisition module, a preprocessing module, a preliminary feature extraction module, a second-level emotion classification module, a deep analysis module and an emotion intensity display module; The data acquisition module is used to acquire multimodal data; wherein the multimodal data includes: text data, picture data, short video data and live audio data; The preprocessing module is used to preprocess the multimodal data; The preliminary feature extraction module is used to perform preliminary feature extraction on the preprocessed multimodal data to generate a basic feature vector; The second-level sentiment classification module is used to perform second-level sentiment classification on the basic feature vector; The deep analysis module is used to perform deep analysis on the second-level sentiment classification results to generate a fine-grained sentiment map; The emotion intensity display module is used to generate an emotion heat map according to the fine-grained emotion map, wherein the emotion heat map reflects the emotion intensity through the depth of color, and displays the emotion heat map through a visual dashboard.
8. The social emotion analysis system based on multimodal data according to claim 7, characterized in that: The preliminary feature extraction module performs preliminary feature extraction on the preprocessed multimodal data, including: MobileNet is used to obtain basic feature vectors from the preprocessed multimodal data; wherein the basic feature vectors include: basic feature vectors for text data: semantic features of syntactic analysis results and word frequency information; basic feature vectors for picture data: image content features of color distribution, texture features, shape features and local area features; basic feature vectors for short video data: feature vectors of each frame, feature vectors of the audio stream, and related information features of the video duration and background music; basic feature vectors for live audio data: audio content features of tone and pitch.
9. The social sentiment analysis system based on multimodal data according to claim 7, characterized in that: The second-level sentiment classification module performs second-level sentiment classification on the basic feature vector, including: Use BiLSTM to perform time series modeling on basic feature vectors to capture the temporal dependency of features; Perform dimensionality reduction on the feature vector after time series modeling to obtain low-dimensional feature representation; Based on the feature vector after dimensionality reduction, the Softmax classifier is used for sentiment classification; if the confidence of the classification result is higher than the preset threshold, the sentiment category is determined; according to the determined sentiment category, a sentiment label data set is generated; using the sentiment label data set, a sentiment analysis model is constructed to achieve sentiment classification in seconds.
Citation Information
Cited By
Image sentiment analysis method fusing multi-dimensional features
CN120656223A
Multi-modal convergence media generation system and method based on dynamic emotion map
CN121030019A
A multi-modal fusion media generation system and method based on a dynamic emotion map
CN121030019B