Multi-modal sentiment analysis method and device based on multi-agent collaboration
By performing feature extraction and mapping of multimodal data packets, the differences between feature vectors are detected and evaluated, and the problem of lack of effective difference detection and evaluation mechanisms in the prior art is solved, and more reliable and accurate sentiment analysis results are achieved.
Patent Information
- Application Number
- CN202510103529.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal emotion analysis method that coordinates multi-agents lacks an effective detection and evaluation mechanism for the degree of difference between the characteristic vectors generated by different agents, resulting in bias in sentiment analysis results.
By clocking synchronization of speech signals, facial expression video frames and physiological parameters, the feature vectors of multimodal data packets are extracted, and mapped with the preset emotional ontology, the degree of difference between multimodal emotional vectors is detected, the conflict detection algorithm based on feature credibility is started for evaluation, and the feature fusion operation is performed based on the evaluation results and weight coefficients is performed to obtain unified emotional recognition results.
Effectively prevent abnormal feature vectors from affecting the fusion results, improve the reliability of multimodal sentiment analysis, and ensure the accuracy and credibility of sentiment analysis results.
Smart Images

Figure CN119908724A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-agent collaboration technology, and in particular to a multimodal sentiment analysis method and device based on multi-agent collaboration. Background Art
[0002] With the development of human-computer interaction technology, multimodal sentiment analysis has been widely used in the fields of intelligent medical care, online education and human-computer interaction. The current mainstream multimodal sentiment analysis methods usually use multiple agents to process data of different modalities separately, such as speech agents responsible for analyzing speech features, facial agents responsible for analyzing expression features, and physiological agents responsible for analyzing physiological signal features, and then simply fuse the processing results of each agent to obtain the final sentiment analysis result. This method performs well when processing single-modal data, but has obvious shortcomings in the multi-agent collaboration process. Due to the lack of an effective detection and evaluation mechanism for the degree of difference between the feature vectors generated by different agents, it is impossible to timely detect anomalies generated during the acquisition or processing of certain modal data. These abnormal feature vectors will be directly brought into the fusion process, which will eventually lead to significant deviations in the sentiment analysis results and affect the reliability of the system. Summary of the invention
[0003] The main purpose of the present invention is to solve the technical problem that the existing multi-agent collaborative multimodal sentiment analysis method lacks an effective detection and evaluation mechanism for the degree of difference between feature vectors generated by different agents, which leads to deviations in sentiment analysis results. A first aspect of the present invention provides a multimodal sentiment analysis method based on multi-agent collaboration, and the multimodal sentiment analysis method based on multi-agent collaboration includes: Perform clock synchronization acquisition and preprocessing on speech signals, facial expression video frames and physiological parameters to obtain multimodal data packets with time sequence marks; Using a voice agent, a facial agent, and a physiological agent to extract features from the multimodal data packet, respectively, and mapping the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; The difference between the vectors in the multimodal sentiment vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started to evaluate and obtain an evaluation result; A feature fusion operation is performed on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and the emotion recognition result is stored in association with the corresponding multimodal original data by timestamp.
[0004] Optionally, in a first implementation of the first aspect of the present invention, the use of a voice agent, a facial agent, and a physiological agent to extract features from the multimodal data packet respectively, and mapping the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space includes: Perform Fourier transform and Mel spectrum analysis on speech signals, and use convolutional neural network to extract phoneme and prosodic features to obtain speech feature vectors; Perform facial key point detection and spatial feature extraction on facial expression video frames to obtain facial expression feature vectors; Perform wavelet transform and frequency domain analysis on physiological parameters to obtain physiological feature vectors; According to the emotion label system in the preset emotion ontology knowledge base, the speech feature vector, the facial expression feature vector and the physiological feature vector are semantically mapped to obtain a multimodal emotion vector in a unified emotion semantic space.
[0005] Optionally, in a second implementation of the first aspect of the present invention, the emotion tag system in the preset emotion ontology knowledge base is used to semantically map the speech feature vector, the facial expression feature vector, and the physiological feature vector to obtain a multimodal emotion vector in a unified emotion semantic space, including: Normalizing the speech feature vector, the facial expression feature vector and the physiological feature vector to obtain a standardized feature vector; Performing a projection transformation on the standardized feature vector according to the basic emotion tags in a preset emotion ontology knowledge base to obtain an initial semantic mapping vector; Performing a nonlinear activation function operation on the initial semantic mapping vector to obtain an activated semantic vector, and performing a combination operation on the activated semantic vector according to a composite emotion rule in a preset emotion ontology knowledge base to obtain a multi-dimensional emotion feature; The multi-dimensional emotional features are subjected to dimensionality reduction processing to obtain a multi-modal emotional vector in a unified emotional semantic space.
[0006] Optionally, in a third implementation of the first aspect of the present invention, the degree of difference between vectors in the multimodal emotion vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started for evaluation, and the evaluation result obtained includes: Constructing a time series sliding window for the multimodal emotion vector, calculating the autocorrelation coefficient and the mutual correlation coefficient of the emotion vector of each mode in the time series sliding window, and obtaining a correlation matrix; Using a preset spectral clustering algorithm to calculate the correlation between the sentiment vectors of each modality according to the correlation matrix, and obtaining a difference index; According to the difference index, a pre-trained anomaly detection network is called to perform comparative analysis on the multimodal sentiment vector to obtain an abnormal vector label, and a wavelet transform is performed on the sentiment vector labeled as abnormal to obtain a local fluctuation coefficient and a global trend coefficient; The credibility score of the sentiment vector marked as abnormal is calculated according to the local fluctuation coefficient and the global trend coefficient using the preset hierarchical feature evaluation rule, and the credibility score is weighted and corrected according to the historical statistical indicators of the multimodal sentiment vector to obtain an evaluation result.
[0007] Optionally, in a fourth implementation of the first aspect of the present invention, the credibility score of the emotion vector marked as abnormal is calculated according to the local fluctuation coefficient and the global trend coefficient using the preset hierarchical feature evaluation rule, and the credibility score is weighted and corrected according to the historical statistical indicators of the multimodal emotion vector, and the evaluation result includes: Performing frequency band decomposition on the local fluctuation coefficient to obtain a multi-scale fluctuation feature, and calculating a signal stability index according to the multi-scale fluctuation feature and a global trend coefficient to obtain a basic credibility value; Performing sliding window statistics on the historical data of the multimodal sentiment vector to obtain a historical fluctuation range; The basic credibility value is interval mapped according to the historical fluctuation range to obtain a correction coefficient, and the basic credibility value is corrected and calculated using the correction coefficient to obtain an evaluation result.
[0008] Optionally, in a fifth implementation of the first aspect of the present invention, the feature fusion operation of the multimodal emotion vector according to the evaluation result and the preset weight coefficient to obtain a unified emotion recognition result, and the emotion recognition result and the corresponding multimodal original data are stored by timestamp association, including: Dynamically adjusting the preset weight coefficient according to the evaluation result to obtain a modified weight coefficient, and using the modified weight coefficient to perform a weighted fusion operation on the multimodal emotion vector to obtain a fused feature vector; Calculating the probability distribution of each emotion category according to the fused feature vector to obtain a unified emotion recognition result; A unified timestamp is added to the emotion recognition result and the corresponding multimodal original data, and the data is stored in a distributed database according to a time series structure to obtain emotion analysis data with traceability marks.
[0009] Optionally, in a sixth implementation of the first aspect of the present invention, dynamically adjusting the preset weight coefficient according to the evaluation result to obtain the modified weight coefficient includes: Calculating the relative credibility ratios between the modes according to the evaluation results to obtain a credibility ratio matrix, and constructing a weight adjustment matrix using the credibility ratio matrix to obtain weight constraint conditions; Performing linear programming to solve the preset weight coefficient according to the weight constraint condition to obtain an initial adjustment coefficient, and applying an exponential smoothing algorithm to the initial adjustment coefficient to obtain a stabilized adjustment parameter with time series continuity; A matrix multiplication operation is performed on the stabilization adjustment parameter and the preset weight coefficient to obtain an adjusted weight value, and a softmax function operation is performed on the adjusted weight value to obtain a modified weight coefficient whose sum is 1.
[0010] A second aspect of the present invention provides a multimodal sentiment analysis device based on multi-agent collaboration, the multimodal sentiment analysis device based on multi-agent collaboration comprising: The data acquisition module is used to perform clock synchronization acquisition and preprocessing of speech signals, facial expression video frames and physiological parameters to obtain multimodal data packets with time sequence marks; A feature mapping module is used to extract features from the multimodal data packet using a voice agent, a facial agent, and a physiological agent, and to map the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; The difference detection module is used to detect the degree of difference between vectors in the multimodal sentiment vector, and when a significant difference is detected, the conflict detection algorithm based on feature credibility is started to evaluate and obtain the evaluation result; The result fusion module is used to perform a feature fusion operation on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and store the emotion recognition result in association with the corresponding multimodal original data by timestamp.
[0011] The above-mentioned multimodal sentiment analysis method and device based on multi-agent collaboration obtains a multimodal data packet with a time sequence mark by performing clock synchronization acquisition and preprocessing of multimodal data. Multiple agents are used to extract features from the multimodal data packet respectively, and the extracted feature vectors are mapped with the preset emotion ontology to obtain a multimodal sentiment vector. The degree of difference between the vectors in the multimodal sentiment vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started to evaluate and obtain an evaluation result. According to the evaluation result and the preset weight coefficient, a feature fusion operation is performed on the multimodal sentiment vector to obtain an emotion recognition result, and the emotion recognition result is stored in association with the corresponding multimodal data by timestamp. The present invention effectively prevents abnormal feature vectors from affecting the fusion result by introducing a difference detection and evaluation mechanism, thereby improving the reliability of modal sentiment analysis.
[0012] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0013] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of a first embodiment of a multimodal sentiment analysis method based on multi-agent collaboration in an embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of a multimodal sentiment analysis device based on multi-agent collaboration in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0017] To facilitate understanding of this embodiment, a multimodal sentiment analysis method based on multi-agent collaboration disclosed in an embodiment of the present invention is first introduced in detail. Figure 1 As shown, the method comprises the following steps: 101. Perform clock synchronization acquisition and preprocessing on voice signals, facial expression video frames, and physiological parameters to obtain multimodal data packets with time sequence marks; In one embodiment of the present invention, it is necessary to consider how to ensure that the three modal data can be accurately synchronized during acquisition. In order to ensure the time consistency of the data, a unified time base should be set for each sensor or device first, usually using a high-precision synchronous clock, such as a GPS clock or an atomic clock, which can provide stable and consistent time tags. For voice signal acquisition, a microphone with a high sampling rate is used and recorded through the internal clock of the acquisition device. Each voice data point should be accompanied by a timestamp at that moment to ensure that the voice signal can be synchronized with other modal data in time. For facial expression video frames, facial images are collected by a high frame rate camera. Each frame of the image will be marked with an accurate timestamp during acquisition. The timestamp is synchronized with the time base of the voice signal to ensure that the timeline of the video frame corresponds to the voice signal. The acquisition of physiological parameters is usually performed in real time through sensors, such as heart rate sensors or skin temperature sensors, through a dedicated data acquisition system, and each time point of these data will also be added with a corresponding timestamp. The clock synchronization of physiological signals can be adjusted by hardware synchronization or by software algorithms so that the timestamps of these data are consistent with those of voice signals and facial expression video frames.
[0018] After data collection is completed, the data of each modality needs to be preprocessed. For speech signals, denoising and gain control must be performed first to remove background noise and environmental interference to ensure the clarity of speech signals. Facial expression video frames need to undergo image enhancement and denoising, which includes but is not limited to adjusting the brightness, contrast and sharpness of the video frames to improve the accuracy of subsequent facial expression recognition. Physiological parameter data also needs to be processed by removing outliers and noise filtering to ensure data accuracy. All preprocessed data will be marked and organized according to their respective timestamps to ensure that each data point has an accurate time stamp. Finally, the data of the three modalities will be sorted by timestamp and merged into a multimodal data packet to form a structured time series data set, in which each data packet contains speech, video frames and physiological signal data, and these data can accurately reflect the emotional state at the same time point.
[0019] 102. Using a voice agent, a facial agent, and a physiological agent to extract features from the multimodal data packet, respectively, and mapping the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; In one embodiment of the present invention, the use of a speech agent, a facial agent, and a physiological agent to extract features from the multimodal data packet respectively, and mapping the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space includes: performing Fourier transform and Mel spectrum analysis on the speech signal, and extracting phonemes and rhythmic features using a convolutional neural network to obtain a speech feature vector; performing facial key point detection and spatial feature extraction on facial expression video frames to obtain a facial expression feature vector; performing wavelet transform and frequency domain analysis on physiological parameters to obtain a physiological feature vector; and semantically mapping the speech feature vector, facial expression feature vector, and physiological feature vector according to an emotion label system in a preset emotion ontology knowledge base to obtain a multimodal emotion vector in a unified emotion semantic space.
[0020] Specifically, the speech signal is first collected by a high-precision microphone. The collected signal usually needs to be Fourier transformed to convert it from the time domain to the frequency domain in order to better analyze its frequency components. In the specific implementation, the fast Fourier transform (FFT) algorithm is used to process the collected speech signal to obtain its spectrum information. Next, the mel spectrum data is further processed using mel-frequency cepstral coefficients (MFCC) to map the spectrum to the mel scale to simulate the auditory perception of the human ear. In this step, a fixed mel filter bank is used to weight the spectrum, and then the weighted spectrum is logarithmically transformed and its cepstral coefficients are extracted to obtain the MFCC features of the speech. Next, these MFCC features are processed using a convolutional neural network (CNN) to extract phoneme and prosodic features. The convolutional neural network extracts features from the MFCC features through multiple layers of convolution kernels, captures the phonemes and prosodic patterns in the speech, such as the pitch and speed of speech, and finally generates a speech feature vector containing phoneme and prosodic features. This vector represents the emotional information in the speech signal and can reflect the speaker's emotional state.
[0021] Specifically, in the feature extraction process of facial expressions, firstly, a video frame of facial expressions is collected by a high frame rate camera, and each frame needs to be detected to locate the facial area. This process usually uses the Haar cascade classifier in OpenCV or the MTCNN algorithm based on deep learning. After the face is detected, the facial key points are calibrated. The geometric features of the face are obtained by detecting facial feature points such as eyes, eyebrows, and mouth. Then, the relative displacement between these feature points is calculated to describe the changes in facial expressions. For example, by calculating the upward degree of the corners of the mouth, the upward degree of the eyebrows, the degree of eye opening, etc., the emotional features in facial expressions can be captured. The extracted facial expression features can be reduced in dimension using PCA (principal component analysis) or LDA (linear discriminant analysis) to obtain a feature vector of facial expression. This vector contains the spatial features reflecting the emotions in the expression and represents the emotional information conveyed by the facial expression at the current moment.
[0022] Specifically, the feature extraction of physiological signals depends on various physiological sensors, such as heart rate sensors, skin galvanic response sensors, etc. The raw data collected by these sensors usually contain large noise and irregular fluctuations, so denoising is required first. When performing wavelet transform on physiological signals, discrete wavelet functions such as Daubechies wavelet or Haar wavelet are used to effectively decompose the signal into components of different frequency bands. This process uses the multi-scale analysis characteristics of wavelet transform to extract the low-frequency and high-frequency components of the signal. In specific implementation, the collected physiological signals are first subjected to discrete wavelet transform to decompose them into wavelet coefficients of multiple scales. By further analyzing these coefficients, emotion-related information such as periodic fluctuations of heartbeats or instantaneous changes in skin galvanic response can be extracted. In addition, when performing frequency domain analysis on the signal, the fast Fourier transform (FFT) is used to convert the signal to the frequency domain to extract the frequency components of heart rate or skin galvanic response, which can reflect emotional fluctuations. For example, rapid changes in heart rate or fluctuations in skin galvanic response may be related to emotional states such as anxiety and fear. The processed physiological signal feature vector contains information such as heart rate and skin electrical response, reflecting the current physiological state and emotional response of the individual.
[0023] Specifically, finally, the feature vectors from different modalities are mapped to the emotion labels in the preset emotion ontology knowledge base. First, the feature vectors of each modality are standardized to ensure that the feature scales of each modality are consistent. Then, according to the emotion label system in the emotion ontology, a specific mapping algorithm is used to associate the feature vectors of each modality with the emotion label. In specific implementation, supervised learning methods such as support vector machines (SVM) or neural networks can be used to learn the mapping relationship between features and emotion labels through training data. Speech feature vectors, facial expression feature vectors, and physiological feature vectors will be input into these models, and after mapping, they will be uniformly mapped to a shared emotion semantic space. This space contains the emotion information of all modal data, and a comprehensive emotion feature vector can be obtained by calculating the similarity between each modality and the emotion label. Finally, these multimodal emotion feature vectors will converge together to form a unified multimodal emotion vector Furthermore, the method of semantically mapping the speech feature vector, facial expression feature vector and physiological feature vector according to the emotion label system in the preset emotion ontology knowledge base to obtain a multimodal emotion vector in a unified emotion semantic space includes: normalizing the speech feature vector, facial expression feature vector and physiological feature vector to obtain a standardized feature vector; projecting the standardized feature vector according to the basic emotion label in the preset emotion ontology knowledge base to obtain an initial semantic mapping vector; performing a nonlinear activation function operation on the initial semantic mapping vector to obtain an activated semantic vector, and combining the activated semantic vector according to the composite emotion rule in the preset emotion ontology knowledge base to obtain a multidimensional emotion feature; and performing dimensionality reduction processing on the multidimensional emotion feature to obtain a multimodal emotion vector in a unified emotion semantic space.
[0024] Specifically, the feature vector of each modality is normalized. Specifically, for each modality feature vector, the mean and standard deviation of the vector are first calculated, and then normalized using the following formula: each eigenvalue minus the mean and then divided by the standard deviation. The purpose of this normalization method is to unify the scale of all eigenvalues, thereby avoiding the large or small numerical range of some features, which leads to deviations in the subsequent feature fusion process. For example, in speech features, the numerical range of pitch and volume may be large, while some features of facial expressions may fluctuate within a smaller range. Through normalization, it can be ensured that the features of all modalities are compared at the same scale, improving the accuracy of subsequent feature mapping.
[0025] Next, emotion labels are mapped to the standardized speech, facial expressions, and physiological feature vectors. The emotion ontology knowledge base contains basic emotion labels, such as "happy", "angry", "sad", etc., and these emotion labels are usually associated with specific feature vectors in a certain way. For example, certain frequency bands, volume, and rhythmic features of speech signals may correspond to certain emotions, while features such as the degree of eye openness and mouth shape in facial expressions may indicate different emotions. To this end, the feature vector of each modality is projected through a preset mapping matrix to map it to the emotion label space. The specific operation of the projection process can be achieved through matrix multiplication. For example, assuming that the speech feature vector is , the facial expression feature vector is , the physiological feature vector is , then each feature vector is multiplied by the corresponding emotion label mapping matrix to obtain the initial emotion representation. Assume that the emotion label matrix is , then the projection operation is: , this operation maps each modal feature to the label space of the emotion ontology to obtain the initial emotion semantic vector. After obtaining the initial semantic mapping vector, it needs to be processed by a nonlinear activation function. The purpose of the nonlinear activation function is to introduce more complex nonlinear relationships so that the emotion vector can better reflect the complexity of real emotions. At this stage, activation functions such as ReLU, siqmoid or tanh are usually used. Taking ReLU as an example, the calculation process is: for each vector element, if the element is greater than 0, retain the value; if it is less than 0, set it to 0. Specifically, in the initial semantic mapping vector of each modality Apply ReLU activation one by one to obtain the activated semantic vector . For example, if some elements in the initial semantic mapping vector represent the intensity of negative emotions, the ReLU function "truncates" them and only retains the positive emotional part, which ensures that the emotional map can more effectively reflect the intensity and tendency of the emotion. Further, based on the composite emotion rules in the emotional ontology knowledge base, the activated semantic vectors are combined. Composite emotion rules mean that some emotional labels can be obtained by combining multiple basic emotions. For example, the emotional ontology may stipulate that "anger" is a combination of the two emotional labels "anger" and "anxiety". In this step, the activated semantic vectors are combined according to these composite rules. This can be achieved by simple addition or weighted averaging. For example, for the emotional label "anger", the activation vectors of its voice, facial expression and physiological signals can be weighted and synthesized to obtain the composite map of the emotion. Specifically, for each emotional label, the activation vectors of each modality are weighted and summed by a given weight, for example: ,in , and γ are preset weight coefficients. In this way, the composite emotional features of each emotional label can be obtained, which characterizes the emotional interaction between different modalities. Finally, the obtained multidimensional emotional features are subjected to dimensionality reduction. The purpose of dimensionality reduction is to map the high-dimensional emotional vector to a low-dimensional space for subsequent emotional classification or recognition. At this stage, methods such as principal component analysis (PCA) or t-SNE can be used. Taking PCA as an example, the covariance matrix of the emotional feature matrix is first calculated, and the eigenvalue decomposition is performed on it, and the first few principal components are selected to form a new low-dimensional space. In this way, the multidimensional feature vector of each emotional label can be converted into a lower-dimensional vector, which not only retains the main information of the emotion but also reduces the computational burden.
[0026] 103. Detect the degree of difference between vectors in the multimodal sentiment vector, and when a significant difference is detected, start a conflict detection algorithm based on feature credibility to perform evaluation and obtain an evaluation result; In one embodiment of the present invention, the degree of difference between vectors in the multimodal emotion vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started for evaluation, and the evaluation result obtained includes: constructing a time series sliding window for the multimodal emotion vector, calculating the autocorrelation coefficient and the mutual correlation coefficient of the emotion vector of each modality in the time series sliding window, and obtaining a correlation matrix; using a preset spectral clustering algorithm to calculate the correlation between the emotion vectors of each modality according to the correlation matrix, and obtaining a difference index; calling a pre-trained anomaly detection network to perform comparative analysis on the multimodal emotion vector according to the difference index to obtain an abnormal vector label, and performing a wavelet transform on the emotion vector marked as abnormal to obtain a local fluctuation coefficient and a global trend coefficient; using a preset hierarchical feature evaluation rule to calculate the credibility score of the emotion vector marked as abnormal according to the local fluctuation coefficient and the global trend coefficient, and performing a weighted correction on the credibility score according to the historical statistical index of the multimodal emotion vector to obtain an evaluation result.
[0027] Specifically, a time series sliding window is first constructed. Assuming that each emotion vector corresponds to a timestamp, a suitable time window size is first selected, such as 100ms. Then, the emotion vectors at each moment and a certain time before and after are gradually intercepted by the sliding window method. For example, in each time window, the emotion vectors of three modes (voice, facial expression, physiological parameters) are collected to obtain a time period data set containing these emotion vectors. The window starts from the starting position of the time axis and slides backward, each time sliding a fixed length (such as 10ms) until the end of the entire data sequence. For each sliding window, the time consistency between the vectors in each modality is analyzed by calculating the autocorrelation coefficients of the emotion vectors of different modalities in the window. The calculated autocorrelation coefficients and mutual correlation coefficients constitute a correlation matrix, which is used to describe the correlation and synchronization between the emotion vectors of different modalities.
[0028] Specifically, next, the correlation between modalities is calculated based on the correlation matrix through the spectral clustering algorithm. First, the correlation matrix is input into the spectral clustering algorithm. The spectral clustering algorithm decomposes the eigenvalues of the matrix to obtain the similarity matrix between the modalities, and then maps it to a low-dimensional space to separate different emotional patterns. Specifically, the algorithm constructs the correlation matrix as a Laplace matrix of the graph, and classifies different emotional vectors into different clusters through the eigenvectors of the graph. The calculated emotional vector cluster reflects the difference index between the emotional vectors of each modality. If an emotional vector fails to be clustered into a reasonable cluster, it means that the vector is greatly different from the emotional vectors of other modalities in some aspects. This difference index will be used for subsequent anomaly detection.
[0029] Specifically, after the difference index is calculated, the pre-trained anomaly detection network is started. The role of this network is to perform deep learning analysis on sentiment vectors with large differences to determine whether they are abnormal sentiment vectors. Specifically, the pre-trained network may be based on LSTM (Long Short-Term Memory Network) or autoencoder. These networks are trained with a large amount of sentiment data and can identify abnormal patterns in sentiment vectors. The input of the anomaly detection network is a multimodal sentiment vector, and the output is an abnormal label. If the sentiment vector shows a significant deviation compared with the historical sample, the sentiment vector is marked as abnormal. The sentiment vector marked as abnormal will be further processed by wavelet transform. Wavelet transform is used to extract frequency domain and time domain features in the sentiment vector to help distinguish local fluctuations from global trends. Specifically, the sentiment vector is decomposed in the time domain and frequency domain to obtain fluctuation coefficients and trend coefficients at multiple scales. The results after wavelet transform help reveal the changing characteristics of the sentiment vector at different scales. For example, the local fluctuation coefficient can reflect the local changes of the sentiment vector, while the global trend coefficient reflects the overall trend of the sentiment vector.
[0030] Specifically, finally, the credibility score of the emotion vector marked as abnormal is calculated by combining the local fluctuation coefficient and the global trend coefficient using the preset hierarchical feature evaluation rules. In the hierarchical evaluation process, different weights are given according to different emotion modes. For emotion vectors with large local fluctuations, their weights in the evaluation can be increased because these emotion changes are usually more influential. The evaluation process combines the weighted calculation of the fluctuation coefficient and the trend coefficient, and weights them according to the preset rules to form a comprehensive credibility score. This credibility score indicates whether the emotion vector belongs to a reliable emotion judgment. If the score is low, it means that the emotion vector may be misjudged due to abnormal fluctuations or external interference, and further correction or elimination is required. In order to further improve the credibility, weighted correction is performed based on the historical data of multimodal emotion vectors. Through statistical analysis of historical data, the fluctuation range and credibility range of each emotion vector under different emotion modes are obtained. According to this historical information, the credibility score of the current emotion vector is adjusted to obtain the final evaluation result.
[0031] Furthermore, the preset hierarchical feature evaluation rules are used to calculate the credibility score of the sentiment vector marked as abnormal according to the local fluctuation coefficient and the global trend coefficient, and the credibility score is weightedly corrected according to the historical statistical indicators of the multimodal sentiment vector, and the evaluation result includes: performing frequency band decomposition on the local fluctuation coefficient to obtain multi-scale fluctuation characteristics, and calculating the signal stability index according to the multi-scale fluctuation characteristics and the global trend coefficient to obtain a basic credibility value; performing sliding window statistics on the historical data of the multimodal sentiment vector to obtain a historical fluctuation range; performing interval mapping on the basic credibility value according to the historical fluctuation range to obtain a correction coefficient, and using the correction coefficient to perform correction calculation on the basic credibility value to obtain an evaluation result.
[0032] Specifically, when implementing the technical solution of processing the local fluctuation coefficient and the global trend coefficient using the preset hierarchical feature evaluation rules and correcting the credibility score with historical data, the local fluctuation coefficient of the emotion vector marked as abnormal is first decomposed by frequency band to obtain multi-scale fluctuation characteristics. The specific operation includes mapping the local fluctuation coefficient of the emotion vector to different frequency bands under the framework of discrete wavelet transform or empirical mode decomposition, and extracting multi-scale fluctuation characteristics that can reflect information such as short-term violent fluctuations or relatively stable sections by quantifying the energy distribution or morphological structure of each frequency band. To achieve this decomposition process, the local fluctuation coefficient can be multi-level filtered in the time domain, first splitting the high-frequency and low-frequency parts layer by layer, and then combining threshold elimination and reconstruction operations, so that various features in the emotion vector can be presented separately at different scales. Subsequently, the fluctuation characteristics of each frequency band generated after the decomposition are operated with the global trend coefficient to calculate a signal stability index. The global trend coefficient can be extracted using polynomial fitting or first-order difference accumulation methods, such as fitting a trend curve on the emotional time series and calculating the slope of the curve between adjacent sample points to obtain a value that characterizes the long-term trend. When these fluctuation characteristics are combined with long-term trend information, a comprehensive stability function can be designed to perform weighted summation or nonlinear combination of the local fluctuations in each frequency band and the stability of the overall trend, output a value that measures the confidence level of the emotional vector, and define this value as the basic credibility value. When the proportion of local high-frequency fluctuations is small and the overall trend is relatively uniform, the basic credibility value is at a high level, indicating that the fluctuation amplitude of the emotional vector is within a reasonable range and is suitable for further evaluation. If the high-frequency component energy accounts for a large proportion or the global trend shows an extreme slope change, the basic credibility value decreases, indicating that the emotional vector has significant anomalies in the time dimension or frequency dimension.
[0033] Specifically, after obtaining the basic credibility value, the historical data of the multimodal sentiment vector is subjected to sliding window statistics to obtain the historical fluctuation range. When implementing this step, a fixed-length observation interval can be set, and the intervals can be slid in sequence on the time axis at a constant step length to intercept the sentiment vectors in each interval and their corresponding credibility records. In each sliding window, the local fluctuation and global trend characteristics of the sentiment vector are processed in the same way, and the basic credibility distribution in the past period is obtained. By performing range analysis, quantile calculation, and mean and variance estimation on these historical basic credibility values, a series of interval boundaries or statistical indicators can be obtained to characterize the credibility fluctuation range of the modality or the group under normal emotional fluctuation conditions. For scenarios with sufficient data support, the sentiment vectors in each sliding window can be decomposed and fitted multiple times to accumulate more complete statistical information, such as the maximum and minimum values of the high-low frequency energy ratio, the distribution kurtosis and skewness of the global trend indicator in multiple windows in the past, etc. These quantitative indicators help to depict the fluctuation amplitude and overall trend of the sentiment vector under historical conditions, thereby providing a comparative benchmark for the credibility score of the current abnormal sentiment vector. This sliding window statistical method can also be used to continuously track the time level, and to judge the degree of difference between the current emotional state and the historical typical state by comparing different time segments. If the local fluctuation or trend slope of the current emotional vector is far beyond the common range in historical window statistics, it will often be reflected as a more obvious correction need in the subsequent steps.
[0034] Specifically, after obtaining the historical fluctuation range, the basic credibility value is interval mapped according to the range to obtain the correction coefficient, and the basic credibility value is corrected and calculated using the correction coefficient to generate the final evaluation result. When implementing this process, several intervals can be delineated within the historical fluctuation range, such as a low credibility interval, a medium credibility interval, and a high credibility interval, and the upper and lower boundaries corresponding to these intervals are compared with the current basic credibility value. If the basic credibility value is in the high credibility interval, the correction coefficient can be made to tend to a value slightly higher than 1, so that the current value maintains a high confidence in the final evaluation link; if the basic credibility value falls in the low credibility interval, the correction coefficient should be given a smaller correction force, suggesting that the credibility of the abnormal emotion vector is relatively limited. In order to make the correction process smooth, the difference between the historical fluctuation range and the current value can be mapped into a continuous correction coefficient based on interpolation or piecewise linear function. For example, by performing linear interpolation within the historical interval range, and calculating the correction coefficient according to the location of the basic credibility value, and then multiplying or nonlinearly combining the coefficient with the original basic credibility value, the final credibility evaluation index is formed. In this way, we can effectively avoid extreme judgments caused by abnormal data at individual moments, and give more reasonable credibility scores when the sentiment vector is close to the historical common state. The final evaluation result is a credibility measure that takes into account the actual situation of the current sentiment vector and the historical fluctuation statistics, which can play a role in screening and identifying abnormal sentiment information in the multimodal collaborative sentiment analysis scenario.
[0035] 104. Perform a feature fusion operation on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and store the emotion recognition result and the corresponding multimodal original data in association with a timestamp.
[0036] In one embodiment of the present invention, the feature fusion operation of the multimodal emotion vector is performed according to the evaluation result and the preset weight coefficient to obtain a unified emotion recognition result, and the emotion recognition result is stored in association with the corresponding multimodal original data with a timestamp, including: dynamically adjusting the preset weight coefficient according to the evaluation result to obtain a corrected weight coefficient, and using the corrected weight coefficient to perform a weighted fusion operation on the multimodal emotion vector to obtain a fused feature vector; calculating the probability distribution of each emotion category according to the fused feature vector to obtain a unified emotion recognition result; adding a unified timestamp identifier to the emotion recognition result and the corresponding multimodal original data, and storing them in a distributed database according to a time series structure to obtain emotion analysis data with traceability marks.
[0037] Specifically, the preset weight coefficient is dynamically adjusted according to the evaluation result to obtain the modified weight coefficient, and the multimodal emotion vector is weighted fused by using the modified weight coefficient to obtain the fused feature vector. The implementation of this step requires pre-defining a set of initial weight coefficients in the system configuration to identify the relative importance of each modality in the emotion analysis. These weight coefficients are usually obtained by statistics of historical data or induction of expert knowledge. After the multimodal emotion vector is evaluated for credibility, an evaluation result is output, which contains the validity or abnormality information of each modality at the current moment or in the current window. The dynamic adjustment process is usually implemented by a method based on matrix operation or optimization solution. The credibility ratio in the evaluation result can be combined with the preset weight coefficient, and a set of constraint equations or linear programming models can be used to solve the new weight distribution. For example, if the speech modality behaves abnormally in this window, the corresponding evaluation result will reduce its credibility ratio, thereby reducing the weight of the speech modality during the adjustment process; if the physiological modality shows high confidence after abnormal detection, its weight will be moderately increased to highlight the contribution of the modality in this fusion. After completing this solution process, the system will generate a modified weight coefficient vector and perform weighted summation or weighted concatenation on the multimodal emotion vectors at the same time in the fusion operation stage. If the weighted summation method is adopted, the calculation formula can be expressed as , in represents the modified weight coefficient, Represents the emotion vector of each modality. Through this weighted fusion, the proportion of different modalities in the final feature vector is precisely controlled, which can not only weaken the low-credibility modality, but also highlight the modality with excellent performance, and finally generate a fused feature vector. After obtaining the fused feature vector, the probability distribution of each emotion category is calculated according to the fused vector to obtain a unified emotion recognition result. When implementing this process, it is necessary to pre-train or set a multi-classification model in the system, such as a neural network based on Softmax output or a Bayesian classifier based on the Gaussian distribution assumption, and use the fused feature vector as the input of the model. If a neural network solution is adopted, the weight parameters are learned through a large number of fused vector samples with emotion labels during the training phase, so that the network can output the predicted probability of each emotion category (such as joy, sadness, fear, etc.) when the fused vector is input. In specific implementation, the fused vector can be processed using a fully connected layer or a convolutional layer. The output layer uses the Softmax function to map the network output to the probability value of each emotion category. The formula example is , in Represents the number of sentiment categories, and is the model parameter corresponding to this category. After calculating the probability distribution of each emotion category, the system will select the most representative emotion label according to the maximum probability principle or set the threshold strategy, and use it as a unified emotion recognition result. In this way, a more reliable and realistic emotion recognition output can be produced under the premise of considering multimodal credibility and dynamic weight adjustment. If the model uses other classification algorithms, such as random forest or support vector machine, the fused feature vector can also be input into the corresponding algorithm, and the discrimination scores of each emotion category can be obtained through tree structure voting or hyperplane partitioning, and then the probability distribution can be obtained through mapping or normalization, and then the same emotion category determination process can be completed.
[0038] Specifically, after generating a unified emotion recognition result, it is necessary to store the result in association with the corresponding multimodal original data with a timestamp to form emotion analysis data with traceability marks. The implementation of this step requires that a unified time mark has been embedded in each modal data stream during the data acquisition and preprocessing stage, for example, through NTP (Network Time Protocol) or hardware clock synchronization, a timestamp consistent with the system master clock is injected into each voice, video frame, and physiological data. After the multimodal emotion vector is fused and the final recognition result is output, the system will create a record in the database containing the emotion recognition result and the original data reference, and set the same timestamp field as the recognition result for the record to keep it accurately corresponding to the source data. If a distributed database is used, it is necessary to add an index field to the data table or document for fast retrieval and backtracking according to the time series structure. In specific implementation, a "emotion analysis result" table can be set up, containing fields such as "timestamp", "emotion category", "emotion probability distribution", etc., and linked to the corresponding original data table through foreign keys or document associations. The original data table should contain the same timestamp information and modal source identifier. In the data writing process, the newly generated emotion recognition results are associated with the multimodal data entries corresponding to the moment through the pre-set writing interface, and a traceability mark is added to identify which batch or batches of original data records the emotion analysis results correspond to in the time series dimension. In this way, the emotion recognition results at a specific moment can be reviewed or verified in the subsequent operation or query phase. For example, by retrieving the emotion recognition results within a certain timestamp range, the corresponding speech waveform, video frame data or physiological sensor reading can be quickly jumped to help analysts confirm whether there are any abnormalities in the process and basis of emotion judgment, thereby improving the traceability and consistency of multimodal emotion analysis.
[0039] Furthermore, the dynamically adjusting the preset weight coefficient according to the evaluation result to obtain the modified weight coefficient includes: calculating the relative credibility ratio between each mode according to the evaluation result to obtain a credibility ratio matrix, and using the credibility ratio matrix to construct a weight adjustment matrix to obtain weight constraints; performing linear programming solution on the preset weight coefficient according to the weight constraints to obtain an initial adjustment coefficient, and applying an exponential smoothing algorithm to the initial adjustment coefficient to obtain a stabilized adjustment parameter with time series continuity; performing matrix multiplication operation on the stabilized adjustment parameter and the preset weight coefficient to obtain an adjusted weight value, and performing a softmax function operation on the adjusted weight value to obtain a modified weight coefficient whose sum is 1.
[0040] Specifically, the relative credibility ratios between the modalities are calculated according to the evaluation results to obtain a credibility ratio matrix, and the credibility ratio matrix is used to construct a weight adjustment matrix to obtain weight constraints. In order to complete this step, it is necessary to maintain a multimodal credibility data structure in the system, which stores the evaluation results of each modality in different situations and its corresponding credibility score. During processing, the algorithm will read the evaluation values of each modality in turn, such as speech modality, facial expression modality, and physiological signal modality, and compare them two by two to calculate the relative credibility ratio. If the evaluation results of two modalities are and , then its relative credibility ratio can be defined as , indicating that in the current context, the modal With modal In order to obtain the credibility ratio matrix, it is necessary to perform the same operation on all modal combinations and record the results in matrix form. For example, if the system contains modes, we can construct a A matrix in which the diagonal elements can be filled with 1 or left empty to indicate that the ratio between the same mode is meaningless. Next, a weight adjustment matrix is constructed based on the matrix to characterize the influence of each mode on the overall weight distribution. The construction process can adopt the idea of graph theory or numerical optimization, and regard the credibility ratio matrix as an adjacency matrix of a weighted graph. In this graph, each mode corresponds to a node, and the weight of each edge is determined by the relative credibility. Then, the minimum spanning tree or shortest path algorithm of the graph is used to generate a preliminary weight adjustment relationship. It is also possible to perform several power operations or normalization operations on the credibility ratio matrix to enhance or weaken the prominent differences between certain modes. It is ensured that there is no extreme imbalance in the subsequent solution. Finally, a weight constraint formula is designed based on a series of threshold judgments or constraints. The formula may contain several inequalities, such as requiring the upper or lower weight limits of certain modes to remain within a reasonable range, or setting a limit of no more than 1 between modes. These inequalities together constitute the weight constraint conditions, providing the range of feasible solution sets for the next linear programming solution process. In this way, the credibility ratio matrix and the weight adjustment matrix are closely linked, so that the subsequent steps can comprehensively consider the evaluation results of different modes and the weight checks and balances between them when allocating the weights of each mode.
[0041] After obtaining the weight constraint, it is necessary to perform linear programming on the preset weight coefficient according to the weight constraint to obtain the initial adjustment coefficient, and apply the exponential smoothing algorithm to the initial adjustment coefficient to obtain a stabilized adjustment parameter with time series continuity. The implementation of linear programming solution can rely on the existing numerical optimization library, for example, using the simplex method or the interior point method to solve the weight allocation problem. The specific process is: first, the preset weight coefficient is used as the initial solution, combined with the weight constraint constructed in the previous stage, to form a typical linear programming model, in which the decision variable is the adjustment amount of each modal weight, and the objective function can be to maximize the overall credibility or minimize the weighted error, or other business-related goals can be set. In the linear programming model, upper and lower bounds are imposed on each decision variable to ensure that the weight will not be adjusted to a negative value or exceed the maximum upper limit, and to ensure that the relative constraints between the modes are met. Through iterative calculation, the numerical solver will find the optimal solution in the solution space that meets the constraints and obtain a set of initial adjustment coefficients. If the credibility evaluation value of some modes is too low, the solver will automatically reduce the proportion of these modes in the overall weight, so that the system pays more attention to the modes with higher credibility during the fusion stage. After obtaining the initial adjustment coefficient, the exponential smoothing algorithm is applied to it to obtain the stable adjustment parameter of the time series continuity. The implementation method of the exponential smoothing algorithm is usually to recursively perform a weighted summation of the smoothing result of the previous moment and the initial adjustment coefficient of the current moment, and let the smoothing parameter Controls the weight of old and new values. For example, ,in Indicates the initial adjustment coefficient at the current moment, represents the smoothing coefficient of the previous moment, is a smoothing factor with a value between 0 and 1. In this way, the stabilization adjustment parameter gradually eliminates random fluctuations while retaining the ability to respond to the latest evaluation results, thereby taking into account both stability and real-time performance in subsequent weight allocation. If necessary, a quadratic smoothing or weighted sliding average mechanism can be superimposed on the exponential smoothing to further reduce the impact of instantaneous mutations on the weights, ensuring that the system can maintain a relatively stable weight allocation policy in different time periods. After obtaining the stabilization adjustment parameter, it is necessary to perform a matrix multiplication operation on the stabilization adjustment parameter and the preset weight coefficient to obtain an adjusted weight value, and perform a softtmax function operation on the adjusted weight value to obtain a modified weight coefficient whose sum is 1. To achieve this process, the preset weight coefficient is first organized into a vector or diagonal matrix form, and then combined with the stabilization adjustment parameter by matrix multiplication. If the preset weight coefficient is a vector , and the stabilization adjustment parameter is , then we can make , where © represents element-by-element multiplication, so that the final weight of each mode will increase or decrease according to its smoothed adjustment coefficient. In some systems, the coupling relationship between modes may be represented in the form of a matrix. Diagonalize it, and let , and then with a diagonal or sparse matrix of the form Multiply to reflect more complex modal interactions. After completing the matrix multiplication, the adjusted weight values need to be processed by the softmax function. The softmax function is defined as soft , which maps a set of real numbers to the interval (0, 1), and the sum of all outputs is 1. By performing softmax processing on the adjusted weight values, the final corrected weight coefficient can be guaranteed satisfy , so that the weight allocation in multimodal sentiment analysis has clear numerical boundaries and comparability. If a certain modality is assigned a significantly increased coefficient in the adjustment stage due to excellent speech estimation results, then after the softmax operation, the weight of the modality will occupy a higher proportion accordingly, thus exerting a greater influence in the subsequent feature fusion. Through this series of operations, dynamic weighting of each modality can be achieved numerically, taking into account the feedback of the evaluation results and the requirements of time series smoothing, and finally outputting a set of modified weight coefficients that can participate in the fusion operation with a total of 1 month.
[0042] In this embodiment, by setting a forward sampling channel and a reverse sampling channel in the source sampling resistor, setting differentiated sampling time windows for different sampling channels, and staggering the time windows of the dual sampling channels, the time period with the strongest common-mode interference can be avoided; at the same time, by mapping the switching speed of the silicon carbide MOS device to a theoretical common-mode interference waveform, and using the waveform to synthesize and process with the bidirectional sampling signal, the common-mode interference component in the sampling signal can be accurately separated, making the synthesized current signal more accurate; on this basis, the current signal is compensated in combination with the on-resistance change caused by the junction temperature change, and a hierarchical protection strategy is used for overcurrent protection. This method effectively overcomes the influence of common-mode interference on current sampling through the staggered design of sampling channel time and the separation of common-mode interference based on switching speed.
[0043] The above describes the multimodal sentiment analysis method based on multi-agent collaboration in the embodiment of the present invention. The following describes the multimodal sentiment analysis device based on multi-agent collaboration in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a multimodal sentiment analysis device based on multi-agent collaboration includes: The data acquisition module 201 is used to perform clock synchronization acquisition and preprocessing on the voice signal, the facial expression video frame and the physiological parameters to obtain a multimodal data packet with a timing mark; A feature mapping module 202 is used to extract features from the multimodal data packet using a voice agent, a facial agent, and a physiological agent, and to map the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; The difference detection module 203 is used to detect the degree of difference between vectors in the multimodal sentiment vector, and when a significant difference is detected, start the conflict detection algorithm based on feature credibility to evaluate and obtain an evaluation result; The result fusion module 204 is used to perform a feature fusion operation on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and store the emotion recognition result and the corresponding multimodal original data in association with a timestamp.
[0044] In an embodiment of the present invention, the multimodal emotion analysis device based on multi-agent collaboration runs the multimodal emotion analysis method based on multi-agent collaboration, and the multimodal emotion analysis device based on multi-agent collaboration obtains a multimodal data packet with a time sequence mark by performing clock synchronization acquisition and preprocessing of multimodal data. Multiple agents are used to extract features from the multimodal data packet respectively, and the extracted feature vectors are mapped with the preset emotion ontology to obtain a multimodal emotion vector. The degree of difference between the vectors in the multimodal emotion vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started to evaluate and obtain an evaluation result. According to the evaluation result and the preset weight coefficient, the multimodal emotion vector is subjected to feature fusion operation to obtain an emotion recognition result, and the emotion recognition result is stored in association with the corresponding multimodal data by timestamp. The present invention effectively prevents abnormal feature vectors from affecting the fusion result by introducing a difference detection and evaluation mechanism, thereby improving the reliability of modal emotion analysis.
[0045] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0046] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0047] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal sentiment analysis method based on multi-agent collaboration, characterized in that: The multimodal sentiment analysis method based on multi-agent collaboration includes: Perform clock synchronization acquisition and preprocessing on speech signals, facial expression video frames and physiological parameters to obtain multimodal data packets with time sequence marks; Using a voice agent, a facial agent, and a physiological agent to extract features from the multimodal data packet, respectively, and mapping the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; The difference between the vectors in the multimodal sentiment vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started to evaluate and obtain an evaluation result; A feature fusion operation is performed on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and the emotion recognition result is stored in association with the corresponding multimodal original data by timestamp.
2. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 1 is characterized in that: The method of using the voice agent, the facial agent and the physiological agent to extract features from the multimodal data packet respectively, and mapping the extracted feature vectors with the preset emotion ontology to obtain the multimodal emotion vector in the unified emotion semantic space includes: Perform Fourier transform and Mel spectrum analysis on speech signals, and use convolutional neural network to extract phoneme and prosodic features to obtain speech feature vectors; Perform facial key point detection and spatial feature extraction on facial expression video frames to obtain facial expression feature vectors; Perform wavelet transform and frequency domain analysis on physiological parameters to obtain physiological feature vectors; According to the emotion label system in the preset emotion ontology knowledge base, the speech feature vector, the facial expression feature vector and the physiological feature vector are semantically mapped to obtain a multimodal emotion vector in a unified emotion semantic space.
3. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 2 is characterized in that: The method of semantically mapping the speech feature vector, the facial expression feature vector and the physiological feature vector according to the emotion label system in the preset emotion ontology knowledge base to obtain a multimodal emotion vector in a unified emotion semantic space includes: Normalizing the speech feature vector, the facial expression feature vector and the physiological feature vector to obtain a standardized feature vector; Performing a projection transformation on the standardized feature vector according to the basic emotion tags in a preset emotion ontology knowledge base to obtain an initial semantic mapping vector; Performing a nonlinear activation function operation on the initial semantic mapping vector to obtain an activated semantic vector, and performing a combination operation on the activated semantic vector according to a composite emotion rule in a preset emotion ontology knowledge base to obtain a multi-dimensional emotion feature; The multi-dimensional emotional features are subjected to dimensionality reduction processing to obtain a multi-modal emotional vector in a unified emotional semantic space.
4. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 1 is characterized in that: The difference between the vectors in the multimodal sentiment vector is detected, and when a significant difference is detected, a conflict detection algorithm based on feature credibility is started to perform evaluation, and the evaluation results obtained include: Constructing a time series sliding window for the multimodal emotion vector, calculating the autocorrelation coefficient and the mutual correlation coefficient of the emotion vector of each mode in the time series sliding window, and obtaining a correlation matrix; Using a preset spectral clustering algorithm to calculate the correlation between the sentiment vectors of each modality according to the correlation matrix, and obtaining a difference index; According to the difference index, a pre-trained anomaly detection network is called to perform comparative analysis on the multimodal sentiment vector to obtain an abnormal vector label, and a wavelet transform is performed on the sentiment vector labeled as abnormal to obtain a local fluctuation coefficient and a global trend coefficient; The credibility score of the sentiment vector marked as abnormal is calculated according to the local fluctuation coefficient and the global trend coefficient using the preset hierarchical feature evaluation rule, and the credibility score is weighted and corrected according to the historical statistical indicators of the multimodal sentiment vector to obtain an evaluation result.
5. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 4 is characterized in that: The preset hierarchical feature evaluation rule is used to calculate the credibility score of the emotion vector marked as abnormal according to the local fluctuation coefficient and the global trend coefficient, and the credibility score is weighted and corrected according to the historical statistical indicators of the multimodal emotion vector, and the evaluation results include: Performing frequency band decomposition on the local fluctuation coefficient to obtain a multi-scale fluctuation feature, and calculating a signal stability index according to the multi-scale fluctuation feature and a global trend coefficient to obtain a basic credibility value; Performing sliding window statistics on the historical data of the multimodal sentiment vector to obtain a historical fluctuation range; The basic credibility value is interval mapped according to the historical fluctuation range to obtain a correction coefficient, and the basic credibility value is corrected and calculated using the correction coefficient to obtain an evaluation result.
6. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 1 is characterized in that: The step of performing a feature fusion operation on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and storing the emotion recognition result in association with the corresponding multimodal original data by timestamp includes: Dynamically adjusting the preset weight coefficient according to the evaluation result to obtain a modified weight coefficient, and using the modified weight coefficient to perform a weighted fusion operation on the multimodal emotion vector to obtain a fused feature vector; Calculating the probability distribution of each emotion category according to the fused feature vector to obtain a unified emotion recognition result; A unified timestamp is added to the emotion recognition result and the corresponding multimodal original data, and the data is stored in a distributed database according to a time series structure to obtain emotion analysis data with traceability marks.
7. The multimodal sentiment analysis method based on multi-agent collaboration according to claim 6 is characterized in that: The dynamically adjusting the preset weight coefficient according to the evaluation result to obtain the modified weight coefficient includes: Calculating the relative credibility ratios between the modes according to the evaluation results to obtain a credibility ratio matrix, and constructing a weight adjustment matrix using the credibility ratio matrix to obtain weight constraint conditions; Performing linear programming to solve the preset weight coefficient according to the weight constraint condition to obtain an initial adjustment coefficient, and applying an exponential smoothing algorithm to the initial adjustment coefficient to obtain a stabilized adjustment parameter with time series continuity; A matrix multiplication operation is performed on the stabilization adjustment parameter and the preset weight coefficient to obtain an adjusted weight value, and a softmax function operation is performed on the adjusted weight value to obtain a modified weight coefficient whose sum is 1.
8. A multi-modal sentiment analysis device based on multi-agent collaboration, characterized in that: The multimodal sentiment analysis device based on multi-agent collaboration includes: The data acquisition module is used to perform clock synchronization acquisition and preprocessing of speech signals, facial expression video frames and physiological parameters to obtain multimodal data packets with time sequence marks; A feature mapping module is used to extract features from the multimodal data packet using a voice agent, a facial agent, and a physiological agent, and to map the extracted feature vectors with a preset emotion ontology to obtain a multimodal emotion vector in a unified emotion semantic space; The difference detection module is used to detect the degree of difference between vectors in the multimodal sentiment vector, and when a significant difference is detected, the conflict detection algorithm based on feature credibility is started to evaluate and obtain the evaluation result; The result fusion module is used to perform a feature fusion operation on the multimodal emotion vector according to the evaluation result and a preset weight coefficient to obtain a unified emotion recognition result, and store the emotion recognition result in association with the corresponding multimodal original data by timestamp.
Citation Information
Cited By
Multi-modal sentiment classification method based on dynamic game strategy
CN120429602A
Street cleanliness real-time evaluation method based on multi-modal data fusion
CN120449111A
Wild animal and plant species identification method
CN120470544A
Cf-PWV synchronous measurement method and system based on multi-modal data fusion
CN120561873A
Artificial intelligence psychological assessment method and device based on multiple modes
CN120600318A