Multi-modal missing data emotion recognition method and system based on dynamic completion

By detecting modality missing data using a dynamic completion method, generating a reliability weight vector, dynamically aligning multimodal data, constructing an evolutionary trajectory graph, and adaptively completing features, the problem of modality missing data in multimodal emotion recognition is solved, improving recognition accuracy and stability.

CN120930079AActive Publication Date: 2025-11-11BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Patent Information

Application Number
CN202511454084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-11-11
Estimated Expiration
2045-10-13

Smart Images

  • Figure CN120930079A_ABST
    Figure CN120930079A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal missing data emotion recognition method and system based on dynamic completion, and relates to the field of emotion recognition, and the method comprises the steps: calculating a modal reliability weight vector, determining a modal combination, carrying out dynamic alignment, constructing an evolution trajectory diagram, predicting missing node features, and carrying out emotion recognition. And dividing a time sequence completion interval based on information entropy distribution to carry out adaptive feature completion, and finally realizing emotion recognition. According to the method, the problem of multi-modal data missing can be effectively solved, the cross-modal feature capture capability and the emotion recognition accuracy are improved, and the system robustness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to emotion recognition technology, and more particularly to an emotion recognition method and system based on dynamic completion for multimodal missing data. Background Technology

[0002] Emotion recognition is an important research area in artificial intelligence, with broad application value in fields such as human-computer interaction, intelligent services, and mental health monitoring. Multimodal emotion recognition, by integrating information from multiple modalities such as speech, text, and vision, can more comprehensively understand and identify human emotional states. However, with the widespread adoption of smart devices and the development of multimodal data acquisition technologies, the challenges faced by multimodal emotion recognition in practical applications are becoming increasingly prominent, particularly the modality gap problem that frequently occurs in real-world environments.

[0003] Existing multimodal emotion recognition methods primarily employ fixed modality combinations or simple completion strategies, making it difficult to effectively handle modality loss issues in complex scenarios. These methods typically assume all modality data is complete and available, or use fixed modality combination strategies, lacking the ability to dynamically detect and adaptively handle modality loss. Furthermore, existing technologies often ignore the temporal dependencies between modalities, failing to accurately capture the dynamic evolution of emotion expression, leading to a significant drop in recognition accuracy when modalities are incomplete. The lack of a mechanism for evaluating modality quality and reliability prevents dynamic adjustment of the contribution weights of each modality in the fusion process based on modality completeness and quality, resulting in a sharp decline in model performance when modalities are missing or of poor quality. Ignoring the temporal correlations and evolutionary patterns between multimodal data makes it difficult to effectively utilize complementary information between modalities for missing data completion, hindering the capture of continuous changes in emotion expression. The lack of adaptive completion strategies for different temporal segments, using a uniform completion method to process all missing data, fails to consider the differences in information entropy of emotion expression across different temporal segments, resulting in poor completion effects and impacting the final emotion recognition accuracy. Summary of the Invention

[0004] This invention provides a method and system for emotion recognition based on dynamic completion of multimodal missing data, which can solve the problems in the prior art.

[0005] A first aspect of this invention provides a method for sentiment recognition of multimodal missing data based on dynamic completion, comprising:

[0006] Acquire multimodal input data, detect missing data locations for each modality, calculate modal information completeness value and quality score, and generate modal reliability weight vector;

[0007] Calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data.

[0008] The temporal features are constructed as nodes in the evolutionary trajectory graph. The temporal dependency strength between nodes is calculated. The nodes are divided into feature groups according to the temporal dependency strength. The predicted features of missing nodes are generated by combining the temporal evolution law of modality combination.

[0009] The information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Within the temporal completion intervals, the predicted features are adaptively completed to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

[0010] Acquire multimodal input data, detect missing data locations for each modality, calculate modal information completeness values ​​and quality scores, and generate modal reliability weight vectors including:

[0011] Acquire multimodal input data including video, audio, and text; calculate the continuity difference value of adjacent data segments in the multimodal input data; mark the difference positions above the difference threshold as data missing positions; and generate a data missing label sequence.

[0012] Calculate the temporal variation features of the non-missing data indicated by the missing data marker sequence, perform feature fusion on the temporal variation features to obtain the temporal correlation degree, and calculate the information completeness value of each modality based on the temporal correlation degree.

[0013] Extract the content features and structural features of the non-missing data, combine them according to preset weights to obtain quality features, and calculate the quality score of each modality based on the quality features;

[0014] The feature importance coefficient is calculated from the information completeness value and quality score. The feature importance coefficient is then normalized to obtain the modality reliability weight vector, which is used to characterize the importance of each modality feature in feature fusion.

[0015] The cross-correlation coefficients between modes in the multimodal input data are calculated. Mode combinations are determined based on the cross-correlation coefficients and reliability weight vectors. Temporal offsets of the mode combinations are detected. The multimodal input data is dynamically aligned based on these temporal offsets. Temporal features are extracted from the dynamically aligned data, including:

[0016] Calculate the temporal entropy value of each modal feature in the multimodal input data, construct an adaptive sampling probability distribution based on the temporal entropy value, resample the multimodal input data according to the adaptive sampling probability distribution to obtain a sampling sequence, and calculate the cross-correlation coefficient between modes in the sampling sequence;

[0017] A dynamic constraint matrix is ​​constructed by combining the cross-correlation coefficient and the time entropy value. The dynamic constraint matrix is ​​multiplied by the reliability weight vector to obtain the modal confidence. A combination threshold is set according to the modal confidence. Modal combinations are constructed by pairing modes that are higher than the combination threshold.

[0018] The temporal offset features are obtained by eigenvalue decomposition of the dynamic constraint matrix, the temporal offset is reconstructed from the temporal offset features, and the sampled sequence is dynamically aligned according to the temporal offset to obtain the aligned sequence.

[0019] The variation features of the aligned sequence are extracted and weighted and fused with the suppression vector to obtain the temporal features.

[0020] Temporal offset features are obtained through eigenvalue decomposition of the dynamic constraint matrix. These features are then reconstructed to obtain the temporal offset. The sampled sequences are then dynamically aligned based on the temporal offset to obtain the aligned sequences, which include:

[0021] The dynamic constraint matrix is ​​decomposed into eigenvalues, the decay curvature of the eigenvalues ​​is calculated to determine the number of principal eigenvalues, the eigenvectors corresponding to the principal eigenvalues ​​are extracted, and they are mapped to the nonlinear time space to obtain the time-series offset features.

[0022] Based on the aforementioned temporal offset features, a multi-scale temporal sliding window is constructed. Local temporal change values ​​are extracted within each sliding window. Probability density estimation is performed on the local temporal change values ​​to obtain the global temporal trend value. The local temporal change values ​​and the global temporal trend value are combined and reconstructed to obtain the temporal offset.

[0023] The relative displacement relationship between modes is calculated based on the time offset, and the sampling sequence is time-corrected to generate an initial alignment sequence.

[0024] Calculate the temporal consistency value and feature correlation value of the initial alignment sequence, adjust the temporal offset according to the consistency value and correlation value, optimize the initial alignment sequence, and obtain the alignment sequence.

[0025] The temporal features are constructed as nodes in an evolutionary trajectory graph. The temporal dependency strength between nodes is calculated, and the nodes are divided into feature groups based on the temporal dependency strength. Predictive features for missing nodes are generated by combining the temporal evolution law of modality combinations, including:

[0026] The temporal features are constructed into nodes of the evolutionary trajectory graph. The length of the time window is dynamically adjusted according to the changing trend of the temporal features. Historical information is extracted within the time window to construct node features.

[0027] The temporal dependency strength between nodes is calculated using multidimensional information entropy and state transition probability. The temporal dependency strength is then corrected by combining the temporal consistency of node features to obtain the final temporal dependency strength.

[0028] Clustering nodes based on temporal dependency strength, determining cluster boundaries based on the similarity of node features, dividing nodes into feature groups, and obtaining the temporal evolution patterns of nodes within the feature groups;

[0029] By analyzing the node features of different feature groups, the temporal evolution law of mode combination is obtained. A co-evolution matrix of feature groups is constructed. The co-evolution matrix is ​​combined with the temporal evolution law to generate the predicted features of missing nodes.

[0030] The predicted features are compared with the actual features to calculate the prediction bias. The co-evolution matrix is ​​then optimized based on the prediction bias, and the predicted features are updated.

[0031] The information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Adaptive completion of the predicted features within the temporal completion intervals yields multimodal fusion features. Based on these multimodal fusion features, the sentiment types of samples are identified, including:

[0032] Calculate the conditional information entropy and joint information entropy of temporal features in the time dimension, combine them to generate an information entropy distribution, identify feature mutation points based on the gradient change of the information entropy distribution, and determine the region between adjacent feature mutation points as the temporal completion interval;

[0033] Extract the frequency distribution and cumulative distribution of time series features within the time series completion interval, construct a probability density function, and generate time series completion constraints based on the change patterns of the probability density function and historical data.

[0034] Based on the temporal completion constraint, the feature importance is calculated, and the predicted features are progressively weighted according to the feature importance. The progressive weighting result is used to adaptively complete the predicted features to generate multimodal fusion features.

[0035] The correlation coefficients and mutual information values ​​of the multimodal fusion features in the time dimension and feature dimension are calculated and combined to obtain the temporal correlation degree. The feature weights are iteratively optimized based on the temporal correlation degree to generate optimized multimodal fusion features. The optimized multimodal fusion features are then mapped to the sentiment feature space for sample sentiment recognition.

[0036] A second aspect of the present invention provides a multimodal missing data sentiment recognition system based on dynamic completion, comprising:

[0037] The first unit is used to acquire multimodal input data, detect the missing data locations of each modality, calculate the modality information completeness value and quality score, and generate a modality reliability weight vector.

[0038] The second unit is used to calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data.

[0039] The third unit is used to construct nodes of the evolutionary trajectory map from temporal features, calculate the temporal dependency strength between nodes, divide nodes into feature groups according to the temporal dependency strength, and generate predicted features for missing nodes by combining the temporal evolution law of modality combination.

[0040] The fourth unit is used to calculate the information entropy distribution of temporal features, divide the temporal completion interval according to the information entropy distribution, and perform adaptive completion on the predicted features within the temporal completion interval to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

[0041] A third aspect of the present invention provides an electronic device, comprising:

[0042] processor;

[0043] Memory used to store processor-executable instructions;

[0044] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0045] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0046] In this embodiment, a reliability weight vector is generated by calculating the modality information completeness value and quality score. The optimal modality combination is then determined by combining the cross-correlation coefficients between modalities, achieving efficient screening of multimodal data and effectively improving the accuracy of emotion recognition in scenarios with missing data. Dynamic alignment using temporal offsets constructs an evolutionary trajectory graph of temporal features. Feature groups are divided according to the strength of temporal dependencies, fully capturing the temporal evolution patterns between different modalities and resolving the mismatch problem of multimodal data in the temporal dimension. Temporal completion intervals are divided based on information entropy distribution, and predictive features are adaptively completed, achieving accurate recovery and fusion of missing modality data. This improves the robustness of the model in complex and variable environments, making the emotion recognition results more reliable. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the multimodal missing data sentiment recognition method based on dynamic completion according to an embodiment of the present invention;

[0048] Figure 2 This is a flowchart of a multimodal fusion and emotion recognition technology based on temporal feature analysis, as described in this embodiment of the invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0051] Figure 1 This is a flowchart illustrating the multimodal missing data sentiment recognition method based on dynamic completion according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0052] Acquire multimodal input data, detect missing data locations for each modality, calculate modal information completeness value and quality score, and generate modal reliability weight vector;

[0053] Calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data.

[0054] The temporal features are constructed as nodes in the evolutionary trajectory graph. The temporal dependency strength between nodes is calculated. The nodes are divided into feature groups according to the temporal dependency strength. The predicted features of missing nodes are generated by combining the temporal evolution law of modality combination.

[0055] The information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Within the temporal completion intervals, the predicted features are adaptively completed to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

[0056] In one optional implementation, multimodal input data is acquired, the missing data locations for each modality are detected, modal information completeness values ​​and quality scores are calculated, and a modal reliability weight vector is generated, including:

[0057] Acquire multimodal input data including video, audio, and text; calculate the continuity difference value of adjacent data segments in the multimodal input data; mark the difference positions above the difference threshold as data missing positions; and generate a data missing label sequence.

[0058] Calculate the temporal variation features of the non-missing data indicated by the missing data marker sequence, perform feature fusion on the temporal variation features to obtain the temporal correlation degree, and calculate the information completeness value of each modality based on the temporal correlation degree.

[0059] Extract the content features and structural features of the non-missing data, combine them according to preset weights to obtain quality features, and calculate the quality score of each modality based on the quality features;

[0060] The feature importance coefficient is calculated from the information completeness value and quality score. The feature importance coefficient is then normalized to obtain the modality reliability weight vector, which is used to characterize the importance of each modality feature in feature fusion.

[0061] In this embodiment, multimodal input data, including video, audio, and text, is acquired. Taking an emotional interaction scenario as an example, video data includes visual information such as user facial expressions and body movements; audio data includes acoustic features such as speech content and intonation variations; and text data is speech-to-text transcription or text content directly input by the user. For the collected data, it is first necessary to detect any missing data locations in each modality.

[0062] For multimodal input data, the continuity difference between adjacent data segments is calculated. For video data, continuity can be measured by calculating the pixel difference or the Euclidean distance of feature vectors between adjacent frames. For example, for facial expression videos, the coordinates of facial key points in each frame are extracted, and the change in key point positions between adjacent frames is calculated. When the change between two frames exceeds a preset threshold (e.g., the average displacement of key points is greater than 10 pixels), it can be determined as a continuity anomaly. For audio data, anomalies can be detected by calculating the acoustic feature differences between adjacent audio segments (e.g., the Euclidean distance of Mel-frequency cepstral coefficients). For text data, the semantic similarity between adjacent text segments can be calculated; a sudden drop in semantic coherence may indicate missing data.

[0063] Locations with differences exceeding a certain threshold are marked as missing data locations. For example, for video data, if the feature difference between two frames is 0.85, and the preset threshold is 0.7, then that location is marked as missing. Similar processing is performed on audio and text data, ultimately generating a sequence of missing data markers. This sequence can be represented as a binary vector, where "1" indicates that the corresponding location has normal data, and "0" indicates that the corresponding location has missing data.

[0064] For non-missing data indicated by missing data marker sequences, temporal variation features are calculated. For video data, temporal features such as the rate, magnitude, and direction of facial expression changes can be extracted; for audio data, temporal features such as pitch change trends, speech rate changes, and volume changes can be extracted; for text data, temporal features such as changes in the distribution of emotional vocabulary and sentence structure changes can be extracted.

[0065] The temporal correlation degree is obtained by fusing the above-mentioned temporal variation features. Feature fusion can be performed using a weighted average method, combining each temporal feature according to a preset weight. For example, for video data, the weights of the facial expression change rate, change amplitude, and change direction can be set to 0.4, 0.3, and 0.3 respectively for weighted fusion.

[0066] The information completeness value for each modality is calculated based on temporal correlation. The information completeness value reflects the continuity and integrity of the modal data. The calculation method is as follows: multiply the temporal correlation by the proportion of data without missing information, and then compare it with a baseline completeness value (e.g., 0.8) to obtain a relative completeness score. For example, if the temporal correlation of video data is 0.75 and the proportion of data without missing information is 0.9, then its information completeness value can be calculated as 0.75 × 0.9 ÷ 0.8 = 0.84. Next, the content features and structural features of the data without missing information are extracted. For video data, content features include facial expression type and intensity, and structural features include the geometric configuration of facial key points; for audio data, content features include the emotional tendency of the speech content, and structural features include the prosodic pattern of the speech; for text data, content features include the frequency of emotional words, and structural features include the complexity of syntactic structure.

[0067] Quality features are obtained by combining content features and structural features according to preset weights. For example, for video data, the weights of content features and structural features can be set to 0.6 and 0.4; for audio data, they can be set to 0.5 and 0.5; and for text data, they can be set to 0.7 and 0.3. The quality score for each modality is calculated based on the quality features, and a mapping function can be used to map the feature values ​​to a score range of 0-1.

[0068] The feature importance coefficient is calculated from the information completeness value and the quality score. The calculation method can be a weighted average of the information completeness value and the quality score, for example, assigning weights of 0.4 and 0.6 respectively. The feature importance coefficient is then normalized to obtain a modality reliability weight vector. Normalization ensures that the sum of the weights for all modalities is 1; for example, the weights for video, audio, and text are 0.4, 0.35, and 0.25, respectively. This weight vector is used to characterize the importance of each modality feature in subsequent feature fusion. Assume that in a multimodal data collection scenario for emotion recognition, the video data has a 2-second freeze at 30 seconds, the audio data has a 1-second silence at 45 seconds, and the text data is basically complete. Using the above method, the information completeness values ​​for video, audio, and text are calculated to be 0.92, 0.95, and 0.98, respectively; and the quality scores are 0.85, 0.80, and 0.90, respectively. These values ​​are combined to calculate the feature importance coefficient, and after normalization, the modality reliability weight vector is [0.32, 0.30, 0.38]. This indicates that in this emotion recognition task, the text modality has the highest reliability and should be given the largest weight, followed by the video modality, and finally the audio modality.

[0069] By dynamically detecting missing locations in multimodal data and evaluating the reliability of each modality, the system's ability to handle incomplete data is improved. Through temporal correlation calculation and quality scoring mechanisms, the validity of each modality's data is accurately assessed. An adaptive weight fusion strategy is adopted, enabling the system to adjust the importance of each modality based on real-time data quality, thus enhancing its robustness in complex environments. This effectively solves the performance degradation problem of traditional fixed-weight methods in scenarios with missing data, improving the accuracy and stability of multimodal emotion recognition.

[0070] In one optional implementation, the cross-correlation coefficients between modes in the multimodal input data are calculated; mode combinations are determined based on the cross-correlation coefficients and reliability weight vectors; the temporal offset of the mode combinations is detected; the multimodal input data is dynamically aligned according to the temporal offset; and temporal features are extracted from the dynamically aligned data, including:

[0071] Calculate the temporal entropy value of each modal feature in the multimodal input data, construct an adaptive sampling probability distribution based on the temporal entropy value, resample the multimodal input data according to the adaptive sampling probability distribution to obtain a sampling sequence, and calculate the cross-correlation coefficient between modes in the sampling sequence;

[0072] A dynamic constraint matrix is ​​constructed by combining the cross-correlation coefficient and the time entropy value. The dynamic constraint matrix is ​​multiplied by the reliability weight vector to obtain the modal confidence. A combination threshold is set according to the modal confidence. Modal combinations are constructed by pairing modes that are higher than the combination threshold.

[0073] The temporal offset features are obtained by eigenvalue decomposition of the dynamic constraint matrix, the temporal offset is reconstructed from the temporal offset features, and the sampled sequence is dynamically aligned according to the temporal offset to obtain the aligned sequence.

[0074] The variation features of the aligned sequence are extracted and weighted and fused with the suppression vector to obtain the temporal features.

[0075] In the specific implementation, the multimodal input data includes video data, audio data, and text data. First, each modal data is preprocessed. Video data is processed using frame extraction techniques to obtain keyframe sequences; audio data is converted into spectral feature sequences through spectral analysis; and text data is converted into text feature sequences through word segmentation and vectorization. The preprocessed modal data are represented as feature vector sequences and stored in system memory for subsequent processing.

[0076] When calculating the temporal entropy of each modality feature in the multimodal input data, a sliding window technique is used to segment the feature sequence. For each modality's feature sequence, the window size is set to 128 time points, with a window overlap rate of 50%. Within each window, the probability density function of the feature distribution is calculated, and estimated by statistically analyzing the frequency of feature values ​​falling into predefined intervals. Taking video features as an example, the feature value range is divided into 32 equally wide intervals, and the frequency of feature values ​​in each interval is calculated and normalized to obtain a probability estimate. The temporal entropy value is calculated using the Shannon entropy formula; a high entropy value indicates that the modality has greater temporal variation and higher information content.

[0077] An adaptive sampling probability distribution is constructed based on the calculated temporal entropy values. Specifically, the temporal entropy values ​​of each modality are normalized to obtain weighting coefficients. For example, if the temporal entropy values ​​for video, audio, and text modalities are 0.85, 0.65, and 0.45, respectively, the normalized weighting coefficients are 0.44, 0.33, and 0.23. These weighting coefficients are used to construct a Gaussian mixture distribution, serving as the probability distribution function for temporal point sampling. Based on this probability distribution function, the original sequence is resampled, with more sampling points obtained in high-entropy regions. In practical applications, a 10-minute multimodal dataset with a sampling rate of 30Hz may be resampled into a sequence of 2000 key temporal points.

[0078] For the resampled sequences, cross-correlation coefficients between modes are calculated. A sliding window technique is used, with a window size of 64 time points. For video and audio mode pairs, the Pearson correlation coefficient is calculated within each window. For example, a cross-correlation coefficient of 0.78 between video motion features and audio volume features within a certain window indicates a high correlation between these two modes within that window period. Cross-correlation coefficients are calculated for all time windows of all mode pairs, forming a cross-correlation coefficient matrix.

[0079] A dynamic constraint matrix is ​​constructed by combining cross-correlation coefficients and temporal entropy values. This matrix is ​​constructed by weighting each element of the cross-correlation coefficient matrix with the temporal entropy value of the corresponding modality pair. The weight parameters are determined to be 0.7 and 0.3 through validation set optimization. For example, if the cross-correlation coefficient for a video-audio modality pair is 0.78 and the corresponding weighted average temporal entropy value is 0.76, then the value of this element in the dynamic constraint matrix is ​​0.774.

[0080] The modal confidence is obtained by multiplying the dynamic constraint matrix by the reliability weight vector. The reliability weight vector is pre-calculated based on the signal-to-noise ratio and integrity metrics of each modality. For example, in a specific application scenario, the reliability weights for video, audio, and text modalities are 0.8, 0.75, and 0.6, respectively. The confidence of each modal pair is calculated through matrix multiplication; for example, the confidence of the video-audio modal pair is 0.619.

[0081] A combination threshold is set based on modality confidence levels. The threshold is set to the mean of all modality confidence levels plus 0.5 standard deviations; in this example, it is 0.58. Modalities with confidence levels above the threshold are paired to construct modality combinations. In this example, the video-audio modality combination is retained because its confidence level of 0.619 is higher than the threshold of 0.58, while the text-audio modality combination is filtered out because its confidence level of 0.51 is lower than the threshold.

[0082] Temporal offset features are obtained based on eigenvalue decomposition of the dynamic constraint matrix. Eigenvalue decomposition is performed on the submatrices of the retained modality combinations' dynamic constraint matrix, extracting the first three eigenvectors as temporal offset features. These eigenvectors are then converted into actual time offsets in milliseconds through a nonlinear mapping. For example, the temporal offset for a video-audio modality combination is determined to be 120 milliseconds, indicating that the video data leads the audio data by 120 milliseconds.

[0083] The sampled sequence is dynamically aligned based on the calculated temporal offset. For video-audio modal combinations, the audio data timeline is shifted by 120 milliseconds to align with the video data time points. The alignment operation is achieved through interpolation calculations to maintain data continuity. The aligned data is stored as an aligned sequence, retaining the index information of the original sample points for easy retrieval of the original data.

[0084] The algorithm extracts variation features from the aligned sequence, focusing primarily on gradient information and local change patterns across different modalities. Inter-frame differences are calculated for video features, spectral change rates for audio features, and semantic transformation strengths for text features. These variation features are then weighted and fused with a pre-trained suppression vector, which reduces the impact of noise and redundant information. For example, the suppression vector assigns a weight of 0.1 to video background noise and a weight of 0.9 to subject motion. The resulting temporal features exhibit higher discriminative power and stability, effectively characterizing the temporal dynamics of multimodal data.

[0085] Through the detailed steps described above, dynamic alignment of multimodal data and effective extraction of temporal features are achieved, enabling the system to better understand and analyze complex data containing multiple information modalities.

[0086] In one optional implementation, time-series offset features are obtained based on eigenvalue decomposition of the dynamic constraint matrix, and the time-series offset is reconstructed to obtain a time-series offset. The sampled sequence is then dynamically aligned based on the time-series offset to obtain an aligned sequence, including:

[0087] The dynamic constraint matrix is ​​decomposed into eigenvalues, the decay curvature of the eigenvalues ​​is calculated to determine the number of principal eigenvalues, the eigenvectors corresponding to the principal eigenvalues ​​are extracted, and they are mapped to the nonlinear time space to obtain the time-series offset features.

[0088] Based on the aforementioned temporal offset features, a multi-scale temporal sliding window is constructed. Local temporal change values ​​are extracted within each sliding window. Probability density estimation is performed on the local temporal change values ​​to obtain the global temporal trend value. The local temporal change values ​​and the global temporal trend value are combined and reconstructed to obtain the temporal offset.

[0089] The relative displacement relationship between modes is calculated based on the time offset, and the sampling sequence is time-corrected to generate an initial alignment sequence.

[0090] Calculate the temporal consistency value and feature correlation value of the initial alignment sequence, adjust the temporal offset according to the consistency value and correlation value, optimize the initial alignment sequence, and obtain the alignment sequence.

[0091] In this embodiment, the dynamic constraint matrix is ​​first decomposed into eigenvalues, the decay curvature of the eigenvalues ​​is calculated to determine the number of principal eigenvalues, the eigenvectors corresponding to the principal eigenvalues ​​are extracted, and mapped to a nonlinear time-series space to obtain time-series offset features. Based on the time-series offset features, a multi-scale time-series sliding window is constructed. Local time-series change values ​​are extracted within each sliding window, and probability density estimation is performed on the local time-series change values ​​to obtain global time-series trend values. The local time-series change values ​​and global time-series trend values ​​are combined to reconstruct the time-series offset. The relative displacement relationship between modes is calculated based on the time-series offset, and the sampled sequence is time-corrected to generate an initial alignment sequence. The time-series consistency value and feature correlation value of the initial alignment sequence are calculated, and the time-series offset is adjusted based on the consistency value and correlation value to optimize the initial alignment sequence and obtain the alignment sequence.

[0092] The process of eigenvalue decomposition for the dynamic constraint matrix first involves acquiring the dataset of sampled sequences to be processed. This dataset contains temporal data of different modalities, such as video and audio data. The sampled sequences exhibit varying degrees of offset along the time axis. When constructing the dynamic constraint matrix, the temporal distance between different sampling points is calculated, forming an n×n square matrix, where n is the length of the sampled sequence. Specifically, for any two points i and j in the sampled sequence, the temporal distance between them is calculated as an element value of the constraint matrix, reflecting the correlation between the two sampling points in the time dimension. For example, for a video-audio dataset, 600 frames of video data and the corresponding audio data can be selected to construct a 600×600 dynamic constraint matrix.

[0093] The constructed dynamic constraint matrix is ​​subjected to eigenvalue decomposition, yielding a set of eigenvalues ​​and corresponding eigenvectors. These eigenvalues ​​are arranged in descending order, and the ratio between adjacent eigenvalues ​​is calculated to obtain the decay curvature of the eigenvalues. By setting a threshold for the decay curvature, the number of principal eigenvalues ​​is determined. In practical applications, the threshold is usually set between 0.1 and 0.2; that is, when the ratio between adjacent eigenvalues ​​is less than this threshold, the subsequent eigenvalues ​​are considered to have a small contribution to the system and can be ignored. For example, for the 600×600 constraint matrix mentioned above, the first 12 eigenvalues ​​may be identified as principal eigenvalues ​​after eigenvalue decomposition.

[0094] The eigenvectors corresponding to these principal eigenvalues ​​are extracted, with each eigenvector representing the projection of the time-series data into the eigenspace. These eigenvectors are then mapped to a nonlinear time-series space to obtain the time-series offset features. The nonlinear mapping is implemented using kernel functions, commonly including Gaussian and polynomial kernels. For this practical case, a Gaussian kernel is chosen with a kernel parameter set to 0.5, mapping the 12 principal eigenvectors to the nonlinear space to obtain a 600×12 dimensional time-series offset feature matrix.

[0095] When constructing a multi-scale temporal sliding window based on temporal offset features, different window sizes are set, such as 10 frames, 20 frames, and 30 frames. For each scale window, the window slides across the temporal offset feature to extract local temporal variation values ​​within each window. These local temporal variation values ​​are obtained by calculating the differences between adjacent frames within the window. For example, for a 10-frame window, calculating the differences between adjacent frames within these 10 frames yields 9 local temporal variation values. Sliding across the entire temporal offset feature yields multiple sets of local temporal variation values.

[0096] Probability density estimation is performed on these local temporal variation values ​​to obtain the global temporal trend value. The probability density estimation uses kernel density estimation, selecting an appropriate bandwidth parameter, such as 0.8, to estimate the distribution of local temporal variation values. Based on the estimated probability density function, the global temporal trend value at each time point is calculated. The local temporal variation values ​​and the global temporal trend value are then weighted and combined with a weight ratio of 6:4 to reconstruct the temporal offset. For 600 frames of data, 600 temporal offsets are ultimately obtained, representing the amount of temporal adjustment required for each frame.

[0097] The relative displacement between modalities is calculated based on the temporal offset. For both video and audio modalities, their relative displacement on the time axis is calculated. For example, if frame 100 of the video corresponds to frame 105 of the audio, the relative displacement between these two frames is 5 frames. Based on the calculated relative displacement, temporal correction is performed on the sampled sequence to generate an initial aligned sequence. Temporal correction is achieved through time interpolation or deletion of sample points, such as linear interpolation or cubic spline interpolation. In practical cases, cubic spline interpolation is used to adjust the video sequence to be aligned with the audio sequence.

[0098] Calculate the temporal consistency and feature correlation values ​​of the initial aligned sequence. The temporal consistency value represents the smoothness of the aligned sequence along the time axis; the standard deviation of temporal changes between adjacent frames is calculated, with a smaller standard deviation indicating higher consistency. The feature correlation value represents the degree of data correlation between different modalities; the cosine similarity between the feature vectors of two modalities is calculated, with higher similarity indicating stronger correlation. Based on the calculated consistency and correlation values, an optimization objective function is constructed with a 7:3 weight ratio, and the temporal offset is adjusted using gradient descent.

[0099] During the optimization process, the learning rate was set to 0.01 and the number of iterations was 100. The iteration was terminated early when the rate of change of the objective function was less than 0.001. The adjusted temporal offset was applied to the initial alignment sequence for further correction to obtain the final alignment sequence.

[0100] By extracting temporal offset features through eigenvalue decomposition of the dynamic constraint matrix, the problem of temporal asynchrony in multimodal data is solved. Multi-scale temporal sliding window analysis is used to analyze local and global temporal changes, accurately capturing temporal offset characteristics. Based on the reconstructed temporal offset, relative displacement correction between modalities is achieved, improving the accuracy of data temporal alignment. The alignment sequence is optimized through a dual evaluation mechanism of temporal consistency and feature correlation, enhancing the temporal matching degree between modalities. This effectively eliminates temporal deviations during multimodal data acquisition, providing a high-quality synchronous data foundation for subsequent feature fusion and sentiment recognition, and improving the overall system performance.

[0101] In one optional implementation, the temporal features are constructed as nodes in an evolutionary trajectory graph. The temporal dependency strength between nodes is calculated, and the nodes are divided into feature groups based on the temporal dependency strength. Predictive features for missing nodes are generated by combining the temporal evolution law of modality combinations, including:

[0102] The temporal features are constructed into nodes of the evolutionary trajectory graph. The length of the time window is dynamically adjusted according to the changing trend of the temporal features. Historical information is extracted within the time window to construct node features.

[0103] The temporal dependency strength between nodes is calculated using multidimensional information entropy and state transition probability. The temporal dependency strength is then corrected by combining the temporal consistency of node features to obtain the final temporal dependency strength.

[0104] Clustering nodes based on temporal dependency strength, determining cluster boundaries based on the similarity of node features, dividing nodes into feature groups, and obtaining the temporal evolution patterns of nodes within the feature groups;

[0105] By analyzing the node features of different feature groups, the temporal evolution law of mode combination is obtained. A co-evolution matrix of feature groups is constructed. The co-evolution matrix is ​​combined with the temporal evolution law to generate the predicted features of missing nodes.

[0106] The predicted features are compared with the actual features to calculate the prediction bias. The co-evolution matrix is ​​then optimized based on the prediction bias, and the predicted features are updated.

[0107] In one embodiment, time-series features are first constructed as nodes in an evolutionary trajectory graph. For a given time-series dataset, multiple feature dimensions are included, each corresponding to a time series. The time window length is dynamically adjusted based on the feature's changing trend. Specifically, the volatility of the time-series data at different time scales is calculated. When the volatility of a feature in recent data exceeds 150% of the historical average, the time window for that feature is shortened to 75% of its original length; when the volatility is below 50% of the historical average, the time window is extended to 125% of its original length. For example, for temperature sensor data, the initial window length is set to 60 minutes. When the temperature volatility in the last 10 minutes is detected to be 3.2℃ (exceeding 150% of the historical average volatility of 2.0℃), the window is adjusted to 45 minutes. Within the determined time window, statistical features of the historical data are extracted as node features, including statistics such as mean, standard deviation, kurtosis, skewness, maximum, and minimum values, as well as the main frequency components after Fourier transform.

[0108] Next, the temporal dependency strength between nodes is calculated, using a method combining multidimensional information entropy and state transition probability to quantify the dependency relationship between nodes. For any two nodes A and B, their feature values ​​are first discretized into a finite number of states, such as dividing the temperature value into several state intervals in 5°C increments. The joint information entropy of nodes A and B, as well as their respective edge information entropies, are calculated to obtain the preliminary temporal dependency strength. Simultaneously, a state transition matrix is ​​constructed to record the probability distribution of node B appearing in each state under different states of node A. For example, when the temperature sensor is in the 25-30°C state, the probability that the humidity sensor will appear in the 60%-70% range at the next moment is 0.72. The information entropy calculation result is combined with the state transition probability to obtain the initial temporal dependency strength. Subsequently, the dependency strength is corrected by considering the temporal consistency of node features. Temporal consistency is obtained by calculating the correlation of the changing trends of the features of two nodes. When the consistency of the changing trends of the two nodes is higher than 75%, their dependency strength is enhanced; when the consistency is lower than 30%, their dependency strength is weakened. Specifically, for a certain pressure sensor and flow sensor, the initial dependency strength is 0.68. Since the consistency of their changing trends reaches 82%, the corrected dependency strength is increased to 0.79.

[0109] Based on the calculated temporal dependency strength, nodes are clustered. An improved spectral clustering algorithm is used, with the temporal dependency strength matrix as input to the similarity matrix. To determine the optimal number of clusters, the silhouette coefficient is used to evaluate the quality of different clustering results, and the clustering scheme with the largest silhouette coefficient is selected. When determining the cluster boundaries, the similarity of node features is introduced as an auxiliary criterion. The cosine similarity between node feature vectors is calculated. When the dependency strength between two nodes is close to the cluster boundary (difference less than 0.1), if the feature similarity is higher than 0.85, they are grouped into the same feature group. For example, 10 sensor nodes in an industrial environment are divided into 3 feature groups after cluster analysis: environmental parameter group (temperature, humidity, air pressure), equipment status group (vibration, noise, current), and product quality group (size, weight, density, strength). For each feature group, by extracting the time series patterns of nodes within the group, periodicity, trend, and abnormal patterns are identified to obtain the temporal evolution law of nodes within the group.

[0110] Further analysis of the correlations between different feature groups reveals the temporal evolution patterns of modal combinations. By calculating the lag correlations between different feature groups, the time delay and intensity of inter-group influence are determined. For example, changes in the environmental parameter group will affect the equipment status group after 30 minutes, with an influence intensity of 0.64; changes in the equipment status group will affect the product quality group after 45 minutes, with an influence intensity of 0.78. Based on this, a co-evolution matrix for feature groups is constructed, where matrix elements represent the influence intensity and time delay between different feature groups. Combining the co-evolution matrix with the internal temporal evolution patterns of each group allows for the generation of predictive features for missing nodes.

[0111] For predicting missing nodes, the feature group to which the node belongs is first determined. Then, by combining real-time data from other nodes within the same group and the inter-group relationships recorded in the co-evolution matrix, a predicted feature value is generated. Taking the missing density sensor data in the product quality group as an example, using the size and weight data from the same group, combined with historical data from the equipment status group 45 minutes prior, the predicted missing density value is 1.42 g / cm³. 3 .

[0112] To verify the accuracy of the prediction, the predicted features were compared with the actual features, and indices such as root mean square error and mean absolute percentage error were calculated. For example, a predicted density value of 1.42 g / cm³ was used. 3 The actual value is 1.45 g / cm³. 3 The relative error is 2.07%. Based on the prediction deviation, the parameters in the co-evolution matrix are dynamically adjusted. When the cumulative prediction error exceeds a preset threshold (e.g., 5%), a matrix optimization process is triggered, updating the matrix elements using gradient descent to gradually bring the prediction results closer to the actual values. The updated co-evolution matrix will be used for subsequent prediction tasks, forming a closed-loop optimization mechanism.

[0113] By constructing an evolutionary trajectory graph, temporal features are visualized as nodes, facilitating the capture of complex temporal patterns. The time window length is dynamically adjusted to adapt to the changing frequency of different features, improving feature extraction accuracy. Multidimensional information entropy and state transition probability are used to calculate temporal dependency strength, accurately quantifying the degree of correlation between nodes. Feature groups are formed based on node clustering to uncover inherent temporal evolution patterns and enhance the collaborative analysis capability between modalities. A collaborative evolution matrix is ​​constructed to guide the prediction of missing node features, achieving intelligent data completion. Prediction model optimization through prediction bias feedback improves the accuracy of multimodal missing data recovery, providing more complete and reliable feature input for emotion recognition.

[0114] like Figure 2 As shown, this embodiment demonstrates the multimodal fusion and emotion recognition technology process based on temporal feature analysis.

[0115] In one optional implementation, the information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Within the temporal completion intervals, the predicted features are adaptively completed to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the samples is identified, including:

[0116] Calculate the conditional information entropy and joint information entropy of temporal features in the time dimension, combine them to generate an information entropy distribution, identify feature mutation points based on the gradient change of the information entropy distribution, and determine the region between adjacent feature mutation points as the temporal completion interval;

[0117] Extract the frequency distribution and cumulative distribution of time series features within the time series completion interval, construct a probability density function, and generate time series completion constraints based on the change patterns of the probability density function and historical data.

[0118] Based on the temporal completion constraint, the feature importance is calculated, and the predicted features are progressively weighted according to the feature importance. The progressive weighting result is used to adaptively complete the predicted features to generate multimodal fusion features.

[0119] The correlation coefficients and mutual information values ​​of the multimodal fusion features in the time dimension and feature dimension are calculated and combined to obtain the temporal correlation degree. The feature weights are iteratively optimized based on the temporal correlation degree to generate optimized multimodal fusion features. The optimized multimodal fusion features are then mapped to the sentiment feature space for sample sentiment recognition.

[0120] This implementation first requires calculating the conditional information entropy and joint information entropy of the temporal features along the time dimension, and then combining them to generate an information entropy distribution. Conditional information entropy reflects the uncertainty of the current feature value relative to past feature values, and can be obtained by calculating the conditional probability distribution of the current feature value given known past feature values. Joint information entropy characterizes the overall uncertainty of feature values ​​across multiple time points, and can be obtained by calculating the joint probability distribution of feature values ​​across multiple time points. Taking audio sentiment features as an example, a sentiment feature sequence of 5 consecutive seconds of audio data is extracted. For each time point t, the conditional information entropy of the feature at that time point and the features at the previous 1 to 3 seconds is calculated, and the joint information entropy of the feature at that time point and the features at the previous 1 to 3 seconds is also calculated. For video features and text features, a similar method is used to calculate their respective conditional information entropy and joint information entropy.

[0121] The conditional information entropy and joint information entropy are combined proportionally, for example, by assigning a weight of 0.6 to the conditional information entropy and 0.4 to the joint information entropy, resulting in a comprehensive information entropy distribution. Feature abrupt change points are identified based on the gradient changes in the information entropy distribution; these points represent locations where the gradient of the information entropy distribution changes significantly. Taking audio emotional features as an example, the first-order difference of the information entropy distribution is calculated. When the absolute value of the difference exceeds a preset threshold (e.g., 0.25), that time point is marked as a feature abrupt change point. In real-world scenarios, such as when a user transitions from a calm state to an excited state, the information entropy of the audio features will change significantly, potentially leading to feature abrupt change points.

[0122] The region between adjacent feature abrupt changes is defined as a temporal completion interval. For example, for 20 seconds of emotional interaction data, if feature abrupt changes are detected at the 5th and 12th seconds, then the interval between 5 and 12 seconds is defined as a temporal completion interval. In multimodal data, multiple temporal completion intervals may exist simultaneously, requiring separate processing for each interval.

[0123] Extract the frequency distribution and cumulative distribution of temporal features within the temporal completion interval. The frequency distribution describes the frequency of feature value occurrence, while the cumulative distribution represents the probability that a feature value is less than or equal to a certain value. Taking facial expression videos as an example, extract facial expression intensity features within the temporal completion interval, statistically analyze the frequency of occurrence of different intensity values, and calculate the cumulative probability of each intensity value. Construct a probability density function, which describes the probability density of a feature value at a certain point. For continuous features, kernel density estimation methods can be used to construct the probability density function; for discrete features, histograms or discrete probability distribution functions can be used.

[0124] Temporal completion constraints are generated based on the variation patterns of the probability density function and historical data. These constraints include the range of eigenvalues, their trends, and boundary conditions. For example, for the feature of speech emotion intensity, historical data analysis can show that under excited emotional states, speech emotion intensity typically fluctuates within the range of 0.7-0.9 with a stable trend; this can be used as a constraint.

[0125] Feature importance is calculated based on temporal completion constraints. Feature importance reflects the contribution of each feature to emotion recognition. Calculation methods can be based on feature stability, discriminative power, and temporal relevance. Stability is calculated using the variance of the feature within the temporal completion interval; a smaller variance indicates greater feature stability. Discriminative power is assessed by evaluating the feature's ability to distinguish different emotion categories, using mutual information or correlation analysis. Temporal relevance is assessed by evaluating the correlation between the feature and historical features. For facial expression videos, for example, the importance of eyebrow position features is 0.8, the importance of mouth corner features is 0.7, and the importance of eye opening / closing features is 0.6.

[0126] The predicted features are progressively weighted based on their importance. Progressive weighting means assigning weights to features in descending order of importance. First, the weight of the eyebrow position feature, which has the highest importance, is adjusted. Then, the weights of the mouth corner position and eye opening / closing features are adjusted sequentially. For missing video frames, the eyebrow position in the missing frame can be predicted based on the eyebrow position features of the preceding and following frames, and assigned a weight of 0.8; then the mouth corner position is predicted and assigned a weight of 0.7; finally, the eye opening / closing feature is predicted and assigned a weight of 0.6.

[0127] A progressively weighted approach is used to adaptively complete the predicted features, generating multimodal fusion features. Adaptive completion refers to dynamically adjusting the completion strategy based on the temporal characteristics and importance of the features. Linear interpolation can be used for slowly changing features, while nonlinear fitting can be used for features that change rapidly. The completed modal features are then fused to generate multimodal fusion features. For example, the completed video, audio, and text features are weighted and fused in a ratio of 0.4:0.3:0.3.

[0128] The correlation coefficients and mutual information values ​​of multimodal fusion features in the time and feature dimensions are calculated, and combined to obtain the temporal correlation degree. The time-dimensional correlation coefficient reflects the change pattern of the fusion features over time, while the feature-dimensional correlation coefficient reflects the dependency relationship between different features. The mutual information value measures the degree of mutual dependence between two variables. For example, the Pearson correlation coefficient of the fusion features at adjacent time points is 0.85, the average correlation coefficient between different features is 0.72, the mutual information value at adjacent time points is 0.68, and the average mutual information value between different features is 0.64. These indicators are combined according to preset weights (e.g., 0.3:0.3:0.2:0.2) to obtain the temporal correlation degree.

[0129] The features are weighted and iteratively optimized based on temporal correlation to generate optimized multimodal fusion features. During the iterative optimization process, the weights of each feature are adjusted to maximize the temporal correlation. For example, if the initial temporal correlation of the fusion features is 0.75, after three rounds of iterative optimization, the weights of the video, audio, and text features are adjusted to 0.45:0.25:0.3, at which point the temporal correlation increases to 0.82.

[0130] The optimized multimodal fusion features are mapped to the sentiment feature space for sample sentiment recognition. The sentiment feature space is a predefined multidimensional space, where each dimension corresponds to a sentiment attribute. Mapping methods can employ nonlinear transformations or deep neural networks. In the sentiment feature space, the distance or similarity between the sample features and the centers of each sentiment category is calculated, and the sample is classified into the sentiment category with the closest distance or the highest similarity. For example, after mapping the fusion features to the sentiment feature space, the similarities with the six basic sentiment categories of "happiness," "sadness," "anger," "fear," "surprise," and "disgust" are calculated to be 0.82, 0.15, 0.05, 0.04, 0.08, and 0.06, respectively. Therefore, the sample is identified as the "happiness" sentiment type.

[0131] By accurately identifying feature mutation points through information entropy distribution, intelligent division of temporal completion intervals is achieved, ensuring the rationality and coherence of feature completion. A progressive weighting strategy driven by feature importance is adopted to improve the accuracy of multimodal data completion. Iterative optimization of feature weighting is achieved through temporal correlation calculation, enhancing the expressive power of multimodal fusion features. The adaptive completion mechanism can dynamically adjust the completion strategy according to the temporal characteristics of features, effectively handling complex emotional change scenarios. Multi-dimensional correlation analysis ensures the temporal consistency of the completed features, improving the accuracy and robustness of emotion recognition, especially maintaining high recognition performance even in cases of missing data.

[0132] A second aspect of this invention provides a multimodal missing data emotion recognition system based on dynamic completion, the system comprising:

[0133] The first unit is used to acquire multimodal input data, detect the missing data locations of each modality, calculate the modality information completeness value and quality score, and generate a modality reliability weight vector.

[0134] The second unit is used to calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data.

[0135] The third unit is used to construct nodes of the evolutionary trajectory map from temporal features, calculate the temporal dependency strength between nodes, divide nodes into feature groups according to the temporal dependency strength, and generate predicted features for missing nodes by combining the temporal evolution law of modality combination.

[0136] The fourth unit is used to calculate the information entropy distribution of temporal features, divide the temporal completion interval according to the information entropy distribution, and perform adaptive completion on the predicted features within the temporal completion interval to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

[0137] A third aspect of the present invention provides an electronic device, comprising:

[0138] processor;

[0139] Memory used to store processor-executable instructions;

[0140] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0141] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0142] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A sentiment recognition method for multimodal missing data based on dynamic completion, characterized in that, include: Acquire multimodal input data, detect missing data locations for each modality, calculate modal information completeness value and quality score, and generate modal reliability weight vector; Calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data. The temporal features are constructed as nodes in the evolutionary trajectory graph. The temporal dependency strength between nodes is calculated. The nodes are divided into feature groups according to the temporal dependency strength. The predicted features of missing nodes are generated by combining the temporal evolution law of modality combination. The information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Within the temporal completion intervals, the predicted features are adaptively completed to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

2. The method according to claim 1, characterized in that, Acquire multimodal input data, detect missing data locations for each modality, calculate modal information completeness values ​​and quality scores, and generate modal reliability weight vectors including: Acquire multimodal input data including video, audio, and text; calculate the continuity difference value of adjacent data segments in the multimodal input data; mark the difference positions above the difference threshold as data missing positions; and generate a data missing label sequence. Calculate the temporal variation features of the non-missing data indicated by the missing data marker sequence, perform feature fusion on the temporal variation features to obtain the temporal correlation degree, and calculate the information completeness value of each modality based on the temporal correlation degree. Extract the content features and structural features of the non-missing data, combine them according to preset weights to obtain quality features, and calculate the quality score of each modality based on the quality features; The feature importance coefficient is calculated from the information completeness value and quality score. The feature importance coefficient is then normalized to obtain the modality reliability weight vector, which is used to characterize the importance of each modality feature in feature fusion.

3. The method according to claim 1, characterized in that, The cross-correlation coefficients between modes in the multimodal input data are calculated. Mode combinations are determined based on the cross-correlation coefficients and reliability weight vectors. Temporal offsets of the mode combinations are detected. The multimodal input data is dynamically aligned based on these temporal offsets. Temporal features are extracted from the dynamically aligned data, including: Calculate the temporal entropy value of each modal feature in the multimodal input data, construct an adaptive sampling probability distribution based on the temporal entropy value, resample the multimodal input data according to the adaptive sampling probability distribution to obtain a sampling sequence, and calculate the cross-correlation coefficient between modes in the sampling sequence; A dynamic constraint matrix is ​​constructed by combining the cross-correlation coefficient and the time entropy value. The dynamic constraint matrix is ​​multiplied by the reliability weight vector to obtain the modal confidence. A combination threshold is set according to the modal confidence. Modal combinations are constructed by pairing modes that are higher than the combination threshold. The temporal offset features are obtained by eigenvalue decomposition of the dynamic constraint matrix, the temporal offset is reconstructed from the temporal offset features, and the sampled sequence is dynamically aligned according to the temporal offset to obtain the aligned sequence. The variation features of the aligned sequence are extracted and weighted and fused with the suppression vector to obtain the temporal features.

4. The method according to claim 3, characterized in that, Temporal offset features are obtained through eigenvalue decomposition of the dynamic constraint matrix. These features are then reconstructed to obtain the temporal offset. The sampled sequences are then dynamically aligned based on the temporal offset to obtain the aligned sequences, which include: The dynamic constraint matrix is ​​decomposed into eigenvalues, the decay curvature of the eigenvalues ​​is calculated to determine the number of principal eigenvalues, the eigenvectors corresponding to the principal eigenvalues ​​are extracted, and they are mapped to the nonlinear time space to obtain the time-series offset features. Based on the aforementioned temporal offset features, a multi-scale temporal sliding window is constructed. Local temporal change values ​​are extracted within each sliding window. Probability density estimation is performed on the local temporal change values ​​to obtain the global temporal trend value. The local temporal change values ​​and the global temporal trend value are combined and reconstructed to obtain the temporal offset. The relative displacement relationship between modes is calculated based on the time offset, and the sampling sequence is time-corrected to generate an initial alignment sequence. Calculate the temporal consistency value and feature correlation value of the initial alignment sequence, adjust the temporal offset according to the consistency value and correlation value, optimize the initial alignment sequence, and obtain the alignment sequence.

5. The method according to claim 1, characterized in that, The temporal features are constructed as nodes in an evolutionary trajectory graph. The temporal dependency strength between nodes is calculated, and the nodes are divided into feature groups based on the temporal dependency strength. Predictive features for missing nodes are generated by combining the temporal evolution law of modality combinations, including: The temporal features are constructed into nodes of the evolutionary trajectory graph. The length of the time window is dynamically adjusted according to the changing trend of the temporal features. Historical information is extracted within the time window to construct node features. The temporal dependency strength between nodes is calculated using multidimensional information entropy and state transition probability. The temporal dependency strength is then corrected by combining the temporal consistency of node features to obtain the final temporal dependency strength. Clustering nodes based on temporal dependency strength, determining cluster boundaries based on the similarity of node features, dividing nodes into feature groups, and obtaining the temporal evolution patterns of nodes within the feature groups; By analyzing the node features of different feature groups, the temporal evolution law of mode combination is obtained. A co-evolution matrix of feature groups is constructed. The co-evolution matrix is ​​combined with the temporal evolution law to generate the predicted features of missing nodes. The predicted features are compared with the actual features to calculate the prediction bias. The co-evolution matrix is ​​then optimized based on the prediction bias, and the predicted features are updated.

6. The method according to claim 1, characterized in that, The information entropy distribution of temporal features is calculated, and temporal completion intervals are divided according to the information entropy distribution. Adaptive completion of the predicted features within the temporal completion intervals yields multimodal fusion features. Based on these multimodal fusion features, the sentiment types of samples are identified, including: Calculate the conditional information entropy and joint information entropy of temporal features in the time dimension, combine them to generate an information entropy distribution, identify feature mutation points based on the gradient change of the information entropy distribution, and determine the region between adjacent feature mutation points as the temporal completion interval; Extract the frequency distribution and cumulative distribution of time series features within the time series completion interval, construct a probability density function, and generate time series completion constraints based on the change patterns of the probability density function and historical data. Based on the temporal completion constraint, the feature importance is calculated, and the predicted features are progressively weighted according to the feature importance. The progressive weighting result is used to adaptively complete the predicted features to generate multimodal fusion features. The correlation coefficients and mutual information values ​​of the multimodal fusion features in the time dimension and feature dimension are calculated and combined to obtain the temporal correlation degree. The feature weights are iteratively optimized based on the temporal correlation degree to generate optimized multimodal fusion features. The optimized multimodal fusion features are then mapped to the sentiment feature space for sample sentiment recognition.

7. A multimodal missing data sentiment recognition system based on dynamic completion, used to implement the method of any one of claims 1-6, characterized in that, include: The first unit is used to acquire multimodal input data, detect the missing data locations of each modality, calculate the modality information completeness value and quality score, and generate a modality reliability weight vector. The second unit is used to calculate the cross-correlation coefficient between modes in the multimodal input data, determine the mode combination based on the cross-correlation coefficient and the reliability weight vector, detect the temporal offset of the mode combination, dynamically align the multimodal input data according to the temporal offset, and extract temporal features from the dynamically aligned data. The third unit is used to construct nodes of the evolutionary trajectory map from temporal features, calculate the temporal dependency strength between nodes, divide nodes into feature groups according to the temporal dependency strength, and generate predicted features for missing nodes by combining the temporal evolution law of modality combination. The fourth unit is used to calculate the information entropy distribution of temporal features, divide the temporal completion interval according to the information entropy distribution, and perform adaptive completion on the predicted features within the temporal completion interval to obtain multimodal fusion features. Based on the multimodal fusion features, the sentiment type of the sample is identified.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Neural network model construction method, time sequence prediction method and device

    CN114511065A

  • Multi-modal sentiment analysis method, system and equipment based on similar modal completion

    CN117540007A

  • Patient emotion analysis device based on multiple modes

    CN119174609A

  • Semantic comprehension driven cross-modal information fusion and retrieval method and system

    CN120448563A

  • Prompt-driven mode-missing-oriented multi-mode sentiment analysis method and system

    CN120524445A

Cited By

  • Pilot cognitive load real-time monitoring system based on multi-modal data

    CN121421473A