Multi-modal perception and video understanding method based on S1iME framework
Through the multimodal perception and video understanding method based on the S1iME framework, the problem of difficulty in fusion and alignment of multimodal information in the existing technology is solved, dynamic memory and self-supervised learning are realized, the adaptability and generalization capabilities of the system are improved, and efficient and robust.
Patent Information
- Application Number
- CN202510135509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimodal perception technologies are difficult to achieve effective information fusion and alignment when dealing with complex scenarios, and lack dynamic memory and self-supervised learning mechanisms, resulting in insufficient adaptability and generalization capabilities in dynamic changeable and unlabeled data scenarios.
A multimodal perception and video understanding method based on the S1iME framework is proposed, and feature alignment and fusion is performed through cross-modal self-attention mechanism, dynamic memory module and time-dependent modeling are introduced, and a self-supervised learning mechanism is combined to optimize model weights and inference strategies.
It realizes dynamic fusion and alignment of multimodal information in complex scenarios, improves the system's understanding of multimodal data, enhances the adaptability and generalization capabilities of the model, and has high efficiency, robustness and wide application potential.
Smart Images

Figure CN120067983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal learning, and particularly to a multimodal perception and video understanding method based on the S1iME framework. Background Art
[0002] In modern society, with the rapid development of artificial intelligence and data processing technologies, multimodal perception and fusion technologies have become one of the hotspots in the field of artificial intelligence research. Multimodal technologies aim to process information from different sensory or data sources (such as vision, audition, text, etc.), and through the deep fusion of these heterogeneous information, achieve complex tasks such as video understanding, sentiment analysis, and event prediction. Existing multimodal perception technologies are widely applied in scenarios such as autonomous driving, intelligent monitoring, medical diagnosis, and human-computer interaction, playing an important role in improving the performance and adaptability of intelligent systems.
[0003] Currently, for the perception and fusion of multimodal data, existing technologies usually rely on traditional deep learning frameworks. These frameworks mainly process data of different modalities through specific network architectures (such as convolutional neural networks, recurrent neural networks, and Transformers), and then fuse the features of each modality at a specific level to achieve joint modeling. Common fusion methods include feature-level fusion, decision-level fusion, and intermediate-layer fusion. Among them, feature-level fusion realizes information fusion by connecting modality features at the input layer, but it is vulnerable to data redundancy and feature conflicts. Decision-level fusion depends on separately modeling each modality and then weighted-combining the decision results of each modality. However, this method is difficult to capture the deep-level correlation information between modalities. Intermediate-layer fusion attempts to combine modality features in the middle layer of the network, which can not only maintain the independence of modalities but also capture their interaction relationships. However, these methods usually require complex feature alignment and data preprocessing steps, and at the same time, they cannot fully adapt to dynamically changing scenarios.
[0004] In the prior art, multimodal alignment and information consistency are also a core challenge. There are significant differences in data distributions and sampling frequencies between different modalities. For example, visual data usually contains high-dimensional spatial information, while text data mainly consists of serialized semantic information. Most existing alignment methods rely on simple matching strategies based on time steps or feature spaces. This static alignment mechanism is difficult to cope with the dynamic changes in complex scenarios. In addition, there are also problems of information redundancy and noise interference in the process of processing multimodal data. Existing technologies often require a large amount of manual design and data annotation to improve the performance of the model, which limits its applicability in data-scarce scenarios.
[0005] In terms of modeling methods, existing multi-modal learning frameworks also have limitations in dealing with complex temporal dependencies. Video understanding tasks usually require the system to capture the associations between events over a long time span, which poses extremely high requirements on the memory and reasoning capabilities of the model. Existing recurrent neural networks (such as LSTM and GRU) can handle temporal dependencies to a certain extent, but still have problems such as vanishing gradients and computational bottlenecks when faced with long sequence data. Although Transformer-based methods can capture long-range dependencies through self-attention mechanisms, their computational complexity increases exponentially with the increase in the length of the input sequence, limiting their application in high-dimensional temporal data. In addition, existing technologies generally lack a dynamic memory mechanism for key objects, events, and emotional information in video data, making it difficult to fully utilize context information for deep reasoning and generation.
[0006] Another key issue is the lack of generalization ability and adaptability in multi-modal perception technologies. Traditional methods are highly sensitive to changes in data distribution and cannot achieve good performance in unlabeled or weakly labeled data scenarios. Although self-supervised learning methods that have emerged in recent years reduce the dependence on manual labeling, most studies are still limited to single-modal tasks. How to effectively apply self-supervised learning mechanisms to multi-modal scenarios remains an unsolved problem. In addition, existing technologies usually have a fixed design for model weights in the training and inference stages, lacking the ability of dynamic optimization and update, which does not match the diversity and temporal changes of data in actual application scenarios.
[0007] Therefore, how to provide a multi-modal perception and video understanding method based on the S1iME framework is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0008] An object of the present invention is to propose a multi-modal perception and video understanding method based on the S1iME framework. The present invention makes full use of technologies such as multi-modal learning, dynamic memory optimization, and self-supervised learning, and details the complete processes of multi-modal feature fusion, cross-modal alignment, dynamic memory update, and reasoning generation, with the following advantages: First, it can achieve dynamic fusion and alignment of multi-modal information in complex scenarios, improving the depth and accuracy of the system's understanding of multi-modal data; Second, by introducing time-dependent modeling and dynamic memory mechanisms, it effectively enhances the model's ability to capture event correlations and emotional changes in long time sequences; Third, combined with self-supervised learning mechanisms, it reduces the dependence on large-scale labeled data and improves the adaptability and generalization ability of the model in weakly labeled or unlabeled scenarios. The implementation method of the present invention has high efficiency, adaptability, and robustness, and can be widely applied to fields such as video content understanding, emotion analysis, and event prediction.
[0009] A multi-modal perception and video understanding method based on the S1iME framework according to an embodiment of the present invention includes the following steps:
[0010] S1. Receive video, audio, and text data, preprocess the data, and generate structured multi-modal data;
[0011] S2. Extract features from the structured multi-modal data to construct a preliminary multi-modal feature representation;
[0012] S3. Based on the preliminary multi-modal feature representation, use the cross-modal self-attention mechanism in the S1iME framework for feature alignment and fusion, and generate a multi-modal feature representation through dynamic interaction modeling;
[0013] S4. Input the multi-modal feature representation into the short-term memory module in the S1iME framework, combine the long-term memory module to dynamically update the historical information across time periods, and generate a time-enhanced multi-modal feature representation;
[0014] S5. Based on the time-enhanced multi-modal feature representation, perform temporal modeling through the inference module in the S1iME framework to generate an inference feature representation;
[0015] S6. Input the inference feature representation into the generation module in the S1iME framework, and generate a multi-modal understanding result by combining the current multi-modal feature representation and the inference feature representation;
[0016] S7. Optimize the S1iME framework through a self-supervised learning mechanism, and use the correlation information between the multi-modal understanding result and the multi-modal data segments for deviation comparison and feature correction;
[0017] S8. Based on the optimized S1iME framework, combine the multi-modal feature representation and the inference feature representation to output the result of video understanding.
[0018] Optionally, the S3 specifically includes:
[0019] S31. For the preliminary multi-modal feature representation, perform linear transformations on the features of the video modality, audio modality, and text modality respectively to generate a query matrix Q, a key matrix K, and a value matrix V;
[0020] S32. Based on the query matrix Q and the key matrix K, calculate the interaction weights between the video modality, audio modality, and text modality, and adopt the self-attention mechanism:
[0021]
[0022] where α ij is the interaction weight of the i-th modality feature with respect to the j-th modality feature, d is the dimension of the feature vector, n is the total number of modality features, exp is the exponential function, Kj is the key vector of the j-th modal feature, K k is the key vector of the k-th modal feature, Q i is the query vector of the i-th modal feature;
[0023] S33. Based on the modal interaction weight α ij and the value vector V j , perform weighted summation and non-linear transformation on the features between modalities to generate a fused multi-modal feature representation:
[0024]
[0025] where F i is the fused i-th modal feature representation, W f and b f are learnable fusion parameter matrices and bias vectors, σ is a non-linear activation function, n is the total number of modal features, and j is an index used to traverse each modal feature and perform calculations in the weighted summation;
[0026] S34. Based on the fused multi-modal feature representation F i , dynamically adjust the feature importance of different modalities to generate an optimized weighted feature representation:
[0027] where the dynamic weight w i is calculated, and the mutual information between modalities is introduced as an additional constraint:
[0028] w i = εW a ·F i + b a + λ·HF i ,F j ;
[0029] where the modal weight w i is applied to the fused feature representation F i to generate an optimized weighted feature representation F' i :
[0030] F' i = w i ·F i + γ·W r ·F i ;
[0031] where ε is a non-linear activation function, W a and b a are learnable parameter matrices and biases, λ is a mutual information weight factor, HF i ,F jis the mutual information between the features of modality i and modality j, γ is the residual coefficient, and W r is the learnable parameter matrix;
[0032] S35. Perform deep interaction and context modeling on the optimized weighted feature representation F' i to capture the global correlation information and temporal dependence relationship between modalities, enhance the context of multimodal features, and generate the context-enhanced feature representation F” i ;
[0033]
[0034] where LayerNorm is the layer normalization operation, Concat is to concatenate the outputs of h attention heads into a complete context representation, is the normalization of the attention weights, is the scaling factor, and W o is the output transformation matrix of the multi-head attention, Q t , K t , V t are the query vector, key vector, and value vector respectively, H t-1 is the hidden state of the previous time step, and W h is the linear transformation matrix of the time-dependent features;
[0035] S36. Map the context-enhanced feature representation F” i to a unified semantic space, and generate the multimodal feature representation U through semantic alignment, normalization, and feature fusion t :
[0036]
[0037] where μ is the non-linear activation function, W u is the mapping matrix of the unified semantic space, b u is the bias term, U i is the unified representation of the i-th modality, w i is the fusion weight of the i-th modality, and n is the total number of modality features.
[0038] Optionally, the S4 specifically includes:
[0039] S41. Input the generated multimodal representation F” i into the short-term memory module to extract the key information of the current time segment:
[0040] M s = τW s ·F” i + b s ;
[0041] Among them, M s is the current time segment feature output by the short-term memory module, W s and b s are the learnable weight matrix and bias of the short-term memory respectively, and τ is the non-linear activation function;
[0042] S42. Dynamically update the long-term memory state based on the historical state stored in the long-term memory module and the current time segment feature M s :
[0043]
[0044] Among them, H t-1 is the long-term memory state of the previous time step, is the weighted intensity of the short-term memory calculated based on the attention mechanism, is the scaling factor, W a is the modal attention weight matrix, GRU is the gated recurrent unit, and H t is the updated long-term memory state;
[0045] S43. Weightedly fuse the short-term memory feature M s and the long-term memory state H t to generate a dynamic multi-modal feature representation containing time dependence:
[0046] F d =α s ·M s +α h ·H t +γ·M s -H t ;
[0047] Among them, F d is the dynamic multi-modal feature representation, α s and α h respectively represent the dynamic weights of the short-term memory and the long-term memory, and γ is the residual coefficient;
[0048] S44. Optimize the modal weights and standardize the dynamic feature representation F d to generate a time-enhanced multi-modal feature representation:
[0049]
[0050] Among them, F n is the time-enhanced multi-modal feature representation, LayerNorm is to standardize the fused features, w i is the weight of the i-th modality, σ is the non-linear activation function, is the linear transformation matrix of the i-th modality, is the bias term for the i-th modality, and n represents the total number of modalities.
[0051] Optionally, the S5 specifically includes:
[0052] S51. Input the generated time-enhanced multi-modal feature representation F n into the inference module as the initial feature representation for temporal modeling;
[0053] S52. Based on the time-enhanced multi-modal feature representation F n , use the gated recurrent unit for temporal modeling to dynamically capture temporal dependencies, and balance the contributions of the current time-step features and historical states through the modal temporal weight optimization mechanism:
[0054] H t = γ · GRU(H t-1 , F n + 1 - γ · F n );
[0055] where γ is the dynamic temporal weight, GRU is the gated recurrent unit, H t-1 is the hidden state of the previous time step, and H t is the hidden state of the current time step;
[0056] S53. Based on the hidden state H t of the current time step and the time-enhanced multi-modal feature representation F n , perform dynamic weight allocation, and generate the time-dynamic multi-modal fusion feature by combining cross-modal relationships:
[0057]
[0058] where M t is the multi-modal fusion feature of the current time step, is to calculate the attention weight of the i-th modality, is the scaling factor, is the modality-specific weight matrix, is the optimization matrix, is the bias vector of the i-th modality, and n represents the total number of modalities;
[0059] S54. Perform residual update and normalization operations on the time-dynamic multi-modal fusion feature M t , and the time-enhanced multi-modal feature representation F n , to generate the optimized multi-modal feature representation F u :
[0060] F u = LayerNorm(M t + F n · W r+b r ;
[0061] wherein, W r is a learnable transformation matrix, LayerNorm is a layer normalization operation, and b r is a bias vector;
[0062] S55. Based on the optimized multi-modal feature representation F u , capture the temporal correlation and semantic relationship between features through the inference module to generate an inference feature representation R t :
[0063] R t = LayerNormσF u ·W p +b p +H t ;
[0064] wherein, σ is a non-linear activation function, W p is a weight matrix, and b p is a bias term.
[0065] Optionally, the S6 specifically includes:
[0066] S61. Input the generated inference feature representation R t into the generation module as the initial feature representation for the multi-modal generation task;
[0067] S62. Based on the input inference feature representation R t , perform fusion processing on it through the cross-modal generation unit to generate a multi-modal fusion feature G t :
[0068]
[0069] wherein, is the normalization of the output after calculating the modality-specific weights for the inference features, is a scaling factor, is the weight matrix of the i-th modality, is the feature optimization matrix of the i-th modality, is the bias term of the i-th modality, σ is a non-linear activation function, and W h is the global transformation matrix, and n represents the total number of modalities;
[0070] S63. Perform semantic mapping for the specific task on the multi-modal fusion feature G t to generate a semantic description feature representation S t :
[0071] S t = DecoderσGt ·W u +b s ;
[0072] where Decoder is the decoding function in the generation module, and b s is the bias term for semantic generation, and W u is the task-specific semantic generation weight matrix;
[0073] S64. Generate an event prediction feature representation E t through temporal modeling based on the multi-modal fusion feature G t :
[0074] E t = GRU(E t-1 , G t );
[0075] where E t is the event prediction feature at the current time step, GRU is the gated recurrent unit for temporal modeling, and E t-1 is the event prediction hidden state at the previous time step;
[0076] S65. Map the multi-modal fusion feature G t to the sentiment space, and generate a sentiment analysis feature representation through dynamic weight allocation, modal feature optimization, and global feature enhancement:
[0077]
[0078] where α i is the dynamic sentiment weight of the i-th modality, is the feature optimization weight matrix of the i-th modality, is the bias term of the i-th modality, and W c is the global sentiment optimization matrix;
[0079] S66. Output the generated semantic description feature representation S t , the event prediction feature representation E t , and the sentiment analysis feature representation A t as the multi-modal understanding result O t .
[0080] Optionally, the specific steps of S7 include:
[0081] S71. Generate a self-supervised learning signal based on the multi-modal understanding result O t and the inference feature representation R t , and construct the association relationship between multi-modal segments:
[0082]
[0083] Among them, L context is the context relevance loss, T is the time step, Sim is the similarity function, and t is the time step;
[0084] S72. Optimize the robustness of the multi-modal understanding result based on positive and negative sample contrast learning. The positive sample is the true understanding result, and the negative sample is generated from the interference segment:
[0085]
[0086] Among them, L contrast is the loss function of contrast learning, ln is the logarithmic function, exp is the exponential function, is the positive sample at the current time step, is the j-th candidate sample, and N is the total number of candidate samples;
[0087] S73. Based on the inference feature representation R t and the multi-modal fusion feature G t , generate the updated inference feature representation through the dynamic correction mechanism:
[0088]
[0089] Among them, LayerNorm is the layer normalization operation, W a is the weight matrix, b k is the bias term, is the updated inference feature representation;
[0090] S74. Combine the updated inference feature representation with the optimized multi-modal feature representation F u to capture the context information through time series modeling and generate the context enhanced feature
[0091]
[0092] Among them, GRU is the gated recurrent unit, H t-1 is the hidden state of the previous time step;
[0093] S75. Based on the context enhanced feature, combine the context relevance loss and the contrast loss to construct the global optimization objective, and optimize the model parameters through the self-supervised learning mechanism.
[0094] The beneficial effects of the present invention are:
[0095] Through the multi-modal perception and video understanding method of the present invention, many defects in the prior art in multi-modal information processing and fusion are overcome, and the adaptability and reasoning depth of the system in complex dynamic scenarios are significantly improved. The S1iME framework proposed by the present invention can achieve dynamic fusion and cross-modal alignment of multi-modal data, solve the problems of information redundancy and feature conflict in traditional methods, and thus improve the effectiveness of information interaction between modalities. In addition, by introducing a dynamic memory module and a time-dependent modeling mechanism, the present invention can capture the time correlation of key objects and events in video content, and thus achieve more accurate event understanding and sentiment analysis on long-time series data.
[0096] The combination of the self-supervised learning mechanism further reduces the dependence on manually labeled data, enabling the present invention to also have excellent performance in weakly labeled or unlabeled data scenarios. This mechanism not only enhances the generalization ability of the system but also significantly improves its adaptability in real-time applications. At the same time, the dynamic optimization module designed by the present invention can flexibly adjust the weights and reasoning strategies of the model according to the characteristics of the input data and task requirements, so as to maintain high processing performance in complex scenarios. This innovative multi-modal perception and fusion method shows excellent robustness and accuracy in multi-task environments such as video understanding, sentiment analysis, and event prediction, providing an effective solution to the multi-modal perception problem in practical applications.
[0097] Compared with traditional methods, the present invention significantly improves the accuracy of multi-modal data fusion, the coherence of reasoning, and the adaptive ability to environmental changes, and has high efficiency, robustness, and broad application potential. Brief Description of the Drawings
[0098] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0099] Figure 1 is a flowchart of the multi-modal perception and video understanding method based on the S1iME framework proposed by the present invention;
[0100] Figure 2 is a schematic structural diagram of the multi-modal feature extraction and fusion module of the multi-modal perception and video understanding method based on the S1iME framework proposed by the present invention. Detailed Embodiments
[0101] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0102] Reference Figure 1-2, A multi-modal perception and video understanding method based on the S1iME framework, comprising the following steps:
[0103] S1. Receive video, audio, and text data, preprocess the data, and generate structured multi-modal data;
[0104] S2. Extract features from the structured multi-modal data to construct a preliminary multi-modal feature representation;
[0105] S3. Based on the preliminary multi-modal feature representation, use the cross-modal self-attention mechanism in the S1iME framework for feature alignment and fusion, and generate a multi-modal feature representation through dynamic interaction modeling;
[0106] S4. Input the multi-modal feature representation into the short-term memory module in the S1iME framework, combine it with the long-term memory module to dynamically update the historical information across time periods, and generate a time-enhanced multi-modal feature representation;
[0107] S5. Based on the time-enhanced multi-modal feature representation, perform temporal modeling through the inference module in the S1iME framework to generate an inference feature representation;
[0108] S6. Input the inference feature representation into the generation module in the S1iME framework, and generate a multi-modal understanding result by combining the current multi-modal feature representation and the inference feature representation;
[0109] S7. Optimize the S1iME framework through a self-supervised learning mechanism, and use the association information between the multi-modal understanding result and the multi-modal data segments for deviation comparison and feature correction;
[0110] S8. Based on the optimized S1iME framework, combine the multi-modal feature representation and the inference feature representation to output the result of video understanding.
[0111] In this embodiment, the specific steps of S3 include:
[0112] S31. For the preliminary multi-modal feature representation, perform linear transformations on the features of the video modality, audio modality, and text modality respectively to generate a query matrix Q, a key matrix K, and a value matrix V;
[0113] S32. Based on the query matrix Q and the key matrix K, calculate the interaction weights between the video modality, audio modality, and text modality, and adopt the self-attention mechanism:
[0114]
[0115] where α ij is the interaction weight of the i-th modality feature with respect to the j-th modality feature, d is the dimension of the feature vector, n is the total number of modality features, exp is the exponential function, and K jis the key vector of the j-th modal feature, K k is the key vector of the k-th modal feature, Q i is the query vector of the i-th modal feature;
[0116] S33. Based on the modal interaction weight α ij and the value vector V j , perform weighted summation and non-linear transformation on the features between modalities to generate a fused multi-modal feature representation:
[0117]
[0118] where F i is the fused i-th modal feature representation, W f and b f are learnable fusion parameter matrices and bias vectors, σ is a non-linear activation function, n is the total number of modal features, and j is an index used to traverse each modal feature and perform calculations in the weighted summation;
[0119] S34. Based on the fused multi-modal feature representation F i , dynamically adjust the feature importance of different modalities to generate an optimized weighted feature representation:
[0120] where the dynamic weight w i is calculated, and the mutual information between modalities is introduced as an additional constraint:
[0121] w i = εW a ·F i + b a + λ·HF i ,F j ;
[0122] where the modal weight w i is applied to the fused feature representation F i to generate an optimized weighted feature representation F' i :
[0123] F' i = w i ·F i + γ·W r ·F i ;
[0124] where ε is a non-linear activation function, W a and b a are learnable parameter matrices and biases, λ is the mutual information weight factor, HF i ,F j is the mutual information between the features of modality i and modality j, and γ is the residual coefficient, Wr is a learnable parameter matrix;
[0125] S35. Perform deep interaction and context modeling on the optimized weighted feature representation F', i capture the global correlation information and temporal dependence relationship between modalities, enhance the context of the multimodal features, and generate a context-enhanced feature representation F'' i ;
[0126]
[0127] where LayerNorm is a layer normalization operation, Concat is to concatenate the outputs of h attention heads into a complete context representation, is the normalization of the attention weights, is the scaling factor, W o is the output transformation matrix of the multi-head attention, Q t , K t , V t are the query vector, key vector, and value vector respectively, H t-1 is the hidden state of the previous time step, W h is the linear transformation matrix of the time-dependent features;
[0128] S36. Map the context-enhanced feature representation F'' i to a unified semantic space, and generate a multimodal feature representation U through semantic alignment, normalization, and feature fusion t :
[0129]
[0130] where μ is a non-linear activation function, W u is the mapping matrix of the unified semantic space, b u is the bias term, U i is the unified representation of the i-th modality, w i is the fusion weight of the i-th modality, and n is the total number of modality features.
[0131] In this embodiment, the S4 specifically includes:
[0132] S41. Input the generated multimodal representation F'' i into the short-term memory module to extract the key information of the current time segment:
[0133] M s = τW s ·F'' i + b s ;
[0134] where M sis the current time segment feature output by the short-term memory module, W s and b s are the learnable weight matrix and bias of the short-term memory respectively, and τ is the non-linear activation function;
[0135] S42. Dynamically update the long-term memory state based on the historical state stored in the long-term memory module and the current time segment feature M s , as follows:
[0136]
[0137] where H t-1 is the long-term memory state at the previous time step, is the weighted intensity of the short-term memory calculated based on the attention mechanism, is the scaling factor, W a is the modal attention weight matrix, GRU is the gated recurrent unit, and H t is the updated long-term memory state;
[0138] S43. Weightedly fuse the short-term memory feature M s and the long-term memory state H t to generate a dynamic multi-modal feature representation containing time dependence:
[0139] F d =α s ·M s +α h ·H t +γ·M s -H t ;
[0140] where F d is the dynamic multi-modal feature representation, α s and α h represent the dynamic weights of the short-term memory and the long-term memory respectively, and γ is the residual coefficient;
[0141] S44. Optimize the modal weights and perform normalization processing on the dynamic feature representation F d to generate a time-enhanced multi-modal feature representation:
[0142]
[0143] where F n is the time-enhanced multi-modal feature representation, LayerNorm is to normalize the fused features, w i is the weight of the i-th modality, σ is the non-linear activation function, is the linear transformation matrix of the i-th modality, is the bias term for the i-th modality, and n represents the total number of modalities.
[0144] In this embodiment, S5 specifically includes:
[0145] S51. Input the generated time-enhanced multi-modal feature representation F n into the inference module as the initial feature representation for temporal modeling;
[0146] S52. Based on the time-enhanced multi-modal feature representation F n , use a gated recurrent unit for temporal modeling to dynamically capture temporal dependencies, and balance the contributions of the current time-step features and historical states through a modal temporal weight optimization mechanism:
[0147] H t = γ · GRU(H t-1 , F n ) + 1 - γ · F n ;
[0148] where γ is the dynamic temporal weight, GRU is the gated recurrent unit, H t-1 is the hidden state of the previous time step, and H t is the hidden state of the current time step;
[0149] S53. Based on the hidden state H t of the current time step and the time-enhanced multi-modal feature representation F n , perform dynamic weight allocation, and generate a time-dynamic multi-modal fusion feature by combining cross-modal relationships:
[0150]
[0151] where M t is the multi-modal fusion feature of the current time step, is to calculate the attention weight of the i-th modality, is the scaling factor, is the modality-specific weight matrix, is the optimization matrix, is the bias vector of the i-th modality, and n represents the total number of modalities;
[0152] S54. Perform residual update and normalization operations on the time-dynamic multi-modal fusion feature M t , and the time-enhanced multi-modal feature representation F n , to generate an optimized multi-modal feature representation F u :
[0153] F u = LayerNorm(M t ) + F n · Wr +b r ;
[0154] where W r is a learnable transformation matrix, LayerNorm is a layer normalization operation, and b r is a bias vector;
[0155] S55. Based on the optimized multi-modal feature representation F u , capture the temporal correlation and semantic relationship between features through the inference module to generate the inference feature representation R t :
[0156] R t = LayerNormσF u ·W p +b p +H t ;
[0157] where σ is a non-linear activation function, W p is a weight matrix, and b p is a bias term.
[0158] In this embodiment, the specific steps of S6 include:
[0159] S61. Input the generated inference feature representation R t into the generation module as the initial feature representation for the multi-modal generation task;
[0160] S62. Based on the input inference feature representation R t , perform fusion processing on it through the cross-modal generation unit to generate the multi-modal fusion feature G t :
[0161]
[0162] where is the normalization of the output after calculating the modality-specific weights of the inference features, is the scaling factor, is the weight matrix of the i-th modality, is the feature optimization matrix of the i-th modality, is the bias term of the i-th modality, σ is a non-linear activation function, and W h is the global transformation matrix, and n represents the total number of modalities;
[0163] S63. Perform semantic mapping for the specific task on the multi-modal fusion feature G t to generate the semantic description feature representation S t :
[0164] S t= DecoderσG t ·W u + b s ;
[0165] Among them, Decoder is the decoding function in the generation module, and b s is the bias term for semantic generation, and W u is the task-specific semantic generation weight matrix;
[0166] S64. Based on the multi-modal fusion feature G t , generate the event prediction feature representation E t through temporal modeling:
[0167] E t = GRU(E t-1 , G t );
[0168] Among them, E t is the event prediction feature at the current time step, GRU is the gated recurrent unit for temporal modeling, and E t-1 is the event prediction hidden state at the previous time step;
[0169] S65. Map the multi-modal fusion feature G t to the sentiment space, and generate the sentiment analysis feature representation through dynamic weight allocation, modal feature optimization, and global feature enhancement:
[0170]
[0171] Among them, α i is the dynamic sentiment weight of the i-th modality, is the feature optimization weight matrix of the i-th modality, is the bias term of the i-th modality, and W c is the global sentiment optimization matrix;
[0172] S66. Output the generated semantic description feature representation S t , the event prediction feature representation E t and the sentiment analysis feature representation A t as the multi-modal understanding result O t .
[0173] In this embodiment, the S7 specifically includes:
[0174] S71. Based on the multi-modal understanding result O t and the inference feature representation R t , generate a self-supervised learning signal, and construct the association relationship between multi-modal segments:
[0175]
[0176] Among them, L context is the context relevance loss, T is the time step, Sim is the similarity function, and t is the time step;
[0177] S72. Optimize the robustness of the multi-modal understanding result based on positive and negative sample contrast learning. The positive sample is the true understanding result, and the negative sample is generated from the interference segment:
[0178]
[0179] Among them, L contrast is the loss function of contrast learning, ln is the logarithmic function, exp is the exponential function, is the positive sample at the current time step, is the j-th candidate sample, and N is the total number of candidate samples;
[0180] S73. Based on the inference feature representation R t and the multi-modal fusion feature G t , generate the updated inference feature representation through the dynamic correction mechanism:
[0181]
[0182] Among them, LayerNorm is the layer normalization operation, W a is the weight matrix, b k is the bias term, is the updated inference feature representation;
[0183] S74. Combine the updated inference feature representation with the optimized multi-modal feature representation F u to capture the context information through time series modeling and generate the context-enhanced feature
[0184]
[0185] Among them, GRU is the gated recurrent unit, and H t-1 is the hidden state of the previous time step;
[0186] S75. Based on the context-enhanced feature, construct the global optimization objective by combining the context relevance loss and the contrast loss, and optimize the model parameters through the self-supervised learning mechanism.
[0187] Example 1:
[0188] To verify the feasibility of the present invention in implementation, the present invention is applied to intelligent monitoring in a public place. The experimental scenario is the public area of a large shopping mall, involving real-time processing and analysis of multi-modal data, including video data, audio data, and text data. The core tasks of the experiment are to identify potential abnormal events, conduct sentiment analysis, and generate video content summaries.
[0189] The main challenges in this scenario lie in how to fuse data from different modalities (such as video, audio, and text), accurately detect abnormal events and emotional changes, and predict the possible development of events. Traditional multi-modal analysis techniques often suffer from information redundancy and modality conflicts during the fusion process, and lack the ability to model time dependence. In addition, existing methods have poor adaptability to large-scale unlabeled data and cannot meet the requirements of real-time and accuracy in public place monitoring.
[0190] In this scenario, the present invention realizes multi-modal perception and video understanding tasks through the following steps: First, collect the video stream of the surveillance camera, the audio data of the microphone, and the log text generated by the system. Then, use the preprocessing module to clean and structure these data into a unified multi-modal input format. Next, respectively extract the visual features of the video, the temporal features of the audio, and the semantic features of the text through the multi-modal feature extraction module, and use the cross-modal attention mechanism of the S1iME framework to align and fuse these features to generate a unified multi-modal representation.
[0191] During the occurrence of an event, for example, at 10:30 in the morning on a certain day, the surveillance camera captured a crowd gathering at the mall entrance, and at the same time, noisy sounds were detected in the audio data, and the text log recorded the description of "urgently dispatch personnel for evacuation". Through the dynamic memory module, the system records the current visual abnormal features, audio changes, and text information in real time, generating a situation description of "a crowd congestion occurred at the mall entrance, and there may be a dispute event". Subsequently, the system predicts that the event may further escalate into a conflict and sends an alarm signal to the management personnel.
[0192] Through experiments, we compared with traditional methods and evaluated the accuracy of abnormal event detection and the precision of sentiment analysis. The experimental data came from 30 days of surveillance records in the shopping mall, 12 hours a day, collecting approximately 500,000 video frames, 240,000 audio segments, and 10,000 text logs. To verify the effect of the present invention, we divided the data into a training set and a test set, accounting for 80% and 20% respectively. The comparative experiment shows that the present invention is significantly superior to traditional methods.
[0193] Table 1 Comparison of abnormal event detection accuracy
[0194]
[0195] Table 2 Comparison of Sentiment Analysis Accuracy
[0196]
[0197] As can be seen from Table 1, the method of the present invention based on the S1iME framework performs far better than traditional methods in the abnormal event detection task. In the processing of unimodal data, the accuracy rates of the present invention in the video modality, audio modality, and text modality reach 92.3%, 89.7%, and 93.5% respectively, all significantly higher than the performances of traditional methods A and B.
[0198] Table 2 shows the accuracy comparison of different methods in the sentiment analysis task. The accuracy rates of the method of the present invention in positive sentiment, negative sentiment, and neutral sentiment classification reach 94.2%, 91.8%, and 90.3% respectively, and the comprehensive average accuracy is 92.1%. In contrast, the average accuracy of traditional method A is only 81.5%, while the average accuracy of traditional method B is 85.1%. The present invention is significantly superior to traditional methods in all sentiment classification tasks, especially in the negative sentiment classification, with a performance 8.6 percentage points higher than that of traditional method B.
[0199] As can be seen from the tabular data, the S1iME framework proposed by the present invention shows significant advantages in the multimodal perception task. In abnormal event detection, through cross-modal alignment and dynamic interaction modeling, a multimodal fusion accuracy rate of 96.1% is achieved, far higher than 85.4% and 88.3% of traditional methods. At the same time, in the sentiment analysis task, the average accuracy of the present invention reaches 92.1%, an increase of 7 percentage points compared to traditional method B, and it performs particularly outstanding in the negative sentiment classification.
[0200] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A multimodal perception and video understanding method based on the S1iME framework, characterized in that: The steps include: S1, receiving video, audio and text data, preprocessing the data, and generating structured multimodal data; S2, extract features from structured multimodal data and construct a preliminary multimodal feature representation; S3. Based on the preliminary multimodal feature representation, the cross-modal self-attention mechanism in the S1iME framework is used to align and fuse features, and multimodal feature representation is generated through dynamic interaction modeling; S4, input the multimodal feature representation into the short-term memory module in the S1iME framework, and combine it with the long-term memory module to dynamically update the historical information across time periods to generate a time-enhanced multimodal feature representation; S5, based on time-enhanced multimodal feature representation, the reasoning module in the S1iME framework is used to perform temporal modeling and generate reasoning feature representation; S6, input the reasoning feature representation into the generation module in the S1iME framework, and generate a multimodal understanding result by combining the current multimodal feature representation and the reasoning feature representation; S7, optimize the S1iME framework through self-supervised learning mechanism, and use the multimodal understanding results and the correlation information between multimodal data fragments to perform deviation comparison and feature correction; S8. Based on the optimized S1iME framework, multimodal feature representation and reasoning feature representation are combined to output the results of video understanding.
2. The multimodal perception and video understanding method based on the S1iME framework according to claim 1 is characterized in that: The S3 specifically includes: S31, for the preliminary multimodal feature representation, linearly transform the features of the video modality, the audio modality, and the text modality respectively to generate a query matrix Q, a key matrix K, and a value matrix V; S32. Based on the query matrix Q and the key matrix K, the interaction weights between the video modality, the audio modality, and the text modality are calculated, and the self-attention mechanism is adopted: Among them, α ij is the interaction weight of the i-th modal feature to the j-th modal feature, d is the dimension of the feature vector, n is the total number of modal features, exp is the exponential function, K j is the key vector of the jth modal feature, K k is the key vector of the kth modal feature, Q i is the query vector of the i-th modal feature; S33, based on modal interaction weight α ij Sum value vector V j , perform weighted summation and nonlinear transformation on the features between modalities to generate a fused multimodal feature representation: Among them, F i is the fused feature representation of the i-th modality, W f and b f is the learnable fusion parameter matrix and bias vector, σ is the nonlinear activation function, n is the total number of modal features, and j is the index used to traverse each modal feature and calculate it in the weighted sum; S34, based on the fused multimodal feature representation F i , dynamically adjust the weights of the features of different modalities and generate optimized weighted feature representations: Among them, the dynamic weight w is calculated i , introducing the mutual information between modes as an additional constraint: Among them, the modal weight w i Acting on the fused feature representation F i , generate the optimized weighted feature representation F' i : F' i =w i ·F i +γ·W r ·F i ; Among them, ε is a nonlinear activation function, W a and b a is the learnable parameter matrix and bias, λ is the mutual information weight factor, H(F i ,F j ) is the mutual information between the features of mode i and mode j, γ is the residual coefficient, W r is the learnable parameter matrix; S35, the optimized weighted feature representation F' i Conduct deep interaction and context modeling to capture global correlation information and temporal dependencies between modalities, perform context enhancement on multimodal features, and generate context enhanced feature representation F” i ; Among them, LayerNorm is the layer normalization operation, Concat is to concatenate the outputs of h attention heads into a complete context representation, is the normalization of attention weights, is the scaling factor, W o is the output transformation matrix of multi-head attention, Q t ,K t ,V t are query vector, key vector and value vector respectively, H t-1 is the hidden state of the previous time step, W h is the linear transformation matrix of time-dependent features; S36, representing the context enhanced feature F" i Mapped to a unified semantic space, a multimodal feature representation U is generated through semantic alignment, normalization and feature fusion. t : Among them, μ is a nonlinear activation function, W u is the mapping matrix of the unified semantic space, b u is the bias term, U i is the unified representation of the i-th mode, w i is the fusion weight of the i-th modality, and n is the total number of modality features.
3. The multimodal perception and video understanding method based on the S1iME framework according to claim 1 is characterized in that: The S4 specifically includes: S41, the generated multimodal representation F" i Enter the short-term memory module to extract the key information of the current time segment: M s =τ(W s ·F” i +b s ); Among them, M s is the current time segment feature output by the short-term memory module, W s and b s are the learnable weight matrix and bias of short-term memory, respectively, and τ is the nonlinear activation function; S42, historical state and current time segment features M stored in long-term memory module s , dynamically update the long-term memory state: Among them, H t-1 is the long-term memory state of the previous time step, To calculate the weighted strength of short-term memory based on the attention mechanism, is the scaling factor, W a is the modality attention weight matrix, GRU is the gated recurrent unit, H t is the updated long-term memory state; S43, the short-term memory feature M s and long-term memory state H t Perform weighted fusion to generate dynamic multimodal feature representations that include time dependencies: F d =a s ·M s +a h ·H t +γ·(M s -H t ); Among them, F d is the dynamic multimodal feature representation, α s and α h denote the dynamic weights of short-term memory and long-term memory respectively, and γ is the residual coefficient; S44, dynamic feature representation F d Perform modal weight optimization and normalization to generate time-enhanced multimodal feature representation: Among them, F n is a time-enhanced multimodal feature representation, LayerNorm is used to normalize the fused features, and w i is the weight of the i-th mode, σ is the nonlinear activation function, is the linear transformation matrix of the ith mode, is the bias term of the i-th mode, and n represents the total number of modes.
4. The multimodal perception and video understanding method based on the S1iME framework according to claim 1 is characterized in that: The S5 specifically includes: S51, the generated time-enhanced multimodal feature representation F n Input into the inference module as the initial feature representation for time series modeling; S52. Time-based enhanced multimodal feature representation F n , using gated recurrent units for timing modeling, dynamically capturing temporal dependencies, and balancing the contribution of the current time step features and historical states through a modal time weight optimization mechanism: H t =γ·GRU(H t-1 ,F n )+(1-γ)·F n ; Among them, γ is the dynamic time weight, GRU is the gated recurrent unit, H t-1 is the hidden state of the previous time step, H t is the hidden state of the current time step; S53, based on the hidden state H of the current time step t and time-enhanced multimodal feature representation F n Perform dynamic weight allocation and combine cross-modal relationships to generate temporal dynamic multimodal fusion features: Among them, M t is the multimodal fusion feature of the current time step, To calculate the attention weight of the i-th modality, is the scaling factor, is the modality-specific weight matrix, To optimize the matrix, is the bias vector of the i-th mode, and n represents the total number of modes; S54, time dynamic multimodal fusion feature M t Perform residual update and normalization operations, and the time-enhanced multimodal feature representation F n , generate optimized multimodal feature representation F u : F u =LayerNorm(M t +F n ·W r +b r ); Among them, W r is a learnable transformation matrix, LayerNorm is a layer normalization operation, and b r is the bias vector; S55. Multimodal feature representation based on optimization F u , through the reasoning module to capture the temporal association and semantic relationship between features, and generate the reasoning feature representation R t : R t =LayerNorm(σ(F u ·W p +b p )+H t ); Among them, σ is a nonlinear activation function, W p is the weight matrix, b p is the bias term.
5. The multimodal perception and video understanding method based on the S1iME framework according to claim 1 is characterized in that: The S6 specifically includes: S61, the generated inference feature is represented by R t Input generation module as the initial feature representation for the multimodal generation task; S62. Input-based reasoning feature representation R t , and fuse them through the cross-modal generation unit to generate the multi-modal fusion feature G t : in, To normalize the output of the inference features after the modality-specific weights are calculated, is the scaling factor, is the weight matrix of the i-th mode, is the feature optimization matrix of the i-th mode, is the bias term of the i-th mode, σ is the nonlinear activation function, W h is the global transformation matrix, n represents the total number of modes; S63, multimodal fusion feature G t Perform semantic mapping for specific tasks and generate semantic description feature representation S t : S t =Decoder(σ(G t ·W u +b s )); Among them, Decoder is the decoding function in the generation module, b s is the bias term for semantic generation, W u Generate weight matrices for task-specific semantics; S64, based on multimodal fusion feature G t , generate event prediction feature representation E through time series modeling t : AND t =GRU(E t-1 ,G t ); Among them, E t The event prediction feature of the current time step, GRU is a gated recurrent unit for time series modeling, E t-1 Predict the hidden state for the event at the previous time step; S65, multimodal fusion feature G t Mapped to the sentiment space, the sentiment analysis feature representation is generated through dynamic weight allocation, modal feature optimization and global feature enhancement: Among them, α i is the dynamic sentiment weight of the i-th modality, Optimize the weight matrix for the features of the i-th mode, is the bias term of the ith mode, W c Optimize the matrix for global sentiment; S66, the generated semantic description feature is represented as S t , event prediction feature representation E t and sentiment analysis feature representation A t Output, as the multimodal understanding result O t .
6. The multimodal perception and video understanding method based on the S1iME framework according to claim 1 is characterized in that: The S7 specifically includes: S71. Based on multimodal understanding results O t and the inference feature representation R t , generate self-supervised learning signals and build associations between multimodal segments: Among them, L context is the context relevance loss, T is the time step, Sim is the similarity function, and t is the time step; S72. Based on positive and negative sample contrast learning, the robustness of multimodal understanding results is optimized. Positive samples are real understanding results, and negative samples are generated from interference fragments: Among them, L contrast is the loss function of contrastive learning, ln is the logarithmic function, exp is the exponential function, is the positive sample of the current time step, is the jth candidate sample, and N is the total number of candidate samples; S73, based on the inference feature representation R t And multimodal fusion feature G t , an updated inference feature representation is generated through a dynamic correction mechanism: Among them, LayerNorm is the layer normalization operation, W a is the weight matrix, b k is the bias term, is the updated inference feature representation; S74, the updated inference feature is represented With the optimized multimodal feature representation F u Combined with capturing contextual information through time series modeling, generating context-enhanced features Among them, GRU is a gated recurrent unit, H t-1 is the hidden state of the previous time step; S75. Based on the context-enhanced features, the global optimization objective is constructed by combining the context-relevance loss and the contrast loss, and the model parameters are optimized through the self-supervised learning mechanism.
Citation Information
Cited By
Logic behavior detection method and device, electronic equipment and computer readable storage medium
CN120725024A
Physical constraint-based multi-mode indoor carbon dioxide concentration prediction method
CN121580357A