Multi-modal video target identification method based on cross-modal spatio-temporal joint learning
By splicing features on the Patch quantity dimension and designing a directional attention mechanism, the problems of insufficient correlation between modes and high computational complexity in multimodal video target recognition are solved, and more efficient feature fusion and recognition accuracy are achieved, especially in scenarios with high real-time requirements.
Patent Information
- Application Number
- CN202510679690.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing multimodal video object recognition method fails to fully explore the spatial and temporal dependence between modes, low feature fusion efficiency, high computational complexity, and difficult to meet the real-time application requirements.
The cross-modal space-time joint learning method is adopted to splice features on the Patch quantity dimension and design a directional attention mechanism to limit the scope of attention calculation, and enhance the space-time dependence and feature fusion efficiency between modes.
It significantly improves the accuracy and real-time nature of multimodal video target recognition, improves feature fusion efficiency by 5% to 10%, and reduces the computational complexity by nearly 3 times. It is suitable for intelligent monitoring and medical image analysis.
Smart Images

Figure CN120236233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal data processing, and relates to a multimodal video object recognition method based on cross-modal spatio-temporal joint learning. Background Art
[0002] Multimodal Video Object Recognition is a video analysis technology that combines various perceptual modal data (such as vision, audio, depth information, radar signals, etc.), aiming to improve the accuracy and robustness of object detection and recognition in complex scenarios by fusing complementary information from different modalities. Its core lies in leveraging the synergistic effect of multimodal data to address the limitations of a single modality in scenarios such as illumination changes, occlusion, or noise interference.
[0003] With the progress of multimodal sensor technology, Multimodal Video Object Recognition has shown important value in fields such as intelligent monitoring, medical image analysis, and autonomous driving. Video data of different modalities often have complementary characteristics. For example, due to its relatively low resolution, depth video has limited ability to capture spatial texture details but shows high sensitivity to the motion trajectory or temporal changes of objects. In contrast, infrared video provides rich spatial information through heat distribution but is difficult to comprehensively reflect dynamic features in scenarios where object motion is not significant. In the field of medical imaging, grayscale ultrasound video is good at presenting the anatomical structure and texture features of objects, while contrast-enhanced ultrasound video reveals hemodynamic information through the temporal changes in the intensity of contrast agents. Since a single modality is difficult to fully characterize object properties, multimodal fusion has become a key approach to improving recognition performance. In the prior art, Multimodal Video Object Recognition is usually achieved in two ways. One way is to extract the features of each modality separately, then fuse them through weighted averaging or feature concatenation, and then input them into a classifier for recognition. Another way is to concatenate multimodal videos along the channel dimension at the input stage, and then extract features through a unified convolutional neural network or Transformer.
[0004] However, these methods have significant deficiencies in practical applications. First, they fail to fully exploit the spatio-temporal dependencies between modalities. For example, the extraction of temporal features of modality A may require the spatial structure of modality B as a guide, while the extraction of spatial features of modality B may depend on the temporal context of modality A. Second, the existing feature fusion methods are often inefficient. Simple concatenation or late fusion easily leads to information redundancy or loss of key features. In addition, when dealing with multi-modal data, the self-attention mechanism based on traditional 3D Transformer usually calculates attention without discrimination for all features, which not only ignores the modality characteristics but also brings a high computational complexity, limiting the feasibility of real-time applications. For example, in invention CN119399670A, a bad video classification method and system based on multi-modal alignment and class balance, the flowchart is as Figure 2 shown. Therefore, the existing technologies fail to accurately model the spatio-temporal dependencies between modalities, with low fusion efficiency and serious information redundancy or loss.
[0005] Therefore, there is a need for a multi-modal video object recognition method that can accurately perform spatio-temporal interaction and reduce computational complexity to solve the above technical problems. Summary of the Invention
[0006] The present invention aims to provide a multi-modal video object recognition method based on cross-modal spatio-temporal joint learning to solve the problems of insufficient correlation between modalities and low feature fusion efficiency in the prior art. The core of the present invention is to propose a novel cross-modal self-attention mechanism. Traditional feature fusion mostly uses concatenation in the channel dimension, which easily leads to premature mixing of modality information. In contrast, the present invention concatenates features in the Patch number dimension, thus preserving the independence between modalities and providing greater flexibility for subsequent cross-modal interaction. In addition, aiming at the modality characteristics, the present invention designs a directional attention mechanism, such that the extraction of temporal features of modality A is accurately guided by the spatial information of modality B, while the extraction of spatial features of modality B is effectively constrained by the temporal information of modality A, thereby changing the calculation method of traditional full-dimensional attention. By restricting the scope of attention calculation, not only the ability to model spatio-temporal dependencies between modalities is enhanced, but also the computational complexity is significantly reduced. Thus, it jointly contributes to more efficient feature fusion and more accurate object recognition in multi-modal video.
[0007] The technical solution adopted by the present invention to solve the technical problems is: a multi-modal video object recognition method based on cross-modal spatio-temporal joint learning, including the following steps:
[0008] Step 1, data preprocessing; Denote the modality videos as modality A and modality B respectively, perform preprocessing on the two modality videos, uniformly adjust the spatial resolution of each frame of the modality videos, and resample the time dimension to frames; Enhance data stability through normalization operations;
[0009] Step 2, Visual feature extraction: Based on the feature extraction network of 2D convolutional neural network, extract low-order visual features from each frame, perform convolutional operations on each frame independently, and extract feature maps;
[0010] Step 3, Cross-modal spatio-temporal joint learning network: Construct a cross-temporal joint learning network based on 3D Transformer to extract high-order features and achieve cross-modal fusion;
[0011] Step 4, Feature fusion and target recognition: Obtain fused features through multi-layer cross-modal self-attention calculation; compress the fused features into feature vectors through global pooling operation, and then output classification results through fully connected layers and Softmax function.
[0012] Preferably, the specific steps of Step 3 are as follows;
[0013] Step 3-1, Concatenate features in the dimension of the number of Patches;
[0014] Step 3-2, Concatenate in the dimension of the number of Patches to form a joint Token sequence;
[0015] Step 3-3, Design a cross-modal directional attention mechanism, propose an attention interaction rule, for the th frame and the th Patch of modality A, the attention of the Query only interacts with the spatial Key and Value of the th frame of modality B, and for the th frame and the th Patch of modality B, the attention of the Query only interacts with the temporal Key and Value of the th frame of modality A;
[0016] Step 3-4, Update features through residual connection and layer normalization.
[0017] More preferably, in Step 1, when uniformly adjusting the spatial resolution of each frame, the resolution is uniformly adjusted to ; the normalization operation normalizes the pixel values to the interval [0, 1].
[0018] More preferably, in Step 2, when performing convolutional operations on each frame independently, the downsampled spatial resolution is .
[0019] More preferably, in Step 3, the feature map is divided into Patches with a size of .
[0020] More preferably, in Step 1, the input video of modality A is , and the input video of modality B is , where is the number of frames, is the spatial resolution, and is the number of channels;
[0021] In step 2, the feature map is:
[0022]
[0023] where is the feature dimension;
[0024] In step 3-1, the feature maps and are divided into Patches, and the Patches are flattened into a Token sequence to obtain:
[0025]
[0026] In step 3-2, the combined Token sequence is: ;
[0027] In step 3-3, in the attention interaction rule, the interaction between modality A attention and modality B is:
[0028]
[0029] where , , , , are learnable projection matrices respectively; is a D-dimensional vector composed of all D elements in the t-th row and i-th column of X_A, is the t-th row of the tensor X_B, that is, a 2D matrix composed of elements with coordinates [t,:,:];
[0030] The interaction between modality B attention and modality A is:
[0031]
[0032] where , , , , is a vector composed of all elements in the t-th row and j-th column of the tensor X_B, that is, a vector composed of all elements with coordinates [t,j,:], is a matrix composed of elements in the j-th column of all rows in the tensor X_A;
[0033] In step 3-4, the updated feature is obtained through residual connection and layer normalization as:
[0034]
[0035] In step 4, the fused feature is , and the feature vector is , and the classification result output by the Softmax function is:
[0036]
[0037] where, W fc represents the weight matrix of the fully connected layer, and b represents the bias vector.
[0038] Preferably, in step 2, the 2D convolutional neural network has 3 layers of convolution, ReLU activation, max pooling, and the convolutional kernel , and the pooling window .
[0039] Preferably, in step 3, the joint Token sequence is input into 4 layers of Transformer with 8 attention heads.
[0040] The beneficial effects of the present invention are:
[0041] 1. The present invention splices features in the dimension of the number of patches, thereby retaining the independence between modalities and providing greater flexibility for subsequent cross-modal interactions.
[0042] 2. The present invention designs a directional attention mechanism according to the modality characteristics, so that the extraction of the temporal features of modality A is accurately guided by the spatial information of modality B, while the extraction of the spatial features of modality B is effectively constrained by the temporal information of modality A, changing the calculation method of the traditional full-dimensional attention; by restricting the range of attention calculation, the present invention not only enhances the spatio-temporal dependence modeling ability between modalities, but also significantly reduces the computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of a multi-modal video object recognition method based on cross-modal spatio-temporal joint learning according to the present invention;
[0044] Figure 2 is a flowchart of the prior art. DETAILED DESCRIPTION OF THE INVENTION
[0045] The following will clearly and completely describe the related technologies in the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] Reference Figure 1 , the technical solution in this embodiment is implemented through the following steps. Each step includes a detailed calculation process and function description, which will be elaborated one by one below.
[0047] Step 1, Data preprocessing:
[0048] To ensure that the input video data is suitable for subsequent feature extraction, the present invention first preprocesses two modalities of videos (denoted as modality A and modality B, such as depth video and infrared video). Let the input video of modality A be , and modality B be , where is the number of frames, is the spatial resolution, and are the number of channels. The preprocessing includes uniformly adjusting the spatial resolution of each frame to , and resampling the time dimension to frames to adapt to the input requirements of the depth convolutional neural network. At the same time, data stability is enhanced through a normalization operation (such as normalizing pixel values to the [0,1] interval). The resolution adjustment uses bilinear interpolation to ensure the retention of details while avoiding significant distortion. Compared with the prior art where directly using the original resolution may lead to inconsistent model inputs or waste of computing resources, the preprocessing step of the present invention standardizes the input format, improves the robustness and efficiency of feature extraction, and lays a solid foundation for subsequent cross-modal fusion.
[0049] Step 2, Visual feature extraction:
[0050] After preprocessing, the present invention uses a feature extraction network based on a 2D convolutional neural network (CNN) to extract low-order visual features from each frame. Modality A usually has a strong ability to represent temporal information. For example, a depth video can capture the movement trajectory of an object; modality B provides significant spatial information. For example, an infrared video reflects the thermal distribution characteristics of an object. Through 2D CNN, convolution operations are performed independently on each frame to extract feature maps:
[0051]
[0052] where is the downsampled spatial resolution, is the feature dimension (usually set to 256). Compared with the prior art method of directly splicing raw video data, the present invention extracts features through preprocessing and an independent CNN, avoiding early confusion of modality information and retaining their respective spatio-temporal characteristics. This method solves the problem of dilution of key features caused by early fusion in traditional methods, provides high-quality input features for subsequent cross-modal interaction, and thus lays a foundation for efficient fusion.
[0053] Step 3, Cross-modal Spatiotemporal Joint Learning Network:
[0054] After extracting the preliminary features, the present invention constructs a cross-temporal and spatial joint learning network based on 3D Transformer to extract high-order features and achieve cross-modal fusion. Traditional methods usually concatenate features along the channel dimension or integrate independently extracted features through late fusion. However, these methods either lead to premature mixing of modal information or fail to fully exploit the spatiotemporal dependencies between modalities, resulting in insufficient representational power of the fused features. The present invention breaks through these limitations through two major innovations, significantly improving the fusion effect.
[0055] First, the present invention abandons the traditional practice of concatenating features along the channel dimension and instead concatenates features in the dimension of the number of patches. Traditional channel concatenation confuses modal information in the early fusion stage, making it difficult to optimize according to modal characteristics, resulting in information redundancy or loss. The present invention divides the feature maps and into patches of size (i.e., each 1×1 area is used as a patch), generating patches, which are flattened into a token sequence to obtain:
[0056]
[0057] Subsequently, a joint token sequence is formed by concatenating in the dimension of the number of patches. This design preserves the independence between modalities, avoids information confusion, and creates conditions for subsequent cross-modal attention calculation. Compared with the feature dilution caused by channel concatenation in the prior art, the patch concatenation method of the present invention ensures the integrity of modal characteristics before fusion, enabling subsequent interactions to more precisely utilize complementary information, thereby enhancing the pertinence and representational power of the fused features.
[0058] Second, the present invention designs a cross-modal directional attention mechanism, breaking through the limitations of the full-dimensional attention of traditional 3D Transformer. Traditional self-attention calculates all tokens without discrimination, ignoring modal characteristics. For example, the temporal information of depth videos and the spatial information of infrared videos cannot be effectively coordinated, and the computational complexity is high. The present invention proposes a unique attention interaction rule according to the characteristics of rich temporal information in modality A and rich spatial information in modality B. For the token (the query of the frame and the th patch in modality A), its attention only interacts with the spatial key and value of the th frame in modality B:
[0059]
[0060] Among them , , , , is a learnable projection matrix. Similarly, the tokens of modality B only interact with the temporal Key and Value at the corresponding positions of modality A:
[0061]
[0062] Among them , , , . The updated features are obtained through residual connection and layer normalization:
[0063]
[0064] This mechanism enables the extraction of temporal features of modality A to be guided by the spatial information of modality B, and the extraction of spatial features of modality B to be constrained by the temporal information of modality A, accurately modeling the spatio-temporal dependence. Compared with traditional full-dimensional attention, the present invention reduces the computational complexity from to , and improves the feature representation ability through targeted interaction, solving the problem of poor fusion effect caused by traditional methods ignoring modality characteristics.
[0065] Step 4, Feature fusion and target recognition:
[0066] After multi-layer cross-modal self-attention calculation, the fused feature is obtained. It is compressed into a feature vector through a global pooling operation (such as average pooling), and then the classification result is output through a fully connected layer and a Softmax function:
[0067]
[0068] Compared with the traditional late fusion method that is prone to information redundancy, the layer-by-layer cross-modal interaction and global pooling of the present invention retain the key features and ensure high-precision classification. This method makes full use of the accurately fused features in step 3 and is significantly better than the rough integration of traditional methods.
[0069] The advantages of the present invention lie in solving the problems of insufficient inter-modal correlation and high computational complexity through data preprocessing, Patch quantity dimension splicing, and cross-modal directional attention mechanism. Experiments show that the recognition accuracy of this method on the multi-modal video dataset has increased by 5% - 10%, and the real-time performance has been improved by nearly 3 times, providing an efficient and accurate solution for intelligent monitoring and medical image analysis.
[0070] Embodiment
[0071] Taking the recognition of pedestrians and vehicles in the night intelligent monitoring scenario as an example, this embodiment elaborates in detail how to achieve the fusion of depth video and infrared video through Steps 1 to 4, highlighting the differences from the prior art and the objective improvement effects.
[0072] Application scenario and input data:
[0073] Under low-light conditions at night, the monitoring system needs to recognize pedestrians and vehicles. The input is depth video and infrared video , the number of frames , the resolution . The depth video captures the movement trajectory (such as the steps of pedestrians), but the spatial details are weak; the infrared video presents the heat distribution (such as the heat of the vehicle engine), but the dynamic information is insufficient. Traditional methods such as channel splicing or post-fusion cannot fully utilize the complementary characteristics, and the recognition accuracy and real-time performance are limited. This embodiment accurately fuses the modal characteristics through cross-modal spatio-temporal joint learning, significantly improving the performance.
[0074] Step 1, Data preprocessing:
[0075] Preprocess the input videos and , adjust the resolution of each frame from to , and use bilinear interpolation to retain details. Subsequently, normalize the pixel values to the [0,1] interval to enhance data stability. Compared with the prior art where the direct use of non-standard resolution leads to inconsistent model inputs, the preprocessing in this embodiment ensures a unified input format, improving the robustness and computational efficiency of feature extraction.
[0076] Step 2, Visual feature extraction:
[0077] Use two groups of 2D CNNs (3-layer convolution, ReLU activation, max pooling, convolution kernel , pooling window ) to process the preprocessed and , and output the feature maps Compared with directly splicing the original data in the traditional way, in this embodiment, the modal characteristics are retained through preprocessing and independent extraction, avoiding information confusion and providing high-quality features for efficient fusion.
[0078] Step 3: Cross-modal spatio-temporal joint learning:
[0079] Divide and into Patches, generating Patches, obtaining the Token sequence and concatenating them in the dimension of the number of Patches to form . This design is superior to traditional channel concatenation, avoiding information confusion. The joint Token sequence is input into a 4-layer Transformer (8 attention heads). The cross-modal directional attention mechanism enables the temporal features of the depth video (such as the pedestrian's steps) to be guided by the spatial information of the infrared video (such as the body contour), and the spatial features of the infrared video are constrained by the temporal information of the depth video. The computational complexity is reduced from to , and the fused feature is more representative. Compared with traditional full-dimensional attention, this embodiment accurately models spatio-temporal dependencies and solves the problem of ignoring modal characteristics.
[0080] Step 4: Feature fusion and target recognition:
[0081] Perform global average pooling on to generate , and output the probabilities of "pedestrian" or "vehicle" through a fully connected layer. Compared with the information redundancy of traditional late fusion, the layer-by-layer interaction in this embodiment retains key features and improves the classification accuracy.
[0082] This embodiment proposes a spatio-temporal feature interaction method different from existing methods, which reduces the computational complexity while enhancing the mutual guidance between temporal features and spatial features, fully utilizes the information characteristics of each modality, and realizes more efficient and accurate multi-modal video target recognition. Therefore, the cross-modal spatio-temporal correlation learning method is the core of this embodiment.
[0083] In this embodiment, by concatenating features in the dimension of the number of Patches, the independence between modalities is retained, avoiding information confusion caused by traditional channel concatenation, and providing flexibility and efficiency for cross-modal attention calculation.
[0084] The cross-modal directional attention mechanism of this embodiment realizes precise spatio-temporal information interaction for modal characteristics by restricting the attention calculation range, significantly improving the pertinence and representativeness of feature fusion.
[0085] In specific implementations, the depth video captures the motion trajectory of the target, such as the walking path of a pedestrian in a surveillance scenario, while the infrared video provides the thermal distribution information of the target, such as the heat profile of a human body or a vehicle. Traditional methods (such as channel splicing) are prone to losing the temporal details of the depth video or the spatial texture of the infrared video during fusion. However, the Patch dimension splicing of the present invention preserves the modal independence, and the cross-modal attention mechanism further enhances the feature complementarity through directional interaction. For example, in a night surveillance scenario, the depth video may be difficult to capture spatial details due to insufficient light, but its temporal information can reflect the target motion; the infrared video provides a clear thermal distribution but lacks dynamic information. The attention mechanism of the present invention enables the temporal features of the depth video to be guided by the spatial heat map of the infrared video to form a more complete motion trajectory representation. At the same time, the spatial features of the infrared video are constrained by the temporal information of the depth video to optimize the boundary detection of static targets. Experimental results show that on a public multi-modal dataset (such as FLIR ADAS), the recognition accuracy of this method reaches 92.5%, which is about 7.5% higher than that of the traditional channel splicing method (about 85%) and about 5.5% higher than that of the late fusion method (about 87%). In addition, due to the limited attention range, the computing time is reduced from about 0.8 seconds per sample of the traditional method to about 0.5 seconds per sample, and the real-time performance is improved by about 37.5%. These improvements verify the significant advantages of the present invention in feature fusion efficiency and target recognition accuracy.
[0086] In summary, through the Patch number dimension splicing and the cross-modal directional attention mechanism, the present invention not only preserves the modal independence but also realizes precise spatio-temporal interaction and reduces the computational complexity. The present invention improves in feature fusion efficiency and target recognition accuracy, especially performing well in scenarios with high real-time requirements.
[0087] It should be emphasized that the above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Any simple modification made to the above embodiments based on the technical essence of the present invention also belongs to the protection scope of the present invention. Other equivalent changes and modifications are still within the scope of the technical solution of the present invention.
Claims
1. A multi-modal video object recognition method based on cross-modal spatio-temporal joint learning, characterized in that It includes the following steps: Step 1, data preprocessing; Denote the modality videos as modality A and modality B respectively, preprocess the two modality videos, uniformly adjust the spatial resolution of each frame of the modality videos, and resample the time dimension to frames; Enhance data stability through normalization operations; Step 2, Visual feature extraction: Based on the feature extraction network of 2D convolutional neural network, extract low-order visual features from each frame, perform convolutional operations on each frame independently, and extract feature maps; Step 3, Cross-modal spatio-temporal joint learning network: Construct a cross-spatio-temporal joint learning network based on 3D Transformer to extract high-order features and achieve cross-modal fusion; Step 4, Feature fusion and target recognition: Obtain fused features through multi-layer cross-modal self-attention calculation; Compress the fused features into feature vectors through global pooling operation, and then output the classification results through the fully connected layer and Softmax function.
2. The multimodal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 1, wherein, The specific steps of Step 3 include the following sub-steps; Step 3-1, Concatenate features in the dimension of the number of patches; Step 3-2, Concatenate in the dimension of the number of patches to form a joint token sequence; Step 3-3: Design a cross-modal directional attention mechanism and propose an attention interaction rule. For the Query of the th frame and the th Patch of modality A, its attention only interacts with the spatial Key and Value of the th frame of modality B. For the Query of the th frame and the th Patch of modality B, its attention only interacts with the temporal Key and Value of the th frame of modality A; Step 3-4, Update features through residual connection and layer normalization.
3. A multimodal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, characterized in that, In step 1, when the spatial resolution of each frame is uniformly adjusted, the resolution is uniformly adjusted to ; The normalization operation normalizes the pixel values to the range of [0, 1].
4. A multimodal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, characterized in that In step 2, when performing convolution operations on each frame independently, the spatial resolution after downsampling is .
5. The multimodal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, wherein In step 3, the feature map is divided into patches of size .
6. The multimodal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, wherein, In the said step 1, the input video of modality A is , and the input video of modality B is , where is the number of frames, is the spatial resolution, and are the number of channels; In Step 2, the feature map is: Among them, is the feature dimension; In step 3-1, the feature maps and are divided into Patches, and the Patches are flattened into a Token sequence, obtaining: In the said step 3-2, the combined Token sequence is: ; In Step 3-3, in the attention interaction rule, the interaction between modality A attention and modality B is: Among them, , , , , are learnable projection matrices respectively; is a D-dimensional vector composed of all D elements in the i-th column and t-th row of X_A, is the t-th row of the tensor X_B, that is, a 2D matrix composed of elements with coordinates [t, :, :]; The interaction between modality B attention and modality A is: Among them, , , , , is the vector composed of all elements in the \(t\)-th row and \(j\)-th column of the tensor \(X_B\), that is, the vector composed of all elements with coordinates \([t, j, :]\). is the matrix composed of the elements in the \(j\)-th column of all rows in the tensor \(X_A\). In Step 3-4, the update of features through residual connection and layer normalization is: In the said step 4, the fusion feature is , and the feature vector is . The classification result output by the Softmax function is: Among them, W fc represents the fully connected layer weight matrix, and b represents the bias vector.
7. A multi-modal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, characterized in that In the said step 2, the 2D convolutional neural network has 3 layers of convolution, ReLU activation, max pooling, and the convolutional kernel , and the pooling window .
8. A multi-modal video object recognition method based on cross-modal spatio-temporal joint learning according to claim 2, characterized in that In Step 3, the joint token sequence is input into 4 layers of Transformer with 8 attention heads.
Citation Information
Patent Citations
Video face emotion recognition method based on frame attention mechanism
CN115393933A
Driver intention prediction method based on multi-dimensional cross-modal information interaction
CN116110018A
Video emotion recognition method based on circular interaction Transform and dimension cross fusion
CN116168324A
Video behavior recognition method based on cross-modal fusion
CN116311525A
Token recombination model-based shallow-deep feature fusion method and system
CN116958765A
Cited By
Human motion posture recognition method and system based on multi-modal data fusion
CN120804842A