A multi-modal video-oriented full-process action recognition method

By combining a temporal shift module and a 2D convolutional neural network, the multimodal data processing flow is optimized, solving the problems of overfitting and computational resource consumption in multimodal action recognition, and achieving efficient and accurate action recognition and improved generalization ability.

CN120032424BActive Publication Date: 2025-11-11NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510074667.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-11-11
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing multimodal action recognition technologies suffer from problems such as overfitting due to the scarcity of multimodal data, huge consumption of computational resources, and low information fusion efficiency when facing real-world problems. In particular, when the number of samples is limited, the model's generalization ability is poor.

Method used

We employ a Temporal Shift Module (TSM) combined with a 2D Convolutional Neural Network (CNN) for multimodal data processing. We optimize the data processing flow through techniques such as dynamic group temporal sampling, group normalization, group batch augmentation, multimodal fusion, pre-trained transfer learning, random weight averaging, and test-time augmentation.

Benefits of technology

It improves the effective mining and fusion of multimodal information, enhances the accuracy of action recognition and the generalization ability of the model, reduces computational costs, and adapts to large-scale practical application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032424B_ABST
    Figure CN120032424B_ABST
Patent Text Reader

Abstract

This invention discloses a full-process action recognition method for multimodal videos. First, it transforms and expands existing data by optimizing augmentation techniques for multimodal data to increase the training scale. The backbone network is pre-trained using a larger RGB dataset and adapted to new tasks through transfer learning. Second, it extracts multimodal spatial features using 2D CNNs and combines them with a temporal shift module to achieve multimodal spatial-temporal feature extraction comparable to 3D CNNs, while improving computational efficiency. Prediction augmentation methods are used to integrate knowledge from the same and different architectures across different training stages, thereby predicting actions from different perspectives and fully utilizing target information. This invention overcomes the problems of data scarcity and overfitting, improves spatiotemporal modeling capabilities, and effectively integrates multimodal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a full-process action recognition method for multimodal videos, and in particular proposes an efficient scheme that combines a temporal shift module (TSM) and spatiotemporal feature extraction, which is applicable to the fields of computer vision and deep learning. Background Technology

[0002] With the significant success of deep neural networks (DNNs) in various tasks, especially in image processing and pattern recognition, single-modal data processing is gradually becoming insufficient to meet the demands of complex real-world scenarios. To improve action recognition and scene understanding, multimodal learning has emerged as a crucial research area and a current research hotspot. By integrating data from multiple sensors (such as RGB images, depth images, and thermal infrared images), multimodal technology can provide more comprehensive information for tasks, thereby more accurately identifying complex dynamic scenes.

[0003] However, despite some progress in multimodal action recognition technology, existing technologies still face a series of challenges in addressing real-world problems. First, the scarcity of multimodal data poses a significant challenge to training deep learning models, especially when the number of samples is limited, leading to overfitting and poor generalization. Second, efficiently processing information from different modalities and effectively fusing features remains a pressing issue. Existing technologies largely rely on computationally intensive 3D convolutional neural networks (CNNs), which, while capable of handling spatiotemporal features, consume enormous computational resources, limiting their application on large-scale datasets. Furthermore, although data augmentation and transfer learning techniques have alleviated these problems to some extent, effective optimization strategies for multimodal data fusion and processing still lack depth. Summary of the Invention

[0004] Purpose of the invention: To address the shortcomings of existing action recognition and multimodal technologies, this invention proposes an efficient multimodal action recognition solution to improve the performance and generalization ability of action recognition systems when facing complex real-world scenarios.

[0005] Technical solution: A full-process action recognition method for multimodal videos, which can be divided into four stages: multimodal data preprocessing stage, multimodal fusion stage, training stage, and inference stage;

[0006] The multimodal data preprocessing includes the following steps:

[0007] Step 11, Dynamic Group Timing: When loading data from different modalities, the data for each modality is evenly divided into multiple groups, and then frames are randomly selected from each group to form a multimodal video frame sequence for training. This aims to improve training efficiency and reduce interference from a large number of redundant frames. The different modalities of data include RGB video data, infrared video data, and depth video data.

[0008] Step 12, group normalization; normalize the data for each modality.

[0009] The specific operation involves subtracting the mean from the data points of each modality and dividing by the variance of the corresponding data points. The purpose is to improve the training speed and stability of the model and reduce overfitting.

[0010] Step 13, Batch Enhancement (TR) Processing: The obtained multimodal video frames are subjected to various efficient batch enhancements, including group multi-scale cropping and group random horizontal flipping (the former randomly crops images at different scales to enhance data diversity and simulate visual effects under different perspectives and fields of view; the latter randomly flips images horizontally to enhance the symmetry of the data, forcing the model to learn more variations and robustness). The purpose is to effectively enhance the data, solve the problem of limited training sample quantity, expand the training dataset, improve the model's generalization ability, and reduce overfitting.

[0011] The multimodal fusion stage includes the following steps:

[0012] Step 21: Feature extraction of multimodal data; Select two multimodal action recognition models, use 2D convolutional neural networks (CNN) to extract multimodal spatial features, and combine them with the time translation module (TSM) of the multimodal action recognition model to achieve multimodal spatiotemporal feature extraction.

[0013] The TSM-Res50 and TSM-Res101 models are used as two multimodal action recognition models. Both TSM-Res50 and TSM-Res101 models are based on the TSM framework.

[0014] Step 22, Multimodal Fusion: The outputs (Logits) of different modalities obtained in Step 21 are weighted and fused to effectively utilize multimodal information. The specific formula is shown below:

[0015] Logits R Logits I Logits D =F θ (Cat(X R X I X D ));

[0016] ;

[0017] The RGB modal video data is denoted as X. R ∈R T×W×H×C Infrared video data is denoted as X. I ∈R T×W×H×C Depth video data is denoted as X D ∈R T×W×H×C T, W, H, and C represent the number of video frames, frame width, frame height, and number of video channels, respectively. θ This represents the model to be trained, parameterized by θ. The function Cat(·) represents the connection operation. β, γ, and α are weight coefficients. Logits R Logits I and Logits D Let J(·) represent the Logits for different modalities, J(·) represent the cross-entropy loss function, and Y be the actual label predicted from the multimodal video data.

[0018] The training phase includes the following steps:

[0019] Step 31, pre-training transfer learning; using pre-trained knowledge to enhance the performance of the multimodal action recognition model in downstream tasks;

[0020] Considering the limited training dataset, when performing large-scale pre-training, we leverage the pre-training knowledge from the Kinetics400 and ImageNet datasets to enhance the performance of the TSM-Res50 model on downstream tasks. For the same purpose, we use the pre-training knowledge from the Something-somethingV2 and ImageNet datasets to enhance the performance of the TSM-Res101 model on downstream tasks. The goal is to reduce training time on the target training set and improve the model's generalization ability.

[0021] Step 32, Model Training Phase: Perform hyperparameter configuration and gradient clipping.

[0022] During the training phases of the TSM-Res50 and TSM-Res101 models, similar and finely tuned hyperparameter configurations and gradient clipping strategies were used to ensure smooth model training. The former involved dividing the input video frame data into 8 groups, setting 30 training epochs, a batch size of 6, an initial learning rate of 0.01, and using a decaying learning rate. Additionally, a momentum of 0.9 was used to accelerate the optimization process, and overfitting was controlled through weight decay with a decay rate of 0.0005. The latter aimed to prevent gradient explosion and ensure that batch normalization was fully computed in each batch, thus guaranteeing the efficiency and stability of model training.

[0023] The reasoning phase includes the following steps:

[0024] Step 41, Preprocessing during the inference phase; A set of enhancement (TR) strategies were selected to preprocess the multimodal video data during the inference phase, including group scaling and group centering techniques. These methods enhance the model's robustness to the input data by scaling and cropping the images at different scales.

[0025] Step 42, Test-Time Enhancement (TTA) techniques, which include operations such as horizontal image flipping. This strategy allows the model to evaluate video content from different perspectives, aiming to improve the model's adaptability to changes in viewpoint.

[0026] Step 43, random weight averaging: Select the best-performing weight sets from the multiple weight sets saved at different stages of the multimodal action recognition model training process, and perform random weight averaging on the weight sets to obtain the optimized multimodal action recognition model weights for subsequent model integration.

[0027] The Stochastic Weight Averaging (SWA) technique selects the three best-performing weight sets from multiple weight sets saved at different stages of the training process of the TSM-Res50 and TSM-Res101 models, and performs SWA processing on them to obtain optimized TSM-Res50 and TSM-Res101 model weights for subsequent model ensemble. The purpose is to obtain more stable model weights for subsequent inference.

[0028] Step 44, Model Integration: By integrating the prediction results of different modalities, multimodal action recognition models with different architectures and pre-trained knowledge are integrated.

[0029] By integrating prediction results from different modalities, model ensembles are performed on TSM-Res50 and TSM-Res101 models with different architectures and pre-trained knowledge. The aim is to further improve model performance and effectively enhance the model's generalization ability.

[0030] Step 45, multi-time sampling; that is, repeatedly sampling multimodal data over time and fusing the inference results obtained from these samples, with the aim of improving the model's recognition accuracy.

[0031] Step 46, full-resolution inference; using multimodal video frames with a resolution of 256×256 as input, the aim is to improve recognition accuracy at the cost of increased computational complexity.

[0032] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the full-process motion recognition method for multimodal video as described above.

[0033] Beneficial effects: Compared with existing technologies, the method proposed in this invention can efficiently process multimodal data, balance the differences between different modalities, further improve the effective mining of multimodal information, achieve efficient multimodal information fusion, and improve the accuracy of multimodal action recognition. Specifically, this is reflected in the following aspects:

[0034] By comprehensively utilizing multimodal data including RGB, depth images, and thermal infrared (TIR) ​​images, efficient and accurate recognition of complex actions is achieved. This represents a breakthrough in recognition accuracy, significantly improving Top-1 accuracy to 99% and Top-5 accuracy to 100%.

[0035] This invention employs 2DCNNs (2D Convolutional Neural Networks) combined with a time-shifting module (TSM), which significantly reduces computational costs and improves processing speed while maintaining high performance, making it more suitable for large-scale practical applications.

[0036] By leveraging data augmentation techniques and pre-trained knowledge from transfer learning, this invention effectively enhances the model's generalization ability, enabling it to better adapt to new tasks and different application scenarios.

[0037] The method used in this invention achieves spatiotemporal feature extraction comparable to 3DCNN while retaining the training efficiency of 2DCNN and optimizing computational efficiency.

[0038] By employing dynamic group time-series sampling technology and group batch processing technology, this invention can more flexibly process data from different sources, improving the adaptability of data processing. This invention can be widely applied in fields such as video surveillance, human-computer interaction, smart homes, and industrial automation, and has broad market application prospects. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating the training of the TSM-Res50 model according to an embodiment of the present invention;

[0040] Figure 2 This is a flowchart illustrating the training of the TSM-Res101 model according to an embodiment of the present invention;

[0041] Figure 3 This is a flowchart of the reasoning stage in an embodiment of the present invention. Detailed Implementation

[0042] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0043] This paper presents a full-process action recognition method for multimodal videos. The method improves computational efficiency by optimizing data augmentation techniques, utilizing a pre-trained backbone network based on transfer learning, and combining multimodal spatial feature extraction and a temporal displacement module (TSM). Furthermore, it employs prediction augmentation methods such as stochastic weighted averaging (SWA), ensemble, and test-time augmentation (TTA) to predict actions from different perspectives and fully utilize target information.

[0044] The specific implementation process of the method is as follows:

[0045] RGB video data, TIR video data, and depth video data

[0046] Step 1, Data Acquisition: Acquire RGB video data, infrared (TIR) ​​video data, and depth video data from multiple sensors.

[0047] Step 2, data preprocessing: The data for each modality is evenly divided into multiple groups, and then frames are randomly selected from each group to form a multimodal video frame sequence for training. The data for each modality is normalized by subtracting the mean from each data point and dividing by the variance of the corresponding data point, ensuring that the data input to the model are within the same numerical range. After preprocessing the training set, the preprocessed training set is obtained.

[0048] Step 3, Group Batch Enhancement (TR), performs group multi-scale cropping and group random horizontal flipping on the multimodal video frames.

[0049] Step 4: Data concatenation. The multimodal data (RGB, TIR, and Depth) are concatenated along the channel dimension into an input tensor containing multimodal information.

[0050] Step 5, Spatiotemporal Feature Extraction: The concatenated input tensor is passed to the TSM-Res50 and TSM-Res101 models for spatiotemporal feature extraction. Then, the Time Translation Module (TSM) is used to capture the motion changes between frames in the time dimension, thereby better understanding the dynamic information of the video.

[0051] Step 6, Multimodal fusion, weighted fusion of the outputs (logits) of different modalities, and optimization of the model using the cross-entropy loss function.

[0052] Step 7: Model pre-training. Pre-trained weights from the Kinetics400, ImageNet, and Something-somethingV2 datasets are used to perform knowledge transfer on the TSM-Res50 and TSM-Res101 models before initialization. The pre-training process for TSM-Res50 and TSM-Res101 is detailed in the appendix. Figure 1 and attached Figure 2 .

[0053] Step 8, Model Fine-tuning: On a task-specific dataset, fine-tune the pre-trained model by adjusting hyperparameters (such as learning rate and batch size) to adapt to the distribution characteristics of the target dataset. Feed the pre-trained training set into the model for training, and use momentum and weight decay during training to accelerate convergence and prevent overfitting. Figure 1 and attached Figure 2 The training set preprocessing process is described in step 2. The input video frame data is divided into 8 groups, with 30 training epochs, a batch size of 6, an initial learning rate of 0.01, a momentum of 0.9 to accelerate the optimization process, and overfitting is controlled by weight decay with a decay rate of 0.0005.

[0054] Step 9: Random weight averaging. During training, save the weight sets for multiple stages, select the three best-performing weight sets for SWA processing, and use them for model fusion in the inference stage.

[0055] Step 10: In the inference stage, the data is first preprocessed, as shown in the attached document. Figure 3 The multimodal video frame preprocessing in the video is detailed in step 2. A multi-temporal sampling strategy is used to perform multiple temporal samplings of the video, for example, [attached]. Figure 3 Two different time samples (i.e., time sample 1 and time sample 2) are used to generate prediction results for different time segments, and these results are then merged to improve accuracy.

[0056] Step 11, Full-resolution input: Adjust the video frames to a resolution of 256×256 and input them into the model for inference at full resolution.

[0057] Step 12, preprocessing during the inference phase: A group enhancement (TR) strategy is selected to preprocess the multimodal video frames, resulting in preprocessed multimodal video frames. The TR strategy includes group scaling and group centering techniques to ensure consistent video scale, eliminate image edge noise, and focus on the main action area. (For example, see attached...) Figure 3 TR1 and TR2 in the text.

[0058] Step 13, Test-Time Enhancement (TTA) strategy: During inference, various enhancement operations, such as horizontal flipping, are performed on the input video frames to obtain prediction results from multiple viewpoints. All enhanced prediction results are then averaged to improve the model's adaptability to different viewpoints and lighting conditions. For example, see attached... Figure 3 The TTA1 and TTA2 in the sample are different because they are two different time samples and the time segments of the samples are also different.

[0059] Step 14, Data Concatenation: The processed multimodal data (RGB, TIR, and Depth) are concatenated along the channel dimension to form an input / output tensor containing multimodal information, as shown in the attached figure. Figure 3 The splicing operation in the middle.

[0060] Step 15, Model Integration: The concatenated data is fed into the TSM-Net50 and TSM-Net101 models respectively. The outputs of TSM-Res50 and TSM-Res101 are then weighted and fused using different model integration strategies to improve the final prediction accuracy and combine the advantages of each model to enhance generalization ability. (See attached diagram.) Figure 3 The data will be input into the TSM-Res50 model and the TSM-Res101 model.

[0061] Step 16, Final Prediction Output: The outputs of all augmentations and multi-time sampling results are fused, i.e., appended... Figure 3 The fusion output of the data is the final action recognition result, i.e., the final output, which is the attached data. Figure 3 The final output in the process.

[0062] The multimodal action recognition system corresponding to the end-to-end action recognition method for multimodal videos is described as follows:

[0063] (1) System Overview and Structural Framework

[0064] This multimodal action recognition system combines RGB, thermal infrared (TIR), and depth image data, utilizing a time-shift module (TSM) and convolutional neural networks of varying depths (such as ResNet50 and ResNet101) to efficiently extract spatiotemporal features. The system mainly comprises the following modules:

[0065] Data preprocessing module: Used for standardization and enhancement of multimodal data.

[0066] Feature extraction module: Spatiotemporal feature extraction is performed using TSM and pre-trained ResNet50 / ResNet101 networks.

[0067] Fusion and Training Module: Data from different modalities are concatenated along the channel dimension and trained using a weighted fusion strategy.

[0068] Inference and Prediction Module: Employs various augmentation and inference strategies, including Test-Time Data Augmentation (TTA), Stochastic Weighted Average (SWA), and model ensemble strategies, to improve prediction accuracy.

[0069] (2) Best practice for data preprocessing

[0070] To ensure the effectiveness of multimodal data and enhance the diversity of the dataset, the following methods were adopted in the preprocessing stage of this system:

[0071] Dynamic group time sampling strategy: Each video is divided into 8 time segments, and one frame is randomly sampled from each time segment to form a multimodal video frame sequence for training. This reduces interference from redundant frames and improves training efficiency.

[0072] Group Batch Augmentation (TR): This includes group multi-scale cropping and group random horizontal flipping. These methods can simulate video content from different perspectives and viewpoints, improving the model's generalization ability under different perspectives.

[0073] (3) Best practice for model training

[0074] During the training phase, the TSM framework, ResNet50, and ResNet101 were used as the base networks, and training was performed using multimodal data input. The specific steps are as follows:

[0075] Pre-training: TSM-ResNet50 and TSM-ResNet101 models were pre-trained using ImageNet, Kinetics400, and Something-somethingV2 datasets to improve the initial state of the models, ensure fast convergence, and achieve better initial performance.

[0076] Multimodal input and fusion: Frames from RGB, thermal infrared, and depth video are concatenated along the channel dimension and fed into the model for training. The model outputs logits for different modalities, which are then weighted and fused (e.g., α=0.2, for the output coefficients of depth data) for the final decision.

[0077] Hyperparameter configuration: Input data is divided into 8 groups, training epochs are 30, batch size is 6, initial learning rate is 0.01, learning rate decays during training, momentum is 0.9. Gradient clipping is used to prevent gradient explosion, and the weight decay coefficient is set to 5e-4.

[0078] (4) Implementation method of the reasoning stage

[0079] To improve the model's predictive performance, a series of enhancement strategies were applied during the inference phase:

[0080] Test-time data augmentation (TTA): By horizontally flipping and other transformations of the image, the model can evaluate video content from different perspectives, improving the model's adaptability to input data.

[0081] Random Weighted Average (SWA): During training, model weights from multiple different stages are saved, and the three weights with the best performance are selected for SWA processing to obtain a more robust prediction model.

[0082] Model ensemble: The prediction results of TSM-Res50 and TSM-Res101 are weighted and fused to further improve the generalization ability of the model.

[0083] Dual temporal sampling: Data is repeatedly sampled from the video during the inference phase to ensure that complete temporal clues are captured and to reduce interference from redundant information.

[0084] Detailed instructions for reproducing the invention: Technicians can implement the system using standard deep learning frameworks (such as PyTorch or TensorFlow) and pre-trained models (ResNet50, ResNet101) following the above implementation method. Data preprocessing, model training, and inference steps can all be implemented using existing open-source libraries (such as OpenCV, NumPy, etc.) without the need for additional exploration, research, or experimentation.

[0085] Obviously, those skilled in the art should understand that the steps of the full-process action recognition method for multimodal video or the modules of the multimodal action recognition system described in the above embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.

Claims

1. A method for end-to-end action recognition in multimodal video, characterized in that, The method includes: a multimodal data preprocessing stage, a multimodal fusion stage, a training stage, and an inference stage; the multimodal data refers to data from different modules, including RGB video data, infrared video data, and depth video data; The multimodal data preprocessing includes the following steps: Step 11, Dynamic Group Timing: When loading data of different modalities, the data of each modality is evenly divided into multiple groups, and frames are randomly selected from each group to form a multimodal video frame sequence for training. Step 12, group normalization; normalize the data for each modality; Step 13, Batch Enhancement Processing; Batch enhancement is performed on the obtained multimodal video frames; The multimodal fusion stage includes the following steps: Step 21: Feature extraction of multimodal data; Select two multimodal action recognition models with different architectures and pre-trained knowledge, use 2D convolutional neural networks to extract multimodal spatial features, and combine them with the time translation module of the multimodal action recognition model to achieve multimodal spatiotemporal feature extraction; The multimodal action recognition model selected is a model based on the TSM framework as the two multimodal action recognition models; Step 22, Multimodal fusion; The feature extraction results of different modalities obtained in step 21 are weighted and fused. The training phase includes the following steps: Step 31, pre-training transfer learning; using pre-trained knowledge to enhance the performance of the multimodal action recognition model in downstream tasks; Step 32: In the training phase of the multimodal action recognition model, perform hyperparameter configuration and gradient clipping. The reasoning phase includes the following steps: Step 41, Preprocessing during the inference stage; Select a set of enhancement strategies to preprocess the multimodal video data, including group scaling and group centering; Step 42, Enhancement during testing; Horizontal flip of the image; Step 43, random weight averaging; select the best-performing weight sets from the multiple weight sets saved at different stages of the multimodal action recognition model training process, and perform random weight averaging on the weight sets to obtain the optimized multimodal action recognition model weights for subsequent model integration. Step 44, Model Integration; By integrating the prediction results of different modalities, multimodal action recognition models with different architectures and pre-trained knowledge are integrated. Step 45, Multi-time sampling; Repeated time sampling is performed on the multimodal data, and the inference results obtained from these samples are fused; Step 46, full-resolution inference; input multimodal video frames with a resolution of 256×256 into multimodal action recognition.

2. The full-process action recognition method for multimodal video according to claim 1, characterized in that, In step 21, the TSM-Res50 model and the TSM-Res101 model based on the TSM framework are selected as two multimodal action recognition models.

3. The full-process action recognition method for multimodal video according to claim 2, characterized in that, The formula for the weighted fusion is: Logits R , Logits I , Logits D =F θ (Cat(X R , X I , X D )); ; The RGB modal video data is denoted as X. R ∈R T×W×H×C Infrared video data is denoted as X. I ∈R T×W×H×C Depth video data is denoted as X D ∈R T×W×H×C T, W, H, and C represent the number of video frames, frame width, frame height, and number of video channels, respectively; F θ The model to be trained is represented by θ parameterization; the function Cat(·) represents the connection operation; β, γ, and α are weight coefficients; Logits R Logits I and Logits D Let J(·) represent the Logits for different modalities, J(·) represent the cross-entropy loss function, and Y be the actual label predicted from the multimodal video data.

4. The full-process action recognition method for multimodal video according to claim 2, characterized in that, In step 31, during large-scale pre-training, the pre-training knowledge of the Kinetics400 and ImageNet datasets is used to enhance the performance of the TSM-Res50 model on downstream tasks, and the pre-training knowledge of the Something-somethingV2 and ImageNet datasets is used to enhance the performance of the TSM-Res101 model on downstream tasks.

5. The full-process action recognition method for multimodal video according to claim 2, characterized in that, In step 32, during the training phase of the multimodal action recognition model, hyperparameter configuration and gradient clipping are performed. In the hyperparameter configuration, the input video frame data is divided into 8 groups, 30 training epochs are set, the batch size is 6, the initial learning rate is 0.01, and a decaying learning rate is used. Momentum (set to 0.9) is used to accelerate the model optimization process. Overfitting is controlled by weight decay, with a decay rate of 0.0005.

6. The full-process action recognition method for multimodal video according to claim 2, characterized in that, In step 43, the three best-performing weight sets among the multiple weight sets saved at different stages of the training process of the TSM-Res50 and TSM-Res101 models are selected, and the three weight sets are subjected to SWA processing to obtain the optimized TSM-Res50 and TSM-Res101 model weights for subsequent model integration.

7. The full-process action recognition method for multimodal video according to claim 1, characterized in that, In step 21, the normalization process for the data of each modality is as follows: subtract the mean from the data points of each modality and divide by the variance of the data points.

8. The full-process action recognition method for multimodal video according to claim 1, characterized in that, In step 13, the obtained multimodal video frames are batch enhanced; the enhancement methods include group multi-scale cropping and group random horizontal flipping; group multi-scale cropping is to randomly crop images at different scales; group random horizontal flipping is to randomly flip images horizontally.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by the processor, they implement the steps of the full-process action recognition method for multimodal video as described in claim 1.