An action recognition method and system based on multimodal fusion

Through improved single-modal feature extraction and cross-modal feature fusion, the Transformer network and attention mechanism are used to solve the problem of high computing complexity in the existing methods, achieving higher recognition accuracy and lower computing consumption, and are suitable for mobile devices.

CN115205979BActive Publication Date: 2025-07-29NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210960093.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-07-29
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The existing multimodal fusion-based action recognition method performs a large amount of unnecessary time modeling before cross-modal interactions, and fails to effectively eliminate redundant information in modal features, resulting in high computational complexity and difficulty in deploying on mobile or low-resource devices.

Method used

The improved single-modal feature extraction method is adopted, parameterless modeling is used using the Transformer network, and meaningful cross-modal feature combinations are extracted through the attention mechanism, and the most valuable information is selected in combination with the Token selection module to fusion to reduce redundant information.

Benefits of technology

This improves the accuracy of action recognition, while significantly reducing computing consumption, making the method more suitable for deployment on mobile or low-resource devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205979B_ABST
    Figure CN115205979B_ABST
Patent Text Reader

Abstract

The present invention discloses an action recognition method and system based on multimodal fusion. The method includes: extracting visual modal data and auditory modal data from an action video; preprocessing the visual modal data and the auditory modal data to obtain a visual modal shallow Token sequence and an auditory modal shallow Token sequence; inputting the visual modal shallow Token sequence into a visual feature extraction network to obtain a visual modal deep Token sequence; inputting the auditory modal shallow Token sequence into an auditory feature extraction network to obtain an auditory modal deep Token sequence; merging the visual modal deep Token sequence and the auditory modal deep Token sequence to obtain a merged Token sequence; inputting the merged Token sequence into a feature fusion network to obtain a Token sequence after fusion and interaction; and inputting the Token sequence after fusion and interaction into a fully connected layer to obtain an action classification result. Compared with existing methods, the present invention has a higher recognition accuracy and lower computational consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of action recognition, and in particular to an action recognition method and system based on multimodal fusion. Background Art

[0002] Action recognition is a key research topic in the fields of computer vision and multimedia, and its application directions include video surveillance, robot-human interaction, video retrieval, sports game analysis, etc. Currently, most mainstream action recognition methods are based on RGB videos, that is, they only focus on visual information and lack research on auditory information. In fact, auditory information is of great significance for action recognition. There are a large number of actions in human life that can be judged by sound, such as whistling and playing musical instruments, and some actions are very difficult to distinguish only by vision, such as "singing" and "speaking". For these actions, we can use audio information for supplementation to obtain a higher recognition accuracy. The goal of multimodal fusion is to make full use of all available information around us, so as to improve the model's perception ability of the external environment, and ultimately approach or even exceed humans.

[0003] Currently, the action recognition methods based on multimodal fusion mainly use Transformer as the backbone network, and there are mainly three fusion methods: early fusion, mid-term fusion, and late fusion. Late fusion is the simplest, and only needs to calculate the average value of the classification scores of independent audio and video networks to obtain the final classification result. Both early fusion and mid-term fusion utilize the characteristic that the Transformer network can process sequences of different lengths, ignoring the problem of the form difference of audio-visual inputs, and connecting the feature vectors of the audio-visual modalities before inputting into the Transformer network or in the middle layer of the network, so that the entire model has the ability to simultaneously sense auditory and visual information. Among them, compared with early fusion, the mid-term fusion model has higher recognition accuracy and lower computational complexity. However, the mid-term fusion model still has two disadvantages: (1) Before cross-modal interaction, the model performs a large amount of unnecessary temporal modeling on video and audio inputs. (2) For cross-modal modeling, it does not consider the large amount of redundant information existing in each modality, but directly connects the features of the two modalities. These two disadvantages cause a sharp increase in the computational complexity of the model, which is not conducive to the deployment of action recognition methods on mobile devices or other devices with low resource consumption. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the present invention provides an action recognition method and system based on multimodal fusion.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for action recognition based on multi-modal fusion, comprising:

[0007] Extracting visual modal data and auditory modal data from an action video;

[0008] Preprocessing the visual modal data and the auditory modal data to obtain a visual modal shallow Token sequence and an auditory modal shallow Token sequence;

[0009] Inputting the visual modal shallow Token sequence into a visual feature extraction network to obtain a visual modal deep Token sequence;

[0010] Inputting the auditory modal shallow Token sequence into an auditory feature extraction network to obtain an auditory modal deep Token sequence;

[0011] Merging the visual modal deep Token sequence and the auditory modal deep Token sequence to obtain a merged Token sequence;

[0012] Inputting the merged Token sequence into a feature fusion network to obtain a Token sequence after fusion and interaction;

[0013] Inputting the Token sequence after fusion and interaction into a fully connected layer to obtain an action classification result.

[0014] Optionally, extracting visual modal data and auditory modal data from an action video, specifically including:

[0015] Dividing the action video into multiple parts;

[0016] Randomly extracting 1 RGB image from each part to obtain visual modal data;

[0017] Extracting audio of a set length from each part;

[0018] Extracting a spectrogram of a set frequency dimension from the audio to obtain auditory modal data.

[0019] Optionally, preprocessing the visual modal data and the auditory modal data, specifically including:

[0020] Dividing both the visual modal data and the auditory modal data into multiple image blocks to obtain visual modal image blocks and auditory modal image blocks;

[0021] Flattening each visual modal image block and each auditory modal image block into a one-dimensional vector to obtain visual modal Tokens and auditory modal Tokens;

[0022] Perform a linear transformation on the visual modality tokens and the auditory modality tokens to obtain an initial visual modality token sequence and an initial auditory modality token sequence;

[0023] Add learnable variables as position information to the initial visual modality token sequence and the initial auditory modality token sequence respectively to obtain a shallow visual modality token sequence and a shallow auditory modality token sequence.

[0024] Optionally, before inputting the shallow visual modality token sequence into the visual feature extraction network and inputting the shallow auditory modality token sequence into the auditory feature extraction network, it further includes;

[0025] Set a classification vector in front of the shallow visual modality token sequence and the shallow auditory modality token sequence respectively, and move the classification vector.

[0026] Optionally, input the merged token sequence into a feature fusion network to obtain a token sequence after fusion and interaction, specifically including:

[0027] Merge the classification vectors in the deep visual modality token sequence and merge the parts other than the classification vectors in the deep visual modality token sequence to obtain a merged deep visual modality token sequence;

[0028] Merge the classification vectors in the deep auditory modality token sequence and merge the parts other than the classification vectors in the deep auditory modality token sequence to obtain a merged deep auditory modality token sequence;

[0029] Merge the merged deep visual modality token sequence and the merged deep auditory modality token sequence to obtain a merged token sequence.

[0030] Optionally, the feature fusion network includes a token selection module.

[0031] The present invention also provides an action recognition system based on multi-modal fusion, including:

[0032] A modality data extraction module for extracting visual modality data and auditory modality data from an action video;

[0033] A preprocessing module for preprocessing the visual modality data and the auditory modality data to obtain a shallow visual modality token sequence and a shallow auditory modality token sequence;

[0034] The first input module is used to input the visual modality shallow Token sequence into the visual feature extraction network to obtain a visual modality deep Token sequence;

[0035] The second input module is used to input the auditory modality shallow Token sequence into the auditory feature extraction network to obtain an auditory modality deep Token sequence;

[0036] The merging module is used to merge the visual modality deep Token sequence and the auditory modality deep Token sequence to obtain a merged Token sequence;

[0037] The third input module is used to input the merged Token sequence into the feature fusion network to obtain a Token sequence after fusion and interaction;

[0038] The fourth input module is used to input the Token sequence after fusion and interaction into the fully connected layer to obtain an action classification result.

[0039] Optionally, the modality data extraction module specifically includes:

[0040] The first partitioning unit is used to partition the action video into multiple parts;

[0041] The first extraction unit is used to randomly extract 1 RGB image from each part to obtain visual modality data;

[0042] The second extraction unit is used to extract audio of a set length from each part;

[0043] The third extraction unit is used to extract a spectrogram of a set frequency dimension from the audio to obtain auditory modality data.

[0044] Optionally, the preprocessing module specifically includes:

[0045] The second partitioning unit is used to partition both the visual modality data and the auditory modality data into multiple image blocks to obtain visual modality image blocks and auditory modality image blocks;

[0046] The flattening unit is used to flatten each visual modality image block and each auditory modality image block into a one-dimensional vector to obtain visual modality Tokens and auditory modality Tokens;

[0047] The linear transformation unit is used to perform a linear transformation on the visual modality Tokens and the auditory modality Tokens once to obtain a visual modality initial Token sequence and an auditory modality initial Token sequence;

[0048] An adding unit for respectively adding learnable variables as position information to the initial Token sequence of the visual modality and the initial Token sequence of the auditory modality to obtain a shallow Token sequence of the visual modality and a shallow Token sequence of the auditory modality.

[0049] Optionally, a Token selection module is included in the feature fusion network.

[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0051] An action recognition method based on multi-modal fusion provided by the present invention includes: extracting visual modality data and auditory modality data from an action video; preprocessing the visual modality data and the auditory modality data to obtain a shallow Token sequence of the visual modality and a shallow Token sequence of the auditory modality; inputting the shallow Token sequence of the visual modality into a visual feature extraction network to obtain a deep Token sequence of the visual modality; inputting the shallow Token sequence of the auditory modality into an auditory feature extraction network to obtain a deep Token sequence of the auditory modality; merging the deep Token sequence of the visual modality and the deep Token sequence of the auditory modality to obtain a merged Token sequence; inputting the merged Token sequence into a feature fusion network to obtain a Token sequence after fusion and interaction; inputting the Token sequence after fusion and interaction into a fully connected layer to obtain an action classification result. The multi-modal based action recognition method proposed by the present invention has a higher recognition accuracy and lower computational consumption compared with existing methods. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the action recognition method based on multi-modal fusion provided by the present invention;

[0054] Figure 2 It is an overall flowchart of the action recognition method based on multi-modal fusion provided by the present invention;

[0055] Figure 3 It is a schematic diagram of the Token selection module provided by the present invention. Detailed Embodiments

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] The present invention proposes a novel action recognition method, which adopts mid-term fusion. The whole method is divided into two stages: a single-modal feature extraction stage and a cross-modal feature fusion stage. In single-modal modeling, the present invention replaces the traditional time modeling method with a parameter-free modeling method with less computational complexity; in cross-modal modeling, the present invention uses an attention mechanism to extract meaningful cross-modal feature combinations to eliminate redundant information.

[0058] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] As Figure 1-2 shown, a multi-modal fusion-based action recognition method of the present invention includes the following steps:

[0060] Step 101: Extract visual modal data and auditory modal data from the action video.

[0061] In practical applications, the present invention divides the video clip into 8 parts, randomly extracts 1 frame of image from each part, and a total of 8 frames of images are used as the visual modal data I v , where the pixel size of each frame of image is 224×224. In addition, the present invention extracts a section of audio with a length of 0.639s from each part, and extracts a logmel spectrogram with a frequency dimension of 128 from the audio. At the same time, the size and step of the Hamming window are adjusted so that the size of the spectrogram is 128×128. The present invention uses these 8 spectrograms as the auditory modal data I a .

[0062] Step 102: Preprocess the visual modal data and the auditory modal data to obtain a visual modal shallow Token sequence and an auditory modal shallow Token sequence.

[0063] In practical applications, according to the introduction in step 101, the size of the visual modal data I v is 8×3×224×224, and the size of the auditory modal data I a is 8×1×128×128. First, each frame of RGB image and logmel spectrogram is divided into 16×16 image blocks, and the number of blocks is 8×(224 / 16) for the visual modality 2= 156 blocks, auditory modality 8×(128 / 16) 2 = 512 blocks. After that, each image block is flattened into a one-dimensional vector, and this vector is called a Token. Among them, the vector dimension of the visual modality is 16×16×3 = 768, and the vector dimension of the auditory modality is 16×16×1 = 256. The Token sequences obtained by merging the Tokens corresponding to all image blocks have dimensions of 1568×768 for the visual modality and 512×256 for the auditory modality respectively. Further, it is necessary to extract the features of each image block. A linear transformation is performed on the Token vector corresponding to each image block in both the visual and auditory modalities, that is, through a fully connected layer, and the output vector dimension is 768. Calculate the Token sequences M v 、M a , whose dimensions are 1568×768 and 512×768 respectively. Such Token sequences cannot be directly input into the Transformer's Encoder module as they lack the position information of the image blocks. Use learnable variables as the position information of the image blocks and add them to M v 、M a to obtain the final initial Token sequences of the visual modality and the initial Token sequences of the auditory modality This sequence contains the feature information and position information of each image block.

[0064] Step 103: Input the shallow Token sequence of the visual modality into the visual feature extraction network to obtain the deep Token sequence of the visual modality.

[0065] Step 104: Input the shallow Token sequence of the auditory modality into the auditory feature extraction network to obtain the deep Token sequence of the auditory modality.

[0066] In practical applications, the shallow Token sequences of the visual and auditory modalities obtained in step 102 are respectively input into independent single-modal feature extraction networks to obtain sequences E v ,E a that can represent the deep features of visual and auditory information.

[0067] The visual feature extraction network and the auditory feature extraction network mentioned in steps 103 - 104 are both networks improved from the single-modal feature extraction network based on Transformer, and do not require time modeling, that is, there is no attention operation on the sequence of image blocks between frames in the Transformer's Encoder module.

[0068] In Before inputting into the Encoder module, the shallow Token sequences of the visual and auditory modalities are recombined, and the shallow Token sequence of each frame is represented separately. Finally, a classification vector S is set before the shallow Token sequence of each frame. c , the shallow Token sequence of the visual modality has a size of 8×197×768, and the shallow Token sequence of the auditory modality has a size of 8×65×768. The shallow Token sequence corresponding to each frame of RGB image / spectrogram is separately input into the Encoder module for self-attention operation. Compared with directly calculating the correlation of all Token vectors, although this greatly reduces the computational complexity, it loses the dynamic information unique to the action. To regain the dynamic information of the action, the classification vectors in the shallow Token sequence of each frame are shifted. Each classification vector S c can be divided into three parts, S c = [g a , g b , g c , where g a / g c represent the parts shifted to the previous frame / next frame respectively, and g b represents the part that does not move. This operation can be expressed by the following formula:

[0069] g a (t) = g a (t - 1)

[0070] g b (t) = g b (t)

[0071] g c (t) = g c (t + 1)

[0072] t = 1, 2,..., T

[0073] After passing through 10 improved Encoder modules, the deep Token sequences E v , E a of the visual and auditory modalities are finally obtained. The dimensional sizes of these two sequences are 8×197×76 and 8×65×768 respectively.

[0074] Step 105: Merge the deep Token sequence of the visual modality and the deep Token sequence of the auditory modality to obtain the merged Token sequence.

[0075] In practical applications, the deep Token sequences E v , Ea Merge into E f 。

[0076] Step 1: Merge the depth token sequences of each frame of each modality. The depth token sequences to be merged are divided into two parts: (1) Merging of classification vectors. Calculate the average value of the classification vectors of each frame to obtain a classification vector that can represent the overall category of the modality. (2) Merging of the parts other than the classification vectors. Concatenate the depth token sequences of 8 frames except for the classification vectors. Finally, merge to obtain E'. v , E' a , with dimensions of 1569×768 and 513×768 respectively.

[0077] Step 2: Merge the token sequences E' v , E' a of the two modalities into a single token sequence E f . Calculate the average of the classification vectors and perform a concatenation operation on the remaining parts. Finally, merge to obtain E f , with a dimension of 2081×76.

[0078] Step 106: Input the merged token sequence into the feature fusion network to obtain the token sequence after fusion and interaction.

[0079] In practical applications, input the E obtained in step 105 f into the feature fusion network to enable full interaction of the information of the two modalities and output the token sequence E' f .

[0080] In the present invention, the feature fusion network is similar to the traditional mid-term fusion method and is also based on the Transformer structure. However, different from mid-term fusion, the feature fusion network used in the present invention includes a token selection module, which selects 16 different token combinations from the token sequence of length 2080. Each token combination is the weighted average of all token sequences, which emphasizes the most meaningful part of the token sequence. The schematic diagram of the token selection module is as Figure 3 shown.

[0081] Among them, the present invention designs a function s i = A i (E f ) to use the cross-modal attention mechanism to select and combine meaningful tokens. The token selection module calculates the weight mapping related to the input E f through learning and multiplies it with E f itself. The present invention will ai (E f ) as a function for generating cross-modal weight mappings, which is composed of MLPs. Function A i (·) can be further written as: s i = A i (E f ) = p(E f ⊙ a i (E f ))), where ⊙ represents dot product, and p(·) represents the global average pooling function. There are a total of 16 functions A i (·) in the present invention. That is to say, through the Token selection module, the present invention can obtain a new Token sequence s = [s1, s2,..., s 16 .

[0082] To further establish the relationship between different modalities, the present invention combines the selected Token sequence S and the classification vector C obtained in step 4 and inputs them into a two-layer Transformer network, and finally outputs E' f = [C 2 ; S 2 .

[0083] Step 107: Input the Token sequence after fusion and interaction into the fully connected layer to obtain the action classification result.

[0084] The multi-modal based action recognition method proposed by the present invention has a higher recognition accuracy and lower computational consumption compared with the existing methods.

[0085] The advantages mainly come from two aspects: (1) The improvement of the single-modal modeling method in steps 103 - 104. The method proposed by the present invention replaces the temporal modeling in the existing method, that is, instead of directly performing Token interaction between frames, it adopts the method of translating the classification vector between different frames to obtain the dynamic information of the action. This method reduces the length of the Transformer input sequence, thereby greatly reducing the computational consumption. (2) The improvement of the feature fusion network in step 106. The feature fusion network of the present invention adds a Token selection module, which selects the most valuable 16 Token combinations, greatly reducing the redundancy of multi-modal data, thereby greatly reducing the computational amount while improving the recognition accuracy.

[0086] The present invention also provides a multi-modal fusion based action recognition system, including:

[0087] A modal data extraction module for extracting visual modal data and auditory modal data from the action video;

[0088] A preprocessing module for preprocessing visual modality data and auditory modality data to obtain a visual modality shallow Token sequence and an auditory modality shallow Token sequence;

[0089] A first input module for inputting the visual modality shallow Token sequence into a visual feature extraction network to obtain a visual modality deep Token sequence;

[0090] A second input module for inputting the auditory modality shallow Token sequence into an auditory feature extraction network to obtain an auditory modality deep Token sequence;

[0091] A merging module for merging the visual modality deep Token sequence and the auditory modality deep Token sequence to obtain a merged Token sequence;

[0092] A third input module for inputting the merged Token sequence into a feature fusion network to obtain a fused and interactive Token sequence;

[0093] A fourth input module for inputting the fused and interactive Token sequence into a fully connected layer to obtain an action classification result.

[0094] Among them, the modality data extraction module specifically includes:

[0095] A first partitioning unit for partitioning an action video into multiple parts;

[0096] A first extraction unit for randomly extracting 1 RGB image from each part to obtain visual modality data;

[0097] A second extraction unit for extracting a set length of audio from each part;

[0098] A third extraction unit for extracting a spectrogram of a set frequency dimension from the audio to obtain auditory modality data.

[0099] Among them, the preprocessing module specifically includes:

[0100] A second partitioning unit for partitioning both the visual modality data and the auditory modality data into multiple image blocks to obtain visual modality image blocks and auditory modality image blocks;

[0101] A flattening unit for flattening each visual modality image block and each auditory modality image block into a one-dimensional vector to obtain visual modality Tokens and auditory modality Tokens;

[0102] A linear transformation unit for performing a linear transformation on the visual modality Tokens and the auditory modality Tokens once to obtain a visual modality initial Token sequence and an auditory modality initial Token sequence;

[0103] An adding unit, configured to separately add learnable variables as position information to an initial Token sequence of a visual modality and an initial Token sequence of an auditory modality, to obtain a shallow Token sequence of the visual modality and a shallow Token sequence of the auditory modality.

[0104] Optionally, a Token selection module is included in the feature fusion network.

[0105] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0106] Specific examples are used herein to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An action recognition method based on multimodal fusion, characterized in that Including: Extracting visual modality data and auditory modality data from an action video; Preprocessing the visual modality data and the auditory modality data to obtain a visual modality shallow Token sequence and an auditory modality shallow Token sequence; Inputting the visual modality shallow Token sequence into a visual feature extraction network to obtain a visual modality deep Token sequence; Inputting the auditory modality shallow Token sequence into an auditory feature extraction network to obtain an auditory modality deep Token sequence; Merging the visual modality deep Token sequence and the auditory modality deep Token sequence to obtain a merged Token sequence; Inputting the merged Token sequence into a feature fusion network to obtain a Token sequence after fusion and interaction; specifically including: merging the classification vectors in the visual modality deep Token sequence and merging the parts other than the classification vectors in the visual modality deep Token sequence to obtain a merged visual modality deep Token sequence; merging the classification vectors in the auditory modality deep Token sequence and merging the parts other than the classification vectors in the auditory modality deep Token sequence to obtain a merged auditory modality deep Token sequence; merging the merged visual modality deep Token sequence and the merged auditory modality deep Token sequence to obtain a merged Token sequence; wherein, the feature fusion network includes a Token selection module, and the Token selection module is selected and combined by using a cross-modal attention mechanism; Inputting the Token sequence after fusion and interaction into a fully connected layer to obtain an action classification result; Wherein, before inputting the visual modality shallow Token sequence into the visual feature extraction network and inputting the auditory modality shallow Token sequence into the auditory feature extraction network, it further includes: Setting a classification vector before the visual modality shallow Token sequence and the auditory modality shallow Token sequence respectively, and moving the classification vector.

2. The action recognition method based on multi-modal fusion according to claim 1, wherein Extracting visual modality data and auditory modality data from an action video, specifically including: Dividing the action video into multiple parts; Randomly extracting 1 RGB image from each part to obtain visual modality data; Extracting audio of a set length from each part; Extracting a spectrogram of a set frequency dimension from the audio to obtain auditory modality data.

3. The action recognition method based on multi-modal fusion according to claim 1, wherein Preprocessing the visual modality data and the auditory modality data, specifically including: Dividing both the visual modality data and the auditory modality data into multiple image blocks to obtain visual modality image blocks and auditory modality image blocks; Flattening each visual modality image block and each auditory modality image block into a one-dimensional vector to obtain visual modality Tokens and auditory modality Tokens; Performing a linear transformation on the visual modality Tokens and the auditory modality Tokens to obtain a visual modality initial Token sequence and an auditory modality initial Token sequence; Add learnable variables as position information to the initial Token sequence of the visual modality and the initial Token sequence of the auditory modality respectively, to obtain a shallow Token sequence of the visual modality and a shallow Token sequence of the auditory modality.

4. An action recognition system based on multimodal fusion, characterized in that, It includes: A modality data extraction module for extracting visual modality data and auditory modality data from an action video; A preprocessing module for preprocessing the visual modality data and the auditory modality data to obtain a shallow Token sequence of the visual modality and a shallow Token sequence of the auditory modality; A first input module for inputting the shallow Token sequence of the visual modality into a visual feature extraction network to obtain a deep Token sequence of the visual modality; A second input module for inputting the shallow Token sequence of the auditory modality into an auditory feature extraction network to obtain a deep Token sequence of the auditory modality; A merging module for merging the deep Token sequence of the visual modality and the deep Token sequence of the auditory modality to obtain a merged Token sequence; A third input module for inputting the merged Token sequence into a feature fusion network to obtain a Token sequence after fusion and interaction; specifically including: merging the classification vectors in the deep Token sequence of the visual modality and merging the parts other than the classification vectors in the deep Token sequence of the visual modality to obtain a merged deep Token sequence of the visual modality; merging the classification vectors in the deep Token sequence of the auditory modality and merging the parts other than the classification vectors in the deep Token sequence of the auditory modality to obtain a merged deep Token sequence of the auditory modality; merging the merged deep Token sequence of the visual modality and the merged deep Token sequence of the auditory modality to obtain a merged Token sequence; wherein, the feature fusion network includes a Token selection module, and the Token selection module is selected and combined using a cross-modal attention mechanism; A fourth input module for inputting the Token sequence after fusion and interaction into a fully connected layer to obtain an action classification result; Wherein, before inputting the shallow Token sequence of the visual modality into the visual feature extraction network and inputting the shallow Token sequence of the auditory modality into the auditory feature extraction network, it further includes: Set a classification vector before the shallow Token sequence of the visual modality and the shallow Token sequence of the auditory modality respectively, and move the classification vector.

5. The action recognition system based on multi-modal fusion according to claim 4, characterized in that, The modality data extraction module specifically includes: A first partitioning unit for partitioning the action video into multiple parts; A first extraction unit for randomly extracting 1 frame of RGB image from each part to obtain visual modality data; A second extraction unit for extracting audio of a set length from each part; A third extraction unit for extracting a spectrogram of a set frequency dimension from the audio to obtain auditory modality data.

6. The action recognition system based on multi-modal fusion according to claim 4, characterized in that The preprocessing module specifically includes: A second partitioning unit, configured to partition both the visual modality data and the auditory modality data into a plurality of image patches, obtaining visual modality image patches and auditory modality image patches; A flattening unit, configured to flatten each visual modality image patch and each auditory modality image patch into a one-dimensional vector, obtaining visual modality tokens and auditory modality tokens; A linear transformation unit, configured to perform a linear transformation on the visual modality tokens and the auditory modality tokens once, obtaining an initial visual modality token sequence and an initial auditory modality token sequence; An addition unit, configured to add learnable variables as position information to the initial visual modality token sequence and the initial auditory modality token sequence respectively, obtaining a shallow visual modality token sequence and a shallow auditory modality token sequence.