Video classification method for carrying out adaptive adjustment on image basic model based on Mangbar network
The basic image model is adapted and adjusted through the Mamba network, which solves the problem of space-time dynamic capture of video data, and realizes efficient video recognition and feature representation, which improves the accuracy and computing efficiency of video recognition.
Patent Information
- Application Number
- CN202510320695.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art has problems such as spatiotemporal modeling fragmentation, high computational complexity, information loss and semantic gap in video data processing, and it is difficult to effectively capture the spatiotemporal dynamics of video data.
The Mamba network is used to adapt and adjust the image basic model, group the autocorrelation characteristics through the self-attention mechanism, and sequence modulation is used by the Mamba network's state space model, and forward propagation and loss function training are carried out in combination with the image basic model to freeze the parameters of the image basic model.
It improves the accuracy and computing efficiency of video recognition, enhances feature representation ability and network generalization ability, reduces computational complexity, and optimizes video feature extraction and model learning efficiency.
Smart Images

Figure CN120388315A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and video understanding, and in particular to a video classification method for adaptively adjusting an image base model based on a Mamba network. Background Art
[0002] Video understanding is a key and challenging task in computer vision. The key to solving this challenge lies in learning effective spatio-temporal representations from video data. Since video data exhibits complex spatio-temporal dynamics, training a video understanding model from scratch is inefficient and requires a large amount of data. In contrast, significant progress has been made in image understanding by using image base models. This has prompted the exploration of applying image base models to video understanding, as image base models provide powerful pre-trained representations and reduce the dependence on training task-specific models from scratch.
[0003] For example, Chinese Patent Application Publication No. CN115272941A discloses a weakly supervised video temporal action detection and classification method and system. However, this method performs video understanding by separately processing spatial information and temporal information, and these designs simplify the complete modeling of the sequence. Therefore, image patches at different spatial and temporal positions can only interact in an indirect and implicit manner. This may not be sufficient to capture the complex dynamics inherent in video data. To effectively capture spatio-temporal dynamics, traditional methods process the entire spatio-temporal sequence through global attention, but due to the high complexity, it leads to memory and speed efficiency problems.
[0004] Therefore, the spatio-temporal coupling characteristics of video data lead to the following problems when directly migrating image models:
[0005] (1) Spatio-temporal modeling fragmentation: Existing methods separate spatial feature extraction from temporal correlation analysis, resulting in limited spatio-temporal interaction;
[0006] (2) Computational complexity explosion: Although the global spatio-temporal attention mechanism can capture long-range dependencies, the computational complexity grows quadratically with the number of frames, making it difficult to process long video sequences.
[0007] (3) Information loss and semantic gap: To reduce computational costs, most methods downsample or compress the video, resulting in the loss of high-frequency motion details and weakening the ability to capture fast actions or microscopic changes.
[0008] In summary, there is currently a lack of a video classification method to solve or partially solve the aforementioned problems. Summary of the Invention
[0009] The object of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a video classification method for adaptively adjusting an image base model based on the Mamba network, so as to solve or partially solve the problems of high computational complexity, low efficiency, and difficulty in capturing complete spatio-temporal dynamics when processing video data.
[0010] The object of the present invention can be achieved by the following technical solutions:
[0011] One aspect of the present invention provides a video classification method for adaptively adjusting an image base model based on the Mamba network, and uses the adaptively adjusted image base model to classify a target video. The process of obtaining the adaptively adjusted image base model includes the following steps:
[0012] Step S1, obtain a video, and perform feature encoding on each frame image in the video by using a pre-trained image base model to obtain a long sequence of video features including frame time information;
[0013] Step S2, group the long sequence of video features to obtain multiple subsequences, and calculate the self-correlation features of each subsequence based on the self-attention mechanism;
[0014] Step S3, use the merged self-correlation features as the input of the Mamba network to obtain an output feature sequence, and perform sequence modulation on the output feature sequence and the long sequence of video features to obtain modulated features;
[0015] Step S4, based on the modulated features, perform forward propagation on each layer of the image base model, and repeat steps S2 - S4 to obtain the final features of the video;
[0016] Step S5, freeze the parameters of the image base model, and based on the final features, use a classifier to obtain a class probability distribution, and train the Mamba network by calculating a loss function to achieve the adaptive adjustment of the image base model.
[0017] As a preferred technical solution, in step S1, the process of obtaining the long sequence of video features includes the following steps:
[0018] Step S101, extract frames from the obtained video to obtain a video frame sequence;
[0019] Step S102, perform image-level encoding on each video frame by using a pre-trained image base model, and add position encoding information based on the time series;
[0020] Step S103, merge the encoding results of all video frames to obtain a long sequence of video features.
[0021] As a preferred technical solution, in step S2, the process of calculating the autocorrelation features of each subsequence includes the following steps:
[0022] Step S201: For the long sequence video features, use a 2D window to segment them to obtain multiple subsequences;
[0023] Step S202: For each subsequence, use the self-attention layer weights pre-trained by the image basic model to calculate its autocorrelation features;
[0024] Step S203: Combine the calculation results of all subsequences to obtain the combined autocorrelation features of the entire long sequence video.
[0025] As a preferred technical solution, in step S3, the process of obtaining the modulated features includes the following steps:
[0026] Step S301: Configure the state space model in the Mamba network so that it can accept the original video sequence as input and output two different output feature sequences as modulation parameters;
[0027] Step S302: Use the sequence modulation function to perform sequence modulation on the output feature sequence and the long sequence video features to obtain the modulated features.
[0028] As a preferred technical solution, the modulation function is:
[0029] SeqMod(x,y1,y2) = x⊙y1 + x + y2
[0030] where SeqMod(x,y1,y2) is the modulated feature, x is the long sequence video feature, and y1 and y2 are two different output feature sequences output by the state space model.
[0031] As a preferred technical solution, in step S4, based on the modulated features, perform forward propagation on each layer of the image basic model, and repeat steps S1 - S4 to obtain the final features of the video, the process includes the following steps:
[0032] Step S401: For the i-th layer of the image basic model, use the weights of its forward propagation part to process the current sequence to complete the adaptation adjustment of the current model layer;
[0033] Step S402: Repeat steps S2 - S401 until the adaptation adjustment of all layers of the current model is completed and the final features of the output video are obtained.
[0034] As a preferred technical solution, step S4 further includes:
[0035] Step S403, perform temporal and spatial averaging operations on the final features of the video.
[0036] As a preferred technical solution, step S5 includes the following steps:
[0037] Step S501, process according to the language names of all categories to be classified through a pre-trained language model to obtain an untrainable category classifier;
[0038] Step S502, calculate the cosine similarity between the final features of the video and each category representative of the classifier, and obtain the category prediction probability distribution according to the similarity;
[0039] Step S503, based on the comparison between the category prediction probability distribution and the true label, calculate the cross-entropy as the classification loss function;
[0040] Step S504, based on the intermediate calculation results of the forward inference in step S4, and compare with the intermediate calculation results of the unadjusted base model, and calculate the mean square error as the distillation loss function;
[0041] Step S505, based on the classification loss function and the distillation loss function, calculate the final loss function through weighted summation;
[0042] Step S506, based on the final loss function, train the Mamba network by the gradient descent method until convergence.
[0043] As a preferred technical solution, the image base model is CLIP or SigLIP.
[0044] Another aspect of the present invention provides a video classification system for adaptively adjusting an image base model based on a Mamba network, for implementing the foregoing video classification method for adaptively adjusting an image base model based on a Mamba network. The video classification system includes:
[0045] An image encoding module, configured to obtain a video, and perform feature encoding on each frame image in the video by using a pre-trained image base model to obtain a long sequence of video features including frame time information;
[0046] An autocorrelation calculation module, configured to group the long sequence of video features to obtain a plurality of subsequences, and calculate the autocorrelation features of each subsequence based on the self-attention mechanism;
[0047] A feature modulation module, configured to use the merged autocorrelation features as the input of the Mamba network to obtain an output feature sequence, and perform sequence modulation on the output feature sequence and the long sequence of video features to obtain modulated features;
[0048] A forward propagation module, which is used to perform forward propagation on each layer of the image base model based on the modulated features, and repeat multiple times to obtain the final features of the video;
[0049] A training module, which is used to freeze the parameters of the image base model, based on the final features, use a classifier to obtain a class probability distribution, and train the Mamba network by calculating a loss function to achieve adaptive adjustment of the image base model.
[0050] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0051] (1) Improve the training accuracy of video recognition: The method for adapting an image base model based on the Mamba network provided by the present invention only trains the Mamba network while freezing the parameters of the pre-trained image base model, and can effectively adjust the video features by introducing only the Mamba network module with a small number of parameters without changing the structure of the original image base model. This method can maximize the use of the weights of the image pre-training model and better capture the spatio-temporal dynamics in the video, thereby significantly improving the accuracy of video recognition.
[0052] (2) Improve the representation ability of features: The present invention divides the long-sequence video features into multiple subsequences (i.e., partitioning), and performs forward propagation (i.e., modulation) on each layer of the image base model based on the modulated features, thereby providing a video recognition framework including a "partitioning and adjustment" stage. This framework first applies window-based spatial local attention to each layer of the video base model, and then injects complete spatio-temporal information through a modulation function. Such a design can not only improve the feature representation ability, but also enhance the generalization ability of the network, making it more adaptable to the changes of different video data.
[0053] (3) Low computational complexity: By introducing the Mamba network and the state space model (SSM), the present invention can efficiently process long-sequence video features, reduce the computational complexity, and significantly improve the processing speed and memory utilization rate.
[0054] (4) Good training effect: Through the weighted training method of the distillation loss function and the classification loss function, combined with the adjustment of the base model and the Mamba network, the present invention can simultaneously optimize the video feature extraction ability and the learning efficiency of the model, and further improve the performance of the video recognition task. Description of the Drawings
[0055] Figure 1 It is a flowchart of the video classification method for adaptively adjusting the image base model based on the Mamba network in the embodiment;
[0056] Figure 2Schematic diagram of adaptively adjusting an image base model based on the Mamba network in an embodiment;
[0057] Figure 3 Schematic diagram of a video classification system for adaptively adjusting an image base model based on the Mamba network in an embodiment;
[0058] Figure 4 Schematic diagram of an electronic device in an embodiment. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment 1
[0061] In view of the problems existing in the foregoing prior art, this embodiment provides a video classification method for adaptively adjusting an image base model based on the Mamba network, aiming to solve the problems of high computational complexity, low efficiency, and difficulty in capturing complete spatio-temporal dynamics in existing methods when processing video data. Refer to Figure 1 and Figure 2 , the method includes the following steps:
[0062] Step S1: Preprocess the video and encode it into long-sequence video features using the image base model.
[0063] Specifically, step S1 may include steps S101 - S103:
[0064] Step S101: Extract frames from the video to obtain a video frame sequence.
[0065] After the video data is input into the system, it needs to be preprocessed to extract key frames and extract features for each video frame. The specific method includes decomposing the video data into frames, extracting each frame image as an independent image unit, and normalizing these images to ensure the consistency and efficiency of the model input.
[0066] Step S102: Use a pre-trained image foundation model to perform image-level encoding on each video frame and add temporal position embedding information to each frame based on its time series.
[0067] Use pre-trained image-based models (such as CLIP, SigLIP, etc.) to perform image-level feature encoding on each frame of the image. The image-based model is a deep network model pre-trained through a large-scale dataset and has powerful feature expression capabilities. After each frame of the image is encoded by the image-based model, a high-dimensional feature vector will be generated.
[0068] Considering the adaptation to the temporal dependencies in video data, temporal position encoding (Temporal Position Embedding) is added to each frame of features. This process can enable the model to effectively distinguish the temporal information of different frames by adding temporal position encoding to each frame of the image, thereby laying the foundation for subsequent spatio-temporal feature modeling. The temporal position encoding can adopt a position encoding method similar to that in the Transformer model, taking the temporal information of the frame as part of the input and embedding it into the features.
[0069] Step S103: Merge the encoding results of all frames to obtain the long-sequence video feature V ∈ R HWT×C 。
[0070] The feature vectors of all video frames are merged into a long-sequence feature matrix V ∈ R HWT×C , where H, W, and T are the spatial dimensions (height and width) and temporal dimension (number of frames) of the video respectively, and C is the dimension of each frame's feature vector. The long-sequence video feature obtained in this way will be used as the input for subsequent processing.
[0071] Step S2: Group the long-sequence video features and calculate the self-correlation features separately within multiple subsequences. Specifically, Step S2 can include Steps S201 - S203:
[0072] Step S201: Split the long-sequence video feature sequence V ∈ R HWT×C using a 2D window of size w*w to obtain multiple subsequences, and the number of its subsequences is N = HWT / w 2 。
[0073] Considering improving the calculation efficiency and reducing the calculation burden of the model, the long-sequence video feature V is split into multiple subsequences. The specific approach is to use a 2D window of size w × w to divide the video feature matrix V into multiple smaller subsequences. The feature dimension of each subsequence remains the same as that of the original video feature. The number of subsequences after splitting is N = HWT / w 2 , making each subsequence a local feature block, thereby being able to reduce the time complexity during each calculation.
[0074] Step S202: For each subsequence, for the $i$-th layer of the image base model, calculate its self-correlation features using the pre-trained self-attention weights of its self-attention layer.
[0075] For each subsequence, the self-attention weights of the $i$-th layer of the image base model are used to calculate the self-correlation features of the subsequence. Since the subsequence is spatially local, using the self-attention mechanism to calculate the self-correlation features can avoid the high computational cost of global calculation while maintaining local features.
[0076] Step S203: Combine the calculation results of all subsequences to obtain the processing result $x\in\mathbb{R}$ of the entire long-sequence video. HWT×C 。
[0077] The self-correlation calculation results of all subsequences are combined to obtain the processing result $x\in\mathbb{R}$ of the entire long-sequence video features, HWT×C and provide the preliminarily processed spatio-temporal features for subsequent adjustment steps.
[0078] Step S3: Process the entire long-sequence video using the Mamba network and modulate the result with the original features using a modulation function.
[0079] To further improve the spatio-temporal adaptability of video features, this method introduces the state space model (SSM) of the Mamba network to adjust the entire long-sequence video features obtained in Step S2. The Mamba network can effectively model long-term dependencies through its state space model and has good computational efficiency and scalability.
[0080] Specifically, Step S3 can include Steps S301 - S302:
[0081] Step S301: Introduce the state space model (state space model, SSM) in the Mamba network as a basic module and adjust the output dimension of its last layer so that it can accept the original video sequence as input and output two different feature sequences as modulation parameters: $y_1, y_2 = \text{SSM}(x)$.
[0082] Adjust the state space model of the Mamba network so that the output dimension of its last layer adapts to the input video sequence. The state space model of the Mamba network accepts the video feature $x$ as input and outputs two feature sequences $y_1$ and $y_2$, which are used as modulation parameters to adjust the original video features.
[0083] Step S302: Design a sequence modulation function (SeqMod) to modulate the original features. Its specific form is: SeqMod(x, y1, y2) = x ⊙ y1 + x + y2, where ⊙ is the Hamilton multiplication symbol.
[0084] Construct a sequence modulation function (SeqMod) to further enhance the expressive power of spatio-temporal information by weighted modulation of the input features. The specific modulation process is as follows:
[0085] SeqMod(x, y1, y2) = x ⊙ y1 + x + y2
[0086] Among them, the symbol ⊙ represents Hamilton multiplication. This operation method can effectively fuse the video features and the modulation parameters output by the Mamba network, inject more spatio-temporal information, and help the model better understand the dynamic changes in video data.
[0087] Step S4: Feed the modulation result into the subsequent forward propagation process of the vision base model, continuously repeat Steps S1 - S4, and finally obtain the feature representation of the current video. Specifically, Step S4 can include Steps S401 - S403:
[0088] Step S401: For the i-th layer of the image base model, use the weights of its feed forward part to process the current sequence and complete the adaptation adjustment of the current model layer.
[0089] The modulated feature x is fed into each layer of the image base model for forward propagation. For each layer, the weights of the feed forward part of this layer are used to adaptively adjust the current feature so that the feature can better integrate into the output of the current layer. This process will be completed layer by layer in turn, gradually adjusting the feature representation of each layer until the features of the entire video pass through the forward propagation of each layer to obtain the final output. The output feature representation of the l-th layer is denoted as y l ∈R HWT .
[0090] Step S402: Continuously repeat Step S2 to Step S401 until the adaptation adjustment for all layers of the current model is completed and the output result y L ∈R HWT .
[0091] The model will not only adapt to video features layer by layer but also inject spatio-temporal information after each layer is completed. This step will continue to iterate until the adjustment process of all model layers is completed, and finally the complete feature representation y of the video is obtained L ∈R HWT×C .
[0092] Step S403: Take the average of the final output result y of the network L in both time and space to obtain the feature representation y of the video o ∈R C .
[0093] For the final output result y L , perform average operations on it in both time and space to obtain the final feature representation y of the video o ∈R C , and this feature representation contains the global spatio-temporal features of the video and can provide effective support for subsequent classification steps
[0094] Step S5: Send the feature representation into a classifier to obtain the class prediction probability distribution, compare it with the true label, and calculate the loss function for training. Specifically, Step S5 can include Steps S501 - S506:
[0095] Step S501: Process according to the language names of all classes to be classified through a pre-trained language model to obtain an untrainable class classifier
[0096] The video feature y after adjustment and forward propagation o is sent into the classifier for class prediction. The classifier first converts the class label into an untrainable class classifier according to the pre-trained language model and in combination with the language name of the video class
[0097] Step S502: Calculate the cosine similarity between the obtained feature and each class representative of the classifier, and obtain the class prediction probability distribution according to the similarity
[0098] According to the video feature y o and the similarity with each class classifier, calculate the cosine similarity, so as to obtain the prediction probability distribution of each class
[0099] Step S503: Compare the prediction probability distribution with the true label, and calculate the cross-entropy as the classification loss function
[0100] Compare the predicted probability distribution with the true label, calculate the cross-entropy loss, and use it as the basic loss function for model training
[0101] Step S504: At the same time, collect the intermediate calculation results of the network forward inference, and compare them with the intermediate calculation results of the basic model without adjustment, and calculate the mean square error between the two as the distillation loss function
[0102] Collect the intermediate calculation results of each layer of the image basic model forward inference in Steps S2 to S4 of the network {y 1 , y 2 , …, y L}, and compare it with the intermediate result of the unadjusted base model, calculate the mean square error between the two as the distillation loss function. This distillation loss can help the model better combine the characteristics of the base model and the Mamba network, and improve its performance in video recognition tasks.
[0103] Step S505: Weightedly sum the classification loss function and the distillation loss function to obtain the final loss function.
[0104] Weightedly sum the classification loss and the distillation loss to obtain the total loss function.
[0105] Step S506: Freeze all the parameters of the base model in the network, only train the parameters of the Mamba network, perform backpropagation on the loss function, and use the gradient descent method to train the network until convergence.
[0106] Train the entire network through backpropagation and the gradient descent method until the model converges.
[0107] In summary, the present method has the following characteristics:
[0108] (1) Improve the training accuracy of video recognition: This method freezes the parameters of most pre-trained image base models. Without changing the structure of the original image base model, it can introduce only the Mamba network module with a small number of parameters to effectively adjust video features. This method can maximize the use of the weights of the image pre-trained model, better capture the spatio-temporal dynamics in the video, and thus significantly improve the accuracy of video recognition.
[0109] (2) Novel video recognition framework: This method provides a unique video recognition framework, including a "partitioning and adjustment" stage. This framework first applies window-based spatial local attention to each layer of the video base model, and then injects complete spatio-temporal information through a modulation function. Such a design can not only improve the feature representation ability but also enhance the generalization ability of the network, making it more adaptable to the changes of different video data.
[0110] (3) Efficient computing method: By introducing the Mamba network and the state space model (SSM), it can achieve efficient processing of long-sequence video features, reduce the computational complexity, and significantly improve the processing speed and memory utilization rate.
[0111] (4) Effective training process: Through the weighted training method of the distillation loss function and the classification loss function, combined with the adjustment of the base model and the Mamba network, it can simultaneously optimize the video feature extraction ability and the learning efficiency of the model, and further improve the performance of the video recognition task.
[0112] Embodiment 2
[0113] On the basis of Embodiment 1, refer to Figure 3, this embodiment provides a video classification system for adaptively adjusting an image base model based on the Mamba network, which is used to implement the video classification method for adaptively adjusting the image base model based on the Mamba network in Embodiment 1. The video classification system includes:
[0114] (1) An image encoding module, configured to obtain a video, and perform feature encoding on each frame image in the video by using a pre-trained image base model to obtain a long-sequence video feature including frame time information.
[0115] (2) An autocorrelation calculation module, configured to group the long-sequence video feature to obtain a plurality of subsequences, and calculate the autocorrelation feature of each subsequence based on the self-attention mechanism.
[0116] (3) A feature modulation module, configured to use the merged autocorrelation feature as the input of the Mamba network to obtain an output feature sequence, and perform sequence modulation on the output feature sequence and the long-sequence video feature to obtain a modulated feature.
[0117] (4) A forward propagation module, configured to perform forward propagation on each layer of the image base model based on the modulated feature, and repeat multiple times to obtain the final feature of the video.
[0118] (5) A training module, configured to freeze the parameters of the image base model, and based on the final feature, use a classifier to obtain a class probability distribution, and train the Mamba network by calculating a loss function to implement the adaptive adjustment of the image base model.
[0119] The present invention provides a new video recognition framework, which can improve the video recognition performance through the adjustment mechanism of the Mamba network without changing the structure of the base model. Compared with traditional methods, the present invention has better spatio-temporal feature modeling ability and computational efficiency, is applicable to large-scale video analysis tasks, and has strong promotion and practical value.
[0120] Embodiment 3
[0121] Based on the foregoing embodiment, this embodiment provides an electronic device, including: one or more processors and a memory, where the memory stores one or more programs, and the one or more programs include instructions for executing the video classification method for adaptively adjusting the image base model based on the Mamba network as described in Embodiment 1.
[0122] As Figure 4 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the aboveFigure 1 The method described above. Of course, in addition to the software implementation, the present invention does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.
[0123] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in forms such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0124] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0125] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A video classification method for adapting and adjusting an image base model based on the Mamba network, characterized in that Classify the target video using the adapted image base model, where the process of obtaining the adapted image base model includes the following steps: Step S1, obtain the video, perform feature encoding on each frame image in the video using the pre-trained image base model, and obtain a long sequence of video features including frame time information; Step S2, group the long sequence of video features to obtain multiple subsequences, and calculate the self-correlation features of each subsequence based on the self-attention mechanism; Step S3, use the merged self-correlation features as the input of the Mamba network to obtain an output feature sequence, and perform sequence modulation on the output feature sequence and the long sequence of video features to obtain the modulated features; Step S4, based on the modulated features, perform forward propagation on each layer of the image base model, and repeat steps S2 - S4 to obtain the final features of the video; Step S5, freeze the parameters of the image base model, based on the final features, use the classifier to obtain the class probability distribution, and train the Mamba network by calculating the loss function to achieve the adaptation adjustment of the image base model.
2. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 1, characterized in that In the above-mentioned step S1, the process of obtaining the long sequence of video features includes the following steps: Step S101, extract frames from the obtained video to obtain a video frame sequence; Step S102, perform image-level encoding on each video frame using the pre-trained image base model, and add position encoding information based on the time series; Step S103, merge the encoding results of all video frames to obtain a long sequence of video features.
3. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 1, characterized in that, In the above-mentioned step S2, the process of calculating the self-correlation features of each subsequence includes the following steps: Step S201, for the long sequence of video features, use a 2D window to split it into multiple subsequences; Step S202, for each subsequence, use the self-attention layer weights pre-trained by the image base model to calculate its self-correlation features; Step S203, merge the calculation results of all subsequences to obtain the merged self-correlation features of the entire long sequence of video.
4. A video classification method for adaptively adjusting an image base model based on a Mamba network according to claim 1, characterized in that In the above-mentioned step S3, the process of obtaining the modulated features includes the following steps: Step S301, configure the state space model in the Mamba network so that it can accept the original video sequence as input and output two different output feature sequences as modulation parameters; Step S302, use the sequence modulation function to perform sequence modulation on the output feature sequence and the long sequence of video features to obtain the modulated features.
5. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 4, characterized in that, The modulation function is: SeqMod(x,y1,y2)=x⊙y1+x+y2 where SeqMod(x,y1,y2) is the modulated feature, x is the long sequence of video features, and y1, y2 are two different output feature sequences output by the state space model.
6. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 1, characterized in that, In the above-mentioned step S4, the process of performing forward propagation on each layer of the image base model based on the modulated features and repeating steps S1 - S4 to obtain the final features of the video includes the following steps: Step S401: For the i-th layer of the image base model, process the current sequence using the weights of its forward propagation part to complete the adaptation adjustment of the current model layer. Step S402: Repeat steps S2 - S401 until the adaptation adjustment for all layers of the current model is completed and the final features of the output video are obtained.
7. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 6, characterized in that, The said step S4 further includes: Step S403: Perform temporal and spatial averaging operations on the final features of the video.
8. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 1, characterized in that, The said step S5 includes the following steps: Step S501: According to the language names of all categories to be classified, process them through a pre-trained language model to obtain an untrainable category classifier. Step S502: Calculate the cosine similarity between the final features of the video and each category representative of the classifier, and obtain the category prediction probability distribution based on the similarity. Step S503: Based on the comparison between the category prediction probability distribution and the true labels, calculate the cross-entropy as the classification loss function. Step S504: Based on the intermediate calculation results in the forward inference in step S4, and compare them with the intermediate calculation results of the unadjusted base model, and calculate the mean squared error as the distillation loss function. Step S505: Based on the classification loss function and the distillation loss function, calculate the final loss function through weighted summation. Step S506: Based on the final loss function, train the Mamba network using the gradient descent method until convergence.
9. A video classification method for adaptively adjusting an image base model based on the Mamba network according to claim 1, characterized in that, The said image base model is CLIP or SigLIP.
10. A video classification system for adaptively adjusting an image base model based on the Mamba network, characterized in that, A video classification system for implementing the video classification method for adaptively adjusting an image base model based on a Mamba network as described in any one of claims 1 - 9, the video classification system includes: An image encoding module, configured to obtain a video, and perform feature encoding on each frame image in the video using a pre-trained image base model to obtain long-sequence video features including frame time information. An autocorrelation calculation module, configured to group the long-sequence video features to obtain multiple subsequences, and calculate the autocorrelation features of each subsequence based on the self-attention mechanism. A feature modulation module, configured to use the merged autocorrelation features as the input of the Mamba network to obtain an output feature sequence, and perform sequence modulation on the output feature sequence and the long-sequence video features to obtain modulated features. A forward propagation module, configured to perform forward propagation on each layer of the image base model based on the modulated features, and repeat multiple times to obtain the final features of the video. A training module, configured to freeze the parameters of the image base model, and based on the final features, use a classifier to reach the category probability distribution, and train the Mamba network by calculating the loss function to achieve the adaptation adjustment of the image base model.
Citation Information
Patent Citations
Weak supervision video time sequence action detection and classification method and system
CN115272941A
Video classification method based on acceleration Transform model
CN114048818A
Video processing method and device based on hybrid expert model, equipment and medium
CN118781518A
Video processing method and device
CN118968373A
Multimode self-supervised mixed Mangbar hyperspectral image classification method
CN119007024A