Video identification and segmentation method based on Mama architecture, storage medium and equipment
By adopting the video recognition and segmentation method based on the Mamba architecture in middle school experimental video recognition and segmentation, the accuracy, timeliness and operability problems in the prior art are solved, and high-precision and high-efficiency experimental step recognition and segmentation are achieved, meeting the real-time needs in teaching scenarios.
Patent Information
- Application Number
- CN202510182567.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art has problems with accuracy, timeliness and operability in the identification and segmentation of middle school experimental videos. Especially when dealing with complex semantic information and long-term videos, it is difficult to meet the real-time requirements in teaching scenarios.
Using a video recognition and segmentation method based on the Mamba architecture, the experimental video is preprocessed, the resolution is reduced and the spatiotemporal features are extracted, combined with an improved time action detector, and the action features of different durations are fused to achieve accurate identification and time segmentation of experimental steps.
It significantly improves the recognition accuracy and robustness of experimental steps, reduces human resource consumption during manual detection, improves analysis and processing efficiency, and meets the real-time needs in teaching scenarios.
Smart Images

Figure CN120107856A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent education technology of artificial intelligence and also belongs to the field of deep learning technology. It relates to the technology of identifying and segmenting experimental videos of middle schools, and mainly relates to a video recognition and segmentation method, storage medium and device based on Mamba architecture. Background Art
[0002] As a key link in cultivating students' scientific literacy, practical ability and innovation ability, experimental teaching occupies an irreplaceable position in the modern education system. It can not only deepen students' understanding of subject knowledge, but also promote students' critical thinking, problem-solving ability and teamwork spirit through hands-on operation. As countries pay more and more attention to STEM (science, technology, engineering and mathematics) education, experimental teaching has become an important part of national curriculum plans and curriculum standards.
[0003] However, in the current experimental teaching practice, especially in middle schools, there are many challenges that affect the quality and effectiveness of teaching. First, due to factors such as large class size and limited teacher resources, traditional manual evaluation methods are difficult to achieve a comprehensive and timely evaluation of students' experimental skills. Secondly, the subjective bias in the manual scoring process may lead to inconsistent scores between different teachers or the same teacher at different time points, which in turn affects the fairness and objectivity of the evaluation results. In addition, with the popularization of online educational resources, more and more teaching activities have begun to be conducted in the form of video recordings. How to effectively use these video materials to assist experimental teaching has become one of the problems that need to be solved urgently.
[0004] In response to the above problems, researchers have begun to explore the use of artificial intelligence technology to improve the evaluation mechanism in the experimental teaching process in recent years. Artificial intelligence technology can provide instant feedback through automatic analysis of students' experimental operation videos, and help establish a standardized evaluation system, thereby significantly improving the quality and efficiency of experimental teaching. However, some existing solutions still face some limitations when applied to middle school experimental videos. For example, many existing action recognition algorithms perform poorly when processing experimental videos containing complex semantic information and continuous long sequences of actions, making it difficult to accurately extract key step information; at the same time, these algorithms are often computationally inefficient and cannot meet the real-time requirements in teaching scenarios.
[0005] In recent years, artificial intelligence technology, especially deep learning models, has made significant progress in image recognition, natural language processing and other fields, and has gradually been applied to the field of education. Among them, the Mamba architecture, as an efficient data processing platform, stands out with its powerful parallel computing capabilities and flexible modular design. The Mamba architecture was originally developed for high-performance computing scenarios, such as large-scale data processing and complex algorithm acceleration. Its features include support for multiple programming languages, easy expansion and optimized memory management, making it very suitable for processing application scenarios containing large amounts of video data. With the development of the Mamba architecture, researchers began to explore its application in video content analysis tasks, especially in situations that require high precision and real-time performance. In practical applications, especially in teaching scenarios, real-time performance is crucial, and previous methods are difficult to meet this requirement, limiting their application effect in actual teaching. Summary of the invention
[0006] The present invention is aimed at the accuracy, timeliness and operability of the recognition and segmentation process of the video in the prior art, and provides a video recognition and segmentation method, storage medium and device based on the Mamba architecture. First, the video to be recognized and segmented is preprocessed to reduce the video resolution and convert it into a standardized format suitable for further analysis. Then, an efficient video backbone network is used to extract the visual spatiotemporal features of the video, and the extracted features are input into the improved time action detector. The time action detector significantly improves the recognition ability of short-term rapid change steps and long-term stable steps by effectively integrating the action features of different durations, and finally outputs the category of each experimental step and the corresponding time segmentation result. The method of the present invention can accurately capture the changes and action features of the experimental steps, significantly improve the accuracy and robustness of the detection model, is both efficient and accurate, and effectively solves the problems of large human resource consumption, unstable detection effect, low efficiency, etc., and has high application value and broad application prospects.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is: a video recognition and segmentation method based on the Mamba architecture, comprising the following steps:
[0008] S1. Data preprocessing: preprocess the input experimental video to be identified and segmented to reduce the video resolution;
[0009] S2, feature extraction: extract spatiotemporal features of the video preprocessed in step S1 through the large-scale video backbone network VideoMAE;
[0010] S3, nonlinear mapping: the spatiotemporal features extracted in step S2 are input into the temporal action detector based on the hollow Mamba, and the features are nonlinearly mapped through the embedding module;
[0011] S4, spatiotemporal information fusion: After the spatiotemporal features of the module are embedded in step S3, spatiotemporal information of different scales is extracted through intra-block parallel components or inter-block parallel components, and the spatiotemporal information of different scales is fused by summing; the intra-block parallel component is based on multi-convolution information fusion inside the void Mamba, and the inter-block parallel component directly fuses the output features of multiple void Mambas;
[0012] S5, identification and segmentation: The multi-scale features fused in step S4 are downsampled in a feature pyramid manner, and the category and time segmentation results of each experimental step are output through the classification head and regression head respectively to achieve identification and segmentation of the experimental steps.
[0013] As an improvement of the present invention, the reduction of video resolution in the step S1 data preprocessing mainly includes: adjusting the short side of the video to 480 pixels, keeping the aspect ratio of the video unchanged, automatically calculating and adjusting the corresponding long side resolution, and ensuring the consistency of the video ratio.
[0014] As another improvement of the present invention, in step S2, before feature extraction, the input video frame is cropped to a resolution of 224×224; during feature extraction, the video is divided into video segments of 16 frames each, the video frame rate is 30 frames per second, and the stride is 4 frames, so that a feature vector is extracted every 4 / 30 seconds.
[0015] As another improvement of the present invention, the workflow of the embedding module in step S3 is as follows: first, the input spatiotemporal features are processed through a 1D convolution layer to extract temporal features; if the input data requires masking, the data is masked before and after the convolution operation, and the role of the mask is to shield or weight specific data positions; after the convolution calculation is completed, a normalization operation is performed, and the normalization layer will process and adjust the dimension of the input data accordingly according to its type; finally, an activation function is applied to introduce nonlinear characteristics.
[0016] As another improvement of the present invention, in step S4, the inter-block parallel component captures features of different scales by parallelizing multiple hole Mamba blocks and adopting a fusion strategy of summation or splicing, wherein the number of convolutions of the hole Mamba is 1;
[0017] The intra-block parallel component captures features of different scales by serially stacking multiple hole Mamba blocks, each hole Mamba block contains N convolution operations, and learns diversified information by parallelizing multiple convolution operations with different expansion rates within the hole Mamba block, and then summing and merging the convolution outputs with different expansion rates; the multiple convolutions share weights.
[0018] As another improvement of the present invention, the hole Mamba block is defined as: Let the input tensor be X∈R B×D×L , where B is the batch size, D is the feature dimension, L is the sequence length, and X represents the output after being processed by the video backbone network and mapped by the embedding module;
[0019] The input X is mapped to a higher dimension through a linear layer, and then divided into four parts along the feature dimension D: X f , Z f , X b , Z b :
[0020]
[0021]
[0022] Among them, X b and Z b is the input for the reverse sequence;
[0023] Feature X f and the reverse sequence of input X b Processed by dilated causal convolution, where X b Reverse along the feature dimension L:
[0024]
[0025] Among them, Conv i Represents a dilated causal convolution with a dilation rate of d i , F is the flip function flip used to reverse the sequence;
[0026] Output D of the dilated causal convolution f and D b After the activation function, it is input into the selective state space model:
[0027] S f ,S b =σ(SSM(D f ,D b ))
[0028] Where σ is the SiLU activation function;
[0029] S f , S b , Z f , Z b Splice, reverse sequence S b and Z b It will be flipped back to the original order and concatenated with the forward sequence, and the final output is mapped back to the original D dimension through a linear layer:
[0030] R=Linear(Concat(S f ,σ(Z f ),F(S b ),σ(Z b )))
[0031] Among them, R represents the output result, and Concat represents the concatenation operation on the feature dimension D.
[0032] As another improvement of the present invention, the atrous causal convolution separates the input time steps according to the dilation factor, and its output is expressed as:
[0033]
[0034] Where d is the dilation factor, k is the size of the convolution kernel, w[k] is the weight of the kth convolution kernel element, x(t) is the input at time t, and p is the padding factor.
[0035] As a further improvement of the present invention, the filling factor p is calculated as follows:
[0036]
[0037] In order to achieve the above-mentioned purpose, the present invention also adopts a technical solution: a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned video recognition and segmentation methods based on the Mamba architecture.
[0038] In order to achieve the above object, the present invention also adopts a technical solution: a device comprising
[0039] A memory for storing instructions;
[0040] The processor is used to execute the instructions so that the device performs the operations of any of the above-mentioned video recognition and segmentation methods based on the Mamba architecture.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention adopts a temporal action detection method based on the Mamba architecture, which can more accurately capture the changes in experimental steps and action characteristics, and significantly improve the accuracy and robustness of the classification model.
[0043] (2) By introducing the void causal convolution and multi-scale information fusion method, the present invention effectively enhances the model's ability to recognize experimental steps of different durations, further improving the accuracy and stability of detection.
[0044] (3) The present invention can reduce the consumption of human resources in the manual detection process, solve the problem that the existing methods do not fully utilize the temporal relationship when processing long time-series videos, and at the same time improve the analysis and processing efficiency. It has high application value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flow chart of the steps of the video recognition and segmentation method based on the Mamba architecture of the present invention;
[0046] Figure 2 It is a comparison diagram of the timing characteristics of the hollow causal convolution and the standard convolution processing in the method of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0048] Example 1
[0049] The video recognition and segmentation method based on the Mamba architecture is mainly used for the segmentation and recognition of experimental steps in offline middle school experimental videos, such as Figure 1 As shown, the following steps are included:
[0050] Step S1: Preprocess the input middle school experiment video, refer to the common temporal action localization method, slightly reduce the video resolution to reduce memory usage and speed up processing, while maintaining sufficient visual information for subsequent analysis.
[0051] The preprocessing step at least includes scaling the short side of the video to adjust the short side to 480 pixels while keeping the aspect ratio of the video unchanged, automatically calculating and adjusting the corresponding long side resolution, and ensuring the consistency of the video content ratio.
[0052] Step S2: Use the large-scale video backbone network VideoMAE to extract the spatiotemporal features of the preprocessed video. During the extraction process, feature vectors are extracted in units of video segments to ensure efficient and fine-grained spatiotemporal feature representation.
[0053] When extracting spatiotemporal features, the video is divided into 16-frame segments, with a frame rate of about 30 frames per second and a stride of 4 frames, so that a feature vector is extracted every 4 / 30 seconds (about 0.1333 seconds). Before extracting features, the input video frames are cropped and adjusted to a resolution of 224×224 to ensure a standardized input size and efficient and fine-grained representation of spatiotemporal features.
[0054] Step S3: Figure 1As shown on the left, the spatiotemporal features extracted in step S2 are input into the temporal action detector based on the hollow Mamba, and the features are nonlinearly mapped through the embedding module to adapt to the subsequent multi-scale information processing.
[0055] The workflow of the embedding module is as follows: first, the input spatiotemporal features are processed through a 1D convolutional layer to extract temporal features; if the input data needs to be masked, the data is masked before and after the convolution operation. The role of the mask is to shield or weight specific data positions; after the convolution calculation is completed, the normalization operation is performed. The normalization layer will process and adjust the dimension of the input data accordingly according to its type; finally, the activation function is applied to introduce nonlinear characteristics.
[0056] Step S4: The spatiotemporal features of the embedded module extract spatiotemporal information of different scales through intra-block parallel components or inter-block parallel components, and then fuse the spatiotemporal information of different scales through a simple summation method. The intra-block parallel components and inter-block parallel components represent two different combination methods. The former relies on the fusion of multi-convolution information inside the void Mamba, and the latter directly fuses the output features of multiple void Mambas.
[0057] like Figure 1 As shown on the left, the two components are implemented in different combinations, specifically divided into internal multi-convolution parallelism and external multi-hole Mamba parallelism. The intra-block parallel component is more suitable for detecting scenes with low overlap, providing higher recall rate and reducing the possibility of missed detection. The inter-block parallel component is more adaptable to the complexity of detection and provides more accurate positioning.
[0058] Inter-block parallel component: This component captures features of different scales by parallelizing multiple hole Mamba blocks and adopting simple fusion strategies (such as summation or splicing). This process is called feature fusion layer. In this process, the splicing method needs to introduce additional linear layers to reduce the dimension, making it complicated. In contrast, direct summation can provide higher efficiency. In order to further improve the ability of action localization, we pass the fused features to a multi-layer feature pyramid. The number of convolutions of the hole Mamba in this component is 1.
[0059] Intra-block parallel component: Unlike inter-block parallel components, intra-block parallel components are implemented by stacking multiple atrous Mamba blocks in series, which makes its structure more concise. Unlike inter-block parallel components, each atrous Mamba block in the intra-block parallel component contains N convolution operations. This component learns diverse information by parallelizing multiple convolution operations with different dilation rates inside the atrous Mamba block, thereby enhancing the representation ability of the module. Subsequently, a simple summation method is used to effectively fuse the convolution outputs with different dilation rates, thereby achieving a better balance between receptive fields. Multiple convolutions in this method share weights, so the performance is improved without increasing the number of parameters.
[0060] like Figure 1 As shown on the right, the core module of the temporal action detector in step S4, Mamba, is formally defined as: Let the input tensor be X∈R B×D×L , where B is the batch size, D is the feature dimension, and L is the sequence length. X represents the output after being processed by the video backbone network and mapped by the embedding module.
[0061] First, the input X is mapped to a higher dimension through a linear layer, and then it is divided into four parts along the feature dimension D: X f , Z f , X b , Z b , the specific operations are as follows:
[0062]
[0063]
[0064] Among them, X b and Z b is the input for the reverse sequence.
[0065] Next, feature X f and the reverse sequence of input X b Processed by dilated causal convolution, where X b Reverse along the feature dimension L. This operation aims to capture the temporal features from the forward and reverse sequences, expressed as follows:
[0066]
[0067] Among them, Conv i Represents a dilated causal convolution with a dilation rate of d i , and F is the flip function flip used to reverse the sequence. Each convolution operation has its own expansion rate d i , allowing the model to capture the temporal dependence of changes at different time scales.
[0068] Output D of the dilated causal convolution f and D b After the activation function is used to enhance nonlinearity, it is input into the selective state space model (SSM). The selective state space model follows the SSM in the original Mamba. This process can be expressed as:
[0069] S f ,S b =σ(SSM(D f ,D b ))
[0070] Where σ is the SiLU activation function.
[0071] Finally, S f , S b , Z f , Z b (The result of the activation function) is spliced. Reverse sequence S b and Z b It will be flipped back to the original order and concatenated with the forward sequence, and the final output is mapped back to the original d dimension through a linear layer. The process can be expressed as:
[0072] R=Linear(Concat(S f ,σ(Z f ),F(S b ),σ(Z b )))
[0073] Among them, R represents our output result, and Concat represents the concatenation operation on the feature dimension D.
[0074] like Figure 2 As shown in Figure 1, the difference between the dilated causal convolution in step S4 and the standard convolution in Mamba is that the dilated causal convolution separates the input time steps according to the dilation factor, allowing the model to cover a wider range of time while maintaining causality. Formally, the output of the dilated causal convolution can be expressed as:
[0075]
[0076] Where d is the dilation factor, k is the size of the convolution kernel, w[k] is the weight of the kth convolution kernel element, and x(t) is the input at time t. The constant p is called the filling factor and is calculated as:
[0077]
[0078] The padding factor p is used to pad the input x before the convolution operation to ensure the correct alignment during the convolution operation. In contrast, the standard causal convolution processes continuous time steps, so its receptive field is significantly smaller. This feature makes the atrous causal convolution particularly suitable for processing long-lasting actions.
[0079] Step S5: Downsample the fused multi-scale features in a feature pyramid manner, and output the category and time segmentation results of each experimental step through the classification head and regression head, so as to achieve accurate identification and segmentation of the experimental steps.
[0080] In this embodiment, the common temporal action localization method is followed in the preprocessing step, and the Adam method is used for stochastic gradient descent, the batch size is configured to 2, the learning rate is set to 0.0002, and the weight decay is set to 0.05. In the embodiment, end-to-end training can be performed directly or a separation method can be used to extract features first and then train the temporal action detector.
[0081] Test Case
[0082] The method of the present invention was applied to a carbon dioxide experimental data set for experiments. The carbon dioxide experimental data set contains 222 groups of videos and 11 steps of producing carbon dioxide from limestone. Due to the complexity of the experimental content and the different durations of each action, the accuracy and sensitivity of recognition and segmentation are challenged. In order to measure the gap between the segmentation results and the actual results, this test example uses the mean average precision (mAP) as the main evaluation indicator for evaluating middle school experimental tasks. mAP is a comprehensive evaluation method that measures the overall performance of the model in action classification and positioning by taking a weighted average of the mean accuracy (AP) of each action category. The higher the mAP value, the better the model performs in identifying action categories and accurately locating action time periods.
[0083] In this test example, the method of the present invention is compared with the most popular temporal action detector. The main comparison indicators include: the average mAP under the frozen video backbone network, the number of parameters of the temporal action detector, and the average training time of the network. The comparison results are shown in Table 1:
[0084] Table 1
[0085] method Average mAP Parameter quantity Average training time Actionformer 76.8 27.6M 10.5S The present invention 79.3 19.8M 7.5S
[0086] As can be seen from the table above, by comparing the average mAP, the present invention has more advantages in action recognition accuracy; at the same time, by comparing the number of model parameters and training time, the present invention has advantages in training efficiency and reasoning efficiency. It can also be seen from the table that the detection accuracy of the present invention on the carbon dioxide dataset is better than that of the popular methods, and it can also achieve significant optimization in parameter efficiency and training time, making it more suitable for practical applications.
[0087] In summary, the video recognition and segmentation method based on the Mamba architecture disclosed in the present invention uses an efficient video backbone network to extract the visual spatiotemporal features of the video, and inputs the extracted features into an improved time action detector, which significantly improves the recognition ability of short-term rapid change steps and long-term stable steps by effectively fusing action features of different durations, and finally realizes the accurate classification and time segmentation of each experimental step in the experimental video. When processing complex experimental video data, the method of the present invention can significantly improve the recognition accuracy and robustness of the experimental steps, effectively solve the problems of large human resource consumption, unstable detection effect, low efficiency, etc., and greatly improve the automation level and evaluation accuracy of experimental teaching. This method has high application value and broad application prospects, especially in the fields of middle school experimental teaching and education evaluation, and is more practical.
[0088] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0089] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.
Claims
1. A video recognition and segmentation method based on the Mamba architecture, characterized in that: The steps include: S1. Data preprocessing: preprocess the input experimental video to be identified and segmented to reduce the video resolution; S2, feature extraction: extract spatiotemporal features of the video preprocessed in step S1 through the large-scale video backbone network VideoMAE; S3, nonlinear mapping: the spatiotemporal features extracted in step S2 are input into the temporal action detector based on the hollow Mamba, and the features are nonlinearly mapped through the embedding module; S4, spatiotemporal information fusion: After the spatiotemporal features of the module are embedded in step S3, spatiotemporal information of different scales is extracted through intra-block parallel components or inter-block parallel components, and the spatiotemporal information of different scales is fused by summing; the intra-block parallel component is based on multi-convolution information fusion inside the void Mamba, and the inter-block parallel component directly fuses the output features of multiple void Mambas; S5, identification and segmentation: The multi-scale features fused in step S4 are downsampled in a feature pyramid manner, and the category and time segmentation results of each experimental step are output through the classification head and regression head respectively to achieve identification and segmentation of the experimental steps.
2. The video recognition and segmentation method based on Mamba architecture as claimed in claim 1, characterized in that: The step S1 of reducing the video resolution in the data preprocessing mainly includes: adjusting the short side of the video to 480 pixels, keeping the aspect ratio of the video unchanged, and automatically calculating and adjusting the corresponding long side resolution to ensure that the video ratio is consistent.
3. The video recognition and segmentation method based on Mamba architecture as claimed in claim 1, characterized in that: In step S2, before feature extraction, the input video frame is cropped to a resolution of 224×224; during feature extraction, the video is divided into video segments of 16 frames each, the video frame rate is 30 frames per second, and the stride is 4 frames, so that a feature vector is extracted every 4 / 30 seconds.
4. The video recognition and segmentation method based on Mamba architecture as claimed in claim 1, characterized in that: In the embedding module of step S3, the input spatiotemporal features are used to extract temporal features through a 1D convolutional layer; After the convolution is completed, normalization is performed, and finally, an activation function is applied to introduce nonlinear characteristics.
5. The video recognition and segmentation method based on Mamba architecture as claimed in claim 1, characterized in that: In step S4, the inter-block parallel component captures features of different scales by parallelizing multiple hole Mamba blocks and adopting a fusion strategy of summation or splicing, wherein the number of convolutions of the hole Mamba is 1; The intra-block parallel component captures features of different scales by serially stacking multiple hole Mamba blocks, each hole Mamba block contains N convolution operations, and learns diversified information by parallelizing multiple convolution operations with different expansion rates within the hole Mamba block, and then summing and merging the convolution outputs with different expansion rates; the multiple convolutions share weights.
6. The video recognition and segmentation method based on Mamba architecture as claimed in claim 4, characterized in that: The hole Mamba block is defined as: Let the input tensor be X∈R B×D×L , where B is the batch size, D is the feature dimension, L is the sequence length, and X represents the output after being processed by the video backbone network and mapped by the embedding module; The input X is mapped to a higher dimension through a linear layer, and then divided into four parts along the feature dimension D: X f , Z f , X b , Z b : Among them, X b and Z b is the input for the reverse sequence; Feature X f and the reverse sequence of input X b Processed by dilated causal convolution, where X b Reverse along the feature dimension L: Among them, Conv i Represents a dilated causal convolution with a dilation rate of d i , F is the flip function flip used to reverse the sequence; the output of the void causal convolution D f and D b After the activation function, it is input into the selective state space model: S f ,S b =σ(SSM(D f ,D b )) Where σ is the SiLU activation function; S f , S b , Z f , Z b Splice, reverse sequence S b and Z b It will be flipped back to the original order and concatenated with the forward sequence, and the final output is mapped back to the original D dimension through a linear layer: R=Linear(Concat(S f ,σ(Z f ),F(S b ),σ(Z b ))) Among them, R represents the output result, and Concat represents the concatenation operation on the feature dimension D.
7. The video recognition and segmentation method based on Mamba architecture as claimed in claim 5, characterized in that: The dilated causal convolution separates the input time steps according to the dilation factor, and its output is expressed as: Where d is the dilation factor, k is the size of the convolution kernel, w[k] is the weight of the kth convolution kernel element, x(t) is the input at time t, and p is the padding factor.
8. The video recognition and segmentation method based on Mamba architecture as claimed in claim 6, characterized in that: The filling factor p is calculated as follows:
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the video recognition and segmentation method based on the Mamba architecture as described in any one of claims 1 to 8 is implemented.
10. A device, characterized in that: include A memory for storing instructions; The processor is used to execute the instruction so that the device performs the operation of the video recognition and segmentation method based on the Mamba architecture as described in any one of claims 1-8.
Citation Information
Patent Citations
Two-channel interaction time convolution network, close-range video action segmentation method, computer system and medium
CN113537232A
Differential feature enhancement-based multi-view middle school experiment step detection method and system
CN119478754A