A video object detection method based on motion vector guided local attention
By employing a motion vector-guided local attention method, the problems of high computational cost and low detection accuracy in video object detection are solved, achieving efficient video object detection and improving detection accuracy and consistency.
Patent Information
- Application Number
- CN202510143847.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing video object detection methods suffer from high computational cost, information redundancy, and low detection accuracy when processing video frames. In particular, the inconsistency between the semantic information of low-order and high-order features affects the model's ability to understand the target, increasing the likelihood of false positives and false negatives.
A motion vector-guided local attention video target detection method is adopted. The motion vector prediction network guides the local attention direction, uses features from adjacent frames to propagate to the current frame, and fuses the transmitted features with the current frame through a feature fusion module, thereby reducing the amount of computation and training cost, while improving the detection accuracy.
It significantly reduces computational load and training costs, improves detection accuracy and efficiency, enhances the model's ability to identify and locate targets, and improves the consistency and accuracy of detection results.
Smart Images

Figure CN120071217B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of target detection, and more particularly, to a video target detection method based on motion vector guided local attention. BACKGROUND
[0002] The purpose of target detection is to locate single or multiple targets and identify their categories. This process has an important influence on subsequent tasks such as target tracking, instance segmentation, behavior recognition and image description. In static images, target detection mainly focuses on identifying the category and location of the target. However, compared with static images, videos provide more abundant visual information, especially in the real world, where monitoring systems, vehicle cameras, social media and wearable devices mostly deal with video data. Videos not only contain the information of static images, but also add the time dimension. Directly applying static image target detection methods to videos for frame-by-frame detection may face some challenges. For example, moving targets may have problems such as occlusion, blurring and defocusing, which affect the detection performance. In addition, there is high similarity between video frames, and there is temporal redundancy. Therefore, in the task of video target detection, the use of time correlation specific to videos to predict inter-frame changes can significantly improve detection efficiency. Therefore, it is of great practical significance to study how to make full use of the temporal information in videos to improve the performance of video target detection on the basis of static image target detection methods.
[0003] In recent years, with the rapid development of deep convolutional neural networks, video target detection methods based on deep convolutional neural networks have made significant achievements and gradually become the main solution in the field of video target detection. This video target detection method integrates temporal information to improve detection performance on the basis of using static image target detection methods. FGFA(Flow-guided feature aggregation) is a flow-guided feature fusion model that uses optical flow information to align the features of multiple adjacent frames to the current frame and adaptively weights the aligned multi-frame features to compensate for the current frame features. LWDN(Locally-weighted deformable neighbors) proposes an adaptive position-sensitive feature propagation method that uses shallow feature maps as input to a weight prediction network to generate position-dependent weights and offsets. Kernel weights are used to determine the weighting of features, while offsets are used to adjust the convolution. The LSTS(Learnable Spatio-Temporal Sampling) method applies deep networks to extract features in key frames, while shallow networks are used for feature extraction in non-key frames. At the same time, this method propagates a memory feature and samples feature points from other frames during feature alignment to establish spatial correspondence between inter-frame features.
[0004] The application discloses a video target detection method based on motion vector guided local attention. In the method, the semantic information inconsistency between low-order features and high-order features affects the model's ability to understand the target, reduces the detection accuracy, increases the misjudgment and missed judgment, and affects the feature fusion effect. The learning sampling point position proposed in the LSTS needs to predict the offset of each sampling point, and the offset needs to be updated through back propagation during training, which increases the calculation complexity and the difficulty of model training. Accordingly, the application provides a video target detection method based on motion vector guided local attention. In order to avoid the semantic information inconsistency between low-order features and high-order features, the framework adopts a recursive video target detection method without key frames, guides the local attention direction through a motion vector prediction network, does not need an additional training stage or a dataset pre-training, thereby propagating the features of adjacent frames to the current frame by using the motion vector, reducing the calculation amount and the training cost, and simultaneously fusing the transferred features with the current frame through a feature fusion module, and transferring the fusion result to the next frame. SUMMARY
[0005] Therefore, the application embodiment provides a video target detection method based on motion vector guided local attention to realize the target detection of a video object.
[0006] In order to realize the above purpose, the application embodiment provides the following scheme.
[0007] A video target detection method based on motion vector guided local attention, characterized in that it comprises the following six steps.
[0008] Step 1. Feature extraction
[0009] The feature extraction network used in the application is a ResNet101 network. According to the size of the feature map extracted by the feature extraction network, the feature extraction network is divided into five layers, wherein the feature map output by layer1 has a size of 1 / 2 of the original picture and a channel number of 64, the feature map output by layer2 has a size of 1 / 4 of the original picture and a channel number of 128, the feature map output by layer3 has a size of 1 / 8 of the original picture and a channel number of 256, the feature map output by layer4 has a size of 1 / 16 of the original picture and a channel number of 512, and the feature map output by layer5 has a size of 1 / 32 of the original picture and a channel number of 1024.
[0010] Step 2. Predicting the motion vector
[0011] The current frame image and the previous frame image are stacked in the channel dimension to input into a motion vector prediction network to predict the motion vector. The network is composed of four groups of convolutions, wherein the first group and the second group of convolutions are respectively composed of convolution layers with channel numbers of 32 and 64, the third group of convolutions is a multi-scale feature fusion layer composed of four convolution layers with expansion rates of 1, 2, 4 and 8, and the fourth group of convolutions predicts the motion vector through a convolution layer with a convolution kernel size of 3*3, and finally obtains a motion vector map with the same resolution as the feature map.
[0012] Step 3. Feature propagation;
[0013] When detecting the current frame, if the current frame is the initial frame of the video frame, image target detection is used for the frame. At the same time, the memory feature is initialized as the feature map of the frame for subsequent feature propagation; if the current frame is not the initial frame of the video frame, the feature extraction network is used to extract the feature, the motion vector prediction network is used to predict the motion vector, and the memory feature is propagated from the previous frame to the current frame in a manner of guiding local attention by the motion vector.
[0014] Step 4. Feature fusion;
[0015] After the memory feature is propagated to the current frame, it needs to be fused with the current frame. The memory feature can compensate the current feature and alleviate the problems such as blur and occlusion that may occur in the current frame. At the same time, the introduction of the current frame feature makes the propagated feature better adapt to the changes of the current scene.
[0016] Step 5. Obtain the detection result;
[0017] The enhanced feature map is input into the target detection task network to obtain the detection result, that is, after extracting the candidate regions through the RPN network and extracting the candidate region features through the ROI pooling, two fully connected layers with a channel number of 1024 are used to further extract the candidate region features, and then fully connected layers with channel numbers of C+1 and 4 are respectively used to predict the confidence that the region of interest belongs to a specific category and the more accurate target position.
[0018] Step 6. Determine whether the detection is completed;
[0019] Determine whether the image sequence is detected, if the detection is completed, output the detection result of the image sequence, if not, return to step 1.
[0020] Preferably, the algorithm adopts a motion vector guided local attention manner for feature alignment.
[0021] Preferably, the algorithm is trained for a total of 240,000 iterations (i.e., max_iteration is 240,000).
[0022] Preferably, the learning rate of the algorithm is 2.5x10 -4 for the first 160,000 iterations during training, and 2.5x10 -5 for the last 80,000 iterations.
[0023] Preferably, the algorithm uses NMS with an IoU threshold of 0.5 to suppress duplicate detection boxes.
[0024] The method effectively propagates the features of adjacent frames to the current frame by introducing a mechanism of motion vector guided attention. Motion vectors provide key information about inter-frame motion. Coarse-grained motion vectors can effectively limit the scope of attention calculation, thereby significantly reducing the computational load and training cost of attention, and further improving the overall detection efficiency. Motion vectors combined with attention replace the optical flow-based deformation operation for semantic information propagation. Compared with the optical flow deformation operation method, the network is smaller, and without additional training phase and dataset pre-training, the method avoids the noise and errors that may occur in optical flow calculation. To further enhance the integration effect of features, the method proposes a feature fusion module, which is specifically used for fusing inter-frame features and combining the features of the current frame with the features transferred from adjacent frames. This fusion not only preserves important information of the current frame, but also introduces context information from adjacent frames. The fused features are then passed to the detection task network, enhancing the model's ability in target recognition and positioning, and also passing the fusion results to the next frame to form memory features. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0026] Figure 1 Faster R-CNN as a single-frame target detection model structure diagram;
[0027] Figure 2 Faster R-CNN as a single-frame target detection model structure diagram;
[0028] Figure 3 Faster R-CNN as a single-frame target detection model structure diagram;
[0029] Figure 4 Faster R-CNN as a single-frame target detection model structure diagram;
[0030] Figure 5 A schematic diagram of the motion vector guided local attention alignment method provided by the embodiments of the present application;
[0031] Figure 6 A schematic diagram of the feature fusion module provided by the embodiments of the present application;
[0032] Figure 7 The training process of the motion vector guided local attention video object detection method provided by the embodiments of the present application. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.
[0034] A video is composed of a sequence of continuous images, so the video object detection can directly use the image object detection method for frame-by-frame detection. However, the frame-by-frame detection not only has a huge amount of calculation, but also does not consider the temporal information of the video frames, so it is difficult to give consistent detection results before and after. To solve the above problems, the video object detection method usually uses the features of adjacent frames to improve the accuracy and consistency of the current frame object detection result. However, densely fusing the features of multiple frames in the adjacent range will increase the amount of calculation and cause information redundancy, in addition, it is also difficult to use the temporal information in the frames farther than the adjacent frames.
[0035] The present application is based on a single-frame object detection model Faster R-CNN, as shown in Figure 1 , the Faster R-CNN is divided into a feature extraction part and an object detection task network, wherein the object detection task network includes an RPN network and a detection network. Therefore, the present application proposes a motion vector guided local attention video object detection, and the model construction is as shown in Figure 2 , a recursive fusion video object detection model. This model uses a simple motion vector prediction network to predict the motion vector for guiding the position of attention calculation, and aligns the features of adjacent frames to the current frame. At the same time, this model also uses the aligned feature map of the previous frame and the current frame feature map for inter-frame feature fusion, and transmits the fusion result to the next frame. Since the predicted is a coarse-grained motion vector, the motion vector guided attention can limit its calculation range, and this design significantly reduces the calculation amount and training cost of attention, thereby improving the overall detection efficiency. This method not only improves the detection accuracy, but also optimizes the use of computing resources.
[0036] AsFigure 2 As shown, the T-th frame in the video is the current frame to be detected, and the T-1-th frame is the previous frame of the current frame. enhance The memory feature is a feature propagated on the time axis, and is used to compensate the features extracted from the previous frame and the current frame. enhance The initial value of the memory feature is the feature of the first frame in the video sequence, and is updated after being propagated to the current frame and fused with the feature of the current frame each time, so that the memory feature contains semantic information of all previous frames of the current frame. When the current frame is detected, the feature F T The entire ResNet101 network is used as the feature extraction network, and the current frame image and the previous frame image are used as the input of the motion vector prediction network to predict the motion vector, and the motion vector is used to guide the alignment of the position calculated by the local attention to obtain the feature F propagated to the current frame, and fused with the current feature map F T to enhance the current feature, and input to the RPN and the detection network for subsequent operations. Meanwhile, the fused feature F enhance is updated.
[0037] The embodiment of the application discloses a video object detection method based on motion vector guided local attention, which is used to complete the target detection of the video object. Referring to Figure 3 The above method comprises at least the following six steps.
[0038] Step 1. Feature extraction
[0039] The feature extraction network used in the application is the ResNet101 network. The image is first input into the feature extraction network in the method, and the feature extraction network is divided into five layers according to the size of the feature map extracted by the feature extraction network. The feature map output by layer1 has a size of 1 / 2 of the original image and 64 channels, the feature map output by layer2 has a size of 1 / 4 of the original image and 128 channels, the feature map output by layer3 has a size of 1 / 8 of the original image and 256 channels, the feature map output by layer4 has a size of 1 / 16 of the original image and 512 channels, and the feature map output by layer5 has a size of 1 / 32 of the original image and 1024 channels.
[0040] Step 2. Predicting the motion vector
[0041] The current frame image and the previous frame image are stacked in the channel dimension and input into the motion vector prediction network to predict the motion vector. The motion vector prediction network is as shown in Figure 4As shown, the network consists of four groups of convolutions, where the first and second groups of convolution layers have channel numbers of 32 and 64, respectively, a convolution kernel size of 3x3, and a stride of 2. The third group of convolutions is a multi-scale feature fusion layer, which consists of four convolution layers with dilation rates of 1, 2, 4, and 8, respectively, each with a kernel size of 3x3, a number of 64, and a stride of 1. The feature maps (dilated_conv_feat1, dilated_conv_feat2, dilated_conv_feat3, dilated_conv_feat4) generated by these convolutions are gradually fused by pixel-wise addition to obtain fusion_feat1, fusion_feat2, and fusion_feat3. Next, the three fused features are stacked with dilated_conv_feat1 in the channel dimension, and a 1x1 convolution layer is used to reduce the dimension to generate the third group of convolution features. The fourth group of convolutions uses a 3x3 convolution layer to predict the motion vector, and the output is down-sampled by a 2x2 pooling layer to obtain a motion vector map with the same resolution as the key frame deep feature. Compared with the optical flow network, this network is smaller, does not require an additional training phase and dataset pre-training, and focuses on guiding the direction of local attention.
[0042] Step 3. Feature propagation
[0043] In the video target detection network proposed in the present application, since the correspondence between the object positions in the inter-frame can be directly found through the motion vector, it is not necessary to calculate the correlation between each query feature point and all key feature points according to the calculation method of the Transformer, nor is it necessary to align the feature map to the current frame by using the time-consuming optical flow. In the process of feature propagation, each query feature point in the present application only needs to calculate the correlation with C key feature points under the guidance of the motion vector, where C is a constant, so the time complexity is only O(HW), which greatly reduces the calculation amount and training cost of attention, and further improves the overall detection efficiency of the network.
[0044] As shown in the figure, this module defines the high-order feature map of the current frame as the query feature map (Q), and the memory feature passed from the previous frame as the key feature map (K) and the value feature map (V). Unlike ordinary attention, the motion vector guided attention takes all position query points in the query feature map (Q) as a query set According to the motion vector information, the corresponding keys of the query points on the key feature map K form a key set where N=HxW represents the total number of feature map positions, D represents the number of feature map channels, and C represents the number of keys corresponding to one query point. For example, in Figure 5 According to the motion vector information, the plane at the upper right corner of the T frame is located at the central position of the (T-1) frame, so the feature map FT The key set corresponding to the query point in the upper right corner is the feature in the central window of the T-1 frame enhanced feature map. Since each query point in the query set only calculates the correlation with the corresponding C (C equals 4r 2 and r is a constant representing the attention radius) keys according to the motion vector information, the calculation complexity of attention will be reduced from O(H 2 W 2 to O(HW).
[0045] The specific process of attention calculation is divided into three steps:
[0046] 1) Calculate the score using dot product similarity:
[0047]
[0048] In the formula, φ(·) represents the embedding layer, Mv x and Mv y represent the motion vectors in the x and y directions, i and j represent the sampled positions in the key feature map, i, j ∈ (-r, r), and r is the attention radius.
[0049] 2) Weight normalization:
[0050]
[0051] In the formula, w x,y is a value between 0 and 1, and the sum of all weights in the attention radius is 1.
[0052] 3) Weighted summation:
[0053]
[0054] In the formula, F′ T represents the alignment of memory features to the features of the current frame by motion vector guided attention.
[0055] Step 4. Feature fusion
[0056] In the feature propagation process, in order to solve the problems such as blur and occlusion that may exist in the current frame, feature fusion is introduced for compensation. As shown in Figure 6 This module mainly consists of three convolutions. Through convolution and softmax operation, a spatial position weight can be generated according to the result of splicing two feature maps , which is split in the channel dimension to obtain and The weights are respectively passed through the new feature maps and the feature maps extracted from the current frame Weighted summation is performed at spatial locations to reduce the impact of unreliable semantic representations on the segmentation of the current frame. This process can be represented by the following formula:
[0057]
[0058] In the formula, F′ T F is the new feature map output by feature propagation. T It is the frame feature map extracted by the feature extraction network from the current frame. This is the enhanced feature map of the current frame. ⊙ represents the Hadamard product, which is the element-wise multiplication of two matrices at the same position. Although the shapes of matrices W1 and W2 are similar to F... T and F′ T The weight matrix is not related in the first dimension, but before performing the Hadamard product, it is automatically copied and expanded along the first dimension to obtain a matrix of shape C×H×W to ensure that the calculation will not be wrong.
[0059] Step 5. Obtain the test results
[0060] The enhanced feature map is input into the object detection task network to obtain the detection results. That is, after extracting candidate regions through the RPN network and extracting candidate region features through ROI pooling, two fully connected layers with 1024 channels are used to further extract features from the candidate regions. Then, fully connected layers with C+1 and 4 channels are used to predict the confidence of the region of interest belonging to a specific category and the more accurate target location.
[0061] Step 6. Determine if the detection is complete.
[0062] Determine whether the image sequence detection is complete. If the detection is complete, output the detection result of the image sequence. If not, return to step 1.
[0063] During the model training phase, such as Figure 7 As shown, three frames I are selected from the video frames. T0 I T1 I T2 , where I T0 As a reference frame preceding the current frame to be detected, I T1 As the current frame to be detected, I T2 As a reference frame following the current frame to be detected. Selected at I T1 Any frame prior to and within K / 2 frames of it is taken as I T0 Selected in I T1 Then, any frame within K / 2 frames is taken as I. T2 For these three images I T0 I T1 I T2extracted features F T0 , F T1 , F T2 , the motion vectors Mv T0 and Mv T2 of T0 to T1 and T2 to T1 are respectively predicted by using a motion vector prediction network T0 , F T2 are propagated to the aligned features F′ T1 and F′ T0 with multi-frame context information of I T2 by using a motion vector guided local attention alignment method T1 , the aligned features are fused with F T1 to obtain the final enhanced features of I The final image features are sent to the subsequent RPN and detection network as the image features to be detected, to generate the bounding box regression score and the class confidence, and to calculate the loss function to update the network parameters.
[0064] Preferably, the algorithm sets the attention radius parameter r to 3 when aligning. The parameters of the ResNet-101 part of the feature extraction network are also pre-trained on the ImageNet dataset, and the parameters of the remaining convolutional layers are initialized by Gaussian distribution. In the training phase, the input image is resized so that the shorter side is 600 pixels. The network parameters are updated using the SGD algorithm. Each GPU contains one batch, and the number of images per batch is set to 2. The model is trained for a total of 240K iterations, and in the first 160K iterations and the last 80K iterations, the learning rates are 2.5x10 -4 and 2.5x10 -5 respectively. In the inference phase, NMS with a 0.5 IoU threshold is used to suppress duplicate detection boxes.
[0065] The model is verified by using ImageNet VID dataset and ImageNet DET dataset. The ImageNet DET dataset consists of 200 target classes, wherein the training set includes 450K images, the verification set includes 20K images, and the test set includes 40K images. The ImageNet VID dataset contains 30 target classes for target detection in videos. The training set, the verification set and the test set respectively include 3862, 555 and 937 completely labeled video clips. These 30 target classes are a subset of a part of the classes in the ImageNet DET image dataset. The model is first trained on the corresponding subsets of the ImageNet VID dataset and the ImageNet DET dataset. Since the true value labels of the ImageNet VID test set are not disclosed, the accuracy of the detection result cannot be evaluated, so the model performance is evaluated on the verification set of the ImageNet VID.
[0066] The mean average precision (mAP) and the algorithm running time are used as quantitative evaluation indicators. The mean average precision (mAP) is the mean value of the average precision of all classes, and the larger the value is, the better. The algorithm running time is the time required for the network model to process a frame of video, and the smaller the value is, the better. Table 1 quantitatively compares the mAP and the average detection time of the algorithm of the method of the present application and the TCN, TPN+LSTM, LWDN, OGEMENT, DFF, FGFA and LSTS models on the ImageNET VID verification set. The results of the TCN model, the TPN+LSTM model, the LWDN model and the OGEMENT model are published in their papers, and the results of the DFF model, the FGFA model and the LSTS model are obtained by running the published paper source code under the same experimental environment and under the experimental setting of the original paper. The mAP of the method proposed in the present application is 77.07%, and the average detection time is 91.6ms. The mAP is higher than that of other models in the table, which is 29.57% higher than that of the TCN model, 8.53% higher than that of the TPN+LSTM model, 0.77% higher than that of the LWDN model, 0.27% higher than that of the OGEMENT model, 8.62% higher than that of the DFF model, 2.21% higher than that of the FGFA model, and 4.54% higher than that of the LSTS model.
[0067] Table 1 Comparison of the method disclosed in the present application and other methods in mAP and Times
[0068] Table 1 Experimental results of the method disclosed in the present application
[0069]
[0070] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Modifications of these embodiments will occur to persons of skill in the art, and, while certain embodiments within the scope of the application are submitted as examples, it is the intent of the inventors that any equivalents of the subject matter described herein be included in the scope of the application. The disclosure is not intended to be limited to the disclosed embodiments, but wants to include any modification of the subject matter as long as it is within the scope of the relative patent claims.
Claims
1.A method for video object detection based on motion vector guided local attention, characterized in that, Comprise the following 6 steps: Step 1. Feature extraction; The feature extraction network is a ResNet101 network, according to the size of the feature map extracted by the feature extraction network, the feature extraction network is divided into 5 layers, wherein the feature map output by layer1 is 1 / 2 of the original picture in height and width and has 64 channels, the feature map of layer2 is 1 / 4 of the original picture in height and width and has 128 channels, the feature map of layer3 is 1 / 8 of the original picture in height and width and has 256 channels, the feature map of layer4 is 1 / 16 of the original picture in height and width and has 512 channels, and the feature map of layer5 is 1 / 32 of the original picture in height and width and has 1024 channels; Step 2. Predict the motion vector; The current frame image and the previous frame image are stacked in the channel dimension and input into the motion vector prediction network to predict the motion vector; the network is composed of four groups of convolutions, wherein the first group and the second group of convolutions are respectively composed of convolution layers with channel numbers of 32 and 64, the third group of convolutions is a multi-scale feature fusion layer composed of four convolution layers with expansion rates of 1, 2, 4 and 8, and the fourth group of convolutions predicts the motion vector through a convolution layer with a convolution kernel size of 7x7, and finally obtains a motion vector map with the same resolution as the feature map; the video target detection method adopts a motion vector guided local attention method for feature alignment; the Faster R-CNN model uses the motion vector predicted by the motion vector prediction network to guide the position of attention calculation, and aligns the features of adjacent frames to the current frame; Step 3. Feature propagation; When detecting the current frame, if the current frame is the initial frame of the video frame, image target detection is used for the frame; at the same time, the memory feature is initialized as the feature map of the frame for subsequent feature propagation; if the current frame is not the initial frame of the video frame, the feature extraction network is used to extract the feature, the motion vector prediction network is used to predict the motion vector, and the motion vector guided local attention method is used to propagate the memory feature from the previous frame to the current frame; Step 4. Feature fusion; After the memory feature is propagated to the current frame, it needs to be fused with the current frame; The memory feature can enhance the current feature and alleviate the blur and occlusion problems that may occur in the current frame; at the same time, the introduction of the current frame feature makes the propagated feature better adapt to the changes of the current scene; Step 5. Obtain the detection result; The enhanced feature map is input into the target detection task network to obtain the detection result; after extracting the candidate regions by the RPN network and extracting the candidate region features by the ROI pooling, two fully connected layers with a channel number of 1024 are used to further extract the features of the candidate regions, and then two fully connected layers with a channel number of C+1 and 4 are used to predict the confidence that the region of interest belongs to a specific category and the more accurate target position respectively; Step 6. Determine whether the detection is completed; Determine whether the image sequence is detected, if the detection is completed, output the detection result of the image sequence, if not, return to step 1. 2.The method of claim 1, wherein, The video target detection method is trained for a total of 240,000 iterations. 3.The method of claim 1, wherein, The video target detection method has a learning rate of 2.5*10 -4 for the first 160,000 iterations during training -5 , and a learning rate of 2.5*10 -5 for the last 80,000 iterations. 4.The method of claim 1, wherein, The video target detection method uses 0.5 IoU threshold NMS to suppress repeated detection boxes.
Citation Information
Patent Citations
Video target segmentation method based on space-time decoupling attention mechanism
CN116416553A
Video target segmentation method based on space-time decoupling attention mechanism
WO2024183024A1