A video object detection method based on a deep convolutional neural network
By dividing video frames into keyframes and non-keyframes, and using feature propagation methods with multi-feeling domain sampling alignment, inter-frame redundancy and occlusion problems in video object detection are solved, and the accuracy and speed of detection are improved.
Patent Information
- Application Number
- CN202310163429.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-02-24
AI Technical Summary
The prior art fails to effectively utilize the time correlation of video in video object detection, resulting in occlusion, blurring and inter-frame redundancy problems, affecting detection efficiency and accuracy.
The video object detection method based on deep convolutional neural network is adopted to divide the video frame into keyframes and non-keyframes. The feature propagation method of multi-sensory domain sampling alignment is used to propagate memory features on the time dimension, compensate for the current frame features through feature fusion, and construct a multi-scale feature pyramid for motion information modeling.
It improves the accuracy and speed of video object detection, maintains real-time detection, and effectively utilizes the video time information.
Smart Images

Figure CN116168327B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and more specifically, to a video object detection method based on a deep convolutional neural network. Background Art
[0002] Object detection aims to locate the positions of single or multiple objects and identify their categories, which has an important impact on subsequent analysis and understanding tasks such as object tracking, instance segmentation, action recognition, and image description. Static image object detection mainly detects the categories and positions of objects from images. Compared with static images, videos can provide richer visual information. In the real world, more video data is received by monitoring systems, camera devices of vehicles, social media networks, and wearable devices. Videos add a time dimension to static images. If the static image object detection method is directly used to detect video frames one by one, problems such as occlusion, blur, and defocus may occur for moving objects in the video, which will affect the detection performance. In addition, the similarity between video frames is relatively high, and there is temporal redundancy. Utilizing the unique temporal correlation in the video to predict the changes between frames in the video object detection task can improve the detection efficiency. Therefore, it is of great significance to study how to improve the performance of the video object detection method by using the unique temporal information of the video based on the static image object detection method.
[0003] In recent years, deep convolutional neural networks have been widely applied to the field of object detection. Video object detection methods based on deep convolutional neural networks incorporate temporal information on the basis of using static image object detection methods to improve the detection performance. FGFA (Flow-guided feature aggregation) is an optical flow-guided feature fusion model that uses optical flow information to align the features of adjacent multiple frames of images to the current frame, and adaptively weights and fuses the aligned multiple-frame features to compensate for the features of the current frame. DFF (Deep feature flow for video recognition) applies a deep network to extract features from sparse key frames, and propagates the deep features of the key frames to dense non-key frames under the guidance of optical flow. Through feature reuse, the speed of the video object detection model is improved. LSTS (Learnable spatio-temporal sampling) applies a deep network to extract features at key frames, uses a shallow network to extract features at non-key frames, and simultaneously propagates a memory feature. During the feature alignment process, feature points are sampled from other frames to establish the spatial correspondence of inter-frame features, and the positions of the sampled feature points are adaptively learned through training.
[0004] The present invention proposes a video object detection method based on a deep convolutional neural network, which divides video frames into key frames and non-key frames, propagates memory features containing historical key frame information in the time dimension, uses different feature propagation methods to propagate the memory features to key frames and non-key frames, and corrects the current frame features through feature fusion; adjusts the model structure based on Faster R-CNN and uses it as the single-frame detection model in the video object detection method proposed by the present invention; applies the single-frame object detection method of the improved Faster R-CNN to the feature reuse and recursive fusion framework, and adopts a feature propagation method of multi-receptive field sampling alignment during feature propagation. By constructing a multi-scale feature pyramid and sampling feature points in multiple receptive fields to model motion information for feature propagation. The distance between key frames is relatively far, and it is easier for the target to change in position or size. For the deep features of key frames with a relatively far distance, sampling alignment with multiple receptive fields is used to propagate the memory features to the current frame and perform fusion, which can more accurately utilize the time information. The non-key frames are relatively close to the key frames and the frame content changes less. For the shallow features of non-key frames with a relatively close distance, a method of sampling alignment with a small receptive field is used for propagation, which can effectively improve the speed. Since deep features are extracted at key frames and shallow features are extracted at non-key frames, updating the memory features at key frames has a higher accuracy rate. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a video object detection method based on a deep convolutional neural network to achieve the object detection of video objects.
[0006] To achieve the above object, the embodiment of the present invention provides the following solution:
[0007] A video object detection method based on a deep convolutional neural network, characterized by comprising the following five steps:
[0008] Step 1. Set key frames;
[0009] Select the -th frame in the video sequence as the first key frame, and then select a frame as a key frame every K frames. Each key frame will use the frames before and the frames after it as non-key frames similar to this key frame.
[0010] Step 2. Determine whether the current frame is a key frame;
[0011] For the current image sequence, it can be expressed as {I t}, t = 1, 2,..., N, and set its key frame index as the following formula:
[0012]
[0013] When the current frame index t satisfies the above formula, it is a key frame; otherwise, it is a non-key frame.
[0014] Step 3. Feature extraction;
[0015] When detecting the current frame, if the current frame is a key frame then use a deep network to extract features If the current frame is a non-key frame then use a shallow network to extract features And further extract features from the shallow features through a lightweight convolutional network as a conversion network to approximate the deep semantic features.
[0016] Step 4. Feature reuse;
[0017] Step 4.1 Feature propagation:
[0018] When detecting the current frame, if the current frame is a key frame, use a deep network to extract features, and propagate them to the current frame through the multi-receptive field sampling alignment method. If the current frame is a non-key frame, use a shallow network to extract features, and propagate the memory features to the current non-key frame through the small receptive field sampling alignment.
[0019] Step 4.2 Adaptive feature fusion:
[0020] After the memory features are propagated to the current frame, feature fusion with the current frame is required. Among them, the memory features can compensate for the current features and alleviate problems such as blur and occlusion that may exist in the current frame, while the addition of the current frame features makes the propagated features more adaptable to the changes in the current scene.
[0021] Step 5. Obtain the detection result;
[0022] In the detection network of the method disclosed by the present invention, after extracting the candidate region features through ROI pooling, use two fully connected layers with 1024 channels to further extract features from the candidate region features, and then use fully connected layers with C + 1 and 4 channels respectively to predict the confidence that the region of interest belongs to a specific category and the more accurate target position.
[0023] Step 6. Determine whether the detection is completed;
[0024] Determine whether the image sequence has been detected. If the detection is completed, output the detection result of the image sequence. If not, return to Step 2.
[0025] Preferably, the algorithm uses the multi-receptive field sampling method for feature alignment.
[0026] Preferably, the algorithm iterates 4 epochs in total.
[0027] Preferably, the initial value of the learning rate is 0.0001, and it decays to 0.1 of the previous value every 2.33 epochs of iteration.
[0028] Preferably, the key frame interval is set to 10.
[0029] The present invention discloses a video object detection method based on a deep convolutional neural network to achieve object detection of video objects. Under the video object detection framework of feature reuse and recursive fusion, the present invention uses a video object detection algorithm for multi-receptive field sampling alignment. The algorithm divides the video into key frames and non-key frames, propagates a memory feature containing historical key frame information in the time dimension, and compensates the current frame feature through feature fusion. For key frames that are far apart, the memory feature is propagated using the multi-receptive field sampling alignment method. By constructing a multi-scale feature pyramid, feature points in multiple receptive fields are sampled to model the motion information, and features in different receptive fields focus on different degrees of displacement and deformation of the target. For non-key frames that are close, the memory feature is propagated using the small receptive field sampling alignment method; while maintaining the speed, the algorithm further improves the detection accuracy. Description of the Drawings
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 It is a schematic diagram of the improved Faster R-CNN single-frame object detection model structure provided by the embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the video object detection method based on a deep convolutional neural network provided by the embodiment of the present invention;
[0033] Figure 3 It is a flowchart of the video object detection method based on a deep convolutional neural network provided by the embodiment of the present invention;
[0034] Figure 4 It is a schematic diagram of the small receptive field and large receptive field structures of the video object detection method based on a deep convolutional neural network provided by the embodiment of the present invention;
[0035] Figure 5 It is a schematic diagram of the large receptive field sampling alignment provided by the embodiment of the present invention;
[0036] Figure 6Schematic diagram of multi-receptive field sampling alignment provided by an embodiment of the present invention;
[0037] Figure 7 Training process of the video object detection method based on a deep convolutional neural network provided by an embodiment of the present invention. Detailed implementation manners
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] The model of the feature extraction network proposed by the present invention is as Figure 1 shown, and the specific parameters are shown in Table 1. It consists of 5 groups of convolutions of ResNet-101 and a dilated convolution layer with a dilation rate of 6. Among them, the 5th group of convolutions is shared on the entire image to extract deeper features. The first convolution layer of this group of convolutions is a dilated convolution with a dilation rate of 2, and the convolution stride is 1, so that the dilated convolution feature map has the same resolution as the feature map of the 4th group of convolutions, and the position information is retained. A dilated convolution with a dilation rate of 6 is used in the dimensionality reduction convolution layer to generate a dilated convolution feature map dilate_conv_feat, further expanding the feature receptive field. Dilate_conv_feat is divided into two parts along the channel. One part is input into the Region Proposal Network (RPN) to generate candidate regions, and the other part is input into the detection network. The candidate regions are generated and the features of the candidate regions are extracted based on the same feature map, retaining the position information when generating the candidate regions and effectively utilizing the features with rich semantic information at the same time.
[0040] Table 1 Model parameters of the feature extraction network
[0041]
[0042]
[0043] In a video, adjacent video frames usually have redundancy in content. This redundancy can be used to reduce the computational amount by feature reuse in the feature extraction stage. Fusing the features of other frames using the inter-frame correlation in the video can compensate for the features of the current frame and improve the detection accuracy. However, densely fusing the features of multiple frames in the adjacent range will increase the computational amount and cause information redundancy. In addition, it is difficult to utilize the temporal information in frames farther away than adjacent frames. Therefore, the present invention proposes multi-receptive field sampling alignment for video object detection and constructs as Figure 2The video object detection model with recursive fusion as shown. The video frames are divided into key frames and non-key frames, and the non-key frames are regarded as redundant frames of the key frames. To reduce the computational complexity, the feature extraction networks for different types of frames are different. The key frames use a deep feature extraction network, and the non-key frames use a shallow feature extraction network. To utilize the inter-frame correlation over a long distance, memory features are introduced and gradually updated among the key frames. In addition, in the shallow feature extraction network, the shallow features on the non-key frames usually cannot obtain good detection results. Therefore, the memory features between the key frames are used to enhance the shallow features on the non-key frames.
[0044] As Figure 2 shown, let the i-th key frame in the video be the frame at time t i , i = 0, 1, 2, …, and the frame at time t i +k be the non-key frame close to the frame at time t i . M is the memory feature propagated on the time axis and is used to compensate the features extracted from the key frames and non-key frames. The initial value of M is the deep feature of the first key frame in the video sequence frame, and it is updated after each propagation to the key frame and fusion with the key frame feature. Therefore, the memory feature contains the information in all the previous key frames of the current frame. When detecting the current frame, if the current frame is a key frame then use the deep network to extract features that is, use the entire ResNet network as the deep network, and propagate M to the current frame through the multi-receptive field sampling alignment method, and fuse it with the current feature for compensation, and input it into the RPN and detection network for subsequent operations. At the same time, update M using the fused feature. If the current frame is a non-key frame use the shallow network to extract features Here, the structure before conv4_3 in the ResNet network is used as the shallow network. Since the shallow features do not contain enough semantic information to model the spatial correspondence, a lightweight convolutional network is used as the conversion network to further extract features from the shallow features to approach the deep semantic features. Propagate the memory feature to the current non-key frame through small receptive field sampling alignment, fuse it with the shallow features of the current frame to compensate the current frame feature, and input it into the RPN and detection network for subsequent operations.
[0045] The embodiment of the present invention discloses a method for video object detection based on a deep convolutional neural network to complete the object detection of video objects. Refer to Figure 3 , the above method at least includes the following 5 steps.
[0046] Step 1. Set key frames
[0047] The The frame is selected as the first key frame, and then one frame is selected as a key frame every K frames. Each key frame will use the previous frame and the subsequent frame as non-key frames similar to this key frame.
[0048] Step 2. Determine whether the current frame is a key frame
[0049] For the current image sequence, it can be expressed as {I t}, where t = 1, 2,..., N. Set its key frame index as the following formula:
[0050]
[0051] When the current frame index t satisfies formula (1), it is a key frame; otherwise, it is a non-key frame.
[0052] Step 3. Feature extraction
[0053] When detecting the current frame, if the current frame is a key frame then use the deep network to extract features If the current frame is a non-key frame then use the shallow network to extract features Since the shallow features do not contain enough semantic information to model the spatial correspondence, a lightweight convolutional network is used as a conversion network to further extract features from the shallow features to approximate the deep semantic features. This conversion network consists of one 1×1 and two 3×3 convolutions, with the number of channels being 256, 256, and 1024 respectively.
[0054] Step 4. Feature reuse
[0055] Step 4.1 Feature propagation:
[0056] In the recursive fusion video object detection model, when propagating the memory features to the current frame for feature compensation, it is necessary to first align the feature maps across frames. Optical flow is difficult to model the motion information between two frames that are far apart. At the same time, the memory features are recursively propagated on the time axis, and the optical flow between the memory features and the current frame cannot be calculated. Therefore, the present invention proposes a propagation method of multi-receptive field sampling alignment. Since the memory features are updated at the time nodes of the key frames, if the current frame is a key frame, due to the large distance from the previous key frame, the target is more likely to change in position or size relative to the previous key frame. Therefore, by expanding the receptive field of the sampling points, the influence brought by the larger displacement and size change of the target is compensated. If the current frame is a non-key frame, then the distance from the previous key frame is relatively close, and the frame content changes little. Therefore, the sampling alignment method is directly used for propagation.
[0057] The present invention first performs a pooling operation on the features to be propagated, down-samples the propagation features by a factor of s to aggregate feature information, samples feature points at a stride of s on the features after aggregating the information, and expands the receptive field of each feature point without changing the computational amount. When s = 1, the receptive field of the small receptive field sampling alignment is as shown in Figure 4 (a). When s = 2, its receptive field relative to the features of this layer is as shown in Figure 4 (b), and the receptive field size is s 2 times that of the original.
[0058] The large receptive field sampling alignment when s takes the value of 2 includes three stages as shown in Figure 5 . Figure 5 The number ① in represents the stage of obtaining the features of the large receptive field. A pooling operation with a stride of 2 and a pooling kernel size of 2×2 is performed on M to aggregate information, and a feature map M with an expanded receptive field of the feature point is obtained ↓2 . Assume that the key frame features are F t . For the convenience of calculating the similarity between the feature points at each position (x, y) in F t and the sampled feature points in M ↓2 , first use nearest neighbor interpolation to upsample M ↓2 by a factor of 2 to restore the same resolution as F t , and obtain (M ↓2 ) ↑2 . The receptive field of each feature point in (M ↓2 ) ↑2 is still increased. Figure 5 The numbers ② and ③ in represent the stage of sampling alignment, which includes two parts: calculating the similarity weights and weighted summation. When calculating the similarity weights, (M ↓2 ) ↑2 and F t are dimensionally reduced through an embedded convolutional neural network to generate φ((M ↓2 ) ↑2 ) and φ(F t ). Sample (2k + 1)×(2k + 1) points in the neighborhood range (4k + 1)×(4k + 1) of the position (x, v) in φ((M ↓2 ) ↑2 ) at a stride of 2, and calculate the cosine similarity between these points and the feature points at the position (x, y) in φ(F t ). If the current position (x, y) is the sampling point of f(F Figure 5 ) in, then calculate its cosine similarity with the sampled feature points in the neighborhood range of the position (x, y) in φ((M t ) ↓2 ) ↑2 ). When performing weighted summation, use the normalized cosine similarity as the weight to perform weighted summation on (M ↓2 )↑2 Weighted sum of feature points sampled within the neighborhood of the position (x, y) in the middle to generate the aligned feature G t 。
[0059] In the sampling alignment with a large receptive field and a step size of s, the calculation formula for the similarity weight is as follows:
[0060]
[0061] In the formula, (x + si, y + sj) represents a certain sampling position within the neighborhood of the position (x, y) in φ((M ↓s ) ↑s ), and φ(·) represents the embedding network. Assuming the resolution of the feature map is W×H, the similarity weight represents the key-frame feature F t at each position (x, y), the cosine similarity between the feature point and the large-receptive-field memory feature (M ↓s ) ↑s within the range of (2sk + 1)×(2sk + 1) of the neighborhood of the position (x, y) in (M x,y (i, j) is the value on the i(2k + 1)+j-th channel at the position (x, y) in the similarity map, representing the cosine similarity between the feature point at the position (x, y) in F t and the feature point at the position (x + si, y + sj) in (M ↓s ) ↑s normalized within the neighborhood range. The calculation formula for the weighted sum is as follows:
[0062]
[0063] In the formula, G t represents the feature after aligning the memory feature to the key frame through large-receptive-field sampling alignment.
[0064] Based on the small-receptive-field sampling alignment and large-receptive-field sampling alignment methods, the method disclosed in the present invention uses the Figure 6 shown multi-receptive-field sampling alignment method to propagate the memory feature for the key frame. In order to sample feature points with multiple receptive fields for feature alignment during the process of propagating features, first construct a multi-scale feature pyramid with three levels, downsample M twice through pooling operations to obtain the 2-fold and 4-fold downsampling results M ↓2 and M ↓4 . To facilitate the calculation of sampling alignment at each level, upsample M ↓2 and M ↓4 by nearest-neighbor interpolation to restore the same resolution as F t , generating (M ↓2 ) ↑2 and (M↓4 ) ↑4 Input M, (M ↓2 ) ↑2 , (M ↓4 ) ↑4 into the sampling alignment module respectively, and generate alignment features through sampling alignments with step sizes of 1, 2, and 4 t where the sampling alignment with a step size of 1 is the ordinary small receptive field sampling alignment, and the sampling alignments with step sizes of 2 and 4 represent the large receptive field sampling alignments when s = 2 and s = 4. The different alignment features generated by using sampling alignments of multiple receptive field features focus on the displacements and size changes of different degrees during the target movement process. and The sampling alignment stage described in the non - key frame propagation module is the small receptive field sampling alignment. Each sampled feature point only includes the feature information at that position, that is, the receptive field of each sampled feature point relative to the feature of this layer is only restricted to the point itself. Assuming 9 points in the sampling neighborhood, the sampling receptive field is as shown in
[0065] (a). If the sampling neighborhood range is expanded, that is, the number of sampled points is increased to expand the receptive field to increase information, the computational amount will also increase accordingly. Figure 4 (a). If the sampling neighborhood range is expanded, that is, the number of sampled points is increased to expand the receptive field to increase information, the computational amount will also increase accordingly.
[0066] Step 4.2 Adaptive feature fusion:
[0067] After the memory feature is propagated to the current frame, it needs to be feature - fused with the current frame. The memory feature can compensate the current feature to solve problems such as blur and occlusion that may exist in the current frame, while the addition of the current frame feature makes the propagated feature more adaptable to the changes in the current scene.
[0068] Adaptive weighted feature fusion inputs the features to be fused into a convolutional neural network to generate adaptive weights respectively, and weights the features to be fused according to the corresponding weights. The weight generation network consists of three convolutional layers, with the convolutional kernel sizes being 1×1, 3×3, and 3×3 respectively, and the corresponding number of convolutional channels being 256, 16, and 1.
[0069] For the key frame I t , the features to be fused are the results of propagating the memory feature M to the current frame through the multi - receptive field sampling alignment method and and the deep - layer feature F of the current frame t . Assuming the scale of the features to be fused is C×M×N, the corresponding fusion weights are adaptively generated by the weight generation network respectively and and and the fusion weights are normalized using the Softmax function, expressed as:
[0070]
[0071]
[0072] In the formula, represents the propagation feature The fusion weight at position (m, n) in, i = {s, m, l}. W t (1, m, n) represents the fusion weight at position (m, n) in the current key-frame feature F t , m = 1, 2,..., M, n = 1, 2,..., N. According to the propagation feature and the current key-frame feature F t , and the corresponding normalized fusion weights and The calculation formula for feature fusion is:
[0073]
[0074] In the formula, is the fusion feature of the current key frame, and Perform per-channel operations on and F t respectively.
[0075] For non-key frame I t+k , the feature to be fused is the result of the memory feature M propagated to the current frame and the shallow feature L of the current frame t+k . Assume respectively pass through the feature fusion network and adaptively generate the corresponding fusion weights and And normalize the fusion weights using the Softmax function, expressed as:
[0076]
[0077]
[0078] In the formula, represents the propagation feature The fusion weight at position (m, n) in, W t+k (1, m, n) represents the fusion weight at position (m, n) in the shallow feature L of the current non-key frame t+k , m = 1, 2,..., M, n = 1, 2,..., N. According to the propagation feature and the shallow feature L of the current non-key frame t+k , and the corresponding normalized fusion weights and The calculation formula for feature fusion is:
[0079]
[0080] In the formula, is the fusion feature of the current non-key frame, and perform channel-by-channel operations on and L t+k respectively.
[0081] Step 5. Obtain the detection result
[0082] After extracting the candidate region features through ROI pooling in the detection network disclosed by the present invention, two fully connected layers with 1024 channels are used to further extract features from the candidate region features, and then fully connected layers with C + 1 and 4 channels are used to predict the confidence that the region of interest belongs to a specific category and the more accurate target position respectively.
[0083] Step 6. Determine whether the detection is completed
[0084] Determine whether the image sequence has been detected. If the detection is completed, output the detection result of the image sequence. If not, return to Step 2.
[0085] In the model training stage, as Figure 7 shown, three frames of images are selected Among them, is used as the previous key frame, is used as the current key frame, is used as the current non-key frame to be detected and similar in content to . Select any frame within K / 2 frames before and after as Select any frame within K / 2 to K frames before as For and extract deep features and Use the multi-receptive field sampling alignment method to propagate to to obtain the aligned features with different receptive fields of the feature points and Fuse the aligned features with to obtain compensation feature At the same time, is also used as the memory feature M. For extract shallow features to generate Use the small receptive field sampling alignment method to propagate M to and fuse to obtain Compensation feature will as the image to be detected The final image features are sent to the subsequent RPN and detection networks. Generate bounding box regression scores and class confidences, and calculate the loss function to update the network parameters.
[0086] Preferably, when sampling and aligning, the parameter k defining the neighborhood range is set to 4, and the key frame interval K is set to 10. The parameters of the ResNet-101 part in the feature extraction network are also pre-trained on the ImageNet dataset, and the parameters of the remaining convolutional layers are initialized by Gaussian distribution. The shallow features of non-key frames are extracted by the convolutional layers before the 3rd group of the 4th stage in ResNet-101. In the training stage, the number of images per batch is set to 2, and the initial learning rate is 10 -4 , the model is trained for a total of 4 epochs, and the learning rate is reduced to 1 / 10 of the original after 2.333 epochs of iteration. After the training is completed, the model parameters are saved.
[0087] The video object detection method based on deep convolutional neural network disclosed in the present invention is based on an improved Faster RCNN single-frame detection model, incorporates temporal information, and adopts a feature alignment method of multi-receptive field sampling alignment. The present invention uses the ImageNet VID dataset and the data subset with the same categories as the ImageNet VID dataset in the ImageNet DET dataset to train and test the model. The ground truth labels of the ImageNet VID test set are not publicly available, and the accuracy of the detection results cannot be evaluated. Therefore, the model performance is evaluated on the validation set of ImageNet VID.
[0088] The present invention uses the mean average precision mAP and the algorithm running time as the quantitative evaluation metrics. mAP is the average of the average precisions AP of detections for each category. The higher the mAP, the higher the accuracy of object detection. Table 2 shows the AP, mAP values, and average time per frame of detections for each category of the method disclosed in the present invention, the feature reuse method DFF, the image-level feature fusion method FGFA, and the image-level sparse fusion method LSTS on the ImageNet VID validation set. It can be seen from Table 2 that the mAP value of the method disclosed in the present invention is 75.49%, and the average detection time is 65.2 ms. Compared with FGFA, the mAP of the present invention is increased by 0.63%, and the detection time is 1 / 5 of it. Compared with the same type of model LSTS, the detection time of the model of the present invention is basically the same as it, and the mAP is 1.67% higher.
[0089] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0090] Table 2 Experimental results of the method disclosed in the present invention
[0091]
Claims
1. A video object detection method based on a deep convolutional neural network, characterized in that, It includes the following 5 steps: Step 1. Set key frames; Select the th frame in the video sequence as the first key frame, and then select one frame as a key frame every K frames. Each key frame will use the frames before and frames after it as non-key frames similar to this key frame; Step 2. Determine whether the current frame is a key frame; For the current image sequence, which is represented as {I t}, where t = 1, 2, …, N, set its key-frame index as follows: When the current frame index t satisfies the above formula, it is a key frame; otherwise, it is a non-key frame; Step 3. Feature extraction; When detecting the current frame, if the current frame is a key frame then use a deep network to extract features If the current frame is a non-key frame then use a shallow network to extract features and further extract features from the shallow features through a lightweight convolutional network as a conversion network to approximate the deep semantic features; Step 4. Feature reuse; Step 4.1 Feature propagation: When detecting the current frame, if the current frame is a key frame, features are extracted using a deep network and propagated to the current frame through the multi-receptive field sampling alignment method; if the current frame is a non-key frame, features are extracted using a shallow network and the memory features are propagated to the current non-key frame through the small receptive field sampling alignment; Step 4.2 Adaptive feature fusion: After the memory features are propagated to the current frame, feature fusion with the current frame is required; among them, the memory features compensate for the current features, alleviating the possible blur and occlusion in the current frame, and the addition of the current frame features makes the propagated features more adaptable to the changes in the current scene; Step 5. Obtain the detection result; After extracting the candidate region features through ROIpooling, the candidate region features are further extracted using two fully connected layers with 1024 channels, and then the confidence that the region of interest belongs to a specific category and the more accurate target position are predicted using fully connected layers with C+1 and 4 channels respectively; Step 6. Determine whether the detection is completed; Determine whether the image sequence detection is completed. If the detection is completed, the detection results of the image sequence are output. If not, return to Step 2.
2. The video object detection method based on a deep convolutional neural network according to claim 1, characterized in that Feature alignment is performed using the multi-receptive field sampling method.
3. The video object detection method based on a deep convolutional neural network according to claim 1, characterized in that, A total of 4 epochs are iterated.
4. The video object detection method based on a deep convolutional neural network according to claim 1, characterized in that The initial value of the learning rate is 0.0001, and it decays to 0.1 of the previous value every 2.33 epochs.
5. The video object detection method based on a deep convolutional neural network according to claim 1, wherein The key frame interval is set to 10.
Citation Information
Patent Citations
Object detection method and system based on dynamic memory and motion perception
CN109191498A
License plate detection method based on knowledge distillation training and space-time joint attention
CN114283402A