A method and device for target detection based on image sequence

By using the correlation information of the image sequence in the object detection method, multiple feature maps are extracted and fused, the problem of degradation of detection performance under interference such as image motion blur and object occlusion is solved, and a more accurate object detection is achieved.

CN116342991BActive Publication Date: 2025-05-06COMP APPL TECH INST OF CHINA NORTH IND GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310287403.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-05-06
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

The existing object detection method has deteriorated detection performance under interference such as image motion blur and object occlusion, and has failed to effectively utilize the correlation information between images before and after the video stream.

Method used

By acquiring the image sequence, multiple backbone feature maps and sequence feature maps of the current time image and image sequence are extracted, and channel aggregation is performed to fuse multiple target feature maps to obtain the target detection results.

Benefits of technology

In the case of interference such as image motion blur and object occlusion, the accuracy and performance of target detection are improved, and are suitable for fields such as smart traffic and aerial remote sensing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342991B_ABST
    Figure CN116342991B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for target detection based on image sequences, the method comprising: obtaining a sequence of images to be detected including an image at a current moment; respectively obtaining multiple backbone feature maps of different sizes and multiple sequence feature maps of corresponding sizes of the image at the current moment and the image sequence; performing channel aggregation based on multiple backbone feature maps and sequence feature maps of corresponding sizes to obtain multiple target feature maps of corresponding sizes; fusing multiple target feature maps to obtain target detection results. The present invention solves the problem that the target detection performance of the target detection method in the prior art is limited under interference such as image motion blur and object occlusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence target detection, and in particular, relates to a target detection method and device based on image sequences. Background Art

[0002] The task of target detection is to detect the target object of interest in a static image (or dynamic video), which can provide good conditions for subsequent target tracking, behavior prediction, etc. The target detection task is an important part of computer vision and is widely used in many fields such as unmanned driving, robot navigation, video surveillance, industrial inspection, aerospace, etc.

[0003] Existing target detection methods have achieved a lot of results, and typical research results include Faster-RCNN, SSD, CenterNet, YOLOv1~YOLOv5, etc. These methods only use the information of a single image, and do not effectively use the correlation information between the image sequence or the previous and next images of the video stream, resulting in the degradation of target detection performance under interference such as image motion blur and object occlusion, which hinders the actual application of target detection systems. How to improve the target detection performance under interference such as image motion blur and object occlusion has become an urgent problem to be solved. Summary of the invention

[0004] In view of the above analysis, the present invention aims to provide a target detection method and device based on image sequence, which is used to solve the problem that the target detection performance of the target detection method in the prior art is limited under the interference of image motion blur, object occlusion, etc.

[0005] The purpose of the present invention is mainly achieved through the following technical solutions:

[0006] In one aspect, the present invention provides a method for object detection based on an image sequence, comprising:

[0007] Acquire a sequence of images to be detected; the image sequence includes an image at a current moment;

[0008] Respectively obtaining a plurality of backbone feature maps of different sizes and a plurality of sequence feature maps of corresponding sizes of the current moment image and the image sequence;

[0009] Perform channel aggregation based on the plurality of backbone feature maps and sequence feature maps of corresponding sizes to obtain a plurality of target feature maps of corresponding sizes;

[0010] The target detection result is obtained by fusing multiple target feature maps.

[0011] Further, backbone feature map extraction, sequence feature map extraction and channel aggregation are performed respectively through the pre-set first backbone network, second backbone network and head network;

[0012] The first backbone network is used to obtain the current moment image and perform multiple feature map size compression and feature extraction on the current moment image in sequence to obtain n backbone feature maps of different sizes;

[0013] The second backbone network is used to perform multiple image size and dimension compression and feature extraction on the input image sequence to obtain n sequence feature graphs of corresponding sizes;

[0014] The head network is used to perform multiple channel aggregations based on the backbone feature map and the sequence feature map to obtain n target feature maps of corresponding sizes, where n is an integer greater than 1.

[0015] Furthermore, n is 3; channel aggregation is performed based on the plurality of backbone feature maps and sequence feature maps of corresponding sizes to obtain a plurality of target feature maps of corresponding sizes, including:

[0016] After compressing and upsampling the number of channels of the third backbone feature map, the third backbone feature map is aggregated with the second backbone feature map, and the feature map after channel aggregation is extracted through the C3 layer to obtain the first fused feature map;

[0017] After compressing the number of channels and expanding the size of the first fusion feature map, performing channel aggregation with the first backbone feature map and the first sequence feature map to obtain a first target feature map;

[0018] Perform channel aggregation based on the first target feature map, the second sequence feature map and the first fusion feature map to obtain a second target feature map;

[0019] Channel aggregation is performed based on the second target feature map, the third sequence feature map and the third backbone feature map to obtain a third target feature map.

[0020] Further, the second backbone network includes 5 SpatialConv3D modules connected in sequence, and 3 TemporalConv3D modules connected to the 3rd, 4th and 5th SpatialConv3D modules respectively;

[0021] The SpatialConv3D module is used to compress the feature map size of the input image sequence;

[0022] The TemporalConv3D module is used to extract image sequence features from the compressed image sequences output by the corresponding SpatialConv3D module, and compress them into one sequence feature map.

[0023] Furthermore, the SpatialConv3D module includes a Conv3d layer, a BatchNorm3d layer and a SiLU layer which are arranged in sequence; by setting the convolution kernel parameters of the Conv3d layer, the size of the input image sequence is changed from N*H i *W i Adjust to N*H j *W j , where N is the number of feature maps in the image sequence, H i , W i is the length and width of the input feature map, H j , W j is the length and width of the output feature map.

[0024] Furthermore, the TemporalConv3D module includes: a Conv3d layer, a BatchNorm3d layer, a SiLU layer and a squeeze layer; the Conv3d layer and the squeeze layer respectively perform feature extraction and dimension compression on the input image sequence, and the size and dimension of the input feature map sequence are reduced from N*H j *W j Adjust to H j *W j , and get a size of H j *W j Sequence feature diagram.

[0025] Furthermore, after performing feature extraction and dimension compression on the feature map sequences of different sizes output by the 3rd, 4th and 5th SpatialConv3D modules respectively through 3 TemporalConv3D modules, the CA attention layer is also used to extract features of interest from the compressed feature maps to obtain the first sequence feature maps, the second sequence feature maps and the third sequence feature maps.

[0026] Further, the first backbone network includes an ImageSub layer, a first feature map compression extraction module, a second feature map compression extraction module, a third feature map compression extraction module and a fourth feature map compression extraction module connected in sequence;

[0027] The ImageSub layer is used to obtain the current moment image in the image sequence and output it to the first feature map compression extraction module;

[0028] The first feature map compression and extraction module includes a primary visual cortex simulation module, a Conv layer and a CA attention layer connected in sequence; it is used to perform bionic visual feature extraction on the input current moment feature map to obtain a bionic visual feature map, and input it into the second feature extraction module for feature compression and feature extraction;

[0029] The second feature map compression extraction module and the third feature map compression extraction module both include a Conv layer, a C3 layer and a CA attention layer connected in sequence; the input feature map is compressed in size through the Conv layer and then input into the corresponding C3 layer for feature extraction to obtain a first backbone feature map and a second backbone feature map; and the first backbone feature map and the second backbone feature map are respectively extracted with features of interest through the corresponding CA attention layer as inputs to the next layer feature map compression extraction module;

[0030] The fourth feature map compression and extraction module includes a Conv layer, a C3 layer, an SPPF layer and a CA attention layer connected in sequence; and is used to perform feature map size compression, feature map spatial information fusion and feature extraction on the input feature map to obtain a third backbone feature map.

[0031] Furthermore, the head network includes four serially connected channel feature fusion modules; wherein,

[0032] The first channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the third backbone feature map, perform channel aggregation with the second backbone feature map, and perform feature extraction on the feature map after channel aggregation through the C3 layer to obtain a first fused feature map;

[0033] The second channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the number of feature map channels of the first fusion feature map, perform channel aggregation with the first backbone feature map and the first sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain a first target feature map;

[0034] The third channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the first target feature map, aggregate the channels of the image obtained by compressing the number of feature map channels of the first fusion feature map and the second sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain the second target feature map;

[0035] The fourth channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the second target feature map, perform channel aggregation with the image obtained by compressing the number of feature map channels of the third backbone feature map and the third sequence feature map, and perform feature extraction on the feature map after channel aggregation through the C3 layer to obtain a third target feature map;

[0036] The first target feature map, the second target feature map and the third target feature map are subjected to feature fusion through the Detect layer to obtain a target detection result.

[0037] On the other hand, there is also provided an electronic device, comprising at least one processor, and at least one memory communicatively connected to the processor;

[0038] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the aforementioned target detection method based on image sequences.

[0039] Beneficial effects of this technical solution:

[0040] The target detection method of the present invention can accurately extract the target information in the image at the current moment through the backbone feature map of the image at the current moment; however, when the image at the current moment is disturbed by motion blur or object occlusion, missed detection or false detection is likely to occur. In view of the missed detection or false detection, the present invention extracts the correlation information between n consecutive images including the image at the current moment through the sequence feature map of the image sequence; through the auxiliary role of the correlation information, the fused feature map can realize accurate target recognition under the interference of image motion blur, object occlusion and other interferences, thereby improving the target detection performance under the interference of image motion blur, object occlusion and other interferences, and can perform more accurate target detection in application scenarios such as the smart transportation field (traffic accident monitoring, road target monitoring) and the aerial remote sensing field (ground and sea target detection).

[0041] Other features and advantages of the present invention will be described in the following description, and part of them will become obvious from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like components throughout the drawings.

[0043] Figure 1 is a flow chart of a target detection method based on an image sequence according to an embodiment of the present invention;

[0044] Figure 2 is a schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0045] Figure 3 This is a structural diagram of an imitation primary visual cortex model according to an embodiment of the present invention;

[0046] Figure 4is a schematic diagram of the structure of the SpatialConv3D module of an embodiment of the present invention;

[0047] Figure 5 It is a schematic diagram of the TemporalConv3D module structure of an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the implementation cases of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0049] An embodiment of the present invention provides a method for object detection based on image sequences, such as Figure 1 As shown, the following steps are included:

[0050] Step S1, obtaining a sequence of images to be detected.

[0051] Specifically, this embodiment sets an ImagesPub layer for reading the image sequence to be detected. The image sequence acquired in this embodiment can be video stream data collected by the camera in real time or a time-continuous picture set stored in a hard disk; and the acquired image sequence includes the current moment image; that is, the image sequence to be detected is an image sequence containing N video frames or N continuous picture sets of images at the current moment. Preferably, the value of N is an integer of 2-5.

[0052] Step S2, respectively obtaining a plurality of backbone feature maps of different sizes and a plurality of sequence feature maps of corresponding sizes for the current moment image and the image sequence.

[0053] Specifically, Figure 2 As shown, this embodiment performs backbone feature map extraction and sequence feature map extraction through a pre-set first backbone network and a second backbone network; wherein the first backbone network is used to obtain the image at the current moment and perform multiple feature map size compression and feature extraction on the image at the current moment in sequence to obtain n backbone feature maps of different sizes; the second backbone network is used to perform multiple image size and dimension compression and feature extraction on the input image sequence to obtain n sequence feature maps of corresponding sizes; preferably, n is set to 3.

[0054] Particularly, in this embodiment, the first backbone network includes an ImageSub layer, a first feature map compression extraction module, a second feature map compression extraction module, a third feature map compression extraction module, and a fourth feature map compression extraction module connected in sequence;

[0055] The ImageSub layer is used to obtain the current moment image in the image sequence and output it to the first feature map compression extraction module;

[0056] The first feature map compression and extraction module includes a primary visual cortex module, a C3 layer, and a CA attention layer connected in sequence; it is used to extract the bionic visual features of the input current moment image to obtain a bionic visual feature map, and input it into the second feature extraction module for feature compression and feature extraction; wherein, if Figure 3 As shown in FIG. 1 , the simulated primary visual cortex module includes a VOneBlock layer, a Conv layer, and a feature fusion layer. The VoneBlock layer is used to extract bionic visual features from the current image to be detected, that is, to extract features from the input image by simulating the information processing mechanism of the human visual perception cortex, and obtain a feature map that is closer to the features after human brain visual processing. The VOneBlock layer extracts features and compresses the size of the input current image of size H*W, and obtains a feature map of size The Conv layer is used to compress the input image to obtain a feature map of the same size as the feature map output by the VOneBlock layer; the feature fusion layer is used to fuse the feature maps output by the VOneBlock layer and the Conv layer, and the fused size is Further, the size of the output of the simulated primary visual cortex module is The feature map of is compressed and interesting features are extracted through the C3 layer and the CA attention layer, and the size of the output of the first feature map compression extraction module is Bionic visual feature map.

[0057] The second feature map compression extraction module and the third feature map compression extraction module both include a Conv layer, a C3 layer, and a CA attention layer connected in sequence; the Conv layer compresses the input feature map and then inputs it into the corresponding C3 layer for feature extraction, and obtains a size of The first backbone feature map and The second backbone feature map of the first backbone feature map is extracted through the corresponding CA attention layer, and the features of interest are extracted from the first backbone feature map and the second backbone feature map respectively as the input of the feature map compression extraction module of the next layer.

[0058] The fourth feature map compression and extraction module includes a Conv layer, a C3 layer, a SPPF layer, and a CA attention layer connected in sequence; it is used to perform feature map size compression, feature map spatial information fusion, and feature extraction on the input feature map, and obtain a size of The third backbone feature map.

[0059] Furthermore, the second backbone network includes 5 SpatialConv3D modules connected in sequence, and 3 TemporalConv3D modules connected to the 3rd, 4th and 5th SpatialConv3D modules respectively; wherein the SpatialConv3D module is used to compress the feature map size of the input image sequence; the TemporalConv3D module is used to extract image sequence features from the compressed image sequences output by the corresponding SpatialConv3D modules, and compress them into one sequence feature map.

[0060] Special, such as Figure 4 As shown in FIG. 1 , the SpatialConv3D module includes a Conv3d layer, a BatchNorm3d layer, and a SiLU layer which are set in sequence; by setting the convolution kernel parameters of the Conv3d layer, the size of the input image sequence is changed from N*H i *W i Adjust to N*H j *W j , and then pass through the BatchNorm3d layer and SiLU layer for normalization and activation operations, and obtain the image sequence after the SpatialConv3D module sequence compression, where N is the number of feature maps in the image sequence, H i , W i is the length and width of the input feature map, H j , W j is the length and width of the output feature map.

[0061] like Figure 5 As shown, the TemporalConv3D module includes a Conv3d layer, a BatchNorm3d layer, a SiLU layer, and a squeeze layer arranged in sequence; first, the size of the SpatialConv3D module output is N*H j *W j The image sequence is convolved through the Conv3d layer to obtain a size of 1*H j *W j For an image of size 1*H j *W j The image is normalized and activated by the BatchNorm3d layer and the SiLU layer, and then input into the squeeze layer for dimension compression, and the output feature map size is adjusted to H j *W j , that is, we get a sheet of size H j *W j Sequence feature diagram.

[0062] Specifically, the adjustment parameters of the Conv3d layer in the SpatialConv3D module can be set to: convolution kernel size k = (1, 3, 3), sliding window step size s = (1, 2, 2); after the second backbone network reads an image sequence of size N*H*W consisting of N images, it passes through 5 SpatialConv3D modules in sequence to compress the image sequence size, and outputs the size of and sequence of images.

[0063] Furthermore, the adjustment parameters of the Conv3d layer in the TemporalConv3D module are set to: convolution kernel size k = (N, 1, 1), sliding window step size s = (1, 1, 1); the image sequences output by the 3rd, 4th and 5th SpatialConv3D modules are subjected to feature extraction and dimension compression by the corresponding TemporalConv3D modules to obtain image sizes of and

[0064] In particular, after extracting features and compressing the dimensions of the feature map sequences of different sizes output by the 3rd, 4th and 5th SpatialConv3D modules respectively, the CA attention layer is used to extract the features of interest from the compressed feature maps to obtain the corresponding sizes of the three backbone feature maps. The first sequence characteristic diagram, The second sequence characteristic diagram and The third sequence characteristic diagram of .

[0065] Step S3: performing channel aggregation based on multiple backbone feature maps and sequence feature maps of corresponding sizes to obtain multiple target feature maps of corresponding sizes.

[0066] In this embodiment, the head network of the improved YOLOv5 model is used to perform multiple channel aggregation based on the backbone feature map and the sequence feature map to obtain n target feature maps of corresponding sizes. In this embodiment, n is 3. Specifically, after the third backbone feature map is compressed and upsampled in the number of channels, it is channel aggregated with the second backbone feature map, and the feature map after channel aggregation is extracted through the C3 layer to obtain the first fused feature map; after the first fused feature map is compressed in the number of channels and expanded in size, it is channel aggregated with the first backbone feature map and the first sequence feature map to obtain the first target feature map; channel aggregation is performed based on the first target feature map, the second sequence feature map and the first fused feature map to obtain the second target feature map; channel aggregation is performed based on the second target feature map, the third sequence feature map and the third backbone feature map to obtain the third target feature map.

[0067] More specifically, the head network of this embodiment includes four serially connected channel feature fusion modules; wherein,

[0068] The first channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the feature map channels of the third backbone feature map, perform channel aggregation with the second backbone feature map, and perform feature extraction on the feature map after channel aggregation through the C3 layer to obtain a first fused feature map;

[0069] The second channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the number of feature map channels of the first fusion feature map, aggregate the channels with the first backbone feature map and the first sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain a size of The first target feature map;

[0070] The third channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the first target feature map, aggregate the image after the number of feature map channels of the first fusion feature map and the second sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain a size of The second target feature map;

[0071] The fourth channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the second target feature map, aggregate the image after the number of feature map channels of the third backbone feature map and the third sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain a size of The third target feature map.

[0072] Step S4: Fusing multiple target feature maps to obtain target detection results.

[0073] Specifically, this embodiment performs feature fusion on the first target feature map, the second target feature map, and the third target feature map through the Detect layer of the YOLOv5 model to obtain a target detection result.

[0074] A second embodiment of the present invention further discloses an electronic device, comprising at least one processor, and at least one memory communicatively connected to the processor;

[0075] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the aforementioned target detection method based on image sequences.

[0076] In summary, an embodiment of the present invention provides a target detection method based on an image sequence, which performs multiple size compressions and feature extractions on the current moment image and the image sequence including the current moment image, respectively, to obtain multiple feature maps of corresponding sizes, and fuses the features of the current moment image and the image sequence based on the improved YOLOv5 model. It can effectively utilize the correlation information between the image sequence or the previous and next images of the video stream, improve the target detection performance under interference such as image motion blur and object occlusion, and achieve more accurate target detection.

[0077] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0078] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A target detection method based on image sequence, characterized in that: include: Acquire a sequence of images to be detected; the image sequence includes an image at a current moment; Respectively obtaining a plurality of backbone feature maps of different sizes and a plurality of sequence feature maps of corresponding sizes of the current moment image and the image sequence; Channel aggregation is performed based on the plurality of backbone feature maps and sequence feature maps of corresponding sizes to obtain a plurality of target feature maps of corresponding sizes; including: The first backbone network, the second backbone network and the head network are used to perform backbone feature map extraction, sequence feature map extraction and channel aggregation respectively; the first backbone network is used to obtain the image at the current moment and perform multiple feature map size compression and feature extraction on the image at the current moment in sequence to obtain n backbone feature maps of different sizes; the second backbone network is used to perform multiple image size and dimension compression and feature extraction on the input image sequence to obtain n sequence feature maps of corresponding sizes; the head network is used to perform multiple channel aggregation based on the backbone feature map and the sequence feature map to obtain n target feature maps of corresponding sizes; The n is 3; after compressing the number of channels and upsampling the third backbone feature map, the third backbone feature map is subjected to channel aggregation with the second backbone feature map, and feature extraction is performed on the feature map after channel aggregation to obtain a first fused feature map; after compressing the number of channels and expanding the size of the first fused feature map, the third backbone feature map and the first sequence feature map are subjected to channel aggregation to obtain a first target feature map; based on the first target feature map, the second sequence feature map and the first fused feature map, channel aggregation is performed to obtain a second target feature map; based on the second target feature map, the third sequence feature map and the third backbone feature map, channel aggregation is performed to obtain a third target feature map; The target detection result is obtained by fusing multiple target feature maps.

2. The target detection method based on image sequence according to claim 1, characterized in that: The second backbone network includes 5 SpatialConv3D modules connected in sequence, and 3 TemporalConv3D modules connected to the 3rd, 4th and 5th SpatialConv3D modules respectively; The SpatialConv3D module is used to compress the feature map size of the input image sequence; The TemporalConv3D module is used to extract image sequence features from the compressed image sequences output by the corresponding SpatialConv3D module, and compress them into one sequence feature map.

3. The target detection method based on image sequence according to claim 2, characterized in that: The SpatialConv3D module includes a Conv3d layer, a BatchNorm3d layer and a SiLU layer which are arranged in sequence; by setting the convolution kernel parameters of the Conv3d layer, the size of the input image sequence is changed from Adjust to , where N is the number of feature maps in the image sequence, , is the length and width of the input feature map, , is the length and width of the output feature map.

4. The target detection method based on image sequence according to claim 2, characterized in that: The TemporalConv3D module includes: a Conv3d layer, a BatchNorm3d layer, a SiLU layer and a squeeze layer; the Conv3d layer and the squeeze layer are used to extract features and compress dimensions of the input image sequence, respectively, and the size and dimension of the input feature map sequence are reduced by Adjust to , Get a size of Sequence feature diagram.

5. The target detection method based on image sequence according to claim 2, characterized in that: After performing feature extraction and dimension compression on the feature map sequences of different sizes output by the 3rd, 4th and 5th SpatialConv3D modules respectively through 3 TemporalConv3D modules, the method also includes extracting features of interest from the compressed feature maps through the CA attention layer to obtain the first sequence feature map, the second sequence feature map and the third sequence feature map.

6. The target detection method based on image sequence according to claim 2, characterized in that: The first backbone network includes an ImageSub layer, a first feature map compression extraction module, a second feature map compression extraction module, a third feature map compression extraction module and a fourth feature map compression extraction module connected in sequence; The ImageSub layer is used to obtain the current moment image in the image sequence and output it to the first feature map compression extraction module; The first feature map compression and extraction module includes a primary visual cortex simulation module, a Conv layer and a CA attention layer connected in sequence; and is used to perform bionic visual feature extraction on the input current moment feature map to obtain a bionic visual feature map, and input the bionic visual feature map into the second feature map compression and extraction module for feature compression and feature extraction; The second feature map compression extraction module and the third feature map compression extraction module both include a Conv layer, a C3 layer and a CA attention layer connected in sequence; the input feature map is compressed in size through the Conv layer and then input into the corresponding C3 layer for feature extraction to obtain a first backbone feature map and a second backbone feature map; and the first backbone feature map and the second backbone feature map are respectively extracted with features of interest through the corresponding CA attention layer as inputs to the next layer feature map compression extraction module; The fourth feature map compression and extraction module includes a Conv layer, a C3 layer, an SPPF layer and a CA attention layer connected in sequence; and is used to perform feature map size compression, feature map spatial information fusion and feature extraction on the input feature map to obtain a third backbone feature map.

7. The target detection method based on image sequence according to claim 1, characterized in that: The head network includes four serially connected channel feature fusion modules; wherein, The first channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the third backbone feature map, perform channel aggregation with the second backbone feature map, and perform feature extraction on the feature map after channel aggregation through the C3 layer to obtain a first fused feature map; The second channel feature fusion module includes a Conv layer, an Upsamle layer, a Concat layer and a C3 layer connected in sequence; it is used to compress and upsample the number of feature map channels of the first fusion feature map, perform channel aggregation with the first backbone feature map and the first sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain a first target feature map; The third channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the first target feature map, aggregate the channels of the image obtained by compressing the number of feature map channels of the first fusion feature map and the second sequence feature map, and extract features from the feature map after channel aggregation through the C3 layer to obtain the second target feature map; The fourth channel feature fusion module includes a Conv layer, a Concat layer and a C3 layer connected in sequence; it is used to compress the number of feature map channels of the second target feature map, perform channel aggregation with the image obtained by compressing the number of feature map channels of the third backbone feature map and the third sequence feature map, and perform feature extraction on the feature map after channel aggregation through the C3 layer to obtain a third target feature map; The first target feature map, the second target feature map and the third target feature map are subjected to feature fusion through the Detect layer to obtain a target detection result.

8. An electronic device, characterized in that: comprising at least one processor, and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the target detection method based on image sequence according to any one of claims 1-7.

Citation Information

Patent Citations

  • Attention mechanism introduced small target detection method, device and equipment

    CN115035563A

  • Visual cortex imitated multi-scale small target detection method, device and equipment

    CN115035565A