A target detection method and device based on convolution of the primary visual cortex

Through the object detection method based on primary visual cortex convolution, through multiple feature map size compression and channel aggregation, the brain visual cortex processing is simulated, which solves the problem of existing methods being susceptible to noise interference, and improves environmental adaptability and detection performance.

CN116403086BActive Publication Date: 2025-07-11COMP APPL TECH INST OF CHINA NORTH IND GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310286272.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-07-11
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

The existing object detection methods are susceptible to noise interference and have poor environmental adaptability.

Method used

The object detection method based on primary visual cortex convolution is adopted, and the information processing process of the primary visual cortex primary visual cortex mechanism is simulated through the feature map size compression and feature extraction of multiple primary visual cortex mechanisms, combined with the primary visual cortex convolution module and head network.

Benefits of technology

It improves the environmental adaptability of the target detection method under noise interference and improves the target detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403086B_ABST
    Figure CN116403086B_ABST
Patent Text Reader

Abstract

The present invention provides an object detection method and device based on convolution of the primary visual cortex. The detection method includes: obtaining an image to be detected; successively performing multiple times of feature map size compression and feature extraction based on the primary visual cortex mechanism on the image to be detected to obtain multiple backbone feature maps of different sizes; performing multiple times of convolution feature extraction and channel aggregation based on the primary visual cortex mechanism on the multiple backbone feature maps to obtain object feature maps of multiple sizes; and fusing the multiple object feature maps to obtain an object detection result. The present invention improves the problems in the prior art that the object detection method is easily interfered by noise and has poor environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence target detection, and in particular relates to a target detection method and device based on convolution of the primary visual cortex. Background Art

[0002] The task of target detection is to detect target objects of interest in static images (or dynamic videos), which can provide good conditions for subsequent target tracking, behavior prediction, etc. The target detection task is an important part of computer vision and is widely used in many fields such as unmanned driving, robot navigation, video surveillance, industrial inspection, and aerospace.

[0003] Existing target detection methods mainly include Faster-RCNN, SSD, CenterNet, and YOLOv1 to YOLOv5, etc. However, existing methods still have problems of being vulnerable to noise interference and poor environmental adaptability. Summary of the Invention

[0004] In view of the above analysis, the present invention aims to provide a target detection method and device based on convolution of the primary visual cortex to improve the problems of vulnerability to noise interference and poor environmental adaptability in existing target detection methods.

[0005] The object of the present invention is mainly achieved by the following technical solutions:

[0006] On the one hand, the present invention provides a target detection method based on convolution of the primary visual cortex, the method comprising:

[0007] Obtain an image to be detected;

[0008] Based on the image to be detected, perform multiple times of feature map size compression and feature extraction of the primary visual cortex mechanism in sequence to obtain multiple backbone feature maps of different sizes;

[0009] Based on multiple backbone feature maps, perform multiple times of convolution feature extraction and channel aggregation of the primary visual cortex mechanism to obtain target feature maps of multiple sizes;

[0010] Fuse multiple target feature maps to obtain a target detection result.

[0011] Further, extract backbone feature maps through a preset backbone network;

[0012] The backbone network includes n + 1 feature map compression and extraction modules including primary visual cortex convolution modules connected in sequence, for performing multiple times of feature map size compression and feature extraction of the primary visual cortex mechanism on the received feature map to obtain n backbone feature maps of different sizes.

[0013] Further, convolutional feature extraction and channel aggregation of the primary visual cortex mechanism are performed through a preset head network;

[0014] The head network includes multiple primary visual cortex convolutional modules and channel aggregation modules; multiple convolutional feature extractions or feature map size compressions of the primary visual cortex mechanism are performed based on multiple backbone feature maps; and channel aggregation is performed by the aggregation module based on the backbone feature map and the feature maps output by the primary visual cortex convolutional modules to obtain n target feature maps; the convolutional feature extraction and feature map size compression are achieved by setting different parameters of the primary visual cortex convolutional modules.

[0015] Further, the primary visual cortex convolutional module includes: a first Conv layer, a VOneBlock layer, a second Conv layer, a third Conv layer, and a feature fusion layer;

[0016] After the first Conv layer, the VOneBlock layer, and the second Conv layer are serially connected in sequence, they are connected in parallel with the third Conv layer;

[0017] The feature fusion layer is used to perform feature fusion on the feature maps output by the second Conv layer and the third conv layer.

[0018] Further, the first Conv layer, the second Conv layer, and the third Conv layer are used to perform convolution operations and channel number adjustments on the received feature maps;

[0019] The VOneBlock layer is used to perform primary visual cortex feature extraction on the received feature maps;

[0020] The feature maps output by the first Conv layer and the third conv layer have the same size and channel number.

[0021] Further, the number of channels of the feature map output by the first Conv layer is 8 channels, 16 channels, or 32 channels.

[0022] Further, when using the primary visual cortex convolutional module to compress the size of the feature map of the primary visual cortex mechanism, the convolutional kernel sizes of the first Conv layer and the third conv layer are set to 3x3, and the sliding window stride is 2 or 4;

[0023] When using the primary visual cortex convolutional module to perform convolutional feature extraction of the primary visual cortex mechanism, the convolutional kernel sizes of the first Conv layer and the third conv layer are set to 1*1, and the sliding window stride is 1.

[0024] Further, n is 3;

[0025] The first to third feature map compression and extraction modules each include a primary visual cortex convolution module and a C3 layer connected in sequence; after the input feature map is size-compressed by the primary visual cortex convolution module respectively, the compressed feature map is input into the corresponding C3 layer for feature extraction to obtain the output of the corresponding feature map compression and extraction module; based on the outputs of the second and third feature map compression and extraction modules, a first backbone feature map and a second backbone feature map are obtained;

[0026] The fourth feature map compression and extraction module includes a primary visual cortex convolution module, an SPP layer, and a C3 layer connected in sequence; it is used to perform feature map size compression, feature map spatial information fusion, and feature extraction on the feature map output by the third feature map compression and extraction module in sequence to obtain a third backbone feature map.

[0027] Further, based on multiple said backbone feature maps, convolutional feature extraction and channel aggregation of the primary cortex mechanism are performed multiple times to obtain target feature maps of multiple sizes, including:

[0028] After performing visual cortex convolutional feature extraction and size expansion on the third backbone feature map, channel aggregation is performed with the second backbone feature map, and feature extraction is performed on the feature map after channel aggregation to obtain a first fusion feature map;

[0029] After performing visual cortex convolutional feature extraction and size expansion on the said first fusion feature map, channel aggregation is performed with the first backbone feature map, and feature extraction is performed on the feature map after channel aggregation to obtain a first target feature map;

[0030] After performing visual cortex convolutional feature extraction on the first target feature map and the first fusion feature map respectively and then performing channel aggregation, and feature extraction is performed on the feature map after channel aggregation to obtain a second target feature map;

[0031] After performing visual cortex convolutional feature extraction on the second target feature map and the third backbone feature map respectively and then performing channel aggregation, and feature extraction is performed on the feature map after channel aggregation to obtain a third target feature map.

[0032] On the other hand, an electronic device is also provided, including at least one processor and at least one memory communicatively connected to the processor;

[0033] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the aforementioned object detection method based on primary visual cortex convolution.

[0034] The beneficial effects of this technical solution:

[0035] 1. The present invention uses the VOneBlock module of the primary visual cortex to perform multiple convolution feature extractions and feature map size compressions on the image to be detected based on the mechanism of the primary visual cortex of the brain, which can more effectively simulate the information processing process of the primary visual cortex of the brain and improve the environmental adaptability of the target detection method under noise interference.

[0036] 2. The present invention uses a preset primary visual cortex convolution module to solve the problem of low processing efficiency of the VOneBlock module in high-computation applications, realizes the function of the VOneBlock module for multi-size and multi-channel feature extraction, and further improves the target detection performance.

[0037] Other features and advantages of the present invention will be described in the following specification, and some will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0038] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0039] Figure 1 is the flowchart of the target detection method based on primary visual cortex convolution according to an embodiment of the present invention;

[0040] Figure 2 is the schematic structural diagram of the target detection process based on primary visual cortex convolution according to an embodiment of the present invention;

[0041] Figure 3 is the structural diagram of the primary visual cortex convolution module according to an embodiment of the present invention. Detailed Embodiments

[0042] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.

[0043] An embodiment of the present invention provides a target detection method based on primary visual cortex convolution, as Figure 1 shown, including the following steps:

[0044] Step S1, obtain the image to be detected.

[0045] Specifically, the image to be detected can be an image taken in any way, such as an image obtained by aerial photography with a drone, including ground vehicles, personnel, mobile equipment, infrastructure, and other objects to be detected.

[0046] Step S2: Based on the image to be detected, perform multiple times of feature map size compression and feature extraction of the primary visual cortex mechanism in sequence to obtain multiple backbone feature maps of different sizes.

[0047] Specifically, extract backbone feature maps through a preset backbone network. The backbone network includes n + 1 feature map compression and extraction modules including primary visual cortex convolution modules connected in sequence, which are used to perform multiple times of feature map size compression and feature extraction of the primary visual cortex mechanism on the received feature map to obtain n backbone feature maps of different sizes.

[0048] More specifically, in this embodiment, the value of n is 3.

[0049] As Figure 2 shown, the first to third feature map compression and extraction modules in the backbone network each include a primary visual cortex convolution module and a C3 layer connected in sequence. After the input feature map is size-compressed by the primary visual cortex convolution module respectively, the compressed feature map is input into the corresponding C3 layer for feature extraction to obtain the output of the corresponding feature map compression and extraction module. Based on the outputs of the second and third feature map compression and extraction modules, the first backbone feature map and the second backbone feature map are obtained.

[0050] The fourth feature map compression and extraction module includes a primary visual cortex convolution module, an SPP layer, and a C3 layer connected in sequence, and is used to perform feature map size compression, feature map spatial information fusion, and feature extraction on the feature map output by the third feature map compression and extraction module in sequence to obtain the third backbone feature map.

[0051] Preferably, as Figure 3 shown, the primary visual cortex convolution module includes: a first Conv layer, a VOneBlock layer, a second Conv layer, a third Conv layer, and a feature fusion layer. Among them, after the first Conv layer, the VOneBlock layer, and the second Conv layer are serially connected in sequence, they are connected in parallel with the third Conv layer. The feature fusion layer is used to perform feature fusion on the feature maps output by the second Conv layer and the third Conv layer. The first Conv layer, the second Conv layer, and the third Conv layer are used to perform convolution operations and channel number adjustment on the received feature maps. The VOneBlock layer is used to perform visual cortex feature extraction on the received feature maps. The sizes and channel numbers of the feature maps output by the first Conv layer and the third Conv layer are the same.

[0052] It should be noted that the VOneBlock layer is a neural network layer constructed based on the primary visual cortex of primates. Using the Gabor filter as the core component, it simulates the information processing mechanism of the human visual perception cortex to extract visual features from the input image, and can obtain a feature map closer to the features after human brain visual processing. In the prior art, due to the computational efficiency limitation of the VOneBlock layer for high computational amounts, usually only the 3-channel color image or 1-channel grayscale image collected can be used for preliminary feature extraction by the VOneBlock layer. In this embodiment, through a preset primary visual cortex convolution module, a first Conv layer is added before the VOneBlock layer to adjust the number of channels of the feature map input to the VOneBlock layer; preferably, the number of channels of the feature map output by the first Conv layer can be set to 8 channels, 16 channels or 32 channels to achieve a computational amount suitable for the VOneBlock layer and balance the problems of feature extraction accuracy and computational efficiency.

[0053] Preferably, when using the primary visual cortex convolution module to compress the size of the feature map of the primary visual cortex mechanism, the convolution kernel sizes of the first Conv layer and the third conv layer are set to 3x3, and the sliding window stride is 2 or 4.

[0054] As a specific embodiment, the sizes of the image to be detected obtained and the feature map obtained after feature extraction can be expressed as Hi×Wi×Ci, where Hi is the height of the feature map, Wi is the width of the feature map, and Ci is the number of channels of the feature map; in this embodiment, the size of the image to be detected received by the first feature map compression and extraction module in the backbone network is H×W×3. The convolutional kernel sizes of the first Conv layer and the third conv layer in the primary visual cortex convolutional module in the first feature map compression and extraction module are set to 3x3, and the sliding window stride is set to 4. After the primary visual cortex convolutional module performs size compression and channel number adjustment, a feature map with a size of H / 4×W / 4×128 is obtained; the obtained feature map is input into the second feature map compression and extraction module after feature extraction through the C3 attention layer; the convolutional kernel sizes of the first Conv layer and the third conv layer in the primary visual cortex convolutional module in the second feature map compression and extraction module are set to 3x3, and the sliding window stride is set to 2. After the primary visual cortex convolutional module performs size compression and channel number adjustment and feature extraction through the C3 layer, a first backbone feature map with a size of H / 8×W / 8×256 is obtained; the first backbone feature map is input into the third feature map compression and extraction module, and successively undergoes size compression, channel number adjustment, and feature extraction through the primary visual cortex convolutional module and the C3 layer, obtaining a second backbone feature map with a size of H / 16×W / 16×512. The second backbone feature map is input into the fourth feature map compression and extraction module, and successively undergoes size compression, channel number adjustment, and feature extraction through the primary visual cortex convolutional module and the C3 layer, obtaining a third backbone feature map with a size of H / 32×W / 32×1024.

[0055] In this embodiment, the VOneBlock module is successfully applied to the middle layer of the backbone network, realizing feature extraction closer to human brain vision multiple times and improving the performance of anti-noise interference.

[0056] Step S3: Based on multiple backbone feature maps, perform convolutional feature extraction and channel aggregation of the primary visual cortex mechanism multiple times to obtain target feature maps of multiple sizes.

[0057] Specifically, perform convolutional feature extraction and channel aggregation of the primary visual cortex mechanism through a preset head network; the head network includes multiple primary visual cortex convolutional modules and channel aggregation modules; perform convolutional feature extraction of the primary visual cortex mechanism or feature map size compression based on multiple backbone feature maps multiple times; and perform channel aggregation on the backbone feature maps and the feature maps output by the primary visual cortex convolutional modules through the aggregation module to obtain n target feature maps; the convolutional feature extraction and feature map size compression are achieved by setting different parameters of the primary visual cortex convolutional module.

[0058] Preferably, when using the primary visual cortex convolution module to extract convolution features of the primary visual cortex mechanism, the convolution kernel sizes of the first Conv layer and the third conv layer are set to 1*1, and the sliding window stride is 1.

[0059] More specifically, as Figure 2 shown, multiple convolution feature extractions and channel aggregations of the primary cortex mechanism are performed based on multiple backbone feature maps to obtain target feature maps of multiple sizes, including:

[0060] Perform visual cortex convolution feature extraction on the third backbone feature map through the primary visual cortex convolution module V1 to obtain a feature map with a size of H / 32×W / 32×512, and further expand the size through the first Upsample layer to obtain a feature map with a size of H / 16×W / 16×512; Aggregate the channels of the feature map with the expanded size and the second backbone feature map through the first Concat layer to obtain a feature map with a size of H / 16×W / 16×1024, and perform feature extraction on the channel-aggregated feature map through the fifth C3 layer to obtain the first fusion feature map with a size of H / 16×W / 16×512.

[0061] Perform visual cortex convolution feature extraction on the first fusion feature map through the primary visual cortex convolution module V6 to obtain a feature map with a size of H / 16×W / 16×256, further expand the size through the second Upsample layer, and aggregate the channels of the feature map with a size of H / 8×W / 8×256 and the first backbone feature map through the second Concat layer to obtain a feature map with a size of H / 8×W / 8×512, and perform feature extraction and channel adjustment on the channel-aggregated feature map through the sixth C3 layer to obtain the first target feature map with a size of H / 8×W / 8×256;

[0062] Compress the size of the first target feature map through the primary visual cortex convolution module V7, and perform visual cortex convolution feature extraction on the first fusion feature map through the sixth primary visual cortex convolution module, both obtaining feature maps with a size of H / 16×W / 16×256. Aggregate the channels of the two obtained feature maps through the third Concat layer to obtain a feature map with a size of H / 16×W / 16×512; Further perform feature extraction on the channel-aggregated feature map through the seventh C3 layer to obtain the second target feature map with a size of H / 16×W / 16×512;

[0063] The second target feature map is size compressed by the primary visual cortex convolution module V8, and the third backbone feature map is subjected to visual cortex convolution feature extraction by the fifth primary visual cortex convolution module, and both feature maps with a size of H / 32×W / 32×512 are obtained. The two feature maps are channel-aggregated by the fourth Concat layer to obtain a feature map with a size of H / 32×W / 32×1024. The feature map after channel aggregation is subjected to feature extraction by the eighth C3 layer to obtain a third target feature map with a size of H / 32×W / 32×1024.

[0064] Step S4: Fuse multiple target feature maps to obtain target detection results.

[0065] Specifically, the three target feature maps are subjected to feature fusion through the Detect layer to achieve target classification and coordinate positioning. This embodiment uses the Detect layer of the existing YOLOv5 model to perform target classification and coordinate positioning to obtain the target detection result.

[0066] In summary, the present invention uses the VOneBlock module of the primary visual cortex to perform multiple convolution feature extractions and feature map size compressions on the image to be detected based on the mechanism of the primary visual cortex of the brain, which can more effectively simulate the information processing process of the primary visual cortex of the brain and improve the environmental adaptability of the target detection method under noise interference. In addition, the primary visual cortex convolution module is creatively proposed to solve the problem of low processing efficiency of the VOneBlock module in high-computation applications, and realize the function of the VOneBlock module to extract multi-size and multi-channel features, further improving the target detection performance.

[0067] A second embodiment of the present invention further discloses an electronic device, comprising at least one processor, and at least one memory communicatively connected to the processor;

[0068] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the target detection method based on primary visual cortex convolution in the aforementioned embodiment.

[0069] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0070] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A target detection method based on convolution of the primary visual cortex, characterized in that Including: Obtain the image to be detected; Based on the image to be detected, perform feature map size compression and feature extraction of the primary visual cortex mechanism multiple times in sequence to obtain multiple backbone feature maps of different sizes, including: extracting backbone feature maps through a preset backbone network; the backbone network includes n + 1 feature map compression and extraction modules each including a primary visual cortex convolution module, which are used to perform feature map size compression and feature extraction of the primary visual cortex mechanism on the received feature map multiple times to obtain n backbone feature maps of different sizes; The primary visual cortex convolution module includes: a first Conv layer, a VOneBlock layer, a second Conv layer, a third Conv layer, and a feature fusion layer; after the first Conv layer, the VOneBlock layer, and the second Conv layer are serially connected in sequence, they are connected in parallel with the third Conv layer; the feature fusion layer is used to perform feature fusion on the feature maps output by the second Conv layer and the third conv layer; the number of channels of the feature map output by the first Conv layer is 8 channels, 16 channels, or 32 channels; when using the primary visual cortex convolution module to compress the size of the feature map of the primary visual cortex mechanism, the convolution kernel size of the first Conv layer and the third conv layer is set to 3x3, and the sliding window step size is 2 or 4; when using the primary visual cortex convolution module to extract the convolution features of the primary visual cortex mechanism, the convolution kernel size of the first Conv layer and the third conv layer is set to , and the sliding window step size is 1; Based on multiple said backbone feature maps, perform convolution feature extraction and channel aggregation of the primary visual cortex mechanism multiple times to obtain target feature maps of multiple sizes; Fuse multiple said target feature maps to obtain the target detection result.

2. The object detection method based on convolution of the primary visual cortex according to claim 1, wherein Perform convolution feature extraction and channel aggregation of the primary visual cortex mechanism through a preset head network; The head network includes multiple primary visual cortex convolution modules and channel aggregation modules; perform convolution feature extraction or feature map size compression of the primary visual cortex mechanism multiple times based on multiple said backbone feature maps; and perform channel aggregation on the backbone feature maps and the feature maps output by the primary visual cortex convolution modules through the aggregation module to obtain n target feature maps; the convolution feature extraction and feature map size compression are achieved by setting different parameters for the primary visual cortex convolution modules.

3. The object detection method based on convolution of the primary visual cortex according to claim 1, characterized in that The first Conv layer, the second Conv layer, and the third Conv layer are used to perform convolution operations and channel number adjustment on the received feature maps; The VOneBlock layer is used to perform visual cortex feature extraction on the received feature maps; The feature map sizes and channel numbers output by the first Conv layer and the third conv layer are the same.

4. The object detection method based on convolution of the primary visual cortex according to claim 1, wherein The n is 3; The 1st - 3rd feature map compression and extraction modules each include a primary visual cortex convolution module and a C3 layer connected in sequence; after respectively performing size compression on the input feature maps through the primary visual cortex convolution module, input the compressed feature maps into the corresponding C3 layer for feature extraction to obtain the output of the corresponding feature map compression and extraction module; based on the outputs of the 2nd feature map compression and extraction module and the 3rd feature map compression and extraction module, obtain the first backbone feature map and the second backbone feature map; The 4th feature map compression and extraction module includes a primary visual cortex convolution module, an SPP layer, and a C3 layer connected in sequence; it is used to perform feature map size compression, feature map spatial information fusion, and feature extraction on the feature map output by the 3rd feature map compression and extraction module in sequence to obtain the third backbone feature map.

5. The object detection method based on convolution of the primary visual cortex according to claim 4, characterized in that, Based on multiple said backbone feature maps, perform convolution feature extraction and channel aggregation of the primary cortex mechanism multiple times to obtain target feature maps of multiple sizes, including: After performing visual cortex convolution feature extraction and size expansion on the third backbone feature map, perform channel aggregation with the second backbone feature map, and perform feature extraction on the feature map after channel aggregation to obtain the first fusion feature map; After performing visual cortex convolution feature extraction and size expansion on the first fusion feature map, channel aggregation is carried out with the first backbone feature map, and feature extraction is performed on the feature map after channel aggregation to obtain the first target feature map; After performing visual cortex convolution feature extraction on the first target feature map and the first fusion feature map respectively and then carrying out channel aggregation, and performing feature extraction on the feature map after channel aggregation to obtain the second target feature map; After performing visual cortex convolution feature extraction on the second target feature map and the third backbone feature map respectively and then carrying out channel aggregation, and performing feature extraction on the feature map after channel aggregation to obtain the third target feature map.

6. An electronic device, characterized in that, Comprising at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the object detection method based on primary visual cortex convolution according to any one of claims 1-5.

Citation Information

Patent Citations

  • Visual cortex imitated multi-scale small target detection method, device and equipment

    CN115035565A

  • Three-dimensional target detection method and system

    CN115775379A