Small target detection method inspired by visual ventral pathway neural mechanism

This method for small target detection, which simulates the neural mechanism of the ventral visual pathway, solves the problem of insufficient accuracy in small target detection in complex backgrounds. By simulating the receptive field and self-feedback regulation mechanism of the visual pathway, it achieves higher detection accuracy and recall.

CN120997473APending Publication Date: 2025-11-21GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510831541.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In complex contexts, small target detection is prone to problems such as background interference and inter-class similarity due to its small size and low boundary contrast. Existing technologies struggle to effectively extract detailed feature information, resulting in insufficient detection accuracy.

Method used

We construct a small target detection method inspired by the neural mechanisms of the ventral visual pathway. By simulating the neuronal mechanisms of the ventral visual pathway, including the frontal network RLV, backbone network, GLRF module, Neck network, and Head network, we employ multi-scale feature fusion and self-feedback attention module SFA to simulate the receptive field and self-feedback regulation mechanism of the visual pathway, extract detailed features, and suppress background interference.

Benefits of technology

It effectively extracts detailed feature information, improving the accuracy and recall of small target detection. In particular, it can better distinguish targets from backgrounds in complex environments, thus enhancing detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention aims to provide a small target detection method inspired by a visual ventral pathway neural mechanism, and the method comprises the following steps: A, constructing a neural network which comprises a front network RLV, a Backbone network module, a GLRF module, a Neck network module, and a Head network module; b, performing feature enhancement on the original image through a front network RLV, inputting the original image into a Backbone network, and sequentially performing four down-sampling processing in the Backbone network to obtain a first down-sampling result, a second down-sampling result, a third down-sampling result and a fourth down-sampling result; wherein the second down-sampling result is extracted through the GLRF module to obtain global and local features of the small target, and the global and local features are input into the Neck network module; second to fourth down-sampling results are respectively input into the Neck network module; and C, performing multi-scale feature fusion on each input result by the Neck network module to obtain two features with different resolutions, and inputting the features into the Head network module for processing to obtain a final result. According to the invention, detail feature information can be effectively extracted, so that a more accurate detection effect is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer image processing, in particular to a small target detection method inspired by the visual ventral pathway neural mechanism. BACKGROUND

[0002] Target detection, as a core research direction in the field of computer vision, together with image classification and semantic segmentation, constitutes the three basic tasks in this field. The core goal is to accurately classify instances in images and locate them through bounding boxes. Aerial image detection, as an important application scenario of target detection, faces unique challenges such as small target size, dense distribution, and complex background, and places higher demands on the accuracy and real-time performance of detection algorithms. In recent years, with the breakthrough of deep learning technology, target detection algorithms based on convolutional neural networks (CNN) have developed rapidly. Current mainstream methods generally adopt a single-stage architecture of "main network-neck-detection head", and through optimizing feature extraction and designing efficient feature fusion strategies, they have made significant progress in performance improvement. Experimental results show that such models perform excellently on the VisDrone dataset for small target detection in aerial images, effectively detecting tiny targets in complex scenes, and providing solid technical support for practical applications such as remote sensing monitoring and intelligent inspection.

[0003] Small target detection is an important and challenging research direction in the field of aerospace. In complex backgrounds, small targets are prone to background interference and inter-class similarity due to their small size and low boundary contrast. Physiological studies have shown that the visual pathway can extract features such as contour, shape, and color through its own neural mechanisms, effectively filtering out background noise and recognizing complex shapes. Therefore, it is necessary to conduct in-depth research on small target detection. SUMMARY

[0004] The present application aims to provide a small target detection method inspired by the visual ventral pathway neural mechanism, which can effectively extract detailed feature information and achieve more accurate detection results.

[0005] The technical solution of the present application is as follows:

[0006] The small target detection method inspired by the visual ventral pathway neural mechanism comprises the following steps:

[0007] A. Construct a neural network, which comprises a pre-network RLV, a Backbone network module, a GLRF module, a Neck network module, and a Head network module.

[0008] The pre-network RLV simulates a simple cell receptive field in the V1 area, four down-samplings are arranged in the Backbone network, the GLRF module simulates a global-local receptive field, and the path aggregation network PAN is arranged in the Neck network module;

[0009] B, the original image is first subjected to feature enhancement by the pre-network RLV and then input into the Backbone network, and sequentially subjected to four down-sampling processes in the Backbone network to obtain first-four down-sampling results;

[0010] The second down-sampling result is subjected to the GLRF module to extract global and local features of small targets, and the global and local features of small targets are input into the Neck network module;

[0011] The second-four down-sampling results are also input into the Neck network module;

[0012] C, after the Neck network module performs multi-scale feature fusion on the input results, two features with different resolutions are obtained and input into the Head network module for processing to obtain a series of detection frames and category data, which are the final results.

[0013] The processing process in the pre-network RLV is as follows:

[0014] The input result is divided into two paths, the first path is converted into a gray image and then subjected to convolution processing, and the obtained convolution results are subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing, respectively, the two results obtained after being spliced by the Concat function are subjected to V1 simulation module processing to obtain the first path processing result;

[0015] In the second path, the input result is first divided into R, G and B single-path color images, the R, G and B single-path color images are subjected to convolution processing respectively to obtain R channel convolution results, G channel convolution results and B channel convolution results, the three convolution results are added two by two to obtain three two-by-two addition results, the spliced result obtained after the three two-by-two addition results are spliced by the Concat function is subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing, respectively, and the two results obtained after being spliced by the Concat function are subjected to V1 simulation module processing to obtain the second path processing result;

[0016] The first path processing result and the second path processing result are spliced by the Concat function to obtain the output result.

[0017] The processing process in the OFF-center Sigmoid module is as follows:

[0018] The input result is respectively processed by 1*1 convolution and 3*3*1 convolution, and the output result is obtained by subtracting the 1*1 convolution result from the 3*3*1 convolution result.

[0019] The processing process in the ON-center Sigmoid module is as follows:

[0020] The input result is respectively processed by 1*1 convolution and 3*3*1 convolution, and the output result is obtained by subtracting the 3*3*1 convolution result from the 1*1 convolution result.

[0021] The V1 simulation module is based on a 3*3 convolution module, and the middle three weights are used to suppress the weights on both sides, and the weight suppression calculation formula is as follows:

[0022]

[0023] Wherein, W is the convolution kernel weight, W diagonal is the adjacent diagonal weight of W j , j∈(1, 2, 3, 4), and α is set to 20.

[0024] The processing process in the GLRF module is as follows:

[0025] The input result is divided into three paths; the first path becomes a gray image, and the gray image is processed by two branches; the first branch is sequentially processed by 3*1*1*1 hollow convolution, 3*1*2*2 hollow convolution and 3*1*5*5 hollow convolution to obtain the first branch result; the second branch is sequentially processed by two-dimensional average pooling, 1*1 convolution and up-sampling to obtain the second branch result; the gray image, the first branch result and the second branch result are processed by the Concat function to obtain the first path result;

[0026] In the second path, the input result is sequentially processed by 3*1*1 convolution, 3*2*1 convolution, 3*1*1 convolution, 1*1 convolution and up-sampling to obtain an up-sampling result, and the input result is subtracted from the up-sampling result, and sequentially processed by the CBAM module, up-sampling, 3*1*1 convolution and down-sampling to obtain the second path result;

[0027] In the third path, the input result is processed by the RB module to obtain the third path result;

[0028] The first path result and the second path result are spliced by the Concat function, and then processed by 1*1 convolution, and the obtained result and the third path result are spliced by the Concat function to obtain the output result.

[0029] The processing process in the RB module is as follows:

[0030] The input result is divided into four paths, each path is processed by 3x3 convolution, the obtained four results are spliced by a Concat function, then processed by 1x1 convolution, the obtained result is added to the input result, and the output result is obtained.

[0031] The processing process in the Neck network module is as follows:

[0032] The fourth down-sampling result is up-sampled, spliced with the third down-sampling result by a Goncat function, the obtained result is up-sampled to obtain an intermediate result

[0033] The intermediate result is spliced with the second down-sampling result and the global and local features of the small target by a Goncat function to obtain a first output result.

[0034] The first output result is down-sampled, spliced with the intermediate result by a Goncat function to obtain a second output result.

[0035] The first output result and the second output result are respectively input into the Head network module.

[0036] The processing process in the Head network module is as follows:

[0037] The first output result and the second output result are respectively processed by a self-feedback attention module SFA, and then input into an ITHead detection head for detection to obtain a series of detection frames and category data, which is the final result.

[0038] The processing process in the self-feedback attention module SFA is as follows:

[0039] The input result is processed by 1x1 convolution to obtain an Output1 result, and the Output1 result is processed by a Conv block module, and the obtained result is multiplied by the Output1 result to obtain an Output2 result.

[0040] The Output1 result is subtracted from the Output2 result to obtain an Output3 result, the Output3 result is processed by a Conv block module, and the obtained result is multiplied by the Output3 result to obtain an Output4 result, and the Output4 result is the output result.

[0041] The processing process in the Conv block module is as follows:

[0042] The input result is processed by 3x3 convolution, 1x1 convolution, 3x3 convolution and Sigmoid function in sequence to obtain the output result.

[0043] The processing process in the ITHead detection head is as follows:

[0044] The input result is divided into two branches, the first branch sequentially passes through 3*1*1 convolution, 3*1*1 convolution and 1*1 convolution processing to obtain a positioning loss; the positioning loss sequentially passes through left-up direction convolution, left-down direction convolution, right-up direction convolution and right-down direction convolution processing, the four direction convolution results obtained are spliced through a Concat function, then sequentially pass through a CBAM module and a Sigmoid function processing to obtain a first branch result; the second branch sequentially passes through 3*1*1 convolution, 3*1*1 convolution and 1*1 convolution processing to obtain a second branch result, the first branch result is multiplied by the second branch result to obtain a classification loss; the positioning loss and the classification loss are the detection result.

[0045] The left-up direction convolution is based on a 3*3 convolution module, and the four weights at the left-up are used to suppress the weights on each side; the left-down direction convolution is based on a 3*3 convolution module, and the four weights at the left-down are used to suppress the weights on each side; the right-up direction convolution is based on a 3*3 convolution module, and the four weights at the right-up are used to suppress the weights on each side; and the right-down direction convolution is based on a 3*3 convolution module, and the four weights at the right-down are used to suppress the weights on each side.

[0046] The calculation formula of the weight suppression is as follows:

[0047]

[0048] Wherein, W is the convolution kernel weight, W diagonal is the adjacent diagonal weight of W j , j is (1, 2, 3, 4), and alpha is set to 20.

[0049] The beneficial effects of the present application are as follows:

[0050] The method of the present application designs a pre-network RLV to simulate the "Retina-LGN-V1" channel, and divides the input image into a gray image and a color image with an RGB three-channel to be transmitted into a feature enhancement module to simulate the cone and rod cells.

[0051] The present application adopts the adjacent weight inhibition method to simulate the V1 area simple cell receptive field, and highlights the excitability of the middle region and suppresses the surrounding area by the adjacent weight inhibition method, thereby simulating the receptive field mechanism of the V1 area simple cell.

[0052] The present application also simulates the global-local receptive field mechanism through the GLRF module, and extracts shape and more delicate color features.

[0053] The present application designs a self-feedback attention module SFA to simulate the V4 area self-feedback regulation mechanism to realize the modulation of feature information, which can effectively suppress irrelevant interference information, separate the target from the background impurities, and provide a basis for subsequent information fusion.

[0054] The present application also designs an ITHead detection head, adopts an anchor-free split detection head, increases a hierarchical modulation structure between Bbox.Loss and Cls.Loss to simulate the orientation selectivity mechanism of the IT area, can make the model better understand the complex shape and extract its features, and provides a basis for extracting non-common features of similar objects between classes.

[0055] The biological inspired remote sensing small target detection network (BRSTD) proposed in the present application can effectively extract detailed feature information, thereby obtaining more accurate detection effect. In addition, the network design of the present application has physiological visual basis support, which is different from general one-stage target detection model, and experiments can show that the XYWA module and the TDSA module proposed by the present application can improve the detection performance. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1Structure diagram of a neural network of embodiment 1 of the present application;

[0057] Figure 2 Structure diagram of a pre-network RLV of embodiment 1;

[0058] Figure 3 Structure diagram of a GLRF module of embodiment 1;

[0059] Figure 4 Structure diagram of a self-feedback attention module SFA of embodiment 1;

[0060] Figure 5 Structure diagram of an ITHead detection head of embodiment 1;

[0061] Figure 6 Structure diagram of four direction convolutions in the ITHead detection head of embodiment 1, wherein (a) left-up direction convolution, (b) right-up direction convolution, (c) left-down direction convolution, (d) right-down direction convolution;

[0062] Figure 7 Comparison diagram of detection effects of the method of embodiment 1 of the present application and prior art; DETAILED DESCRIPTION

[0063] The present application will be described in detail below with reference to the accompanying drawings and embodiments.

[0064] Embodiment 1

[0065] The small target detection method inspired by the visual ventral pathway neural mechanism comprises the following steps:

[0066] A, constructing a neural network, as shown in Figure 1 The neural network comprises a pre-network RLV, a Backbone network module, a GLRF module, a Neck network module and a Head network module.

[0067] The pre-network RLV simulates the simple cell receptive field of V1 area, four down-samplings are arranged in the Backbone network, the GLRF module simulates the global-local receptive field, and the Neck network module is provided with a path aggregation network PAN.

[0068] B, the original image is first input into the pre-network RLV for feature enhancement, then into the Backbone network, and sequentially through four down-sampling processes in the Backbone network to obtain first-four down-sampling results.

[0069] The second down-sampling result is input into the GLRF module to extract the global and local features of the small target, and the global and local features of the small target are input into the Neck network module.

[0070] The second to fourth down-sampling results are also input into the Neck network module respectively.

[0071] As shown in Figure 2 , the processing procedure in the pre-network RLV is as follows:

[0072] The input result is divided into two paths, the first path is converted into a gray image, and then is subjected to convolution processing, and the obtained convolution result is subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing respectively, and then the two results obtained are spliced by a Concat function, and then are subjected to V1 simulation module processing to obtain the first path processing result.

[0073] In the second path, the input result is first divided into R, G and B single-path color images, and then the R, G and B single-path color images are subjected to convolution processing respectively to obtain R channel convolution result, G channel convolution result and B channel convolution result, and then the three convolution results are added two by two to obtain three two-by-two addition results, and then the three two-by-two addition results are spliced by a Concat function to obtain a spliced result, and then the spliced result is subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing respectively, and then the two results obtained are spliced by a Concat function, and then are subjected to V1 simulation module processing to obtain the second path processing result.

[0074] The first path processing result and the second path processing result are spliced by a Concat function to obtain an output result.

[0075] The processing procedure in the OFF-center Sigmoid module is as follows:

[0076] The input result is subjected to 1x1 convolution and 3x3x1 convolution processing respectively, and then the 3x3x1 convolution result is subtracted from the 1x1 convolution result to obtain an output result.

[0077] The processing procedure in the ON-center Sigmoid module is as follows:

[0078] The input result is subjected to 1x1 convolution and 3x3x1 convolution processing respectively, and then the 1x1 convolution result is subtracted from the 3x3x1 convolution result to obtain an output result.

[0079] The V1 simulation module is based on a 3x3 convolution module, and the middle three weights are used to suppress the weights on both sides respectively, and the calculation formula of the weight suppression is as follows:

[0080]

[0081] Wherein, W is the convolution kernel weight, W diagonal is W jthe adjacent diagonal weight of the 2x2 filter, j e (1, 2, 3, 4), where a is set to 20;

[0082] As shown in Figure 3 The processing process in the GLRF module is as follows:

[0083] The input result is divided into three paths; the first path becomes a gray image, and the gray image is processed through two branches; the first branch is sequentially processed through 3x1x1x1 empty convolution, 3x1x2x2 empty convolution, and 3x1x5x5 empty convolution to obtain the first branch result; the second branch is sequentially processed through two-dimensional average pooling, 1x1 convolution, and up-sampling to obtain the second branch result; the gray image, the first branch result, and the second branch result are processed through a Concat function to obtain the first path result;

[0084] In the second path, the input result is sequentially processed through 3x1x1 convolution, 3x2x1 convolution, 3x1x1 convolution, 1x1 convolution, and up-sampling to obtain an up-sampling result; the input result is subtracted from the up-sampling result, and then sequentially processed through a CBAM module, up-sampling, 3x1x1 convolution, and down-sampling to obtain the second path result;

[0085] In the third path, the input result is processed through an RB module to obtain the third path result;

[0086] The first path result and the second path result are concatenated through a Concat function and then processed through 1x1 convolution; the obtained result and the third path result are concatenated through a Concat function to obtain the output result.

[0087] The processing process in the RB module is as follows:

[0088] The input result is divided into four paths, and each path is processed through 3x3 convolution; the obtained four results are concatenated through a Concat function and then processed through 1x1 convolution; the obtained result is added to the input result to obtain the output result.

[0089] The processing process in the Neck network module is as follows:

[0090] The fourth down-sampling result is up-sampled and concatenated with the third down-sampling result through a Goncat function; the obtained result is up-sampled to obtain an intermediate result

[0091] The intermediate result, the second down-sampling result, and the global and local features of the small target are concatenated through a Goncat function to obtain a first output result;

[0092] The first output result is down-sampled and concatenated with the intermediate result through a Goncat function to obtain a second output result;

[0093] The first output result and the second output result are respectively input into a Head network module.

[0094] The Neck network module performs multi-scale feature fusion on each input result to obtain two features with different resolutions, which are input into the Head network module for processing to obtain a series of detection frames and category data, i.e., the final result.

[0095] The processing process in the Head network module is as follows:

[0096] The first output result and the second output result are respectively input into an ITHead detection head for detection after being processed by a self-feedback attention module SFA to obtain a series of detection frames and category data, i.e., the final result.

[0097] As shown in Figure 4 The processing process in the self-feedback attention module SFA is as follows:

[0098] The input result is processed by 1x1 convolution to obtain an Output1 result, and the Output1 result is processed by a Conv block module, and the obtained result is multiplied by the Output1 result to obtain an Output2 result.

[0099] The Output1 result is subtracted from the Output2 result to obtain an Output3 result, and the Output3 result is processed by the Conv block module, and the obtained result is multiplied by the Output3 result to obtain an Output4 result, and the Output4 result is the output result.

[0100] The processing process in the Conv block module is as follows:

[0101] The input result is sequentially processed by 3x3 convolution, 1x1 convolution, 3x3 convolution, and Sigmoid function to obtain the output result.

[0102] As shown in Figure 5 The processing process in the ITHead detection head is as follows:

[0103] The input result is divided into two branches, the first branch sequentially passes through 3x1x1 convolution, 3x1x1 convolution, 1x1 convolution processing, and positioning loss is obtained; the positioning loss sequentially passes through left-up direction convolution, left-down direction convolution, right-up direction convolution, and right-down direction convolution processing, and the four direction convolution results obtained after being spliced by the Concat function sequentially pass through the CBAM module and the Sigmoid function processing, and the first branch result is obtained; the second branch sequentially passes through 3x1x1 convolution, 3x1x1 convolution, and 1x1 convolution processing, and the second branch result is obtained, and the classification loss is obtained after the first branch result and the second branch result are multiplied; the positioning loss and the classification loss are the detection result.

[0104] As shown in Figure 6 the left-up direction convolution is based on a 3x3 convolution module, and the four weights at the left-up are used to suppress the weights on each side respectively; the left-down direction convolution is based on a 3x3 convolution module, and the four weights at the left-down are used to suppress the weights on each side respectively; the right-up direction convolution is based on a 3x3 convolution module, and the four weights at the right-up are used to suppress the weights on each side respectively; the right-down direction convolution is based on a 3x3 convolution module, and the four weights at the right-down are used to suppress the weights on each side respectively;

[0105] The calculation formula of weight suppression is as follows:

[0106]

[0107] Wherein, W is the convolution kernel weight, W diagonal is the adjacent diagonal weight of W j , j∈(1,2,3,4), and α is set to 20 here.

[0108] Example 2

[0109] For the quantitative performance evaluation of the final aerial small target image, the performance measurement standard widely used in the target detection field is adopted, and the specific evaluation is shown in formula (9) and (10).

[0110]

[0111] Wherein, AP represents the integral of R(Recall) on P(precision), the confidence threshold is from 0 to 1, mAP represents the average AP value of all classes in the data set, and N is the number of classes. The higher the value of mAP, the stronger the detection performance of the model.

[0112] Table 1 summarizes the ablation experiment data of embodiment 1 on the unmanned aerial vehicle aerial photography data set (VisDrone). From the effect of the experiment, in the comparison with the baseline, when our designed RLV, GLRF, SFA, ITHead are added alone, the model performance is obviously higher than the baseline performance; when the four modules are added at the same time, the model performance is better than the performance when any module is used alone.

[0113] Table 1 shows the ablation results on the VisDrone dataset. The best performance is indicated in bold.

[0114]

[0115] Embodiment 3

[0116] As shown in the experimental comparison chart of FIG. 7, the method of embodiment 1 of the present application (RSVDet) is compared with the prior art methods YOLOv9c (FIG. 7(a)) and EMAattention (FIG. 7(b)) for detection. The results clearly show that, compared with the EMAattention and YOLOv9c algorithms, RSVDet can identify more small target objects; especially in a complex background, the EMAattention algorithm is prone to miss detection of small targets, while RSVDet, with the unique advantage of ventral pathway to identify objects, accurately captures the subtle features of small targets, effectively reduces the loss of feature information, significantly improves the recall rate of small target detection, and fully proves the effectiveness and superiority of the method of the present application in the task of small target detection.

Claims

1. A small object detection method inspired by the visual ventral pathway neural mechanism, characterized in that, It comprises the following steps: A. Constructing a neural network, wherein the neural network comprises a pre-network RLV, a Backbone network module, a GLRF module, a Neck network module, and a Head network module; The pre-network RLV simulates the simple cell receptive field of the V1 area, the Backbone network is internally provided with four down-samplings, the GLRF module simulates the global-local receptive field, and the Neck network module is provided with a path aggregation network PAN; B. The original image is first subjected to feature enhancement by the pre-network RLV and then input into the Backbone network, and sequentially subjected to four down-sampling processes in the Backbone network to obtain first-four down-sampling results; The second down-sampling result is subjected to the GLRF module to extract the global and local features of small targets, and the global and local features of small targets are input into the Neck network module; The second-four down-sampling results are also input into the Neck network module; C. The Neck network module performs multi-scale feature fusion on the input results to obtain two features with different resolutions, which are input into the Head network module for processing to obtain a series of detection frames and category data, i.e., the final result.

2. The small target detection method inspired by the ventral visual pathway neural mechanism according to claim 1, wherein: The processing process in the pre-network RLV is as follows: The input result is divided into two paths, the first path is converted into a grayscale image and then subjected to convolution processing, and the obtained convolution result is subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing, respectively, the two results obtained are spliced by a Concat function, and then subjected to V1 simulation module processing to obtain the first path processing result; In the second path, the input result is first divided into R, G, and B single-path color images, the R, G, and B single-path color images are subjected to convolution processing, respectively, to obtain R channel convolution result, G channel convolution result, and B channel convolution result, the three convolution results are added two by two, the three two-by-two addition results are spliced by a Concat function, and then subjected to OFF-center Sigmoid module and ON-center Sigmoid module processing, respectively, the two results obtained are spliced by a Concat function, and then subjected to V1 simulation module processing to obtain the second path processing result; The first path processing result and the second path processing result are spliced by a Concat function to obtain the output result.

3. The small target detection method inspired by the ventral visual pathway neural mechanism according to claim 2, wherein: The processing process in the OFF-center Sigmoid module is as follows: The input result is subjected to 1x1 convolution and 3x3x1 convolution processing, respectively, the 3x3x1 convolution result is subtracted from the 1x1 convolution result to obtain the output result; The processing process in the ON-center Sigmoid module is as follows: The input result is processed by 1*1 convolution and 3*3*1 convolution respectively, and the output result is obtained by subtracting the result of 3*3*1 convolution from the result of 1*1 convolution. 4.The small target detection method inspired by visual ventral pathway neural mechanism according to claim 2, wherein: The V1 simulation module is based on a 3*3 convolution module, and the middle three weights are used to suppress the weights on both sides respectively, and the weight suppression calculation formula is as follows: where W is a convolution kernel weight, W diagonal is an adjacent diagonal weight of W j , j ∈ (1, 2, 3, 4), and a is set to 20 here. 5.The small target detection method inspired by visual ventral pathway neural mechanism according to claim 1, wherein: The processing process in the GLRF module is as follows: The input result is divided into three paths; the first path becomes a gray image, and the gray image is processed by two branches; the first branch is processed by 3*1*1*1 hollow convolution, 3*1*2*2 hollow convolution and 3*1*5*5 hollow convolution in turn, and the first branch result is obtained; The second branch is processed by two-dimensional average pooling, 1*1 convolution and up-sampling in turn, and the second branch result is obtained; The gray image, the first branch result and the second branch result are processed by the Concat function, and the first path result is obtained; In the second path, the input result is processed by 3*1*1 convolution, 3*2*1 convolution, 3*1*1 convolution, 1*1 convolution and up-sampling in turn, and the up-sampling result is obtained, and the input result is subtracted from the up-sampling result, and then processed by the CBAM module, up-sampling, 3*1*1 convolution and down-sampling in turn, and the second path result is obtained; In the third path, the input result is processed by the RB module, and the third path result is obtained; The first path result and the second path result are spliced by the Concat function, and then processed by 1*1 convolution, and the obtained result is spliced with the third path result by the Concat function, and the output result is obtained. 6.The small target detection method inspired by visual ventral pathway neural mechanism according to claim 5, wherein: The processing process in the RB module is as follows: The input result is divided into four paths, and each path is processed by 3*3 convolution, and the four obtained results are spliced by the Concat function, and then processed by 1*1 convolution, and the obtained result is added to the input result, and the output result is obtained. 7.The small target detection method inspired by visual ventral pathway neural mechanism according to claim 1, wherein: The processing process in the Neck network module is as follows: The fourth down-sampling result is up-sampled, and then spliced with the third down-sampling result by the Goncat function, and the obtained result is up-sampled to obtain an intermediate result The intermediate result, the second down-sampling result and the global and local features of the small target are spliced by the Goncat function to obtain a first output result; The first output result is down-sampled, and then spliced with the intermediate result by the Goncat function to obtain a second output result; The first output result and the second output result are input into the Head network module respectively. 8.The small target detection method inspired by visual ventral pathway neural mechanism according to claim 1, wherein: The processing process in the Head network module is as follows: The first output result and the second output result are respectively input into the ITHead detection head after being processed by the self-feedback attention module SFA, and a series of detection frames and category data are obtained, which are the final results. 9.The small target detection method inspired by visual ventral pathway neural mechanism of claim 8, wherein: The processing process in the self-feedback attention module SFA is as follows: The input result is processed by 1*1 convolution to obtain an Output1 result, and the Output1 result is processed by a Conv block module, and the obtained result is multiplied by the Output1 result to obtain an Output2 result; An Output3 result is obtained by subtracting the Output2 result from the Output1 result, and the Output3 result is processed by a Conv block module, and the obtained result is multiplied by the Output3 result to obtain an Output4 result, and the Output4 result is the output result; The processing process in the Conv block module is as follows: The input result is sequentially processed by 3*3 convolution, 1*1 convolution, 3*3 convolution and Sigmoid function to obtain the output result. 10.The small target detection method inspired by visual ventral pathway neural mechanism of claim 8, wherein: The processing process in the ITHead detection head is as follows: The input result is divided into two branches, the first branch is sequentially processed by 3*1*1 convolution, 3*1*1 convolution and 1*1 convolution to obtain a positioning loss, the positioning loss is sequentially processed by left-up direction convolution, left-down direction convolution, right-up direction convolution and right-down direction convolution, the four direction convolution results are spliced by a Concat function, and then sequentially processed by a CBAM module and a Sigmoid function to obtain a first branch result; the second branch is sequentially processed by 3*1*1 convolution, 3*1*1 convolution and 1*1 convolution to obtain a second branch result, and the first branch result is multiplied by the second branch result to obtain a classification loss; the positioning loss and the classification loss are the detection results. The left-up direction convolution is based on a 3*3 convolution module, and the left-up four weights are used to suppress the weights on both sides respectively; the left-down direction convolution is based on a 3*3 convolution module, and the left-down four weights are used to suppress the weights on both sides respectively; the right-up direction convolution is based on a 3*3 convolution module, and the right-up four weights are used to suppress the weights on both sides respectively; the right-down direction convolution is based on a 3*3 convolution module, and the right-down four weights are used to suppress the weights on both sides respectively; The calculation formula of the weight suppression is as follows: where W is a convolution kernel weight, W diagonal is an adjacent diagonal weight of W j , j ∈ (1, 2, 3, 4), and a is set to 20 here.