Small target detection method and device based on full convolutional network, and medium

Through the small object detection method of a full convolutional network, feature extraction and multi-scale enhancement are used for residual neural network and recursive modules, combined with focus loss function training, the feature extraction and sample imbalance problems in small object detection are solved, and the detection accuracy and effect are improved.

CN120495810APending Publication Date: 2025-08-15JIANGSU MINGXIAO INTELLIGENT TRANSPORTATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328131.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect small targets in complex contexts, and it is difficult to select and extract features, and the positive and negative samples are unbalanced, resulting in low detection accuracy.

Method used

The small object detection method of a full convolutional network is adopted, including residual neural network, recursive module, reconstruction network and detection head. Through feature extraction, multi-scale enhancement and reconstruction, it is trained in combination with the focus loss function to improve the feature extraction accuracy.

Benefits of technology

Without increasing training parameters, expand network depth, improve small object detection accuracy, enhance target and background contrast, solve the problem of positive and negative sample imbalance, and improve detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495810A_ABST
    Figure CN120495810A_ABST
Patent Text Reader

Abstract

The invention relates to a small target detection method and device based on a full convolutional network, and a medium, and the method comprises the following steps: obtaining a to-be-detected image, inputting the to-be-detected image into a small target detection model based on the full convolutional network, and outputting a small target detection result, wherein the small target detection model based on the full convolutional network comprises a residual neural network, a recursion module, a reconstruction network and a detection head which are connected in sequence; the execution process of the small target detection model based on the full convolutional network comprises the following steps: based on the to-be-detected image, performing feature extraction by using a residual neural network to obtain image features; the image features are input into a recursion module for different-scale feature enhancement, and multi-scale enhancement features are obtained; performing multi-scale reconstruction on the enhanced features by adopting a reconstruction network to obtain multi-scale reconstruction features; and detecting the multi-scale reconstruction features by using a detection head, and outputting a small target detection result. Compared with the prior art, the method has the advantages of sufficient extraction precision, improvement of small target detection precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a small target detection method, device and medium based on a fully convolutional network. Background Art

[0002] As an important branch of computer vision, object detection plays a crucial role in numerous fields, including intelligent surveillance, aerospace, and industrial inspection. In recent years, it has garnered significant attention and has been the subject of extensive research. Small object detection is a key issue in this field. Because small objects occupy a small number of pixels in a video or image, their associated features are less distinct. Furthermore, current object detection algorithms are mostly based on convolutional neural networks (CNNs). After multiple convolution and pooling operations, a large amount of small object feature information is discarded, resulting in a high incidence of false positives and missed detections. Therefore, small object detection has always been a difficult problem in object detection. In real-world scenarios, most objects appear as small objects, making the study of small object detection a crucial and challenging task.

[0003] With the development and advancement of technology, the devices we interact with in our daily lives are becoming increasingly intelligent. Image recognition, as a key element of smart devices, has also seen significant development and progress. Automatically adapting solutions to specific objects is an essential capability for smart devices, and a highly accurate image recognition algorithm greatly enhances the ability of smart devices to locate objects. Image recognition is a technology that uses a computer to analyze and calculate an image to identify the desired target or object. The algorithm can be divided into three parts: target detection, target classification, and outputting the target's identity through database matching.

[0004] As the primary component of image recognition, target detection plays a crucial role. Successful target detection is essential for identification, and the accuracy of target detection significantly impacts the accuracy of final target recognition. Therefore, target detection has long been a hot topic in the field of computer vision. Infrared small target detection and tracking is a key technology for infrared guidance, and is crucial in aerospace applications such as small celestial body detection, missile guidance, and battlefield reconnaissance. Due to the long detection range of infrared small targets, the target's image size is very small, occupying only a few dozen or even a few pixels on the imaging plane. This greatly increases the difficulty of small target detection, primarily due to the following factors: weak target signals lacking characteristic information such as target texture, shape, and size; high target motion and maneuverability, making it difficult to obtain information such as speed and direction; and uneven background grayscale distribution, as well as interference from random noise and high-brightness backgrounds.

[0005] Current technologies mainly focus on detecting larger objects in images, but the detection of small objects is easily overlooked. There are currently several designs for small object detection:

[0006] 1) Median filtering method. The maximum median filter proposed by Deshpand et al. effectively suppresses both the fluctuating background information and the edge texture information of the scene by performing a differential operation between the infrared image and the filtered image. However, this method is only effective for small targets with a high signal-to-noise ratio. Top-Hat is a practical nonlinear background estimation method. Its detection effect on small targets depends on the size and shape of the structural element. However, under long-distance imaging conditions, it is impossible to obtain prior information about small targets and select a unified structural element.

[0007] 2) Frequency domain-based small target detection method. The frequency domain-based small target detection method converts the image from the spatial domain to the frequency domain through Fourier transform, then uses a high-pass filter to filter it, and finally performs an inverse Fourier transform to obtain the predicted image. Yang et al. regard point targets as high-frequency components of the image and propose an adaptive Butterworth high-pass filter (BHPF). Hilliard et al. use a low-pass IIR (infinite impulse response, IIR) filter to predict clutter, which is suitable for multi-target situations. Reed et al. use the frequency domain optimal three-dimensional linear matching filter technology to detect moving point targets in image sequences when the speed, background clutter and noise are known.

[0008] 3) A method based on multi-scale local contrast measurement. This method considers the distribution differences between target edges and cloud edges and uses the minimum product in the diagonal direction as the final enhancement result. This method can enhance both bright and dark targets and achieves good results in cloud edge removal. However, its detection performance degrades in complex backgrounds and strong clutter interference.

[0009] The problem of a large number of false alarms in the detection results of small targets in complex environments indicates that the features manually extracted by these traditional algorithms are not sufficient, while deep learning algorithms have powerful feature extraction and information abstraction capabilities. The difficulty of applying deep learning to small target detection in infrared images lies in the fact that the targets are small and lack contour feature information, which brings great difficulties to the design of target detection networks. In addition, the diversity and complexity of the background in which the targets are located, as well as the fluctuations in the grayscale and size of the targets themselves, also increase the difficulty of detection. By summarizing and analyzing the current research status of small target detection, it is found that the research field of small target detection mainly has the following problems: (1) Feature selection and extraction problems. At present, whether it is saliency detection of a single image or small target detection of multiple images, they are mostly based on underlying features such as color, space and texture, and are often only effective for a certain type of specific image. The extraction and selection of these underlying features are all done manually, which is a task that requires certain professional knowledge and is extremely labor-intensive. The quality of the selection results mainly depends on subjective experience. Although deep feature extraction methods can be automatically obtained, they are relatively slow, which affects the efficiency of saliency detection. (2) There is a serious imbalance between positive and negative samples. Summary of the Invention

[0010] The purpose of the present invention is to provide a small target detection method, device and medium based on a fully convolutional network to improve the detection accuracy of small targets.

[0011] The purpose of the present invention can be achieved by the following technical solutions:

[0012] A small target detection method based on a fully convolutional network includes the following steps:

[0013] Acquire a test image, input it into a small target detection model based on a fully convolutional network, and output a small target detection result, wherein the small target detection model based on the fully convolutional network includes a residual neural network, a recursive module, a reconstruction network, and a detection head connected in sequence;

[0014] The execution process of the small target detection model based on the fully convolutional network includes:

[0015] Based on the image to be tested, a residual neural network is used to perform feature extraction to obtain image features;

[0016] Inputting the image features into a recursive module to perform feature enhancement at different scales to obtain multi-scale enhanced features;

[0017] Using a reconstruction network to perform multi-scale reconstruction on the enhanced features to obtain multi-scale reconstruction features;

[0018] A detection head is used to detect the multi-scale reconstruction features and output a small target detection result.

[0019] Furthermore, the residual neural network includes an interest point detection module and a plurality of stacked residual blocks, the interest point detection module is connected after the last residual block, each residual block includes a plurality of cascaded basic residual units, and the step of obtaining image features includes:

[0020] Setting a sliding step size, wherein the sliding step size is smaller than a sampling window size, and the sampling window size is a convolution kernel size of a residual neural network;

[0021] Based on the image to be tested and the sliding step size, a basic residual unit is used to extract primary features, and a residual block is used to combine the primary features to extract multi-level features, and finally high-level features are obtained through multiple stacked residual blocks;

[0022] The high-level features are input into an interest point detection module to detect all interest points, and all interest points constitute image features.

[0023] Furthermore, the basic residual unit includes two convolutional layers and a batch normalization layer and an activation function layer connected after each convolutional layer. The operation process of the basic residual unit is expressed as:

[0024]

[0025] in:

[0026] F(x,W)=W2σ(W1x)

[0027] Where, is the output of the basic residual unit, h(x) is an identity mapping, h(x) = x, x is the input of the basic residual unit, W is a set of weight parameters, σ represents the ReLU activation function, F(x,W) is the residual mapping to be learned, W1 and W2 are the weight parameters of the two convolutional layers respectively.

[0028] Furthermore, the execution steps of the interest point detection module include:

[0029] Based on the color enhancement detection algorithm, the color saliency of the high-level features is enhanced to generate a saliency matrix M ′ , wherein the color enhancement detection algorithm performs color saliency enhancement through a conversion function, and the saliency matrix M ′ Expressed as:

[0030]

[0031] in:

[0032]

[0033] Where p is the probability function, I xis the input image, I x ′ is the color enhanced image, g is the conversion function, ω(x,y) is the Gaussian weight function, g(I x ) and g(I y ) is the gradient component after color enhancement;

[0034] According to the significance matrix M ′ , filter out all points of interest;

[0035] Select an interest point as the starting position, use the starting position as a reference, select the interest point with the smallest angle value according to the preset direction and connect them, then repeat this step with the selected interest point as the next starting point until it is connected to the starting position to form a closed area as the image feature.

[0036] Furthermore, the recursive module includes multiple recursive layers, some of which are connected to some of the convolutional layers in the residual neural network, and are used to iteratively perform the convolution operation of the convolutional layer using the recursive layer to enhance features of different scales and obtain multi-scale enhanced features, wherein the operation expression of the d-th recursive layer is:

[0037] H d =g(G d-1 )=max(0,W*H d-1 + b) = g d-1 (H1) = g d (H)

[0038] in:

[0039] g(H)=max(0,W*H+b)

[0040] H1=max(0,WX5+b)+X1=g 1 (X5)

[0041] H2=max(0,WH1+b)=g 2 (X5)

[0042] H3=max(0,WH2+b)+X3=g 3 (X5)

[0043] H4=max(0,WH3+b)=g d (X5)

[0044] f2(X5)=H4

[0045] Where H d is the output of the d-th recurrent layer, g dis the d iterations of the convolution function, W is the weight, b is the bias parameter, H is the input of the recursive layer, H1, H2, H3, and H4 are the inputs of the 1st to 4th recursive layers respectively, X5 represents the features of the fifth convolutional layer during feature extraction, and f2 is the parameter of the second convolution branch in the reconstruction network.

[0046] Furthermore, the reconstruction network includes multiple branch networks with parallel convolution structures, each branch network is connected to the output end of different recursive layers in the recursive module, and each branch network performs convolution processing on the enhanced features of different scales output by different recursive layers to obtain reconstruction features of different scales, thereby forming multi-scale reconstruction features. The operation process of each branch network is expressed as follows:

[0047]

[0048] Where, is the output of the d-th branch network, f3 is the parameter of the third branch network in the reconstruction network, g d is the d iterations of the convolution function in the recursive layer, X5 is the parameter of the second branch network in the reconstruction network, f1 is the parameter of the first branch network in the reconstruction network, and x is the input of the branch network.

[0049] Furthermore, the expression for the detection head to detect the multi-scale reconstruction feature is:

[0050]

[0051] Where, is the prediction result, w d is the weight of the reconstructed network.

[0052] Furthermore, the small target detection model based on the fully convolutional network is trained using a focus loss function, and the expression of the focus loss function is:

[0053]

[0054] Where FL is the focal loss function, y is the true label of the sample, is the prediction result of the sample, α is the balance factor, which is used to balance the number ratio of positive and negative samples. is the modulation coefficient used to control the weights of different samples.

[0055] The present invention also provides an electronic device comprising: one or more processors; a memory; and one or more programs stored in the memory, wherein the one or more programs include instructions for executing the small target detection method based on a fully convolutional network as described above.

[0056] The present invention also provides a computer-readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the small target detection method based on a fully convolutional network as described above.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] (1) In the small target detection model of the present invention, based on the traditional residual neural network for feature extraction, the recursive module and the reconstruction network are connected, and the network depth of the model is expanded without adding additional training parameters. By using the recursive layer to iteratively perform the convolution operation of the convolution layer and using the reconstruction network to perform multi-scale feature reconstruction, the features can be fully extracted, more accurate features can be obtained, and the detection accuracy of small targets can be improved.

[0059] (2) The reconstruction network of the present invention includes multiple branch networks with parallel convolution structures, each of which is connected to the output end of a different recursive layer in the recursive module. The parallel convolution structure simulates the multi-scale operation of the traditional algorithm, which is beneficial to enhancing the contrast between the target and the background in a complex environment, thereby improving the feature extraction accuracy of small targets.

[0060] (3) The present invention adopts the focus loss function for model training. The focus loss function is an improvement of the cross entropy loss function based on the pixel point. It can not only check each pixel individually, but also add a modulation coefficient to the cross entropy loss function. It can control the weights of easy-to-classify samples and difficult-to-classify samples, allowing the model to focus more on difficult and misclassified samples, and introduce a balance factor α to solve the problem of imbalance in the number of positive and negative samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Schematic diagram of the method flow of the present invention;

[0062] Figure 2 This is a structural diagram of the small target detection model based on the fully convolutional network of the present invention. DETAILED DESCRIPTION

[0063] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0064] This embodiment provides a small target detection method based on a fully convolutional network. The method is implemented based on a small target detection model of a fully convolutional network. Figure 2 As shown in Figure 2, the small target detection model includes a residual neural network, a recursive module, a reconstruction network, and a detection head connected in sequence. Figure 1 and Figure 2 , the method comprises the following steps:

[0065] The first step is to set the sliding step size.

[0066] The sliding step size is set to be slightly smaller than the sampling window size (the convolution kernel size of the residual neural network) to avoid the loss of target edge information caused by window boundary cutting. By selecting the sampling window size, the number of times the network searches the entire image can be controlled. The calculation formula is:

[0067]

[0068] W and H represent the width and height of the infrared image respectively; S represents the sampling window size; stride is the search step size.

[0069] The second step is to use the residual neural network to perform feature extraction and obtain image features.

[0070] The residual neural network (as a feature extraction network) includes an interest point detection module and multiple stacked residual blocks. The interest point detection module is connected after the last residual block. Each residual block includes multiple cascaded basic residual units. Basic residual units are used to extract primary features. These residual blocks are stacked together to form a very deep network. The residual blocks combine the primary features to extract multi-level features, and finally obtain high-level features through multiple stacked residual blocks. Let the input be x and the basic residual unit be represented as:

[0071]

[0072] Where: is the output of the residual unit; h(x) is an identity mapping: h(x) = x; W is a set of weight parameters; σ represents the ReLU activation function; F(x,W) is the residual mapping to be learned. For a basic residual unit with two convolutional layers stacked on top of it, we have:

[0073] F(x,W)=W2σ(W1x)

[0074] Where: W1 and W2 are the weight parameters of the two convolutional layers respectively.

[0075] By stacking these structures, a 152-layer network is constructed. Then, in the interest point detection module, a detection algorithm based on color enhancement is proposed to calculate the interest points through a conversion function, which requires a conversion function g.

[0076]

[0077] The transformation achieved by function g is called color saliency enhancement. Once function g is found, the saliency of color enhancement can be calculated. Substituting function g into the new saliency matrix M can be obtained. ′ Detect points of interest.

[0078]

[0079] By traversing and comparing, we find and connect the points of interest that meet the requirements. Using the color enhancement detection algorithm to obtain all points of interest, we first find a starting point among all the points of interest. This starting point should be located at the bottom right of all the points of interest. Secondly, we search for the point of interest with the smallest angle in a certain direction. We connect the starting point to this point of interest. Finally, we repeat the above process until we connect it to the starting position, forming a closed area. This image feature is then used as the target candidate area feature.

[0080] The third step is to use the recursive module to enhance features of different scales and obtain multi-scale enhanced features.

[0081] This example adds a recursive block to the model. This block contains 16 recursive layers, each of which uses the same convolution parameters. This prevents the model from adding additional parameters when performing convolution operations within the recursive layers. For a recursive block with D recursive layers, we use the same weight W and bias parameter b for all convolution operations. Define g as the convolution function of a single recursive layer in the recursive block, and H as the input to the recursive layer.

[0082] g(H)=max(0,W*H+b)

[0083] The output of the recurrent module at the d-th recurrent layer is:

[0084] H d =g(H d-1 )=max(0,W*H d-1 + b) = g d-1 (H1) = g d (H)

[0085] Where: g d represents d iterations of the function.

[0086] Secondly, using skip connections in the residual neural network, we connect the first and third convolutional layers in the residual neural network to the first and third recurrent layers in the recurrent module, respectively. This provides information supervision for these two recurrent layers to mitigate the effects of vanishing or exploding gradients. Since all recurrent convolutional layers in the recurrent module share the same convolution parameters, only one set of convolution kernel parameters, W,b, needs to be trained for the recurrent module, which accelerates model convergence.

[0087] H1=max(0,WX5+b)+X1=g 1 (X5)

[0088] H2=max(0,WH1+b)=g 2 (X5)

[0089] H3=max(0,WH2+b)+X3=g 3 (X5)

[0090] H4=max(0,WH3+b)=g d (X5)

[0091] f2(X5)=H4

[0092] The fourth step is to use the reconstruction network to perform multi-scale reconstruction of the enhanced features.

[0093] The inter-layer outputs of all recursive layers in the recursive module are used as inputs to different branches of the parallel convolutional structure of the reconstruction network. In the reconstruction network, the parallel convolutional structure has four branches, all of which share the same set of network branch parameters. Each branch undergoes four layers of convolution operations. Because the information input of different branches comes from the output of recursive layers at different network depths, the receptive field of the image varies. This means that the different branches of the parallel convolutional structure are equivalent to multi-scale representation.

[0094]

[0095] Where: d = 1, 2, 3, 4.

[0096] Afterwards, shared filters are used to fuse the outputs of the four branch networks to generate a unified and compact feature representation, namely the multi-scale reconstruction feature.

[0097] Step 5: Use the detection head to predict small targets.

[0098] Using the detection head, the multi-scale reconstruction features are weighted w in the depth direction d Weighted average, this parameter is obtained during network training. The prediction result of the model is:

[0099]

[0100] Finally, the pixel-based cross entropy loss function is introduced as the most commonly used loss function in semantic segmentation problems. This loss function examines each pixel individually, and its mathematical expression is:

[0101]

[0102] Where: y is the label of the real sample, It is the prediction result of the reconstructed network (the value is between 0 and 1).

[0103] In order to reduce the contribution of easy-to-classify samples and make the network pay more attention to difficult-to-classify samples, this embodiment introduces the Focal function (FL), which adds a modulation coefficient to the cross entropy loss function. It can control the weights of easy-to-classify samples and difficult-to-classify samples, allowing the model to focus more on difficult and misclassified samples, and introduce a balance factor α to balance the uneven ratio of the number of positive and negative samples.

[0104]

[0105] Where: α = 0.25, γ = 2.

[0106] In summary, the embodiments of the present invention adopt sliding window sampling training in the field of small target detection. It is derived from the idea of nested structure in traditional target detection algorithms based on human visual characteristics. A fully convolutional network using recursive convolutional layers is designed to expand the network depth of the model without adding additional training parameters. The multiple branch networks of the parallel convolution structure of the network simulate the multi-scale operation of the traditional algorithm, which is conducive to enhancing the contrast between the target and the background in complex environments. In addition, a combination of multiple loss functions is designed to combat the problem of severe imbalance between positive and negative samples.

[0107] Example 2

[0108] This embodiment provides an electronic device, comprising: one or more processors; a memory; and one or more programs stored in the memory, wherein the one or more programs include instructions for executing the small target detection method based on a fully convolutional network as described in the above-mentioned embodiment 1.

[0109] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0110] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0111] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0112] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0114] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0115] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A small target detection method based on a fully convolutional network, characterized in that: The following steps are involved: Acquire a test image, input it into a small target detection model based on a fully convolutional network, and output a small target detection result, wherein the small target detection model based on the fully convolutional network includes a residual neural network, a recursive module, a reconstruction network, and a detection head connected in sequence; The execution process of the small target detection model based on the fully convolutional network includes: Based on the image to be tested, a residual neural network is used to perform feature extraction to obtain image features; Inputting the image features into a recursive module to perform feature enhancement at different scales to obtain multi-scale enhanced features; Using a reconstruction network to perform multi-scale reconstruction on the enhanced features to obtain multi-scale reconstruction features; A detection head is used to detect the multi-scale reconstruction features and output a small target detection result.

2. The small target detection method based on a fully convolutional network according to claim 1, characterized in that: The residual neural network includes an interest point detection module and multiple stacked residual blocks, the interest point detection module is connected after the last residual block, each residual block includes multiple cascaded basic residual units, and the step of obtaining image features includes: Setting a sliding step size, wherein the sliding step size is smaller than a sampling window size, and the sampling window size is a convolution kernel size of a residual neural network; Based on the image to be tested and the sliding step size, a basic residual unit is used to extract primary features, and a residual block is used to combine the primary features to extract multi-level features, and finally high-level features are obtained through multiple stacked residual blocks; The high-level features are input into an interest point detection module to detect all interest points, and all interest points constitute image features.

3. The small target detection method based on a fully convolutional network according to claim 2, characterized in that: The basic residual unit includes two convolutional layers and a batch normalization layer and an activation function layer connected after each convolutional layer. The operation process of the basic residual unit is expressed as follows: in: F(x,W)=W2σ(W1x) Where, is the output of the basic residual unit, h(x) is an identity mapping, h(x) = x, x is the input of the basic residual unit, W is a set of weight parameters, σ represents the ReLU activation function, F(x,W) is the residual mapping to be learned, W1 and W2 are the weight parameters of the two convolutional layers respectively.

4. The small target detection method based on a fully convolutional network according to claim 2, characterized in that: The execution steps of the interest point detection module include: Based on the color enhancement detection algorithm, the color saliency of the high-level features is enhanced to generate a saliency matrix M ′ , wherein the color enhancement detection algorithm performs color saliency enhancement through a conversion function, and the saliency matrix M ′ Expressed as: in: Where p is the probability function, I x is the input image, I x ′ is the color enhanced image, g is the conversion function, ω(x,y) is the Gaussian weight function, g(I x ) and g(I y ) is the gradient component after color enhancement; According to the significance matrix M ′ , filter out all points of interest; Select an interest point as the starting position, use the starting position as a reference, select the interest point with the smallest angle value according to the preset direction and connect them, then repeat this step with the selected interest point as the next starting point until it is connected to the starting position to form a closed area as the image feature.

5. The small target detection method based on a fully convolutional network according to claim 1, characterized in that: The recursive module includes multiple recursive layers, some of which are connected to some of the convolutional layers in the residual neural network, and are used to iteratively perform convolution operations of the convolutional layers using the recursive layers to enhance features of different scales and obtain multi-scale enhanced features, wherein the operation expression of the d-th recursive layer is: H d =g(H d-1 )=max(0,W*H d-1 +b)=g d-1 (H1)=g d (H) in: g(H)=max(0,W*H+b) H1=max(0,WX5+b)+X1=g 1 (X5) <h2 style=";text-align:left;direction:ltr">H2 = max(0,WH1 + b) = g<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> (X5) <h2 style=";text-align:left;direction:ltr">H3 = max(0,WH2 + b) + X3 = g<h2 style=";text-align:left;direction:ltr"> 3 <h2 style=";text-align:left;direction:ltr"> (X5) H4=max(0,WH3+b)=g 4 (X5) f2(X5)=H4 Where H d is the output of the d-th recurrent layer, g d is the d iterations of the convolution function, W is the weight, b is the bias parameter, H is the input of the recursive layer, H1, H2, H3, and H4 are the inputs of the 1st to 4th recursive layers respectively, WX5 represents the features of the fifth convolutional layer during feature extraction, and f2 is the parameter of the second branch network in the reconstruction network.

6. The small target detection method based on a fully convolutional network according to claim 1, characterized in that: The reconstruction network includes multiple branch networks with parallel convolution structures. Each branch network is connected to the output end of different recursive layers in the recursive module. Each branch network performs convolution processing on the enhanced features of different scales output by different recursive layers to obtain reconstruction features of different scales, thereby forming multi-scale reconstruction features. The operation process of each branch network is expressed as follows: Where, is the output of the d-th branch network, f3 is the parameter of the third branch network in the reconstruction network, g d is the d iterations of the convolution function in the recursive layer, X5 is the parameter of the second branch network in the reconstruction network, f1 is the parameter of the first branch network in the reconstruction network, and x is the input of the branch network.

7. The small target detection method based on a fully convolutional network according to claim 6, characterized in that: The expression for the detection head to detect the multi-scale reconstruction features is: Where, is the prediction result, w d is the weight of the reconstructed network.

8. The small target detection method based on a fully convolutional network according to claim 1, characterized in that: The small target detection model based on the fully convolutional network is trained using a focus loss function, which is expressed as: Where FL is the focal loss function, y is the true label of the sample, is the prediction result of the sample, α is the balance factor, which is used to balance the number ratio of positive and negative samples. is the modulation coefficient used to control the weights of different samples.

9. An electronic device, characterized in that: include: one or more processors; Memory; and One or more programs stored in a memory, wherein the one or more programs include instructions for executing the small target detection method based on a fully convolutional network as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the small target detection method based on a fully convolutional network as described in any one of claims 1-8.