Small target detection method guided by fluid mechanics
By introducing fluid mechanics-guided detail reconstruction blocks into the small object detection model, the problem of low accuracy of small object detection in the prior art is solved, and clearer edge structure details extraction and higher detection accuracy are achieved.
Patent Information
- Application Number
- CN202310941571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-07-28
AI Technical Summary
When existing small object detection methods detect small objects in complex backgrounds, the accuracy rate is low, making it difficult to extract the shape, size and direction characteristics of small objects, resulting in unclear details of edge structure.
Using a small object detection model based on fluid mechanics guidance, a model including feature extraction blocks, feature reconstruction blocks and prediction blocks is constructed. The feature reconstruction block uses a fluid mechanic-guided detail reconstruction block to reconstruct the high-level features to accurately reconstruct the shape, size and direction features of the small object.
Through the fluid mechanics-guided detail reconstruction block, the model can more clearly extract the edge structure details of small targets, significantly improving the accuracy of small target detection.
Smart Images

Figure CN116994091B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically relates to a small target detection method guided by fluid mechanics, which can be used for aerial image analysis and autonomous driving systems. Background Art
[0002] Small target detection aims to detect small targets from complex backgrounds, and there is a wide range of demands in many fields such as medical diagnosis, autonomous driving, and drone navigation. With the continuous development of deep learning technology, the accuracy of small target detection has been improved to a certain extent. Given the characteristics of small targets with fewer pixels and blurred shape features in images, it is difficult to extract fine target features from complex backgrounds, resulting in misdetection and missed detection of targets. The performance of small target detection algorithms based on deep learning still needs to be further improved and optimized.
[0003] For example, in the patent document "A Small Target Detection Method Based on High-Low Level Feature Fusion" (Patent Application No.: CN202211557960.9, Publication No.: CN116229135A) applied by Northwestern Polytechnical University, a target detection neural network based on a preset multi-scale fusion strategy is used to perform small target detection on a target image. Among them, a high-low level feature fusion module in the target detection neural network is used to realize the fusion of low-level features and high-level features; the high-low level feature fusion module is used to preprocess the low-level features and high-level features respectively to make the dimensions of the low-level features and high-level features the same; the preprocessed low-level features are first subjected to channel compression and then channel expansion to retain useful information of small targets and filter out useless information to obtain channel compressed and expanded features; the channel compressed and expanded features are multiplied by the preprocessed high-level features, and the product result is added to the preprocessed high-level features as the high-low level fusion features. In this invention, as the number of network layers increases during the training process, features such as the shape, size, and direction of the target are easily lost, resulting in unclear detailed features of the small target edge structure extracted by the network and low detection accuracy. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a small target detection method guided by fluid mechanics to solve the technical problem of low detection accuracy existing in the existing small target detection methods.
[0005] To achieve the above object, the technical solution adopted by the present invention includes the following steps:
[0006] (1) Obtain a training sample set and a test sample set:
[0007] Get K small target images, and label the small targets in each small target image, then combine M small target images and their corresponding labels into a training sample set R1, and combine the remaining KM small target images and their corresponding labels into a test sample set E1, where K ≥ 500.
[0008] (2) Construct a small target detection model O based on fluid mechanics guidance:
[0009] Construct a small target detection model O including a Stem block, N feature extraction blocks, N-1 feature reconstruction blocks and a prediction block connected in sequence, and the output end of the nth feature extraction block is also connected to the input end of Nn feature reconstruction blocks, wherein the feature reconstruction block includes a cascaded first convolution block and a fluid mechanics-guided detail reconstruction block, as well as a cascaded deconvolution layer and a second convolution block, and the output end of the fluid mechanics-guided detail reconstruction block is connected to the output end of the second convolution block; the fluid mechanics-guided detail reconstruction block includes a first branch composed of a cascaded horizontal Gaussian elliptic operator and a nonlinear activation layer, and a second branch composed of a cascaded vertical Gaussian elliptic operator and a nonlinear activation layer, the output ends of the two branches are respectively connected to the output end of the first convolution block by multiplication, and then added; wherein N≥2;
[0010] (3) Initialization parameters:
[0011] The number of initialization iterations is t, the maximum number of iterations is T, T ≥ 1000, and the small target detection model O of the tth iteration t The weight and bias parameters in are w t , b t , and let t = 0, O t =O;
[0012] (4) Training the small target detection model O:
[0013] Randomly select L training samples with replacement from the training sample set R1 as the input of the small target detection model O for forward propagation to obtain L small target detection results, where 1≤L≤M;
[0014] (5) Update the parameters of the small target detection model:
[0015] The L small target detection results obtained in step (4) are used to detect the small target detection model O t The weight and bias parameter w t , b t Update and get the network model O of this iteration t ; Determine whether t≥T holds. If so, obtain the trained small target detection model O*. Otherwise, set t=t+1 and execute step (4);
[0016] (6) Obtain the small target detection result:
[0017] Use the test sample set E1 as the input of the trained small target detection model O* for forward propagation to obtain the small target detection results corresponding to K-M test samples.
[0018] Compared with the prior art, the present invention has the following advantages:
[0019] The small target detection model constructed by the present invention includes a feature reconstruction block. During the training of the model and the process of obtaining the small target detection result, the feature reconstruction block uses low-level features to reconstruct high-level features. The included hydrodynamic-guided detail reconstruction block combines hydrodynamic theory to accurately reconstruct the shape, size, and orientation features of small targets, thereby obtaining more detailed information on the edge structure of small targets. Experimental results show that the present invention effectively improves the accuracy of small target detection. Brief Description of the Drawings
[0020] Figure 1 It is a flowchart for the implementation of the present invention;
[0021] Figure 2 It is a schematic structural diagram of the small target detection model based on hydrodynamic guidance adopted by the present invention;
[0022] Figure 3 It is a schematic structural diagram of the feature extraction block adopted by the present invention;
[0023] Figure 4 It is a schematic structural diagram of the feature reconstruction block adopted by the present invention. Specific Embodiments
[0024] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments.
[0025] Refer to Figure 1 , the present invention includes the following steps:
[0026] (1) Obtain the training sample set and the test sample set:
[0027] Obtain 1000 small target images included in the IRSTD-1k dataset, annotate the small targets in each small target image, and then form the training sample set R1 with 600 small target images and their corresponding labels, and form the test sample set E1 with the remaining 400 small target images and their corresponding labels;
[0028] (2) Construct a small target detection model O based on hydrodynamic guidance:
[0029] Refer to Figure 2, a further description is made of the small target detection model based on hydrodynamic guidance adopted in the present invention.
[0030] Construct a model including a Stem block, three feature extraction blocks, two feature reconstruction blocks, and a prediction block connected in sequence. The output end of the first feature extraction block is also connected to the input end of the second feature reconstruction block, and the output end of the second feature extraction block is also connected to the input end of the first feature reconstruction block. Among them, the Stem block includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a pooling layer connected in sequence; the three feature extraction blocks adopt the ResNet-20 structure, and each feature extraction block includes a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, and a ninth convolutional layer connected in sequence; the feature reconstruction block includes a cascaded first convolutional block and a hydrodynamic guidance detail reconstruction block, as well as a cascaded transposed convolutional layer and a second convolutional block. The output end of the hydrodynamic guidance detail reconstruction block is connected to the output end of the second convolutional block. The convolutional block includes a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer connected in sequence. The hydrodynamic guidance detail reconstruction block includes a first branch arranged in parallel, which is composed of a cascaded horizontal Gaussian elliptical operator and a first non-linear activation layer, and a second branch composed of a cascaded vertical Gaussian elliptical operator and a second non-linear activation layer. After the output ends of the two branches are multiplied and connected to the output end of the first convolutional block respectively, they are added and connected; the Head module includes a thirteenth convolutional layer, a normalization layer, a third non-linear activation layer, a random dropout layer, a fourteenth convolutional layer, and an interpolation layer connected in sequence;
[0031] The specific parameter settings are as follows: the convolutional kernel size of the first convolutional layer is 3*3, the stride is 2, and the padding is 1; the convolutional kernel sizes of the second convolutional layer and the third convolutional layer are both 3*3, the stride is 1, and the padding is 1; the convolutional kernel sizes of the convolutional layers in the first feature extraction block are all 3*3, the stride is 1, and the padding is 1; the convolutional kernel sizes of the convolutional layers in the second feature extraction block and the third feature extraction block are both 3*3, the stride is 2, and the padding is 1; the convolutional kernel sizes of the tenth convolutional layer and the eleventh convolutional layer are both 1*1, the stride is 1, used to adjust the number of channels, and the output number of channels is the output number of channels of the Nth feature extraction block; the convolutional kernel sizes of the twelfth convolutional layer and the thirteenth convolutional layer are 3*3, the stride is 1, and the padding is 1; the convolutional kernel size of the fourteenth convolutional layer is 1*1, the stride is 1; the pooling layer adopts max pooling; the convolutional kernel sizes of the transposed convolutional layers are all 4*4, the stride is 2, and the padding is 1; the first non-linear activation layer and the second non-linear activation layer both adopt the Sigmoid function, and the third non-linear activation layer adopts the ReLU function; the normalization layer adopts layer normalization; the interpolation layer adopts the interpolate operation, and the output spatial size is the size of the input image;
[0032] Among them, the theory in fluid mechanics is used to design the fluid mechanics-guided detail reconstruction block. The transformation distribution of pixels through the network layer can be similar to the flow of particles in a hydrodynamic system. The feature image reconstruction process is regarded as a near-fluid motion process. Under the guidance of the horizontal Gaussian elliptical operator, the vertical Gaussian elliptical operator, and the loss function, the pixels move towards the distribution targeted at the high-resolution image, gradually forming clear edge structure details.
[0033] (3) Initialize parameters:
[0034] Initialize the iteration number as t, the maximum iteration number as T = 1000, and the weights and bias parameters of the small target detection model O in the t-th iteration are w t and b t respectively, and let t = 0, O t = O; t
[0035] (4) Train the small target detection model O:
[0036] Randomly select 32 training samples from the training sample set R1 with replacement as the input of the small target detection model O for forward propagation:
[0037] (4a) The input image size is 512*512, the number of channels is 3, and the Stem block downsamples each image to obtain a preprocessed feature map with a size of 128*128 and a channel number of 16. The first feature extraction block extracts features from the feature map to obtain a local feature map with a size of 128*128 and a channel number of 16. The second feature extraction block extracts features from the feature map to obtain a local feature map with a size of 64*64 and a channel number of 32. The third feature extraction block extracts features from the feature map to obtain a local feature map with a size of 32*32 and a channel number of 64;
[0038] (4b) The first feature reconstruction block uses the local feature map to perform feature image reconstruction on the local feature map to obtain a feature reconstruction feature map with a size of 64*64 and a channel number of 64. The second feature reconstruction block uses the local feature map to perform feature image reconstruction on the feature reconstruction feature map to obtain a feature reconstruction feature map with a size of 128*128 and a channel number of 64;
[0039] Among them, in the first feature reconstruction block, the first convolutional block processes the local feature map Adjust the number of channels and extract features. The fluid - mechanics - guided detail reconstruction block reconstructs the edge detail features of the output feature map of the first convolution block to obtain an edge - detail - reconstructed feature map The transposed convolution layer upsamples the local feature map The second convolution block adjusts the number of channels and extracts features from the output feature map of the transposed convolution layer to obtain an upsampled feature map The edge - detail - reconstructed feature map is element - wise added to the upsampled feature map to obtain a feature - reconstructed feature map
[0040] In the second feature reconstruction block, the first convolution block adjusts the number of channels and extracts features from the local feature map The fluid - mechanics - guided detail reconstruction block reconstructs the edge detail features of the output feature map of the first convolution block to obtain an edge - detail - reconstructed feature map The transposed convolution layer upsamples the feature - reconstructed feature map The second convolution block adjusts the number of channels and extracts features from the output feature map of the transposed convolution layer to obtain an upsampled feature map The edge - detail - reconstructed feature map is element - wise added to the upsampled feature map to obtain a feature - reconstructed feature map
[0041] Among them, the specific process of the fluid - mechanics - guided detail reconstruction block reconstructing the edge detail features of the output feature map of the first convolution block to obtain an edge - detail - reconstructed feature map can be expressed as:
[0042]
[0043]
[0044]
[0045] U h =[cos0°, - sin0°; sin0°, cos0°] (4)
[0046]
[0047]
[0048]
[0049] U v= [cos90°, -sin90°; sin90°, cos90°] (8)
[0050]
[0051] where s1(·) is the first non - linear activation layer, s2(·) is the second non - linear activation layer, G h (·) is the Gaussian elliptical operator in the horizontal direction, G v (·) is the Gaussian elliptical operator in the vertical direction, is the coordinates of adjacent pixels and the central pixel in the image Σ h is the covariance matrix in the horizontal direction, U h is the orthogonal matrix in the horizontal direction, Λ h is the diagonal matrix in the horizontal direction, Σ v is the covariance matrix in the vertical direction, U v is the orthogonal matrix in the vertical direction, Λ v is the diagonal matrix in the vertical direction, σ1, σ2, σ3, σ4 are elongation rate parameters,
[0052] (4c) The prediction block performs object detection on the feature reconstruction feature map and obtains 32 small object detection results.
[0053] (5) Update the parameters of the small object detection model:
[0054] Based on the detection results of the 32 small objects obtained in step (4), update the weights, bias parameters w t and b t of the small object detection model O t to obtain the network model O t of this iteration:
[0055] (5a) Using the Dice loss function L Dice and the cross - entropy loss function L CE , calculate the loss value L t of the small object detection model through the generated 32 small object detection results and the labels of the corresponding 32 images:
[0056]
[0057]
[0058] L t = L Dice +L CE (12)
[0059] where rl represents the label category of each pixel in the input image, p l represents the probability that each pixel in the small target detection result belongs to the label category, ε represents the correction factor, and its value is any real number within the range of (0, 0.1). Its function is to prevent the denominator of the fraction from being zero;
[0060] (5b) Calculate L through the chain rule t for the weight parameter ω t and the bias parameter b t partial derivatives and Finally, according to for ω t , b t perform the update:
[0061]
[0062]
[0063] where ω t , b t represent the weights and bias parameters of all learnable parameters of O t , w t ', b t ' represent the update results of ω t , b t , α represents the learning rate; judge whether t≥T holds. If so, obtain the trained small target detection model O*, otherwise, set t = t + 1 and execute step (4).
[0064] (6) Obtain the small target detection result:
[0065] Use the test sample set E1 as the input of the trained small target detection model O* for forward propagation to obtain the small target detection results corresponding to 400 test samples.
[0066] The following combines simulation experiments to illustrate the technical effects of the present invention
[0067] Simulation conditions, content, and result analysis:
[0068] The hardware platform for the simulation experiment is: the processor is an Intel(R) Core i9-9900K CPU with a main frequency of 3.5 GHz, the memory is 32 GB, and the graphics card is an NVIDIA GeForce RTX 2080Ti. The software platform for the simulation experiment is: the Ubuntu 16.04 operating system, the python version is 3.7, and the Pytorch version is 1.7.1.
[0069] Compare the detection performance of the patent document "A Small Target Detection Method Based on High-Level and Low-Level Feature Fusion" (Patent Application No.: CN202211557960.9, Publication No.: CN116229135A) and the present invention using the Intersection over Union (IoU) and accuracy evaluation metrics. The IoU of the small target detection result of the existing method is 61.86%, and the accuracy is 94.22%. The IoU of the small target detection result of the present invention is 65.37%, and the accuracy is 97.76%. Compared with the prior art, the detection accuracy of the present invention has been significantly improved.
Claims
1. A small target detection method guided by fluid mechanics, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Obtain K small target images, label the small targets in each small target image, and then form a training sample set R1 with M small target images and their corresponding labels. Form a test sample set E1 with the remaining K - M small target images and their corresponding labels, where K ≥ 500. (2) Construct a small target detection model O guided by fluid mechanics: Construct a small target detection model O including a Stem block, N feature extraction blocks, N - 1 feature reconstruction blocks, and a prediction block connected in sequence, and the output end of the nth feature extraction block is also connected to the input end of N - n feature reconstruction blocks. Among them, the feature reconstruction block includes a cascaded first convolution block and a fluid mechanics-guided detail reconstruction block, as well as a cascaded deconvolution layer and a second convolution block. The output end of the fluid mechanics-guided detail reconstruction block is connected to the output end of the second convolution block; the fluid mechanics-guided detail reconstruction block includes a first branch arranged in parallel composed of a cascaded horizontal direction Gaussian ellipse operator and a non-linear activation layer, and a second branch composed of a cascaded vertical direction Gaussian ellipse operator and a non-linear activation layer. The output ends of the two branches are respectively multiplied and then added to the output end of the first convolution block; where N≥2; (3) Initialize the parameters: The number of initialization iterations is t, the maximum number of iterations is T, T ≥ 1000, and the small target detection model O of the tth iteration t The weight and bias parameters in are w t 、b t , and let t = 0, O t =O; (4) Train the small target detection model O: Randomly and with replacement select L training samples from the training sample set R1 as the input of the small target detection model O for forward propagation to obtain L small target detection results, where 1≤L≤M; (5) Update the parameters of the small target detection model: The weights, bias parameters w t and b t of the small target detection model O t are updated based on the L small target detection results obtained in step (4) to obtain the network model O t for this iteration; determine whether t ≥ T holds. If so, obtain the trained small target detection model O*. Otherwise, set t = t + 1 and execute step (4); (6) Obtain the small target detection results: Use the test sample set E1 as the input of the trained small target detection model O* for forward propagation to obtain the small target detection results corresponding to K - M test samples.
2. The small target detection method based on hydrodynamic guidance according to claim 1, wherein The small target detection network model O described in step (2), where: The Stem block includes a plurality of convolution layers and a pooling layer connected in sequence; the feature extraction block includes a plurality of convolution layers connected in sequence; the prediction block includes a convolution layer, a normalization layer, a non-linear activation layer, a random dropout layer, a convolution layer, and an interpolation layer connected in sequence.
3. The small target detection method based on hydrodynamic guidance according to claim 1, characterized in that The training of the small target detection model O described in step (4) is implemented as follows: (4a) The Stem block downsamples each image to obtain L preprocessed feature maps, and N feature extraction blocks extract features from each preprocessed image to obtain L local feature maps; (4b) N feature reconstruction blocks use their corresponding low-level feature maps to perform feature image reconstruction on the high-level feature maps to obtain L feature reconstruction feature maps. Among them, the corresponding low-level feature map in the first feature reconstruction block is the output feature map of the N - 1th feature extraction block, and the high-level feature map is the output feature map of the Nth feature extraction block. The corresponding low-level feature maps in the 2nd to Nth feature reconstruction blocks are the output feature maps of the N - nth feature extraction blocks, and the high-level feature maps are the output feature maps of the previous-level feature reconstruction blocks; (4c) The prediction block performs target detection on each feature reconstruction feature map output by the Nth feature reconstruction block to obtain L small target detection results.
4. The small target detection method based on hydrodynamic guidance according to claim 3, characterized in that, The specific process of the N feature reconstruction blocks described in step (4b) using their corresponding low-level feature maps to perform feature image reconstruction on the high-level feature maps is as follows: In the nth feature reconstruction block, the first convolutional block adjusts the number of channels of the low-level feature map and extracts features, and the fluid dynamics-guided detail reconstruction block reconstructs the edge detail features of the output feature map of the first convolutional block to obtain an edge detail reconstruction feature map The deconvolution layer upsamples the high-level feature map, and the second convolutional block adjusts the number of channels of the output feature map of the deconvolution layer and extracts features to obtain an upsampled feature map Edge detail reconstruction feature map and the upsampled feature map are element-wise added to obtain L feature reconstruction feature maps Among them, the fluid mechanics-guided detailed reconstruction block reconstructs the edge detail features of the output feature map of the first convolutional block to obtain an edge detail reconstruction feature map The specific process can be expressed as: U h = [cos 0°, -sin 0°; sin 0°, cos 0°] U v = [cos90°, -sin90°; sin90°, cos90°] Among them, s1(·) is the first non-linear activation layer, s2(·) is the second non-linear activation layer, G h (·) is the Gaussian elliptical operator in the horizontal direction, G v (·) is the Gaussian elliptical operator in the vertical direction, is the coordinates of adjacent pixels and the central pixel in the image , Σ h is the covariance matrix in the horizontal direction, U h is the orthogonal matrix in the horizontal direction, Λ h is the diagonal matrix in the horizontal direction, Σ v is the covariance matrix in the vertical direction, U v is the orthogonal matrix in the vertical direction, Λ v is the diagonal matrix in the vertical direction, and σ1, σ2, σ3, σ4 are elongation rate parameters.
5. The small target detection method based on hydrodynamic guidance according to claim 1, characterized in that The update of the parameters of the small target detection model described in step (5) is implemented as follows: (5a) Using the Dice loss function $L$ Dice and the cross-entropy loss function $L$ CE , calculate the loss value $L$ of the small object detection model through the generated $L$ small object detection results and the labels of the corresponding $L$ images t : L t = L Dice + L CE where r l represents the label category of each pixel in the input image, p l represents the probability that each pixel in the small target detection result belongs to the label category, and ε represents the correction factor; (5b) Calculate L by the chain rule t For the weight parameter ω t and the bias parameter b t partial derivatives and Finally, according to for ω t and b t perform the update: Among them, ω t , b t Indicates O t The weights and bias parameters of all learnable parameters, w t ', b t ' indicates ω t , b t The update result of , α represents the learning rate.
Citation Information
Patent Citations
Small target detection method based on high and low layer feature fusion
CN116229135A
Online detection and identification method for rare animal protection
CN110837768A
Method for determining a set of optical imaging functions for three-dimensional flow measurement
US20120274746A1