Micro-target detection model training method and device, and electronic equipment
Patent Information
- Application Number
- CN202410641487.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-05-22
AI Technical Summary
[0004]本申请提供一种微小目标检测模型训练方法、装置及电子设备,用以解决现有技术中存在的检测效果不好的问题
[0054]The small target detection model training method, apparatus, and electronic equipment provided in this application obtain a reconstructed image by decoding the downsampled features of the training image. This improves the representation of the downsampled features and increases their expressive power, while also facilitating comparison with the training image, thus enhancing the model training effect. By determining target parameters that characterize the mutual information between the reconstructed image and the small target, the mutual information between the reconstructed image and the small target is maximized. This allows the model to filter background regions and focus on the differences between the small target and the background, thereby enhancing the model's feature discrimination of small targets and ultimately improving the model's ability to recognize small targets.
Smart Images

Figure CN118506154B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus and electronic device for training a small target detection model. Background Technology
[0002] Tiny target detection refers to the accurate location and identification of tiny objects in complex scenes, and can be applied to scenarios such as driver assistance, traffic management, and maritime rescue.
[0003] Current technology still uses traditional object detectors for the detection of small targets, but the detection effect is poor because the number of small targets is small and the pixel count is low. Summary of the Invention
[0004] This application provides a method, apparatus, and electronic device for training a small target detection model, in order to solve the problem of poor detection performance in the prior art.
[0005] Firstly, this application provides a method for training a small target detection model, including:
[0006] Determine the training images and their training labels, where the training labels are labels used to mark small targets in the training images;
[0007] The downsampling features of the training image are decoded to obtain the reconstructed image, wherein the size of the reconstructed image and the size of the training image meet the preset image size requirements;
[0008] Based on the reconstructed image and the training labels of the training image, the target parameters are determined. The target parameters are used to characterize the mutual information between the small targets in the reconstructed image and the training labels.
[0009] Based on the target parameters and the initial loss function, the updated loss function is obtained. The initial loss function is determined based on the first relationship between the training image and the downsampled features, and the second relationship between the training labels and the downsampled features.
[0010] Based on the updated loss function, the model to be trained is adjusted to obtain the target model.
[0011] In this application, when there are two or more downsampling features,
[0012] Decoding the downsampled features of the training image yields a reconstructed image, including:
[0013] The downsampled features are then stitched together to obtain the stitched result;
[0014] Based on the splicing result, a first upsampling result and a second upsampling result are obtained. The first upsampling result is obtained by inputting the splicing result into the first bilinear interpolation operation layer for upsampling processing. The second upsampling result is obtained by inputting the second integrated result into the second bilinear interpolation operation layer for upsampling processing. The second integrated result is obtained based on the residual feature extraction result and the convolution result. The residual feature extraction result is obtained by inputting the convolution result into the residual feature extraction module. The convolution result is obtained by inputting the splicing result into the first convolutional layer.
[0015] Based on the first upsampling result and the second upsampling result, the first integrated result is obtained;
[0016] The first integration result is input into the linear layer for channel dimension compression to obtain the compressed result;
[0017] The compressed result is input into a pixel blending layer for reconstruction processing to obtain the reconstructed image.
[0018] In this application, the downsampling features are spliced together to obtain the spliced result, including:
[0019] The downsampled features are input into the deconvolution layer for restoration processing to obtain the restored result;
[0020] The restored result is input into the third bilinear interpolation layer for upsampling processing to obtain the third upsampling result;
[0021] The third upsampling result is then stitched together to obtain the stitched result.
[0022] In this application, based on the splicing results, a first upsampling result and a second upsampling result are obtained, including:
[0023] The splicing result is input into the first bilinear interpolation layer for upsampling processing to obtain the first upsampling result;
[0024] The concatenated result is input into the first convolutional layer for convolution processing to obtain the convolution result;
[0025] The convolution result is input into the residual feature extraction module for residual feature extraction processing to obtain the residual feature extraction result;
[0026] The residual feature extraction results and the convolution results are integrated to obtain a second integrated result;
[0027] The second integration result is input into the second bilinear interpolation layer for upsampling processing to obtain the second upsampling result.
[0028] In this application, an updated loss function is obtained based on the target parameters and the initial loss function, including:
[0029] Based on the training labels, a binary mask for the tiny target is determined. The binary mask is used to identify the location information of the tiny target in the training image.
[0030] The training image is segmented using a binary mask to obtain images of small targets and background images;
[0031] The target parameters are updated based on the image of the small target to obtain the first updated target parameters;
[0032] The initial loss function is adjusted based on the first update target parameter to obtain the updated loss function.
[0033] In this application, the initial loss function is adjusted according to the first update target parameter to obtain the updated loss function, including:
[0034] Obtain the reconstructed target image corresponding to the small target image, and the reconstructed background image of the background image, wherein the reconstructed target image and the reconstructed background image are images in the reconstructed image;
[0035] The downsampling features of the target image and the downsampling features of the background image are decoded separately to obtain the reconstructed target image corresponding to the downsampling features of the target image and the reconstructed background image corresponding to the downsampling features of the background image.
[0036] Based on the reconstructed target image, the reconstructed background image, and the small target image, the first updated target parameters are updated to obtain the second updated target parameters;
[0037] The initial loss function is adjusted based on the second update target parameter to obtain the updated loss function.
[0038] In this application, the second update target parameter satisfies:
[0039] I(φ(F;X) f )=∥∥φ(F f )-X f ∥∥1+∥φ(F b )∥1,
[0040] Wherein, I(φ(F); X f ) is the second update target parameter, φ(F) f To reconstruct the target image, X f For the target image, φ(F) b () is used to reconstruct the background image.
[0041] In this application, the update loss function satisfies:
[0042]
[0043] in, To update the loss function, Let λ be the initial loss function and λ be the adjustment coefficient.
[0044] Secondly, this application provides a training device for a small target detection model, comprising:
[0045] The first determining module is used to determine the training image and the training label of the training image, wherein the training label is a label for marking the small targets in the training image;
[0046] The decoding module is used to decode the downsampling features of the training image to obtain the reconstructed image, wherein the size of the reconstructed image and the size of the training image meet the preset image size requirements.
[0047] The second determining module is used to determine the target parameters based on the reconstructed image and the training labels of the training image. The target parameters are used to characterize the mutual information between the small targets in the reconstructed image and the training labels.
[0048] The module is used to obtain the updated loss function based on the target parameters and the initial loss function. The initial loss function is determined based on the first relationship between the training image and the downsampled features, and the second relationship between the training labels and the downsampled features.
[0049] The adjustment module is used to adjust the model to be trained based on the updated loss function to obtain the target model.
[0050] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0051] The memory stores the instructions that the computer executes;
[0052] The processor executes computer execution instructions stored in memory to implement the method provided in this application.
[0053] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of this application.
[0054] The small target detection model training method, apparatus, and electronic equipment provided in this application obtain a reconstructed image by decoding the downsampled features of the training image. This improves the representation of the downsampled features and increases their expressive power, while also facilitating comparison with the training image, thus enhancing the model training effect. By determining target parameters that characterize the mutual information between the reconstructed image and the small target, the mutual information between the reconstructed image and the small target is maximized. This allows the model to filter background regions and focus on the differences between the small target and the background, thereby enhancing the model's feature discrimination of small targets and ultimately improving the model's ability to recognize small targets. Attached Figure Description
[0055] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0056] Figure 1 A schematic diagram illustrating a scenario for training a small target detection model, as provided in an embodiment of this application.
[0057] Figure 2 A flowchart illustrating a method for training a small target detection model provided in an embodiment of this application;
[0058] Figure 3 This is a schematic diagram of the decoding process provided in an embodiment of this application;
[0059] Figure 4 This is a schematic diagram of the structure of a training device for a small target detection model provided in an embodiment of this application;
[0060] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0061] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0062] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0063] To clearly understand the technical solution of this application, the solutions of the prior art will be described in detail first.
[0064] Tiny target detection refers to the accurate location and identification of tiny targets in complex scenes, and can be applied to scenarios such as driving assistance, traffic management, and maritime rescue.
[0065] Current technology still uses traditional object detectors for the detection of small targets, but the detection effect is poor because the number of small targets is small and the pixel count is low.
[0066] To address the aforementioned issue of poor detection performance, the inventors discovered in their research that target parameters can be used to maximize the mutual information between the reconstructed image and the target image, thereby improving the discriminative power of the model under training for small targets and thus enhancing the model's detection performance for small targets.
[0067] Among these, "small targets" can be vehicles, ships, or aircraft in aerial photographs, or people in surveillance images. A small target can refer to image content smaller than a preset size, such as 16×16 pixels. For example, in an aerial photograph containing several vehicles, if a vehicle's pixel size is smaller than 16×16, then that vehicle can be identified as a small target.
[0068] In other implementations, a small target can also refer to image content whose area occupies less than a preset percentage. For example, the preset percentage can be 1%. An aerial photograph containing several vehicles has a pixel size of 400×400. One of the vehicles has a pixel size of 16×16, and the pixel occupancy of that vehicle is 0.16%. Since 0.16% is less than the preset percentage of 1%, the vehicle can be identified as a small target.
[0069] The following describes the application scenarios of the micro-target detection model training provided in the embodiments of this application.
[0070] Figure 1 This is a schematic diagram illustrating a scenario for training a small target detection model, as provided in an embodiment of this application. Figure 1As shown, the execution entity of this small target detection model training method can be a server. The server can be a mobile phone, tablet, computer, or other device. This embodiment does not impose any particular restrictions on the implementation method of the execution entity, as long as the execution entity can determine the training image and its training label; decode the downsampling features of the training image to obtain a reconstructed image; determine the target parameters based on the reconstructed image and the training label; obtain an updated loss function based on the target parameters and the initial loss function, wherein the initial loss function is determined based on the first relationship between the training image and the downsampling features, and the second relationship between the training label and the downsampling features; and adjust the model to be trained based on the updated loss function to obtain the target model.
[0071] Figure 2 This is a flowchart illustrating a method for training a small target detection model according to an embodiment of this application, as shown below. Figure 2 As shown, the method includes:
[0072] S201. Determine the training image and the training label of the training image, wherein the training label is a label used to mark the small targets in the training image.
[0073] The training images can refer to images containing small targets, such as aerial photographs of vehicles or ships.
[0074] Training labels can refer to the labels attached to small targets such as vehicles or ships in the training images. In some embodiments, training labels can be displayed on the training images as line boxes.
[0075] S202. Decode the downsampling features of the training image to obtain the reconstructed image, wherein the size of the reconstructed image and the size of the training image meet the preset image size requirements.
[0076] Downsampling features can refer to the features obtained after downsampling the training images. In this embodiment, downsampling features can also be called intermediate features, high-level features, etc.
[0077] Downsampling refers to the process of reducing the resolution of an image. Specifically, downsampling can include image processing based on an image pyramid.
[0078] Decoding can refer to the process of extracting features from downsampled data and converting them back into the original data.
[0079] Reconstructed images can refer to the results obtained after decoding downsampled features.
[0080] The preset image size requirement can mean that the size of the reconstructed image is the same as the size of the training image, which makes it easier to construct a loss function based on the reconstructed image and the training image.
[0081] In this embodiment of the application, when there are two or more downsampling features...
[0082] Decoding the downsampled features of the training image yields a reconstructed image, including:
[0083] The downsampled features are then stitched together to obtain the stitched result;
[0084] Based on the splicing result, a first upsampling result and a second upsampling result are obtained. The first upsampling result is obtained by inputting the splicing result into the first bilinear interpolation operation layer for upsampling processing. The second upsampling result is obtained by inputting the second integrated result into the second bilinear interpolation operation layer for upsampling processing. The second integrated result is obtained based on the residual feature extraction result and the convolution result. The residual feature extraction result is obtained by inputting the convolution result into the residual feature extraction module. The convolution result is obtained by inputting the splicing result into the first convolutional layer.
[0085] Based on the first upsampling result and the second upsampling result, the first integrated result is obtained;
[0086] The first integration result is input into the linear layer for channel dimension compression to obtain the compressed result;
[0087] The compressed result is input into a pixel blending layer for reconstruction processing to obtain the reconstructed image.
[0088] The concatenation process refers to merging downsampled features into a new feature. Specifically, each downsampled feature can be processed to make its size consistent, and then concatenated along the channel dimension of the downsampled features to obtain the concatenated result.
[0089] For example, suppose we have two features A and B, with shapes A:[batch_size,height,width,channels_A] and B:[batch_size,height,width,channels_B], where batch_size represents the batch size, height and width represent the height and width of the feature, and channels_A and channels_B represent the number of channels in features A and B, respectively. If we want to concatenate A and B, we concatenate these two features along the channel dimension to obtain a new feature C with the shape C:[batch_size,height,width,channels_A+channels_B].
[0090] The splicing result can refer to the result after splicing downsampled features.
[0091] In this embodiment of the application, the downsampling features are spliced to obtain a splicing result, including:
[0092] The downsampled features are input into the deconvolution layer for restoration processing to obtain the restored result;
[0093] The restored result is input into the third bilinear interpolation layer for upsampling processing to obtain the third upsampling result;
[0094] The third upsampling result is then stitched together to obtain the stitched result.
[0095] Among them, the deconvolutional layer is a type of layer in deep learning neural networks, used to implement the restoration operation. The deconvolutional layer maps the downsampled features to a larger output feature map, increasing the spatial dimension of the downsampled features.
[0096] Restoration processing can refer to the process of processing downsampled features through deconvolution layers.
[0097] A bilinear interpolation layer refers to the process of interpolating discrete pixel values in an image using bilinear interpolation to obtain more accurate pixel values. By using bilinear interpolation, smooth pixel value estimation can be achieved in image processing, improving the accuracy and quality of image processing.
[0098] The bilinear interpolation operation layer may include a first bilinear interpolation operation layer, a second bilinear interpolation operation layer, and a third bilinear interpolation operation layer.
[0099] The restoration result is upsampled by a third bilinear interpolation layer to improve its resolution, resulting in a third upsampled result. The restoration result is determined based on downsampling features and a deconvolution layer.
[0100] The splicing result can refer to the result obtained after splicing multiple third-party upsampling results.
[0101] The first upsampling result can refer to the result after upsampling the splicing result through the first bilinear interpolation operation layer.
[0102] The second upsampling result can refer to the result after upsampling the second integrated result through the second bilinear interpolation operation layer.
[0103] The second integration result can refer to the result after integrating the residual feature extraction result and the convolution result.
[0104] Integration processing refers to combining feature information from different levels or sources to improve the model's understanding and representation of input data, thereby enhancing the model's performance and generalization ability. Integration processing can be achieved through different methods, including concatenation, addition, multiplication, attention mechanisms, etc.
[0105] 1. Concatenation: This method combines features from different levels or sources according to certain rules to form a longer feature vector. It is commonly used for multi-scale feature fusion or multi-modal information fusion.
[0106] 2. Addition: Adding different features together directly makes it easier for the model to learn the correlation between features, which helps to improve the model's representation ability.
[0107] 3. Multiplication: Multiplying different features element by element can introduce interactive information and help learn the correlation between features.
[0108] 4. Attention Mechanism: By learning attention weights, the importance of different features is dynamically adjusted, enabling the model to pay more attention to features that are useful for the current task.
[0109] The residual feature extraction result can refer to the result after the residual feature extraction module performs residual feature extraction processing on the convolution result.
[0110] The residual feature extraction module may include a second convolutional layer, a nonlinear layer, and a third convolutional layer.
[0111] The first, second, and third convolutional layers can be different convolutional layers. For example, the parameters of the first convolutional layer can be kernel=1, channel=128, stride=1; the parameters of the second convolutional layer can be kernel=3, channel=16, stride=1; and the parameters of the third convolutional layer can be kernel=3, channel=128, stride=1.
[0112] A convolutional layer is a commonly used layer type in deep learning neural networks, used to perform convolution operations on input data.
[0113] A nonlinear layer is a type of layer in a neural network that uses a nonlinear function to perform a nonlinear transformation on the linear output, enabling the neural network to learn and represent more complex patterns and relationships.
[0114] Nonlinear functions can include rectified linear units (ReLU), sigmoid functions, hyperbolic tangent functions (Tanh), etc.
[0115] The convolution result can refer to the result obtained after performing convolution processing on the spliced result through the first convolutional layer.
[0116] The first integration result can refer to the result after integrating the first upsampling result and the second upsampling result.
[0117] A linear layer is a basic layer type in deep learning neural networks, also known as a fully connected layer or a dense layer. In this embodiment, the linear layer uses a built-in convolutional neural network to perform channel dimension compression. The parameters of the convolutional neural network can be kernel=1, channel=12, and stride=1.
[0118] PixelShuffle can refer to a deep learning-based upsampling method that learns the mapping relationship between low-resolution and high-resolution images, upsamples low-resolution images into high-resolution images, and reconstructs the compressed results through the pixel shuffle layer, resulting in better reconstruction effects and improved computational efficiency.
[0119] In this embodiment of the application, obtaining the first upsampling result and the second upsampling result based on the splicing result includes:
[0120] The splicing result is input into the first bilinear interpolation layer for upsampling processing to obtain the first upsampling result;
[0121] The concatenated result is input into the first convolutional layer for convolution processing to obtain the convolution result;
[0122] The convolution result is input into the residual feature extraction module for residual feature extraction processing to obtain the residual feature extraction result;
[0123] The residual feature extraction results and the convolution results are integrated to obtain a second integrated result;
[0124] The second integration result is input into the second bilinear interpolation layer for upsampling processing to obtain the second upsampling result.
[0125] S203. Based on the reconstructed image and the training labels of the training image, determine the target parameters. The target parameters are used to characterize the mutual information between the small targets in the reconstructed image and the training labels.
[0126] Mutual information can be used to measure the interdependence or correlation between two random variables. It represents the amount of information contained in one random variable about another, or the reduced uncertainty of one random variable due to knowledge of the other.
[0127] By using the mutual information between the reconstructed image and the small target as the target parameter, and then adjusting the initial loss function through the target parameter, the mutual information between the reconstructed image and the small target can be maximized, thereby improving the highly discriminative representation of the small target.
[0128] The target parameters can be expressed as: I(φ(F); X f ), where φ(.) represents the decoding process, F represents the downsampling feature, and X f For images of tiny targets, X f It can also be called micro-target features, foreground features, or foreground information. Other images in the training image besides micro-targets can be called background images, background features, or background information.
[0129] I(φ(F;X) f ) can be approximated as I(φ(F) f ); X f )+I(φ(F f ); 0), because small target features and background features are mutually exclusive in space. The first term I(φ(F) f ); X f The term I(φ(F) indicates activation of foreground features and reduction of foreground feature degradation during feature extraction. f ) ; 0) is used to remove noise from background features and to distinguish foreground features from background features. When I(φ(F f When 0)=0, it is equivalent to ∥φ(F) b )∥1=0,F b This is the background image.
[0130] S204. Based on the target parameters and the initial loss function, the updated loss function is obtained, wherein the initial loss function is determined based on the first relationship between the training image and the downsampled features, and the second relationship between the training label and the downsampled features.
[0131] The initial loss function can refer to the loss function determined based on the first relation and the second relation, where the first relation can refer to the mutual information between the training image and the downsampled features, and the second relation can refer to the mutual information between the training label and the downsampled features.
[0132] The initial loss function can be expressed as:
[0133]
[0134] Where X is the training image, F is the downsampling feature, and y GT Training labels, δ is the adjustment coefficient.
[0135] The update loss function can be expressed as:
[0136]
[0137] Where λ is the adjustment coefficient.
[0138] In this embodiment of the application, the updated loss function is obtained based on the target parameters and the initial loss function, including:
[0139] Based on the training labels, a binary mask for the tiny target is determined. The binary mask is used to identify the location information of the tiny target in the training image.
[0140] The training image is segmented using a binary mask to obtain images of small targets and background images;
[0141] The target parameters are updated based on the image of the small target to obtain the first updated target parameters;
[0142] The initial loss function is adjusted based on the first update target parameter to obtain the updated loss function.
[0143] Binary masks can be used to identify the location of small targets by assigning a 1 to the position of the target and a 0 to other positions. A binary mask, also known as a foreground mask, can be expressed as follows:
[0144] M i,j =1[(i,j)∈B]
[0145] Where M i,j Let be the binary mask for position (i,j). The indicator function 1[(i,j)∈B] indicates that if a position (i,j) belongs to a small target, its value is 1, otherwise it is 0.
[0146] Methods for determining a binary mask may include determining the location information of tiny targets in a training image based on training labels. The location information can be represented as the range of pixels where the tiny target is located. Based on the location information of the tiny target, a binary mask can be determined, converting image information into data information, which is convenient for computer recognition and calculation.
[0147] Segmentation processing can refer to the process of dividing the content of a training image into a small target image and a background image using a binary mask.
[0148] A micro-target image can refer to an image composed of pixels containing a micro-target.
[0149] The process of updating the target parameters based on the image of the small target to obtain the first updated target parameters can be described as follows:
[0150] First, the target parameters are expanded and expressed as:
[0151] I(φ(F;X) f )=H(φ(F))-H(φ(F)∣X f )
[0152] Where H(·) is the information entropy.
[0153] Then, Represented as a small target image, X f Replace with Obtain the first update target parameters:
[0154]
[0155] Where M is the binary mask and X is the training image. This represents the Hadamard product.
[0156] Adjusting the initial loss function based on the first update target parameter, the resulting updated loss function can be expressed as:
[0157]
[0158] In this embodiment of the application, adjusting the initial loss function according to the first update target parameter to obtain the updated loss function includes:
[0159] Obtain the reconstructed target image corresponding to the small target image, and the reconstructed background image of the background image, wherein the reconstructed target image and the reconstructed background image are images in the reconstructed image;
[0160] Based on the reconstructed target image, the reconstructed background image, and the target image, the first updated target parameters are updated to obtain the second updated target parameters;
[0161] The initial loss function is adjusted based on the second update target parameter to obtain the updated loss function.
[0162] The methods for obtaining the target image and the background image for reconstruction may include:
[0163] Obtain the target image downsampling features of the small target image, and the background image downsampling features of the background image;
[0164] The downsampling features of the target image and the downsampling features of the background image are decoded separately to obtain the reconstructed target image corresponding to the downsampling features of the target image and the reconstructed background image corresponding to the downsampling features of the background image.
[0165] Alternatively, the reconstructed image can be decoupled using a binary mask to obtain the reconstructed target image and the reconstructed background image.
[0166] The process of updating the first target parameters based on the reconstructed target image, the reconstructed background image, and the target image to obtain the second target parameters can be as follows:
[0167] Because of the first update target parameters It is implicit. Alternatively, methods from super-resolution tasks can be used to minimize the l1 norm distance, such that φ(F) and... The distance between them is minimized. Then, the generated binary mask is used to separate the supervision for refining discriminative features, with the following formula:
[0168]
[0169] Since the training image can be decoupled into a small target image and a background image through a binary mask, φ(F) can be expressed as φ(F) = φ(F) f )+φ(F b ), where F f To reconstruct the target image, F b To reconstruct the background image, the first update target parameters are then updated to obtain the second update target parameters. In this embodiment, the second update target function satisfies:
[0170] I(φ(F;X) f )=∥∥φ(F f )-X f ∥∥1+∥φ(F b )∥1.
[0171] The initial loss function is adjusted based on the second update target parameter to obtain the updated loss function.
[0172] In this embodiment of the application, the update loss function satisfies:
[0173]
[0174] S205. Adjust the model to be trained according to the updated loss function to obtain the target model.
[0175] The model to be trained can refer to a detector, an algorithm or model used to detect specific targets, objects, or features in an image. These detectors are commonly used in computer vision tasks such as object detection, object recognition, and face detection. Detectors play a crucial role in the field of image vision, automatically identifying specific targets or features in images and providing a foundation for subsequent analysis, recognition, or decision-making. Image vision detectors include:
[0176] 1. Object Detector: Used to detect target objects in an image, such as pedestrians, vehicles, and animals. Object detectors typically output the target's location, bounding box, and category information.
[0177] 2. Face Detector: Specifically designed to detect facial regions in images, typically used in applications such as face recognition and facial expression analysis.
[0178] 3. Feature Point Detector: Used to detect key feature points in an image, such as corner points and edge points. Feature point detectors are commonly used in tasks such as image registration and object tracking.
[0179] 4. Edge Detector: Used to detect edge information in an image, helping to extract the image's contour and structural features.
[0180] The object detector can include a backbone network, a neck, and a detection head. The backbone network is the main feature extraction part of the model, used to extract low-level features (such as edges and textures) and high-level features (such as object shape and structure) from the image. It can be composed of a deep convolutional neural network (CNN). The choice of backbone network can depend on the complexity of the task and the input image. The backbone network can include:
[0181] ResNet (Residual Network): Due to its residual connection structure, it is possible to train very deep networks.
[0182] VGGNet (Visual Geometry Group Network): It has a simple, repetitive structure and is suitable for small datasets.
[0183] MobileNet (Lightweight Network): A lightweight network designed specifically for mobile devices and embedded systems.
[0184] EfficientNet: Improves model efficiency by automating the search for network structures.
[0185] The neckline is an intermediate layer between the backbone network and the head. Its role is to further perform feature fusion, context enhancement, and other operations based on the features extracted by the backbone network. The neckline structure can include convolutional layers, pooling layers, attention mechanisms, etc. The neckline allows the model to perceive targets at different scales and provides more contextual information. The neckline structure can include:
[0186] FPN (Feature Pyramid Network): Used to process feature maps at different scales, helping the network detect targets at different scales.
[0187] ASPP (Atrous Spatial Pyramid Pooling): This method uses dilated convolutions with different sampling rates to increase the receptive field, and is used for image segmentation tasks.
[0188] SAM (Spatial Attention Module): Introduces an attention mechanism that enables the network to automatically focus on regions of interest.
[0189] The head is the output part of the model, responsible for the final task prediction. For object detection tasks, the head can include a classification head (for predicting the object category) and a regression head (for predicting the location of the object bounding box). For image segmentation tasks, the head can be a segmentation network that outputs a class label for each pixel.
[0190] Classification Head: The softmax function can be used to map features to a class distribution to predict the class of each target.
[0191] Regression Head: Outputs the coordinates of the target bounding box (such as the coordinates of the top left and bottom right corners of the bounding box), used to locate the target.
[0192] The design of the head is closely related to the specific task. For example, in object detection, common head structures include the RPN (Region Proposal Network) in Faster R-CNN and the ROI Align network in Fast R-CNN.
[0193] In some implementations, training images can be input into the backbone network of the target detector, and then the downsampled features output from the neck can be input into the decoding unit for decoding to obtain a reconstructed image. Then, the small target images in the training images can be segmented using training labels. The updated loss function can be obtained by using the reconstructed image and the small target images. Finally, the learning parameters in the detector can be adjusted by using the updated loss function to obtain the target model.
[0194] This application provides a method for training a small target detection model. It obtains a reconstructed image by decoding the downsampled features of the training image. This improves the representation of the downsampled features, increases their expressive power, and facilitates comparison with the training image, thus enhancing the model's training effect. The method also segments the training image using a binary mask to obtain a small target image, mimicking how humans filter background regions and focus on the differences between the small target and the background when perceiving small targets. This enhances the model's feature discrimination ability for small targets. Finally, the initial loss function is adjusted based on the target parameters determined from the reconstructed image and the small target image, thereby improving the model's ability to recognize small targets.
[0195] Figure 3 This is a schematic diagram of the decoding process provided in the embodiments of this application, such as... Figure 3 As shown, the decoding process is as follows:
[0196] Multiple downsampled features are input into the corresponding deconvolution layers to obtain the restored results;
[0197] The restoration result is input into the third bilinear interpolation layer to obtain the third upsampling result corresponding to each downsampling feature;
[0198] The individual third upsampling results are then concatenated to obtain the concatenated result.
[0199] The concatenation results are input into the first convolutional layer and the first bilinear interpolation layer respectively to obtain the convolution result and the first upsampling result;
[0200] The convolution result is input into the residual module to obtain the residual feature extraction result;
[0201] The residual feature extraction results and the convolution results are integrated to obtain the second integrated result;
[0202] The second integration result is input into the second bilinear interpolation layer to obtain the second upsampling result;
[0203] The second upsampling result is combined with the first upsampling result to obtain the first integrated result;
[0204] The first integration result is input into the linear layer to obtain the compression result;
[0205] The compressed result is input into the pixel blending layer to obtain the reconstructed image.
[0206] Figure 4 This is a schematic diagram of the structure of a small target detection model training device provided in an embodiment of this application, as shown below. Figure 4 As shown, the device includes a first determining module 401, a decoding module 402, a second determining module 403, a obtaining module 404, and an adjusting module 405, wherein:
[0207] The first determining module 401 is used to determine the training image and the training label of the training image, wherein the training label is a label for marking the small targets in the training image;
[0208] The decoding module 402 is used to decode the downsampling features of the training image to obtain the reconstructed image, wherein the size of the reconstructed image and the size of the training image meet the preset image size requirements.
[0209] The second determining module 403 is used to determine target parameters based on the reconstructed image and the training labels of the training image. The target parameters are used to characterize the mutual information between the small targets in the reconstructed image and the training labels.
[0210] Module 404 is used to obtain an updated loss function based on the target parameters and the initial loss function, wherein the initial loss function is determined based on the first relationship between the training image and the downsampled features, and the second relationship between the training labels and the downsampled features;
[0211] The adjustment module 405 is used to adjust the model to be trained according to the updated loss function to obtain the target model.
[0212] In this embodiment of the application, the decoding module 402 is further configured to:
[0213] The downsampled features are then stitched together to obtain the stitched result;
[0214] Based on the splicing result, a first upsampling result and a second upsampling result are obtained. The first upsampling result is obtained by inputting the splicing result into the first bilinear interpolation operation layer for upsampling processing. The second upsampling result is obtained by inputting the second integrated result into the second bilinear interpolation operation layer for upsampling processing. The second integrated result is obtained based on the residual feature extraction result and the convolution result. The residual feature extraction result is obtained by inputting the convolution result into the residual feature extraction module. The convolution result is obtained by inputting the splicing result into the first convolutional layer.
[0215] Based on the first upsampling result and the second upsampling result, the first integrated result is obtained;
[0216] The first integration result is input into the linear layer for channel dimension compression to obtain the compressed result;
[0217] The compressed result is input into a pixel blending layer for reconstruction processing to obtain the reconstructed image.
[0218] In this embodiment of the application, the decoding module 402 is further configured to:
[0219] The downsampled features are input into the deconvolution layer for restoration processing to obtain the restored result;
[0220] The restored result is input into the third bilinear interpolation layer for upsampling processing to obtain the third upsampling result;
[0221] The third upsampling result is then stitched together to obtain the stitched result.
[0222] In this embodiment of the application, the decoding module 402 is further configured to:
[0223] The splicing result is input into the first bilinear interpolation layer for upsampling processing to obtain the first upsampling result;
[0224] The concatenated result is input into the first convolutional layer for convolution processing to obtain the convolution result;
[0225] The convolution result is input into the residual feature extraction module for residual feature extraction processing to obtain the residual feature extraction result;
[0226] The residual feature extraction results and the convolution results are integrated to obtain a second integrated result;
[0227] The second integration result is input into the second bilinear interpolation layer for upsampling processing to obtain the second upsampling result.
[0228] In this embodiment of the application, module 404 is further configured to:
[0229] Based on the training labels, a binary mask for the tiny target is determined. The binary mask is used to identify the location information of the tiny target in the training image.
[0230] The training image is segmented using a binary mask to obtain images of small targets and background images;
[0231] The target parameters are updated based on the image of the small target to obtain the first updated target parameters;
[0232] The initial loss function is adjusted based on the first update target parameter to obtain the updated loss function.
[0233] In this embodiment of the application, module 404 is further configured to:
[0234] Obtain the reconstructed target image corresponding to the small target image, and the reconstructed background image of the background image, wherein the reconstructed target image and the reconstructed background image are images in the reconstructed image;
[0235] The downsampling features of the target image and the downsampling features of the background image are decoded separately to obtain the reconstructed target image corresponding to the downsampling features of the target image and the reconstructed background image corresponding to the downsampling features of the background image.
[0236] Based on the reconstructed target image, the reconstructed background image, and the small target image, the first updated target parameters are updated to obtain the second updated target parameters;
[0237] The initial loss function is adjusted based on the second update target parameter to obtain the updated loss function.
[0238] In this embodiment of the application, module 404 is further configured to:
[0239] I(φ(F;X) f )=∥∥φ(F f )-X f ∥∥1+∥φ(F b )∥1,
[0240] Wherein, I(φ(F); X f ) is the second update target parameter, φ(F) f To reconstruct the target image, X f For the target image, φ(F) b () is used to reconstruct the background image.
[0241] In this embodiment of the application, module 404 is further configured to:
[0242]
[0243] in, To update the loss function, Let λ be the initial loss function and λ be the adjustment coefficient.
[0244] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 50 includes:
[0245] The electronic device 50 may include one or more processors 501 with processing cores, one or more memory 502s of computer-readable storage media, communication components 503, and other components. The processor 501, memory 502, and communication components 503 are connected via a bus 504.
[0246] In the specific implementation process, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to execute the above-mentioned small target detection model training method.
[0247] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0248] In the above Figure 5 In the illustrated embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0249] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0250] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0251] In some embodiments, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the steps in any of the above-described small target detection model training methods.
[0252] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0253] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0254] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the small target detection model training methods provided in embodiments of this application.
[0255] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0256] According to one aspect of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium.
[0257] Since the instructions stored in the storage medium can execute the steps in any of the small target detection model training methods provided in the embodiments of this application, the beneficial effects that any of the small target detection model training methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0258] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0259] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for training a small target detection model, characterized in that, include: Determine training images and training labels for the training images, wherein the training labels are labels used to mark small targets in the training images; The downsampling features of the training image are decoded to obtain a reconstructed image, wherein the size of the reconstructed image and the size of the training image meet a preset image size requirement; When there are two or more downsampling features, the decoding process of the downsampling features of the training image to obtain the reconstructed image includes: The downsampling features are concatenated to obtain a concatenated result. Based on the concatenated result, a first upsampling result and a second upsampling result are obtained. The first upsampling result is obtained by inputting the concatenated result into a first bilinear interpolation layer for upsampling processing. The second upsampling result is obtained by inputting a second integrated result into a second bilinear interpolation layer for upsampling processing. The second integrated result is obtained based on the residual feature extraction result and the convolution result. The residual feature extraction result is obtained by inputting the convolution result into a residual feature extraction module. The convolution result is obtained by inputting the concatenated result into a first convolutional layer. Based on the first upsampling result and the second upsampling result, a first integrated result is obtained. The first integrated result is input into a linear layer for channel dimension compression processing to obtain a compressed result. The compressed result is input into a pixel blending layer for reconstruction processing to obtain a reconstructed image. Based on the reconstructed image and the training labels of the training image, target parameters are determined, wherein the target parameters are used to characterize the mutual information between the reconstructed image and the minute targets in the training labels; An updated loss function is obtained based on the target parameters and the initial loss function, wherein the initial loss function is determined based on a first relationship between the training image and the downsampled features, and a second relationship between the training labels and the downsampled features; Based on the updated loss function, the model to be trained is adjusted to obtain the target model; The step of obtaining the updated loss function based on the target parameters and the initial loss function includes: Based on the training labels, a binary mask for the micro-target is determined, which is used to identify the location information of the micro-target in the training image. The training image is then segmented based on the binary mask to obtain a micro-target image and a background image. The target parameters are updated based on the micro-target image to obtain a first updated target parameter. The initial loss function is adjusted based on the first updated target parameter to obtain an updated loss function.
2. The method according to claim 1, characterized in that, The process of concatenating the downsampled features to obtain the concatenated result includes: The downsampled features are input into a deconvolution layer for restoration processing to obtain the restoration result; The restoration result is input into the third bilinear interpolation layer for upsampling processing to obtain the third upsampling result; The third upsampling result is then spliced together to obtain the spliced result.
3. The method according to claim 1, characterized in that, The step of obtaining the first upsampling result and the second upsampling result based on the splicing result includes: The splicing result is input into the first bilinear interpolation layer for upsampling processing to obtain the first upsampling result; The splicing result is input into the first convolutional layer for convolution processing to obtain the convolution result; The convolution result is input into the residual feature extraction module for residual feature extraction processing to obtain the residual feature extraction result; The residual feature extraction result and the convolution result are integrated to obtain a second integrated result; The second integration result is input into the second bilinear interpolation layer for upsampling processing to obtain the second upsampling result.
4. The method according to claim 1, characterized in that, The step of adjusting the initial loss function according to the first update target parameter to obtain the updated loss function includes: Obtain the reconstructed target image corresponding to the tiny target image and the reconstructed background image of the background image, wherein the reconstructed target image and the reconstructed background image are images in the reconstructed image; The target image downsampling features and the background image downsampling features are decoded respectively to obtain the reconstructed target image corresponding to the target image downsampling features and the reconstructed background image corresponding to the background image downsampling features; Based on the reconstructed target image, the reconstructed background image, and the tiny target image, the first updated target parameters are updated to obtain the second updated target parameters; The initial loss function is adjusted based on the second update target parameter to obtain the updated loss function.
5. The method according to claim 4, characterized in that, The second update target parameter satisfies: , in, For the second update target parameter, For the reconstructed target image, The target image, The reconstructed background image.
6. The method according to claim 4, characterized in that, The update loss function satisfies: , Among them, the To update the loss function, The initial loss function is... This is the adjustment coefficient.
7. A training device for a small target detection model, characterized in that, include: The first determining module is used to determine a training image and a training label for the training image, wherein the training label is a label used to mark small targets in the training image; A decoding module is used to decode the downsampling features of the training image to obtain a reconstructed image, wherein the size of the reconstructed image and the size of the training image meet a preset image size requirement. When there are two or more downsampling features, the decoding module is further used to: concatenate the downsampling features to obtain a concatenation result; obtain a first upsampling result and a second upsampling result based on the concatenation result, wherein the first upsampling result is obtained by inputting the concatenation result into a first bilinear interpolation layer for upsampling processing, and the second upsampling result is obtained by inputting a second integrated result into a second bilinear interpolation layer for upsampling processing, the second integrated result being obtained based on residual feature extraction results and convolution results, wherein the residual feature extraction result is obtained by inputting the convolution result into a residual feature extraction module, and the convolution result is obtained by inputting the concatenation result into a first convolutional layer; obtain a first integrated result based on the first upsampling result and the second upsampling result; input the first integrated result into a linear layer for channel dimension compression processing to obtain a compressed result; and input the compressed result into a pixel blending layer for reconstruction processing to obtain a reconstructed image. The second determining module is used to determine target parameters based on the reconstructed image and the training labels of the training image, wherein the target parameters are used to characterize the mutual information between the reconstructed image and the minute targets in the training labels; The module is configured to obtain an updated loss function based on the target parameters and the initial loss function, wherein the initial loss function is determined based on a first relationship between the training image and the downsampled features, and a second relationship between the training labels and the downsampled features; The adjustment module is used to adjust the model to be trained according to the updated loss function to obtain the target model; The obtaining module is also used for: Based on the training labels, a binary mask for the micro-target is determined, which is used to identify the location information of the micro-target in the training image. The training image is then segmented based on the binary mask to obtain a micro-target image and a background image. The target parameters are updated based on the micro-target image to obtain a first updated target parameter. The initial loss function is adjusted based on the first updated target parameter to obtain an updated loss function.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-6.