A Background Zeroing Mosaic Data Augmentation Method for Small Target Datasets

Through the background zeroing Mosaic data enhancement method, the problem of complex background interference in small object detection is solved. Through the selective local background retention and grid target pasting methods, the performance of small object detection and the generalization ability of the model are improved.

CN116310515BActive Publication Date: 2025-07-01SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310138285.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-07-01
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

In small object detection, the detection performance is degraded due to the small object data set's small data volume, the number of targets, and the characteristics of the targets are easily interfered with in complex backgrounds.

Method used

A background zeroing Mosaic data enhancement method is proposed, which reduces background interference through selective local background zeroing and cropping, and solves the problem of unnaturally overlapping small targets through grid target pasting.

Benefits of technology

It effectively increases the data volume of small target data sets and the detection accuracy of the model, reduces background interference, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310515B_ABST
    Figure CN116310515B_ABST
Patent Text Reader

Abstract

The present invention proposes a background-zeroing Mosaic data augmentation method for small target datasets. This data augmentation method first makes a copy of the original training set, performs a background-zeroing operation with selective local background retention on one of the original training sets, then obtains a small target set through cropping with selective local background retention, then obtains a target-pasted training set through the grid target pasting method, mixes the other original training set with the target-pasted training set, and finally performs Mosaic data augmentation on it. The present invention increases the proportion of effective pixels of small targets in the image, enabling small targets to have the opportunity not to be overwhelmed by complex backgrounds when feature extraction is performed. At the same time, the present invention overcomes the disadvantage of unnatural overlap of small targets caused by means such as cropping and pasting in previous small target datasets, and can effectively improve the recognition accuracy of small target datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning object detection, and relates to a method for zeroing the background and performing Mosaic data augmentation for a small object dataset. Background Art

[0002] With the combination of deep learning theory and practice, object detection technology has also achieved rapid development. The purpose of the object detection task is to obtain the position and type of objects in an image, and it has currently been applied to multiple fields such as object tracking. However, in small object detection, due to the small amount of data in the small object dataset, the small number of objects, and the small proportion of effective pixels occupied by the objects in the image, it is easy to be interfered by a large amount of background during feature extraction, resulting in the loss of small object features, making small object detection face great difficulties and challenges.

[0003] In the data-driven field of deep learning, a better dataset often leads to a more satisfactory network model. The literature "Alexey Bochkovskiy; Chien-Yao Wang; Hong-Yuan Mark Liao. YOLOv4: Optimal Speed and Accuracy of Object Detection[J]. 2020," proposed using the Mosaic data augmentation method to perform random flipping, scaling, color gamut change, etc. on four random pictures, and then randomly arranging and splicing them into one picture according to the positions of upper left, upper right, lower left, and lower right. This not only increases the data diversity, enhances the model robustness, but also improves the small object detection performance. And because of the normalization operation that calculates the data of four pictures at one time, it reduces the memory requirement of the model. The literature "KISANTAL M, WOJNA Z, MURAWSKI J, et al. Augmentation for small object detection[EB / OL]. (2019-02-19)[2019-02-19]." proposed a method of copy augmentation for problems such as the small area covered by small objects, the lack of diversity in the appearance positions, and the intersection over union between the detection box and the ground truth box being much smaller than the expected threshold. By repeatedly copying and pasting small objects in the image to increase the number of training samples of small objects, the detection performance of small objects is improved. Although the above methods alleviate the problem of information loss of small objects to a certain extent, they ignore the interference of complex backgrounds on small object feature extraction. Summary of the Invention

[0004] 1. Object of the Invention:

[0005] The object of the present invention is to propose a background-zeroing Mosaic data augmentation method for small target datasets, which is used to double the data volume and the number of small targets in the small target dataset, increase the proportion of effective pixels of small targets in the image, so that small targets have the opportunity not to be overwhelmed by complex backgrounds when feature extraction is performed. At the same time, it overcomes the disadvantage that small target datasets in the past have unnatural overlaps of small targets caused by means such as cropping and pasting.

[0006] 2. Technical solution:

[0007] A background-zeroing Mosaic data augmentation method for small target datasets, characterized by including the following steps:

[0008] Step 1: Determine all types of small targets in the small target dataset according to the definition of small targets. Divide all types of small targets into small targets class_NO that do not depend on the local background and small targets class_YES that depend on the local background. Let I be the original training set, I0 be the background-zeroing training set obtained by background-zeroing through selective local background retention from I, and S be the small target set obtained by cropping from I0 through selective local background retention, where S contains small target sets of all types. is the target-pasting training set obtained by the grid target pasting method from I0 and S. is obtained by I and through merging to obtain the background-zeroing augmented training set. is obtained by through Mosaic data augmentation to obtain the background-zeroing Mosaic data augmentation training set;

[0009] Among them, background-zeroing through selective local background retention means: According to the type of the corresponding small target, the center point, width, and height of the target box in the corresponding image for each data label, the width and height of the target box of the small target whose small target type belongs to class_YES are magnified by n times. According to the center point, width, and height of each small target box in each image, the corresponding small target box area can be obtained. In each image, all areas outside the small target box areas are regarded as the background of the image, and the pixel values of the background are set to zero. Each new image and the corresponding label form a new dataset.

[0010] Among them, the cropping of selective local background retention means: according to the category of the small target and the center point, width, and height of the target box obtained for each data label in the corresponding image, the width and height of the target box of the small target whose small target category belongs to class_YES are enlarged by n times. According to the center point, width, and height of each small target box in each image, the corresponding small target box area can be obtained. Lock the area of each small target box obtained for each image. If the small target category belongs to class_YES, set the pixel values of all other target box areas within the area of the small target box that are not the original target box area of the small target to zero, and then crop out the locked area of each small target box and put it into the set of small targets of the corresponding category in S, so as to obtain all the cropped small targets;

[0011] Among them, the grid target pasting method means: divide the background zero image to be enhanced into a grid of M×N, the size of a single standard grid is m_grid×n_grid, and judge whether the sum of the pixel values in each standard grid is zero. If it is zero, randomly select a small target image of a certain category from the set of small targets and paste it in the grid, so that the center point of the small target image lands at a random position within a certain range of the grid. And define the width and height of the randomly selected small target box as the width and height of the small target image, and the center point as the generated random landing position. Reduce the width and height of the target box of the small target whose small target category belongs to class_YES in the randomly selected small target by n times. According to the information such as the randomly selected small target category and the center point, width, and height of the randomly selected small target box, generate the corresponding small target label, so as to obtain all the small target labels pasted in the corresponding image;

[0012] Among them, merging means: a simple superposition of two data sets to achieve the expansion of the data set;

[0013] Among them, Mosaic data augmentation means: randomly select four pictures, and perform random operations such as flipping, scaling, and color gamut change on each picture, and then randomly arrange and splice them into one picture according to the positions of top left, top right, bottom left, and bottom right, crop the out-of-bounds part of the picture and transform the corresponding labels;

[0014] Step 2: Duplicate I once to obtain two original training sets;

[0015] Step 3: Perform background zeroing with selective local background retention on the images of one of the I's to obtain I0;

[0016] Step 4: Crop out all small targets from I0 through the cropping of selective local background retention and save them in S;

[0017] Step 5: Randomly select small targets from S and paste them into the image of I0 through the grid target pasting method, and create corresponding labels to obtain

[0018] Step 6: Merge I with to obtain

[0019] Step 7: Perform Mosaic data augmentation on to obtain

[0020] 3. Innovation points:

[0021] Compared with the Mosaic data augmentation method, the first six steps of the present invention are all different;

[0022] Generally speaking, the present patent introduces the following methods and ideas:

[0023] (1) Aiming at the interference problem of a large number of backgrounds on the recognition of small targets, a background zeroing method of selective local background retention is adopted to zero the pixel values of a large number of invalid backgrounds;

[0024] (2) Aiming at the loss of effective local background information when cropping small targets that rely on local backgrounds, a cropping method of selective local background retention is adopted to crop the effective local background and small targets together;

[0025] (3) Aiming at the problem of unnatural overlap of small targets during the pasting process, the grid target pasting method is adopted to divide the image of the small target to be pasted into grid regions, and it is judged whether a small target should be pasted in the grid by the sum of pixel values in each grid region.

[0026] 4. Beneficial effects:

[0027] The present invention discloses a background zeroing Mosaic data augmentation method for small target datasets. By selectively retaining the local background for background zeroing, the interference of the background on small targets is effectively reduced. By selectively retaining the local background for cropping, the effective local background information of small targets that rely on local backgrounds is retained. By the grid target pasting method, the problem of unnatural overlap of small targets during the pasting process is effectively solved. This method improves the detection accuracy of the model while enhancing the generalization ability of the model. Description of the drawings

[0028] Figure 1Flowchart of the background zeroing Mosaic data augmentation method for small target datasets. Input the original training set, make a copy of the original training set to obtain two copies of the original training set; perform the background zeroing operation with selective local background retention on one of the original training sets to obtain the background zeroing training set; through the cropping with selective local background retention, crop out all small targets from the background zeroing training set to obtain a set of small targets of all types; through the grid target pasting method, randomly select small targets from the small target set and paste them into the images of the background zeroing training set, and create corresponding labels to obtain the target pasting training set; merge the target pasting training set with the original training set to obtain the background zeroing enhanced training set; perform Mosaic data augmentation on the background zeroing enhanced training set to obtain the background zeroing Mosaic data augmentation training set;

[0029] Figure 2 It is a comparison diagram before and after the background zeroing operation with selective local background retention for an image. The upper part of the figure is the image before the operation and the drawn target labels, and the lower part is the image after the operation and the drawn target labels. It can be seen that the pixel values of most of the invalid backgrounds in the image after the operation become zero, and the local backgrounds of small targets that rely on the local background are retained;

[0030] Figure 3 It is a comparison diagram before and after the cropping operation with selective local background retention for an image. The upper part of the figure is the background zeroing image to be cropped, and the lower part is the cropped small targets. It can be seen that the cropped small targets include small targets that do not rely on the local background and small targets that rely on the local background, and among them, the small targets that rely on the local background retain the local background of the non-target area;

[0031] Figure 4 It is a comparison diagram before and after the grid target pasting method operation for an image. The upper part of the figure is the background zeroing image and the small target image to be pasted, and the lower part is the image after the target pasting. It can be seen that there are no other targets within a certain range around the target after the target pasting, and there are no targets overlapping with it. Detailed implementation method

[0032] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The embodiments use a self-made dataset of the cigarette packing workshop, and the dataset label format is the YOLO dataset format. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0033] See Figure 1, the present invention provides a background-zeroing Mosaic data augmentation method for small target datasets, and the specific implementation is as follows:

[0034] Step 1: Since the images in the self-made wrapping workshop dataset are all large images of 1920×1080 pixels, targets with a target width less than 0.1 times the image width or a target height less than 0.1 times the image height in the image are defined as small targets. Thus, the small targets in the image include ten categories: bench, material box, unknown collar, closed collar, open collar, unknown mask, worn mask, unworn mask, long hair, and mobile phone. Among them, the bench, material box, unknown mask, worn mask, unworn mask, long hair, and mobile phone are classified as class_NO, and the unknown collar, closed collar, and open collar are classified as class_YES. Let I be the original training set of the self-made wrapping workshop, I0 be the background-zeroing training set obtained by background zeroing of I through selective local background retention, S be the small target set obtained by cropping I0 through selective local background retention, where S contains small target sets of all categories. is the target pasting training set obtained by the grid target pasting method from I0 and S. is from I and is the background-zeroing enhanced training set obtained by merging. is from is the background-zeroing Mosaic data augmentation training set obtained by Mosaic data augmentation.

[0035] Step 2: Copy I once to obtain two original training sets.

[0036] Step 3: Perform background zeroing with selective local background retention on the images of one of the I's, see Figure 2 : Calculate the center point (cx, xy) and width-height wh of the corresponding small target bounding box according to each data label class cx_normal cy_normal w_normal h_normal of each image in I and the width-height w_image h_image of the image. The calculation formula is: Double the width-height of the bounding boxes of small targets with class belonging to class_YES. According to the center point and width-height of each small target bounding box in each image, the corresponding small target bounding box area can be obtained. In each image, all areas outside the small target bounding box areas are regarded as the background of the image, and the pixel values of the background are set to zero. Each new image and the corresponding label form a new dataset.

[0037] Step 4: Crop the small targets in I0 with selective local background retention, see Figure 3: Calculate the center point (cx, xy) and width-height (w, h) of the corresponding small target bounding box based on each data label class cx_normal, cy_normal, w_normal, h_normal of each image in I0 and the width-height w_image, h_image of the image. The calculation formula is as follows: Double the width-height of the bounding boxes of the small targets whose class belongs to class_YES. Based on the center point, width-height of each small target bounding box in each image, the corresponding small target bounding box area can be obtained. Lock the area of each small target bounding box obtained from each image. If the class belongs to class_YES, set the pixel values of all other bounding box areas within the small target bounding box area except the original bounding box area of the small target to zero, then crop the locked area of each small target bounding box and put it into the small target set of the corresponding category in S, thus obtaining all the cropped small targets.

[0038] Step Five: Through the grid target pasting method, refer to Figure 4 , randomly select small target images from S and paste them into the images in I0 and create corresponding labels: Divide each image in I0 into a grid of M×N, the size of a single standard grid is m_grid×n_grid, and judge whether the sum of pixel values in each standard grid is zero. If it is zero, randomly select a small target image of a certain category from the small target set and paste it into the grid, making the center point of the small target image land at a random position within a certain range of the grid. The random position within a certain range of the grid refers to the position (m_c_grid, n_c_grid) based on the grid center. And define the width-height of the randomly selected small target bounding box as the width-height of the small target image, and the center point as the generated random landing position. Shrink the width-height of the bounding boxes of the small targets whose small target category belongs to class_YES among the randomly selected small targets by half. Calculate the label class cx_normal, cy_normal, w_normal, h_normal of the small target based on the randomly selected small target category class, the center point (cx, xy) of the randomly selected small target bounding box, the width-height w, h, and the width-height w_image, h_image of the image. The calculation formula is as follows: Thus, all the small target labels pasted in the corresponding image are obtained, and then a new dataset is obtained.

[0039] Step Six: Merge I with . The merge means a simple superposition of the two datasets, thus realizing the expansion of the dataset and obtaining a new dataset.

[0040] Step 7: Perform Mosaic data augmentation: Randomly select four pictures, perform random flipping, scaling, and color gamut change operations on each picture, then randomly arrange and splice them into one picture in the positions of upper left, upper right, lower left, and lower right. Determine the boundaries of the spliced picture and crop the out-of-bounds part of the picture. At the same time, transform the corresponding labels to obtain

Claims

1. A background-zeroing Mosaic data augmentation method for small target datasets, characterized in that It includes the following steps: Step 1: Determine all types of small targets in the small target dataset according to the definition of small targets. Divide all types of small targets into small targets class_NO that do not depend on local background and small targets class_YES that depend on local background. Let I be the original training set, I0 be the background-zeroed training set obtained by setting the background to zero through selective local background retention from I, and S be the small target set obtained by cropping from I0 through selective local background retention. Among them, S contains small target sets of all types. is the target-pasted training set obtained by the grid target pasting method from I0 and S. is obtained from I and is the background-zeroed enhanced training set obtained by merging. is obtained from is the background-zeroed Mosaic data enhanced training set obtained by Mosaic data augmentation. Among them, setting the background to zero with selective local background retention means: according to each data label, the type of the corresponding small target, the center point, width, and height of the target box in the corresponding image are obtained. The width and height of the target box of the small target whose small target type belongs to class_YES are magnified by n times. According to the center point, width, and height of each small target box in each image, the corresponding small target box area can be obtained. In each image, all areas outside the small target box areas are regarded as the background of the image, and the pixel values of the background are set to zero. Each new image and the corresponding label obtained form a new data set; Among them, cropping with selective local background retention means: according to each data label, the type of the corresponding small target, the center point, width, and height of the target box in the corresponding image are obtained. The width and height of the target box of the small target whose small target type belongs to class_YES are magnified by n times. According to the center point, width, and height of each small target box in each image, the corresponding small target box area can be obtained. Each small target box area obtained in each image is locked. If the small target type belongs to class_YES, the pixel values of all other target box areas within the small target box area that are not the original target box area of the small target are set to zero, and then each locked small target box area is cropped and placed in the small target set of the corresponding type in S, so as to obtain all the cropped small targets; Among them, the grid target pasting method means: dividing the background-zeroed image to be enhanced into a grid of M×N, the size of a single standard grid is m_grid×n_grid, and judging whether the sum of the pixel values in each standard grid is zero. If it is zero, a small target image of a certain type is randomly selected from the small target set and pasted in the grid, so that the center point of the small target image lands at a random position within the grid, and it is defined that the width and height of the randomly selected small target box are the width and height of the small target image, and the center point is the generated random landing position. The width and height of the target box of the small target whose small target type belongs to class_YES among the randomly selected small targets are reduced by n times. According to the randomly selected small target type and the center point, width, and height information of the randomly selected small target box, the corresponding small target label is generated, so as to obtain all the small target labels pasted in the corresponding image; Among them, merging means: a simple superposition of two data sets to achieve the expansion of the data set; Among them, Mosaic data augmentation means: randomly selecting four pictures, performing random flipping, scaling, and color gamut change operations on each picture, and then randomly arranging and splicing them into a picture in the upper left, upper right, lower left, and lower right positions, cropping the out-of-bounds part of the picture and transforming the corresponding label; Step 2: Copy I once to obtain two original training sets; Step 3: Perform background zeroing with selective local background retention on the images of one of the I's to obtain I0; Step 4: Crop all small targets from I0 through selective local background retention and save them in S; Step 5: Randomly select small targets from S and paste them into the image of I0 by the grid target pasting method, and create corresponding labels to obtain Step 6: Combine I with to obtain Step 7: For perform Mosaic data augmentation to obtain

Citation Information

Patent Citations

  • Small target detection algorithm based on improved YOLOv5

    CN114241548A

  • Targeted data augmentation using neural style transfer

    US20180373999A1