A practical data enhancement method, device, equipment and medium
Through the adjustment and filling technology of binning and splicing pictures, the problem of important areas being masked in small target detection is solved, more effective data enhancement is achieved, the probability of false detection and missed detection is reduced, and the output of any size is supported.
Patent Information
- Application Number
- CN202411543585.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing data augmentation methods can easily lead to the masking of important areas in small-object detection, resulting in the network being unable to learn important features that distinguish the target from the background, which leads to false detection and missed detection.
By obtaining the image size in the dataset, the sub-dataset is obtained by binning, and the stitching pictures are obtained for adjustment and filling, and after synthesising the preprocessed pictures, it is deleted to obtain a new sample picture.
Ensure that important areas are not masked and support arbitrary size output, avoiding the problems of missed detection and missed detection in small object detection.
Smart Images

Figure CN119418153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a practical data enhancement method, a practical data enhancement device, a computer device and a computer-readable storage medium. Background Art
[0002] At present, the existing data expansion methods include: (1) The more commonly used geometric transformation methods are: flipping, rotation, cropping, scaling, translation, and jittering. It is worth noting that in some specific tasks, when using these methods, the main label data needs to be transformed. For example, if flipping is used in target detection, the GT box needs to be adjusted accordingly. (2) The more commonly used pixel transformation methods are: adding salt and pepper noise, Gaussian noise, Gaussian blur, adjusting HSV contrast, adjusting brightness, saturation, histogram equalization, adjusting white balance, etc. (3) Several data enhancement methods suitable for detection tasks: cutout, random erasing, grid masking, and mosaic.
[0003] Among them, cutout: randomly cut out part of the area in the sample and fill it with 0 pixel value; RandomErasing: randomly erase rectangles with random proportions and sizes, and fill them with randomly generated values or pixel means on ImageNet.
[0004] GridMas: Generates a mask with the same resolution as the original image, and then multiplies the mask with the original image to get an image. Figure 1 The value of the gray area is 1, and the value of the black area is 0. This achieves information dropping in a specific area, which can essentially be understood as a regularization method. GridMask corresponds to 4 parameters, [x, y, r, d], and the settings of the four parameters are as follows Figure 2 As shown. Figure 2 As can be seen, [r] represents the proportion of the original image information retained. [d] determines the size of a dropped square, and the values of parameters [x] and [y] are somewhat random.
[0005] Mosaic data enhancement uses four pictures to stitch the four pictures together. Each picture has its corresponding GT frame. After stitching the four pictures together, a new picture is obtained, and the boundary (GT, Ground Truth) frame corresponding to this picture is also obtained. The specific steps are as follows: First, randomly select the coordinates (x, y) of the picture stitching reference point, and randomly select four pictures; then adjust the size and scale the four pictures according to the reference points, and place them in the upper left, upper right, lower left, and lower right positions of the specified large picture; according to the size transformation method of each picture, the mapping relationship is mapped to the picture label; finally, according to the specified square horizontal and vertical coordinates, the large picture is stitched together, and the detection frame coordinates that exceed the boundary are processed.
[0006] The cutout method and Random Erasing methods both randomly select the selected areas, which easily results in the situation where all important parts are covered; the GridMask method will at most partially cover important areas, and it will almost certainly cover part of the area, using arranged square areas for masking. In the detection of small targets, the above three methods will all have important areas that can distinguish the target partially or completely covered. For small targets, since the size of the target itself is too small (the target area is less than 32×32 pixels), if the important feature area is covered, the network will not be able to learn the important features that can distinguish the target from the background, which will lead to a certain probability of false detection and missed detection.
[0007] In addition, the mosaic method only supports the expanded images to be square in size, that is, the width and height are consistent; and in the expanded images, some targets will appear to be only a small part of the original target. For small targets, the size of the target itself is too small. If the target in the expanded image is only a part of the original small target, this will mislead the network model to a certain extent, resulting in incomplete learned features, and finally there will be a certain probability of false detection and missed detection, affecting the training effect of the model.
[0008] Therefore, it is necessary to provide a practical data enhancement method, a practical data enhancement device, a computer device and a computer-readable storage medium to effectively solve the above problems. Summary of the invention
[0009] The present invention provides a method for practical data enhancement, a device for practical data enhancement, a computer device and a computer-readable storage medium.
[0010] The embodiment of the present invention provides a method for practical data enhancement, including:
[0011] Get the size of each picture in the dataset that needs to be expanded;
[0012] Based on the size of the image, bin the data set to obtain N sub-data sets;
[0013] Obtaining a spliced picture in the mth sub-data set; the number of the spliced pictures is greater than 1 and less than or equal to the number of pictures in the data set;
[0014] Obtaining a target image size and a center point, and adjusting and filling the stitched image based on the target image size and the center point to synthesize into a pre-processed image;
[0015] The preprocessed image is deleted based on the filled area in the preprocessed image to obtain a new sample image.
[0016] Preferably, the data set is binned based on the size of the picture to obtain N sub-data sets, including:
[0017] sorting the images based on aspect ratios of the images;
[0018] Based on the aspect ratio of the image, obtaining N adjacent aspect ratio data intervals;
[0019] Based on the aspect ratio of each of the pictures and the data interval, N sub-data sets are obtained.
[0020] Preferably, obtaining the target image size includes:
[0021] Determine whether the aspect ratio of the stitched image is greater than or equal to 1;
[0022] If yes, the target image width is set to a preset multiple of the first standard width and the target image height is set to a preset multiple of the first standard height;
[0023] If not, setting the target image width to a preset multiple of the second standard width and the target image height to a preset multiple of the second standard height;
[0024] The first standard width is greater than or equal to the first standard height, and the second standard width is less than the second standard height.
[0025] Preferably, the number of stitched pictures is equal to 4.
[0026] Preferably, the step of adjusting and filling the stitched images based on the target image size and the center point to synthesize them into a pre-processed image includes:
[0027] Obtaining an adjustment coefficient based on the size of the stitched image and the size of the target image;
[0028] Adjust the size of the stitched image based on the adjustment coefficient;
[0029] Get the padding base;
[0030] Obtaining a first filling size based on the size of the stitched image, the target image size, and the filling base;
[0031] Filling the stitched image based on the first filling size;
[0032] Stitching the filled stitched pictures into a first pre-processed picture;
[0033] The first preprocessed picture is padded based on the target picture size to obtain a preprocessed picture.
[0034] Preferably, in the step of filling the spliced image based on the first filling size,
[0035] If the stitched image is located at the upper left, the filling direction of the stitched image includes the upper and left sides thereof;
[0036] If the stitched image is located at the upper right corner, the filling direction of the stitched image includes the upper side and the right side thereof;
[0037] If the stitched image is located at the lower left, the filling direction of the stitched image includes the lower part and the left part;
[0038] If the stitched image is located at the lower right, the filling direction of the stitched image includes the lower side and the right side thereof.
[0039] Preferably, the deleting the pre-processed image based on the filling area to obtain a new sample image includes:
[0040] Based on the filling area boundary of the spliced image, the pre-processed image is deleted to obtain the new sample image.
[0041] A device for practical data enhancement is also provided, comprising:
[0042] The first acquisition module is used to obtain the size of each picture in the data set to be expanded;
[0043] A binning module, used to bin the data set based on the size of the image to obtain N sub-data sets;
[0044] A second acquisition module is configured to acquire a spliced picture in the mth sub-data set; the number of the spliced pictures is greater than 1 and less than or equal to the number of pictures in the data set;
[0045] A preprocessing module, used for obtaining a target image size and a center point, and adjusting and filling the stitched image based on the target image size and the center point to synthesize into a preprocessed image;
[0046] The deletion module is used to delete the pre-processed image based on the filled area in the pre-processed image to obtain a new sample image.
[0047] Furthermore, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above-mentioned methods when executing the computer program.
[0048] Furthermore, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0049] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:
[0050] The practical data enhancement method, the practical data enhancement device, the computer equipment and the computer-readable storage medium provided by the embodiments of the present invention are aimed at the cutout, Random Erasing and GridMask methods, which partially or completely cover the important areas that can distinguish the target and the background. The present invention will overcome this shortcoming so that the important areas will not be covered. Moreover, for the mosaic method that only supports square size and the expanded target is a part of the original small target, the method of the present invention can be set to output at any size (that is, it not only supports square size, but also supports rectangular size). In addition, the target in the picture before expansion can still appear completely in the expanded picture. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are some embodiments of the present invention, not all embodiments. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor.
[0052] Figure 1 An operation diagram of a GridMask method provided by one embodiment of the present invention;
[0053] Figure 2 A schematic diagram of a GridMask provided for one embodiment of the present invention;
[0054] Figure 3A flowchart of a practical data enhancement method provided by an embodiment of the present invention;
[0055] Figure 4 An operation diagram of a practical data enhancement method provided by an embodiment of the present invention;
[0056] Figure 5a1-5c A schematic diagram of a picture processing process of a practical data enhancement method provided in some embodiments of the present application;
[0057] Figure 6a-6f A schematic diagram of a picture processing process of a practical data enhancement method provided for other embodiments of the present application;
[0058] Figure 6g-6h A schematic diagram of a picture processing process of a practical data enhancement method provided in some other embodiments of the present application;
[0059] Figure 7 A schematic diagram of a module of a practical data enhancement device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0062] Based on the problems existing in the prior art, the embodiments of the present invention provide a method for practical data enhancement, an apparatus for practical data enhancement, a computer device and a computer-readable storage medium.
[0063] Figure 3 A flowchart of a practical data enhancement method provided for one embodiment of the present invention. Figure 4 The following is a schematic diagram of the process operation of a practical data enhancement method provided by an embodiment of the present invention. Figure 3 and 4 As shown, an embodiment of the present invention provides a method for practical data enhancement, including:
[0064] Step S11, obtaining the size of each picture in the data set to be expanded.
[0065] In a specific implementation, this solution is a small target detection algorithm, which is aimed at data with an area less than 32×32 pixels. The data in the data set is image data. The size of each image can include the width, height, aspect ratio, etc. of the image.
[0066] Step S12: bin the data set based on the size of the image to obtain N sub-data sets.
[0067] In some examples, step S12 includes:
[0068] Step S121, sorting the pictures based on the aspect ratio.
[0069] Step S122, based on the aspect ratio, obtain N adjacent aspect ratio data intervals.
[0070] Step S123, obtaining N sub-data sets based on the aspect ratio and data interval of each picture.
[0071] In a specific implementation, the aspect ratios of all images are sorted, and each image in the data set is binned based on the corresponding aspect ratio, that is, the aspect ratio data is divided into multiple adjacent value intervals, the number of bins is N, and each interval is [R i , R i+1 ], i=1, 2, ... N, and then the aspect ratio is between [R i , R i+1 ] are placed in the same folder, so that N sub-datasets can be obtained.
[0072] Step S13, obtaining the spliced pictures in the mth sub-data set; the number of spliced pictures is greater than 1 and less than or equal to the number of pictures in the data set.
[0073] In some examples, the number of stitched pictures is equal to 4. Four pictures in the mth sub-dataset are randomly obtained as stitched pictures.
[0074] In a specific implementation, the following operations are performed on each of the N sub-datasets: the number of pictures in the mth sub-dataset is L m , m takes values of (1, 2, ..., N). For i from 1 to L m , get the i-th picture from the m-th sub-dataset, then randomly select 3 different pictures from the m-th sub-dataset, and randomly shuffle the order of these 4 pictures.
[0075] Step S14, obtaining the target image size and center point, and adjusting and filling the stitched image based on the target image size and center point to synthesize into a pre-processed image.
[0076] In some examples, obtaining the target image size includes: determining whether the aspect ratio of the stitched image is greater than or equal to 1; if so, setting the width of the target image to a preset multiple of the first standard width and the height of the target image to a preset multiple of the first standard height. If not, setting the width of the target image to a preset multiple of the second standard width and the height of the target image to a preset multiple of the second standard height. The first standard width is greater than or equal to the first standard height, and the second standard width is less than the second standard height. The target image is a stitched image.
[0077] In a specific implementation, when the aspect ratio of the mth sub-data set is greater than or equal to 1, it is sufficient to set the data with the first standard width greater than the first standard height. The standard width and standard height (width1, height1) of the image can be set to (320, 240), the preset multiple is 1, and the width and height (width, height) of the target image are (320, 240). When the aspect ratio of the mth sub-data set is less than 1, the standard width and standard height (width1, height1) of the image can be set to (240, 320), the preset multiple is 1, and the width and height (width, height) of the target image are (240, 320).
[0078] For example, when the number of stitched images is 4, a new sample image (img) with a shape size of (height × 2, width × 2, 3) and a value of 114 is generated. Specifically, use the opencv library to generate an RGB image (that is, a 3-channel image) with a width and height that are twice the standard width and height. The pixel value is 114 or any other value from 0 to 255. It is generally recommended to set it to 114 (grayscale value). This image may be recorded as img (here is an explanation of the shape of img. When the standard width and height are 320 and 240 respectively, the preset multiple is 1. When the width and height of the target image are 320 and 240 respectively, the width and height of img are 640 and 480 respectively, and it is a 3-channel image).
[0079] Select the center point (x center ,y center ). As shown in formulas (1) and (2).
[0080] ,
[0081] Among them, random.uniform(x, y) is a Python syntax, which means randomly generating a real number in the range [x, y]; int(x) is a Python syntax, which means rounding the real number x. The purpose of selecting the center point is to select the intersection point of the four images. The coordinates of the intersection point are generated by the above formula, and the randomness is increased in turn to further enrich the image. Width and height are the width and height of the target image respectively.
[0082] In some examples, in step S14, the stitched images are adjusted and filled based on the target image size and the center point to be synthesized into a pre-processed image, including:
[0083] Step S141, obtaining an adjustment coefficient based on the size of the stitched image and the target image size.
[0084] Step S142, adjusting the size of the stitched image based on the adjustment coefficient.
[0085] Step S143, obtaining the filling base.
[0086] Step S144, obtaining a first filling size based on the size of the stitched image, the target image size and the filling base.
[0087] Step S145, filling the stitched image based on the first filling size.
[0088] Step S146, stitching the filled stitched pictures into a first pre-processed picture.
[0089] Step S147, padding the first preprocessed image based on the target image size to obtain a preprocessed image.
[0090] In a specific implementation, the adjustment coefficient is calculated as shown in formula (3).
[0091] ,
[0092] Among them, ratio is the adjustment coefficient (resize), width origin and height origin The actual width and height of the stitched image. width and height are the width and height of the target image.
[0093] If the width and height of the target image are 240 and 320 respectively, calculate 240 divided by height origin The value and 320 divided by width origin, and then select the smallest of the two values as the adjustment factor (resize), and then adjust the original spliced image to a width and height of int (width origin × ratio) and int(height origin × ratio). Here, the same resize is used for both width and height to keep the aspect ratio of the image unchanged before and after resizing, that is, to keep the objects in the original image from being deformed to the greatest extent possible.
[0094] The size of the stitched image is determined based on the adjustment coefficient. See formulas (4) and (5) for details.
[0095] ,
[0096] Among them, resize imgwidth To adjust the width of the stitched image, resize imgheight width is the height of the stitched image after resizing. origin and height origin are the actual width and height of the stitched image respectively. ratio is the adjustment coefficient. × represents multiplication operation, and int is the rounding function.
[0097] Then, the padding cardinality is obtained, for example, the padding cardinality is 8. Then, based on the size of the stitched image, the size of the target image, and the padding cardinality, the first padding size is obtained. For details, see formulas (6) and (7).
[0098] ,
[0099] The first filling size includes the first filling width pad w and the first fill height pad h . numpy.mod(x, y) is a function in python that represents the remainder after x is divided by y. width origin and height origin The actual width and height of the stitched image respectively. imgwidth To adjust the width of the stitched image, resize imgheight The height of the stitched image after resizing.
[0100] Then, each stitched image is filled based on the first filling size. When the stitched image is filled, if the stitched image is located in the upper left, the filling direction of the stitched image includes the upper and left sides thereof; if the stitched image is located in the upper right, the filling direction of the stitched image includes the upper and right sides thereof; if the stitched image is located in the lower left, the filling direction of the stitched image includes the lower and left sides thereof; if the stitched image is located in the lower right, the filling direction of the stitched image includes the lower and right sides thereof.
[0101] Then, the four padded stitched images are stitched into a first pre-processed image based on the center point. After comparing the overall size of the first pre-processed image with the image size (2×width, 2×height in this embodiment), the first pre-processed image is padded to obtain a pre-processed image.
[0102] Step S15, deleting the preprocessed image based on the filled area in the preprocessed image to obtain a new sample image.
[0103] In some examples, step S15 includes deleting the pre-processed image based on the filling area boundary of the stitched image to obtain a new sample image.
[0104] In some embodiments, the width width of the target image is set to 320, the height height of the target image is set to 240, and the padding base stride is set to 8.
[0105] like Figure 5a1-5a4 (Synthesis of 4 pictures) shows that the first spliced picture (upper left): the original width is 480, the original height is 384 (the aspect ratio is 1.25); according to formula (3), the specific value ratio = 0.625 is substituted, and the width of the spliced picture adjusted based on the adjustment coefficient is 300, and the height is 240. In the first padding size, the padding width pad_w = 4 and the padding height pad_h = 0, indicating that 4 rows need to be filled in the width direction (that is, the left and right directions are filled. In order to make the 4 pictures seamlessly connected, the present invention only fills 4 rows on the left), and no filling is required in the height direction. In this way, the width and height of the picture after adjustment and padding are 304 and 240 respectively.
[0106] Similarly, the second stitched picture can be obtained, the original width is 492, the original height is 384 (the aspect ratio is 1.28125), and the calculated ratio = 0.625, pad_w = 4 and pad_h = 0, indicating that the adjusted stitched picture (width and height are 308 and 240 respectively) needs to be padded with 4 rows in the width direction (that is, padded in the left and right directions. In order to make the 4 pictures seamlessly connected, the present invention only fills 4 rows on the right), and does not need to be padded in the height direction. In this way, the width and height of the picture after adjustment and padding are 312 and 240 respectively.
[0107] Similarly, the third stitched image (lower left) has an original width of 512 and an original height of 384 (aspect ratio of 1.333333). After calculation, ratio = 0.625, pad_w = 0 and pad_h = 0, indicating that the adjusted stitched image (width and height are 320 and 240 respectively) does not need to be padded in the width and height directions. In this way, the width and height of the image after adjustment and padding are 320 and 240 respectively.
[0108] Similarly, the fourth stitched image (lower right) has an original width of 500 and an original height of 384 (aspect ratio of 1.3020833). After calculation, ratio = 0.625, pad_w = 0 and pad_h = 0, indicating that the adjusted stitched image (width and height are 312 and 240 respectively) does not need to be padded in the width and height directions. In this way, the width and height of the image after adjustment and padding are 312 and 240 respectively.
[0109] When the center point (x center ,y center ) is (320, 230), the stitched image of the four images after adjustment and filling is as follows Figure 5b shown.
[0110] Remove all padding areas and get the final image. Figure 5c (width and height are 606 and 468 respectively).
[0111] In some other embodiments, the width of the target image is set to 480, the height of the target image is set to 320, and the padding base stride is set to 8.
[0112] like Figure 6a As shown in the figure, the original width of the first stitched image (upper left) is 480, and the original height is 384 (the aspect ratio is 1.25). According to formula (3), the specific value ratio = 0.8333334 is substituted, and the width of the stitched image after adjustment based on the adjustment coefficient is 400, and the height is 320. In the first padding size, the padding width pad_w = 0 and the padding height pad_h = 0, indicating that no padding is required in the width direction and the height direction. In this way, the width and height of the image after adjustment and padding are 400 and 320 respectively.
[0113] like Figure 6bAs shown, the second stitched picture can be obtained in the same way, the original width is 492, the original height is 384 (the aspect ratio is 1.28125), and the calculated ratio = 0.8333334, pad_w = 6 and pad_h = 0, indicating that the adjusted stitched picture (width and height are 410 and 320 respectively) needs to be padded with 6 rows in the width direction (that is, padded in the left and right directions. In order to make the 4 pictures seamlessly connected, the present invention only fills 6 rows on the right), and no padding is required in the height direction. In this way, the width and height of the picture after adjustment and padding are 416 and 320 respectively.
[0114] like Figure 6c As shown, the third stitched image (lower left) can be obtained by the same logic. The original width is 512, the original height is 384 (the aspect ratio is 1.333333), and the calculated ratio is 0.8333334, pad_w = 5 and pad_h = 0, which means that the adjusted stitched image (width and height are 427 and 320 respectively) is padded with 5 rows in the width direction and does not need to be padded in the height direction. In this way, the width and height of the image after adjustment and padding are 432 and 320 respectively.
[0115] like Figure 6d As shown, the fourth stitched image (lower right) can be obtained by the same logic. The original width is 500, the original height is 384 (the aspect ratio is 1.3020833), and the calculated ratio is 0.8333334, pad_w = 7 and pad_h = 0, which means that the adjusted stitched image (width and height are 417 and 320 respectively) is padded with 7 rows in the width direction and does not need to be padded in the height direction. In this way, the width and height of the adjusted and padded image are 424 and 320 respectively.
[0116] When the center point (x center ,y center ) is (472, 316), the stitched image of the four images after adjustment and filling is as follows Figure 6e shown.
[0117] Remove all padding areas and get the final image. Figure 6f (width and height are 808 and 634 respectively).
[0118] In addition, when the center point (x center ,y center ) is (477, 322), the four images after adjustment and filling are combined into a single image (width and height are 960 and 640 respectively) as shown below Figure 6g shown.
[0119] Remove all padding areas and get the final image. Figure 6h(width and height are 808 and 636 respectively).
[0120] The present application also provides a practical data enhancement device, such as Figure 7 As shown, it includes: a first acquisition module 71 is used to obtain the size of each picture in the data set that needs to be expanded. A binning module 72 is used to bin the data set based on the size of the picture to obtain N sub-data sets. A second acquisition module 73 obtains the spliced picture in the mth sub-data set; the number of spliced pictures is greater than 1 and less than or equal to the number of pictures in the data set. The preprocessing module 74 is used to obtain the target picture size and center point, and adjust and fill the spliced picture based on the target picture size and center point to synthesize it into a preprocessed picture. The deletion module 75 is used to delete the preprocessed picture based on the filled area in the preprocessed picture to obtain a new sample picture.
[0121] The device for practical data enhancement provided in the present application can implement any embodiment of the practical data enhancement method described above.
[0122] Furthermore, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the task scheduling method in any of the above embodiments are implemented.
[0123] Furthermore, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the task scheduling method in any of the above embodiments are implemented.
[0124] The practical data enhancement method, the practical data enhancement device, the computer equipment and the computer-readable storage medium provided by the embodiments of the present invention are aimed at the cutout, Random Erasing and GridMask methods, which partially or completely cover the important areas that can distinguish the target and the background. The present invention will overcome this shortcoming so that the important areas will not be covered. Moreover, for the mosaic method that only supports square size and the expanded target is a part of the original small target, the method of the present invention can be set to output at any size (that is, it not only supports square size, but also supports rectangular size). In addition, the target in the picture before expansion can still appear completely in the expanded picture.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A practical data enhancement method, characterized in that: include: Get the size of each picture in the dataset that needs to be expanded; Based on the size of the image, bin the data set to obtain N sub-data sets; Obtaining the spliced image in the mth sub-dataset; The number of stitched pictures is greater than 1 and less than or equal to the number of pictures in the data set; The number of stitched pictures is equal to 4; Obtaining a target image size and a center point, and adjusting and filling the stitched image based on the target image size and the center point to synthesize into a pre-processed image; Deleting the preprocessed image based on the filled area in the preprocessed image to obtain a new sample image; The step of adjusting and filling the stitched image based on the target image size and the center point to synthesize the stitched image into a pre-processed image includes: Obtaining an adjustment coefficient based on the size of the stitched image and the size of the target image; Adjust the size of the stitched image based on the adjustment coefficient; Get the padding base; Obtaining a first filling size based on the size of the stitched image, the target image size, and the filling base; Filling the stitched image based on the first filling size; Stitching the filled stitched pictures into a first pre-processed picture; The first preprocessed picture is padded based on the target picture size to obtain a preprocessed picture.
2. The method for practical data enhancement according to claim 1, characterized in that: The data set is binned based on the size of the image to obtain N sub-data sets, including: sorting the images based on aspect ratios of the images; Based on the aspect ratio of the image, obtaining N adjacent aspect ratio data intervals; Based on the aspect ratio of each of the pictures and the data interval, N sub-data sets are obtained.
3. The method for practical data enhancement according to claim 1, characterized in that: The step of obtaining the target image size includes: Determine whether the aspect ratio of the stitched image is greater than or equal to 1; If yes, the target image width is set to a preset multiple of the first standard width and the target image height is set to a preset multiple of the first standard height; If not, setting the target image width to a preset multiple of the second standard width and the target image height to a preset multiple of the second standard height; The first standard width is greater than or equal to the first standard height, and the second standard width is less than the second standard height.
4. The method for practical data enhancement according to claim 1, characterized in that: In the step of filling the spliced image based on the first filling size, If the stitched image is located at the upper left, the filling direction of the stitched image includes the upper and left sides thereof; If the stitched image is located at the upper right corner, the filling direction of the stitched image includes the upper side and the right side thereof; If the stitched image is located at the lower left, the filling direction of the stitched image includes the lower part and the left part; If the stitched image is located at the lower right, the filling direction of the stitched image includes the lower side and the right side thereof.
5. The method for practical data enhancement according to claim 1, characterized in that: The deleting the pre-processed image based on the filling area to obtain a new sample image includes: Based on the filling area boundary of the spliced image, the pre-processed image is deleted to obtain the new sample image.
6. A device for practical data enhancement, characterized in that: include: The first acquisition module is used to obtain the size of each picture in the data set to be expanded; A binning module, used to bin the data set based on the size of the image to obtain N sub-data sets; A second acquisition module is used to acquire a spliced image in the mth sub-dataset; The number of stitched pictures is greater than 1 and less than or equal to the number of pictures in the data set; The number of stitched pictures is equal to 4; A preprocessing module, used for obtaining a target image size and a center point, and adjusting and filling the stitched image based on the target image size and the center point to synthesize into a preprocessed image; A deletion module, configured to delete the preprocessed image based on the filled area in the preprocessed image to obtain a new sample image; The step of adjusting and filling the stitched image based on the target image size and the center point to synthesize the stitched image into a pre-processed image includes: Obtaining an adjustment coefficient based on the size of the stitched image and the size of the target image; Adjust the size of the stitched image based on the adjustment coefficient; Get the padding base; Obtaining a first filling size based on the size of the stitched image, the target image size, and the filling base; Filling the stitched image based on the first filling size; Stitching the filled stitched pictures into a first pre-processed picture; The first preprocessed picture is padded based on the target picture size to obtain a preprocessed picture.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image processing method and device, image processing model training method and device and image mode recognition method and device
CN113327195A
Data enhancement method, device and equipment for image-text cross-modal model and storage medium
CN115203375A