Detection target image generation method based on image restoration
Through the detection target image generation method based on image repair, high-quality expanded detection images are generated, which solves the problem of insufficient sample quality in the existing technology small and medium-sized sample detection tasks, and significantly improves the detection and recognition accuracy.
Patent Information
- Application Number
- CN202311576998.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
When facing the small sample problem of detection tasks, the image sample generation and expansion method leads to the background and foreground targets being treated equally, and the training data quality of the detected image cannot be effectively improved, thereby affecting the detection and recognition accuracy.
Using an image repair-based detection target image generation method, the original detection image is converted into a three-channel sandwich image, and the generator and discriminator generate an expanded detection image until the discriminator cannot distinguish the generated target image from the reference truth image, so as to achieve high-quality sample generation expansion.
By generating high-quality expanded detection images, the scale and quality of detection and identification data are improved, and the detection and identification accuracy is reflected in the mAP@0.5 accuracy improvement ranges on the test data set are 7.11%, 9.13% and 4.6%, respectively.
Smart Images

Figure CN120047317A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to a detection target image generation method based on image restoration. Background Art
[0002] At present, the small sample problem is one of the core problems that machine learning, especially deep learning technology, has not yet completely overcome. In essence, the small sample problem is caused by the mismatch between the small number of samples and the powerful fitting ability of deep neural networks. Therefore, its solutions are also divided into two categories: one is data-oriented, which improves the quantity and quality of data by generating and expanding, thereby converting the small sample recognition problem into a common recognition problem; the other is model-oriented algorithm, which improves the performance of the algorithm on small sample training data by reasonably designing and changing the feature representation structure. However, the small sample problem for detection tasks is more complicated than the small sample problem for classification problems. The integrated image sample generation and expansion treats background and foreground targets equally, which is not conducive to improving the quality of training data for detection images, and will ultimately affect the detection and recognition accuracy. Summary of the invention
[0003] The present invention provides a detection target image generation method based on image restoration, which can solve the problems in the prior art.
[0004] The present invention provides a method for generating a detection target image based on image restoration, wherein the method comprises:
[0005] Input the original detection image;
[0006] Generate an expanded detection image according to the original detection image;
[0007] Use the expanded detection images to learn and train the detection and recognition algorithm model;
[0008] The detection and recognition test is performed on the image to be detected according to the trained detection and recognition algorithm model.
[0009] Preferably, generating an expanded detection image according to the original detection image includes:
[0010] Convert the text annotations in the original detection image into image annotations of the same size;
[0011] Combine the detection image containing the image-shaped annotation and the original detection image into a three-channel sandwich image, and use the three-channel sandwich image as the input of the generator;
[0012] Generate a target image using the generator;
[0013] Use a discriminator to judge the authenticity of the generated target image and the reference true-value image until the discriminator can no longer distinguish the generated target image from the reference true-value image.
[0014] Preferably, each original detection image includes multiple targets to be detected, and each target to be detected includes position information and category information. The position information is a four-value array describing the target detection box, and the category information is an integer category code agreed in advance.
[0015] Preferably, converting the annotation in text form in the original detection image into an annotation in image form of the same size includes:
[0016] Generate a full-zero image matrix of the same size as the original detection image as the background encoding map;
[0017] For the position information, based on the four-value array, use a solid rectangle box for the detection box rectangle of the target to be detected in the background encoding map as the position encoding;
[0018] For the category information, select n gray levels from 0 to 255 to represent n target categories to be detected, where n represents the number of target types to be detected in the entire dataset;
[0019] Assign the solid rectangle box of the target to be detected the gray level corresponding to the category.
[0020] Preferably, on the basis of an image in the form of a three-channel sandwich as the input of the generator, add Gaussian noise of a predetermined intensity as the input of the generator.
[0021] Preferably, the generator is a generator based on the U-net framework.
[0022] Preferably, the generator includes a self-attention layer and a residual convolution module.
[0023] Preferably, in the process of generating the augmented detection image, it includes a generative adversarial loss function, a global point-to-point calculation function, and a bounding box loss function.
[0024] Through the above technical solutions, an augmented detection image can be generated according to the input original detection image, realizing high-quality sample generation and augmentation for detection images to improve the scale and quality of detection and recognition data; then, the detection and recognition algorithm model can be learned and trained according to the augmented detection image, and the detection and recognition test of the image to be detected can be carried out according to the trained model, thereby improving the detection and recognition accuracy. Brief Description of the Drawings
[0025] The accompanying drawings included are used to provide a further understanding of the embodiments of the present invention, which form a part of the specification, illustrate the embodiments of the present invention, and, together with the written description, explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0026] Figure 1 The flowchart of a method for generating a detection target image based on image inpainting according to an embodiment of the present invention is shown;
[0027] Figure 2 The schematic diagram of the conversion coding from text to an equal-sized image of the annotation format according to an embodiment of the present invention is shown;
[0028] Figure 3 The schematic diagram of a specified target generation framework based on image inpainting according to an embodiment of the present invention is shown;
[0029] Figure 4 The schematic diagram of the calculation range of the bounding box loss of the generated target according to an embodiment of the present invention is shown;
[0030] Figure 5 The schematic diagram of the generated image effect of the target at a specified position and category according to an embodiment of the present invention is shown. Detailed implementation manners
[0031] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way restrictive of the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0032] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0033] Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0034] Figure 1 The flowchart of a method for generating a detection target image based on image inpainting according to an embodiment of the present invention is shown.
[0035] Among them, this method is applicable to the improvement requirements of target detection and recognition accuracy in various situations where only small-sample image data is available, including but not limited to target detection and recognition under imaging single sensors such as visible light, infrared, SAR, and multimodal image sensors formed by their mutual combination. In addition, the present invention is not limited to the application field of detection and recognition. Based on the principle of the present invention, it can be widely applied to pattern application fields such as semantic segmentation, instance segmentation, and target tracking.
[0036] As Figure 1 shown, an embodiment of the present invention provides a method for generating a detection target image based on image inpainting, where the method includes:
[0037] S100, input the original detection image;
[0038] S102, generate an augmented detection image according to the original detection image;
[0039] S104, use the augmented detection image to learn and train the detection and recognition algorithm model;
[0040] S106, perform detection and recognition tests on the image to be detected according to the trained detection and recognition algorithm model.
[0041] Through the above technical solution, an augmented detection image can be generated according to the input original detection image, realizing high-quality sample generation and augmentation for detection images to improve the scale and quality of detection and recognition data; then, the detection and recognition algorithm model can be learned and trained according to the augmented detection image, and the image to be detected can be detected and recognized according to the trained model, thereby improving the detection and recognition accuracy.
[0042] Among them, the model training process of S104 and the detection and recognition test process of S106 can adopt the existing methods in the prior art. To avoid confusing the present invention, they will not be elaborated herein.
[0043] According to an embodiment of the present invention, generating an augmented detection image from an original detection image includes:
[0044] Converting the annotation in text form in the original detection image into an annotation in image form of the same size, that is, annotation encoding;
[0045] Combining the detection image (mask) containing the annotation in image form with the original detection image to form an image in a three-channel sandwich form (background), and using the image in the three-channel sandwich form as the input of the generator;
[0046] Generating a target image using a generator (Generator);
[0047] Using a discriminator (Discriminator) to judge the authenticity of the generated target image and a reference ground truth image until the discriminator can no longer distinguish the generated target image from the reference ground truth image.
[0048] Among them, the present invention does not impose constraints on the fixed parameters of the discriminator, that is, any form of discriminator can be applied in the present invention.
[0049] According to an embodiment of the present invention, each original detection image includes multiple targets to be detected. Each target to be detected includes position information and category information. The position information is a four-value array describing the target detection frame (bounding box), and the category information is an integer category code agreed in advance.
[0050] Among them, the annotation encoding is to convert the above two pieces of information from text to image.
[0051] According to an embodiment of the present invention, converting the annotation in text form in the original detection image into an annotation in image form of the same size includes:
[0052] Generating a full-zero image matrix of the same size as the original detection image as the background encoding map;
[0053] For the position information, based on the four-value array, the detection frame rectangle of the target to be detected in the background encoding map is used as a solid rectangle frame as the position encoding;
[0054] For the category information, select n gray levels from 0 to 255 to represent n target categories to be detected, where n represents the number of target types to be detected in the entire dataset;
[0055] Assigning the solid rectangle frame of the target to be detected to the gray level corresponding to the category.
[0056] Thus, the detection annotation encoding in the form of an image can be completed, such as Figure 2 shown.
[0057] According to an embodiment of the present invention, as Figure 3 shown, on the basis of using an image in the form of a three-channel sandwich as the input of the generator, Gaussian noise with a predetermined intensity is added as the input of the generator.
[0058] Thus, by introducing Gaussian noise with a fixed intensity on the basis of the three-channel input, the randomness can be enhanced.
[0059] According to an embodiment of the present invention, the generator is a generator based on the U-net framework.
[0060] According to an embodiment of the present invention, the generator includes a self-attention layer and a residual convolution module.
[0061] Thus, the feature expression performance of the generator can be enhanced, and the generation of the target image from the target-encoded image with a specified position and a specified category can be realized through the generator.
[0062] According to an embodiment of the present invention, in the process of generating the augmented detection image, a generative adversarial loss function, a global point-to-point calculation function, and a bounding box loss function are included.
[0063] Among them, for the generative adversarial loss function, that is, the GAN loss, the form of wgan can be adopted to obtain better convergence ability; for the global point-to-point loss function, since there is a ground truth image in the target generation task, the MSE loss between the generated image and the ground truth image can be used to measure their similarity; for the bounding box loss function, since the generated content is only within the target detection box, most of the global point-to-point loss is the similarity of the background. In order to focus on the generation quality within the target detection box, a hollow frame ring can be marked around the detection box, and the loss calculation is performed within the range of this ring to measure the integration degree of the generated target and the background, as Figure 4 shown.
[0064] In the actual application and deployment stage after training is completed, only the generator needs to be retained to realize the generation of the target to be detected with a specified category in a patched manner at a specified position on the background image, thereby greatly expanding the sample quantity and quality of the detection image, as Figure 5 shown. In Figure 5 , the leftmost column is the original image, and the rightmost three columns are all corresponding augmented images (including the target points with specified positions and specified categories).
[0065] The present invention is used for high-quality detection image target generation and data augmentation, which can greatly alleviate the problem of low accuracy caused by a small number of training samples. For example, the present invention is based on a training dataset of 128 infrared images with 9 common types of targets, and a detection algorithm comparison experiment is carried out on 640 test images without intersection. After training on the original training dataset, based on three classic detection and recognition algorithms, Faster RCNN, SSD, and RentinaNet, the mAP@0.5 accuracy on the test dataset is 51.09%, 27.79%, and 25.38% respectively. After the detection targets are generated by the present invention and the training dataset is augmented, the mAP@0.5 accuracy of Faster RCNN, SSD, and RentinaNet on the test dataset is 58.2%, 36.92%, and 29.98% respectively, and the accuracy improvement ranges are 7.11%, 9.13%, and 4.6% respectively.
[0066] As can be seen from the above embodiments, compared with the prior art of directly copying and pasting all the contents within the detection frame of the target to be detected in the figure to a random position of the image to achieve target sample augmentation, the detection target image generation method based on image repair of the present invention realizes high-quality fusion of the target and the background, rather than direct texture mapping fusion. Therefore, it has better detection target generation quality, and can bring higher accuracy improvement for detection and recognition data augmentation.
[0067] Those skilled in the art should understand that although the above embodiments of the present invention take the multi-target detection and recognition of infrared images as examples, those skilled in the art can make corresponding changes and modifications to these examples according to the present invention. Therefore, the present invention covers other applications including the processing of imaging sensor signals such as visible light, infrared, SAR, etc., as well as semantic segmentation, instance segmentation, target tracking, etc. with similar principles.
[0068] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by orientation words such as "front, back, up, down, left, right", "horizontal, vertical, perpendicular, horizontal" and "top, bottom" is usually based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description. Without contrary explanation, these orientation words do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the protection scope of the present invention; the orientation words "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0069] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "upper" etc. may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" the other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations of the spatial relative descriptions used herein will be made accordingly.
[0070] In addition, it should be noted that the use of terms such as "first", "second" etc. to define components is only for the convenience of differentiating the corresponding components. Without additional statements, the above terms have no special meanings, and thus should not be construed as limiting the protection scope of the present invention.
[0071] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating a detection target image based on image inpainting, characterized in that, the method includes: inputting an original detection image; generating an augmented detection image according to the original detection image; using the augmented detection image to learn and train a detection recognition algorithm model; performing a detection recognition test on the image to be detected according to the trained detection recognition algorithm model.
2. The method according to claim 1, characterized in that, generating an augmented detection image according to the original detection image includes: converting the annotation in text form in the original detection image into an annotation in image form of the same size; forming an image in a three-channel sandwich form by combining the detection image containing the annotation in image form with the original detection image, and using the image in the three-channel sandwich form as the input of the generator; using the generator to generate a target image; using the discriminator to judge the authenticity of the generated target image and a reference ground truth image until the discriminator cannot distinguish the generated target image from the reference ground truth image.
3. The method according to claim 2, characterized in that, each original detection image includes multiple targets to be detected, each target to be detected includes position information and category information, the position information is a four-value array describing the target detection frame, and the category information is an integer category code agreed in advance.
4. The method according to claim 3, characterized in that, converting the annotation in text form in the original detection image into an annotation in image form of the same size includes: generating a full-zero image matrix of the same size as the original detection image as a background encoding map; for the position information, based on the four-value array, taking the detection frame rectangle of the target to be detected in the background encoding map as a solid rectangle frame as the position encoding; for the category information, selecting n gray levels from 0 to 255 to represent n target categories to be detected, where n represents the number of target types to be detected in the entire dataset; assigning the solid rectangle frame of the target to be detected to the gray level corresponding to the category.
5. The method according to claim 4, characterized in that, on the basis that the image in the three-channel sandwich form is used as the input of the generator, adding Gaussian noise with a predetermined intensity as the input of the generator.
6. The method according to claim 5, characterized in that, the generator is a generator based on the U-net framework.
7. The method according to claim 6, characterized in that, the generator includes a self-attention layer and a residual convolution module.
8. The method according to claim 7, characterized in that, in the process of generating the augmented detection image, it includes a generative adversarial loss function, a global point-to-point calculation function, and a bounding box loss function.