Data generation method, model training method and related device

The image generation diffusion model generates and filters the target synthetic image based on the mask image and the target label, which solves the problem of poor data set quality in the prior art and improves the training effect of the image processing model.

CN120071047APending Publication Date: 2025-05-30BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, data sets for model training are generated by copy-paste methods, which have poor quality and are difficult to generate naturally fusion data sets when background images are complex or objects are obstructed.

Method used

The image generation diffusion model generates target synthetic images based on the mask image and the target label, and filters them with aesthetic scores and image quality scores to improve the training data quality of the image processing model.

Benefits of technology

The generated target synthetic images have higher authenticity and improved image quality and accuracy, thus improving the training effect of the image processing model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071047A_ABST
    Figure CN120071047A_ABST
Patent Text Reader

Abstract

The invention relates to a data generation method, a model training method and a related device, and the data generation method comprises the steps: obtaining a target composite image for training an image processing model through an image generation diffusion model according to a mask image and a target label; the target composite image is generated through the mask image, so that the authenticity of the target composite image can be improved; in the process of generating the target composite image by the image generation diffusion model, the natural fusion of the target label and the mask image is beneficial to reducing the randomness of the composite image, so that the image quality and accuracy of the target composite image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technologies, and in particular, to a data generation method, a model training method, and related devices. Background Art

[0002] Object detection and segmentation technologies play an important role in the field of computer vision and have a wide range of applications in various applications, including autonomous driving, video surveillance, and object recognition. Traditional object detection methods usually rely on large-scale labeled datasets, so sometimes they may encounter situations where resources are limited or the cost is high.

[0003] In related technologies, a dataset for model training is generated by the method of copy and paste. However, the quality of the dataset generated by this method is poor. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a data generation method, a model training method, and related devices to solve the technical problems existing in related technologies.

[0005] To achieve the above purpose, in a first aspect, the present disclosure provides a data generation method, including: Using an image generation diffusion model, according to a mask image and a target label, to obtain a target synthetic image for training an image processing model, where the target label is the label of the image expected to be generated.

[0006] Optionally, the step of using an image generation diffusion model, according to a mask image and a target label, to obtain a target synthetic image for training an image processing model includes: Using an image generation diffusion model, according to a mask image and a target label, to obtain an initial synthetic image; Filtering the initial synthetic image to obtain the target synthetic image for training the image processing model.

[0007] Optionally, the step of filtering the initial synthetic image to obtain the target synthetic image for training the image processing model includes: Filtering the images in the initial synthetic image whose aesthetic score and / or image quality score is less than or equal to a preset threshold to obtain the target synthetic image for training the image processing model.

[0008] Optionally, filtering the images in the initial synthetic image whose aesthetic score and image quality score are less than or equal to a preset threshold to obtain the target synthetic image for training the image processing model includes: Filtering the images in the initial synthetic image whose aesthetic score is less than or equal to a preset threshold to obtain candidate synthetic images; Filter out the images in the candidate synthetic images with an image quality score less than or equal to a preset threshold to obtain the target synthetic images for training the image processing model.

[0009] Optionally, the aesthetic score is obtained by an image aesthetics classifier, and / or the image quality score is obtained by a target detector.

[0010] Optionally, the target detector is trained with non-synthetic images labeled with target detection labels.

[0011] Optionally, the mask image is obtained by the following method: Determine non-synthetic images labeled with target detection labels; Add a mask of a preset shape to the region corresponding to the target detection label in the non-synthetic image to obtain the mask image.

[0012] In a second aspect, the present disclosure provides a model training method, including: Obtain synthetic images and non-synthetic images, where the synthetic images are obtained by an image generation diffusion model according to a mask image and a target label, and the target label is the label of the image expected to be generated by the image generation diffusion model; Train an image processing model according to the synthetic images and non-synthetic images.

[0013] Optionally, the image processing model includes a target detection model and / or a target segmentation model.

[0014] Optionally, the training of the image processing model according to the synthetic images and non-synthetic images includes: Extract the synthetic images and the non-synthetic images according to a preset sampling probability, and train the image processing model.

[0015] Optionally, the method further includes: Determine a preset value, where the preset value is used to judge the value for extracting the synthetic image or the non-synthetic image; The training of the image processing model by extracting the synthetic images and the non-synthetic images according to a preset sampling probability includes: When the preset value is less than the preset sampling probability, extract the synthetic image to train the image processing model, and when the preset value is greater than or equal to the preset sampling probability, extract the non-synthetic image to train the image processing model.

[0016] Optionally, the training of the image processing model according to the synthetic images and non-synthetic images includes: When the first score output by the region proposal network in the image processing model is greater than the first threshold, the regions belonging to the foreground and labeled with target labels in the synthetic image or the non-synthetic image are ignored when calculating the loss, where the first score is used to characterize the probability that the regions labeled with target labels belong to the foreground in the synthetic image or the non-synthetic image.

[0017] Optionally, training the image processing model according to the synthetic image and the non-synthetic image includes: When the second score output by the detection head in the image processing model is greater than the second threshold, the regions belonging to the foreground and labeled with target labels in the synthetic image or the non-synthetic image are ignored when calculating the loss, where the second score is used to characterize the probability that the regions labeled with target labels belong to the foreground in the synthetic image or the non-synthetic image.

[0018] In a third aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the methods provided in the first aspect and the second aspect of the present disclosure are implemented.

[0019] In a fourth aspect, the present disclosure provides an electronic device, including: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of any of the methods provided in the first aspect and the second aspect of the present disclosure.

[0020] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the methods provided in the first aspect and the second aspect of the present disclosure are implemented.

[0021] Through the above technical solutions, by inputting the mask image and the target label into the image generation diffusion model, a target synthetic image for training image processing can be obtained, where the target label is the label of the image to be generated, and the target synthetic image can be multiple synthetic images. Generating the target synthetic image through the mask image can improve the authenticity of the target synthetic image; during the process of the image generation diffusion model generating the target synthetic image, the natural fusion of the target label and the mask image helps to reduce the randomness of the synthetic image, and thus can improve the image quality and accuracy of the target synthetic image.

[0022] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure, but do not limit the present disclosure. In the accompanying drawings: Figure 1 is a schematic diagram of data augmentation in the related art.

[0024] Figure 2 is a schematic diagram showing a data generation method according to an exemplary embodiment of the present disclosure.

[0025] Figure 3 is a schematic diagram showing a model training method according to an exemplary embodiment of the present disclosure.

[0026] Figure 4 is a flowchart showing a model training method according to an exemplary embodiment of the present disclosure.

[0027] Figure 5 is a schematic diagram showing a data generation device according to an exemplary embodiment of the present disclosure.

[0028] Figure 6 is a schematic diagram showing a model training device according to an exemplary embodiment of the present disclosure.

[0029] Figure 7 is a block diagram of an electronic device according to an exemplary embodiment. Detailed Description

[0030] The following details the specific embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustrating and explaining the present disclosure and do not limit the present disclosure.

[0031] Currently, generating synthetic data for training has been proven to make up for the problem of insufficient labeled data. In the related art, as Figure 1 shown, data augmentation is performed by a copy-paste technique that copies and pastes target object instances in one image into another image. This method first randomly selects Image Ⅰ and Image Ⅱ from the selected dataset in the data augmentation stage S2, and performs random scale jittering and random horizontal flipping on each image. Then, a subset of target objects is randomly selected in Image Ⅰ, and this subset is pasted at a random position in Image Ⅱ to obtain a new image. Then, by adjusting the true annotation of the new image, the mask and bounding box of the target object that is partially occluded are updated, thereby generating an image with copy-paste data augmentation.

[0032] In the training stage S3, the instance segmentation model is trained using the data-augmented dataset, and different types of models can be selected, such as Mask R-CNN, EfficientNetB7-FPN, etc. In addition, the method also includes a self-training method for semi-supervised learning. The supervised model is trained using the augmented dataset, and then the model is used to train the non-augmented dataset to generate a pseudo-label dataset. Finally, the target object instances are pasted into the pseudo-label dataset and the augmented dataset to generate a pasted dataset, and this dataset is used to train the instance segmentation model to further improve the performance.

[0033] However, the inventors found that when generating a dataset for model training by the method in the related art, the generated dataset may have inappropriate annotation adjustments, which may cause the dataset to not naturally fuse with the corresponding background image; and when generating a dataset through steps such as randomly selecting images, randomly jittering, and randomly flipping horizontally, the quality of the generated dataset may have a certain degree of uncertainty; and when there are too many or complex background images, such as when object images are occluded or overlapped, the corresponding dataset cannot be generated, or the generated dataset does not conform to the object relationships in the real world.

[0034] In view of this, the present disclosure provides a data generation method, a model training method, and related devices to solve the technical problems existing in the related art.

[0035] As Figure 2 shown, Figure 2 is a schematic diagram showing a data generation method according to an exemplary embodiment of the present disclosure. Referring to Figure 2 , it includes: S201: Using an image generation diffusion model, according to the mask image and the target label, obtain a target synthetic image for training an image processing model, where the target label is the label of the image expected to be generated.

[0036] Through the above technical solution, by inputting the mask image and the target label into the image generation diffusion model, a target synthetic image for training image processing can be obtained, where the target label is the label of the image expected to be generated, and the target synthetic image can be multiple synthetic images. Generating the target synthetic image through the mask image can improve the authenticity of the target synthetic image; during the process of the image generation diffusion model generating the target synthetic image, the natural fusion of the target label and the mask image helps to reduce the randomness of the synthetic image, and thus can improve the image quality and accuracy of the target synthetic image.

[0037] To enable those skilled in the art to better understand the data generation method provided by the present disclosure, the above steps will be described in detail with examples below.

[0038] Exemplarily, a mask image can be a special image used to control a processing area or process in image processing. Among them, the mask image can be a mask image of any shape. For example, it can be a mask image of regular shapes such as rectangles and circles, or an image of an irregular shape. The target composite image generated according to the mask image can improve the authenticity of the target composite image.

[0039] A target label can be a label of an image that a user expects to generate. In the image generation diffusion model, the target label is used to fill the masked area in the mask image. In the embodiments of the present disclosure, when the target label and the mask image are input into the image generation diffusion model, a target composite image for training an image processing model can be obtained. Among them, the image processing model can be a target detection model and an image segmentation model. In this regard, the embodiments of the present disclosure do not make specific limitations.

[0040] By adopting an advanced image generation diffusion model, synthetic instances of the same category are generated in the masked area to provide training data for the image processing model. Specifically, given a mask image and its corresponding target label, they are input into the image generation diffusion model. The image generation diffusion model generates a realistic target composite image based on the mask image, the bounding box of the masked area in the mask image, and the target label. This process can ensure that the generated objects are spatially consistent with the background while maintaining the overall semantics and spatial structure of the original image.

[0041] The target composite image generated by the above technical solution can improve the authenticity of the generated objects and help the image processing model better learn how to fill in the missing areas during the training process, thereby improving the effect of missing area repair. In addition, during the process of the image generation diffusion model generating the target composite image, the natural fusion of the target label and the mask image helps to reduce the randomness of the composite image, and thus can improve the image quality and accuracy of the target composite image.

[0042] In a possible manner, obtaining a target composite image for training an image processing model by the image generation diffusion model according to the mask image and the target label includes: Obtaining an initial composite image by the image generation diffusion model according to the mask image and the target label; Filtering the initial composite image to obtain the target composite image for training the image processing model.

[0043] It should be understood that the initial synthesized image can be an image obtained by directly inputting the mask image and the target label into an image generation diffusion model. Then, the initial synthesized image is filtered to remove the images with relatively low quality in the initial synthesized image, obtaining the target synthesized image. Among them, filtering the initial synthesized image can be image set filtering of the initial synthesized image or instance-level filtering of the initial synthesized image. In this regard, the embodiments of the present disclosure do not make specific limitations. By filtering the initial synthesized image, the images with poor quality or incorrect generation in the initial synthesized image can be removed, obtaining the target synthesized image.

[0044] By filtering the initial synthesized image, the quality of the target synthesized image can be improved. Furthermore, when the target synthesized image is used to train an image processing model, the training effect of the image processing model can be improved.

[0045] In a possible manner, filtering the initial synthesized image to obtain the target synthesized image for training an image processing model includes: Filtering the images in the initial synthesized image with an aesthetic score and / or an image quality score less than or equal to a preset threshold to obtain the target synthesized image for training an image processing model.

[0046] It should be understood that the aesthetic score can be the degree of aesthetic attraction of an image. The image quality score can be a score used to evaluate the quality of the initial synthesized image. When the aesthetic score of the initial synthesized image is less than the preset threshold, it can represent that the initial synthesized image does not reach the average degree of people's favorite images, and this initial synthesized image needs to be filtered out. Among them, filtering can be performed through a CLIP (Contrastive Language–Image Pre-training) + MLP (Multilayer Perceptron) architecture, and the model in the CLIP + MLP architecture can be the LAION-Aesthetics Predictor v2 model (a model for predicting the aesthetic quality of images). In this regard, the embodiments of the present disclosure do not make specific limitations. When the image quality score of the initial synthesized image is less than the preset threshold, it can represent that the quality of the initial synthesized image is not good, and this initial synthesized image needs to be filtered out.

[0047] By removing the images in the initial synthesized image with an aesthetic score and / or an image quality score less than or equal to the preset threshold to obtain the target synthesized image, the defects and noises in the target synthesized image can be reduced. Furthermore, the quality of the target synthesized image can be improved, and when this target synthesized image is used to train an image processing model, the image processing model can be better generalized to real-world scenarios.

[0048] In a possible way, filtering the images in the initial synthetic images whose aesthetic scores and image quality scores are less than or equal to a preset threshold to obtain target synthetic images for training an image processing model, including: Filtering the images in the initial synthetic images whose aesthetic scores are less than or equal to a preset threshold to obtain candidate synthetic images; Filtering the images in the candidate synthetic images whose image quality scores are less than or equal to a preset threshold to obtain the target synthetic images for training the image processing model.

[0049] It should be understood that when filtering the initial synthetic images to obtain the target synthetic images, in the embodiments of the present disclosure, the images in the initial synthetic images whose aesthetic scores are less than or equal to the preset threshold can be filtered out first to obtain candidate synthetic images, and then the images in the candidate synthetic images whose image quality scores are less than the preset threshold are filtered out to obtain the target synthetic images.

[0050] By removing the images in the initial synthetic images whose aesthetic scores are less than or equal to the preset threshold and the images whose image quality scores are less than or equal to the preset threshold, the target synthetic images can be obtained, and the initial synthetic images generated with low quality can be effectively removed, thereby improving the quality of the target synthetic images.

[0051] In a possible way, the aesthetic score is obtained by an image aesthetic classifier, and / or the image quality score is obtained by a target detector.

[0052] It should be understood that the aesthetic score in the initial synthetic images can be obtained by an image aesthetic classifier, where the image aesthetic classifier can be a CLIP+MLP architecture model, and the embodiments of the present disclosure do not make specific limitations thereto. The image quality score can be obtained by a target detector, and the target detector can be trained with non-synthetic images labeled with target detection labels.

[0053] In the actual processing process, an image-level filtering is performed by using a pre-trained classifier based on the CLIP+MLP architecture, and the classifier is trained to score the aesthetic attractiveness of the initial synthetic images to obtain the aesthetic scores. By setting a preset threshold corresponding to the aesthetic score, the initial synthetic images with aesthetic scores lower than the preset threshold are discarded, and the corresponding annotations in the discarded initial synthetic images are also removed to obtain candidate synthetic images.

[0054] After that, the candidate synthetic images can be input into the target detector for training, and the candidate synthetic images are determined according to the relationship between the training results of the target detector and the preset threshold. Through two-layer filtering of the pre-trained classifier based on the CLIP+MLP architecture and the target detector, the target synthetic images can be obtained, and the low-quality initial synthetic images can be removed, thereby improving the quality of the target synthetic images.

[0055] In a possible way, the target detector is trained with non-synthetic images labeled with target detection labels.

[0056] It should be understood that the target detector can be trained with non-synthetic images labeled with target detection labels. Among them, the non-synthetic images labeled with target detection labels can be obtained from the Coco dataset (Common Objects in Context, a large-scale image recognition dataset) and the lvis dataset (Large Vocabulary Instance Segmentation, a large-scale instance segmentation dataset).

[0057] In a possible way, the mask image is obtained by the following method: Determine non-synthetic images labeled with target detection labels; Add a mask of a preset shape to the region corresponding to the target detection label in the non-synthetic image to obtain the mask image.

[0058] It should be understood that non-synthetic images labeled with target detection labels are determined in the Coco dataset or the lvis dataset, and the non-synthetic images can be real labeled image data. Then, a mask of a preset shape can be added to the non-synthetic image and on the region corresponding to the target detection label to obtain the mask image. Among them, the preset shape can be a regular shape or an irregular shape.

[0059] By adding a mask of a preset shape to the non-synthetic image, these masked images are input into an image generation diffusion model, enabling it to learn to generate realistic target synthetic images given the mask. It can enable the image generation diffusion model to learn to reasonably fill different mask regions, thereby improving its performance in missing region repair.

[0060] Through the above technical solution, the mask image and the target label are input into the image production diffusion model to obtain an initial synthetic image. After filtering the initial synthetic image by aesthetic scoring and image quality scoring, the target synthetic image can be obtained, which can improve the authenticity of the target synthetic image; and in the process of the image generation diffusion model generating the target synthetic image, the natural fusion of the target label and the mask image helps to reduce the randomness of the synthetic image, and further can improve the image quality and accuracy of the target synthetic image.

[0061] In addition, the inventors also found that when the dataset generated by the method in the related technology is used to train an image processing model, due to the low quality of the dataset, it may be affected by noise when training the image processing model, resulting in poor effects of the trained image processing model.

[0062] As shown Figure 3 in Figure 3 the figure, it is a schematic diagram showing a model training method according to an exemplary embodiment of the present disclosure. Referring to Figure 3 , it includes: S301: Obtain a synthetic image and a non-synthetic image, where the synthetic image is obtained by an image generation diffusion model according to a mask image and a target label, and the target label is the label of the image expected to be generated by the image generation diffusion model; S302: Train an image processing model according to the synthetic image and the non-synthetic image.

[0063] Through the above technical solution, the image processing model is trained with a synthetic image and a non-synthetic image, and the synthetic image is generated by an image generation diffusion model, which can balance the synthetic image and the non-synthetic image, improve the quality of the dataset for training the image processing model, and further improve the performance of the image processing model. At the same time, the effect of the trained image processing model can be improved.

[0064] To enable those skilled in the art to better understand the model training method provided by the present disclosure, the above steps will be described in detail with examples below.

[0065] Exemplarily, the image processing model can be an object detection model or an object segmentation model. In this regard, the embodiments of the present disclosure do not make specific limitations. Among them, the image processing model can be a Faster R-CNN model (a two-stage object detection model), a Mask R-CNN model (a model that simultaneously performs object detection and instance segmentation), and a Centernet2 model (a two-stage object detection model based on probability interpretation). In this regard, the embodiments of the present disclosure do not make specific limitations.

[0066] In a possible manner, the image processing model includes an object detection model and / or an object segmentation model.

[0067] It should be understood that the object detection model is a model that can be used to identify target objects in an image. Among them, the object detection model can be a Faster R-CNN model or a Centernet2 model. In this regard, the embodiments of the present disclosure do not make specific determinations.

[0068] A target segmentation model can be used to detect a target object and simultaneously perform a segmentation mask branch on the target object. Among them, the target segmentation model can be a Mask R-CNN model, a SOLOv2 model (Segmenting Objects by Locations version 2, an object segmentation model based on location), or a segmentation model based on the Detectron2 framework (an open-source framework for computer vision tasks). In this regard, the embodiments of the present disclosure do not make specific limitations. In the embodiments of the present disclosure, the image processing model can be a target detection model and / or a target detection model. In this regard, the embodiments of the present disclosure do not make specific limitations.

[0069] In a possible way, training the image processing model according to the synthetic image and the non-synthetic image includes: Extracting the synthetic image and the non-synthetic image according to a preset sampling probability, and training the image processing model.

[0070] It should be understood that in the embodiments of the present disclosure, when training the image processing model according to the synthetic image and the non-synthetic image, the image processing model can be processed according to a preset sampling probability. Specifically, a data set for training can be extracted from the synthetic image and the non-synthetic image according to a preset sampling probability, and the extracted data set can be used to train the image processing model, thereby improving the performance of the image processing model.

[0071] In a possible way, the method further includes: Determining a preset value, where the preset value is used to judge the value of extracting the synthetic image or the non-synthetic image; The training the image processing model by extracting the synthetic image and the non-synthetic image according to a preset sampling probability includes: When the preset value is less than the preset sampling probability, extracting the synthetic image to train the image processing model, and when the preset value is greater than or equal to the preset sampling probability, extracting the non-synthetic image to train the image processing model.

[0072] It should be understood that the preset value can be used to judge whether the extracted image belongs to the synthetic image or the non-synthetic image. When the preset value is less than the preset sampling probability, a batch of multiple images can be extracted from the synthetic image to train the image processing model. When the preset value is greater than or equal to the preset sampling probability, a batch of multiple images can be extracted from the non-synthetic image to train the image processing model.

[0073] In the embodiments of the present disclosure, when training an image processing model, the image processing model can be trained with different images in multiple batches. Before training each batch, a corresponding preset value can be given by the user. During training, the preset value can be compared with a preset sampling probability, and according to the comparison result, the images of the corresponding batch can be extracted to train the image processing model.

[0074] In the actual operation process, during the model training stage, synthetic images and non-synthetic images are introduced to train the image processing model. When training the image processing model with the data corresponding to multiple batches, the batches of synthetic images and non-synthetic images are alternately used in multiple batches through the preset sampling probability and the preset value. Specifically, assuming that the preset sampling probability is 0.2, for the preset value corresponding to each batch, when the preset value is less than 0.2, multiple images are extracted from the synthetic images to train the image processing model; when the preset value is greater than or equal to 0.2, multiple images are extracted from the non-synthetic images to train the processing model. Thus, during the training process of the image processing model, information from two data sources can be fully learned. And actual annotations with poor quality or incorrect generation can be filtered out.

[0075] In a possible way, training the image processing model according to the synthetic images and non-synthetic images includes: When the first score output by the region proposal network in the image processing model is greater than the first threshold, when calculating the loss, the regions belonging to the foreground and labeled with target labels in the synthetic image or the non-synthetic image are ignored, where the first score is used to represent the probability that the region labeled with the target label belongs to the foreground in the synthetic image or the non-synthetic image.

[0076] It should be understood that the initial synthetic images with poor quality are filtered through aesthetic scoring and image quality scoring, but some noise may still be introduced, such as the case where the prediction quality of the target detector itself is poor. Through aesthetic scoring and image quality scoring, the corresponding actual annotations in the initial synthetic images with poor quality or incorrect generation can be removed, but the instances with poor quality in the target synthetic images cannot be removed. And the image generation diffusion model may generate incorrect images of multiple object category instances, resulting in a lack of annotations. Therefore, in the embodiments of the present disclosure, when training the image processing model according to the synthetic images and non-synthetic images, the foreground in the non-synthetic images can also be judged, and the images can be post-processed according to the foreground.

[0077] In the actual processing, an ignoring function can be introduced during training. Specifically, for the Region Proposal Network (RPN) in an image processing model, if the first score of its foreground category is higher than the first threshold, then when calculating the loss, the regions belonging to the foreground and labeled with target labels in the synthetic image or non-synthetic image can be ignored. Thus, training can be carried out even in the presence of some regions of poor quality, and the performance of the image processing model will not be degraded due to predicting the objects at these positions.

[0078] In a possible way, the training of the image processing model according to the synthetic image and non-synthetic image includes: When the second score output by the detection head in the image processing model is greater than the second threshold, the regions belonging to the foreground and labeled with target labels in the synthetic image or the non-synthetic image are ignored when calculating the loss, where the second score is used to characterize the probability that the region labeled with the target label belongs to the foreground in the synthetic image or the non-synthetic image.

[0079] It should be understood that for the detection head in the image processing model, if the second score of its foreground category is higher than the second threshold, then when calculating the loss, the regions belonging to the foreground and labeled with target labels in the synthetic image or non-synthetic image can be ignored. Thus, training can be carried out even in the presence of some regions of poor quality, and the performance of the image processing model will not be degraded due to predicting the objects at these positions.

[0080] Through the above technical solution, a target synthetic image is generated by an image generation diffusion model. In the image generation diffusion model, the mask image and the target label are used as inputs, and new instances of the same type can be generated centered on the scene in the image to fill the annotation area, obtaining the target synthetic image. Furthermore, the generated target synthetic image can be made more realistic, with higher scene diversity, and can provide more challenging training data.

[0081] By filtering out low-quality images in the initial synthetic images through aesthetic scoring and image quality scoring, potential defects in the target synthetic image can be reduced. When filtering out low-quality images by aesthetic scoring, it can be processed by an aesthetic classifier, and thus the initial synthetic images with lower aesthetic scores can be discarded. When filtering out low-quality images according to the image quality scoring, it can be processed by a target detector, and thus the initial synthetic images with poor quality or incorrectness can be filtered out to obtain the target synthetic image. The possible defects and noises in the target synthetic image can be reduced, thereby improving the quality of the target synthetic image and simultaneously improving the generalization of the image generation diffusion model to real-world scenes.

[0082] According to a preset sampling probability and a preset value, training an image processing model with a dataset composed of synthetic images and non-synthetic images in multiple batches can maximize the utilization of the information of synthetic images and non-synthetic images, and can effectively fuse data from different sources during the training process, thereby improving the performance of the image processing model. Moreover, when the trained image processing model is used to predict long-tail datasets and low-data-volume scenarios, the prediction accuracy can be improved, and the influence of synthetic images and non-synthetic images can be balanced during the training process.

[0083] By ignoring the regions belonging to the foreground and labeled with target labels in the synthetic images or the non-synthetic images in the PRN network and the detection head during the training process, it can be ensured that the image processing model is more robust when processing synthetic images and non-synthetic images, and thus defects and error instances can be reduced.

[0084] Refer to Figure 4 , Figure 4 which is a schematic flowchart showing a model training method according to an exemplary embodiment of the present disclosure. As Figure 4 shown, the flowchart of the model training method includes the following steps.

[0085] S401: Prepare a mask image and target labels.

[0086] S402: Use an image generation diffusion model to generate an initial synthetic image, and filter the initial synthetic image through aesthetic scoring and image quality scoring to obtain a target synthetic image.

[0087] S403: Extract a dataset from synthetic images and non-synthetic images according to a preset sampling probability.

[0088] S404: Use the extracted dataset to train a target detection model and / or a target segmentation model.

[0089] In the actual processing process, this model training method can be embedded in the training process of an object detection model and / or an object segmentation model. A plug-and-play data generation and synthetic image utilization framework can be realized. Furthermore, components such as a generator, filtering technology, model architecture, and training strategy can be flexibly updated as needed. And in the data generation method, non-synthetic images can be extracted from the LVIS dataset and / or the Coco dataset, and synthetic images and non-synthetic images can be introduced into the image processing model, which can improve the performance of the image processing model. Specifically, when the image processing model is the Faster R-CNN model, after training the Faster R-CNN model through the above model training method, the box AP and mask AP in the Faster R-CNN model are increased by 1.5 and 2.5 respectively. When the image processing model is the Mask R-CNN model, after training the Mask R-CNN model through the above model training method, the box AP and mask AP in the Mask R-CNN model are increased by 1.5 and 4.7 respectively. When the image processing model is the Centernet2 model, after training the Centernet2 model through the above model training method, the box AP and mask AP in the Centernet2 model are increased by 0.9 and 2.8 respectively.

[0090] The specific implementation manners of the above process steps have been described in detail by way of examples above and will not be repeated here. In addition, it should be understood that for the above system embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the action sequence described above. Secondly, those skilled in the art should also know that the embodiments described above are preferred embodiments, and the steps involved are not necessarily essential to the present disclosure.

[0091] Through the above technical solution, synthetic images are generated based on an image generation diffusion model, a masked image, and object labels. During the generation process, low-quality synthetic images are removed through filtering by aesthetic score and image quality score. By combining synthetic images and non-synthetic images and performing model training with a preset sampling probability and foreground ignoring, the performance of an object detection model and / or an object segmentation model can be improved.

[0092] Based on the same concept, this embodiment also discloses a data generation device, as Figure 5 shown, Figure 5 is a schematic diagram of a data generation device 500 shown according to an exemplary embodiment of the present disclosure. Referring to Figure 5 it includes: The first image generation module 501 is configured to obtain a target synthetic image for training an image processing model based on a mask image and a target label through an image generation diffusion model, where the target label is a label of an image to be generated.

[0093] Optionally, the first image generation module 501 includes: A second image generation module, configured to obtain an initial synthetic image based on a mask image and a target label through an image generation diffusion model; A filtering module, configured to filter the initial synthetic image to obtain the target synthetic image for training the image processing model.

[0094] Optionally, the filtering module is configured to: Filter images in the initial synthetic image whose aesthetic score and / or image quality score is less than or equal to a preset threshold to obtain the target synthetic image for training the image processing model.

[0095] Optionally, the filtering module includes: A first filtering sub-module, configured to filter images in the initial synthetic image whose aesthetic score is less than or equal to a preset threshold to obtain candidate synthetic images; A second filtering sub-module, configured to filter images in the candidate synthetic images whose image quality score is less than or equal to a preset threshold to obtain the target synthetic image for training the image processing model.

[0096] Optionally, the aesthetic score is obtained through an image aesthetics classifier, and / or the image quality score is obtained through a target detector.

[0097] Optionally, the target detector is trained through non-synthetic images labeled with target detection labels.

[0098] Optionally, the data generation device further includes: A first determination module, configured to determine non-synthetic images labeled with target detection labels; An addition module, configured to add a mask with a preset shape to the region corresponding to the target detection label in the non-synthetic image to obtain the mask image.

[0099] Based on the same concept, this embodiment also discloses a model training device, as Figure 6 shown, Figure 6 is a schematic diagram of a model training device 600 shown according to an exemplary embodiment of the present disclosure. Referring to Figure 6 , it includes: An image acquisition module 601, configured to acquire a synthetic image and a non-synthetic image, wherein the synthetic image is obtained by an image generation diffusion model according to a mask image and a target label, and the target label is a label of an image expected to be generated by the image generation diffusion model; A training module 602, configured to train an image processing model according to the synthetic image and the non-synthetic image.

[0100] Optionally, the image processing model includes an object detection model and / or an object segmentation model.

[0101] Optionally, the training module 602 is configured to: Extract the synthetic image and the non-synthetic image according to a preset sampling probability, and train the image processing model.

[0102] Optionally, the model training device 600 further includes: A second determination module, configured to, wherein the preset value is used to determine a value for extracting the synthetic image or the non-synthetic image; The training module 602 is configured to: When the preset value is less than the preset sampling probability, extract the synthetic image to train the image processing model, and when the preset value is greater than or equal to the preset sampling probability, extract the non-synthetic image to train the image processing model.

[0103] Optionally, the training module 602 is configured to: When a first score output by a region proposal network in the image processing model is greater than a first threshold, ignore regions in the synthetic image or the non-synthetic image that belong to the foreground and are labeled with the target label during loss calculation, where the first score is used to represent the probability that a region labeled with the target label belongs to the foreground in the synthetic image or the non-synthetic image.

[0104] Optionally, the training module 602 is configured to: When a second score output by a detection head in the image processing model is greater than a second threshold, ignore regions in the synthetic image or the non-synthetic image that belong to the foreground and are labeled with the target label during loss calculation, where the second score is used to represent the probability that a region labeled with the target label belongs to the foreground in the synthetic image or the non-synthetic image.

[0105] Based on the same concept, this embodiment also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data generation method and the model training method disclosed in this embodiment are implemented.

[0106] Based on the same concept, this embodiment also discloses an electronic device, including: A memory storing a computer program thereon; A processor configured to execute the computer program in the memory to implement the steps of the data generation method and the model training method disclosed in this embodiment.

[0107] Figure 7 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. As Figure 7 shown, the electronic device 700 may include: a processor 701, a memory 702. The electronic device 700 may also include one or more of a multimedia component 703, an input / output (I / O) interface 704, and a communication component 705.

[0108] Among them, the processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in the above data generation method and model training method. The memory 702 is used to store various types of data to support the operation of the electronic device 700. These data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data, such as contact data, sent and received messages, pictures, audio, video, and so on. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The multimedia component 703 may include a screen and an audio component. Among them, the screen can be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signal can be further stored in the memory 702 or sent through the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, and the above other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.

[0109] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above data generation method and model training method.

[0110] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above data generation method and model training method are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 702 including program instructions, and the above program instructions may be executed by the processor 701 of the electronic device 700 to complete the above data generation method and model training method.

[0111] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0112] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination methods.

[0113] Furthermore, any combination can be made between different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. A data generation method, characterized in that: include: The target synthetic image for training the image processing model is obtained according to the mask image and the target label by using the image generation diffusion model, wherein the target label is the label of the image expected to be generated.

2. The data generation method according to claim 1, characterized in that: The method of generating a diffusion model through an image to obtain a target synthetic image for training an image processing model according to a mask image and a target label includes: The initial synthetic image is obtained according to the mask image and the target label through the image generation diffusion model; The initial synthetic image is filtered to obtain the target synthetic image for training the image processing model.

3. The data generation method according to claim 2, characterized in that: The filtering the initial synthetic image to obtain the target synthetic image for training the image processing model includes: The images whose aesthetic scores and / or image quality scores are less than or equal to a preset threshold value in the initial synthesized images are filtered to obtain the target synthesized image for training the image processing model.

4. The data generation method according to claim 3, characterized in that: Filtering the images whose aesthetic scores and image quality scores are less than or equal to a preset threshold in the initial synthetic images to obtain a target synthetic image for training the image processing model, comprising: Filtering images whose aesthetic scores are less than or equal to a preset threshold in the initial synthesized images to obtain candidate synthesized images; The images whose image quality scores are less than or equal to a preset threshold are filtered out from the candidate synthetic images to obtain the target synthetic image for training the image processing model.

5. The data generation method according to claim 3 or 4, characterized in that: The aesthetic score is obtained by an image aesthetic classifier, and / or the image quality score is obtained by an object detector.

6. The data generation method according to claim 5, characterized in that: The object detector is trained using non-synthetic images annotated with object detection labels.

7. The data generation method according to claim 1, characterized in that: The mask image is obtained by the following method: Determine non-synthetic images annotated with object detection labels; A mask of a preset shape is added to a region of the non-synthesized image corresponding to the target detection label to obtain the mask image.

8. A model training method, characterized in that: include: Acquire a synthetic image and a non-synthetic image, wherein the synthetic image is obtained by an image generation diffusion model according to a mask image and a target label, and the target label is a label of an image expected to be generated by the image generation diffusion model; An image processing model is trained based on the synthetic images and the non-synthetic images.

9. The model training method according to claim 8, characterized in that: The image processing model includes a target detection model and / or a target segmentation model.

10. The model training method according to claim 8, characterized in that: The training of the image processing model according to the synthetic image and the non-synthetic image comprises: The synthetic image and the non-synthetic image are extracted according to a preset sampling probability, and the image processing model is trained.

11. The model training method according to claim 10, characterized in that: The method further comprises: Determining a preset value, wherein the preset value is used to determine a value for extracting the synthetic image or the non-synthetic image; The extracting the synthetic image and the non-synthetic image according to a preset sampling probability and training the image processing model includes: When the preset value is less than the preset sampling probability, the synthetic image is extracted to train the image processing model; when the preset value is greater than or equal to the preset sampling probability, the non-synthetic image is extracted to train the image processing model.

12. The model training method according to claim 8, characterized in that: The training of the image processing model according to the synthetic image and the non-synthetic image comprises: When a first score output by a region proposal network in the image processing model is greater than a first threshold, an area in the synthetic image or the non-synthetic image that belongs to the foreground and is labeled with a target label is ignored when calculating the loss, wherein the first score is used to characterize the probability that the area labeled with the target label in the synthetic image or the non-synthetic image belongs to the foreground.

13. The model training method according to any one of claims 8 to 11, characterized in that: The training of the image processing model according to the synthetic image and the non-synthetic image comprises: When a second score output by the detection head in the image processing model is greater than a second threshold, the area in the synthetic image or the non-synthetic image that belongs to the foreground and is marked with the target label is ignored when calculating the loss, wherein the second score is used to characterize the probability that the area marked with the target label in the synthetic image or the non-synthetic image belongs to the foreground.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

15. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 13.

16. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.