A Small-Sample Object Detection Method Based on High-Quality Synthetic Image Data
By decoupling and selectively combining foreground and background in synthetic data using segmentation and thresholding, the method addresses the issue of low-quality synthetic data in small-sample target detection, improving model performance and generalization.
Patent Information
- Application Number
- CN202510025054.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the case of insufficient data, the existing small sample object detection method has poor robust performance and insufficient generalization capabilities, especially when facing new categories, it is difficult to make full use of limited data for effective learning. In addition, traditional data augmentation methods have synthetic data quality problems, resulting in a degradation of model performance.
By introducing segmentation technology to decouple the prospect-background, and combining the score threshold screening strategy, high-quality prospects and appropriate backgrounds are selected for reorganization, building high-quality synthetic data sets, and improving the generalization ability of the model in a small sample environment.
It significantly improves the model's detection ability and generalization performance of new category targets, especially when there are very few samples and low data quality, effectively alleviating the overfitting problem and maintaining high detection accuracy and robustness.
Smart Images

Figure CN119810592B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning and object detection, and specifically relates to a method for optimizing synthetic image data, which is used to improve the performance of a few-shot object detection (FSOD) model. This method uses a basic image segmentation model to decouple the foreground and background of the generated target image, and selects high-quality foregrounds through a score threshold screening strategy to recombine with appropriate base-class images, thereby obtaining diverse synthetic training data. Compared with traditional data augmentation methods or random synthesis methods, this method can solve the problem of insufficient data and effectively improve the model's detection ability for new-class objects. Background Art
[0002] With the rapid development of deep learning technology, breakthrough achievements have been made in object detection tasks in the field of computer vision, especially in application scenarios such as pedestrian detection, face detection, and vehicle detection. Traditional object detection methods usually rely on a large amount of labeled training data, which is a huge challenge in practical applications, especially in fields where data acquisition is difficult, such as medical imaging, autonomous driving, and wildlife monitoring. To address the above challenges, few-shot object detection has emerged, which aims to correctly identify and locate new classes through a small number of samples.
[0003] Current mainstream few-shot object detection methods can be roughly divided into two categories: one is to enhance the model's ability to extract features from a small number of samples through model architecture design, which includes transfer learning and meta-learning methods. Among them, meta-learning methods update model parameters by simulating task learning and using limited support samples, while transfer learning methods adapt to new-class samples by fine-tuning pre-trained models; the other is to expand the training set through data augmentation techniques to increase data diversity. However, these methods still face problems such as poor model robustness and insufficient generalization ability. Especially when facing new classes, it is still difficult to effectively utilize limited data for effective learning.
[0004] Few-shot object detection methods based on data augmentation techniques play a crucial role in improving model performance. By performing normalization transformations on the original data or generating synthetic data, data augmentation methods can effectively address the problem of insufficient data. However, existing data augmentation operations typically rely on basic image transformations [Yun S, Han D, Oh S J, etal. Cutmix: Regularization strategy to train strong classifiers withlocalizable features[C] / / Proceedings of the IEEE / CVF international conferenceon computer vision. 2019: 6023-6032.] or directly adding synthetic images to the training data, which has certain limitations. Especially when the synthetic data cannot accurately reflect the features of new class objects, it may lead to a decline in model performance. Therefore, how to select or construct high-quality synthetic data and effectively integrate it into the training process remains the main challenge in current few-shot object detection.
[0005] In recent years, generative foundation models (such as DALL·E
Ramesh A, Dhariwal P, Nichol A, etal. Hierarchical text-conditional image generation with clip latents[J].arXiv preprint arXiv:2204.06125, 2022, 1(2): 3.
Lin S, Wang K,Zeng X, et al. Explore the power of synthetic data on few-shot objectdetection[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition. 2023: 638-647.
Rombach R,Blattmann A, Lorenz D, et al. High-resolution image synthesis with latentdiffusion models[C] / / Proceedings of the IEEE / CVF conference on computervision and pattern recognition. 2022: 10684-10695.
[0006] To address the above problems, the present invention proposes a high-quality synthetic data augmentation method that combines segmentation for decoupling and foreground-background screening and recombination. This method generates generated images of new categories through a generative base model DALL·E, decouples the foreground and background of the generated images using a segmentation base model SEEM, and then combines a fractional threshold screening strategy to select high-quality new category samples from the decoupled foreground and recombine them with the matching background. Compared with traditional data augmentation methods or random synthesis methods, the synthetic data augmentation method based on segmentation decoupling and screening recombination can more accurately screen out representative and authentic generated data, and improve the diversity and adaptability of synthetic data through the recombination of foreground and background, thereby significantly enhancing the model's detection ability for new category targets, especially showing superior performance and generalization ability in the few-shot scenario. Summary of the Invention
[0007] The present invention addresses the problem of data scarcity in few-shot object detection (FSOD) and proposes an innovative synthetic data selection and utilization framework to improve the generalization ability and detection accuracy of the model in few-shot learning scenarios. Existing few-shot object detection methods usually rely on standard data augmentation methods, but these methods often lead to performance degradation due to synthetic data quality issues. The present invention effectively improves the performance of few-shot object detection models by introducing a decoupling and selection strategy to synthesize high-quality data, especially in the case of extremely few samples and low data quality.
[0008] In few-shot object detection (FSOD), the base classes are the targets with a large number of samples and known categories in the training set, while the new classes are the unknown category targets with a small number of samples or even only a few samples. For example, the base classes may include common vehicles (such as cars, trucks) or animals (such as dogs, cats), while the new classes may be rare species (such as foxes) or objects in special scenarios (such as drones).
[0009] The first innovation of the present invention is to propose a foreground-background decoupling module based on segmentation technology. This module uses a segmentation base model (such as SEEM [Zou X, Yang J, Zhang H, et al. Segment everything everywhere all at once[J]. Advances in Neural Information Processing Systems, 2024, 36.]) to finely segment the synthetic image, and selects the highest-confidence mask that obeys the threshold range from several obtained masks as the best foreground, and preferentially selects synthetic samples with clear edges and complete shapes. This process effectively reduces the interference of irrelevant backgrounds and ensures the quality of the synthetic image. The second innovation is to propose a synthesis module, which further improves the quality and diversity of the synthetic image by performing a fixed-threshold screening on the synthetic foreground image through an image encoder. Finally, the synthesis module combines the high-quality foreground image with the base class image that meets the requirements of diversity and consistency to enhance sample diversity and improve the localization ability of the detection model.
[0010] The technical solution of the present invention
[0011] A small-sample object detection algorithm based on high-quality synthetic image data, the steps are as follows:
[0012] Step 1: Synthetic data generation and foreground-background decoupling
[0013] 1.1. Synthetic data generation
[0014] To generate new object images, we input a small number of new class names into ChatGPT to let it generate detailed text prompts (prompts) according to each new class. Then, we feed these generated prompts into the DALL·E model to utilize its image generation ability. By using the multi-round dialogue ability of ChatGPT, the text prompts can be generated and gradually optimized to make them more suitable for generating prominent and clear target images. In this way, we can ensure that each generated image contains a clear object and provide higher feasibility and accuracy for the subsequent foreground-background decoupling process.
[0015] 1.2. Foreground-background decoupling
[0016] After generating the synthetic image, we use a foreground-background decoupling module based on segmentation technology to accurately separate the object and background in the synthetic image. To solve the problems of background interference and boundary artifacts in traditional methods, the present invention combines SEEM technology to optimize the decoupling process.
[0017] The steps of the algorithm are as follows:
[0018] 1.2.1. Generate the initial mask
[0019] Input the generated image into a segmentation base model (such as SEEM) to generate multiple foreground candidate masks . These masks represent different segmentation regions, and each mask is accompanied by a corresponding confidence score .
[0020] 1.2.2. Confidence screening
[0021] First, set a confidence threshold (such as 0.5 or 0.7). If a mask passes the screening, only retain the mask with a confidence score higher than the threshold ; if there are multiple masks that pass the screening, select the mask with the highest confidence. Denote this screened mask matrix as
[0022] 1.2.3. Extraction of the foreground image
[0023] In this technical solution, the extraction of the foreground image is achieved by applying the mask . This mask is a matrix with a value range of 0 - 1, where the area with a value of 1 represents the corresponding image pixels, and the area with a value of 0 is the background. The mask matrix is multiplied element - by - element with the generated image to obtain the foreground image . This step ensures that the foreground extracted from the generated image has clear object boundaries and minimizes background interference, providing a reliable basis for subsequent synthetic data
[0024] Step 2: Screening and combination of the foreground and background
[0025] 2.1. Foreground screening
[0026] The quality of the generated foreground image is crucial for the effectiveness of the synthetic data. To ensure that the selected foreground has high quality and representativeness, the present invention adopts a screening strategy based on feature similarity to screen out high - quality foreground samples through feature matching with real samples. The specific steps are as follows
[0027] 2.1.1. Feature extraction
[0028] Use an image encoder (such as CLIP) to extract features from the generated foreground image and the foreground image of the same - type target in the real dataset . The extracted feature vectors are respectively denoted as and .
[0029] 2.1.2. Similarity Calculation
[0030] Calculate the feature similarity between the generated foreground image and the real sample , where represents the cosine similarity formula
[0031] 2.1.3. Threshold Screening
[0032] Set the similarity threshold (such as 0.7), and filter out the set of foreground samples whose similarity is greater than or equal to this threshold :
[0033]
[0034] 2.1.4. Diversity Constraint
[0035] To ensure the diversity of foreground samples and avoid over-concentration of the filtered samples, the following strategies can be adopted
[0036] (1) Feature coverage strategy: Through clustering analysis or dimensionality reduction visualization, preferentially select samples with a wider coverage range in the feature space to ensure diversity
[0037] (2) Class balance strategy: Ensure that the number of filtered samples for each class is relatively balanced, and avoid a too high proportion of samples in a specific class, thereby improving the fairness and representativeness of the data distribution
[0038] 2.2. Background Screening
[0039] Select images suitable as backgrounds from the basic dataset. The quality of the background images has an important impact on the diversity and authenticity of the synthesized images. Specifically, suitable background images should conform to two properties
[0040] (1) Diversity: When screening background images, to ensure coverage of various environmental types (such as indoor, urban streets, natural landscapes, etc.) and increase the diversity of the background. The control of the scene distribution can be achieved through clustering analysis of the background images for diverse selection
[0041] (2) Consistency: The background images should have a certain semantic or visual consistency with the foreground objects to avoid unreasonable combinations affecting the authenticity of the synthesized images. For example, when the foreground object is a vehicle, the background should be selected as scenes such as streets and parking lots, rather than natural landscapes or indoor scenes. By calculating the feature similarity between the foreground and the candidate background images and selecting the background with a higher degree of feature matching with the foreground, the rationality and credibility of the synthesized images can be further improved
[0042] 2.3. Combination of Foreground and Background
[0043] The screened foreground images are recombined with the background that meets the requirements of diversity and consistency to generate diverse synthetic images. To avoid conflicts between the foreground and the background, a position-aware strategy is adopted to adjust the position and scale of the foreground. The specific steps are as follows:
[0044] 2.3.1. Synthetic Image Generation and Diversity Enhancement
[0045] The adjusted foreground images are randomly pasted onto the background images to generate synthetic images :
[0046]
[0047] wherein, represents the paste operation, represents the placement position of the foreground object. To further improve the diversity of the synthetic data, we randomly adjust the rotation angle, brightness, color and other attributes of the foreground images during the synthesis process. These transformations increase the diversity of the synthetic images and enhance the generalization ability of the model.
[0048] 2.3.2. Quality Inspection
[0049] The generated synthetic images are subject to quality inspection to remove samples that do not meet the requirements and ensure the validity of the synthetic data.
[0050] Step 3: Construction of the Synthetic Dataset and Model Training
[0051] 3.1. Construction of the Synthetic Dataset
[0052] The high-quality synthetic images generated through the above steps are used to construct a synthetic dataset, which is then fused with the real basic class dataset: the synthetic data and the real data are mixed in a certain ratio (such as 10:1) to form a diverse training set.
[0053] 3.2. Model Training
[0054] We use the mainstream object detection model Faster R-CNN [Ren S, He K, Girshick R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE transactions on pattern analysis and machine intelligence, 2016, 39(6): 1137-1149.] for training and use an end-to-end optimization method to improve the model's detection ability for new category targets. To better train the model, we use a jointly optimized loss function to optimize the classification loss and the bounding box regression loss respectively. Among them, the classification loss is used to measure the prediction accuracy of the model for target categories, while the bounding box regression loss is used to measure the prediction accuracy of the model for target positions and sizes. By optimizing these two loss terms simultaneously, the detection performance of the model on new category targets can be effectively improved.
[0055] Advantages of the present invention: First, it can improve the generalization ability of the model. By introducing a decoupling and screening strategy, the present invention can make full use of the generated synthetic data and effectively alleviate the overfitting problem caused by data scarcity. In particular, through foreground-background decoupling processing and threshold screening strategy for synthetic images, the generalization ability of the model in a small sample environment is effectively improved. The present invention can maintain a high detection accuracy in a variety of complex scenarios, enhancing the robustness of the model in different scenarios and insufficient data volumes.
[0056] Second, it can optimize the quality of synthetic data. Compared with traditional data augmentation methods, the present invention finely selects high-quality foreground and background through two modules and then recombines them, thereby ensuring that the used synthetic images have higher authenticity and representativeness. This high-quality synthetic data not only effectively improves the diversity of training data but also helps the model better learn the key features of target objects. Especially in the case of scarce samples, this method can better make up for the problem of insufficient data and improve the overall performance of the model.
[0057] The invention can be flexibly adapted to a variety of models. The framework design of the present invention has high flexibility and can be seamlessly integrated into a variety of models. This not only makes it applicable to existing small sample detection methods but also can be effectively combined with the latest fine-tuning-based model architectures, thereby fully exploiting the potential of the models and maximizing the performance. Thanks to this flexible adaptability, the present invention provides a wider application scenario and technical support for small sample object detection.
[0058] Finally, significantly improve the detection accuracy in the cases of few-shot and extremely few-shot learning. In the scenario of extremely few-shot learning, traditional detection methods often suffer from performance degradation due to insufficient data. However, the present invention significantly improves the performance of the model in few-shot learning by screening high-quality synthetic data and selectively applying it. Especially in the case of an extremely small number of training samples, the present invention can effectively alleviate the problems brought by data scarcity, greatly improve the detection accuracy, and demonstrate strong advantages under limited data conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 FIG. is an overall flowchart of a few-shot object detection method based on high-quality synthetic image data.
[0060] Figure 2 FIG. is an application example diagram of a few-shot object detection method based on high-quality synthetic image data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The following further describes the specific embodiments of the present invention in combination with the drawings and technical solutions.
[0062] Step 1: Data Generation and Preprocessing
[0063] First, input a small number of new class names into ChatGPT and let it generate detailed text prompts (prompts) for each class. Input class: "bird" ChatGPT initial output: "A sky-blue little bird stands on a tree stump with a rock background." Adjust the prompt: "Please remove the rock in the background, set the background to a single color or blur it, and retain the details of the bird and the tree stump." Optimized output: "The feathers of the bird are a bright sky blue. It stands on a wooden stake, its claws tightly gripping the stake. The background is a soft brown, simple and blurred, highlighting the main body of the bird and making the picture appear peaceful and natural." Use a generative base model (such as DALL·E) to generate images of the target class. During the generation process, the quality of the output images can be optimized by adjusting model parameters (such as the number of diffusion steps or generation diversity). Ensure that the generated images have clear foregrounds and backgrounds that are easy to segment to meet the requirements of subsequent tasks.
[0064] Step 2: Foreground-Background Decoupling
[0065] Use SEEM to perform fine segmentation on the generated target images to achieve the decoupling of foreground objects and the background. Multiple masks are generated during the segmentation process, and each mask represents a possible foreground region, thus providing multiple segmentation options. Perform confidence scoring on all the masks of the segmentation results, and filter out the mask with the highest score according to the confidence must be greater than 0.8. Prioritize masks with complete boundaries and high confidence to effectively reduce background interference and highlight the key features of the target, laying a foundation for subsequent processing.
[0066] Step 3: Foreground Screening
[0067] Use CLIP to extract the feature vectors of the foreground images, and at the same time extract the feature vectors of the corresponding category samples in the real dataset. Calculate the matching degree between the feature vectors through the cosine similarity formula, and evaluate the proximity between the generated foreground and the real samples in the feature space. Set the similarity threshold ≥ 0.7 to screen out the foreground images with higher feature similarity to the real samples. The screened foreground samples not only have high authenticity but also retain good diversity in the feature space, thus enhancing the data quality.
[0068] Step 4: Background Screening
[0069] Perform clustering analysis on the background images and divide them into different scene categories: urban, indoor, natural, etc. Select appropriate scene categories, and then select diverse images from each scene category to ensure that the background data covers a variety of environmental types, thereby enhancing the diversity and representativeness of the synthetic data.
[0070] Step 5: Foreground and Background Recombination
[0071] Detect the positions of the basic class objects in the background images, and adjust the scaling ratio and placement positions of the foreground images according to the detection results to avoid an intersection over union (IoU) higher than 0.1 between the foreground and background objects. Through this position-aware strategy, the feature interference between the foreground and background objects can be effectively reduced. Subsequently, paste the adjusted foreground images onto the background images to generate new synthetic images. During the pasting process, the brightness, color, and rotation angle of the foreground objects can be randomly adjusted to further enhance the diversity and authenticity of the synthetic data.
[0072] Step 6: Synthetic Dataset Construction
[0073] Fuse the generated high-quality synthetic images with the real basic class dataset and construct a new training dataset by mixing them in a 1:1 ratio. During the dataset fusion process, it is necessary to ensure that the class distribution of the synthetic samples and the real samples is balanced to avoid disproportionate ratios of certain category samples. Conduct strict quality inspections on the generated synthetic dataset, and eliminate images with serious distortion, foreground-background conflicts, or invalid features to ensure the overall quality of the training data.
[0074] Step 7: Model Training
[0075] The model is trained using the mainstream object detection framework Faster R-CNN. By mixing real samples and synthetic samples, the classification ability and object localization ability of the model are gradually optimized. During the training process, a joint loss function is adopted, combining the classification loss and the bounding box regression loss, to improve the detection performance of the model for objects of new categories. At the same time, by adjusting hyperparameters (such as learning rate, weight decay, etc.), the adaptability of the model to synthetic data is improved. This process fully covers all key steps from data generation, screening, recombination to model training, providing a systematic solution to the small-sample object detection problem. Precise text prompts are generated through ChatGPT, and high-quality images are generated in combination with generative foundation models (such as DALL·E). At the same time, technologies such as SEEM and CLIP are used for foreground-background decoupling and feature similarity screening to ensure the authenticity and diversity of the data. In addition, background clustering analysis and position-aware strategies further optimize the diversity and synthesis quality of the data, and the strategy of training with a mixture of real data and synthetic data effectively improves the detection ability of the model for objects of new categories. Generally speaking, this process has a clear structure, strong technological innovation, and proposes effective solutions to the problems of insufficient data and synthetic data quality. However, the implementation process may have relatively high requirements for computing resources and algorithm performance, and there is still room for optimizing efficiency.
Claims
1. A small-sample object detection method based on high-quality synthetic image data, characterized in that, The steps are as follows: Step 1: Synthetic data generation and foreground-background decoupling Utilize Generate text prompts for new categories, where the new categories are target categories with few or even unknown sample quantities; Feed the text prompts into the model to generate synthetic images; Input the synthetic image into the segmentation base model to generate multiple foreground candidate masks, where each foreground candidate mask represents a different segmentation region and is associated with a corresponding confidence score; Set a confidence threshold; if there is a foreground candidate mask that passes the screening, only retain the foreground candidate masks with confidence scores higher than the confidence threshold T; if there are multiple foreground candidate masks that pass the screening, then select the foreground candidate mask with the highest confidence score, and denote the filtered foreground candidate mask matrix as , in this mask matrix, the area with the value of represents the corresponding image pixels, and the area with the value of is the background; the mask matrix is multiplied element-wise with the synthesized image to obtain the foreground image ; Step 2: Selection and combination of foreground and background Use an image encoder for the generated foreground image and the foreground images of the same type of objects in the real dataset to extract features, and the extracted feature vectors are respectively denoted as and ; Calculate the feature similarity between the generated foreground image and the true sample , denotes the cosine similarity; Set the similarity threshold and filter out the set of foreground images whose feature similarity is greater than or equal to the similarity threshold , ; To ensure the diversity of foreground images, the following strategies are adopted: (1) Select foreground images with a wider coverage range in the feature space through clustering analysis or dimensionality reduction visualization; (2) Class balance strategy: Ensure that the number of selected samples for each class is relatively balanced; Select suitable background images from the basic dataset. The background images should cover various environmental types and have semantic or visual consistency with the foreground images; The screened foreground image is recombined with the background image that meets the requirements of diversity and consistency to generate diverse synthetic images; Adopt a position-aware strategy to adjust the position and scale of the foreground images; Randomly paste the adjusted foreground image onto the background image to generate a synthetic image , and randomly adjust the rotation angle, brightness, and color of the foreground image during the synthesis process; Perform quality inspection on the generated synthetic images and remove those that do not meet the requirements; Step 3: Construction of the synthetic dataset and model training The synthetic image generated through the above steps Construct a synthetic dataset and fuse it with the real base class dataset: the synthetic data and the real data are mixed in proportion to form a diverse training set; use an object detection model for training.
Citation Information
Patent Citations
Image synthesis method and device, electronic equipment and computer readable storage medium
CN111626919A
Method and system for harmonizing synthetic image based on foreground reference image
CN115205544A