A small sample high-quality carrier target detection data set generation method
By using the Stable Diffusion generative large model and specific techniques, the problem of generating target detection datasets for vehicles with small sample sizes was solved, achieving the generation of high-quality datasets and the accuracy of model training, ensuring the accuracy and consistency of vehicle posture and morphological features.
Patent Information
- Application Number
- CN202411682838.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing technologies struggle to generate high-quality target detection datasets for specific vehicle domains, especially when the number of samples is limited, leading to issues such as overfitting and unrealistic scene representation.
By employing a Stable Diffusion generative large model and combining LoRA, ControlNet, and inpaint techniques, a high-quality vehicle target detection dataset is generated through screening, annotation, and feature solidification steps, ensuring the accuracy and consistency of vehicle attitude and morphological features.
It achieves the generation of high-quality vehicle target detection datasets with a small number of samples, ensuring the diversity and representativeness of the data, improving the accuracy and efficiency of model training, and seamlessly integrating vehicles with the background to enhance detection performance.
Smart Images

Figure CN119919751B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target detection data sets, in particular to a small sample high-quality vehicle target detection data set generation method. BACKGROUND
[0002] In the field of target detection, in order to achieve good detection results, a large number of rich training data samples are often needed, but in actual implementation, for a specific vehicle field, it is often difficult to obtain a large number of rich data samples, which has the characteristics of high shooting cost and difficult scene construction.
[0003] Some open source large model data sets often collect massive network pictures as their training sets, which can obtain a large number of samples, but cannot control the scene shooting angle and cannot construct specific types of vehicle data. It is difficult to achieve good results in the detection of specific vehicles.
[0004] Some existing small sample training methods, such as small sample transfer training and incremental learning, are trying to fine-tune or transfer a pre-trained network with a small number of samples to achieve specific target detection capabilities, but due to insufficient samples, overfitting problems often occur, which cannot meet the requirements of target detection indicators in real-world scenarios.
[0005] Some existing sample generation techniques, such as using the network to generate samples or using the cv scheme to manufacture pseudo samples and then a large number of maps, are indeed effective in improving target detection performance, but due to the lack of real scenes, the model often has over-detection problems. SUMMARY
[0006] The purpose of the present application is to overcome the shortcomings of the prior art and provide a small sample high-quality vehicle target detection data set generation method. The Stable Diffusion generative large model technology is used to construct a set of vehicle targets that can be seamlessly embedded in any scene, and specific vehicle types are specified to accurately control the orientation of the vehicle. Only a small number of samples are needed to achieve high-quality data set generation, solving the pain points of data sample collection and construction in specific vehicle fields.
[0007] Technical scheme: The small sample high-quality vehicle target detection data set generation method disclosed by the present application comprises the following steps:
[0008] S1: Select a vehicle model, and select 10 to 100 vehicle images with different attitude angles for the selected vehicle model;
[0009] S2: Label the selected vehicle images to obtain text description information containing vehicle features;
[0010] S3: solidify the morphological features of the selected vehicle model into the input conditions of the Stable Diffusion model, including: loading a pre-trained Stable Diffusion model as a base model, setting training parameters, using vehicle images to train the base model using LoRA, obtaining LoRA model parameters with vehicle morphological features after LoRA training, and fusing the LoRA model parameters with the base model;
[0011] S4: solidify the vehicle posture features into the input conditions of the Stable Diffusion model, including: selecting a vehicle image with a specified posture angle as a reference image for the vehicle posture, extracting the line drawing features of the vehicle from the reference image using line drawing extraction technology, applying ControlNet technology to solidify the line drawing features into the input conditions of the Stable Diffusion, and guiding the Stable Diffusion model to generate images using the line drawing features as input conditions;
[0012] S5: select a target image, create a position for vehicle embedding in the target image using inpaint technology, and generate mask information for local drawing;
[0013] S6: load the text description information, vehicle morphological features, vehicle posture features, and mask information into the Stable Diffusion model, wherein: the text description is converted into a text vector that guides the style of the generated image through a CLIP encoder, the vehicle morphological features are introduced through a LoRA module to activate the morphological features of the vehicle, the posture features are input through a ControlNet module to fix the vehicle posture, and the mask information is input through an inpaint module to define the generated area; the Stable Diffusion model generates a high-quality image according to the input conditions, wherein the position, posture and embedding area of the vehicle are consistent with the target background image, and are naturally integrated with the target background image; output the vehicle image.
[0014] Further improve the above technical solutions, the selection criteria for the vehicle image include: selecting vehicle images containing front view, side view, back view, top view, bottom view, oblique front view and oblique rear view; each posture angle has at least one representative image; the vehicle in the image is in the center of the view angle, without obvious occlusion; the image has good resolution and clarity.
[0015] Further, the annotation processing of the vehicle image includes: collecting and cleaning image data, removing unqualified images; adjusting all images to a uniform size while maintaining the proportion of the vehicle; using an annotation model for preliminary annotation, manually reviewing and correcting the annotation results; selecting the labels generated by preliminary annotation to remove redundant labels; using annotation tools for detailed annotation, including the bounding box and key point information of the vehicle, and ensuring accuracy through multiple rounds of review; storing the label information in a structured form in the database, managing versions and recording modification history.
[0016] Further, the method of solidifying the morphological characteristics of the vehicle model as input conditions for Stable Diffusion includes: selecting a pre-trained Stable Diffusion 1.5 model as a base model; setting training parameters, including learning rate, batch size, training rounds, optimizer, and training size; using vehicle image data to train the base model using LoRA, after training, extracting specific vehicle features and fusing them with the base model; using the extracted features to generate test images, the sampling method is DPM++ 2M Karras, the iteration step is 22-28 times, and the Lora weight is 0.7; manually reviewing the generated images to verify the morphological characteristics of the vehicle meet the expected requirements; if not, the LoRA training parameters or training set need to be adjusted and retrained.
[0017] Further, the method of solidifying the pose features of the vehicle as input conditions for Stable Diffusion includes: selecting a vehicle image with a specified pose angle as a reference image; extracting vehicle line features using Anyline or Canny technology; applying ControlNet technology to solidify the line features as model input conditions; verifying whether the generated image accurately reflects the expected pose features, and adjusting the control parameters of ControlNet according to the verification results.
[0018] Further, the method of integrating the vehicle into a specific image includes: determining the position and size of the vehicle in the target image; using inpaint to create a vehicle embedding position; using BrushNet for edge processing; adjusting brightness, contrast, and color to match the background; inputting the Stable Diffusion model for image generation and optimization.
[0019] Further, the Stable Diffusion model generates a high-quality image according to the input conditions, which includes:
[0020] The descriptive text input by the user is obtained for describing the features of the vehicle, and the descriptive text input by the user is processed by a CLIP encoder to generate a corresponding text vector as a text input condition of the generation model; if the user does not provide a text description, a pre-labeled text description information containing the features of the vehicle is used as a default input condition;
[0021] Load the fused Stable Diffusion model;
[0022] Select a vehicle pose reference graph with a specified perspective, extract the line features of the vehicle from the reference graph using a line extraction technique, and input the extracted line features into the ControlNet module;
[0023] Select a target background image, use inpainting technology to mark the vehicle embedding area in the image, and generate a mask graph to define the local range of vehicle generation;
[0024] Load the text input condition, the pose features of the vehicle, the embedding position information, and the fused Stable Diffusion model into the BrushNet module;
[0025] The BrushNet module is used to enhance the edge processing of the mask area to ensure a natural transition between the vehicle and the background image;
[0026] Adjust the sampler parameters, iteration steps, redraw amplitude and high-definition refinement parameters of the generation model according to the target requirements;
[0027] Generate an image by inputting conditions to drive the Stable Diffusion model;
[0028] The model generates a vehicle image that is fused with the background image in the mask area while maintaining the consistency of the characteristics and pose of the vehicle.
[0029] Further, it further includes the step of manually reviewing the generated image, and the manual review includes:
[0030] Verify whether the morphological features of the generated vehicle meet the expected requirements;
[0031] Verify whether the pose and orientation of the generated vehicle image are consistent with the reference graph or line drawing;
[0032] Verify the fusion effect of the vehicle and the background in the generated image.
[0033] Advantages: Compared with the prior art, the advantages of the present application are that the present application ensures data diversity and representativeness in the data screening process, covers all key pose angles, and high-resolution and high-definition images ensure the accuracy of subsequent processing and labeling.
[0034] The initial labeling is performed using BLIP2 or other algorithms, and manual review and correction are performed. The detailed information of the label includes the bounding box, key points, etc. The label is concise and clear and has been reviewed for multiple rounds to provide high-quality and accurate labeled vehicle image data to support model training and analysis. The uniform size and high-quality labeling ensure the consistency and accuracy of model training.
[0035] A pre-trained Stable Diffusion 1.5 model is selected as the basis, and LoRA technology is used for fine-tuning. The features are extracted and fused with the original Stable Diffusion large model to quickly and efficiently adapt to specific vehicle features and improve the ability of the model to generate specific vehicle images. Low-rank matrix decomposition improves the training efficiency and model adaptability.
[0036] A specific angle of the vehicle image is selected as a reference, and Anyline, Canny, and other line extraction technologies are used to extract line features. ControlNet technology is applied to solidify the line features as input conditions for Stable Diffusion. This improves the accuracy of feature extraction and ensures the consistency and accuracy of the generated image in terms of pose. ControlNet technology enhances the controllability and stability of the generated image.
[0037] Inpaint technology is used to create a vehicle embedding position in the target image, and BrushNet is used for fine edge processing. The Stable Diffusion model is used to generate an image, adjust the brightness, contrast, and color, and seamlessly integrate the vehicle with the background to generate a natural, realistic, and highly consistent image. Fine edge processing and image optimization improve the visual quality and realism of the image. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a flowchart of the method for generating a small-sample high-quality vehicle target detection data set according to the present application. DETAILED DESCRIPTION
[0039] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings, but the scope of protection of the present application is not limited to the described embodiments.
[0040] Example 1: A method for generating a small-sample high-quality vehicle target detection data set as shown in Figure 1 The method includes the following steps:
[0041] 1. Small sample vehicle data screening: 10 to 100 different pose angle vehicle images need to be screened. These images should cover as many poses as possible to ensure data diversity and representativeness. The required pose angles should include but are not limited to: front view, side view (left, right), back view, top view (up), bottom view (down), oblique front view (left front, right front), oblique rear view (left rear, right rear). Each image should have good resolution and clarity, clearly showing the appearance and details of the vehicle. The background diversity of each image is not required, and a simple and clean background should be chosen to facilitate subsequent data processing and analysis. The required samples can be selected from existing databases, public image resources or actual photos. Screening criteria: at least one representative image for each pose angle to ensure sample coverage; the vehicle in the image should be in the center of the view and should not be significantly obscured.
[0042] 2. Vehicle image annotation processing: Vehicle image annotation is the process of accurately identifying and classifying vehicle image data. The specific steps are as follows. First, collect and clean the image data, remove blurred, incomplete or interfering images, and ensure data quality. Then, adjust all images to a uniform size of 512x512 pixels, keeping the vehicle proportion unchanged to avoid distortion. In the label backstepping link, use the annotation model BLIP2 or other algorithms to preliminarily annotate the image, and perform manual review and correction to ensure accuracy and consistency. Next, in the label selection stage, select the most relevant and meaningful labels from the preliminary annotation and remove redundant labels to make each image label concise and clear. Use professional annotation tools to annotate the image in detail, including vehicle bounding box, key points, etc. information, and pass through multiple rounds of review to ensure accuracy. After annotation, store the label information in a structured form in the database and perform version management, record modification history for traceability and recovery. Finally, through consistency check and sample verification, ensure the quality and reliability of the annotated data. Through these steps, a high-quality, accurately annotated vehicle image dataset can be obtained to support subsequent model training, analysis and application.
[0043] 3. Method for solidifying specific vehicle features (referring to morphological features such as shape, size, structure, painting, etc. possessed by specific vehicle models) as input conditions of Stable Diffusion. In order to solidify specific vehicle features as input conditions of Stable Diffusion, first select a pre-trained model suitable for vehicle feature extraction as the base model, such as the Stable Diffusion model. Use LoRA (Low-Rank Adaptation) technology to approximate the neural network weight matrix through low-rank matrix decomposition, thereby speeding up the training speed and improving the adaptability of the model. In the selection of training parameters, set appropriate learning rate 2e-4, batch size 1-4, training rounds (10 to 20 rounds), and select appropriate optimizer AdamW, base model realisticVisionV60B1_v51VAE, dim and Alpha: 64 / 32, single repeat 15, training size 540x540 pixels.
[0044] After training, extract the vehicle features and fuse them with the original Stable Diffusion large model to obtain a model that can generate specific images; finally, test and optimize the model using the extracted features. Through these steps, high-quality vehicle image generation can be achieved, and the use of LoRA model improves the efficiency of feature extraction, ensuring the adaptability and flexibility of the model. After training, generate test images and conduct manual review. Sampling method for generating test images: DPM++ 2M Karras; number of iterations: 22-28; high-definition algorithm: R-ESRGAN_4X; high-definition iteration: 3-7; redraw amplitude: 0.3; Lora weight: 0.7-0.9.
[0045] 4. In order to solidify the attitude features of the vehicle as the input conditions of Stable Diffusion, the specific method is as follows:
[0046] First, collect and select reference images of the vehicle with different pose angles, which should cover a variety of different pose angles, such as front, side, back, etc. Then, use line drawing extraction techniques (such as Anyline, Canny) to extract the line drawing features of the vehicle from these reference images. These line drawings can accurately reflect the shape and pose of the vehicle. Next, apply ControlNet technology to solidify these line drawing features as input conditions for Stable Diffusion. ControlNet is an extended model that can control and guide the details of image generation during the generation process. By using the extracted vehicle line drawing features as the input of ControlNet, it can ensure that the generated image retains the expected pose features. Finally, generate images through the Stable Diffusion model to verify whether the generated images accurately reflect the expected vehicle pose features. Through the above steps, the pose features of the vehicle can be effectively solidified as the input conditions of Stable Diffusion, ensuring that the generated vehicle images have high consistency and accuracy in pose. Using line drawing extraction techniques and ControlNet not only improves the accuracy of feature extraction, but also enhances the controllability and stability of generated images.
[0047] 5、Method for integrating vehicles into specific images, in order to seamlessly integrate vehicles into specific images, first obtain the characteristics and pose of the specific vehicle based on steps 3 and 4, which are used as input conditions for the Stable Diffusion model. Then, select a target image and determine the position and size of the vehicle in it, use inpaint technology to create a vehicle embedding position in the target image, ensuring a natural background transition. Input the processed image into the Stable Diffusion model to further optimize the overall effect and ensure seamless integration of the vehicle with the background. If the background is complex, BrushNet can be used for more detailed edge processing. Embed the generated vehicle image in the specified position of the target image and adjust the brightness, contrast and color to make it consistent with the background.
[0048] 6、Finally, verify and optimize the generated image to ensure that the position, pose and features of the vehicle meet the expectations, thereby generating high-quality data samples. Through these steps, the use of inpaint and BrushNet technology combined with the Stable Diffusion model can accurately control the features and pose of the vehicle, generating natural, realistic and highly consistent images.
[0049] The application is based on an AIGC method, which realizes precise control of the generated target through control and lora, and realizes local and precise redrawing of the original image through Brushnet and inpaint technology, thereby realizing automatic generation of large-scale high-quality carrier data set based on small sample data.
[0050] Embodiment 2: Tank sample production implementation case.
[0051] 1. Tank data screening: 80 clear tank images of different attitude angles of one type need to be screened. These images cover: front view, side view (left, right), back view, overhead view (up), downward view (down), oblique front view (left front, right front), and oblique rear view (left rear, right rear).
[0052] 2. Text annotation of the tank data screened in step 1, the specific steps are as follows: first, remove blurred, incomplete or interfering images to ensure data quality. Then, adjust all images to a uniform size, considering that the shape of the tank is mostly rectangular, uniformly crop and resize to 960x540 pixels, keeping the carrier proportion unchanged to avoid distortion. In the label backstepping link, use the wd-v1-4-convnextv2-tagger-v2 model to preliminarily annotate these samples, then perform manual review and correction, and eliminate some inaccurate tags. In the label selection stage, the most relevant and meaningful labels are selected from the preliminary annotation, and redundant labels are removed, so that the labels of each image are simple and clear, and the accuracy is ensured after multiple rounds of review.
[0053] 3. Solidify specific vehicle features as input conditions for stable diffusion. To solidify specific vehicle features as input conditions for Stable Diffusion, first select a pre-trained model suitable for vehicle feature extraction as the base model. Here, the stabeldiffusion1.5 model is selected as the base model. Using LoRA (Low-Rank Adaptation) technology, the neural network weight matrix is approximated through low-rank matrix decomposition, and a fast training iteration is performed to obtain the weight of an additional network. In terms of training parameters, including learning rate, batch size, training rounds, optimizer, base model, dim and Alpha, single repeat, and training size, the base model is LoRA trained using vehicle image data. After training, extract specific vehicle features and fuse them with the base model; set appropriate learning rate 2e-4, batch size 1-4, training rounds (50-100 rounds), and select appropriate optimizer AdamW, base model realisticVisionV60B1_v51VAE, dim and Alpha: 64 / 32, single repeat 15, and training size 540*540.
[0054] The generated image is manually reviewed by the reviewer, who should sample from the generated data to verify the morphological features of the generated vehicle against the expected requirements; if not, the LoRA training parameters or training set need to be adjusted for re-weighted training.
[0055] The steps for generating vehicle feature images using the Stable Diffusion framework are as follows:
[0056] Step one: Provide descriptive text from the user, or automatically generate descriptive text from images using a label annotation model (such as BLIP2), and convert it into text conditions through a clip encoder. Here, the annotation text used in the LoRA training of the vehicle image in the S2 step can be used;
[0057] Step two: Load the Stable Diffusion model after LoRA feature fusion, and input the preprocessed text into the trained Stable Diffusion model;
[0058] Step three: Adjust the model and sampler parameters, perform inference, and generate images. The model will generate corresponding images based on the input parameter conditions;
[0059] Step four: Output the generated image to the user's specified device or storage path.
[0060] 4、To solidify the pose features of the vehicle as input conditions for Stable Diffusion, the specific method is as follows:
[0061] First, select vehicle images of specific angles as reference pictures, which should cover various different pose angles such as front, side, back, etc. Here, the Anyline line drawing extraction technology is used to extract the line drawing features of the vehicle from these reference pictures. These line drawings can accurately reflect the shape and pose of the vehicle. Next, apply ControlNet technology to solidify these line drawing features as input conditions for Stable Diffusion. By taking the extracted vehicle line drawing features as the input of ControlNet, it can ensure that the generated image retains the expected pose features. Use ControlNet to solidify these pose features as input conditions for the generation model; finally, generate images through the Stable Diffusion model to verify whether the generated images accurately reflect the expected vehicle pose features. Through the above steps, the pose features of the vehicle can be effectively solidified as input conditions for Stable Diffusion, ensuring that the generated vehicle images have high consistency and accuracy in pose. Using line drawing extraction technology and ControlNet not only improves the accuracy of feature extraction, but also enhances the controllability and stability of generated images.
[0062] Verify whether the generated image accurately reflects the expected pose features. The verification personnel should sample from the generated data to verify whether the generated vehicle image is consistent with the input reference image or the vehicle pose and orientation of the line drawing. If not, you need to adjust the control parameters of ControlNet, refer to the official instructions (https: / / github.com / lllyasviel / ControlNet), or check whether the generated line drawing image can clearly reflect the vehicle pose information.
[0063] 5、The method of integrating the vehicle into a specific image, in order to seamlessly integrate the vehicle into a specific image, first obtain the characteristics and pose of the specific vehicle based on steps 3 and 4, which are used as input conditions for the Stable Diffusion model. Next, select a target image and determine the position and size of the vehicle in it. Use the inpaint technology to create a vehicle embedding position in the target image, ensuring a natural transition in the background. Here, BrushNet is used for more detailed edge processing. The entire workflow is connected through the Stable Diffusion network for image generation. Finally, verify and optimize the generated image to ensure that the position, pose, and features of the vehicle meet expectations, resulting in high-quality data samples.
[0064] Determine the position and size of the vehicle in the target image; create a vehicle embedding position using inpaint, which can use the tool's built-in functions (usually the Stable Diffusion framework has official adaptations, or create a corresponding mask image yourself; use BrushNet for edge processing, please refer to the official instructions of BrushNet (https: / / github.com / TencentARC / BrushNet) for specific usage; adjust brightness, contrast, and color to match the background.
[0065] 6. The steps of using the Stable Diffusion framework to generate high-quality vehicle target images are as follows:
[0066] Step one: the user inputs a descriptive text, which is converted into a text input condition through a clip encoder, or uses a label annotation model (such as BLIP2) to automatically generate a descriptive text from an image in Step 2;
[0067] Step two: load the Stable Diffusion model after LoRA feature fusion;
[0068] Step three: select a vehicle pose reference image and extract vehicle line features through Anyline or Canny technology;
[0069] Step four: the vehicle line features are input into the controlnet module to obtain vehicle pose features as input conditions;
[0070] Step five: create a vehicle embedding position using inpaint;
[0071] Step six: input the text input condition, vehicle pose feature input condition, and inpaint information into the fused Stable Diffusion model through the brushnet module for unified loading.
[0072] Step seven: adjust the model and sampler parameters, execute inference, and generate images. The image model will generate corresponding images based on the input parameter conditions.
[0073] As described above, although the present application has been shown and described with reference to specific preferred embodiments, it is not to be construed as being in any way limited to the details shown and described. Various modifications in form and detail can be made without departing from the spirit and scope of the application as defined by the appended claims.
Claims
1. A method for generating a small sample high-quality vehicle target detection dataset, characterized in that, The method comprises the following steps: S1: Select a vehicle model, and screen 10 to 100 vehicle images with different attitude angles for the selected vehicle model; S2: Label the screened vehicle images to obtain text description information containing vehicle characteristics; S3: Solidify the morphological characteristics possessed by the selected vehicle model into the input conditions of the Stable Diffusion model, including: loading a pre-trained Stable Diffusion model as a base model, setting training parameters, using the vehicle image to perform LoRA training on the base model, obtaining LoRA model parameters with vehicle morphological characteristics after LoRA training, and fusing the LoRA model parameters with the base model; S4: Solidify the vehicle attitude characteristics into the input conditions of the Stable Diffusion model, including: selecting a vehicle image with a specified attitude angle as a reference image of the vehicle attitude, extracting the line drawing features of the vehicle from the reference image by using a line drawing extraction technology, and applying a ControlNet technology to solidify the line drawing features into the input conditions of the Stable Diffusion, and using the line drawing features as the input conditions to guide the Stable Diffusion model to generate an image; S5: Select a target image, create a position for vehicle embedding in the target image by using an inpaint technology, and generate mask information for local drawing; S6: Load the text description information, vehicle morphological characteristics, vehicle attitude characteristics, and mask information into the Stable Diffusion model, wherein: the text description is converted into a text vector guiding the generation of image style by a CLIP encoder, the vehicle morphological characteristics are introduced by a LoRA module to activate the morphological characteristics of the vehicle, the attitude characteristics are input by a ControlNet module to fix the vehicle attitude, and the mask information is input by an inpaint module to define the generation area; the Stable Diffusion model generates a high-quality image according to the input conditions, wherein the position, attitude, and embedding area of the vehicle are consistent, and are naturally integrated with the target background image; and output the vehicle image.
2. The method of claim 1, wherein, The screening criteria for the vehicle image include: screening vehicle images containing front view, side view, back view, top view, bottom view, oblique front view, and oblique rear view; at least one representative image for each attitude angle; the vehicle in the image is in the center of the view angle, without obvious occlusion; the image has good resolution and clarity.
3. The method of claim 1, wherein, The labeling process of the vehicle image comprises: Collect and clean up image data, and remove unqualified images; Adjust all images to a uniform size while keeping the vehicle proportion unchanged; Use a labeling model to perform preliminary labeling, and manually review and correct the labeling results; Select the labels generated by the preliminary labeling, and remove redundant labels; Use a labeling tool to perform detailed labeling, including the bounding box of the vehicle and key point information, and ensure accuracy through multiple rounds of review; Store the label information in a structured form in a database, perform version management, and record the modification history.
4. The method of claim 1, wherein, The method for solidifying the morphological characteristics of the vehicle model into the input conditions of Stable Diffusion includes: Selecting a pre-trained Stable Diffusion 1.5 model as the base model; Setting training parameters, including learning rate, batch size, training rounds, optimizer, and training size; Using vehicle image data to perform LoRA training on the base model, after training, extracting specific vehicle features and fusing them with the base model; Using the extracted features to generate test images, the sampling method is DPM++ 2M Karras, the iteration step is 22-28 times, and the Lora weight is 0.7; Artificially reviewing the generated images to verify whether the morphological characteristics of the vehicle meet the expected requirements; if not, the LoRA training parameters or training set need to be adjusted and retrained.
5. The method of claim 1, wherein, The method for solidifying the attitude features of the vehicle into the input conditions of Stable Diffusion includes: Selecting a vehicle image with a specified attitude angle as a reference image; Extracting vehicle line features using Anyline or Canny technology; Applying ControlNet technology to solidify the line features into model input conditions; Verify if the generated image accurately reflects the expected attitude features, and adjust the control parameters of ControlNet according to the verification results.
6. The method of claim 1, wherein, The Stable Diffusion model generates a high-quality image according to the input conditions, which includes: Obtaining user input descriptive text to describe the features of the vehicle, and processing the user input descriptive text through a CLIP encoder to generate a corresponding text vector as the text input condition of the generation model; if the user does not provide text description, use pre-labeled text description information containing vehicle features as the default input condition; Load the fused Stable Diffusion model; Select a vehicle attitude reference image with a specified perspective, extract the line features of the vehicle from the reference image using line extraction technology, and input the extracted line features into the ControlNet module; Select a target background image, use inpaint technology to mark the vehicle embedding area in the image, generate a mask image to define the local range of vehicle generation; Load the text input condition, vehicle attitude features, embedding position information, and fused Stable Diffusion model into the BrushNet module; The BrushNet module is used to enhance the edge processing of the mask area to ensure a natural transition between the vehicle and the background image; Adjust the sampler parameters, iteration steps, redraw amplitude, and high-definition refinement parameters of the generation model according to the target requirements; Generate an image by inputting conditions to drive the Stable Diffusion model; The model generates a vehicle image that is fused with the background image in the mask area, while maintaining the consistency of the vehicle's characteristics and attitude.
7. The method of claim 6, wherein, It also includes the step of manually reviewing the generated image, which includes: Verify whether the morphological characteristics of the generated vehicle meet the expected requirements; Verify whether the attitude and orientation of the generated vehicle image are consistent with the reference image or line drawing; Verify the blending effect of the vehicle and the background in the generated image.
Citation Information
Patent Citations
Paper longitude graph digital restoration method based on convolutional neural network and diffusion model
CN117649365A
Lora-based style migration tampering detection data set generation method
CN118052705A