Image generation method, storage medium, electronic device, and vehicle

By adjusting the parameters of the diffusion generation model and using mask-guided feature fusion, the problems of accuracy and controllability in target object generation in intelligent driving algorithms were solved, achieving efficient and low-cost image generation and improving the training effect of intelligent driving models.

CN122492845APending Publication Date: 2026-07-31BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BYD CO LTD
Filing Date
2026-03-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the research and testing of intelligent driving algorithms, existing technologies suffer from high costs and long time consumption in collecting real-vehicle road data, and it is difficult to obtain image samples of dangerous and extremely rare scenarios through on-site shooting. Existing image synthesis methods cannot accurately control the generation position of target objects, have limited applicable scenarios, and have weak practicality and versatility.

Method used

Based on the feature data of the target object in the reference image, the parameters of the diffusion generation model are adjusted, special prompts are designed using DreamBooth technology, and the model weight parameters are jointly optimized by the reconstruction loss function and the category prior retention loss function. Mask images are generated by combining user interaction or semantic segmentation, and mask-guided feature fusion is performed to achieve accurate and controllable generation of the target object.

Benefits of technology

It enables precise and controllable generation of target objects in intelligent driving scenarios, improves the accuracy and controllability of image generation, reduces data acquisition costs, and enhances model training efficiency and scenario applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492845A_ABST
    Figure CN122492845A_ABST
Patent Text Reader

Abstract

This application relates to an image generation method, storage medium, electronic device, and vehicle. The method includes: adjusting the parameters of a diffusion generation model based on feature data of a target object in a reference image; and generating a target image containing the target object within a preset area of ​​the original image based on the adjusted diffusion generation model. Therefore, this application uses real-world intelligent driving background images as the basis for generation, improving the accuracy, controllability, and scene applicability of intelligent driving image generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive technology, and more particularly to an image generation method, storage medium, electronic device, and vehicle. Background Technology

[0002] In the process of developing and testing intelligent driving algorithms, a large number of image samples containing abnormal targets such as raindrops from the lens, dirt and grime, water accumulation on the road, and cracks in the road surface are needed. Currently, there are two major challenges: first, the cost of collecting real vehicle road data is high and time-consuming, resulting in significant R&D investment; second, image samples of dangerous or extremely rare scenarios are difficult to obtain through on-site shooting.

[0003] Existing image synthesis methods mostly rely on 3D modeling, scene rendering, and physical simulation. These methods require specialized modeling tools to construct scene models and particle simulators to simulate effects such as raindrops. The overall process is complex and demands high levels of equipment and human expertise. Such methods typically only generate general rain-like environmental images and cannot precisely control the generation positions of raindrops, dirt, cracks, etc., on existing real-world background images. Their applicability is limited, and they rely on large amounts of rendering data for model training. This makes it difficult to quickly and flexibly generate abnormal scene images that meet the needs of autonomous driving with a small number of samples, resulting in weak practicality and versatility. Summary of the Invention

[0004] This application provides an image generation method, a storage medium, an electronic device, and a vehicle, which can solve at least one technical problem in the prior art.

[0005] Accordingly, embodiments of this application provide an image generation method, the method comprising:

[0006] Based on the feature data of the target object in the reference image, the parameters of the diffusion generation model are adjusted; based on the adjusted diffusion generation model, a target image containing the target object is generated within a preset area of ​​the original image.

[0007] In one embodiment of this application, adjusting the parameters of the diffusion generation model based on the feature data of the target object in the reference image includes: determining prompt information representing the target object based on the feature data, wherein the feature data includes at least one of visual morphological features, physical attribute features, and semantic category features; and adjusting the trainable weight parameters of the diffusion generation model based on the prompt information and the reference image.

[0008] In one embodiment of this application, determining the prompt information representing the target object based on the feature data includes: generating an object category item for the prompt information based on the semantic category features; determining an identifier item associated with the visual morphological features and the physical attribute features of the target object; and combining the identifier item with the object category item to obtain the prompt information.

[0009] In one embodiment of this application, adjusting the trainable weight parameters of the diffusion generation model includes: adjusting the trainable weight parameters of the diffusion generation model by jointly using a reconstruction loss function and a class prior preservation loss function.

[0010] In one embodiment of this application, generating a target image containing the target object within a preset region of the original image based on the adjusted diffusion generation model includes: inputting the original image into the adjusted diffusion generation model and generating an initial feature map of the original image; generating a mask image based on the preset region, the mask image being used to identify the region in the original image where the target object needs to be generated; and performing mask-guided feature fusion on the initial feature map based on the mask image to obtain the target image containing the target object.

[0011] In one embodiment of this application, generating a mask image based on the preset region includes: obtaining the location information of the preset region, wherein the location information is determined by user interaction or a semantic segmentation model; and generating a mask image that matches the size of the original image based on the location information.

[0012] In one embodiment of this application, the step of performing mask-guided feature fusion on the initial feature map based on the mask image to obtain a target image containing the target object includes: the adjusted diffusion generation model determining the generation region and generation boundary of the target object according to the mask image; controlling the generation region to be consistent with the region represented by the mask image, and dynamically adapting the generation boundary according to the distribution of the target object.

[0013] Accordingly, embodiments of this application provide a computer-readable storage medium including a computer program, which, when run on a computer device, causes the computer device to perform the image generation method described in any of the preceding claims.

[0014] Accordingly, this application provides an electronic device, including: a memory storing a computer program thereon; and a processor for executing the computer program in the memory to implement the image generation method described above.

[0015] Accordingly, embodiments of this application provide a vehicle, including: the electronic device described above; or, a processor, the processor being used to execute the image generation method described above.

[0016] This application provides an image generation method, storage medium, electronic device, and vehicle. By referencing the feature data of the target object in the image, the parameters of the diffusion generation model are adjusted, enabling the model to accurately learn the unique features of the target object. The adjusted model is then used to generate an image containing the target object within a preset area of ​​the original image. This achieves accurate and controllable generation of the target object, ensuring a high degree of matching between the generated result and the target object while effectively limiting the generation area, thereby improving the accuracy, controllability, and scenario applicability of intelligent driving image generation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating one embodiment of the image generation method of this application;

[0019] Figure 2 This is a flowchart illustrating an implementation method of step S100 of this application;

[0020] Figure 3 This is a schematic diagram of an embodiment of the present application, based on reference image one;

[0021] Figure 4 This is a flowchart illustrating one embodiment of step S110 of this application;

[0022] Figure 5 This is a flowchart illustrating one embodiment of step S120 of this application;

[0023] Figure 6 This is a schematic diagram of one implementation method of the generation process of the diffusion generation model of this application;

[0024] Figure 7 This is a flowchart illustrating one embodiment of step S200 of this application;

[0025] Figure 8 This is a flowchart illustrating one embodiment of step S220 of this application;

[0026] Figure 9 This is a schematic diagram of one embodiment of the image generation process of this application;

[0027] Figure 10 This is a schematic diagram of the structure of the electronic device of this application. Detailed Implementation

[0028] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be noted that the following embodiments are merely illustrative of the present application and do not limit its scope. Similarly, the following embodiments are only some, not all, embodiments of the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0029] It should be understood that the terms "upper," "lower," "left," "right," "front," "back," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or relative positional relationship shown in the accompanying drawings. They are used solely for the convenience of describing this application and for simplification, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Unless otherwise specified, the above-mentioned orientational descriptions can be flexibly set in practical applications, provided that the relative positional relationships shown in the accompanying drawings are satisfied.

[0030] The terms "first" and "second" are configured for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0031] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "communication" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection. They can refer to a direct connection or an indirect connection through an intermediate medium, or a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0032] In embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, article, or apparatus that includes that element.

[0033] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0034] The following is a detailed analysis of the proposed solution with reference to the accompanying drawings:

[0035] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the image generation method of this application, as shown below. Figure 1 As shown, the image generation method provided in this application includes the following steps:

[0036] S100, based on the feature data of the target object in the reference image, adjusts the parameters of the diffusion generation model.

[0037] It is understood that the image generation method of this application is mainly aimed at intelligent driving application scenarios and target objects that need to be generated, which includes at least two abnormal scenarios that are crucial to the performance of perception algorithms: abnormal camera scenarios and abnormal road surface scenarios, as described in detail below:

[0038] Specifically, the abnormal camera scenarios refer to visual obstructions caused by external environmental factors during actual vehicle driving, parking, or road condition monitoring. These scenarios directly affect the safety decisions of the intelligent driving system. For example, when driving in rain or on muddy roads, raindrops, mud, dust, and dirty water stains may adhere to the lens surface of the vehicle's surround-view camera or fisheye camera. These obstructions can partially or completely block key information such as traffic participants, lane lines, and traffic lights in the images captured by the camera, causing the perception model to be unable to accurately identify the environment and even make incorrect decisions. Therefore, this application identifies raindrops and dirt on the lens as core generation targets to simulate and train the camera state detection model, enabling it to accurately determine whether the lens is obstructed and trigger safety protection mechanisms (such as intelligent driving function deactivation).

[0039] Furthermore, the abnormal road surface scenarios specifically address the interference caused by the ground environment in determining drivable areas during vehicle parking, low-speed driving, or path planning. These scenarios directly affect the accuracy and stability of intelligent driving systems, especially automatic parking functions. For example, parking lots or urban roads may have cracks, potholes, water accumulation, mud, and other road conditions. These areas visually differ from normal ground, easily causing drivable area judgment models to misclassify them as impassable areas, thus limiting the vehicle's driving path or causing parking failure. Therefore, this application identifies water accumulation and ground cracks as another core generation target to enrich training samples and improve the model's robustness in recognizing complex ground environments.

[0040] Understandably, this application limits the image generation task to the two core requirements of safety and functional implementation in intelligent driving, ensuring that each generated image can directly serve the specific perception model training task and maximize the practical value of the data.

[0041] In the above implementation, by clearly defining the two types of core abnormal scenarios and target objects in the intelligent driving scenario, the subsequent image generation task has a clear focus, which can accurately match the needs of intelligent driving perception algorithm training for special scenario data, avoid generating image content that is irrelevant to actual application, and improve data effectiveness and model training efficiency.

[0042] Of course, in other implementations, to adapt to more complex intelligent driving operating environments, the target object and the corresponding abnormal scenario can be expanded to the following categories:

[0043] 1. Scenarios with abnormal lighting and weather, including but not limited to strong backlight, tunnel shadows, heavy rain, dense fog, sandstorms, blizzards, and low light at night, which can easily lead to a decrease in image quality, an increase in noise, or abnormal contrast.

[0044] 2. Dynamic object interference scenarios, including but not limited to pedestrians or vehicles crossing the field of vision, debris splashing on the road, and objects on the vehicle shaking and obstructing the view. These scenarios will cause temporary obstruction of the field of vision or visual interference.

[0045] 3. System and positioning anomalies, including but not limited to sensor window smudges / fog, positioning signal drift, map data lag, etc.

[0046] The method of this application is also applicable to image generation in the aforementioned extended abnormal scenarios, which can effectively supplement perception training data under all working conditions and further improve the robustness of the intelligent driving system.

[0047] Furthermore, a small number of reference images are acquired, and the parameters of the diffusion generation model are adjusted based on the feature data of the target object in the reference images.

[0048] Please combine further Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of step S100 of this application, as shown below. Figure 2 Step S100 further includes the following sub-steps:

[0049] S110, determine the prompt information representing the target object based on feature data.

[0050] In this embodiment of the application, the feature data of the target object includes at least one of visual morphological features, physical attribute features, and semantic category features, as detailed below:

[0051] 1. Visual morphological features: These features are the intuitive representation of the target object in the image, including the object's geometric outline (such as the direction of a crack or the spherical outline of a raindrop), spatial distribution (such as the location of dirty patches), surface texture (such as ripples on the surface of water), and lighting effects (such as the shadow cast by the object on the background). They are the direct basis for the appearance of the image generated by the model.

[0052] 2. Physical property features: These features define the essential properties of the target object and its interaction with the scene, including the object's material (liquid water, solid asphalt cracks), physical state (adhesion, dripping, closure), and the object's fit / interaction with carriers such as vehicle-mounted cameras and road surfaces, ensuring that the generated image conforms to physical laws and has a sense of realism.

[0053] 3. Semantic category features: These features are general semantic classifications of the target object, including the macro-category to which the object belongs (such as raindrop, dirt, crack). Their role is to provide the model with basic semantic priors, guide the model to learn the core attributes of the target object, and avoid generating semantically incorrect content.

[0054] It is understandable that the above three types of feature data together constitute a complete feature description of the target object. Through the binding and guidance of prompt information, the model can efficiently learn all the features of the target object using only the consulted sample reference image, thereby achieving realistic and controllable image generation.

[0055] Furthermore, unlike traditional methods that require constructing large-scale, high-precision 3D scene models, generating massive amounts of raindrop / stain samples using complex particle physics simulators, and performing rendering and synthesis to construct paired datasets, this embodiment only requires a small number of real reference images containing the target object, such as 3 to 5 images.

[0056] Further integration Figure 3 , Figure 3 This is a schematic diagram of an embodiment of the present application, as shown in Figure 1. Figure 3The reference images in this application can be photographs of actual vehicles taken in real-world environments, or high-quality photographs of similar objects downloaded from public networks or data platforms. The key requirement is that the image content must clearly and completely contain the target object to be generated, such as the shape of raindrops on the lens surface, the texture distribution of dirt and grime, the structural features of ground cracks, and the reflective effect of water on the ground.

[0057] In the above embodiments, data preparation can be completed with a small number of real images, which greatly reduces the cost and time of image data acquisition. There is no need for complex modeling, rendering and simulation tools, which simplifies the data preprocessing process. At the same time, using real images as a reference can ensure the authenticity and practicality of the subsequently generated images.

[0058] Furthermore, based on the aforementioned feature data, we determine the dedicated prompts for DreamBooth fine-tuning.

[0059] Understandably, this step is a crucial preparatory step for model fine-tuning using DreamBooth technology. Its core is designing prompts that uniquely and efficiently bind the feature data of the target object. DreamBooth is a low-resource image fine-tuning technique. Its core mechanism involves strongly binding a custom text identifier to the feature data of the target object in the input reference image, thereby teaching the diffusion model to recognize this new object. In this embodiment, DreamBooth can teach a StableDiffusion model to recognize a new object using only a small number of reference images of a specific object, such as 3-5 images, and enable it to generate an infinite number of new styles and angles of this object.

[0060] Furthermore, the format and design of the prompts adopt the standard DreamBooth prompt word format, which includes custom identifiers, object category items, and categories (i.e., identifier + category name).

[0061] Please refer to further information. Figure 4 , Figure 4 This is a flowchart illustrating an embodiment of step S110 of this application, as shown below. Figure 4 Step S110 of this application, as shown, includes the following sub-steps:

[0062] S111, Based on semantic category features, generate object category items for prompt information.

[0063] In this embodiment, semantic parsing is first performed on the target object in the reference image to obtain the semantic category features corresponding to the target object. These semantic category features characterize the object category to which the target object belongs, such as raindrops, stains, cracks, or puddles. Based on these semantic category features, an object category item is determined and generated in the prompt information. This object category item is a standard category name used to characterize the category to which the target object belongs.

[0064] S112, Identify the identifiers associated with the visual morphological features and physical attribute features of the target object.

[0065] To avoid semantic conflicts and confusion caused by using existing common vocabulary, this implementation adopts a specialized design for the identifier: the identifier uses a random string that does not have a common meaning in existing English dictionaries, rather than existing common English words. By using a random string without common semantics, it is possible to avoid the identifier forming a fixed semantic association unrelated to the target object during the pre-training stage of the diffusion model. This ensures that during subsequent model fine-tuning, the identifier can form a unique and exclusive binding relationship with the features of the target object, avoiding interference with existing semantics and improving the accuracy of model fine-tuning.

[0066] S113, combine the identifier item with the object category item to obtain the prompt information.

[0067] The object category items obtained in the above steps are concatenated and combined to form a prompt message that meets the requirements for fine-tuning the diffusion model. This prompt message is used to adjust the parameters of the diffusion generation model. In this embodiment, the prompt message adopts the standard DreamBooth prompt word format of "identifier item + object category item".

[0068] In a specific application scenario, taking the generation of raindrops attached to a lens as an example: if the general term "raindrops" is directly used as the prompt information, the pre-trained model may already contain the general semantics of "raindrops," causing the fine-tuned model to be unable to accurately learn the unique visual and physical features of the raindrops in the current reference image. Therefore, this implementation uses a random string without general meaning (e.g., "asdfg") as the identifier, concatenating it with the object category "raindrops" to obtain the specific prompt information: "asdfgraindrops".

[0069] During the fine-tuning phase of the diffusion generation model, the identifier term 'asdfg' is deeply bound to the visual morphological features, physical property features, and semantic category features of raindrops in the reference image, ensuring that the identifier term uniquely corresponds to the target raindrop features in this fine-tuning. In subsequent image generation, the model can use the dedicated prompt information to call upon the unique object features bound to the identifier term for image generation, thereby achieving accurate, stable, and unique representation of the target object. This avoids generation deviations caused by semantic interference from common vocabulary, significantly improving the model's learning accuracy and generation stability for specific target objects.

[0070] In the above implementation, by designing prompt words that combine dedicated identifiers with category names, the diffusion model can quickly and accurately learn the unique features of the target object, avoid generation deviations caused by conflicts between identifiers and existing vocabulary, and improve the model's learning accuracy and generation stability for specific target objects.

[0071] S120, based on the prompt information and reference image, adjusts the trainable weight parameters of the diffusion generation model.

[0072] Among them, trainable weight parameters refer to the set of optimizable parameters in the diffusion generative model, which consists of the weights and bias terms of the neural network layers. During the model training process, this set of parameters is updated through the backpropagation algorithm of the loss function, so that the model can learn and represent the features of the target object.

[0073] Specifically, a small number of pre-prepared reference images containing the target object, along with the constructed dedicated prompt information, are used as training data and input into the preset StableDiffusion diffusion generation model to initiate the model fine-tuning training process based on DreamBooth technology. In this embodiment, the number of reference images can be 3 to 5, completing model customization with a small sample size. This model adjustment is not a simple fit to the model, but rather uses two complementary loss functions for joint constraints to ensure optimal model fine-tuning. Specifically, the trainable weight parameters of the diffusion generation model are iteratively adjusted by jointly using the reconstruction loss function and the class prior preservation loss function.

[0074] Please refer to further information. Figure 5 , Figure 5 This is a flowchart illustrating an embodiment of step S120 of this application, as shown below. Figure 5 Step S120 further includes the following sub-steps:

[0075] S121 adjusts the trainable weight parameters of the diffusion generation model by jointly using the reconstruction loss function and the category prior preservation loss function.

[0076] The reconstruction loss function is used to bind the identifier to the visual morphological features and physical attribute features of the target object, and to bind the object category to the semantic category features of the target object; the category prior preservation loss function uses the semantic features corresponding to the object category as constraints to regularize the update process of the trainable weight parameters of the diffusion generation model.

[0077] The first type is Reconstruction Loss. This function binds the identifiers in the prompt information to the visual morphological and physical properties of the target object, and simultaneously binds the object category in the prompt information to the semantic category of the target object. By minimizing the reconstruction loss, the model is forced to learn and memorize the detailed features of the target object in the reference image, ensuring a unique mapping between the identifiers and the feature data of the target object. This guarantees that subsequent image generation maintains a high degree of consistency with the features of the reference image.

[0078] The second type is Class Prior Preservation Loss. This function uses the semantic features corresponding to the object's category as constraints to regularize the update process of the trainable weight parameters of the diffusion generation model. Class Prior Preservation Loss is a key design feature based on DreamBooth fine-tuning, used to avoid overfitting the model on a small number of reference images and prevent the model from only remembering the currently input sample image, thus losing its general generation ability for objects of different shapes within the same category. By introducing Class Prior Preservation Loss, the model learns the unique features of the target object while retaining the general semantic prior knowledge of the corresponding object category learned during pre-training, ensuring the diversity and generalization performance of the model's generated results. See also... Figure 6 , Figure 6 This is a schematic diagram of one implementation method of the diffusion generation model of this application.

[0079] Furthermore, through multiple rounds of iterative training, the trainable weight parameters of the diffusion generation model were continuously optimized, ultimately resulting in a customized diffusion generation model. This model possesses the ability to accurately identify and generate specified target objects, including raindrops, lens dirt, road surface cracks, and road surface water. Compared to the original general diffusion model, the adjusted model can generate diverse images that highly match the target objects in the reference image in terms of appearance, texture, scale, and lighting, and are also well-suited to the background of intelligent driving scenarios.

[0080] Through the DreamBooth fine-tuning process described above, the general diffusion generation model can quickly acquire the ability to generate specific objects required for intelligent driving scenarios, achieve efficient model customization under conditions of a small number of samples, take into account both the accuracy and diversity of object generation, and retain the original general generation capabilities of the model, thereby improving the overall image generation effect and scene applicability.

[0081] S200, based on the adjusted diffusion generation model, generates a target image containing the target object within a preset area of ​​the original image.

[0082] Please combine further Figure 7 , Figure 7 This is a flowchart illustrating an implementation method of step S200 of this application, as shown below. Figure 7 As shown, step S200 includes the following sub-steps:

[0083] S210: Input the original image into the adjusted diffusion generation model and generate the initial feature map of the original image.

[0084] This step provides a realistic and complete scene base for the image generation stage. After the model is fine-tuned, it needs to be provided with a specific, editable intelligent driving scene background to naturally embed the target object into the real environment.

[0085] The selected original background image must be a real, normally captured image from an intelligent driving scenario, meaning the image must not contain raindrops, dirt, cracks, puddles, or other objects that the invention aims to generate. For example, it could be a surround-view camera image taken in clear weather that includes a complete road and lane lines, or a clean, flat parking lot surface image.

[0086] Understandably, the selected original background image serves as the canvas for subsequent local generation operations. It provides a reasonable physical environment and visual context for the generated target objects. For example, generated raindrops need to fall on the lens glass, generated puddles need to appear on the ground, and their lighting, reflection, and perspective relationships need to be consistent with the lighting conditions, viewpoint, and material of the background image. By using realistic background images, the visual plausibility and realism of the final generated anomalous scene images are ensured, enabling them to perfectly simulate data collected from real vehicles and provide high-quality samples for training the perception model.

[0087] Furthermore, the original image is input into the adjusted diffusion generation model to generate an initial feature map of the original image. In this embodiment, the adjusted diffusion generation model (i.e., the StableDiffusion model fine-tuned by DreamBooth) first extracts features from the input original image, mapping the original image to a feature space to obtain an initial feature map that can characterize the basic features of the original image, such as background, texture, and lighting, laying the foundation for the subsequent generation and fusion of target objects.

[0088] In the above embodiments, using real intelligent driving background images as the basis for generation can ensure the scene authenticity and rationality of the final synthesized image, making the generated image closer to the data collected from the actual vehicle, and improving the effectiveness of the image for training the perception model.

[0089] S220, Generate a mask image based on a preset region. The mask image is used to identify the region in the original image where the target object needs to be generated.

[0090] Please refer to further information. Figure 8 , Figure 8 This is a flowchart illustrating an embodiment of step S220 of this application, as shown below. Figure 8 As shown, step S220 includes the following sub-steps:

[0091] S221, Obtain the location information of the preset area. The location information is determined through user interaction or semantic segmentation model.

[0092] In this embodiment, user interaction operations may include the user manually specifying the region in the original image where the target object needs to be generated by selecting with a mouse or marking with a touch, and obtaining the coordinates, range and other location information of the region; the semantic segmentation model can automatically perform semantic analysis on the original image, identify and output the location information of the preset region (such as the lens area of ​​a vehicle lens or the road surface area), realize the automatic acquisition of the preset region location, and improve the efficiency and accuracy of operation.

[0093] Furthermore, this step is a key technical means to achieve controllable generation location, and its purpose is to accurately delineate the region on the background image where the target object needs to be generated. In a specific embodiment of this application, SAM (SegmentAnythingModel) can be used as the region specification tool. In a specific embodiment, the user can provide instructions to the SAM model through simple interactive methods, such as clicking the location where the target object should appear with the mouse, or selecting the approximate area where the target object may exist with the mouse. Based on its powerful image segmentation capabilities, the SAM model can automatically identify and segment continuous regions semantically related to the target object according to the user's interactive instructions, and generate a binary mask image.

[0094] In the generated mask image, white areas represent user-specified target areas where the target object (e.g., raindrops, dirt, cracks, puddles) needs to be generated; black areas represent non-target areas where the original background content needs to be preserved. This application's masking mechanism transforms abstract positional control requirements into image instructions understandable by the model, providing precise spatial constraints for subsequent inpainting operations. The introduction of the SAM model makes the region specification process fast, accurate, and convenient, greatly improving the flexibility and controllability of the generation scheme.

[0095] S222, Generate a mask image that matches the size of the original image based on the location information.

[0096] Furthermore, the mask image has the same resolution and size as the original image. The difference in pixel values ​​clearly distinguishes between the preset area and the non-preset area. For example, the preset area corresponds to the first pixel value (such as 255), and the non-preset area corresponds to the second pixel value (such as 0), thereby accurately identifying the range of the target object to be generated and providing clear regional constraints for subsequent feature repair.

[0097] In the above implementation, the SAM model is used to quickly and accurately generate a mask for a specified region, enabling flexible control over the generation position. The operation is simple and the positioning is accurate, solving the problem that traditional image generation methods cannot control the generation region and position, and meeting the needs of intelligent driving scenarios for the generation of controllable abnormal regions.

[0098] S230, based on the mask image, perform mask-guided feature fusion on the initial feature map to obtain a target image containing the target object.

[0099] This step is the core execution step for generating the final target scene image, and its core technology is Inpainting (image restoration / completion). Inpainting, or image restoration / completion, refers to a technique that generates or fills content only within a specified area of ​​an image. Unlike traditional methods that generate or reconstruct the entire image, Inpainting technology, based on the image's original background, contextual information, and a user-specified region mask, reconstructs, fills, or generates content only within the specific area defined by the mask, while maintaining the overall image structure.

[0100] Furthermore, the adjusted diffusion generation model determines the generation region and generation boundary of the target object based on the mask image, controls the generation region to be consistent with the region represented by the mask image, and dynamically adapts the generation boundary according to the distribution of the target object.

[0101] Specifically, the generated region remains consistent with the region represented by the mask image, indicating that the diffusion generation model only performs feature generation and filling operations on the target object within the preset area identified by the mask, without modifying the original image background area outside the mask, thus ensuring the integrity and realism of the original background. The dynamic adaptation of the generated boundary to the distribution of the target object means that when performing mask-guided feature fusion and content filling, the diffusion generation model does not generate according to fixed, rigid rectangular boundaries. Instead, it adaptively adjusts the generated boundary of the target object based on its shape, density, distribution, and the background structure of the original image. This allows the edges of the generated target object to transition naturally with the surrounding background and conform to the actual distribution of the scene, avoiding harsh, fragmented boundaries and inconsistencies with the scene. Consequently, the final generated target object is more realistic and natural in shape, distribution, and edge transitions compared to the original image, improving the scene consistency and visual realism of the generated image.

[0102] Specifically, the original background image, the binary mask image, and the finely tuned customized diffusion model are input into the Inpainting generation process. Based on the mask's instructions, the model generates content only in white areas. During generation, the model utilizes the target object features learned by the finely tuned model (such as the shape of raindrops, the texture of dirt, the structure of cracks, and the reflection of puddles), combined with contextual information from the background image (such as lighting, viewpoint, and material properties), to intelligently fill content within the target area. The generated target object must not only conform to its own visual characteristics but also achieve a natural and seamless integration with the background image in terms of texture, lighting, scale, and perspective.

[0103] Furthermore, the final output is a complete image of an abnormal scene in intelligent driving, such as... Figure 9 , Figure 9 This is a schematic diagram of one embodiment of the image generation process of this application, as shown below. Figure 9 While preserving the authenticity of the original background, the image is realistically and naturally integrated into the target object at a specified precise location, achieving a controllable position, realistic details, and diverse styles.

[0104] In the above implementation, local area generation is achieved through inpainting, which ensures the realism and diversity of the generated target object without destroying the original background scene, greatly improving the realism and usability of the overall image, while also improving generation efficiency and avoiding the problem of detail distortion caused by full-image generation.

[0105] Finally, the generated large number of diverse and high-quality images of abnormal intelligent driving scenarios are used as training data and input into the core perception model of the intelligent driving system for training.

[0106] Specifically, for applications involving abnormal camera scenarios, a large number of generated images containing raindrops / dirt on the lens are used to train the camera state detection model. The model's task is to analyze images captured by the camera in real time to determine whether there are obstructions on the lens surface. By introducing diverse generated samples, the model can learn lens occlusion features under different shapes, densities, and lighting conditions, thus significantly improving its recognition accuracy and robustness in practical applications. When the model accurately identifies a lens obstruction, it promptly sends a signal to the decision-making module of the intelligent driving system, triggering corresponding safety protection strategies (such as reducing vehicle speed, restricting functions, or prompting the driver to clean the lens) to ensure driving safety.

[0107] For applications involving abnormal road surface scenarios: A large number of generated images containing cracks / water accumulation on the ground are used to train a drivable area identification model. The model's task is to identify areas where vehicles can safely pass through the images during parking or low-speed driving scenarios. Through rich samples of abnormal ground surfaces, the model can learn the visual features of areas such as cracks and water accumulation, accurately distinguishing them from normal ground surfaces. This effectively reduces the likelihood of misclassifying abnormal areas as drivable, thereby lowering the false detection rate and improving the success rate of functions such as automatic parking, as well as the user experience.

[0108] It is understandable that using the image generation method of this application to train the model can effectively reduce the dependence on expensive real vehicle data, significantly shorten the model development cycle, and effectively improve the model's performance in complex environments.

[0109] In the above embodiments, the parameters of the diffusion generation model are adjusted by referring to the feature data of the target object in the image, so that the model can accurately learn the unique features of the target object. The adjusted model is then used to generate an image containing the target object in a preset area of ​​the original image. This enables accurate and controllable generation of the target object. While ensuring that the generated result is highly matched with the target object, the generation area is effectively limited, thereby improving the accuracy, controllability and scene applicability of intelligent driving image generation.

[0110] This application also provides an electronic device, such as... Figure 10 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0111] The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 10 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0112] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing computer programs and / or modules stored in the memory 302, and by calling data stored in the memory 302. Optionally, the processor 301 may include one or more processing cores; optionally, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 301.

[0113] The memory 302 can be used to store computer programs and modules. The processor 301 executes various functional applications and vehicle control by running the computer programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one computer program required for a function (such as voltage control), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include memory electronics to provide the processor 301 with access to the memory 302.

[0114] The electronic device also includes a power supply 303 that supplies power to the various components. Optionally, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0115] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0116] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 302 according to the following instructions, and the processor 301 runs the computer programs stored in the memory 302 to realize various functions, such as:

[0117] Based on the feature data of the target object in the reference image, the parameters of the diffusion generation model are adjusted; based on the adjusted diffusion generation model, a target image containing the target object is generated within a preset area of ​​the original image.

[0118] Therefore, the electronic device provided in this application adjusts the parameters of the diffusion generation model by referring to the feature data of the target object in the image, so that the model can accurately learn the unique features of the target object, and use the adjusted model to generate an image containing the target object in a preset area of ​​the original image. This can achieve accurate and controllable generation of the target object, and while ensuring that the generated result is highly matched with the target object, it effectively limits the generation area, thereby improving the accuracy, controllability and scene applicability of intelligent driving image generation.

[0119] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the detailed description of the image generation method above, which will not be repeated here.

[0120] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the image generation methods provided in embodiments of this application. For example, the computer program can execute the following steps:

[0121] Based on the feature data of the target object in the reference image, the parameters of the diffusion generation model are adjusted; based on the adjusted diffusion generation model, a target image containing the target object is generated within a preset area of ​​the original image.

[0122] Therefore, the storage medium provided in this application embodiment adjusts the parameters of the diffusion generation model by referencing the feature data of the target object in the image, enabling the model to accurately learn the unique features of the target object, and using the adjusted model to generate an image containing the target object in a preset area of ​​the original image. This achieves accurate and controllable generation of the target object, effectively limiting the generation area while ensuring that the generated result highly matches the target object, thereby improving the accuracy, controllability, and scene applicability of intelligent driving image generation.

[0123] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the previous embodiments, which will not be repeated here.

[0124] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0125] Since the computer program stored in the storage medium can execute the steps of any of the image generation methods provided in the embodiments of this application, the beneficial effects that any of the image generation methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0126] This application also provides a vehicle that includes the aforementioned electronic device or processor, the processor being used in real time for the image generation method described in any of the above embodiments.

[0127] The above implementation method adjusts the parameters of the diffusion generation model by referencing the feature data of the target object in the image, enabling the model to accurately learn the unique features of the target object. The adjusted model is then used to generate an image containing the target object within a preset area of ​​the original image. This achieves accurate and controllable generation of the target object, ensuring a high degree of matching between the generated result and the target object while effectively limiting the generation area, thereby improving the accuracy, controllability, and scene applicability of intelligent driving image generation.

[0128] The foregoing has provided a detailed description of a vehicle image generation method, storage medium, electronic device, and vehicle provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only configured to help understand the method and core ideas of this application. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

[0129] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. An image generation method, characterized in that, The method includes: Based on the feature data of the target object in the reference image, the parameters of the diffusion generation model are adjusted. Based on the adjusted diffusion generation model, a target image containing the target object is generated within a preset area of ​​the original image.

2. The image generation method according to claim 1, characterized in that, The step of adjusting the parameters of the diffusion generation model based on the feature data of the target object in the reference image includes: Based on the feature data, prompt information representing the target object is determined, wherein the feature data includes at least one of visual morphological features, physical attribute features, and semantic category features; Based on the prompt information and the reference image, the trainable weight parameters of the diffusion generation model are adjusted.

3. The image generation method according to claim 2, characterized in that, The step of determining the prompt information representing the target object based on the feature data includes: Based on the semantic category features, the object category item of the prompt information is generated; Identify the identifiers associated with the visual morphological features and physical attribute features of the target object; The prompt information is obtained by combining the identifier item with the object category item.

4. The image generation method according to claim 2, characterized in that, The adjustment of the trainable weight parameters of the diffusion generation model includes: The trainable weight parameters of the diffusion generation model are adjusted by jointly using the reconstruction loss function and the category prior preservation loss function.

5. The image generation method according to claim 1, characterized in that, The method of generating a target image containing the target object within a preset area of ​​the original image based on the adjusted diffusion generation model includes: The original image is input into the adjusted diffusion generation model to generate an initial feature map of the original image; A mask image is generated based on the preset region, and the mask image is used to identify the region in the original image where the target object needs to be generated; Based on the mask image, the initial feature map is subjected to mask-guided feature fusion to obtain a target image containing the target object.

6. The image generation method according to claim 5, characterized in that, The process of generating a mask image based on the preset region includes: The location information of the preset area is obtained, and the location information is determined through user interaction or a semantic segmentation model. A mask image matching the size of the original image is generated based on the location information.

7. The image generation method according to claim 5, characterized in that, The step of performing mask-guided feature fusion on the initial feature map based on the mask image to obtain a target image containing the target object includes: The adjusted diffusion generation model determines the generation region and generation boundary of the target object based on the mask image; The generated region is controlled to be consistent with the region represented by the mask image, and the generated boundary is dynamically adapted according to the distribution of the target object.

8. A computer-readable storage medium, characterized in that, Includes a computer program, which, when run on a computer device, causes the computer device to perform the image generation method according to any one of claims 1 to 7.

9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the image generation method according to any one of claims 1-7.

10. A vehicle, characterized in that, The vehicles include: The electronic device as described in claim 9; Alternatively, a processor, said processor being configured to perform the image generation method according to any one of claims 1-7.