Method and device for generating display image, electronic equipment and storage medium
By using a method and device for generating display images, a target LoRa model is used to generate images consistent with the actual object, which solves the problem of inconsistent image generation of large models under fixed object position and size, and improves the efficiency and quality of advertising production.
Patent Information
- Application Number
- CN202510667470.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, large models cannot generate images in which objects are consistent with the actual objects when the object position and size are fixed, resulting in time-consuming and costly advertising production and limited creativity.
By obtaining the target image and text prompt of the target object, a base map containing the target image is generated using the preset large model and the target LoRa model corresponding to the target object, and the composition and character corrections are performed on the base map. Finally, the details are redrawn and corrected to generate the final display image.
The consistency between the target object in the generated image and the actual object is achieved, which improves the efficiency and quality of advertising production, reduces the dependence on professional designers, shortens the production cycle and provides unlimited creative possibilities.
Smart Images

Figure CN120635232A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for generating display images, an electronic device, and a storage medium. Background Art
[0002] In the advertising business, various factors often limit the specific display materials (e.g., merchandise) required for promotion. Advertisers and creative professionals must adhere to a meticulous production process to ensure both quality and effectiveness. First, team members must develop creative ideas and formulate an advertising strategy based on the product's characteristics and market positioning. Next, designers create a visual plan based on this strategy, selecting elements such as color, layout, and imagery to ensure the ad's appeal and reach.
[0003] This process places a significant demand on manual effort. High-quality advertising production relies on the designer's professional skills and aesthetic judgment, which often necessitates lengthy composition work, as well as repeated revisions and proofreading. The advertising industry's high demands for creativity and design detail mean that even minor modifications can require significant manpower and time. Such a complex and demanding process naturally presents bottlenecks. Time and cost are the primary challenges of current advertising production. A shortage of specialized personnel often limits both the speed of production and the breadth of creative output. Furthermore, aesthetic fatigue is another issue: creative exhaustion can occur after advertising producers work on the same theme for extended periods of time.
[0004] In sharp contrast, the technology of text-based graphics (i.e., models that generate images based on text descriptions) based on large AI models has attracted widespread attention in recent years. Although AI-based advertising production also has problems, such as the generated images may not fully meet advertising standards or be not detailed enough, its advantage is that it reduces dependence on professional designers, allowing users without a design background to quickly generate visual content of a certain quality. It can also greatly shorten the production cycle, reduce costs, and to some extent provide unlimited creative possibilities. For problematic result images, only professionals need to make appropriate modifications, which is much less time-consuming and less dependent than a purely manual production process.
[0005] However, when the product's appearance is fixed, relying solely on a large model workflow has limited results. For example, relying solely on prompts (text prompts) cannot ensure that the generated product is consistent with the specified product, the product cannot be reasonably placed on an object at any size, and the product cannot coexist reasonably with people in the background.
[0006] This shows that the related art has a problem that large models cannot make objects in images consistent with actual objects. Summary of the Invention
[0007] The present application provides a method and device, an electronic device and a storage medium for generating a display image, so as to at least solve the problem in the related art that a large model cannot generate a reasonable image based on an object when the position and size of the object are fixed.
[0008] According to one aspect of an embodiment of the present application, a method for generating a display image is provided, comprising:
[0009] Acquire a target image of a target object;
[0010] According to the target image and the text prompt, and through a target model corresponding to the target object, a base map containing the target image and constrained by the text prompt is generated, wherein the target model includes a preset large model and a target lora model corresponding to the target object;
[0011] Correcting the composition of the base image and the characters in the base image to obtain an intermediate image;
[0012] The intermediate image is redrawn and corrected in detail to obtain a final output refined image, and the refined image is used as the display image of the target object.
[0013] Optionally, as in the aforementioned method, generating a base map containing the target image and constrained by the text prompt based on the target image and the text prompt and using a target model corresponding to the target object includes:
[0014] Determine the target model corresponding to the target object among all candidate models according to the type of the target object;
[0015] The target image and the text prompt are input into the target model to generate a base map containing the target image and constrained by the text prompt, wherein the base map includes the target image and a background image.
[0016] Optionally, as in the aforementioned method, inputting the target image and the text prompt into the target model to generate a base map containing the target image and constrained by the text prompt includes:
[0017] Obtaining edge detection information of the target image by executing an edge detection algorithm on the target image;
[0018] The target image, the edge detection information and the text prompt are input into a flux model integrated with the target lora model to generate the base map containing the target image and satisfying the text prompt and preset edge and position constraints.
[0019] Optionally, as in the aforementioned method, the step of correcting the composition and characters of the base image to obtain the intermediate image includes:
[0020] Perform feature mapping of the image structure and the person's posture in the base image to obtain base image features;
[0021] Extracting facial features from the base image to obtain facial features;
[0022] The text prompt, the base image, the base image features and the facial features are used as input data and processed by a preset image generation model and a target lora model corresponding to the target object to obtain the intermediate image.
[0023] Optionally, as in the aforementioned method, redrawing and correcting the details of the intermediate image to obtain a final output refined image as the display image of the target object includes:
[0024] Redrawing and correcting the details of the intermediate image to obtain a result image;
[0025] The face and picture details of the result image are processed to obtain the refined image, and the refined image is used as the display image of the target object.
[0026] Optionally, as in the aforementioned method, redrawing and correcting the details of the intermediate image to obtain a result image includes:
[0027] performing a first redrawing correction operation on the intermediate image to obtain a processed intermediate image;
[0028] The encoding of the processed intermediate image is input into the shallow space, and a second redrawing correction operation is performed on the unreasonable area in the processed intermediate image to obtain the result image.
[0029] Optionally, as in the aforementioned method, processing the face and picture details of the result image to obtain the refined image, and using the refined image as the display image of the target object includes:
[0030] Correcting faces that do not meet preset requirements in the result image using a face restoration algorithm to obtain a face-corrected result image;
[0031] The face correction result image is processed for detail clarification by a high-definition magnification algorithm to obtain the refined image, which is used as the display image of the target object.
[0032] According to another aspect of the embodiments of the present application, there is also provided an apparatus for generating a display image, including:
[0033] An acquisition module, used for acquiring a target image of a target object;
[0034] A base map generation module is configured to generate a base map containing the target image and constrained by the text prompt based on the target image and the text prompt, and using a target model corresponding to the target object, wherein the target model includes a preset large model and a target lora model corresponding to the target object;
[0035] an intermediate image generating module, configured to modify the composition of the base image and the characters in the base image to obtain an intermediate image;
[0036] The refined image generation module is used to redraw and correct the details of the intermediate image to obtain a final output refined image, and use the refined image as the display image of the target object.
[0037] According to another aspect of the embodiments of the present application, an electronic device is also provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein the memory is used to store computer programs; and the processor is used to execute the method steps in any of the above embodiments by running the computer program stored on the memory.
[0038] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the method steps in any of the above embodiments when run.
[0039] In an embodiment of the present application, a method is adopted in which a target image of a target object is obtained; based on the target image and text prompts, and through a target model corresponding to the target object, a base map containing the target image and constrained by the text prompts is generated, wherein the target model includes a preset large model and a target lora model corresponding to the target object; the composition of the base map and the characters in the base map are corrected to obtain an intermediate map; the intermediate map is redrawn and corrected in detail to obtain a final output refined map, and the refined map is used as a display image of the target object. Since the base map is obtained based on the target lora model corresponding to the target object, the target object in the generated base map can be made consistent with the actual object, and finally the target object in the refined map can be made consistent with the actual object, thereby achieving the technical effect of effectively improving the consistency between the image of the target object in the final display image and the actual image of the target object, thereby solving the problem in the related art that the large model cannot make the object consistent with the actual object in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 is a schematic diagram of a hardware environment for an optional method of generating a display image according to an embodiment of the present application;
[0043] Figure 2 is a flowchart of an optional method for generating a display image according to an embodiment of the present application;
[0044] Figure 3 This is a flow chart of another optional method for generating a display image according to an embodiment of the present application.
[0045] Figure 4 is a flowchart of another optional method for generating a display image according to an embodiment of the present application;
[0046] Figure 5 is a structural block diagram of an optional device for generating a display image according to an embodiment of the present application;
[0047] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0049] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0050] According to one aspect of the embodiment of the present application, a method for generating a display image is provided. Optionally, in this embodiment, the method for generating a display image can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal 1402 and a server 1404. Figure 1 As shown, server 1404 is connected to terminal 1402 via a network, and can be used to provide services (such as game services, application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 1404.
[0051] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, or a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal may be, but is not limited to, a PC, a mobile phone, a tablet computer, or the like.
[0052] The method for generating a display image according to the embodiment of the present application can be executed by a server, a terminal, or both. The method for generating a display image according to the embodiment of the present application can be executed by a terminal or a client installed thereon.
[0053] Taking the method of generating a display image in this embodiment executed by a server as an example, Figure 2 A method for generating a display image provided in an embodiment of the present application includes:
[0054] Step S202: Acquire a target image of the target object.
[0055] The method for generating a display image in this embodiment can be applied to scenarios where a corresponding image needs to be generated for a specific object, for example, generating a corresponding advertising image when the object is a commodity; generating a corresponding usage image when the object is a tool (e.g., a fire extinguisher), etc. It can also be applied to scenarios where images of other types of objects are generated. In this embodiment of the application, the method for generating a display image is described using the example of generating a corresponding advertising image when the object is a commodity. The method for generating a display image is also applicable to other types of objects, unless there is a conflict.
[0056] Specifically, it is possible to determine the object for which a corresponding image needs to be generated. Therefore, upon obtaining the requirement, a target image of the target object can be determined. The target object can be any actual product, such as sunscreen, food, or daily necessities. The target image can be obtained by cropping a picture containing the target object along the edges of the target object.
[0057] Step S204: Generate a base map containing the target image and constrained by the text prompt based on the target image and the text prompt, and through a target model corresponding to the target object, wherein the target model includes a preset large model and a target lora model corresponding to the target object.
[0058] That is, after obtaining the target image and text prompt, in this embodiment, a base map containing the target image and constrained by the text prompt is generated by using the target model corresponding to the target object. The target LoRa model in the target model can be a LoRa model trained using training samples containing the target object. Furthermore, the base map generated by the target model can keep the shape of the target object in the base map consistent with the target image.
[0059] As an optional implementation, the following steps can be used to generate a base map containing a target image and constrained by text prompts based on the target image and text prompts, and through a target model corresponding to the target object: among all candidate models, determine the target model corresponding to the target object according to the type of the target object; input the target image and text prompts into the target model to generate a base map containing a target image and constrained by the text prompts, wherein the base map includes the target image and the background image. After obtaining the target image, the target model can be determined among all candidate models based on the correspondence between the preset object type and the candidate model. Furthermore, all candidate models include a preset large model, but there are differences in the LoRa models integrated in the preset large model. The target LoRa model can be a LoRa layer added to the large model. The LoRa layer contains a small number of trainable parameters, which are intended to adapt to specific task requirements through fine-tuning without changing most of the weights of the large model. Furthermore, text prompts can be generated based on the requirements of the desired image. These text prompts can be textual words describing information such as objects in the image, such as: "bottles, indoors, blurry, cup, no humans, depth of field, blurry background, chair, table, plant, still life." Furthermore, the text prompt generation process can utilize any existing open-source method, such as WD Tagger, DeepseekVL, Florence2, or manual description, simply describing the desired content of the final image. After obtaining the target image and text prompts, the target image and text prompts can be input into the target model, which then generates a base map containing the target object according to the requirements of the text prompts. The text prompts can be used to generate a background image within the base map. That is, the base map contains both the target object image and the background image, and the position and size relationship between the target object image and the background image can be set based on the text prompts.
[0060] As an optional implementation, as in the above method, the basic lora model can be trained by the following method:
[0061] Step 2102: Acquire multiple pictures of the object placed on any object using a preset acquisition method.
[0062] Alternatively, you can grab pictures from the internet of various objects placed on arbitrary objects, such as staged photos, advertisements, posters, and close-ups. Furthermore, each of the multiple pictures is required to have high image clarity (for example, the product and background are clearly visible) and high image quality (for example, preferably photos taken with professional photography equipment, with image pixels higher than a preset pixel (for example, 12 million pixels, 20 million pixels, etc.)). The object type can be relatively common and easily recognized by the model automatically, such as lipstick, gift boxes, cups, plates, perfume, etc.
[0063] Step 2104 : Generate a description word corresponding to each picture, wherein the description word corresponding to each picture is used to describe the object in each picture.
[0064] Specifically, after acquiring the images, a description word corresponding to each image (i.e., prompt generation) can be generated as follows: for each image, a full sentence or single word can be used to describe the object in the image, such as: "bottles, indoors, blurry, cup, no humans, depth of field, blurry background, chair, table, plant, still life." This process can use any existing open source method, such as WD Tagger, DeepseekVL, Florence2, or manual description, as long as it can clearly describe the content of the image. This is not limited here.
[0065] Step 2106: Use multiple pictures and the description words corresponding to each picture as input data to train the initial LoRa model to obtain a basic LoRa model.
[0066] Specifically, after obtaining multiple pictures and the descriptive words corresponding to each picture, the multiple pictures and the descriptive words corresponding to each picture can be used as input data for training and input into the initial LoRa model to train the initial LoRa model and obtain the basic LoRa model.
[0067] By using the method of this embodiment, the initial LoRa model can be trained quickly to obtain a basic LoRa model.
[0068] As an optional implementation, as in the above method, the method further includes the following steps:
[0069] Step 2202: Obtain LoRa training samples for training the basic LoRa model.
[0070] Specifically, a LoRa training sample for training a basic LoRa model can be obtained from a preset training sample set. The basic LoRa model can be a LoRa model that has not been trained with images corresponding to the target object.
[0071] As an optional implementation, as in the aforementioned method, the following steps can be used to obtain LoRa training samples for training the basic LoRa model:
[0072] Step 2302 : Obtain a minimum edge clipping of the target object, wherein the minimum edge clipping is an image corresponding to the target object in the picture containing the target object.
[0073] That is, the image containing the target object can be cropped according to the edge of the target object to obtain the minimum edge cropping.
[0074] Step 2304: Determine at least one object input image containing the outline of the target object based on minimum edge clipping, at least one aspect ratio, and an image proportion range, wherein the aspect ratio is the ratio between the length and width of the object input image, and the image proportion range is the ratio of the area of the minimum circumscribed rectangular frame of the minimum edge clipping to the area of the object input image.
[0075] After the minimum edge clipping is obtained, the object input image can be determined according to the aspect ratio of the minimum circumscribed rectangular frame corresponding to the minimum edge clipping and the image proportion range.
[0076] Specifically, at least one aspect ratio and image proportion range may be determined by:
[0077] Calculate the area of the minimum bounding rectangle of the minimum edge clipping product At least one aspect ratio of the object input image can be preset. For example, when setting three aspect ratios: 600:800, 600:600, and 800:600; in addition, the minimum and maximum ranges of the minimum bounding box of the product to the entire image area (i.e., the image proportion range) can be set at the same time, for example: [1 / 16, 1 / 3]. The following loop is executed, and each loop obtains an object input image with a minimum bounding box at a random position and size in the object input image:
[0078] Loop 1: Select the image size of the object input image in turn [[600,800],[600,600],[800,600]];
[0079] Loop 2: Randomly select the minimum bounding rectangle that occupies the entire image area;
[0080] A) Assuming the image size is h, w, and the area ratio (i.e., any value within the image ratio range) is ratio, the area of the minimum bounding rectangle in the image is calculated as area product_in_image =h*w*ratio;
[0081] B) Assume that the actual length of the target object is h product_src , width is w product_src , area is area product , calculate the length of the minimum bounding rectangle in the object input image (h product_in_image ) and width (w product_in_image ) is as follows:
[0082]
[0083] C) Calculate the range of horizontal and vertical coordinates that can be randomly placed in the image: the horizontal coordinate range [0,ww product_in_image ], vertical coordinate range [0,hh product_in_image ], randomly generate a coordinate point in this range, scale the product image and place it to obtain the object input image.
[0084] Step 2306: Obtain a LoRa training sample based on the object input image, the basic LoRa model, and the sample requirement text description.
[0085] Specifically, after obtaining the object input image, the object input image can be processed according to the basic LoRa model and the sample requirement text description to obtain LoRa training samples. Exemplarily, the LoRa training samples can be obtained as follows: The object input image is passed to ControlNet (a module for multi-scale feature extraction) through an input image operation. The user-provided sample requirement text description is received as guidance information for the task of generating LoRa training samples. The object input image and the sample requirement text description are input to the ControlNet module, which extracts multi-scale feature maps. These feature maps capture the key structural information of the input image and provide details at different resolution levels. The sampler gradually generates images from noise based on the large model with LoRa fine-tuning parameters (i.e., the basic LoRa model), the feature maps, and the structural information provided by the object input image. During each denoising step, the sampler references the feature maps from ControlNet and the object input image itself to ensure that the generated image not only meets the requirements of the sample requirement text description but also accurately reflects the structural characteristics of the input object image. Finally, the latent space representation generated by the sampler is decoded into a specific image format to obtain candidate LoRa training samples. The final lora training samples can be obtained by further screening the candidate lora training samples, or can be all candidate lora training samples.
[0086] As an optional implementation, as in the above method, the following method can also be used to obtain LoRa training samples for training the basic LoRa model:
[0087] Step 2402: Acquire multiple pictures including the target object.
[0088] In this embodiment, the target object may be a person, an animal, or other object that needs to appear in the same image as the target object.
[0089] When the target object is a person, multiple pictures including the target object can be collected online, for example, 100 pictures including people. Furthermore, the multiple pictures including the target object are required to be of high quality (preferably professional photography), with the person, scene, and details clearly visible.
[0090] Step 2404 : When the target object includes the designated object, generate a combined image in which the target object is placed on the designated object.
[0091] That is to say, the above-mentioned target objects also include designated objects. The designated objects in this embodiment can be objects on which the target objects can be placed, such as sofas, tables, beds, various floors, benches, stones, piers, window sills, wooden stakes, etc.
[0092] When the target object includes a designated object, a composite image can be generated in which the target object is placed at a reasonable position on the designated object. Alternatively, the target object can be automatically or manually placed at a reasonable position in any of multiple images including the target object. Furthermore, the position and size of the target object in the image including the target object can be randomly set.
[0093] Step 2406: Obtain a LoRa training sample based on the combined image, the basic LoRa model, and the sample requirement text description.
[0094] Specifically, after obtaining the combined image, it can be processed according to the basic LoRa model and the sample requirement text description to obtain LoRa training samples. Exemplarily, the LoRa training samples can be obtained as follows: The combined image is passed to ControlNet (a module for multi-scale feature extraction) by inputting an image. The user-provided sample requirement text description is received as guidance information for the task of generating LoRa training samples. The combined image and the sample requirement text description are input to the ControlNet module, which extracts multi-scale feature maps. These feature maps capture key structural information of the input image and provide details at different resolution levels. Based on the large model with LoRa fine-tuning parameters (i.e., the basic LoRa model), the feature maps, and the structural information provided by the combined image, the sampler gradually generates images from noise. During each denoising step, the sampler references the feature maps from ControlNet and the combined image itself to ensure that the generated image meets the requirements of the sample requirement text description and accurately reflects the structural characteristics of the input combined image. Finally, the latent space representation generated by the sampler is decoded into a specific image format to obtain the final LoRa training samples. These are candidate LoRa training samples. The final lora training samples can be obtained by further screening the candidate lora training samples, or can be all candidate lora training samples.
[0095] Furthermore, after obtaining the above-mentioned candidate LoRa training samples, the above-mentioned candidate LoRa training samples can be screened, and the result images with poor effects (such as samples with poor fusion of target objects and target objects) are eliminated to obtain the final LoRa training samples.
[0096] In the process of obtaining LoRa training samples in steps 2302 to 2306, the input image is only input into ControlNet, and Canny (i.e., edge detection algorithm) is mainly used to constrain the target object itself from being too distorted.
[0097] In the process of obtaining LoRa training samples using steps 2402 to 2406, when the target object includes a person, the input image is not only input into ControlNet (Canny constrains the image content not to change significantly, and OpenPose constrains the person not to change significantly), but is also encoded and input into the sampler's latent space (constraining the entire image not to change significantly), so that the large model can modify only the target object placement area as much as possible to obtain a better fusion effect.
[0098] Step 2204: Generate a description word corresponding to each LORA training sample, wherein the description word corresponding to each LORA training sample is used to describe the object in the LORA training sample.
[0099] After obtaining the Lora training samples, in order to train the basic Lora model, it is also necessary to obtain the descriptive words corresponding to each Lora training sample. This can be achieved by using the following method: for each Lora training sample, use a whole sentence or word to describe the object and other information in each Lora training sample, such as: "solo, short_hair, shirt, black_hair, 1boy, closed_mouth, white_shirt, upper_body_focus, collared_shirt, blurry, black_eyes, depth_of_field, blurry_background, realistic, dress_shirt, white_with_blue_shades, light_skin, indoor_setting, warm_colors, yellow_ambient_lighting, warm_tone". This process can use any existing open source method, such as WDTagger, DeepseekVL, etc., or manual description, as long as it can clearly describe the content of the image, which is not limited here.
[0100] Step 2206: Use the LoRa training samples and the description words corresponding to each LoRa training sample as input data to train the basic LoRa model to obtain the target LoRa model.
[0101] Specifically, after obtaining all LoRa training samples and the descriptive words corresponding to each LoRa training sample, multiple LoRa training samples and the descriptive words corresponding to each LoRa training sample can be used as input data for training and input into the basic LoRa model to train the basic LoRa model and obtain the target LoRa model.
[0102] By using the method of this embodiment, the initial LoRa model can be trained quickly to obtain a basic LoRa model.
[0103] As an optional implementation, as in the aforementioned method, the descriptive words corresponding to each LoRa training sample can be generated by the following method: each LoRa training sample is described in a whole sentence or word format to obtain the descriptive words corresponding to each LoRa training sample.
[0104] Specifically, image description word generation can be to use a whole sentence or word description method for each LoRa training sample to describe the object in the image and other information. For example, when using word description, the description word corresponding to one of the LoRa training samples can be: "solo,short_hair,shirt,black_hair,1boy,closed_mouth,white_shirt,upper_body_focus,collared_shirt,blurry,black_eyes,depth_of_field,blurry_background,realistic,dress_shirt,white_with_blue_shades,light_skin,indoor_setting,warm_colors,yellow_ambient_lighting,warm_tone". This process can use any existing open source method, such as WD Tagger, DeepseekVL, etc., or manual description, just need to clearly describe the content of the image.
[0105] Furthermore, at the front of the descriptive words corresponding to the obtained LoRa training samples, the LoRa model wake-up word is added. It can be any artificial word, as long as it is not in the existing word library and has no specific meaning. It is used to identify the characteristics of the model, that is, all LoRa training samples in the training set have this wake-up word, indicating that the LoRa training samples have commonalities (the specific commonalities are understood by the training process model itself). For example, the above wake-up word can be written as: prsunscreen1, where pr is the abbreviation of the target object, sunscreen is the specific target object, and 1 is the number. In addition, other methods can be used to generate corresponding wake-up words, which are not limited here.
[0106] As an optional implementation, as in the aforementioned method, the following steps can be used to train the basic LORA model using the LORA training samples and the descriptive words corresponding to each LORA training sample as input data to obtain the target LORA model:
[0107] Step 2502: Use the LORA training samples and the description words corresponding to each LORA training sample as input data to train the initial LORA model to obtain a trained LORA model.
[0108] Specifically, the prepared LoRa training samples and the descriptors corresponding to each LoRa training sample are input into the model as input data. In this process, the LoRa training samples are used to generate feature representations, while the descriptors corresponding to each LoRa training sample provide guidance information to help the specified model learn how to generate or adjust the image content based on the given description. Optionally, an appropriate loss function (such as contrast loss) can be defined to measure the difference between the generated image features and the target descriptors. Use an optimization algorithm (such as Adam) to update the parameters in the initial LoRa model, gradually reduce the loss value, and improve model performance. Repeat the above process multiple times to iterate training.
[0109] Step 2504: Test the specified model using preset description words to obtain test results, where the specified model includes a preset large model and a trained lora model.
[0110] Specifically, after constructing a complete model structure including a preset large model and a trained lora model. This model combines the powerful expressive power of the large model and the fine-tuning effect of the lora model for specific tasks. In this embodiment, a set of preset descriptors is used to test the specified model including the preset large model and the trained lora model. For each preset descriptor, the specified model attempts to generate a corresponding image and compares it with the expected result. And by evaluating the quality of the generated image, a combination of automatic evaluation indicators (such as BLEU, CIDEr, etc.) and manual review can be used to comprehensively examine the performance of the specified model and obtain the test results corresponding to the specified model.
[0111] Step 2506: If the test results meet the preset requirements, the trained LoRa model is determined as the target LoRa model.
[0112] That is to say, if the test results show that the specified model can efficiently and accurately generate high-quality images based on the preset descriptors, the training is considered successful, and the current trained LoRa model is marked as the target LoRa model, which can be used for subsequent applications or deployments.
[0113] Step 2508: If the test result does not meet the preset requirements, continue training the trained lora model.
[0114] In other words, if the test results do not meet expectations, it means that there is still room for improvement in the specified model. At this time, you can adjust the training strategy based on the feedback, such as increasing the number of training rounds, modifying hyperparameters, expanding the training data set, etc., and then continue to train the trained LoRa model.
[0115] Through the cyclic iteration method disclosed in this embodiment, a target LoRa model can be finally obtained that can perform well on specific tasks and has strong generalization ability, so as to ensure that the generated target model can effectively capture and reproduce complex visual concepts.
[0116] Step S206: Correct the composition of the base image and the characters in the base image to obtain an intermediate image.
[0117] That is to say, in order to obtain the intermediate image, it is necessary to correct the composition of the base image and the characters in the base image. Correcting the composition of the base image may include but is not limited to correcting the size relationship and position between objects, and correcting the characters may include correcting abnormalities in the characters.
[0118] In step S208 , the intermediate image is redrawn and corrected in detail to obtain a final output refined image, which is used as a display image of the target object.
[0119] Specifically, after obtaining the intermediate image, further processing is required, and the details of the intermediate image (for example, image quality, texture, etc.) need to be redrawn and corrected to obtain a refined image, and the refined image is used as the final display image of the target object.
[0120] This embodiment obtains a base map based on the target lora model corresponding to the target object, so that the target object in the generated base map can be consistent with the actual object, and finally the target object in the refined image can be consistent with the actual object, thereby achieving the technical effect of effectively improving the consistency between the image of the target object in the final display image and the actual image of the target object, thereby solving the problem in the related art that the large model cannot make the object in the image consistent with the actual object.
[0121] like Figure 3 As shown, as an optional implementation, as in the aforementioned method, the target image and text prompt can be input into the target model to generate a base map containing the target image and constrained by the text prompt through the following steps:
[0122] Step S302 : Obtain edge detection information of the target image by executing an edge detection algorithm on the target image.
[0123] Specifically, we can use Canny edge detection in ControlNet as an edge detection algorithm to identify and constrain the shape of the target image's circumscribed edges and their relative position within the entire image. In addition to shape, Canny edge detection can also be used to determine the target image's specific position within the base image, allowing people or other elements to be appropriately positioned around the product.
[0124] Step S304: Input the target image, the edge detection information, and the text prompt into a flux model integrated with the target lora model to generate the base map that includes the target image and satisfies the text prompt and preset edge and position constraints.
[0125] Specifically, you can load a pre-trained Flux model (a model optimized for a specific task or style, which excels in certain areas (such as Asian face generation)). The Flux model provides a strong foundation for generating high-quality images that meet specific stylistic requirements. Then, you load a target LoRa model that has been fine-tuned for specific needs. The self-trained target LoRa model is used to further refine the generated results, such as ensuring that the generated initial image meets the requirements of the actual application scenario, including but not limited to correct placement and reasonable object combination.
[0126] Combining the constraints of text prompts and edge detection information, the flux model starts to work, identifying the area where the person and the target image should be located, completing the preliminary composition of the overall image, and obtaining the base map.
[0127] Through the method of this embodiment, based on the input text prompts and edge detection information constraints, the areas where people and target objects are required to be located can be automatically identified, and the composition generation of the overall image can be completed. At the same time, the target LoRa model can ensure that the target product is placed on a reasonable object and is not left suspended in the air.
[0128] like Figure 4 As shown, as an optional implementation, the composition and characters of the base image can be corrected by the following steps to obtain the intermediate image:
[0129] Step S402 : performing feature mapping of the image structure and the person's posture in the base image to obtain base image features.
[0130] Specifically, after obtaining the base map, the base map can be input into the ControlNet module for feature mapping. The specific execution content may include sending the fine-tuned base map into two different ControlNet sub-modules for control: 1. Input the fine-tuned base map into the Canny module: extract the first base map sub-feature that reflects the edge information of the fine-tuned base map, which is used to keep the composition structure of the entire image unchanged. 2. Input the fine-tuned base map into the OpenPose module: extract the second base map sub-feature that reflects the key points of the human body of the characters in the fine-tuned base map, which is used to ensure that the posture and position of the characters are correct and to prevent the generation of redundant or deformed characters. In this way, the base map features including the first base map sub-feature and the second base map sub-feature can be obtained. Therefore, Canny control can be used to ensure that the final image is consistent with the original image in terms of overall composition. Figure 1 The control network can be used to control the image structure and pose of the character, such as background layout and object position. OpenPose control is also used to prevent the addition of additional characters or limb errors (such as multiple hands or twisted legs) during the redrawing process. This is particularly suitable for product display images containing people. It is worth noting that ControlNet is a "conditional injection" mechanism that does not change the output distribution of the main model. Instead, it guides the sampling process through additional feature maps. In this embodiment, the simultaneous use of the Canny module and the OpenPose module facilitates the dual constraints of image structure and character pose.
[0131] Step S404: extract facial features from the base image to obtain facial features.
[0132] Specifically, the base image can be input into the FaceID module, which extracts the face embedding vector (faceembedding) and uses it as facial features. The facial features can be used to guide the model to prioritize the facial features of the face during the generation process. Since the Stable Diffusion v1.5 model performs poorly in generating Asian faces, using the base image from flux as the FaceID input can "inject" the prior knowledge of Asian faces from the initial processed base image generated and processed by flux into the generation process of the Stable Diffusion v1.5 model, thereby improving the quality of the generated faces.
[0133] In step S406, the text prompt, base image, base image features and facial features are used as input data and processed by a preset image generation model and a target lora model corresponding to the target object to obtain an intermediate image.
[0134] Specifically, load and run an image generation model based on, for example, the Stable Diffusion v1.5 model, and at the same time load and apply the target lora model of user-defined training for fine-tuning the generation effect. The Stable Diffusion v1.5 model provides basic image generation capabilities. The target lora model performs lightweight parameter fine-tuning on the image generation model to make the generation result more in line with a specific style, product type or regional preference (such as Asian race, a certain clothing style, etc.). In this embodiment, the target lora model only modifies part of the weight matrix of the image generation model, so the training cost is low and the deployment is convenient. Further, multiple different target lora models can be combined to superimpose different styles or attributes (for example: Asian face + shirt style + high-definition detail enhancement).
[0135] After obtaining the user-provided text prompts, base map, Canny and OpenPose feature maps (i.e., base map features) provided by ControlNet, and facial features extracted by FaceID in the aforementioned steps, the text prompts, base map, base map features, and facial features can be used as input data and processed through the preset image generation model and the target lora model corresponding to the target object to obtain the intermediate image.
[0136] The intermediate image obtained by the method of this embodiment can achieve the following beneficial effects: making the details in the image clearer, especially the facial features of the characters and the texture of clothing, etc.; adjusting the overall color tone of the image to make it closer to the style of the product image (such as bright, soft, high contrast, etc.); and correcting problems such as facial deformation and body disproportion that may exist in the base image.
[0137] As an optional implementation, the intermediate image can be redrawn and corrected in detail through the following steps to obtain a final output refined image, which can be used as the display image of the target object:
[0138] The intermediate image is redrawn and corrected in detail to obtain a result image. As an optional implementation, the intermediate image can be redrawn and corrected in detail to obtain a result image through the following steps:
[0139] A first redrawing and correction operation is performed on the intermediate image to obtain a processed intermediate image. The encoding of the processed intermediate image is input into the shallow space, and a second redrawing and correction operation is performed on the unreasonable areas in the processed intermediate image to obtain the final image. Specifically, a flux model-based workflow can be used to redraw and correct the details of the intermediate image. The encoding of the intermediate image is input into the latent space. By appropriately reducing the denoise parameter, the majority of the final image content can be derived from the intermediate image. The flux model automatically redraws and corrects unreasonable areas, including faces, hands, posture, and object details. The latent space is a low-dimensional space in which each point represents a potential representation of the original high-dimensional data (such as an image). Once the intermediate image is obtained, it is first converted into a vector or feature representation in the latent space through a pre-trained encoder. For the intermediate image, its encoding can capture key image features such as shape, color, and texture, so that these characteristics can be preserved in the subsequent generation process. Furthermore, in diffusion models or similar generative frameworks, denoising is a key step, which involves recovering a clear signal or image from noisy data. By adjusting the denoising parameter (i.e., the denoise parameter), the degree of similarity between the generated image and the input encoding can be controlled. When the denoise parameter is set low, less randomness is introduced during the generation of the new image, so the generated image is closer to the input encoding (i.e., the intermediate image), which allows the final result image to retain more content from the intermediate image.
[0140] The face and picture details of the result image are processed to obtain a refined image, and the refined image is used as the display image of the target object. As an optional implementation method, the face and picture details of the result image can be processed to obtain a refined image, and the refined image is used as the display image of the target object through the following steps: using a face repair algorithm, the face in the result image that does not meet the preset requirements is corrected to obtain a face-corrected result image; using a high-definition magnification algorithm, the face-corrected result image is processed to clarify the details to obtain a refined image, and the refined image is used as the display image of the target object. Specifically, the face repair algorithm can be used to correct faces that do not meet the preset requirements (for example, faces with messy features or that do not meet normal requirements); in addition, a high-definition magnification algorithm can be used to make the picture details clearer and richer.
[0141] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0142] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM (Read-Only Memory, Read-Only Memory) / RAM (Random Access Memory, Random Access Memory), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0143] According to another aspect of the embodiments of the present application, a device for generating a display image for implementing the above-mentioned method for generating a display image is also provided. Figure 5 is a structural block diagram of an optional device for generating a display image according to an embodiment of the present application, such as Figure 5 As shown, the device may include:
[0144] An acquisition module 51 is used to acquire a target image of a target object;
[0145] A base map generation module 52 is configured to generate a base map containing the target image and constrained by the text prompt based on the target image and the text prompt, and using a target model corresponding to the target object, wherein the target model includes a preset large model and a target lora model corresponding to the target object;
[0146] An intermediate image generation module 53 is used to modify the composition of the base image and the characters in the base image to obtain an intermediate image;
[0147] The refined image generation module 54 is used to redraw and correct the details of the intermediate image to obtain a final output refined image, and use the refined image as the display image of the target object.
[0148] It should be noted that the acquisition module 51 in this embodiment can be used to execute the above-mentioned step S202, the base map generation module 52 in this embodiment can be used to execute the above-mentioned step S204, the intermediate image generation module 53 in this embodiment can be used to execute the above-mentioned step S206, and the refined image generation module 54 in this embodiment can be used to execute the above-mentioned step S208.
[0149] Through the above module, the base map is obtained by obtaining the target lora model corresponding to the target object, so that the target object in the generated base map can be kept consistent with the actual object, and finally the target object in the refined image can be kept consistent with the actual object, thereby achieving the technical effect of effectively improving the consistency between the image of the target object in the final display image and the actual image of the target object, thereby solving the problem in the related technology that the large model cannot make the object in the image consistent with the actual object.
[0150] In addition to the above modules, the device in this embodiment may also include a module for executing any method in any of the above embodiments of the method for generating a display image.
[0151] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. Figure 1 The hardware environment shown can be implemented through software or hardware, wherein the hardware environment includes a network environment.
[0152] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned method of generating a display image is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0153] According to another embodiment of the present application, there is also provided an electronic device, including: Figure 6As shown, the electronic device may include: a processor 1501 , a communication interface 1502 , a memory 1503 and a communication bus 1504 , wherein the processor 1501 , the communication interface 1502 , and the memory 1503 communicate with each other via the communication bus 1504 .
[0154] Memory 1503, used for storing computer programs;
[0155] The processor 1501 is configured to execute the program stored in the memory 1503 to implement the following steps:
[0156] Step S202: Acquire a target image of the target object.
[0157] Step S204: Generate a base map containing the target image and constrained by the text prompt based on the target image and the text prompt, and through a target model corresponding to the target object, wherein the target model includes a preset large model and a target lora model corresponding to the target object.
[0158] Step S206: Correct the composition of the base image and the characters in the base image to obtain an intermediate image.
[0159] In step S208 , the intermediate image is redrawn and corrected in detail to obtain a final output refined image, which is used as a display image of the target object.
[0160] Optionally, in this embodiment, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. This communication bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, the figure shows only one thick line, but this does not imply that there is only one bus or only one type of bus. The communication interface is used for communication between the electronic device and other devices.
[0161] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0162] As an example, the memory 1503 may include, but is not limited to, the acquisition module 51, base image generation module 52, intermediate image generation module 53, and refined image generation module 54 of the apparatus for generating a display image. Furthermore, the memory 1503 may also include, but is not limited to, other modules and units of the apparatus for generating a display image, which will not be described in detail in this example.
[0163] The above-mentioned processor can be a general-purpose processor, which can include but is not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0164] An embodiment of the present application further provides a computer-readable storage medium, the storage medium including a stored program, wherein the method steps of the above method embodiment are executed when the program is run.
[0165] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0166] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0167] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0168] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0169] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0170] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the purpose of the solution provided in this embodiment.
[0171] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0172] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for generating a display image, characterized in that: include: Acquire a target image of a target object; According to the target image and the text prompt, and through a target model corresponding to the target object, a base map containing the target image and constrained by the text prompt is generated, wherein the target model includes a preset large model and a target lora model corresponding to the target object; Correcting the composition of the base image and the characters in the base image to obtain an intermediate image; The intermediate image is redrawn and corrected in detail to obtain a final output refined image, and the refined image is used as the display image of the target object.
2. The method according to claim 1, characterized in that The step of generating a base map including the target image and constrained by the text prompt based on the target image and the text prompt and using a target model corresponding to the target object includes: Determine the target model corresponding to the target object among all candidate models according to the type of the target object; The target image and the text prompt are input into the target model to generate a base map containing the target image and constrained by the text prompt, wherein the base map includes the target image and a background image.
3. The method according to claim 2, characterized in that The step of inputting the target image and the text prompt into the target model to generate a base map containing the target image and constrained by the text prompt includes: Obtaining edge detection information of the target image by executing an edge detection algorithm on the target image; The target image, the edge detection information and the text prompt are input into a flux model integrated with the target lora model to generate the base map containing the target image and satisfying the text prompt and preset edge and position constraints.
4. The method according to claim 1, wherein The step of correcting the composition and characters of the base image to obtain an intermediate image includes: Perform feature mapping of the image structure and the person's posture in the base image to obtain base image features; Extracting facial features from the base image to obtain facial features; The text prompt, the base image, the base image features and the facial features are used as input data and processed by a preset image generation model and a target lora model corresponding to the target object to obtain the intermediate image.
5. The method according to claim 1, characterized in that The detailed redrawing and correction of the intermediate image to obtain a final output refined image as the display image of the target object includes: Redrawing and correcting the details of the intermediate image to obtain a result image; The face and picture details of the result image are processed to obtain the refined image, and the refined image is used as the display image of the target object.
6. The method according to claim 5, characterized in that The step of redrawing and correcting the details of the intermediate image to obtain a result image includes: performing a first redrawing correction operation on the intermediate image to obtain a processed intermediate image; The encoding of the processed intermediate image is input into the shallow space, and a second redrawing correction operation is performed on the unreasonable area in the processed intermediate image to obtain the result image.
7. The method according to claim 5, characterized in that The processing of the face and picture details of the result image to obtain the refined image, and using the refined image as the display image of the target object includes: Correcting faces that do not meet preset requirements in the result image using a face restoration algorithm to obtain a face-corrected result image; The face correction result image is processed for detail clarification by a high-definition magnification algorithm to obtain the refined image, which is used as the display image of the target object.
8. A device for generating a display image, characterized in that: include: An acquisition module, used for acquiring a target image of a target object; A base map generation module is configured to generate a base map containing the target image and constrained by the text prompt based on the target image and the text prompt, and using a target model corresponding to the target object, wherein the target model includes a preset flux model and a target lora model corresponding to the target object; an intermediate image generating module, configured to modify the composition of the base image and the characters in the base image to obtain an intermediate image; The refined image generation module is used to redraw and correct the details of the intermediate image to obtain a final output refined image, and use the refined image as the display image of the target object.
9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 7 by running the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.