Display object image generation method and device, electronic equipment and storage medium
By generating the image of the exhibit through the target model and combining the preset large model with the target LoRa model, the problem that the large model cannot generate reasonable images is solved, and the reasonable display and background coexistence of the exhibit at any size are achieved, which improves the accuracy and efficiency of image generation.
Patent Information
- Application Number
- CN202510667469.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-12
AI Technical Summary
The large models in the existing technology cannot generate reasonable images corresponding to the exhibits. Especially when the appearance of the product is fixed, the generated image cannot be reasonably placed at any size and cannot coexist reasonably with the people in the background.
By determining the target model corresponding to the target exhibit, including the preset large model and the target LoRa model, the target exhibit image and text prompt are input to generate the target image, and the LoRa model is used for fine-tuning to generate a reasonable exhibit image.
It significantly improves the ability of generated target images to accurately display the target exhibits, solves the problem that large models cannot generate reasonable images, reduces dependence on professional designers, shortens the production cycle and reduces costs.
Smart Images

Figure CN120635231A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for generating an image of an exhibit, an electronic device, and a storage medium. Background Art
[0002] In the advertising business, various factors often limit the specific display materials (e.g., merchandise) required for promotion. Advertisers and creative professionals must adhere to a meticulous production process to ensure both quality and effectiveness. First, team members must develop creative ideas and formulate an advertising strategy based on the product's characteristics and market positioning. Next, designers create a visual plan based on this strategy, selecting elements such as color, layout, and imagery to ensure the ad's appeal and reach.
[0003] This process places a significant demand on manual effort. High-quality advertising production relies on the designer's professional skills and aesthetic judgment, which often necessitates lengthy composition work, as well as repeated revisions and proofreading. The advertising industry's high demands for creativity and design detail mean that even minor modifications can require significant manpower and time. Such a complex and demanding process naturally presents bottlenecks. Time and cost are the primary challenges of current advertising production. A shortage of specialized personnel often limits both the speed of production and the breadth of creative output. Furthermore, aesthetic fatigue is another issue: creative exhaustion can occur after advertising producers work on the same theme for extended periods of time.
[0004] In sharp contrast, the technology of text-based graphics (i.e., models that generate images based on text descriptions) based on large AI models has attracted widespread attention in recent years. Although AI-based advertising production also has problems, such as the generated images may not fully meet advertising standards or be not detailed enough, its advantage is that it reduces dependence on professional designers, allowing users without a design background to quickly generate visual content of a certain quality. It can also greatly shorten the production cycle, reduce costs, and to some extent provide unlimited creative possibilities. For problematic result images, only professionals need to make appropriate modifications, which is much less time-consuming and less dependent than a purely manual production process.
[0005] However, when the product's appearance is fixed, relying solely on a large model workflow has limited results. For example, relying solely on prompts (text prompts) cannot ensure that the generated product is consistent with the specified product, the product cannot be reasonably placed on an object at any size, and the product cannot coexist reasonably with people in the background.
[0006] Therefore, there is a problem in the related art that large models cannot generate reasonable images corresponding to the exhibits. Summary of the Invention
[0007] The present application provides a method and device for generating an image of an exhibit, an electronic device, and a storage medium to at least solve the problem in the related art that a large model cannot generate a reasonable image corresponding to the exhibit.
[0008] According to one aspect of an embodiment of the present application, a method for generating an image of an exhibit is provided, comprising:
[0009] Determine the target display object for which images need to be generated;
[0010] Determining a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit;
[0011] The target exhibit image and the text prompt are input into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and a background image.
[0012] Optionally, as in the aforementioned method, the method further includes:
[0013] Get the lora training samples used to train the basic lora model;
[0014] Generate a descriptive word corresponding to each LORA training sample, wherein the descriptive word corresponding to each LORA training sample is used to describe the object in the LORA training sample;
[0015] The lora training samples and the description words corresponding to each lora training sample are used as input data to train the basic lora model to obtain the target lora model.
[0016] Optionally, as in the aforementioned method, before obtaining the LoRa training samples for training the basic LoRa model, the method further includes:
[0017] Use a preset acquisition method to obtain multiple images of display objects placed on any object;
[0018] generating a description word corresponding to each image, wherein the description word corresponding to each image is used to describe an object in each image;
[0019] The initial LoRa model is trained using the multiple pictures and the description words corresponding to each picture as input data to obtain the basic LoRa model.
[0020] Optionally, as in the aforementioned method, obtaining LoRa training samples for training the basic LoRa model includes:
[0021] Obtaining a minimum edge clipping of the target exhibit, wherein the minimum edge clipping is an image corresponding to the target exhibit in a picture containing the target exhibit;
[0022] Determining at least one exhibit input image containing the outline of the target exhibit based on the minimum edge cropping, at least one aspect ratio, and an image proportion range, wherein the aspect ratio is the ratio between the length and width of the exhibit input image, and the image proportion range is the ratio of the area of the minimum circumscribed rectangular frame of the minimum edge cropping to the area of the exhibit input image;
[0023] The LoRa training sample is obtained according to the display input image, the basic LoRa model and the sample requirement text description.
[0024] Optionally, as in the aforementioned method, obtaining LoRa training samples for training the basic LoRa model includes:
[0025] Acquire multiple images including the target object;
[0026] In a case where the target object includes a designated object, generating a combined image of the target exhibit being placed on the designated object;
[0027] The LoRa training sample is obtained according to the combined image, the basic LoRa model and the sample requirement text description.
[0028] Optionally, as in the aforementioned method, generating a description word corresponding to each lora training sample includes:
[0029] Each of the Lora training samples is described in a whole sentence or word manner to obtain the description word corresponding to each of the Lora training samples.
[0030] Optionally, as in the aforementioned method, the training of the basic LORA model using the LORA training samples and the descriptive words corresponding to each LORA training sample as input data to obtain the target LORA model includes:
[0031] The LORA training samples and the description words corresponding to each LORA training sample are used as input data to train the initial LORA model to obtain a trained LORA model;
[0032] The specified model is tested by preset descriptive words to obtain a test result, wherein the specified model includes a preset large model and the trained lora model;
[0033] When the test result meets the preset requirements, the trained LORA model is determined as the target LORA model;
[0034] If the test result does not meet the preset requirement, continue to train the trained lora model.
[0035] According to another aspect of the embodiments of the present application, there is also provided a device for generating an image of an exhibit, comprising:
[0036] A first determination module is used to determine a target exhibit for which an image needs to be generated;
[0037] A second determining module is configured to determine a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit;
[0038] The image generation module is used to input the target exhibit image and text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and a background image.
[0039] According to another aspect of the embodiments of the present application, an electronic device is also provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein the memory is used to store computer programs; and the processor is used to execute the method steps in any of the above embodiments by running the computer program stored on the memory.
[0040] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the method steps in any of the above embodiments when run.
[0041] In an embodiment of the present application, a method for generating an exhibit image based on LoRa is adopted, wherein the method includes: determining a target exhibit for which an image needs to be generated; determining a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target LoRa model corresponding to the target exhibit; inputting the target exhibit image and a text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and a background image. Since the target image corresponding to the target exhibit is generated by the target model corresponding to the target exhibit, and the target model includes a preset large model and a target LoRa model corresponding to the target exhibit, it is possible to achieve a technical effect that the target image obtained can accurately display the target exhibit while being based on a large model solution, thereby solving the problem that the large model in the related art cannot generate a reasonable image corresponding to the exhibit. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 is a schematic diagram of a hardware environment for an optional method for generating an exhibit image according to an embodiment of the present application;
[0045] Figure 2 is a flow chart of an optional method for generating an image of an exhibit according to an embodiment of the present application;
[0046] Figure 3 is a flowchart of an optional method for generating an image of an exhibit according to another embodiment of the present application;
[0047] Figure 4 is a flowchart of an optional method for generating an image of an exhibit according to another embodiment of the present application;
[0048] Figure 5 is a flowchart of an optional method for generating an image of an exhibit according to another embodiment of the present application;
[0049] Figure 6is a flowchart of an optional method for generating an image of an exhibit according to another embodiment of the present application;
[0050] Figure 7 is a flowchart of an optional lora training sample generation method according to another embodiment of the present application;
[0051] Figure 8 is a structural block diagram of an optional display image generating device according to an embodiment of the present application;
[0052] Figure 9 This is a structural block diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0054] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, exhibit or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, exhibits or devices.
[0055] According to one aspect of the embodiment of the present application, a method for generating an image of an exhibit is provided. Optionally, in this embodiment, the above-mentioned method for generating an image of an exhibit can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal 1402 and a server 1404. Figure 1 As shown, server 1404 is connected to terminal 1402 via a network, and can be used to provide services (such as game services, application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 1404.
[0056] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, or a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal may be, but is not limited to, a PC, a mobile phone, a tablet computer, or the like.
[0057] The method for generating an exhibit image according to the embodiment of the present application can be executed by a server, a terminal, or both. The method for generating an exhibit image according to the embodiment of the present application can also be executed by a client installed on the terminal.
[0058] Taking the method of generating an exhibit image in this embodiment executed by a server as an example, Figure 2 A method for generating an image of an exhibit provided in an embodiment of the present application includes the following steps:
[0059] Step S202: Determine the target exhibit for which an image needs to be generated.
[0060] The display image generation method of this embodiment can be applied to scenarios where a corresponding image needs to be generated for a specific display object, for example, when the display object is a commodity, a corresponding advertising image needs to be generated; when the display object is a tool (e.g., a fire extinguisher), a corresponding usage image needs to be generated, etc. It can also be applied to scenarios where images of other types of displays need to be generated. In the embodiments of this application, the display image generation method is described using the example of generating a corresponding advertising image when the display object is a commodity. The display image generation method is also applicable to other types of displays, unless there is any conflict.
[0061] In general, it is possible to determine what kind of exhibits need to generate corresponding images. Therefore, once the requirements are obtained, the target exhibits can be determined. The target exhibits can be any actual product, such as sunscreen, food, daily necessities, etc.
[0062] Step S204: determining a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit.
[0063] Specifically, after obtaining the target exhibit, the target model can be determined from all candidate models based on the correspondence between the preset exhibit and the candidate models. Furthermore, all candidate models contain a preset large model, but there are differences in the LoRa models integrated in the preset large model. The target LoRa model can be a LoRa layer added to the large model. The LoRa layer contains a small number of trainable parameters, which is designed to adapt to specific task requirements through fine-tuning without changing most of the weights of the large model.
[0064] Step S206 : Input the target exhibit image and the text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and the background image.
[0065] Specifically, after determining the target object, an image of the target object can be obtained from the image containing the target object by methods such as cropping. Furthermore, text prompts can be generated based on the requirements of the generated image. These text prompts can be words that describe information such as objects in the image, such as "bottles, indoors, blurry, cup, no humans, depth of field, blurry background, chair, table, plant, still life." Furthermore, the text prompts can be generated using any existing open-source method, such as WD Tagger, DeepseekVL, Florence2, or manual description, as long as they clearly describe the content of the desired final image.
[0066] After obtaining the target exhibit image and the text prompt, these can be input into the target model. The target model can then generate a target image containing the target exhibit in accordance with the text prompt. The text prompt can be used to generate a background image within the target image. In other words, the target image contains both the target exhibit image and the background image, and the position and size relationship between the target exhibit image and the background image can be set based on the text prompt.
[0067] In this embodiment, a target image corresponding to the target exhibit is generated by a target model corresponding to the target exhibit, and the target model includes a preset large model and a target LoRa model corresponding to the target exhibit. Thus, while a solution based on the large model is achieved, the target LoRa model can significantly improve the technical effect that the final target image can accurately display the target exhibit, thereby solving the problem in the related art that the large model cannot generate a reasonable image corresponding to the exhibit.
[0068] like Figure 3 As shown, as an optional implementation, as in the above method, the basic lora model can be trained by the following method:
[0069] Step S302: A preset acquisition method is used to acquire multiple pictures of the exhibit being placed on any object.
[0070] Alternatively, you can grab images from the internet of various exhibits placed on arbitrary objects, such as staged photos, advertisements, posters, and close-ups. Furthermore, each of the multiple images is required to have high image clarity (for example, the product and background are clearly visible) and high image quality (for example, preferably photographed with professional photography equipment, with image pixels higher than a preset pixel count (for example, 12 megapixels, 20 megapixels, etc.)). The exhibits can also be common items that are easily automatically recognized by the model, such as lipstick, gift boxes, cups, plates, perfume, etc.
[0071] Step S304 : generating description words corresponding to each picture, wherein the description words corresponding to each picture are used to describe the objects in each picture.
[0072] Specifically, after acquiring the images, a description word corresponding to each image (i.e., prompt generation) can be generated as follows: for each image, a full sentence or single word can be used to describe the objects and other information in the image, such as: "bottles, indoors, blurry, cup, no humans, depth of field, blurry background, chair, table, plant, still life." This process can use any existing open source method, such as WD Tagger, DeepseekVL, Florence2, or manual description, as long as it can clearly describe the content of the image. This is not limited here.
[0073] Step S306: Use multiple pictures and the description words corresponding to each picture as input data to train the initial LoRa model to obtain a basic LoRa model.
[0074] Specifically, after obtaining multiple pictures and the descriptive words corresponding to each picture, the multiple pictures and the descriptive words corresponding to each picture can be used as input data for training and input into the initial LoRa model to train the initial LoRa model and obtain the basic LoRa model.
[0075] By using the method of this embodiment, the initial LoRa model can be trained quickly to obtain a basic LoRa model.
[0076] like Figure 4 As shown, as an optional implementation, as the above method, the method further includes the following steps:
[0077] Step S402: Obtain LoRa training samples for training the basic LoRa model.
[0078] Specifically, a LoRa training sample for training a basic LoRa model can be obtained from a preset training sample set. The basic LoRa model can be a LoRa model that has not been trained with images corresponding to the target display object.
[0079] like Figure 5 As shown, as an optional implementation, as in the aforementioned method, the following steps can be used to obtain LoRa training samples for training the basic LoRa model:
[0080] Step S502 : obtaining a minimum edge clipping of the target exhibit, wherein the minimum edge clipping is an image corresponding to the target exhibit in the picture containing the target exhibit.
[0081] That is, the picture containing the target exhibit may be cropped according to the edge of the target exhibit to obtain minimum edge cropping.
[0082] Step S504 : Determine at least one exhibit input image containing the outline of the target exhibit based on minimum edge clipping, at least one aspect ratio, and an image proportion range, wherein the aspect ratio is the ratio between the length and width of the exhibit input image, and the image proportion range is the ratio of the area of the minimum bounding rectangle of the minimum edge clipping to the area of the exhibit input image.
[0083] After the minimum edge clipping is obtained, the display object input image can be determined according to the aspect ratio of the minimum circumscribed rectangular frame corresponding to the minimum edge clipping and the image proportion range.
[0084] Specifically, at least one aspect ratio and image proportion range may be determined by:
[0085] Calculate the area of the minimum bounding rectangle of the minimum edge clipping product At least one aspect ratio of the display input image can be preset, for example, when setting three aspect ratios: 600:800, 600:600, 800:600; in addition, the minimum and maximum ranges of the minimum bounding box of the product to the entire image area (i.e., the image proportion range) can be set at the same time, for example: [1 / 16, 1 / 3]. The following loop is executed, and each loop obtains an display input image with a minimum bounding box at a random position and size in the display input image:
[0086] Loop 1: Select the image size of the display input image in turn [[600,800],[600,600],[800,600]];
[0087] Loop 2: Randomly select the minimum bounding rectangle that occupies the entire image area;
[0088] A) Assuming the image size is h, w, and the area ratio (i.e., any value within the image ratio range) is ratio, the area of the minimum bounding rectangle in the image is calculated as area product_in_image =h*w*ratio;
[0089] B) Assume that the actual length of the target display is h product_src , width is w product_src , area is area product , calculate the length of the minimum bounding rectangle in the display input image (h product_in_image ) and width (w product_in_image ) is as follows:
[0090]
[0091] C) Calculate the range of horizontal and vertical coordinates that can be randomly placed in the image: the horizontal coordinate range [0,ww product_in_image ], vertical coordinate range [0,hh product_in_image ], randomly generate a coordinate point in this range, scale the product image and place it to obtain the display input image.
[0092] Step S506: Obtain a LoRa training sample based on the display input image, the basic LoRa model, and the sample requirement text description.
[0093] Specifically, after obtaining the display input image, the display input image can be processed according to the basic LoRa model and the sample requirement text description to obtain LoRa training samples. For example, the LoRa training samples can be obtained as follows: Figure 7As shown, the product image (i.e., the aforementioned display input image) is preprocessed and passed to the ControlNet (a module for multi-scale feature extraction) through the input image operation. The user-provided text description of the sample requirement (i.e., the prompt in the image) is used as guidance information for the task of generating LoRa training samples. The display input image and the sample requirement text description are input to the ControlNet module, which extracts multi-scale feature maps. These feature maps capture the key structural information of the input image and provide details at different resolution levels. The sampler gradually generates images from noise based on the large model with LoRa fine-tuning parameters (i.e., the base LoRa model), the feature maps, and the structural information provided by the display input image. During each denoising step, the sampler references the feature maps from the ControlNet and the display input image itself to ensure that the generated image not only meets the requirements of the sample requirement text description but also accurately reflects the structural characteristics of the input display image. Finally, the latent space representation generated by the sampler is decoded into a specific image format to obtain candidate LoRa training samples. The final lora training samples can be obtained by further screening the candidate lora training samples, or can be all candidate lora training samples.
[0094] like Figure 6 As shown, as an optional implementation, as in the above method, the lora training samples for training the basic lora model can also be obtained by the following method:
[0095] Step S602: Acquire multiple pictures including the target object.
[0096] In this embodiment, the target object may be a person, an animal, or other object that needs to appear in the same image as the target display object.
[0097] When the target object is a person, multiple pictures including the target object can be collected online, for example, 100 pictures including people. Furthermore, the multiple pictures including the target object are required to be of high quality (preferably professional photography), with the person, scene, and details clearly visible.
[0098] Step S604 : When the target object includes the designated object, a combined image is generated in which the target exhibit is placed on the designated object.
[0099] That is to say, the above-mentioned target objects also include designated objects. The designated objects in this embodiment can be objects on which the target exhibits can be placed, such as sofas, tables, beds, various floors, benches, stones, piers, window sills, wooden stakes, etc.
[0100] If the target object includes a designated object, a composite image can be generated in which the target exhibit is placed at a reasonable location on the designated object. Alternatively, the target exhibit can be automatically or manually placed at a reasonable location in any of multiple images that include the target object. Furthermore, the position and size of the target exhibit in the image that includes the target object can be randomly set.
[0101] Step S606: Obtain a LoRa training sample based on the combined image, the basic LoRa model, and the sample requirement text description.
[0102] Specifically, after obtaining the combined image, the combined image can be processed according to the basic LoRa model and the sample requirement text description to obtain a LoRa training sample. For example, the LoRa training sample can be obtained as follows: Figure 7 As shown, the product image (i.e., the combined image) is preprocessed and passed to the ControlNet (a module for multi-scale feature extraction) through the input image operation. The user-provided text description of the sample requirements is used as guidance for the task of generating LoRa training samples. The combined image and the sample requirements text description are input to the ControlNet module, which extracts multi-scale feature maps. These feature maps capture the key structural information of the input image and provide details at different resolution levels. The sampler gradually generates images from noise based on the large model with LoRa fine-tuning parameters (i.e., the base LoRa model), the feature maps, and the structural information provided by the combined image. During each denoising step, the sampler references the feature maps from the ControlNet and the information from the combined image itself to ensure that the generated image not only meets the requirements of the sample requirements text description but also accurately reflects the structural characteristics of the input combined image. Finally, the latent space representation generated by the sampler is decoded into a specific image format to obtain the final LoRa training sample. The final LoRa training sample can be obtained by further screening the candidate LoRa training samples or by combining all candidate LoRa training samples.
[0103] Furthermore, after obtaining the above-mentioned candidate LoRa training samples, the above-mentioned candidate LoRa training samples can be screened, and the result images with poor effects (such as samples in which the target display and the target object are not well integrated) are eliminated to obtain the final LoRa training samples.
[0104] In the process of obtaining LoRa training samples in steps S502 to S506, the input image is only input into ControlNet, and Canny (i.e., edge detection algorithm) is mainly used to constrain the target display object itself from being too distorted.
[0105] In the process of obtaining LoRa training samples using steps S602 to S606, when the target object includes a person, the input image is not only input into ControlNet (Canny constrains the image content not to change significantly, and OpenPose constrains the person not to change significantly), but is also encoded and input into the sampler's latent space (constraining the entire image not to change significantly), so that the large model can modify only the target display placement area as much as possible to obtain a better fusion effect.
[0106] Step S404: Generate a description word corresponding to each LORA training sample, wherein the description word corresponding to each LORA training sample is used to describe the object in the LORA training sample.
[0107] After obtaining the Lora training samples, in order to train the basic Lora model, it is also necessary to obtain the descriptive words corresponding to each Lora training sample. This can be achieved by using the following method: for each Lora training sample, use a whole sentence or word to describe the object and other information in each Lora training sample, such as: "solo, short_hair, shirt, black_hair, 1boy, closed_mouth, white_shirt, upper_body_focus, collared_shirt, blurry, black_eyes, depth_of_field, blurry_background, realistic, dress_shirt, white_with_blue_shades, light_skin, indoor_setting, warm_colors, yellow_ambient_lighting, warm_tone". This process can use any existing open source method, such as WDTagger, DeepseekVL, etc., or manual description, as long as it can clearly describe the content of the image, which is not limited here.
[0108] Step S406: Use the LoRa training samples and the description words corresponding to each LoRa training sample as input data to train the basic LoRa model to obtain the target LoRa model.
[0109] Specifically, after obtaining all LoRa training samples and the descriptive words corresponding to each LoRa training sample, multiple LoRa training samples and the descriptive words corresponding to each LoRa training sample can be used as input data for training and input into the basic LoRa model to train the basic LoRa model and obtain the target LoRa model.
[0110] By using the method of this embodiment, the initial LoRa model can be trained quickly to obtain a basic LoRa model.
[0111] As an optional implementation, as in the aforementioned method, the descriptive words corresponding to each LoRa training sample can be generated by the following method: each LoRa training sample is described in a whole sentence or word format to obtain the descriptive words corresponding to each LoRa training sample.
[0112] Specifically, image description word generation can be to use a whole sentence or word description for each LoRa training sample to describe the objects in the image and other information. For example, when using word description, the description word corresponding to one of the LoRa training samples can be: "solo,short_hair,shirt,black_hair,1boy,closed_mouth,white_shirt,upper_body_focus,collared_shirt,blurry,black_eyes,depth_of_field,blurry_background,realistic,dress_shirt,white_with_blue_shades,light_skin,indoor_setting,warm_colors,yellow_ambient_lighting,warm_tone". This process can use any existing open source method, such as WD Tagger, DeepseekVL, etc., or manual description, as long as the content of the image is clearly described.
[0113] Furthermore, at the front of the descriptive words corresponding to the obtained LoRa training samples, a LoRa model wake-up word is added. It can be any artificial word, as long as it is not in the existing word library and has no specific meaning. It is used to identify the characteristics of the model, that is, all LoRa training samples in the training set have this wake-up word, indicating that the LoRa training samples have commonalities (the specific commonalities are understood by the training process model itself). For example, the above wake-up word can be written as: prsunscreen1, where pr is the abbreviation of the target display, sunscreen is the specific target display, and 1 is the number. In addition, other methods can be used to generate corresponding wake-up words, which are not limited here.
[0114] As an optional implementation, as in the aforementioned method, the following steps can be used to train the basic LORA model using the LORA training samples and the descriptive words corresponding to each LORA training sample as input data to obtain the target LORA model:
[0115] Step 1002: Using the LORA training samples and the description words corresponding to each LORA training sample as input data, the initial LORA model is trained to obtain a trained LORA model.
[0116] Specifically, the prepared LoRa training samples and the descriptors corresponding to each LoRa training sample are input into the model as input data. In this process, the LoRa training samples are used to generate feature representations, while the descriptors corresponding to each LoRa training sample provide guidance information to help the specified model learn how to generate or adjust the image content based on the given description. Optionally, an appropriate loss function (such as contrast loss) can be defined to measure the difference between the generated image features and the target descriptors. Use an optimization algorithm (such as Adam) to update the parameters in the initial LoRa model, gradually reduce the loss value, and improve model performance. Repeat the above process multiple times to iterate training.
[0117] Step 1004: Test the specified model using preset description words to obtain test results, wherein the specified model includes a preset large model and a trained lora model.
[0118] Specifically, after constructing a complete model structure including a preset large model and a trained lora model. This model combines the powerful expressive power of the large model and the fine-tuning effect of the lora model for specific tasks. In this embodiment, a set of preset descriptors is used to test the specified model including the preset large model and the trained lora model. For each preset descriptor, the specified model attempts to generate a corresponding image and compares it with the expected result. And by evaluating the quality of the generated image, a combination of automatic evaluation indicators (such as BLEU, CIDEr, etc.) and manual review can be used to comprehensively examine the performance of the specified model and obtain the test results corresponding to the specified model.
[0119] Step 1006: If the test results meet the preset requirements, the trained LoRa model is determined as the target LoRa model.
[0120] That is to say, if the test results show that the specified model can efficiently and accurately generate high-quality images based on the preset description words, the training is considered successful, and the current trained LoRa model is marked as the target LoRa model, which can be used for subsequent applications or deployments.
[0121] Step 1008: If the test result does not meet the preset requirements, continue training the trained lora model.
[0122] In other words, if the test results do not meet expectations, it means that there is still room for improvement in the specified model. At this time, you can adjust the training strategy based on the feedback, such as increasing the number of training rounds, modifying hyperparameters, expanding the training data set, etc., and then continue to train the trained LoRa model.
[0123] Through the cyclic iteration method disclosed in this embodiment, a target LoRa model can be finally obtained that can perform well on specific tasks and has strong generalization ability, so as to ensure that the generated target model can effectively capture and reproduce complex visual concepts.
[0124] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0125] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software display, which is stored in a storage medium (such as ROM (Read-Only Memory, Read-Only Memory) / RAM (Random Access Memory, Random Access Memory), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the method described in each embodiment of the present application.
[0126] According to another aspect of the embodiments of the present application, there is also provided an exhibit image generating device for implementing the above-mentioned exhibit image generating method. Figure 8 is a structural block diagram of an optional display image generating device according to an embodiment of the present application, such as Figure 8 As shown, the device may include:
[0127] A first determining module 81 is used to determine a target exhibit for which an image needs to be generated;
[0128] The second determining module 82 is used to determine a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit;
[0129] The image generation module 83 is used to input the image of the target exhibit and the text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and the background image.
[0130] It should be noted that the first determination module 81 in this embodiment can be used to execute the above step S202, the second determination module 82 in this embodiment can be used to execute the above step S204, and the image generation module 83 in this embodiment can be used to execute the above step S206.
[0131] Through the above module, the target image corresponding to the target exhibit is generated by the target model corresponding to the target exhibit, and the target model includes a preset large model and a target LoRa model corresponding to the target exhibit. Therefore, while the solution is based on the large model, the target LoRa model can significantly improve the technical effect that the final target image can accurately display the target exhibit, thereby solving the problem in the related art that the large model cannot generate a reasonable image corresponding to the exhibit.
[0132] In addition to the above modules, the device in this embodiment may also include a module for executing any method in any of the above embodiments of the method for generating an image of an exhibit.
[0133] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. Figure 1 The hardware environment shown can be implemented through software or hardware, wherein the hardware environment includes a network environment.
[0134] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned method for generating an exhibit image is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0135] According to another embodiment of the present application, there is also provided an electronic device, including: Figure 9 As shown, the electronic device may include: a processor 1501 , a communication interface 1502 , a memory 1503 and a communication bus 1504 , wherein the processor 1501 , the communication interface 1502 , and the memory 1503 communicate with each other via the communication bus 1504 .
[0136] Memory 1503, used for storing computer programs;
[0137] The processor 1501 is configured to execute the program stored in the memory 1503 to implement the following steps:
[0138] Step S202: Determine the target exhibit for which an image needs to be generated.
[0139] Step S204: determining a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit.
[0140] Step S206 : Input the target exhibit image and the text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and the background image.
[0141] Optionally, in this embodiment, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. This communication bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, the figure shows only one thick line, but this does not imply that there is only one bus or only one type of bus. The communication interface is used for communication between the electronic device and other devices.
[0142] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0143] As an example, the memory 1503 may include, but is not limited to, the first determination module 81, the second determination module 82, and the image generation module 83 of the display image generation apparatus. Furthermore, the memory 1503 may also include, but is not limited to, other modules and units of the display image generation apparatus, which will not be described in detail in this example.
[0144] The above-mentioned processor can be a general-purpose processor, which can include but is not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0145] An embodiment of the present application further provides a computer-readable storage medium, the storage medium including a stored program, wherein the method steps of the above method embodiment are executed when the program is run.
[0146] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0147] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0148] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent exhibits, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software exhibit. The computer software exhibit is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.
[0149] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0151] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the purpose of the solution provided in this embodiment.
[0152] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0153] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for generating an image of an exhibit, characterized in that: include: Determine the target display object for which images need to be generated; Determining a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit; The target exhibit image and the text prompt are input into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and a background image.
2. The method according to claim 1, characterized in that The method further comprises: Get the lora training samples used to train the basic lora model; Generate a descriptive word corresponding to each LORA training sample, wherein the descriptive word corresponding to each LORA training sample is used to describe the object in the LORA training sample; The basic LORA model is trained using the LORA training samples and the description words corresponding to each LORA training sample as input data to obtain the target LORA model.
3. The method according to claim 2, characterized in that Before obtaining the lora training samples for training the basic lora model, the method further includes: Use a preset acquisition method to obtain multiple images of display objects placed on any object; generating a description word corresponding to each image, wherein the description word corresponding to each image is used to describe an object in each image; The initial LoRa model is trained using the multiple pictures and the description words corresponding to each picture as input data to obtain the basic LoRa model.
4. The method according to claim 2, characterized in that The obtaining of LoRa training samples for training the basic LoRa model includes: Obtaining a minimum edge clipping of the target exhibit, wherein the minimum edge clipping is an image corresponding to the target exhibit in a picture containing the target exhibit; Determining at least one exhibit input image containing the outline of the target exhibit based on the minimum edge cropping, at least one aspect ratio, and an image proportion range, wherein the aspect ratio is the ratio between the length and width of the exhibit input image, and the image proportion range is the ratio of the area of the minimum circumscribed rectangular frame of the minimum edge cropping to the area of the exhibit input image; The LoRa training sample is obtained according to the display input image, the basic LoRa model and the sample requirement text description.
5. The method according to claim 2, characterized in that The obtaining of LoRa training samples for training the basic LoRa model includes: Acquire multiple images including the target object; In a case where the target object includes a designated object, generating a combined image of the target exhibit being placed on the designated object; The LoRa training sample is obtained according to the combined image, the basic LoRa model and the sample requirement text description.
6. The method according to claim 2, characterized in that Generating a description word corresponding to each lora training sample includes: Each lora training sample is described in a whole sentence or word manner to obtain the description word corresponding to each lora training sample.
7. The method according to any one of claims 2 to 6, characterized in that The method of training the basic LORA model using the LORA training samples and the descriptive words corresponding to each LORA training sample as input data to obtain the target LORA model includes: The LORA training samples and the description words corresponding to each LORA training sample are used as input data to train the initial LORA model to obtain a trained LORA model; The specified model is tested by preset descriptive words to obtain a test result, wherein the specified model includes a preset large model and the trained lora model; When the test result meets the preset requirements, the trained LORA model is determined as the target LORA model; If the test result does not meet the preset requirement, continue to train the trained lora model.
8. A device for generating an image of an exhibit, characterized in that: include: A first determination module is used to determine a target exhibit for which an image needs to be generated; A second determining module is configured to determine a target model corresponding to the target exhibit, wherein the target model includes a preset large model and a target lora model corresponding to the target exhibit; The image generation module is used to input the target exhibit image and text prompt into the target model to generate a target image corresponding to the target exhibit, wherein the target image includes the target exhibit image and a background image.
9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 7 by running the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.