Image generation method and system, electronic equipment and storage medium

By converting user inputs to format-compatible data and retrieving high-quality samples from a database, the method addresses input mismatches, improving image generation efficiency and quality.

CN120318351APending Publication Date: 2025-07-15ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358220.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When generating images, the existing image processing model has inefficient generation efficiency because the user input does not meet the input conditions.

Method used

Convert the initial input information into target input information, match it with the image processing model, and retrieve high-quality input information samples in the database for analysis to generate scene images.

Benefits of technology

The efficiency of image generation is improved and the problem of low generation efficiency is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318351A_ABST
    Figure CN120318351A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and system, electronic equipment and a storage medium, and relates to the technical field of large model technology and image generation. The method comprises the following steps: acquiring initial input information; converting the initial input information into target input information; in a database, at least one input information sample is retrieved, the similarity between the input information sample and the target input information is larger than a similarity threshold value, the database at least comprises information samples of multiple modes, and the information samples are used for representing attribute information of a display object sample and attribute information of a scene where the display object sample is displayed; the input information sample is analyzed through the image processing model, a scene image is obtained, and the scene image is used for representing the display result of the object to be displayed in the target scene in which the object to be displayed is displayed. The technical problem of low image generation efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of large model technology and image generation technology. Specifically, it relates to a method, system, electronic device, and storage medium for generating images. Background Art

[0002] Currently, when using an image processing model to generate an image, the user's input is usually directly used as the input information of the image processing model. However, the input information of the image processing model (especially involving multi-modal data, such as images and texts) will affect the quality and efficiency of image generation.

[0003] However, in the above method, the user's input usually does not meet the input conditions of the image processing model, and unreasonable input information will lead to low efficiency of the image processing model in generating images, thus there is a problem of low image generation efficiency.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present application provide a method, system, electronic device, and storage medium for generating images, so as to at least solve the technical problem of low image generation efficiency.

[0006] According to one aspect of the embodiments of the present application, a method for generating an image is provided. The method may include: obtaining initial input information, where the initial input information is used to represent at least one object to be displayed; converting the initial input information into target input information, where the information format of the target input information matches that of the image processing model; retrieving in a database at least one input information sample whose similarity to the target input information is greater than a similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; using the image processing model to analyze the input information samples to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0007] According to another aspect of the embodiments of the present application, a method for generating an image is provided. The method may include: in response to an input instruction on an operation interface, displaying initial input information on the operation interface, where the initial input information is used to represent at least one object to be displayed; in response to an image generation instruction on the operation interface, displaying a scene image corresponding to the initial input information on the operation interface, where the scene image is used to represent the display result of the object to be displayed in the initial input information in a target scene, the scene image is obtained by analyzing an input information sample using an image processing model, the input information sample is an information sample in a database whose similarity to the target input information is greater than a similarity threshold, the database includes at least multiple modalities of information samples, the information sample is used to represent the attribute information of a display object sample and the attribute information of the scene in which the display object sample is displayed, the information quality of the input information sample is higher than that of the target input information, the target input information is obtained by converting the initial input information, and the information format of the target input information matches the image processing model.

[0008] According to another aspect of the embodiments of the present application, a method for generating an image is provided. The method may include: obtaining initial input information from an e-commerce platform, where the initial input information is used to represent at least one product to be displayed; converting the initial input information into target input information, where the information format of the target input information matches the image processing model; retrieving at least one input information sample in a database whose similarity to the target input information is greater than a similarity threshold, where the database includes at least multiple modalities of information samples, the information sample is used to represent the attribute information of a display product sample and the attribute information of the scene in which the display product sample is displayed, the information quality of the input information sample is higher than that of the target input information; analyzing the input information sample using an image processing model to obtain a scene image, where the scene image is used to represent the display result of the product to be displayed in a target scene; and sending the scene image to the e-commerce platform.

[0009] According to another aspect of the embodiments of the present application, a method for generating an image is provided. The method may include: obtaining initial input information by invoking a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter is the initial input information, and the initial input information is used to represent at least one object to be displayed; converting the initial input information into target input information, where the information format of the target input information matches that of the image processing model; retrieving in a database at least one input information sample whose similarity to the target input information is greater than a similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the sample of the displayed object and the attribute information of the scene where the sample of the displayed object is located, and the information quality of the input information sample is higher than that of the target input information; analyzing the input information sample by using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is located; and outputting the scene image by invoking a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the scene image.

[0010] According to another aspect of the embodiments of the present application, an image generation system is provided. The system may include: a client for uploading initial input information, where the initial input information is used to represent at least one object to be displayed; a server for converting the initial input information into target input information, where the information format of the target input information matches that of the image processing model; retrieving in a database at least one input information sample whose similarity to the target input information is greater than a similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the sample of the displayed object and the attribute information of the scene where the sample of the displayed object is located, and the information quality of the input information sample is higher than that of the target input information; analyzing the input information sample by using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is located; and sending the scene image to the client for display.

[0011] According to another aspect of the embodiments of the present application, a computing device is further provided, including: a memory storing an executable program; a processor for running the program, where when the program runs, it executes the methods in the various embodiments of the present application.

[0012] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; a processor connected to the memory through a bus for running the program, where when the program runs, it executes the methods in the various embodiments of the present application.

[0013] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, which includes a stored executable program. When the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present application.

[0014] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, which implements the methods in the various embodiments of the present application when executed by a processor.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, which implements the methods in the various embodiments of the present application when executed by a processor.

[0016] According to another aspect of the embodiments of the present application, there is also provided a computer program, which implements the methods in the various embodiments of the present application when executed by a processor.

[0017] In the embodiments of the present application, initial input information is obtained, where the initial input information is used to represent at least one object to be displayed; the initial input information is converted into target input information, where the information format of the target input information matches that of the image processing model; in the database, at least one input information sample whose similarity to the target input information is greater than the similarity threshold is retrieved. The database includes at least information samples of multiple modalities. The information samples are used to represent the attribute information of the display object samples and the attribute information of the scene where the display object samples are displayed. The information quality of the input information samples is higher than that of the target input information; the input information samples are analyzed by using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene. That is to say, in the embodiments of the present application, the initial input information is converted into target input information, and the information format of the target input information matches that of the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model and thus affect the efficiency of the scene image generated by the image processing model. Therefore, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0018] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Description of the Drawings

[0019] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a schematic diagram of an application scenario of a method for generating an image according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a method for generating an image according to an embodiment of the present application;

[0022] Figure 3 is a flowchart of another method for generating an image according to an embodiment of the present application;

[0023] Figure 4 is a flowchart of another method for generating an image according to an embodiment of the present application;

[0024] Figure 5 is a flowchart of yet another method for generating an image according to an embodiment of the present application;

[0025] Figure 6 is a schematic diagram of a recommended mode for indoor studio shooting in homeworks according to an embodiment of the present application;

[0026] Figure 7 is a schematic diagram of a custom mode for indoor studio shooting in homeworks according to an embodiment of the present application;

[0027] Figure 8 is a schematic diagram of an input image according to an embodiment of the present application;

[0028] Figure 9 is a schematic diagram of a generated image result in a recommended mode according to an embodiment of the present application;

[0029] Figure 10 is a schematic diagram of a generated image result in a custom mode according to an embodiment of the present application;

[0030] Figure 11 is a flowchart of a user input enhancement method according to an embodiment of the present application;

[0031] Figure 12 is a schematic structural diagram of a U2Net network according to an embodiment of the present application;

[0032] Figure 13 is a schematic diagram of a scene graph including commodities according to an embodiment of the present application;

[0033] Figure 14Schematic diagram of obtaining a scene graph caption according to an embodiment of the present application;

[0034] Figure 15 Schematic diagram of a commodity bounding box according to an embodiment of the present application;

[0035] Figure 16 Schematic diagram of generating an image based on text according to an embodiment of the present application;

[0036] Figure 17 Schematic diagram of generating an image based on an image according to an embodiment of the present application;

[0037] Figure 18 Schematic diagram of an optional image output result;

[0038] Figure 19 Schematic diagram of another optional image output result;

[0039] Figure 20 Schematic diagram of an image generation system according to an embodiment of the present application;

[0040] Figure 21 Schematic diagram of an image generation device according to an embodiment of the present application;

[0041] Figure 22 Schematic diagram of another image generation device according to an embodiment of the present application;

[0042] Figure 23 Schematic diagram of another image generation device according to an embodiment of the present application;

[0043] Figure 24 Schematic diagram of yet another image generation device according to an embodiment of the present application;

[0044] Figure 25 Structural block diagram of a computing device according to an embodiment of the present application;

[0045] Figure 26 Structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0046] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0047] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0048] The technical solution provided by this application is mainly implemented by using large model technology. Here, the large model refers to a deep learning model with a large number of model parameters, which usually can contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model can also be called the Foundation Model. Through large-scale unlabeled corpus for pre-training of the large model, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.

[0049] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned with a small number of samples, so that the large model can be applied to different tasks. For example, the large model can be widely applied in fields such as natural language processing (NLP), computer vision, speech processing, etc. Specifically, it can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0050] First, some nouns or terms that appear during the description of the embodiments of this application are subject to the following explanations:

[0051] The Stable Diffusion model is a text-to-image generation model based on Latent Diffusion Models (abbreviated as LDMs). Latent Diffusion Models are variants of diffusion models. Diffusion models convert images into Gaussian white noise by gradually adding noise, and then further train a denoising model to gradually recover the images from the noise;

[0052] A Caption refers to the descriptive text of an image, video, or other media content;

[0053] Fused-modal Retrieval is a technique that fuses data of different modalities (such as images and text) for retrieval.

[0054] According to an embodiment of the present application, a method for generating an image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0055] Considering that the number of model parameters of large models is huge and the computing resources of mobile terminals are limited, the above method provided by the embodiments of the present application can be applied to, for example Figure 1 the application scenarios shown, but not limited thereto. Figure 1 It is a schematic diagram of an application scenario of a method for processing dialogue information according to an embodiment of the present application. In the application scenario shown in Figure 1 a large model is deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smartphones, tablets, laptops, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. An operation interface for obtaining initial input information can be deployed on the graphical user interface of the client device. The client device 20 can interact with the user through the graphical user interface to call the large model, thereby implementing the method provided by the embodiments of the present application.

[0056] In the embodiments of the present application, the system composed of a client device and a server may perform the following steps: corresponding operations may be performed in the operation interface on the client device 20 to obtain initial input information. The client device may obtain the initial input information and send it to the server through the network. After receiving the initial input information, the server may perform the following steps: Step S102, obtain the initial input information, where the initial input information is used to represent at least one object to be displayed; Step S104, convert the initial input information into target input information, where the information format of the target input information matches the image processing model; Step S106, in the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; Step S108, analyze the input information samples using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0057] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided by the embodiments of the present application can also be applied to a model all-in-one machine. In an alternative embodiment, multiple models are built into the model all-in-one machine, and one model can be selected for adjustment as needed. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided by the embodiments of the present application. In another alternative embodiment, a trained model is built into the large model all-in-one machine. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided by the embodiments of the present application.

[0058] Furthermore, when it is necessary to train a model in a target scene, the client can also upload its own dataset, which is sent by the client to the server, so that the server can adjust the pre-trained model with this dataset to obtain a model in the target scene and then deploy it to the production environment. To facilitate the adjustment of the model, the server can provide complete adjustment tools, development frameworks, and processes, and support multiple adjustment strategies, so that the adjusted model can better adapt to different field applications and achieve high customization.

[0059] Under the above operating environment, the present application provides a method for processing dialogue information as Figure 2 shown. Figure 2 is a flowchart of a method for processing dialogue information according to an embodiment of the present application, as Figure 2 shown, and the method may include the following steps:

[0060] Step S202, obtain the initial input information.

[0061] In the technical solution provided in step S202 of the present application, the initial input information is used to represent at least one object to be displayed.

[0062] In this embodiment, the initial input information can be obtained. The initial input information can be information input by the user to describe the object to be displayed. The initial input information can include, but is not limited to, multi-modal information such as text and images. The initial input information can be used to represent at least one object to be displayed. The object to be displayed can be an entity object that the user expects to see in the image to be generated. For example, the entity can be a commodity, and the commodity can be a sofa, a floor lamp, etc. This is only for illustration and does not specifically limit the type of the object to be displayed.

[0063] Optionally, the user can provide the initial input information through various interaction methods such as uploading a file (such as an image) and inputting text in the graphical user interface, and then the initial input information input by the user can be obtained.

[0064] For example, the user can upload an image of a sofa in the graphical user interface and attach text description information: "White sofa, with a wooden coffee table beside it". The above image of the sofa and the text description information can constitute the initial input information. The above initial input information can be obtained, and the initial input information can be used to represent the object to be displayed that the user expects to display, that is, a white sofa and a wooden coffee table.

[0065] Step S204, convert the initial input information into target input information.

[0066] In the technical solution provided in step S204 of the present application, the information format of the target input information matches the image processing model.

[0067] In this embodiment, after obtaining the initial input information, the obtained initial input information can be converted into target input information. The target input information can be input information obtained by converting the initial input information. For example, when the initial input information is multi-modal information including text and images, the target input information can be a multi-modal representation obtained by encoding the multi-modal information, and the information format of the target input information matches the image processing model. For example, the information format of the target input information can be a vector format. This is only for illustration and does not specifically limit the information format of the target input information.

[0068] Optionally, the information format of the target input information matches the information format required by the image processing model. For example, when the information format of the target input information is a vector format, since the image processing model usually operates in a vector space, the information format required by the input information of the image processing model is a vector format, which is consistent with the information format of the target input information. Further, the dimension of the vector format of the target input information matches the dimension of the input space of the image processing model. For example, if the vector representation of the image processing model in the vector space is 512-dimensional, then the dimension of the vector format of the target input information should also be adjusted to a vector of the same dimension to ensure that the image processing model can correctly understand and process the input information.

[0069] Optionally, after obtaining the initial input information, if the obtained initial input information contains information of different modalities (such as an image and text description information input by the user), the initial input information can be converted into the target input information through a General Multimodal Encoder (GME) model. The converted target input information is then a multimodal representation, that is, the information of different modalities is fused into a common representation form.

[0070] Optionally, for the above-mentioned converted target input information being a multimodal representation, it may mean that the information format of the converted target input information is a semantically consistent vector representation, that is, the converted target input information can maintain the same semantics as the initial input information. Since the semantically consistent vector representation can reflect the association between the image and text description information input by the user, and the image and text description information input by the user can represent the user's intention, the image processing model can generate an image that conforms to the user's intention based on the semantically consistent vector representation.

[0071] For example, the initial input information is "a white sofa with orange cushions on it". After converting the initial input information into the target input information, the converted target input information can capture the semantic relationship between the "white sofa" and the "orange cushions". Even in vector form, it can enable the image processing model to understand and generate the corresponding image content.

[0072] Step S206: In the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold.

[0073] In the technical solution provided in step S206 of the present application, the database includes at least information samples of multiple modalities. The information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed. The information quality of the input information samples is higher than that of the target input information.

[0074] In this embodiment, after converting the initial input information into target input information, at least one input information sample with a similarity greater than the similarity threshold to the target input information can be retrieved from the database. The database can be a pre-constructed multi-modal high-quality database, which can at least include information samples of multiple modalities (that is, multi-modal: including scene image samples showing object samples, background description information of scene image samples, style information of scene image samples, spatial information of scene image samples, layout information of scene image samples, etc.). The information sample can be used to represent the attribute information of the object sample to be shown and the attribute information of the scene where the object sample is shown.

[0075] In this embodiment, the attribute information of the object sample to be shown can at least include the spatial information and layout information of the object sample to be shown. The object sample to be shown can be an entity object sample shown in the scene image sample in the database. For example, the object sample to be shown can be a commodity sample, and the commodity sample can be a sofa sample, etc. Here, it is only for illustration, and no specific restrictions are imposed on the type of the object sample to be shown. The spatial information of the object sample to be shown can be the environmental information where the object sample to be shown is located. The layout information of the object sample to be shown can be used to represent the coordinate information of the actual position of the object sample to be shown. For example, when the object sample to be shown is a commodity sample, the attribute information of the commodity sample can at least include the commodity spatial information and the commodity layout information. When the commodity sample is a sofa sample, the commodity spatial information can be the living room, and the commodity layout information can be the bounding box information of the sofa sample.

[0076] In this embodiment, the attribute information of the scene where the object sample to be shown is located can at least include the scene graph style and the scene graph description of the scene where the object sample to be shown is located. The scene where the object sample to be shown is located can be the environment or background where the object sample to be shown (such as a sofa sample) is placed. The scene graph style can be the design style or visual style adopted when generating or showing the scene image sample. The scene graph description can be the detailed text description of the scene image sample. For example, when the object sample to be shown is a commodity sample and the commodity sample is a sofa sample, the scene where the sofa sample is located can be the living room, and the scene graph style of this scene can be styles such as log and modern. The scene graph description can be "There is a round coffee table placed in front of the sofa. The coffee table is placed on a carpet with a textured pattern, and there are books, decorative objects and other small items placed on the coffee table. In the background of the sofa, there are artworks with abstract shapes and sculptural features placed on the shelf. There is a floor lamp placed on the left side of the sofa, and the window of the room where the sofa is located has thin curtains, and sunlight can penetrate into the room".

[0077] It should be noted that the above is only for illustration, and no specific restrictions are imposed on the attribute information of the object sample to be shown and the attribute information of the scene where the object sample to be shown is located.

[0078] In this embodiment, the similarity threshold can be a critical value preset according to the actual situation for measuring the similarity between the target input information and the input information sample. The information quality of the input information sample is higher than that of the target input information, that is, compared with the target input information, the information quality of the input information sample is significantly enhanced. The input information sample can be reasonable input information of an information processing model retrieved from a multimodal high-quality database that conforms to the user's intention, and can also be referred to as a retrieval sample or a retrieval result. Since the target input information is obtained by converting the initial input information input by the user, and the target input information usually does not meet the input conditions of the information processing model, while the input information sample meets the input conditions of the information processing model, that is, in this embodiment, by enhancing the user input, the target input information that does not meet the input conditions of the information processing model is prevented from affecting the image generation effect of the information processing model, thereby improving the generation efficiency of the image.

[0079] Optionally, after converting the initial input information into the target input information, the similarity between the target input information and the information samples in the multimodal high-quality database can be calculated, and the information sample with a similarity greater than the similarity threshold (i.e., the highest similarity) can be selected as the input information sample.

[0080] For example, after calculating the similarity between the target input information and the information samples in the multimodal high-quality database, the similarity between each information sample and the target input information can be obtained. Further, the obtained similarities can be sorted, and the similarity between one of the information samples (for example, the k-th information sample) and the target input information can be used as the similarity threshold. There are k information samples in the multimodal high-quality database whose similarity with the target input information is greater than the above similarity threshold, so that the k information samples with the highest similarity to the target input information can be selected from the multimodal high-quality database as the input information samples.

[0081] For example, when the initial input information is a product image and a text description, during the process of retrieving information samples in the multimodal high-quality database, the similarity between the target input information (i.e., the multimodal representation) after converting the initial input information and various constraint conditions, such as product spatial information, product layout information, scene graph style, etc. (it can also be unconstrained, and the text is an optional input, only based on the input product image), can be calculated for the information samples in the multimodal high-quality database, and the information sample with the highest similarity can be selected as the retrieval result.

[0082] Step S208: Analyze the input information sample by using an image processing model to obtain a scene image.

[0083] In the technical solution provided in step S208 of the present application, the image processing model is obtained by training a generative model, and the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0084] In this embodiment, after retrieving at least one input information sample with a similarity greater than the similarity threshold with the target input information in the database, the image processing model can be used to analyze the input information sample to obtain a scene image. Among them, the image processing model can be obtained by training a generative model, and the generative model can be a model that learns and generates new data similar to the training data. For example, the generative model can be a Stable Diffusion model, and the image processing model can be a text-to-image model based on Stable Diffusion, or a image-to-image model based on Stable Diffusion, etc. This is only for illustration and does not specifically limit the type of the image processing model.

[0085] In this embodiment, the scene image can be used to represent the display result of the object to be displayed in the target scene where it is displayed. For example, when the object to be displayed is a sofa, the target scene can be a living room, and the scene image can be an image showing the visual effect of the sofa in the living room environment. For example, the scene image can be such an image: a gray fabric sofa is placed in the center of a bright modern living room, the background wall is in a light blue tone, there is a round coffee table in natural wood color in front, and magazines and a cup of coffee are placed on the coffee table; a pot of green plants is placed in a corner of the living room, warm natural light shines through the window, and the curtains flutter gently. This scene image shows the application effect of the sofa in the modern living room, not only showing the sofa itself, but also providing users with an intuitive feeling and visual imagination of the product use through the environment and layout of the living room where the sofa is located.

[0086] Optionally, after retrieving the input information sample from the database, since the similarity between the input information sample and the target input information is the highest, and the input information sample can be used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed, therefore, the information processing model can be used to analyze the input information sample, that is, the information processing model generates a scene image according to various model input conditions represented by the above input information sample.

[0087] For example, when the retrieved input information sample is used to represent various attribute information of the sofa sample, and the commodity space information of the sofa sample is the living room and the scene graph style of the scene where the sofa sample is located is the log style, the information processing model can be used to analyze the input information sample, that is, the information processing model generates at least a scene image representing that the commodity space information of the sofa is the living room and the scene graph style of the scene where the sofa is located is the log style according to the input conditions such as the commodity space information, commodity layout information, scene graph style, and scene graph description of the sofa sample.

[0088] Through the above steps S202 to S208 of this application, initial input information is obtained, where the initial input information is used to represent at least one object to be displayed; the initial input information is converted into target input information, where the information format of the target input information matches the image processing model; in the database, at least one input information sample whose similarity to the target input information is greater than the similarity threshold is retrieved, where the database includes at least information samples of multiple modalities, and the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; the input information samples are analyzed using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed. That is to say, in the embodiment of this application, the initial input information is converted into target input information, and the information format of the target input information matches the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model and thus affect the efficiency of the scene image generated by the image processing model. Therefore, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0089] The method of converting the initial input information into target input information in the above embodiment will be further introduced below.

[0090] As an optional implementation manner, the initial input information at least includes an initial object image, and the image content of the initial object image includes the object to be displayed. Step S204 of converting the initial input information into target input information includes: at least encoding the initial object image to obtain the target input information, where the information format of the target input information is a vector format.

[0091] In this embodiment, the initial input information may at least include an initial object image, and the initial object image may be an image containing the object to be displayed input by the user, that is, the image content of the initial object image may include the object to be displayed. After obtaining the initial input information, the initial object image included in the initial input information can be encoded to obtain the target input information, where the information format of the target input information may be a vector format.

[0092] Optionally, the initial input information may only include the initial object image. After obtaining the initial object image, the GME model can be used to convert the initial object image into target input information, and the information format of the target input information is in vector format.

[0093] The method of encoding at least the initial object image in the above embodiment to obtain the target input information will be further introduced below.

[0094] As an alternative embodiment, the initial input information further includes text description information. The text description information is used to describe the generation intention of the scene image to be generated. Encoding at least the initial object image to obtain the target input information includes: respectively inputting the initial object image and the text description information into a multi-modal encoding model for encoding to obtain the target input information, where the multi-modal encoding model is obtained by training a large language model.

[0095] In this embodiment, the initial input information may further include text description information. The text description information may be the description information of the scene where the object to be displayed is located input by the user. The text description information can be used to describe the generation intention of the scene image to be generated. After obtaining the initial input information, the initial object image and the text description information included in the initial input information can be respectively input into a multi-modal encoding model for encoding to obtain the target input information. Among them, the multi-modal encoding model is obtained by training a large language model, and the multi-modal encoding model can also be called the General Multi-modal Encoding GME model. The GME model can be used to process single-modal or combined-modal inputs and generate a unified vector representation.

[0096] Optionally, the initial input information may include the initial object image and the text description information. After obtaining the initial object image and the text description information, the GME model can be used to convert the initial object image and the text description information of different modalities into a semantically consistent vector representation.

[0097] In this embodiment, the GME model is used to convert the initial object image and the text description information of different modalities into a semantically consistent vector representation, which can associate information of different modalities, thereby improving the accuracy of scene image generation.

[0098] The method of respectively inputting the initial object image and the text description information into a multi-modal encoding model for encoding to obtain the target input information in the above embodiment will be further introduced below.

[0099] As an alternative implementation, the initial object image and the text description information are respectively input into a multi-modal encoding model for encoding to obtain target input information, including: in response to the background image in the initial object image being a non-solid color background image, converting the background image in the initial object image into a solid color background image to obtain a converted initial object image; respectively inputting the converted initial object image and the text description information into the multi-modal encoding model for encoding to obtain target input information.

[0100] In this embodiment, after obtaining the initial object image and the text description information, when the background image in the initial object image is a non-solid color background image, the background image in the initial object image can be converted into a solid color background image through a matting technique to obtain a converted initial object image. Further, the GME model can be used to convert the text description information and the converted initial object image into target input information respectively.

[0101] For example, when the non-solid color background image is a non-transparent background image, the initial object image can be a non-transparent base map. When the solid color background image is a transparent background image, the initial object image can be a transparent base map. When the initial object image input by the user is a non-transparent base map containing a commodity, the non-transparent base map can be converted into a transparent base map through a commodity matting technique, and then the GME model can be used to convert the text description information input by the user and the converted transparent base map into semantically consistent vector representations respectively.

[0102] In order to avoid the background interference of the initial object image uploaded by the user, in this embodiment, when the initial object image is a non-transparent base map, the non-transparent base map is first converted into a transparent base map, and then the GME model is used to convert the text description information and the transparent base map into semantically consistent vector representations respectively, so as to reduce the background interference and significantly improve the accuracy of scene image generation.

[0103] Next, a further introduction is made to the method of retrieving at least one input information sample in the database whose similarity with the target input information is greater than the similarity threshold in the above embodiment.

[0104] As an alternative implementation, step S206 of retrieving at least one input information sample in the database whose similarity with the target input information is greater than the similarity threshold includes: retrieving multiple initial input information samples in the database whose similarities with the target input information are respectively greater than the similarity threshold; determining multiple input information samples among the multiple initial input information samples, where the difference information between every two input information samples among the multiple input information samples is greater than the difference information threshold.

[0105] In this embodiment, after converting the initial input information into target input information, a plurality of initial input information samples with a similarity greater than the similarity threshold to the target input information can be retrieved from the database. Further, among the plurality of initial input information samples, a plurality of input information samples can be determined. Among them, the initial input information sample can be a complete input condition sample retrieved from the database. The input information sample can be an input condition sample of the image processing model determined from the plurality of initial input information samples, and the difference information between every two input information samples among the plurality of input information samples is greater than the difference information threshold. The difference information can also be referred to as the inter-class deviation, and the difference information threshold can be a critical value set in advance according to the actual situation for measuring the difference information between every two input information samples.

[0106] Optionally, since the output result of the image processing model is affected by the input conditions of the image processing model, in order to ensure the difference between the output results, it is usually necessary to retrieve a plurality of input conditions with obvious differences. Therefore, in this embodiment, a plurality of initial input information samples are first retrieved from the multi-modal high-quality database, that is, a plurality of complete input conditions are retrieved. Further, the inter-class deviation between the plurality of initial input information samples is calculated, and a plurality of input information samples with a larger inter-class deviation are selected from the plurality of initial input information samples, so as to obtain the input conditions of the image processing model.

[0107] For example, M complete input conditions (M pieces of data) can be retrieved from the multi-modal high-quality database. Further, the inter-class deviation between the M pieces of data is calculated, and N (M > N) pieces of data with a larger inter-class deviation are selected as the input conditions of the image processing model.

[0108] In this embodiment, by calculating the inter-class deviation between the plurality of initial input information samples and selecting a plurality of input information samples with a larger inter-class deviation from the plurality of initial input information samples as the input conditions of the image processing model, the diversity of the output results of the image processing model is ensured, and the generation of homogeneous results is avoided.

[0109] As an alternative implementation manner, the method further includes: obtaining at least one scene image sample, where the scene image sample is used to represent the display result of the display object sample in the corresponding scene, and the image quality of the scene image sample is higher than the quality threshold; determining information samples based on the scene image sample to obtain a database.

[0110] In this embodiment, at least one scene image sample can be obtained. Further, based on the obtained scene image sample, an information sample can be determined, and then a database can be obtained. The scene image sample can be used to represent the display result of the display object sample in the corresponding scene, and the image quality of the scene image sample is higher than the quality threshold, which can be a critical value set in advance according to the actual situation for measuring the image quality of the scene image sample. The information sample can be a sample obtained by parsing the scene image sample and can be used as the data in the database.

[0111] Optionally, after obtaining at least one scene image sample, the obtained scene image sample can be parsed to obtain an information sample, which can be the data in the multi-modal high-quality database to be constructed. For example, by parsing an obtained scene image sample, a piece of data in the multi-modal high-quality database to be constructed can be obtained. Further, based on the obtained multiple information samples, a multi-modal high-quality database can be constructed.

[0112] For example, when the scene image sample is used to represent the display result in the scene where the commodity is located, the scene image sample can be parsed, and the obtained information sample can at least include multi-dimensional information such as the scene image sample, the transparent background image of the commodity, the unified multi-modal representation, the scene graph style, the commodity space information, the scene graph Caption, and the commodity layout information. This information sample can be used as a piece of data in the multi-modal high-quality database to be constructed. When the data volume reaches a certain scale, a multi-modal high-quality database can be obtained.

[0113] One piece of data in the multi-modal high-quality database in this embodiment contains multi-dimensional information. Among them, the unified multi-modal representation, the scene graph style, the commodity space information, and the commodity layout information can be used in the retrieval process of the commodity image input by the user. The scene image sample, the scene graph style, the commodity layout information, and the scene graph Caption will be used as the recommendation results and the input conditions of the image processing model, thereby significantly improving the accuracy of image retrieval and generation.

[0114] Next, the method of determining the information sample and obtaining the database based on the scene image sample in this embodiment will be further introduced.

[0115] As an alternative implementation manner, determining the information sample and obtaining the database based on the scene image sample includes: generating a text description information sample based on the scene image sample, where the text description information sample is used to describe the image content of the scene image sample; determining the information sample and obtaining the database based on the text description information sample and the scene image sample.

[0116] In this embodiment, during the construction of the database, after obtaining at least one scene image sample, it is necessary to understand the obtained scene image sample to generate a text description information sample. Further transforming the scene image sample and the generated text description information sample can obtain a unified multimodal representation, and then the information sample can be determined to obtain the database. Among them, the text description information sample is used to describe the image content of the scene image sample and can also be called the scene graph Caption.

[0117] Optionally, since the text description information sample can be used for the image generation background filling task, a multimodal technology that only describes the background content of the scene image sample can be used to generate the text description information sample, and the ability to only describe the background content of the scene image sample can be achieved through means such as seed data production, seed data filtering, and model distillation.

[0118] Optionally, during the seed data production process, for the scene image sample, first, a first prompt word designed manually is used to call the multimodal generation model to obtain a complete and detailed description of the scene image sample. To further obtain the background content that only describes the scene image sample, images in which the main object (such as a commodity) is occluded are obtained through technical means such as image segmentation and image processing. The complete and detailed description of the scene image sample, the second prompt word, and the image in which the main object is occluded are used as the input of the multimodal generation model together to obtain the text description information sample that only describes the background content of the scene image sample output by the multimodal generation model.

[0119] It should be noted that for the Caption task without unconditional restrictions, the results of the multimodal generation model are highly credible. However, due to the increase in prompt word conditions, the obtained seed data (image pairs: the image in which the main object is occluded and the background description of the scene image sample) has a relatively low proportion of hallucinations or errors. Therefore, it is necessary to further filter the seed data.

[0120] Optionally, the background description of the scene image sample is formatted into descriptions of different regions around the main object, and then the relevance between the descriptions of different regions and each region of the image is calculated respectively. When the relevance calculation result of a certain region is low, it is considered that the quality of the seed data is unqualified, and thus the filtering of the seed data is completed.

[0121] Optionally, after obtaining high-quality seed data, model distillation is performed on the bidirectional pre-trained multimodal model to improve the effect of the background description of the scene image sample. After completing the model distillation, the distilled model is evaluated, and the distilled model whose performance meets the conditions is selected. Then, the multimodal ability of the distilled model whose performance meets the conditions is used to achieve the goal of only describing the background content of the scene image sample.

[0122] This embodiment realizes the ability to only describe the background content of the scene image sample through means such as seed data production, seed data filtering, and model distillation, generating a text description information sample of the background content of the scene image sample, avoiding repeated descriptions of the main commodity, and making the generated text description information sample focus on providing rich background information for generating images of the commodity.

[0123] The method for determining information samples and obtaining a database based on the text description information sample and the scene image sample in this embodiment will be further introduced below.

[0124] As an optional implementation manner, the information format includes a vector format. Determining information samples and obtaining a database based on the text description information sample and the scene image sample includes: obtaining image information of the display object sample from the scene image sample; generating an object sample image of the display object sample from the image information of the display object sample, where the image content of the object sample image includes the display object sample on a solid color background image; respectively converting the object sample image and the text description information sample into feature vectors in vector format; and determining at least the feature vectors in vector format as information samples to obtain a database.

[0125] In this embodiment, the information format may include a vector format. After obtaining the scene image sample, image information of the display object sample can be obtained from the obtained scene image sample. Further, the image information of the display object sample is used to generate an object sample image of the display object sample. After generating the text description information sample and the object sample image, the generated text description information sample and the object sample image can be respectively converted into feature vectors in vector format, and finally at least the feature vectors in vector format are determined as information samples to obtain a database. Among them, the image content of the object sample image may include the display object sample on a solid color background image. For example, the solid color background image may be a transparent background image. When the display object sample is a commodity sample, the object sample image may also be referred to as a commodity transparent background image.

[0126] Optionally, after obtaining the scene image sample, image information of each display object sample can be identified and separated from the obtained scene image sample based on techniques such as image segmentation or object detection. After obtaining the image information of the display object sample, a transparent (Alpha) channel can be added to the image information so that the background image of the display object sample is a solid color background image, thereby obtaining an object sample image.

[0127] For example, the obtained scene image sample can be a complete image containing one or more display objects (such as commodities). For example, an image sample of a living room can include commodities such as a sofa, a coffee table, and a carpet. After obtaining the scene image sample, the image information of each commodity can be identified and separated from the scene image sample based on techniques such as image segmentation or object detection. After obtaining the image information of the commodity, an Alpha channel can be added to the image information of the commodity so that the background image of the commodity is a transparent background image, thereby obtaining a transparent background image of the commodity.

[0128] Optionally, after generating the text description information sample and the object sample image, the GME model can be used to convert the text description information sample and the object sample image in different modalities into feature vectors in vector format, that is, into a unified multi-modal representation, and then at least determine the multi-modal representation as the information sample to obtain a database.

[0129] By converting the text description information sample and the object sample image into a unified multi-modal representation, this embodiment can achieve efficient and accurate cross-modal information retrieval, so that the input information sample with the highest similarity to the user input and meeting the input conditions of the image processing model can be retrieved from the constructed database.

[0130] Next, a further introduction will be made to the method of at least determining the feature vector in vector format as the information sample to obtain a database in the above embodiment.

[0131] As an optional implementation manner, at least determining the feature vector in vector format as the information sample to obtain a database includes: determining the feature vector in vector format and at least one of the following information as the information sample to obtain a database: the scene style of the scene image sample, the spatial type sample of the space where the display object sample is located in the corresponding scene, the layout information sample of the layout of the display object sample in the corresponding scene, the object sample image, the scene image sample, and the text description information sample.

[0132] In this embodiment, after the object sample image and the text description information sample are respectively converted into feature vectors in vector format, the feature vectors in vector format, as well as the scene style of the scene image sample, the spatial type sample showing the space where the object sample is located in the corresponding scene, the layout information sample showing the layout of the object sample in the corresponding scene, the object sample image, the scene image sample, and the text description information sample can be determined as information samples to obtain a database. Among them, the scene style can be the artistic or design style reflected by the scene image sample, and can also be called the scene graph style. For example, the scene style can be log, modern, etc. Here is only an example, and the type of the scene style is not specifically limited. The spatial type sample can be used to represent the specific position and spatial relationship of the object sample in the corresponding scene. For example, when the object sample is a commodity, the spatial type sample can also be called commodity spatial information. When the commodity is a sofa, the spatial type sample can be the living room. Here is only an example, and the spatial type sample is not specifically limited. The layout information sample can be used to represent the bounding box of the object sample in the corresponding scene. For example, when the object sample is a commodity, the layout information sample can also be called commodity layout information. When the commodity is a sofa, the layout information sample can be the coordinates of the bounding box of the sofa.

[0133] Optionally, the pre-trained self-supervised learning visual model can extract image features from the scene image sample, and then input the extracted image features into a Multilayer Perceptron (MLP for short) for classification to obtain the scene style and the spatial type sample. That is, the scene style and the spatial type sample can be obtained by training the corresponding classifiers respectively. For different classification tasks, a small number of samples need to be manually labeled to fine-tune the classifier. When the verification accuracy reaches a certain threshold, the classifier can be used for data classification to obtain the scene style and the spatial type sample.

[0134] Optionally, the layout information sample may include the location where the object sample is located and the description of the object positions in the background around the object sample. The acquisition of the description of the object positions in the background around the object sample is the same as that of the text description information sample, which will not be elaborated here. Since the object sample image obtained by cropping the object sample (i.e., the commodity transparent background image) contains transparency channel information, the location where the object sample is located can be obtained according to the transparency channel information. The transparent Alpha channel can be separated from the commodity transparent background image, the pixel values of the Alpha channel can be traversed, and the coordinates of the non-transparent pixels can be recorded. Based on the recorded coordinates of the non-transparent pixels, the bounding box of the object sample can be calculated, so as to obtain the layout information sample.

[0135] For example, when the display object sample is a product sample, multi-dimensional information such as the obtained scene image sample, product transparent background image, text description information sample, multi-modal representation converted from the product transparent background image and text description information sample, scene graph style, product spatial information, and product layout information can be used as an information sample (i.e., one piece of data). When the data volume reaches a certain scale, a multi-modal high-quality database can be obtained.

[0136] This embodiment constructs a multi-modal high-quality database with multi-dimensional information as information samples, which can ensure a high degree of matching between the retrieval result and the user input, thereby improving the quality and efficiency of image generation.

[0137] The method of using an image processing model to analyze the input information sample to obtain a scene image in this embodiment will be further introduced below.

[0138] As an alternative implementation, in step S208, using an image processing model to analyze the input information sample to obtain a scene image includes: using the initial input information to adjust the text description information sample in the input information sample to obtain an adjusted input information sample, where the initial input information is used to describe the generation intention of the scene image to be generated, the text description information is used to describe the image content of the scene image sample in the input information sample, and the adjusted input information sample matches the generation intention; generating a prompt information from the adjusted input information sample, where the prompt information is used to guide the image processing model in the process of analyzing the adjusted input information sample; and using the prompt information to guide the image processing model to analyze the adjusted input information sample to obtain a scene image.

[0139] In this embodiment, after retrieving at least one input information sample with a similarity greater than the similarity threshold to the target input information in the database, the text description information sample in the input information sample can be adjusted using the initial input information to obtain an adjusted input information sample. Further, the adjusted input information sample can be used to generate prompt information, and then the generated prompt information can be used to guide the image processing model to analyze the adjusted input information sample to obtain a scene image. Among them, the initial input information can be used to describe the generation intention of the scene image to be generated, and this generation intention is the user intention. The text description information sample can be a sample that only describes the background content of the scene image sample. The prompt information (Prompt) can be used to guide the image processing model in the process of analyzing the adjusted input information sample and can also be called a prompt word.

[0140] Optionally, after retrieving the input information sample with the highest similarity to the target input information from the multi-modal high-quality database, the retrieved input information sample is usually only close to the user's intention (i.e., the user input) and does not meet the user's needs. Therefore, in order to fully express the user's intention, this embodiment further aligns the user's intention (i.e., rewrites the Prompt).

[0141] Optionally, using the initial input information, the text description information sample in the input information sample can be adjusted to obtain an adjusted input information sample. That is, the retrieved sample (input information sample) is aligned to the user input (initial input information).

[0142] Optionally, after rewriting the scene graph Caption (text description information sample) in the retrieved sample into the user intention alignment result (adjusted input information sample) according to the user input, the rewritten user intention alignment result can be used to generate a prompt message Prompt, and then the prompt message Prompt can be used as the input condition of the image processing model to guide the image processing model to analyze the adjusted input information sample to obtain a scene image.

[0143] This embodiment rewrites the scene graph Caption in the retrieved sample into the user intention alignment result according to the user input, and uses the rewritten user intention alignment result as the input condition of the image processing model, so that the image processing model can better understand and meet the specific needs of the user, and makes the generated image highly consistent with the user intention in function and semantics, thereby improving the user experience.

[0144] The method of using the initial input information to adjust the text description information sample in the input information sample to obtain an adjusted input information sample in this embodiment will be further introduced below.

[0145] As an optional implementation manner, using the initial input information to adjust the text description information sample in the input information sample to obtain an adjusted input information sample includes: inputting the initial input information and the input information sample into a vision-language model, and using the vision-language model to adjust the text description information sample in the input information sample according to the initial input information to obtain an adjusted input information sample, where the vision-language model is obtained by training a multi-modal large model.

[0146] In this embodiment, after retrieving at least one input information sample from the database whose similarity to the target input information is greater than the similarity threshold, the initial input information and the input information sample can be input into the vision-language model, and the vision-language model is used to adjust the text description information sample in the input information sample according to the initial input information to obtain an adjusted input information sample. Among them, the vision-language model can be obtained by training a multimodal large model, and can also be called a fine-tuned vision-language model.

[0147] Optionally, the fine-tuned vision-language model is used to adjust the text description information sample in the input information sample according to the initial input information to obtain an adjusted input information sample. Among them, the fine-tuned vision-language model combines the powerful capabilities of language understanding and visual understanding. The moderate parameter scale of this model enables this model to achieve a good balance between performance and resource requirements. After instruction fine-tuning, it can align the retrieved samples to the user input.

[0148] In this embodiment, by fine-tuning the vision-language model, the retrieved samples can be aligned to the user input, and the text description and image information input by the user can be understood more accurately. Then, the description of the retrieved samples is rewritten to ensure that the input of the image processing model is consistent with the user's intention, Figure 1 so as to improve the accuracy of image generation.

[0149] In the embodiment of the present application, the initial input information is converted into target input information, and the information format of the target input information matches that of the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, this input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may lead to the input not meeting the input conditions of the image processing model and then affecting the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0150] The embodiment of the present application also provides a method for generating an image, Figure 3 which is a flowchart of another method for generating an image according to the embodiment of the present application. As Figure 3 shown, the method may include the following steps:

[0151] Step S302, in response to an input instruction on the operation interface, display the initial input information on the operation interface.

[0152] In the technical solution provided in step S302 of the present application above, the initial input information is used to represent at least one object to be displayed.

[0153] In this embodiment, in response to an input instruction on the operation interface, initial input information can be displayed on the operation interface. Among them, the operation interface can be the operation interface displayed on the client device. For example, it can be an interface for interaction between the user and a computer or software system, and can be a graphical user interface, a command line interface, etc. Here is only an example, and the form of the operation interface is not specifically limited.

[0154] In this embodiment, the input instruction can be an instruction for controlling the operation interface to display the initial input information. For example, it can be an instruction input by the user on the operation interface. The user can input text through the keyboard, or click on a button or option on the operation interface through the mouse, or generate an input instruction by dragging the initial input information to a specified area. The initial input information can be used to represent at least one object to be displayed.

[0155] In an alternative embodiment, the user can open the operation interface deployed on the client device and input the initial input information on the operation interface to generate an input instruction. According to the generated input instruction, the corresponding initial input information can be loaded and the initial input information can be displayed on the operation interface.

[0156] For example, when the user uploads an image of a sofa on the graphical user interface and attaches a text description: "White sofa, with a wooden coffee table beside it", an input instruction can be generated. According to the generated input instruction, the corresponding initial input information can be loaded and the initial input information can be displayed on the operation interface, that is, the above-mentioned image of the sofa and the text description are displayed, and the initial input information can be used to represent the objects to be displayed expected by the user, that is, a white sofa and a wooden coffee table.

[0157] Step S304: In response to an image generation instruction acting on the operation interface, display a scene image corresponding to the initial input information on the operation interface.

[0158] In the technical solution provided in step S304 of the present application above, the scene image is used to represent the display result of the object to be displayed in the initial input information in the target scene where it is displayed. The scene image is obtained by analyzing the input information sample using an image processing model. The image processing model is obtained by training a generative model. The input information sample is an information sample in the database whose similarity to the target input information is greater than the similarity threshold. The database includes at least information samples of multiple modalities. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information. The target input information is obtained by converting the initial input information, and the information format of the target input information matches that of the image processing model.

[0159] In this embodiment, in response to an input instruction on the operation interface, after the initial input information is displayed on the operation interface, in response to an image generation instruction acting on the operation interface, a scene image corresponding to the initial input information can be displayed on the operation interface. Among them, the image generation instruction can be an instruction for generating an image from the initial input information displayed on the operation interface. The user can generate an image generation instruction by clicking on the initial input information displayed on the operation interface. This is only an example for illustration, and no specific limitation is imposed on the image generation instruction.

[0160] Optionally, after the initial input information is displayed on the operation interface, if the displayed initial input information contains information of different modalities (for example, an image and a text description input by the user), the initial input information can be converted into target input information through a GME model. The converted target input information is a multi-modal representation, that is, information of different modalities is fused into a common representation form.

[0161] Optionally, after the initial input information is converted into target input information, a similarity calculation can be performed between the target input information and information samples in a multi-modal high-quality database, and the information sample with the highest similarity is selected as the input information sample.

[0162] Optionally, after the input information sample is retrieved from the database, since the similarity between the input information sample and the target input information is the highest, and the input information sample can be used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed, an information processing model can be used to analyze the input information sample, that is, the information processing model generates a scene image according to various model input conditions represented by the above input information sample.

[0163] Through steps S302 to S304 of the present application above, in response to an input instruction on the operation interface, the initial input information is displayed on the operation interface, where the initial input information is used to represent at least one object to be displayed; in response to an image generation instruction acting on the operation interface, a scene image corresponding to the initial input information is displayed on the operation interface, where the scene image is used to represent the display result of the object to be displayed in the initial input information in the target scene where it is displayed. The scene image is obtained by analyzing the input information sample using an image processing model. The input information sample is an information sample in the database whose similarity with the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information. The target input information is obtained by converting the initial input information, and the information format of the target input information matches that of the image processing model.

[0164] That is to say, in the embodiment of the present application, in response to an image generation instruction acting on the operation interface, a scene image corresponding to the initial input information is displayed on the operation interface. The scene image is obtained by analyzing the input information sample using an image processing model. The input information sample is the sample with the highest similarity to the target input information converted from the initial input information retrieved from the database. Since the information format of the target input information matches the image processing model, and the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model, thereby affecting the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0165] Next, a further introduction is made to the method of the above-mentioned embodiment of responding to an image generation instruction acting on the operation interface and displaying a scene image corresponding to the initial input information on the operation interface.

[0166] As an optional implementation manner, the initial input information at least includes an initial object image, and the image content of the initial object image includes the object to be displayed. Step S304, in response to an image generation instruction acting on the operation interface, displaying a scene image corresponding to the initial input information on the operation interface includes: in response to a recommended mode selection instruction acting on the operation interface, entering the recommended mode; in the recommended mode, in response to an image generation instruction acting on the operation interface, displaying a scene image on the operation interface, where the target input information in the recommended mode is obtained by encoding the initial object image, and the information format of the target input information is a vector format.

[0167] In this embodiment, the initial input information may at least include an initial object image, and the image content of the initial object image may include the object to be displayed. After displaying the initial input information on the operation interface in response to an input instruction acting on the operation interface, in response to a recommended mode selection instruction acting on the operation interface, the recommended mode can be entered. Further, in the recommended mode, in response to an image generation instruction acting on the operation interface, a scene image can be displayed on the operation interface. Among them, the recommended mode selection instruction may be an instruction for the user to select to enter the recommended mode, and the recommended mode may be a mode for quickly generating a scene image. The target input information in the recommended mode may be obtained by encoding the initial object image, and the information format of the target input information may be a vector format.

[0168] Optionally, the initial input information may only include the initial object image. After the initial input information is displayed on the operation interface, in response to a recommended mode selection instruction acting on the operation interface, the recommended mode can be entered. Further, in the recommended mode, in response to an image generation instruction acting on the operation interface, the GME model can be used to convert the initial object image into target input information, and the information format of the target input information is a vector format, so as to retrieve the input information sample with the highest similarity to the target input information from the database, and the image processing model is used to analyze the input information sample to obtain a scene image, and then the scene image can be displayed on the operation interface.

[0169] The method of displaying the scene image corresponding to the initial input information on the operation interface in response to the image generation instruction acting on the operation interface in this embodiment will be further introduced below.

[0170] As an alternative implementation manner, the initial input information includes an initial object image and text description information. The image content of the initial object image includes the object to be displayed, and the text description information is used to describe the generation intention of the scene image to be generated. Step S304, in response to an image generation instruction acting on the operation interface, displaying the scene image corresponding to the initial input information on the operation interface, includes: in response to a custom mode selection instruction acting on the operation interface, entering the custom mode; in the custom mode, in response to an image generation instruction acting on the operation interface, displaying the scene image on the operation interface, where the target input information in the custom mode is obtained by encoding the initial object image and the text description information into a multi-modal encoding model respectively, the information format of the target input information is a vector format, and the multi-modal encoding model is obtained by training a large language model.

[0171] In this embodiment, the initial input information may include an initial object image and text description information. The image content of the initial object image may include the object to be displayed, and the text description information may be used to describe the generation intention of the scene image to be generated. After the initial input information is displayed on the operation interface in response to an input instruction acting on the operation interface, in response to a custom mode selection instruction acting on the operation interface, the custom mode can be entered. Further, in the custom mode, in response to an image generation instruction acting on the operation interface, the scene image can be displayed on the operation interface. Among them, the custom mode selection instruction may be an instruction for the user to select to enter the custom mode, and the custom mode may be a mode that supports the user to input text to express the intention. The target input information in the custom mode may be obtained by encoding the initial object image and the text description information into a multi-modal encoding model respectively, the information format of the target input information may be a vector format, and the multi-modal encoding model may be obtained by training a large language model.

[0172] Optionally, the initial input information may include an initial object image and text description information. After the initial input information is displayed on the operation interface, a custom mode can be entered in response to a custom mode selection instruction acting on the operation interface. Further, in the custom mode, in response to an image generation instruction acting on the operation interface, the GME model can be used to convert the initial object image and text description information in different modalities into a semantically consistent vector representation, so as to retrieve the input information sample with the highest similarity to the semantically consistent vector representation from the database, and an image processing model can be used to analyze the input information sample to obtain a scene image, and then the scene image can be displayed on the operation interface.

[0173] As an alternative embodiment, the method further includes: in response to an information selection instruction acting on the operation interface, displaying constraint condition information on the operation interface, where the input information sample is an information sample in the database that has a similarity greater than a similarity threshold with the target input information and satisfies the constraint condition information.

[0174] In this embodiment, in response to an information selection instruction acting on the operation interface, constraint condition information can be displayed on the operation interface. Wherein, the input information sample is an information sample in the database that has a similarity greater than a similarity threshold with the target input information and satisfies the constraint condition information. For example, when the display object is a commodity, the constraint condition information may include the scene graph style, commodity space information, commodity layout information, etc. This is only an example and does not specifically limit the content of the constraint condition information.

[0175] In the embodiment of the present application, in response to an image generation instruction acting on the operation interface, a scene image corresponding to the initial input information is displayed on the operation interface. The scene image is obtained by analyzing the input information sample using an image processing model. The input information sample is the sample with the highest similarity to the target input information converted from the initial input information retrieved from the database. Since the information format of the target input information matches the image processing model, and the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input not to meet the input conditions of the image processing model, thereby affecting the efficiency of the scene image generated by the image processing model, and thus achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0176] The embodiment of the present application also provides a method for generating an image. Figure 4 It is a flowchart of another method for generating an image according to the embodiment of the present application. As Figure 4 shown, the method may include the following steps:

[0177] Step S402, obtain initial input information from an e-commerce platform.

[0178] In the technical solution provided in step S402 of the present application, the initial input information is used to represent at least one product to be displayed.

[0179] Step S404, convert the initial input information into target input information.

[0180] In the technical solution provided in step S404 of the present application, the information format of the target input information matches the image processing model.

[0181] In this embodiment, after obtaining the initial input information from the e-commerce platform, the obtained initial input information can be converted into target input information.

[0182] Optionally, after obtaining the initial input information from the e-commerce platform, if the obtained initial input information contains information of different modalities (for example, an image and a text description input by the user), the initial input information can be converted into target input information through the GME model, and the converted target input information is a multi-modal representation, that is, information of different modalities is fused into a common representation form.

[0183] Step S406, retrieve at least one input information sample in the database whose similarity to the target input information is greater than the similarity threshold.

[0184] In the technical solution provided in step S406 of the present application, the database includes at least information samples of multiple modalities. The information samples are used to represent the attribute information of the product sample to be displayed and the attribute information of the scene where the product sample is displayed. The information quality of the input information sample is higher than that of the target input information.

[0185] In this embodiment, after converting the initial input information into target input information, at least one input information sample whose similarity to the target input information is greater than the similarity threshold can be retrieved in the database.

[0186] Optionally, after converting the initial input information into target input information, the similarity between the target input information and the information samples in the multi-modal high-quality database can be calculated, and the information sample with the highest similarity can be selected as the input information sample.

[0187] Step S408, analyze the input information sample by using an image processing model to obtain a scene image.

[0188] In the technical solution provided in step S408 of the present application, the image processing model is obtained by training a generative model. The scene image is used to represent the display result of the product to be displayed in the target scene where it is displayed.

[0189] In this embodiment, after retrieving at least one input information sample in the database whose similarity to the target input information is greater than the similarity threshold, an image processing model can be used to analyze the input information sample to obtain a scene image.

[0190] Optionally, after retrieving the input information sample from the database, since the similarity between the input information sample and the target input information is the highest, and the input information sample can be used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed, therefore, an information processing model can be used to analyze the input information sample, that is, the information processing model generates a scene image according to various model input conditions represented by the above input information sample.

[0191] Step S410: Send the scene image to the e-commerce platform.

[0192] In this embodiment, after using the image processing model to analyze the input information sample to obtain a scene image, the scene image can be sent to the e-commerce platform so that users can view or download the generated scene image.

[0193] Through the above steps S402 to S410 of the present application, obtain initial input information from the e-commerce platform, where the initial input information is used to represent at least one product to be displayed; convert the initial input information into target input information, where the information format of the target input information matches the image processing model; in the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold, where the database includes at least information samples of multiple modalities, and the information samples are used to represent the attribute information of the display product sample and the attribute information of the scene where the display product sample is displayed, and the information quality of the input information sample is higher than that of the target input information; use the image processing model to analyze the input information sample to obtain a scene image, where the scene image is used to represent the display result of the product to be displayed in the target scene where it is displayed; send the scene image to the e-commerce platform.

[0194] That is to say, the embodiment of the present application converts the initial input information obtained from the e-commerce platform into target input information, and the information format of the target input information matches that of the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input not to meet the input conditions of the image processing model and further affect the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0195] The embodiment of the present application also provides a method for generating an image. Figure 5 It is a flowchart of another method for generating an image according to the embodiment of the present application, as Figure 5 shown. The method may include the following steps:

[0196] Step S502, obtain the initial input information by calling the first interface.

[0197] In the technical solution provided in step S502 of the present application above, the first interface includes a first parameter, and the parameter value of the first parameter is the initial input information, and the initial input information is used to represent at least one object to be displayed.

[0198] Step S504, convert the initial input information into target input information.

[0199] In the technical solution provided in step S504 of the present application above, the information format of the target input information matches that of the image processing model.

[0200] In this embodiment, after obtaining the initial input information by calling the first interface, the obtained initial input information can be converted into target input information.

[0201] Optionally, after obtaining the initial input information by calling the first interface, if the obtained initial input information contains information of different modalities (for example, an image and a text description input by the user), the initial input information can be converted into target input information through the GME model, and the converted target input information is a multi-modal representation, that is, the information of different modalities is fused into a common representation form.

[0202] Step S506, in the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold.

[0203] In the technical solution provided in step S506 of the present application, the database includes at least information samples of multiple modalities. The information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information.

[0204] In this embodiment, after converting the initial input information into the target input information, at least one input information sample with a similarity greater than the similarity threshold to the target input information can be retrieved from the database.

[0205] Optionally, after converting the initial input information into the target input information, the similarity between the target input information and the information samples in the multi-modal high-quality database can be calculated, and the information sample with the highest similarity can be selected as the input information sample.

[0206] Step S508: Analyze the input information sample using an image processing model to obtain a scene image.

[0207] In the technical solution provided in step S508 of the present application, the image processing model is obtained by training a generative model, and the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0208] In this embodiment, after retrieving at least one input information sample with a similarity greater than the similarity threshold to the target input information from the database, the input information sample can be analyzed using an image processing model to obtain a scene image.

[0209] Optionally, after retrieving the input information sample from the database, since the similarity between the input information sample and the target input information is the highest, and the input information sample can be used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed, therefore, the information processing model can be used to analyze the input information sample, that is, the information processing model generates a scene image according to the multiple model input conditions represented by the above input information sample.

[0210] Step S510: Output the scene image by calling the second interface.

[0211] In the technical solution provided in step S510 of the present application, the second interface includes a second parameter, and the parameter value of the second parameter includes the scene image.

[0212] In this embodiment, after analyzing the input information sample using an image processing model to obtain a scene image, the scene image can be output by calling the second interface.

[0213] Through the above steps S502 to S510 of this application, the initial input information is obtained by calling the first interface. Among them, the first interface includes a first parameter, and the parameter value of the first parameter is the initial input information, and the initial input information is used to represent at least one object to be displayed; the initial input information is converted into target input information, where the information format of the target input information matches the image processing model; in the database, at least one input information sample whose similarity to the target input information is greater than the similarity threshold is retrieved. Among them, the database includes at least information samples of multiple modalities. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information; the image processing model is used to analyze the input information sample to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene; the scene image is output by calling the second interface. Among them, the second interface includes a second parameter, and the parameter value of the second parameter includes the scene image.

[0214] That is to say, the embodiment of this application converts the initial input information obtained by calling the first interface into target input information, and the information format of the target input information matches the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may lead to the input not meeting the input conditions of the image processing model and then affecting the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0215] Next, the technical solutions of the embodiments of the present disclosure will be further introduced by way of preferred embodiments.

[0216] Currently, due to the wide application of the diffusion model in the generation model and the maturity of the diffusion model application and multi-modal technology, the application of artificial intelligence generated content (AIGC) has also made progress. For example, AI painting under various conditions, 3D model / voice / video generation, etc.

[0217] Visual Language Models (VLM) are technologies that integrate computer vision and natural language processing in the field of artificial intelligence. By understanding and generating associations between vision (such as images, videos) and language (such as text, speech), VLM enables multi-modal information processing and interaction.

[0218] Retrieval-Augmented Generation (RAG) is a technique that uses information from private or proprietary data sources to assist in text generation. It combines a retrieval model (designed to search large datasets or knowledge bases) and a generation model (such as a large language model LLM, which uses the retrieved information to generate a readable text response). By adding background information from more data sources and supplementing the original knowledge base of the LLM through training, retrieval-augmented generation can improve the relevance of the search experience and enhance the output of large language models without the need to retrain the model.

[0219] During the process of generating scene images, the input conditions of the image processing model will directly affect the image generation effect of the image processing model. However, users usually cannot accurately input the reasonable conditions required by the image processing model. Therefore, enhancing user input is an urgent problem to be solved.

[0220] To solve the above problems, this embodiment constructs a multi-modal high-quality database by understanding product images and scene images, and further combines technologies such as fusion-modal retrieval and Prompt rewriting to construct the user input into diverse and reasonable input conditions that conform to the user's intention, so as to solve problems such as difficult user input and alignment of user input.

[0221] In an optional example, the advantage of existing image-to-image generation techniques is that the edges do not deform, but they cannot align with the user's intention, resulting in problems such as unreasonable layout, low aesthetics, and homogenization of the generated images.

[0222] To solve the above problems, this application proposes a user input enhancement technical solution. This method supports different levels of merchants to quickly generate diverse scene graphs and custom generate scene graphs that meet the intentions. Based on the construction of a multi-modal high-quality database, through unified multi-modal representation and retrieval, it solves problems such as difficulties for users to construct input conditions, unappealing and highly homogeneous default-configured generated images, and further aligns user intentions with the retrieval results to improve the performance of user intentions in the generated image results. This method uses a multi-modal encoding model for unified multi-modal representation and retrieval, fully considering the user's multi-modal input conditions; uses a vision-language model, with the user input and retrieval results as the model inputs, to obtain the input of the image generation model that aligns with the intention; obtains the model input conditions through retrieval enhancement to solve problems such as difficult user input and homogeneous results, thereby improving the image generation effect.

[0223] This method can be applied to the recommendation mode or custom mode for indoor studio shooting in homeworks. Among them, the recommendation mode is suitable for quick generation, and the custom mode supports users to input text and express intentions. Figure 6 It is a schematic diagram of a recommendation mode for indoor studio shooting in homeworks according to an embodiment of the present application, as Figure 6 shown. The recommendation mode supports displaying the product main body, product type, placement space, scene style, aspect ratio, and composition method. Among them, the product main body is an image of a sofa, and options such as upload or history are provided to upload the image of the product main body. The product type is a sofa, the placement space is the living room, the scene styles include intelligent recommendation, log, modern, Nordic, French, luxury, Italian, and neo-Chinese, etc., and it also supports viewing more scene styles. The aspect ratio is 1:1, and the composition method is intelligent composition.

[0224] Figure 7 It is a schematic diagram of a custom mode for indoor studio shooting in homeworks according to an embodiment of the present application, as Figure 7 shown. The custom mode supports displaying the product main body, product type, scene description, number of pictures, aspect ratio, and composition method. Among them, the product main body and product type are Figure 6 the same. The scene description part shows prompt words such as "Please supplement scene requirements, for example, input [carpet], and this element will try to appear in the picture", and there are various types of prompt words available for selection below the scene description, such as those related to space like dining room, against the wall, on the floor, hallway cabinet, on the carpet, etc., those related to decoration like fruit plate, curtain, chair, bedside table, wooden floor, etc., and those related to atmosphere like cool tone, creamy style, warm tone, vintage tone, etc., and users can replace the above prompt words according to their personal preferences. The number of pictures is 4, the aspect ratio is 1:1, and the composition method is free composition.

[0225] This method belongs to the image - to - image generation ability for product images. Users can directly generate diverse scene images by only inputting a product transparent background image or a scene image containing the product; they can also customize the scene layout to generate scene images that meet the user's intentions. This method can automatically construct diverse, reasonable, and user - intention - compliant input conditions for the image - processing model, and these input conditions will directly affect the output effect of the image - processing model.

[0226] Figure 8 It is a schematic diagram of an input image according to an embodiment of the present application. As Figure 8 shown, the image input by the user is an image of a sofa, and the text input by the user can be "There is a fishing lamp on the left side of the sofa, a round marble coffee table in front of the sofa, and a decorative painting behind the sofa."

[0227] Figure 9 It is a schematic diagram of the image - generation result in a recommended mode according to an embodiment of the present application. As Figure 9 shown, it can be seen the image - generation result using this method in the recommended mode. In the recommended mode, images can be generated quickly, and the user input does not contain text.

[0228] Figure 10 It is a schematic diagram of the image - generation result in a custom mode according to an embodiment of the present application. As Figure 10 shown, it can be seen the image - generation result using this method in the custom mode. The custom mode is an advanced version of the recommended mode. When the user has no clear idea, they can directly use the recommended mode to quickly generate images. When the user has a clear demand or idea, they can express it in the custom mode. This method will fully consider the user's intention, enhance the user input, and finally reflect it in the image - generation result. The following mainly describes this method in the custom mode.

[0229] Figure 11 It is a flowchart of a user - input enhancement method according to an embodiment of the present application. As Figure 11 shown, the process of this user - input enhancement method can include the following steps:

[0230] Step S1101, obtain the text and image input by the user.

[0231] Step S1102, determine whether the image is a product transparent background image.

[0232] In the above steps, after obtaining the text and image input by the user, first determine whether the input image is a product transparent background image. If the image is not a product transparent background image, then proceed to step S1103; otherwise, proceed to step S1104.

[0233] Step S1103, perform matte extraction on the product.

[0234] In the above steps, when the input image is not a transparent background image of the product, product matting is used to obtain the transparent background image of the product. Optionally, the user can directly upload the transparent background image of the product, or directly upload the scene image with the product. After product matting, the transparent background image of the product can be obtained.

[0235] Optionally, nested U-structure is used for salient object detection in product matting (Going Deeper with NestedU-Structure for Salient Object Detection, abbreviated as U2Net network). Figure 12 It is a schematic structural diagram of a U2Net network according to an embodiment of the present application. The U2Net network is proposed for the salient object detection task (SalientObject Detection, abbreviated as SOD). The salient object detection task is similar to the semantic segmentation task. The salient object detection task is a binary classification task, and its task is to segment the objects or regions in the picture. Therefore, it only includes two categories: foreground and background. Figure 13 It is a schematic diagram of a scene image containing a product according to an embodiment of the present application. Figure 13 After the scene image shown is subjected to product matting, an image of the main content of the product with a transparent channel can be obtained, that is, Figure 8 the transparent background image of the product obtained by product matting as shown.

[0236] Step S1104, obtain the transparent background image of the product.

[0237] Step S1105, calculate the multi-modal representation of the transparent background image of the product and the text.

[0238] In the above steps, after obtaining the transparent background image of the product, the GME model can be used to calculate the unified multi-modal representation of the transparent background image of the product and the text or scene description input by the user.

[0239] Optionally, in order to avoid interference from the background of the image uploaded by the user, only the transparent background image of the product is used for calculation during the process of calculating the features of the product image. This embodiment uses the GME model to convert data of different modalities (the image and text description input by the user) into vector representations with consistent semantics.

[0240] Currently, large language models perform well in natural language processing tasks. With the emergence of multimodal large language models, they can be used to process multimodal data. The GME model used in this embodiment is an architecture for improving multimodal retrieval performance, which can significantly enhance the effect of multimodal retrieval and achieve more accurate and efficient cross-modal information retrieval. The training of the GME model mainly adopts a method combining self-supervised learning and contrastive learning. Self-supervised learning enables the GME model to have basic multimodal understanding and generation capabilities, and further conducts contrastive learning through the relationship between similar and dissimilar samples to optimize the semantic representation ability of the model.

[0241] Optionally, the GME model may include components such as a multimodal encoder, a joint embedding space, and a cross-modal alignment mechanism. Among them, the multimodal encoder is the core encoder of the multimodal large language model, which can process multiple modal inputs such as text and images simultaneously. Through a unified encoder, the GME model can convert data of different modalities into semantically consistent vector representations, eliminating the representation differences between different modalities. The joint embedding space component can be used to map data of different modalities into a shared embedding space during the training process through methods such as contrastive learning. In this space, similar contents (regardless of their original modalities) have higher similarity, facilitating cross-modal retrieval. The cross-modal alignment mechanism component can be used to further improve the alignment effect between different modalities. The GME model introduces a cross-modal alignment mechanism, and by optimizing the loss function, the semantic relationship between different modalities can be better captured and maintained. Through the GME model in this embodiment, user input can be converted into a semantically consistent vector representation (for example, a 1*1536-dimensional vector).

[0242] Step S1106, construct a high-quality multimodal database.

[0243] Optionally, during the construction of the high-quality multimodal database, it is first necessary to understand the scene graph, that is, obtain the scene graph Caption, and further obtain the multimodal representation based on the scene graph and the scene graph Caption. Taking the above input image (scene graph) as an example, the following introduces how to parse a scene graph into a piece of data in the high-quality multimodal database.

[0244] After commodity matting and calculating the unified multimodal representation, the data information that can be obtained includes: the scene graph, the transparent background image of the commodity, and the unified multimodal representation (in the actual process, it is necessary to obtain the transparent background image of the commodity and the scene graph Caption first to obtain the unified multimodal representation). The information that still needs to be further obtained includes: the scene graph style, the commodity space information, the scene graph Caption, and the commodity layout information. The following further introduces the methods for obtaining the scene graph style, the commodity space information, the scene graph Caption, and the commodity layout information.

[0245] First, the acquisition of the scene graph style and commodity spatial information can be achieved by training classifiers, and separate classifiers are trained for the scene graph style and commodity spatial information respectively. In the image classification task, a pre-trained self-supervised learning vision (Distillation with No Labels Version 2, abbreviated as DINOv2) model can be used to extract image features, and then the extracted image features are input into a multi-layer perceptron MLP for classification. For different classification tasks, a small number of samples need to be manually labeled to fine-tune the classifier. When the validation accuracy reaches a certain threshold, the classifier can be used for data classification to obtain the scene graph style and commodity spatial information.

[0246] It should be noted that the DINOv2 model is an algorithm for self-supervised learning from unlabeled images in the field of computer vision. Dinov2 is good at extracting high-level features from images, and these features can perform well in a wide range of downstream tasks. The working principle of the DINOv2 model can include self-supervised learning, Vision Transformer (abbreviated as ViT) model, visual Figure 1 consistency and contrastive loss. Among them, the DINOv2 model is trained through self-supervised learning and does not require labeled data. The DINOv2 model uses the distillation technique, where a large model is used as the teacher model, which can provide learning objectives for the student model. The DINOv2 model uses the Vision Transformer ViT as the basic architecture. ViT receives an image as input, divides the image into small image regions (patches), and then encodes these patches to finally generate a series of feature representations. When the DINOv2 model is trained, different forms of data augmentation (such as scaling, cropping, color jitter, etc.) are adopted to generate multiple views of the same image. Further, the model can be trained to maintain consistent feature representations across different views. The DINOv2 model uses contrastive loss to make the model learn to make the features between different views of the same image close and separate the features between views of different images. Optionally, by loading a pre-trained DINOv2 model for image similarity retrieval, the input image is preprocessed to extract feature vectors (for example, 257*768-dimensional feature vectors), and then the similarity between the feature vectors is calculated.

[0247] Secondly, for the acquisition of the scene graph Caption, since the scene graph Caption can be used for the image generation from image background filling task, it can be realized by a multimodal technology that only describes the background content of the scene graph. By leveraging the multimodal capabilities of the multimodal generation model, methods for seed data production, seed data filtering, and model distillation can be designed to provide the multimodal technology capabilities that only describe the background content of the scene graph.

[0248] Figure 14 is a schematic diagram of obtaining the scene graph Caption according to an embodiment of the present application. As Figure 14 shown, during the seed data production process, for indoor scene images, first, a first prompt word designed manually is used to call the multimodal generation model to obtain a complete and detailed description of the scene graph. For the Caption task without unconditional restrictions, the results of the multimodal generation model are highly credible. To further obtain the Caption that only describes the background content of the scene graph, images with the main object occluded are obtained through technical means such as image segmentation and image processing. The complete and detailed description of the scene graph, the second prompt word, and the images with the main object occluded are jointly used as the input of the multimodal generation model to obtain the Caption that only describes the background content of the scene graph output by the multimodal generation model.

[0249] Although the background description of the scene graph in the above steps is input by the detailed description of the scene graph, due to the increase in the prompt word condition restrictions, there are hallucinations or errors in a relatively low proportion in the obtained seed data (image pairs: images with the main object occluded and the background description of the scene graph). Therefore, it is necessary to further filter the seed data. The background description of the scene graph is formatted into descriptions of different regions around the main object, and the correlations between the descriptions of different regions and each region of the image can be calculated respectively. When the correlation calculation result of a certain region is low, it is considered that the quality of the seed image data pair is unqualified, thus completing the filtering of the seed data. After obtaining high-quality seed data, model distillation is performed on the bidirectional pre-trained multimodal model to enable the bidirectional pre-trained multimodal model to achieve better results in the scene graph background description task. After completing the model distillation, the distilled model is evaluated, and a model whose performance meets the conditions is selected, and its multimodal capabilities are used to achieve the goal of only describing the background content of the scene graph. For example, the finally obtained scene graph Caption can be "There is a round coffee table placed in front of the sofa. The coffee table is placed on a carpet with a textured pattern, and there are books, decorative objects, and other small items placed on the coffee table. In the background of the sofa, there are artworks with abstract shapes and sculptural features placed on the shelf. There is a floor lamp placed on the left side of the sofa, and the windows of the room where the sofa is located have thin curtains, and sunlight can penetrate into the room."

[0250] Finally, for obtaining the product layout information, the product layout information mainly includes two parts: the location where the product image is located and the description of the positions of the objects in the background around the product image. Among them, the description of the positions of the objects in the background around the product image can be achieved through the method for obtaining the Caption of the above-mentioned scene graph. The following mainly introduces the method for obtaining the location where the product image is located.

[0251] Since the product transparent background image obtained by product matte extraction contains alpha channel information, therefore, the location where the product is located can be obtained according to the alpha channel information. Figure 15 It is a schematic diagram of a product bounding box according to an embodiment of the present application. As Figure 15 shown, the alpha channel can be separated from the product transparent background image, the pixel values of the alpha channel are traversed, the coordinates of the non-transparent pixels are recorded, and based on the recorded coordinates of the non-transparent pixels, the bounding box information of the product image is calculated. For example, the obtained bounding box information can be (74, 368, 733, 624) and (74, 368, 733, 624). Among them, the first coordinate can be used to represent the coordinates of the upper left corner and the lower right corner of the bounding box, and the second coordinate can be used to represent the normalized coordinates.

[0252] To sum up, a piece of data containing information in various dimensions such as scene graph, product transparent background image, unified multi-modal representation, scene graph style, product spatial information, product layout information, and scene graph Caption can be obtained. Among them, the unified multi-modal representation, scene graph style, product spatial information, and product layout information can be used in the retrieval process of the user's product image, and the scene image, scene graph style, product layout information, and scene graph Caption will be used as the recommended results. When the amount of data reaches a certain scale, a multi-modal high-quality database can be obtained.

[0253] Step S1107, obtain the multi-modal retrieval result.

[0254] In the above steps, after calculating the unified multi-modal representation of the product transparent background image and the text or scene description, the multi-modal representation can be used to perform retrieval in the multi-modal high-quality database to obtain the multi-modal retrieval result.

[0255] Optionally, during the data retrieval process, according to the user input (image and text, the text is an optional input) and various constraint conditions (such as style, space, position layout, or it can also be unconditional, only one product image), the similarity calculation is performed on the data samples in the database, and the sample with the highest similarity is selected as the retrieval result.

[0256] Optionally, after obtaining the multi-modal retrieval results (scene graph style, product space information, scene graph Caption product layout information), the retrieval results can be further aligned with the user's intention and used as the model input conditions. A complete set of input conditions all originate from the same high-quality scene graph, thus ensuring the integrity and rationality among the input conditions.

[0257] Since the image generation results of the model are largely affected by the input conditions, to ensure the differences among the output results, usually multiple significantly different input conditions need to be retrieved. To solve this problem, this embodiment first retrieves M complete input conditions (M pieces of data), further calculates the inter-class deviation among these M pieces of data, and selects N (M > N) pieces of data with larger deviations as the model input conditions.

[0258] Step S1108, align the user's intention.

[0259] In the above steps, after obtaining multiple sets of retrieval results, calculating the inter-class deviation among the multiple sets of retrieval results, and using several sets of data with larger deviations as the final retrieval results to ensure the differences among the input conditions, in order to fully align the user's intention, it is also necessary to use a vision-language model to perform Prompt rewriting on the retrieval results.

[0260] Optionally, the samples with the highest similarity are obtained from the multi-modal high-quality database through a unified multi-modal representation. Usually, these samples are close to the user's intention (user input text or scene description Prompt) but do not meet the user's needs. To fully express the user's intention, it is also necessary to further align the user's intention (i.e., Prompt rewriting).

[0261] Optionally, the Prompt rewriting method used in this embodiment is realized by fine-tuning a vision-language model. The fine-tuned vision-language model is a multi-modal large model that focuses on vision-language tasks, has the ability to process image and text information, and can generate corresponding outputs according to instructions. The fine-tuned vision-language model combines the powerful capabilities of language understanding and vision understanding. The moderate parameter scale of this model achieves a good balance between performance and resource requirements and performs well in practical applications after instruction fine-tuning. After performing instruction fine-tuning on the model, it is possible to align the retrieved samples with the user input.

[0262] Optionally, according to the user input, rewrite the scene graph caption in the retrieved sample to the result aligned with the user intention (the [] area is the marked rewritten result, and the result before rewriting can be seen in the above scene graph caption), and use the rewritten retrieved sample as the model input condition (i.e., Prompt): "There is a [round marble coffee table] placed in front of the sofa. The coffee table is placed on a carpet with textured graphics, and there are books, decorative objects, and other small items placed on the coffee table. In the background of the sofa, there is a [decorative painting] with abstract shapes and sculptural features placed on the shelf. There is a [fishing lamp] placed on the left side of the sofa. The window of the room where the sofa is located has thin curtains, and sunlight can penetrate into the room."

[0263] Step S1109, obtain the input condition.

[0264] In the above steps, after using the vision-language model to perform Prompt rewriting on the retrieval results, the above-mentioned multiple sets of model input conditions with differences can be respectively used for generating images of goods.

[0265] Step S1110, generate images of goods.

[0266] In the above steps, multiple scene images containing goods that are aligned with the user intention, have a reasonable layout, high aesthetics, and are differentiated can be obtained.

[0267] Optionally, StableDiffusion, as a text-to-image model, not only allows creators to edit the generated images, but also this model is open-source and can run on a consumer-grade Graphics Processing Unit (GPU for short). StableDiffusion is a diffusion model based on latent space. It introduces text condition in U2Net to generate images based on text. The core of StableDiffusion comes from latent diffusion. Conventional diffusion models are pixel-based generation models, while latent diffusion is a latent-based generation model. StableDiffusion first uses an autoencoder to compress the image into the latent space, then uses the diffusion model to generate the latent variables of the image, and finally sends them into the decoder module of the autoencoder to obtain the generated image. Since the latent space of the image is smaller than the pixel space of the image, the computational efficiency of the latent-based diffusion model is higher.

[0268] Figure 16 It is a schematic diagram of generating an image based on text according to an embodiment of the present application, as Figure 16As shown, the input text can be "an astronaut riding a horse". According to the text encoder, text feature vectors or text embeddings are extracted from the input text. At the same time, a random noise is initialized through a Random Number Generator (RNG for short). Further, the text embeddings and the noise are input into the diffusion model to generate the denoised Latent. Finally, the denoised Latent is input into the decoder module of the autoencoder to obtain the generated image.

[0269] Image-to-Image is an extension of the Text-to-Image function. The principle of Image-to-Image is that given a reference image, a certain amount of Gaussian noise can be added to the reference image (performing the diffusion process) to obtain a noisy image, and then the noisy image is denoised based on the diffusion model to generate a new image, and this image is basically consistent with the input image in terms of structure and layout.

[0270] Figure 17 is a schematic diagram of generating an image based on an image according to an embodiment of the present application, as Figure 17 shown. Compared with Text-to-Image, the noise Latent is obtained by adding Gaussian noise to the Latent after the initial image is encoded by the autoencoder (i.e., noise + Latent). At the same time, it should be noted that the number of steps in the denoising process is the same as the number of steps in the noise addition process, so as to generate the required noise-free image.

[0271] Figure 18 is a schematic diagram of an optional image generation result, Figure 19 is a schematic diagram of another optional image generation result, Figure 18 and Figure 19 are both results generated by manually selecting different styles twice. Comparing Figure 9 with Figure 10 of this embodiment, it can be seen that the image generation results obtained in this embodiment have significant advantages in terms of user intention alignment, layout rationality, aesthetics, and homogenization.

[0272] In this embodiment, by using a vision-language model to align the user's intention, the problem that the user's intention expression does not meet the input conditions of the image processing model is solved; by constructing a multi-modal high-quality database and a unified multi-modal representation, the problems such as unreasonable layout and insufficient aesthetics of the generated scene images are solved; in the process of obtaining the model input conditions, data with large inter-class deviations can be selected from the recommended multiple groups of data to ensure the diversity of product image generation, thus solving the problem of homogenization of the generated scene images.

[0273] This embodiment uses a multi-modal encoding model to perform unified multi-modal representation and retrieval of user input to obtain input conditions that fully consider the multi-modal data of the user input; uses a vision-language model to obtain the input of the image generation model that aligns with the user's intention (i.e., Prompt); and solves problems such as difficult user input and homogenization of image generation results by calculating the correlation between the user input and the multi-modal high-quality database, thereby improving the image generation effect.

[0274] In the embodiment of the present application, the initial input information is converted into target input information, and the information format of the target input information matches that of the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input not to meet the input conditions of the image processing model and further affect the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0275] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0276] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0277] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0278] The embodiment of the present application also provides an image generation system. Figure 20 It is a schematic diagram of an image generation system according to an embodiment of the present application. As Figure 20 shown, the image generation system 2000 may include: a client 2002 and a server 2004.

[0279] The client 2002 is used to upload initial input information, where the initial input information is used to represent at least one object to be displayed.

[0280] The server 2004 is used to convert the initial input information into target input information, where the information format of the target input information matches the image processing model; in the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold, where the database includes at least multiple modalities of information samples, and the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; use the image processing model to analyze the input information samples to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene; and send the scene image to the client for display.

[0281] In the image generation system 2000, initial input information is uploaded through the client 2002, where the initial input information is used to represent at least one object to be displayed. The server 2004 converts the initial input information into target input information, where the information format of the target input information matches the image processing model; in the database, at least one input information sample with a similarity greater than the similarity threshold to the target input information is retrieved, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; the image processing model is used to analyze the input information samples to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene; and the scene image is sent to the client for display.

[0282] That is to say, the image generation system converts the initial input information into target input information through the server, and the information format of the target input information matches the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model and thus affect the efficiency of the scene image generated by the image processing model, thereby achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0283] According to an embodiment of the present application, there is also provided an image generation device for implementing the above Figure 2 shown image generation method.

[0284] Figure 21 is a schematic diagram of an image generation device according to an embodiment of the present application, as Figure 21 shown. The image generation device 2100 may include: a first acquisition unit 2102, a first conversion unit 2104, a first retrieval unit 2106, and a first generation unit 2108.

[0285] The first acquisition unit 2102 is used to acquire initial input information, where the initial input information is used to represent at least one object to be displayed.

[0286] The first conversion unit 2104 is used to convert the initial input information into target input information, where the information format of the target input information matches the image processing model.

[0287] The first retrieval unit 2106 is configured to retrieve at least one input information sample in a database, the similarity between which and the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, where the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information.

[0288] The first generation unit 2108 is configured to analyze the input information sample by using an image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0289] It should be noted here that the above first acquisition unit 2102, first conversion unit 2104, first retrieval unit 2106, and first generation unit 2108 correspond to steps S202 to S208 in the above embodiment. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be part of a device and can run in the server 10 provided in the above embodiment.

[0290] In the image generation device, the first acquisition unit 2102 acquires initial input information, where the initial input information is used to represent at least one object to be displayed. The first conversion unit 2104 converts the initial input information into target input information, where the information format of the target input information matches that of the image processing model. The first retrieval unit 2106 retrieves at least one input information sample in a database, the similarity between which and the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, where the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information. The first generation unit 2108 analyzes the input information sample by using an image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

[0291] That is to say, in the embodiment of the present application, the initial input information is converted into target input information, and the information format of the target input information matches the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model and thus affect the efficiency of the scene image generated by the image processing model. Therefore, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0292] According to an embodiment of the present application, there is also provided an image generation device for implementing the above Figure 3 image generation method shown.

[0293] Figure 22 is a schematic diagram of another image generation device according to an embodiment of the present application, as shown in Figure 22 shown. The image generation device 2200 may include: a first display unit 2202 and a second display unit 2204.

[0294] The first display unit 2202 is configured to respond to an input instruction on the operation interface and display the initial input information on the operation interface, where the initial input information is used to represent at least one object to be displayed.

[0295] The second display unit 2204 is configured to respond to an image generation instruction on the operation interface and display a scene image corresponding to the initial input information on the operation interface, where the scene image is used to represent the display result of the object to be displayed in the initial input information in the target scene to be displayed. The scene image is obtained by analyzing the input information sample using an image processing model. The input information sample is an information sample in the database whose similarity to the target input information is greater than a similarity threshold. The database includes at least multiple modalities of information samples. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene in which the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information. The target input information is obtained by converting the initial input information, and the information format of the target input information matches the image processing model.

[0296] It should be noted here that the above first display unit 2202 and second display unit 2204 correspond to steps S302 to S304 in the above embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above module or unit may be a hardware component or a software component stored in a memory and processed by one or more processors. The above module may also be part of a device and may run in the server 10 provided in the above embodiment.

[0297] In the image generation device, the first display unit 2202 responds to an input instruction on the operation interface and displays initial input information on the operation interface, where the initial input information is used to represent at least one object to be displayed. The second display unit 2204 responds to an image generation instruction on the operation interface and displays a scene image corresponding to the initial input information on the operation interface. The scene image is used to represent the display result of the object to be displayed in the initial input information in the target scene to be displayed. The scene image is obtained by analyzing an input information sample using an image processing model. The input information sample is an information sample in the database whose similarity to the target input information is greater than a similarity threshold. The database includes at least multiple modalities of information samples. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene where the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information. The target input information is obtained by converting the initial input information, and the information format of the target input information matches that of the image processing model.

[0298] That is to say, in the embodiment of the present application, in response to an image generation instruction on the operation interface, a scene image corresponding to the initial input information is displayed on the operation interface. The scene image is obtained by analyzing an input information sample using an image processing model. The input information sample is the sample with the highest similarity to the target input information converted from the initial input information retrieved from the database. Since the information format of the target input information matches that of the image processing model, and the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may cause the input to not meet the input conditions of the image processing model, thereby affecting the efficiency of the scene image generated by the image processing model, and thus achieving the technical effect of improving the image generation efficiency and solving the technical problem of low image generation efficiency.

[0299] According to the embodiment of the present application, there is also provided an image generation device for implementing the Figure 4 image generation method shown above.

[0300] Figure 23Schematic diagram of another image generation device according to an embodiment of the present application, as Figure 23 shown, the image generation device 2300 may include: a second acquisition unit 2302, a second conversion unit 2304, a second retrieval unit 2306, a second generation unit 2308, and a sending unit 2310.

[0301] The second acquisition unit 2302 is configured to acquire initial input information from an e-commerce platform, where the initial input information is used to represent at least one product to be displayed.

[0302] The second conversion unit 2304 is configured to convert the initial input information into target input information, where the information format of the target input information matches that of the image processing model.

[0303] The second retrieval unit 2306 is configured to retrieve at least one input information sample in the database whose similarity to the target input information is greater than a similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the displayed product samples and the attribute information of the scenes where the displayed product samples are displayed, and the information quality of the input information samples is higher than that of the target input information.

[0304] The second generation unit 2308 is configured to analyze the input information samples using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the product to be displayed in the target scene where it is displayed.

[0305] The sending unit 2310 is configured to send the scene image to the e-commerce platform.

[0306] It should be noted here that the above-mentioned second acquisition unit 2302, second conversion unit 2304, second retrieval unit 2306, second generation unit 2308, and sending unit 2310 correspond to steps S402 to S410 in the above embodiment. The instances and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of the device and may run in the server 10 provided in the above embodiment.

[0307] In the image generation device, the second acquisition unit 2302 acquires initial input information from an e-commerce platform, where the initial input information is used to represent at least one product to be displayed. The second conversion unit 2304 converts the initial input information into target input information, where the information format of the target input information matches that of the image processing model. The second retrieval unit 2306 retrieves at least one input information sample from a database, the similarity between which and the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, and the information samples are used to represent the attribute information of the displayed product samples and the attribute information of the scenes where the displayed product samples are located. The information quality of the input information samples is higher than that of the target input information. The second generation unit 2308 analyzes the input information samples using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the product to be displayed in the target scene. The sending unit 2310 sends the scene image to the e-commerce platform.

[0308] That is to say, in the embodiment of the present application, the initial input information obtained from the e-commerce platform is converted into target input information, and the information format of the target input information matches that of the image processing model. Then, at least one input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may lead to the input not meeting the input conditions of the image processing model and further affecting the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0309] According to the embodiment of the present application, there is also provided an image generation device for implementing the above Figure 5 shown image generation method.

[0310] Figure 24 FIG. is a schematic diagram of another image generation device according to the embodiment of the present application. As Figure 24 shown, the image generation device 2400 may include: a third acquisition unit 2402, a third conversion unit 2404, a third retrieval unit 2406, a third generation unit 2408, and an output unit 2410.

[0311] The third acquisition unit 2402 is configured to acquire initial input information by invoking a first interface. The first interface includes a first parameter, and the parameter value of the first parameter is the initial input information. The initial input information is used to represent at least one object to be displayed.

[0312] A third conversion unit 2404, configured to convert the initial input information into target input information, where the information format of the target input information matches the image processing model.

[0313] A third retrieval unit 2406, configured to retrieve at least one input information sample in a database, where the similarity between the retrieved input information sample and the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, and the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are located. The information quality of the input information samples is higher than that of the target input information.

[0314] A third generation unit 2408, configured to analyze the input information sample by using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is located.

[0315] An output unit 2410, configured to output the scene image by invoking a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the scene image.

[0316] It should be noted here that the above-mentioned third acquisition unit 2402, third conversion unit 2404, third retrieval unit 2406, third generation unit 2408, and output unit 2410 correspond to steps S502 to S510 in the above-mentioned embodiment. The functions, implementation examples, and application scenarios of the five modules are the same as those of the corresponding steps, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules may also be part of a device and may run in the server 10 provided in the above-mentioned embodiment.

[0317] In the image generation device, the third acquisition unit 2402 obtains initial input information by invoking a first interface. The first interface includes a first parameter, and the parameter value of the first parameter is the initial input information, which is used to represent at least one object to be displayed. The third conversion unit 2404 converts the initial input information into target input information, where the information format of the target input information matches the image processing model. The third retrieval unit 2406 retrieves at least one input information sample from the database whose similarity to the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, and the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed. The information quality of the input information samples is higher than that of the target input information. The third generation unit 2408 analyzes the input information samples using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene. The output unit 2410 outputs the scene image by invoking a second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the scene image.

[0318] That is to say, in the embodiments of the present application, the initial input information obtained by invoking the first interface is converted into target input information, and the information format of the target input information matches the image processing model. Then, the input information sample with the highest similarity to the target input information is retrieved from the database. Since the information quality of the retrieved input information sample is higher than that of the target input information, the input information sample is used as the input of the image processing model, avoiding directly using the target input information as the input, which may lead to the input not meeting the input conditions of the image processing model and further affecting the efficiency of the scene image generated by the image processing model. Thus, the technical effect of improving the image generation efficiency is achieved, and the technical problem of low image generation efficiency is solved.

[0319] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the above embodiments, but are not limited to the schemes provided in the above embodiments.

[0320] Embodiments of the present application can provide a computing device. Figure 25 is a structural block diagram of a computing device according to an embodiment of the present application, as Figure 25 shown. The computing device 2500 may include: one or more (only one is shown in the figure) processors 2502, a memory 2504, a storage controller, and a peripheral interface.

[0321] The above-mentioned computing device can be understood as an integrated intelligent terminal, including but not limited to servers, desktop computers, personal computers (PCs for short), model all-in-ones, etc. Moreover, the computing device may be pre-installed with the models described in the above embodiments of the present application.

[0322] Specifically, the computing device may be pre-installed with various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device may support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effects of the model based on model evaluation tools), etc. In other product forms, the computing device can also create applications based on the model, provide API invocation capabilities, and can call the model into the created application through the API interface, while providing application management tools to realize the control of the application.

[0323] Furthermore, the computing device may further include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.

[0324] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to terminal A through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and their combinations.

[0325] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.

[0326] Embodiments of the present application may provide an electronic device. Figure 26 It is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 26 shown, the electronic device may include: an input / output device 2602; a memory 2604, and a processor 2606. Among them, the processor 2606 is connected to the input / output device 2602 and the memory 2604 through a bus 2608.

[0327] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories can be connected to the terminal A through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0328] The processor can call the executable program stored in the memory through a transmission device to execute the method described in any one of the above embodiments.

[0329] Those of ordinary skill in the art can understand that the structure shown in the figure is only schematic. The computing device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (abbreviated as MID), a PAD and other terminal devices. This figure does not limit the structure of the above computing device. For example, the computing device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0330] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory ROM, a random access memory RAM, a magnetic disk, or an optical disc, etc.

[0331] Embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.

[0332] Optionally, in this embodiment, the above storage medium may be located in the computing device.

[0333] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method described in any one of the above embodiments.

[0334] The embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method provided in the above embodiment.

[0335] The embodiment of the present application also provides a computer program product. Optionally, the above computer program product may include a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium may be used to store a computer program, and when the computer program is executed by a processor, it implements the method provided in the above embodiment.

[0336] The embodiment of the present application also provides a computer program. Optionally, in this embodiment, when the above computer program is executed by a processor, it implements the method provided in the above embodiment.

[0337] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0338] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0339] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0340] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0341] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memory ROM, random access memory RAM, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0342] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for generating an image, characterized in that, Including: Obtain initial input information, where the initial input information is used to represent at least one object to be displayed; Convert the initial input information into target input information, where the information format of the target input information matches that of the image processing model; In the database, retrieve at least one input information sample whose similarity to the target input information is greater than the similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; Use the image processing model to analyze the input information sample to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed.

2. The method according to claim 1, wherein The initial input information includes at least an initial object image, and the image content of the initial object image includes the object to be displayed. Converting the initial input information into target input information includes: At least encode the initial object image to obtain the target input information, where the information format of the target input information is a vector format.

3. The method according to claim 2, wherein The initial input information further includes text description information, which is used to describe the generation intention of the to-be-generated scene image. At least encoding the initial object image to obtain the target input information includes: Input the initial object image and the text description information into a multi-modal encoding model for encoding respectively to obtain the target input information, where the multi-modal encoding model is obtained by training a large language model.

4. The method according to claim 3, characterized in that Input the initial object image and the text description information into a multi-modal encoding model for encoding respectively to obtain the target input information, including: In response to the background image in the initial object image being a non-solid-color background image, convert the background image in the initial object image into a solid-color background image to obtain the converted initial object image; Input the converted initial object image and the text description information into the multi-modal encoding model for encoding respectively to obtain the target input information.

5. The method according to claim 1, wherein In the database, retrieving at least one input information sample whose similarity to the target input information is greater than the similarity threshold includes: In the database, retrieve multiple initial input information samples whose similarities to the target input information are respectively greater than the similarity threshold; Among the multiple initial input information samples, determine multiple input information samples, where the difference information between every two input information samples among the multiple input information samples is greater than the difference information threshold.

6. The method according to claim 1, characterized in that The method further includes: Obtain at least one scene image sample, where the scene image sample is used to represent the display result of the display object sample in the corresponding scene, and the image quality of the scene image sample is higher than the quality threshold; Based on the scene image sample, determine the information sample to obtain the database.

7. The method according to claim 6, characterized in that, Based on the scene image samples, determining the information samples to obtain the database includes: Generating text description information samples based on the scene image samples, where the text description information samples are used to describe the image content of the scene image samples; Determining the information samples based on the text description information samples and the scene image samples to obtain the database.

8. The method according to claim 7, wherein The information format includes a vector format. Based on the text description information samples and the scene image samples, determining the information samples to obtain the database includes: Obtaining the image information of the display object samples from the scene image samples; Generating the object sample images of the display object samples from the image information of the display object samples, where the image content of the object sample images includes the display object samples on a solid-color background image; Converting the object sample images and the text description information samples into feature vectors in the vector format respectively; Determining at least the feature vectors in the vector format as the information samples to obtain the database.

9. The method according to claim 8, wherein Determining at least the feature vectors in the vector format as the information samples to obtain the database includes: Determining the feature vectors in the vector format, and at least one of the following information as the information samples to obtain the database: the scene style of the scene image samples, the spatial type samples of the space where the display object samples are located in the corresponding scene, the layout information samples of the layout of the display object samples in the corresponding scene, the object sample images, the scene image samples, and the text description information samples.

10. The method according to any one of claims 1 to 9, characterized in that, Analyzing the input information samples using the image processing model to obtain a scene image, including: Adjusting the text description information samples in the input information samples using the initial input information to obtain the adjusted input information samples, where the initial input information is used to describe the generation intention of the to-be-generated scene image, the text description information samples are used to describe the image content of the scene image samples in the input information samples, and the adjusted input information samples match the generation intention; Generating prompt information from the adjusted input information samples, where the prompt information is used to guide the process of the image processing model analyzing the adjusted input information samples; Guiding the image processing model to analyze the adjusted input information samples using the prompt information to obtain the scene image.

11. The method according to claim 10, characterized in that, Adjusting the text description information samples in the input information samples using the initial input information to obtain the adjusted input information samples, including: Inputting the initial input information and the input information samples into a vision-language model, and using the vision-language model to adjust the text description information samples in the input information samples according to the initial input information to obtain the adjusted input information samples, where the vision-language model is obtained by training a multi-modal large model.

12. A method for generating an image, characterized in that, Including: In response to an input instruction on the operation interface, initial input information is displayed on the operation interface, where the initial input information is used to represent at least one object to be displayed. In response to an image generation instruction on the operation interface, a scene image corresponding to the initial input information is displayed on the operation interface, where the scene image is used to represent the display result of the object to be displayed in the target scene represented by the initial input information. The scene image is obtained by analyzing an input information sample using an image processing model. The input information sample is an information sample in the database whose similarity to the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities. The information sample is used to represent the attribute information of the display object sample and the attribute information of the scene in which the display object sample is displayed. The information quality of the input information sample is higher than that of the target input information. The target input information is obtained by converting the initial input information, and the information format of the target input information matches that of the image processing model.

13. The method according to claim 12, characterized in that, The initial input information includes at least an initial object image, and the image content of the initial object image includes the object to be displayed. Displaying the scene image corresponding to the initial input information on the operation interface in response to an image generation instruction on the operation interface includes: In response to a recommended mode selection instruction on the operation interface, enter the recommended mode. In the recommended mode, in response to the image generation instruction on the operation interface, the scene image is displayed on the operation interface, where the target input information in the recommended mode is obtained by encoding the initial object image, and the information format of the target input information is a vector format.

14. The method according to claim 12, wherein The initial input information includes an initial object image and text description information. The image content of the initial object image includes the object to be displayed, and the text description information is used to describe the generation intention of the scene image to be generated. Displaying the scene image corresponding to the initial input information on the operation interface in response to an image generation instruction on the operation interface includes: In response to a custom mode selection instruction on the operation interface, enter the custom mode. In the custom mode, in response to the image generation instruction on the operation interface, the scene image is displayed on the operation interface, where the target input information in the custom mode is obtained by encoding the initial object image and the text description information respectively into a multi-modal encoding model, the information format of the target input information is a vector format, and the multi-modal encoding model is obtained by training a large language model.

15. The method according to claim 12, characterized in that The method further includes: In response to an information selection instruction on the operation interface, constraint condition information is displayed on the operation interface, where the input information sample is an information sample in the database whose similarity to the target input information is greater than the similarity threshold and that satisfies the constraint condition information.

16. A method for generating an image, characterized in that, Including: Obtain initial input information from an e-commerce platform, where the initial input information is used to represent at least one product to be displayed; Convert the initial input information into target input information, where the information format of the target input information matches that of an image processing model; Retrieve at least one input information sample in a database whose similarity to the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of a displayed product sample and the attribute information of the scene where the displayed product sample is displayed, and the information quality of the input information sample is higher than that of the target input information; Analyze the input information sample using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the product to be displayed in the target scene where it is displayed; Send the scene image to the e-commerce platform.

17. A method for generating an image, characterized in that, Including: Obtain initial input information by calling a first interface, where the first interface includes a first parameter whose parameter value is the initial input information, and the initial input information is used to represent at least one object to be displayed; Convert the initial input information into target input information, where the information format of the target input information matches that of an image processing model; Retrieve at least one input information sample in a database whose similarity to the target input information is greater than a similarity threshold. The database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of a displayed object sample and the attribute information of the scene where the displayed object sample is displayed, and the information quality of the input information sample is higher than that of the target input information; Analyze the input information sample using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed; Output the scene image by calling a second interface, where the second interface includes a second parameter whose parameter value includes the scene image.

18. An image generation system, characterized in that, Including: A client for uploading initial input information, where the initial input information is used to represent at least one object to be displayed; A server for converting the initial input information into target input information, where the information format of the target input information matches that of an image processing model; retrieving at least one input information sample from a database whose similarity to the target input information is greater than a similarity threshold, where the database includes at least information samples of multiple modalities, the information samples are used to represent the attribute information of the display object samples and the attribute information of the scenes where the display object samples are displayed, and the information quality of the input information samples is higher than that of the target input information; analyzing the input information samples using the image processing model to obtain a scene image, where the scene image is used to represent the display result of the object to be displayed in the target scene where it is displayed; and sending the scene image to the client for display.

19. A computing device, characterized in that, Comprising: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the method according to any one of claims 1 to 17.

20. An electronic device, characterized in that, Comprising: A memory storing an executable program; A processor connected to the memory through a bus for running the program, where when the program runs, it executes the method according to any one of claims 1 to 17.

21. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 17.

22. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 17.