Image generation method and program product
By combining multiple initial control conditions to generate target control conditions, and guiding the image generation model with injection parameters and initial prompt information, the problem of poor image generation effect in the prior art is solved, and more natural and effective image generation is achieved.
Patent Information
- Application Number
- CN202411345237.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-25
AI Technical Summary
When combining multiple Control Nets in the prior art, control failure and unnatural image formation may occur, resulting in poor image generation effect.
By acquiring multiple initial control conditions, combining them into target control conditions, determining the injection parameters based on the target control conditions, and guiding the image generation model to generate the image to be generated using the injection parameters and initial prompt information.
This has achieved the improvement of image generation effect and solved the problem of poor image generation effect.
Smart Images

Figure CN118864653B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to large model technology and image processing fields, and in particular to an image generation method and a program product. Background Art
[0002] At present, in order to flexibly control the generated image, multiple control conditions are usually injected into the model (ControlNet) and used in combination, and the image is completed through the combined Control Net. However, when combining multiple Control Nets, the above method is prone to control failure and unnatural images, and there is a technical problem of poor image generation effect.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present invention provide a method for generating an image and a program product to at least solve the technical problem of poor image generation effect.
[0005] According to one aspect of an embodiment of the present application, a method for generating an image is provided. The method may include: obtaining multiple initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent the attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; combining multiple of the initial control conditions to obtain the target control conditions of the image generation model; based on the target control conditions, determining the injection parameters of the image generation model, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; using the injection parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0006] According to another aspect of the embodiment of the present application, another method for generating an image is also provided. The method may include: obtaining a control image from a client; converting the control image into an initial control condition corresponding to the image to be generated, wherein the initial control condition is used to represent the attribute information of the image to be generated and is used to guide the image generation model to generate the image to be generated; combining multiple initial control conditions to obtain the target control condition of the image generation model; based on the target control condition, determining the injection parameter of the image generation model, wherein the injection parameter is used to characterize the degree of influence of the target control condition on the attribute information; using the injection parameter and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0007] According to another aspect of the embodiment of the present application, another method for generating an image is also provided. The method may include: displaying a display object in a virtual store on an e-commerce platform on a presentation screen of an operation interface; generating and displaying an image to be generated including the display object in response to a generation instruction acting on the operation interface; wherein the image to be generated is generated by guiding an image generation model using injection parameters and initial prompt information of the image to be generated, the initial prompt information is used to prompt the image content of the image to be generated under the attribute information of the image to be generated, the injection parameters are determined based on the target control conditions, and are used to characterize the degree of influence of the target control conditions on the attribute information, the target control conditions are obtained by combining multiple initial control conditions, the initial control conditions are used to represent the attribute information, and are used to guide the image generation model to generate the image to be generated.
[0008] According to another aspect of the embodiments of the present application, a computer terminal is further provided, including: a memory storing an executable program; and a processor for running the program, wherein the method in each embodiment of the present application is executed when the program is running.
[0009] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in each embodiment of the present application.
[0010] According to another aspect of the embodiments of the present application, a computer program product is also provided, including a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0011] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0012] According to another aspect of the embodiments of the present application, a computer program is also provided, and when the computer program is executed by a processor, the methods in the various embodiments of the present application are implemented.
[0013] In an embodiment of the present application, multiple initial control conditions corresponding to the image to be generated are obtained, wherein the initial control conditions are used to represent the attribute information of the image to be generated and to guide the image generation model to generate the image to be generated; multiple initial control conditions are combined to obtain the target control conditions of the image generation model; based on the target control conditions, the injection parameters of the image generation model are determined, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; the image generation model is guided to generate the image to be generated using the injection parameters and the initial prompt information of the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information. That is, in an embodiment of the present application, multiple initial control conditions are obtained, multiple initial control conditions are combined to obtain the target control conditions, and the injection parameters are determined based on the target control conditions. By using the injection parameters, multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected, and using the injection parameters and the initial prompt information to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the effect of image generation and solving the technical problem of poor image generation effect.
[0014] It is easy to notice that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0016] Figure 1 is a schematic diagram of an application scenario of an image generation method according to an embodiment of the present application;
[0017] Figure 2 is a flow chart of a method for generating an image according to an embodiment of the present application;
[0018] Figure 3 is a flow chart of another method for generating an image according to an embodiment of the present application;
[0019] Figure 4 is a flow chart of another method for generating an image according to an embodiment of the present application;
[0020] Figure 5 is a schematic diagram of a combination of multiple image control conditions according to an embodiment of the present application;
[0021] Figure 6 is a schematic diagram of an initial control condition combination according to an embodiment of the present application;
[0022] Figure 7 is a schematic diagram of a feature injection according to an embodiment of the present application;
[0023] Figure 8 is a schematic diagram of a model training according to an embodiment of the present application;
[0024] Fig. 9 is a schematic diagram of generating an image to be generated according to an embodiment of the present application;
[0025] Fig.10 It is a hardware structure block diagram of a computer terminal (or mobile device) according to an image generation method of an embodiment of the present application;
[0026] Fig.11 is a schematic diagram of an image generating device according to an embodiment of the present application;
[0027] Fig.12 is a schematic diagram of another image generating device according to an embodiment of the present application;
[0028] Fig.13 is a schematic diagram of another image generating device according to an embodiment of the present application;
[0029] Fig.14 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] The technical solution provided in this application is mainly implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than 10 trillion model parameters. The large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as large-scale language model (Large Language Model, referred to as LLM), multi-modal pre-training model, etc.
[0033] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned through a small number of samples so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields, and can be specifically applied to computer vision tasks such as visual question answering (VQA), image description (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiment of the present application, data processing through an image generation model in an image generation scenario is used as an example for explanation.
[0034] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0035] The control condition injection model can be an image deformation algorithm based on control points, which can also be called a control network. It can be a network that adds image control signals based on the Wensheng graph model. It can realize image-based control generation by training an encoder. The control image can include edge lines, key point postures of the human body, etc., which can be used to achieve fine control of the generated image.
[0036] Multiple ControlNet combinations can refer to multiple image control conditions being used to control image generation simultaneously;
[0037] Silent control conditions may refer to image control conditions without significant control object signals, such as areas without edge lines or key points of the human body;
[0038] High-frequency information can refer to the parts or details of an image that change dramatically or rapidly;
[0039] Conservatism can refer to the ability of a model to remain stable and robust when making predictions or generating data;
[0040] Feature combination can refer to the process of combining or merging features to generate new features;
[0041] Feature injection can refer to the process of injecting additional features or information into the model to help the model learn and predict better;
[0042] The deep learning text-to-image generation model (Stable Diffusion) can be used to generate detailed images based on text and can be a potential diffusion model.
[0043] According to an embodiment of the present application, a method for generating an image is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] Considering the huge number of model parameters of the large model and the limited computing resources of the mobile terminal, the above-mentioned image generation method provided in the embodiment of the present application can be applied to the following Figure 1 The application scenarios shown are not limited to these. Figure 1 is a schematic diagram of an application scenario of an image generation method according to an embodiment of the present application, such as Figure 1 As shown, in Figure 1 In the application scenario shown, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiment of the present application.
[0045] In an embodiment of the present application, a system composed of a client device and a server can execute the following steps: Step S102, obtaining multiple initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent the attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; Step S104, combining multiple initial control conditions to obtain target control conditions of the image generation model; Step S106, determining injection parameters of the image generation model based on the target control conditions, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; Step S108, using the injection parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0046] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device.
[0047] Under the above operating environment, this application provides Figure 2 The method for generating the image shown. Figure 2 FIG. 1 is a flow chart of a method for generating an image according to an embodiment of the present application. Figure 2 As shown, the method may include the following steps:
[0048] Step S202, obtaining a plurality of initial control conditions corresponding to the image to be generated.
[0049] In the technical solution provided in the above step S202 of the present application, the above image to be generated can be an image to be generated by the image generation model under the guidance of the initial control conditions, for example, it can be a beautifully made product creative picture, etc. The above initial control conditions can be used to represent the attribute information of the image to be generated, and to guide the image generation model to generate the image to be generated. It can be a condition for guiding the image generation model to generate a specific image to be generated. It can be a control signal (or a combination of image control signals), which can be a user input and used to control the depth, action posture, edge and other information of the image generation model. It can be a signal of the type of image, text, etc. For example, it can be a depth map, edge detection map, posture key point map, etc. It should be noted that this is only for example, and there is no specific restriction on the type of initial control conditions. The above attribute information can be used to characterize the boundary information of different areas in the image to be generated, the action information, depth information, posture information of different objects, etc. It should be noted that this is only for example, and there is no specific restriction on the type of attribute information. The above-mentioned image generation model can be used to generate an image to be generated based on text or an image, and can be a potential diffusion model, for example, it can be a deep learning text-to-image generation model, or an image generation model based on a generative adversarial network (GAN). It should be noted that this is only an example, and there is no specific restriction on the type of image generation model.
[0050] Optionally, by using multiple initial control conditions, information such as the shape and edge of the elements in the graphics to be generated can be controlled.
[0051] For example, if a landscape painting needs to be generated, the edge map can be used as the initial control condition. The image generation model can obtain the initial control condition input by the user. The attribute information of the image to be generated can be controlled by the initial control condition. By obtaining multiple initial control conditions, the attribute information can be controlled from different dimensions. For example, the edge map and the depth map can be used as the initial control condition, and the edge and depth of the area or person in the image to be generated can be controlled by the initial control condition.
[0052] Step S204, combining multiple initial control conditions to obtain target control conditions of the image generation model.
[0053] In the technical solution provided in step S204 of the present application, multiple initial control conditions can be obtained, and the multiple initial control conditions can be combined to obtain the target control condition. The target control condition can be a combination of multiple initial control conditions, for example, the combination condition is "according to the depth Figure 1 , and the edge Figure 2 Generate the image to be generated".
[0054] Optionally, in order to guide the image generation model to generate an image to be generated with a specific combination, multiple initial control conditions may be combined to obtain a target control condition, and the image generation model may be guided by the target control condition to obtain the image to be generated.
[0055] For example, to obtain multiple initial control conditions, the combination ratio of different initial control conditions can be pre-set, and the multiple initial control conditions can be combined according to the combination ratio to obtain the target control condition. It should be noted that this is only an example, and there is no specific limitation on the combination of multiple initial control conditions. The process of combining multiple initial control conditions to obtain the target control condition should be within the scope of protection of this application.
[0056] Step S206, determining the injection parameters of the image generation model based on the target control conditions.
[0057] In the technical solution provided in the above step S206 of the present application, the above injection parameters can be used to characterize the degree of influence of the target control conditions on the attribute information in the generated image to be generated.
[0058] Optionally, the target control condition is transformed to obtain an injection parameter, which can be used to guide the image generation model to generate the image to be generated.
[0059] Step S208, using the injected parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0060] In the technical solution provided in the above step S208 of the present application, the above initial prompt information (prompt) can be used to guide the image content under the attribute information of the image generation model, which can be text information, image information, etc. For example, it can be a text description, a label, a keyword, etc. For example, the initial prompt information can be "a sunny beach with a parasol and a deck chair". The above image content can include objects, environments, areas, etc. in the image to be generated. For example, it can be the action of the object under the depth information. It should be noted that this is only an example, and there is no specific limitation on the type of initial prompt information and the information contained in the image content.
[0061] Optionally, based on the target control condition, an injection parameter of the image generation model is determined, and the injection parameter is combined with the initial prompt information to guide the image generation model to generate the image to be generated.
[0062] For example, multiple initial control conditions corresponding to the image to be generated are obtained, and the multiple initial control conditions are combined to obtain the target control conditions of the image generation model. According to the target control conditions, the injection parameters that need to be adjusted by the image generation model are determined. For example, if the target control condition is a green cat, the injection parameters may include color and shape. Furthermore, the injection parameters and the initial prompt information of the image to be generated are used to guide the image generation model to generate the image to be generated.
[0063] When combining multiple initial control conditions, it is necessary to determine the ratio of the various initial control signals to be combined to obtain a suitable target control signal, and inject the target control signal into the initial prompt information according to the determined ratio to guide the image generation model to generate the image to be generated according to the injected initial prompt information. However, the above method only relies on the user's experiment to determine the predicted ratio, which may make the image generation model unable to effectively balance the different initial control conditions, and unable to effectively balance the target control conditions and the initial prompt information, thereby leading to the technical problem of poor image generation effect. Therefore, in this embodiment, in order to solve the above problem, by determining a balanced feature combination and feature injection, based on the target control condition, the injection parameters of the image generation model are determined, and the injection parameters are combined with the initial prompt information to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0064] Through the above steps S202 to S210, multiple initial control conditions corresponding to the image to be generated are obtained, wherein the initial control conditions are used to represent the attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; multiple initial control conditions are combined to obtain the target control conditions of the image generation model; based on the target control conditions, the injection parameters of the image generation model are determined, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; the injection parameters and the initial prompt information of the image to be generated are used to guide the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information. That is, in the embodiment of the present application, multiple initial control conditions can be obtained, multiple initial control conditions can be combined to obtain the target control conditions, and the injection parameters can be determined based on the target control conditions. By using the injection parameters, multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected, and using the injection parameters and the initial prompt information to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the effect of image generation and solving the technical problem of poor image generation effect.
[0065] The above method of this embodiment is further introduced below.
[0066] As an optional implementation, step S206, based on the target control condition, determines the injection parameters of the image generation model, including: obtaining the original feature vector of the initial prompt information; and determining the injection parameters using the target feature vector and the original feature vector of the target control condition.
[0067] In this embodiment, the original feature vector may be a vector representation corresponding to an image, or a vector representation corresponding to a text, and may be used to represent the content of the original prompt information. For example, it may be an encoder (U-Net) feature, and may be used to represent the content of the original prompt information. It should be noted that this is only an example, and there is no specific restriction on the type and expression of the original feature vector.
[0068] Optionally, the original feature vector of the initial prompt information can be obtained by feature extraction and the target feature vector and the original feature vector of the target control condition can be used to determine the injection parameter. The injection parameter can be used to adjust the image to be generated.
[0069] If the user only sets the ratio of the target feature vector to be injected into the original feature vector according to the needs, there will be a technical problem of poor image generation effect. In this embodiment, in order to solve the technical problem of poor image generation effect, the original feature vector of the initial prompt information is determined, and the injection ratio of the target feature vector to the original feature vector is determined by using the relationship between the original feature vector and the target feature vector. The target feature vector is adjusted by using the injection ratio to obtain an injection parameter, and the injection parameter is injected into the original feature vector to adjust the image to be generated based on the original feature vector.
[0070] Optionally, in this embodiment, feature injection is used to dynamically adjust the image to be generated by the image generation model, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0071] For example, the initial prompt information is obtained, and the feature extraction is performed on the initial prompt information to obtain the original feature vector of the initial prompt information. The original feature vector can be the pixel value of the image or the word vector representation of the text. It should be noted that there is no specific restriction on the type of the original feature vector. Furthermore, the association relationship between different target feature vectors, original feature vectors and injection parameters can be constructed in advance. Based on the association relationship, the injection parameters corresponding to the current target feature vector and the original feature vector can be determined. For example, the sum between the target feature vector and the original feature vector can be determined in advance, and they form acute angles with the target feature vector and the original feature vector, respectively. Under this condition, the corresponding injection parameters can be determined. It should be noted that there is no specific restriction on the method of determining the injection parameters and the type of the original feature vector, and it can be selected according to actual needs.
[0072] One method of determining injection parameters is further described below.
[0073] As an optional implementation, the target feature vector and the original feature vector of the target control condition are used to determine the injection parameters, including: determining first weight data based on the target feature vector and the original feature vector, wherein the first weight data is used to represent the degree of influence of the target feature vector on the original feature vector; and using the first weight data, converting the target feature vector into an injection parameter.
[0074] In this embodiment, the first weight data is used to represent the influence of the target feature vector on the original feature vector, and can be used to determine the adjustment degree of the target feature vector on the image to be generated.
[0075] Optionally, first weight data may be determined based on the target feature vector and the original feature vector, and the target feature vector may be transformed using the first weight data to obtain the injection parameters.
[0076] Optionally, in order to balance the relationship between the injected target feature vector and the original feature vector, the first weight data may be determined based on the target feature vector and the original feature vector. Further, the product between the first weight data and the target feature vector may be determined to obtain the injection parameter. The first weight data may be determined by the following formula ( ):
[0077]
[0078]
[0079] Optionally, the injection parameters may be determined according to the following formula;
[0080]
[0081] As an optional implementation, step S208, using the injection parameters and the initial prompt information of the image to be generated, guides the image generation model to generate the image to be generated, including: injecting the injection parameters into the initial prompt information to obtain target prompt information; using the target prompt information, guides the image generation model to generate the image to be generated.
[0082] In this embodiment, the target prompt information can be determined by the following formula ( ):
[0083]
[0084] It should be noted that the above is only one of the methods for determining injection parameters and is not limited thereto. As long as the injection parameters are determined based on the target feature vector and the original feature vector, they should be within the protection scope of this application.
[0085] Optionally, by utilizing target prompt information, the image generation model may be guided to generate the image to be generated.
[0086] As an optional implementation, injecting the injection parameter into the initial prompt information to obtain the target prompt information includes: injecting the injection parameter into the original feature vector of the initial prompt information to obtain the injected original feature vector, wherein the injected original feature vector is used to characterize the target prompt information.
[0087] In this embodiment, the injection parameter may be injected into the initial prompt information to obtain an injected original feature vector, and the injected original feature vector may be used to determine the target prompt information.
[0088] For example, the initial prompt information is determined to be "generate a landscape painting", and the target control condition is determined to be "generate an ink painting style", "mountains and flowing water" and other elements. Based on the original feature vector of the initial prompt information and the target feature vector of the target control condition, the first weight data is determined, and the target control condition is converted based on the first weight data to obtain an injection parameter, and the injection parameter is injected into the original feature vector of the initial prompt information to obtain the injected original feature vector, which can be used to characterize the target prompt information and can be used to guide the image generation model to generate a landscape painting with ink painting style and mountains and flowing water elements.
[0089] As an optional implementation, step S204, combining multiple initial control conditions to obtain target control conditions of the image generation model, includes: determining the feature vector of the initial control conditions; combining the feature vectors of multiple initial control conditions to obtain the target feature vector; determining the target control condition corresponding to the target feature vector.
[0090] In this embodiment, the feature vector of the initial control condition is determined, and the feature vectors of multiple initial control conditions are combined to obtain a target feature vector. The target feature vector can be used to characterize the target control condition and can be used to determine the result after combining multiple initial control conditions. The above feature vector can be used to characterize at least one feature corresponding to the initial control condition, can be used to characterize the color, texture, shape of the image, or can be used to characterize the theme of the text, the frequency of occurrence of words, etc., can be binary coded, or a feature vector set, can be used It should be noted that this is only an example and there is no specific limitation on the content and form of expression of the feature vector.
[0091] Optionally, after obtaining multiple initial control conditions, feature extraction can be performed on the initial control conditions to obtain feature vectors of the initial control conditions, and multiple feature vectors corresponding to the multiple initial control conditions can be obtained. Feature vectors output by the same neural network layer can be combined to obtain a target feature vector, and the target control condition can be obtained using the target feature vector. It should be noted that the method for determining the feature vector of the initial control condition is not specifically limited here.
[0092] In this embodiment, after obtaining multiple initial control conditions, the combination ratio of the multiple initial control conditions that meets the requirements can be determined, that is, the combination ratio of multiple Control Nets can be determined. After the initial control conditions are obtained, the matching initial control conditions are processed using Control Net to obtain feature vectors of the multiple initial control conditions. The multiple feature vectors of the multiple initial control conditions can be combined according to the combination ratio to obtain the final target control vector.
[0093] When combining multiple Control Nets, it is necessary to determine the ratio of the various initial control signals to be combined. In the sampling stage, if the initial control signals are combined as plug-ins based solely on user experiments, the image generation model may not be able to effectively balance the different initial control signals, which may lead to a technical problem of poor image generation. Therefore, in this implementation, in order to solve the above problem, by determining the combination ratio of multiple Control Nets, balanced feature combination and feature injection are achieved, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0094] The method of combining the characteristic vectors of multiple initial control conditions is further explained below.
[0095] As an optional implementation, the characteristic vectors of multiple initial control conditions are combined to obtain a target characteristic vector, including: determining second weight data corresponding to the characteristic vector, wherein the second weight data is used to represent the importance of the characteristic vector to the determined target control condition; and combining the characteristic vectors of multiple initial control conditions according to the second weight data to obtain the target characteristic vector.
[0096] In this embodiment, the second weight data corresponding to the feature vector is determined. The second weight data can be used to determine the degree of influence of the feature vector on the target control condition. According to the second weight data, the feature vectors of multiple initial control conditions can be combined to obtain the target feature vector. The second weight data can be used express, , They can be used to represent different feature vectors respectively.
[0097] Optionally, the second weight data corresponding to the feature vectors can be determined respectively, and the multiple feature vectors can be combined according to the second weight data to obtain the target feature vector. The above second weight data can be determined based on the feature vector, or can be determined in advance according to the control type of the initial control condition. It should be noted that the method for determining the second weight data is not specifically limited here.
[0098] The following further describes how to determine the second weight based on the feature vector.
[0099] As an optional implementation, determining the second weight data corresponding to the feature vector includes: selecting a target number of feature vectors from multiple feature vectors; calling a combination strategy to determine the second weight data corresponding to the selected feature vectors, wherein the combination strategy is used to represent the association relationship between multiple feature vectors.
[0100] In this embodiment, the above-mentioned target number can be a number pre-set according to actual needs, for example, it can be 2. It should be noted that there is no specific restriction on the number of targets here. The above-mentioned combination strategy can be a pre-set strategy, which can be used to express the association relationship between multiple feature vectors. For example, it can be used to limit the sum of the feature vectors finally obtained by multiple feature vectors to form an acute angle with each feature vector. It should be noted that the combination strategy can be changed according to actual conditions, and there is no specific restriction on the type of combination strategy here.
[0101] Optionally, for feature combinations, in order to balance different initial control signals in the sampling stage, the characteristic vector of the initial control signal can be determined, and by utilizing a combination strategy, each of the multiple characteristic vectors is controlled to form an acute angle with the sum of the multiple characteristic vectors, thereby achieving the purpose of balancing the multiple initial control signals before injecting the target characteristic vector into the image generation model, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0102] For example, assuming that the number of targets is 2, two feature vectors can be selected from multiple feature vectors, and a combination strategy can be called to determine the second weight data corresponding to the two feature vectors when the angle between the two feature vectors and the sum of the two feature vectors is an acute angle. It should be noted that the sum of the two second weight data corresponding to the two feature vectors is 1.
[0103] In this embodiment, a target number of feature vectors can be selected, that is, multiple feature vectors can be grouped into multiple sets containing two feature vectors. According to the combination strategy, the second weight data corresponding to each set containing two feature vectors can be determined respectively. According to the second weight data, two feature vectors in the same set are added to obtain the added feature vector, and the feature vectors corresponding to multiple sets are obtained. The feature vectors corresponding to the multiple sets are determined according to the combination strategy, and the weight data is added according to the weight data to obtain the target feature vector. Alternatively, two feature vectors can be selected first, and the second weight data corresponding to the two feature vectors can be calculated. According to the second weight data, the selected two feature vectors are combined to obtain the combined feature vector, and then a feature vector is reselected, and the selected feature vector is combined with the combined feature vector. The combination of multiple feature vectors is completed in the above manner, and the feature vector finally obtained by the combination can be determined as the target feature vector. It should be noted that the way in which multiple feature vectors are combined can be selected according to actual conditions, and no specific restrictions are made here.
[0104] For example, the combination strategy can be set in advance to control the sum of the two feature vectors of the target number, and the angle between each feature vector is an acute angle. Further, the feature vectors of the target number are selected from multiple feature vectors; according to the combination strategy, the second weight data ( ), the second weight data can be determined by the following formula:
[0105]
[0106]
[0107] It should be noted that the above , Can be used to characterize feature vectors. As the selected feature vector changes, , The object of representation also changes.
[0108] Furthermore, by performing weighted summation on the two feature vectors using the calculated second weight data, the result after weighted summation can be obtained by the following formula ( ):
[0109]
[0110] in, , can be used to represent the feature vector, and , The corresponding eigenvectors are the same.
[0111] That is, in this embodiment, the second weight data can be calculated according to the above formula, and then according to The weighted processing is performed on the multiple feature vectors to obtain the weighted results of the multiple feature vectors. According to the above process, the weighted summation is performed on the multiple feature vectors to obtain the final target feature vector.
[0112] As an optional implementation, the characteristic vectors of multiple initial control conditions are combined according to the second weight data to obtain a target characteristic vector, including: combining the characteristic vectors of the target number according to the second weight data corresponding to the characteristic vectors of the target number to obtain a first target characteristic vector; deleting the characteristic vectors of the target number from multiple characteristic vectors, and selecting any characteristic vector from the multiple characteristic vectors after deletion; using the first target characteristic vector and any selected characteristic vector, determining the second weight data corresponding to any selected characteristic vector; combining the first target characteristic vector and any selected characteristic vector according to the second weight data corresponding to any selected characteristic vector to obtain the target characteristic vector.
[0113] In this embodiment, a target number of feature vectors are selected from multiple feature vectors, and the target number of feature vectors are combined to obtain a first target feature vector. Further, the target number of feature vectors are deleted from the multiple feature vectors, and any feature vector is reselected from the deleted multiple feature vectors; the first target feature vector and any selected feature vector are used to determine the second weight data corresponding to any selected feature vector; according to the second weight data corresponding to any selected feature vector, the first target feature vector and any selected feature vector are combined to obtain the target feature vector.
[0114] Optionally, a target number of feature vectors are selected from multiple feature vectors, and the target number of feature vectors are combined according to second weight data to obtain a first target feature vector, a feature vector is reselected, the second weight data corresponding to the selected feature vector is determined, and the selected feature vector and the first target feature vector are combined according to the second weight data to further obtain the target feature vector.
[0115] For example, if there are three feature vectors, two feature vectors may be selected at random and combined to obtain a first target feature vector, and the first target feature vector and the remaining feature vectors may be combined to obtain a target feature vector.
[0116] As an optional implementation, the first target feature vector and any selected feature vector are combined according to the second weight data corresponding to any selected feature vector to obtain the target feature vector, including: a combining step, according to the second weight data corresponding to any selected feature vector, the first target feature vector and any selected feature vector are combined to obtain a second target feature vector, wherein the second target feature vector and the target feature vector are different; a deleting step, from multiple feature vectors, deleting any selected feature vector, and re-selecting any feature vector from at least one feature vector after deletion; a looping step, using the second target feature vector and any feature vector selected again, determining the second weight data corresponding to any feature vector selected again, and combining the second target feature vector and any selected feature vector according to the second weight data corresponding to any feature vector selected again to obtain a third target feature vector; repeating the deleting step and the looping step until all the multiple feature vectors are combined, and determining the third target feature vector determined based on the feature vector selected last time as the target feature vector.
[0117] In this embodiment, the target number of feature vectors are combined according to the second weight data corresponding to the target number of feature vectors to obtain a first target feature vector; the target number of feature vectors are deleted from multiple feature vectors, and any feature vector is selected from the multiple feature vectors after deletion; the second weight data corresponding to any selected feature vector is determined using the first target feature vector and any selected feature vector, and then the combination step, the deletion step and the loop step are sequentially implemented to obtain the final target feature vector.
[0118] Optionally, according to the second weight data corresponding to the feature vectors of the target number, the feature vectors of the target number are combined to obtain a first target feature vector; the feature vectors of the target number are deleted from multiple feature vectors, and any feature vector is selected from the multiple feature vectors after deletion; the second weight data corresponding to any selected feature vector is determined using the first target feature vector and any selected feature vector. Further, a combining step can be implemented: according to the second weight data corresponding to any selected feature vector, the first target feature vector and any selected feature vector are combined to obtain a second target feature vector; after determining the second target feature vector, a deleting step can be implemented. , delete any selected feature vector from multiple feature vectors, and select any feature vector again from at least one feature vector after deletion, and perform a loop step, that is, use the second target feature vector and any feature vector selected again to determine the second weight data corresponding to any feature vector selected again, and combine the second target feature vector and any selected feature vector according to the second weight data corresponding to any feature vector selected again to obtain a third target feature vector; repeat the deletion step and the loop step until all multiple feature vectors are combined, and determine the third target feature vector determined based on the feature vector selected last time as the target feature vector.
[0119] Optionally, the characteristic vectors of multiple initial control conditions are determined to obtain multiple characteristic vectors. If there are only two characteristic vectors, the two characteristic vectors can be combined using the second weight data, and the combined characteristic vector is determined as the target characteristic vector. If there are three characteristic vectors, two characteristic vectors can be selected first, and the two characteristic vectors can be combined according to the second weight data corresponding to the selected characteristic vectors to obtain a first target characteristic vector, and the first target characteristic vector and the unselected characteristic vector are combined to obtain the target characteristic vector.
[0120] Optionally, if there are more than three eigenvectors, two eigenvectors can be selected first, and the two eigenvectors can be combined according to the second weight data corresponding to the selected eigenvectors to obtain a first target eigenvector, and the two selected eigenvectors from the multiple eigenvectors can be deleted, and a eigenvector can be reselected, and the second weight data of the reselected eigenvector can be determined. According to the second weight data, the reselected eigenvector and the first target eigenvector can be combined to obtain a second target eigenvector. Further, the deletion step can be performed again on the multiple eigenvectors, that is, a selected eigenvector can be deleted from the multiple eigenvectors, and a eigenvector can be selected again from the remaining eigenvectors. According to the above steps, all eigenvectors in the multiple eigenvectors are combined to obtain the final target eigenvector. At this time, the target eigenvector obtained is a eigenvector after balancing different initial control conditions.
[0121] For example, assuming that the number of targets is 2 and the number of feature vectors is 4. Select 2 feature vectors from the 4 feature vectors and calculate the second weight data according to the following formula:
[0122]
[0123]
[0124] According to the second weight data, the two selected feature vectors are combined to obtain the first target feature vector, the two feature vectors that have been calculated among the four feature vectors are deleted, and a feature vector is arbitrarily selected from the remaining two feature vectors, and the second weight data of the selected feature vector is determined. According to the second weight data, the feature vector and the first target feature vector are combined to obtain the second target feature vector, and the second weight data of the last remaining feature vector is determined. According to the second weight data, the second target feature vector and the last remaining feature vector are combined to obtain the target feature vector.
[0125] As an optional implementation, determining the feature vector of the initial control condition includes: determining the control type of the initial control condition, wherein the control type is used to represent the control content of the initial control condition; calling a control condition injection model that matches the control type, wherein the control condition injection model is trained using enhanced control condition samples, and the control condition samples are used to adjust the attribute information samples of the image samples to be generated; and inputting the initial control condition into the corresponding control condition injection model to obtain a feature vector.
[0126] In this embodiment, the above control type can be used to represent the control content of the initial control condition, which can include the control of action posture, depth, image edge, and soft edge. It should be noted that there is no specific limitation on the control type here. The above control condition injection model can be trained using enhanced control condition samples, can be a ControlNet network, or can be a neural network architecture for conditional image generation, which is used to control the generation process of the image to be generated using additional initial control conditions (e.g., guidance signals).
[0127] Optionally, an enhanced control condition sample is obtained, and a control condition injection model is obtained by training the enhanced control condition sample, wherein different types of enhanced control condition samples can be used to train control condition injection models corresponding to different control types. After the initial control condition is obtained, the control type of the initial control condition can be determined, and a control condition injection model matching the control type can be called, and the initial control condition can be input into the control condition injection model to obtain a feature vector of the initial control condition.
[0128] In this embodiment, control condition injection models corresponding to different control types are trained separately. In the sampling stage (i.e., the model testing stage), the trained control condition injection models can be combined according to different initial control conditions and used as plug-ins, and the trained control condition injection models are used to process the corresponding initial control conditions, thereby ensuring flexibility and modularity in the data processing process, so that various control conditions can be seamlessly integrated, thereby achieving the purpose of enhancing the image generation capability of the model, achieving the technical effect of improving the effect of image generation, and solving the technical problem of poor image generation effect.
[0129] For example, the initial control condition "a hand-drawn landscape sketch" is obtained, and the control type of the initial control condition is determined to be "edge map" or "contour map". A pre-trained control condition injection model that can generate realistic images under the guidance of the edge map is obtained. The control condition injection model can be trained using a large number of image data sets with edge maps (that is, enhanced control condition samples). The enhanced control condition samples can enable the control condition injection model to learn how to use the edge map to generate the corresponding image to be generated. Furthermore, the initial control condition can be input into the corresponding control condition injection model to obtain a feature vector.
[0130] As an optional implementation, the method may also include: obtaining an image sample to be generated, wherein the image sample to be generated includes multiple regions; calling a preprocessing module of a sub-control condition injection model to preprocess the image sample to be generated to obtain a control condition sample, and determining segmentation mask data of the image sample to be generated, wherein the segmentation mask data is used to characterize edge contours of different regions; using the segmentation mask data, enhancing the control condition sample to obtain an enhanced control condition sample.
[0131] In this embodiment, the image sample to be generated (Original Image) may include multiple regions, and there are edge transition parts (i.e., boundaries) between different regions. During the image processing process, there may be silent control conditions (also referred to as silent signals) in the boundary regions. In this embodiment, in order to avoid the influence of silent control conditions on the control condition injection model, during the training process of the sub-control condition injection model, the control condition samples can be enhanced using segmentation mask data to obtain enhanced control condition samples, and the sub-control condition injection model can be trained using the enhanced control condition samples to obtain a trained control condition injection model. The control condition injection model can well identify different regions in the initial control conditions, thereby improving the image quality of the generated image to be generated.
[0132] Optionally, the above-mentioned segmentation mask data can be a specific type of annotation data that can be used to characterize the edge contours of different regions, can be applied to image segmentation tasks, can be used to characterize the label corresponding to each pixel in the image, and the label can be used to determine the object or category represented by the pixel, and can be used to characterize the position or shape of different objects or regions in the image.
[0133] Optionally, segmentation mask data of the image sample to be generated is obtained by manual annotation, semi-automatic annotation, automatic segmentation algorithm, or a hybrid method of automatic segmentation and manual correction. For example, the image can be automatically annotated through a segmentation network to obtain the segmentation mask data. It should be noted that this is only an example and there is no specific limitation on the method of obtaining the segmentation mask.
[0134] For example, during the training process of the sub-control condition injection model, a sample of an image to be generated can be obtained, and the sample of the image to be generated can include multiple regions or multiple objects. The preprocessing module of the sub-control condition injection model can be called, and the preprocessing module can preprocess the sample of the image to be generated to obtain a control condition sample. Furthermore, the segmentation mask data of the image sample to be generated can be determined, wherein the segmentation mask data is used to characterize the edge contours of different regions; using the segmentation mask data, the control condition sample can be enhanced to obtain an enhanced control condition sample. The above-mentioned enhancement processing can be to use the segmentation mask as an additional information layer of the image during the image processing process. It should be noted that this is only an example, and there is no specific limitation on the method of enhancement processing.
[0135] During the training process of the ControlNet model, there will be silent control conditions, such as edge condition signals. The image areas corresponding to these silent control conditions are usually blurred or lack high-frequency information. Therefore, data bias will occur during the training process, causing the model to suppress high-frequency information in the generated image. In a single-control scenario, this suppression is beneficial for strict generation, but there will be certain problems in a multi-ControlNet combination scenario. For example, when two initial control conditions coexist in an area, one has high-frequency information and the other is a silent control condition, a conflict will occur. Under the influence of the silent control condition, the model may unnecessarily suppress the high-frequency information in the image to be generated. In order to solve the above problem, in this embodiment, a data enhancement method is used to balance the distribution of areas lacking control conditions through data enhancement to help the model generate high-frequency information in the silent control condition area, thereby rebalancing the distribution of areas lacking control conditions.
[0136] Optionally, the implementation can apply segmentation mask data of the image samples to be generated on the initial control conditions to increase the diversity of the image regions corresponding to the silent control conditions. Through the above method, the image regions corresponding to the silent control conditions can be patched to help the model generate high-frequency information in the silent control condition regions, thereby reducing data deviations in the training process, achieving the technical effect of improving the image generation effect, and solving the technical problem of poor image generation effect.
[0137] As an optional implementation, the method may also include: adding noise information to the image sample to be generated to obtain a noise sample; calling the sub-image generation model corresponding to the image generation model to convert the noise sample and the enhanced control condition sample to obtain a generated image sample, wherein the sub-image generation model at least includes a sub-control condition injection model; determining the target model parameters of the sub-control condition injection model based on the image sample to be generated, the noise sample, the generated image sample, and the enhanced control condition sample; and using the target model parameters to adjust the initial model parameters of the sub-control condition injection model to obtain a control condition injection model.
[0138] In this embodiment, the noise information may be noise point information, illumination information, blur information, color distortion information, etc. By adding noise information to the image samples to be generated, the diversity and robustness of the image samples to be generated may be increased.
[0139] Optionally, the sub-graph generation model may be a pre-built diffusion model, which may include a sub-control condition injection model and a codec (eg, an original U-Net).
[0140] Optionally, in the process of training the control condition injection model, a sample of an image to be generated can be obtained, and the sample of the image to be generated can be a training sample, and can be an original image used in the training of the control condition injection model. Noise information can be added to the sample of the image to be generated to obtain a noise sample. A sub-image generation model corresponding to the image generation model can be called, and the noise sample and the enhanced control condition sample can be converted using the sub-image generation model to obtain a generated image sample (Image x t ). Furthermore, based on the image samples to be generated, the noise samples, the generated image samples and the enhanced control condition samples, the target model parameters of the sub-control condition injection model can be determined, and the initial model parameters of the sub-control condition injection model can be updated using the target model parameters to obtain a trained control condition injection model.
[0141] For example, obtain the image sample to be generated, add noise information to the image sample to be generated, and obtain the noise sample (Image x t). Set the initial prompt information used in the training process, which can be a prompt word (prompt) and a generation time (time). Get the control condition sample, such as the edge map (Canny Edge), add segmentation mask data (Mask) to the control condition sample, and get the enhanced control condition sample (Masked Canny Edge c), call the codec (including encoder and decoder) in the sub-image generation model, process the noise sample, prompt word, generation time, and the data output by the sub-control condition injection model, and get the generated image sample (S t ). Based on the image samples to be generated, the noise samples, the generated image samples and the enhanced control condition samples, the target model parameters of the sub-control condition injection model can be determined, and the initial model parameters of the sub-control condition injection model can be updated using the target model parameters to obtain a trained control condition injection model.
[0142] The following further describes how the target model parameters of the sub-control condition injection model can be determined based on the to-be-generated image samples, the noise samples, the generated image samples and the enhanced control condition samples.
[0143] Since the sub-control condition injection model is adjusted by using less training data to obtain a trained control condition injection model, when combining multiple control condition injection models, it is necessary to consider enhancing the conservatism of the score function to ensure the stability and performance of the trained control condition injection model.
[0144] Optionally, after determining the output data 808, the Jacobian matrix corresponding to the model is calculated, and then an estimation method (eg, Hutchinson) is used to estimate the parameters in the Jacobian matrix, thereby completing an unbiased estimation of the loss function.
[0145] In this embodiment, a conditional score function corresponding to the sub-image generation model can be constructed in advance ( ):
[0146]
[0147] right An unbiased estimate can be made to obtain the estimated conservative loss:
[0148]
[0149] Among them, x t Can be used to characterize noise samples, S t Can be used to characterize output data. It can be used to characterize the Jacobian matrix associated with noise samples and output data.
[0150] Optionally, the Jacobian matrix of the sub-image generation model ( ) can be decomposed into two parts: the Jacobian matrix of the original U-Net (which can include the encoder and decoder) ( ) and the additional Jacobian matrix introduced by Control Net ( ):
[0151]
[0152] Since there is a huge gap between the amount of training data of the original U-Net and Control Net, this embodiment converts the Jacobian matrix ( ) and the additional Jacobian matrix introduced by Control Net ( ).
[0153] Optionally, based on the responsibility of conservatism, in the process of equipping the U-Net model with Control Net, the conservatism of the original Jacobian matrix is controlled by the original model parameters of U-Net. At the same time, the original model parameters of Control Net are mainly responsible for managing the additional conservatism introduced by Control Net, which can be described as:
[0154]
[0155] Among them, ϕ can be used to characterize the initial model parameters of the sub-control condition injection model.
[0156] Furthermore, we can ignore Not included The part (that is, excluding the estimated conservative loss corresponding to the original U-Net in the sub-image generation model) is used to obtain the loss function for optimizing the sub-control condition injection model:
[0157]
[0158] The gradient of the loss function can be set equal to the gradient of the estimated conservative loss:
[0159]
[0160] Due to computational limitations, it is still desirable to remove , therefore, the simplified loss function for Control Net optimization can be obtained:
[0161]
[0162] Optionally, a simplified loss function is used to process the generated image samples, noise samples, generated image samples and enhanced control condition samples to determine the target model parameters of the sub-control condition injection model, thereby obtaining a trained control condition injection model.
[0163] Optionally, based on the above content, The Frobenius norm is uniformly constrained by M, so that if the simplified loss is zero, the original loss is also zero.
[0164]
[0165] Optionally, during the model training process, the model parameters of the original U-Net remain unchanged, and the original model parameters of the sub-control condition injection model are continuously updated.
[0166] In an embodiment of the present application, multiple initial control conditions are obtained, combined to obtain target control conditions, and injection parameters are determined based on the target control conditions. By utilizing the injection parameters, the multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected. The injection parameters and initial prompt information are used to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0167] The present application embodiment also provides another method for generating an image for the image generation scenario. Figure 3 is a flowchart of another method for generating an image according to an embodiment of the present application, such as Figure 3 As shown, the method may include the following steps:
[0168] Step S302: Acquire a control image from the client.
[0169] In the technical solution provided in the above step S302 of the present application, a control image can be obtained from the client, and the control image can be data such as a depth image and an edge image. It should be noted that this is only an example and there is no specific limitation on the type of the control image.
[0170] For example, a control image input by a user through a display interface of a client may be obtained, or a control image transmitted by a user through a web page of a client may be obtained. It should be noted that this is only an example, and the method of obtaining the control image is not specifically limited.
[0171] Step S304: converting the control image into initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent attribute information of the image to be generated and to guide the image generation model to generate the image to be generated.
[0172] In the technical solution provided in the above step S304 of the present application, the control image can be converted by using a preprocessing module to obtain the initial control conditions corresponding to the image to be generated. For example, the control image can be converted by using a preprocessing module corresponding to the depth image to obtain a depth image, or the control image can be converted by using a preprocessing module corresponding to the edge image to obtain an edge image. It should be noted that this is only an example, and there is no specific limitation on the method of obtaining the initial control conditions.
[0173] Step S306, combining multiple initial control conditions to obtain target control conditions for the image generation model.
[0174] Step S308, determining the injection parameters of the image generation model based on the target control condition, wherein the injection parameters are used to characterize the degree of influence of the target control condition on the attribute information.
[0175] Step S310, using the injected parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0176] In an embodiment of the present application, a control image is obtained from a client; the control image is converted into initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent the attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; a plurality of initial control conditions are combined to obtain target control conditions of the image generation model; based on the target control conditions, injection parameters of the image generation model are determined, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; the injection parameters and the initial prompt information of the image to be generated are used to guide the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0177] According to an embodiment of the present application, a method for generating an image of a virtual store in an e-commerce platform is also provided. Figure 4 is a flowchart of another method for generating an image according to an embodiment of the present application, such as Figure 4 As shown, the method may include the following steps:
[0178] Step S402: displaying display objects in the virtual store on the e-commerce platform on the presentation screen of the operation interface.
[0179] In the technical solution provided in step S402 of the present application, the e-commerce platform may be an online e-commerce platform. The virtual store may be a store in a shopping software, for example, an online store. The display object may be a commodity in a virtual store, or a commodity in a commodity pool of a virtual store, for example, all clothing, jewelry, accessories, etc. in an online store. This is only an example, and no specific limitation is imposed on the type of display object. The operation interface may be a display interface for interacting with a user, and may be used to display a display object and generate a ready-to-be-generated image.
[0180] Step S404, in response to a generation instruction on the operation interface, generating and displaying an image to be generated including a display object.
[0181] In the technical solution provided in the above step S404 of the present application, the above generation instruction can be used to generate an image to be generated including a display object, and can be triggered by a user.
[0182] Optionally, when the user wants to obtain the image to be generated of the display object, the user can use the operation interface to trigger the generation instruction. In response to the generation instruction, the system obtains multiple initial control conditions corresponding to the image to be generated, combines the multiple initial control conditions, obtains the target control conditions of the image generation model, determines the injection parameters that need to be adjusted for the image generation model according to the target control conditions, and uses the injection parameters and the initial prompt information of the image to be generated to guide the image generation model to generate the image to be generated.
[0183] For example, when a user wants to quickly generate an advertising image containing a display object, the user can input initial prompt information, multiple initial control conditions, and a display object through the operation interface, and trigger a generation instruction. In response to the generation instruction, the system can obtain multiple initial control conditions input by the user, combine multiple initial control conditions, obtain the target control conditions of the image generation model, determine the injection parameters that need to be adjusted for the image generation model according to the target control conditions, and use the injection parameters and the obtained initial prompt information to guide the image generation model to generate the image to be generated, which can be an advertising image (or creative image) that meets the user's needs. Furthermore, the user can make further adjustments to the generated image to be generated according to their own needs.
[0184] In this embodiment, display objects in a virtual store on an e-commerce platform are displayed on a presentation screen of an operation interface, and an image to be generated including the display object is generated and displayed in response to a generation instruction on the operation interface, wherein the image to be generated is generated by guiding an image generation model using injection parameters and initial prompt information of the image to be generated, the initial prompt information is used to prompt the image content of the image to be generated under the attribute information of the image to be generated, the injection parameters are determined based on target control conditions, and are used to characterize the degree of influence of the target control conditions on the attribute information, the target control conditions are obtained by combining multiple initial control conditions, the initial control conditions are used to represent the attribute information, and are used to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0185] At present, in order to enable users (such as advertisers) to flexibly control the generated images, or to create exquisite creative images for products, a method for controlling the combination of multiple images is provided. Figure 5 is a schematic diagram of a combination of multiple image control conditions according to an embodiment of the present application, such as Figure 5 As shown, Figure 5 The left side shows a combination of various image control conditions, and the right side shows the image results generated under a combination of various image control conditions. Figure 5 On the right, the foreground pixels of the vase are kept, and the image is completed through Control Net control; the shapes of elements such as tables and plants are controlled by control conditions, and the edges of elements are controlled by edge Control Net. However, when combining multiple control conditions, the above method is prone to control failure and unnatural images, which results in poor image quality.
[0186] To solve the above problems, in this embodiment, an optimization method for the cultural image model under multi-image control conditions is proposed. This method achieves the purpose of improving the effectiveness of control conditions and the quality of mapping through three effective strategies, namely data enhancement, balanced feature combination and feature injection, and increasing the conservative loss function. It achieves the technical effect of improving the mapping effect and solves the technical problem of poor mapping effect.
[0187] The following is a further introduction to an optimization method of a cultural graph model under multi-image control conditions proposed in an embodiment of the present application.
[0188] As an optional embodiment, control condition injection models corresponding to different control types are trained separately.
[0189] Optionally, in the sampling phase (i.e., the model testing phase), the trained control condition injection model can be combined according to different initial control conditions and used as a plug-in. The trained control condition injection model is used to process the corresponding initial control conditions, thereby ensuring flexibility and modularity in the data processing process, allowing various control conditions to be seamlessly integrated, thereby achieving the purpose of enhancing the image generation capability of the model.
[0190] During the training process of the Control Net model, there will be silent control conditions, such as edge condition signals. The image areas corresponding to these silent control conditions are usually blurred or lack high-frequency information. Therefore, data bias will occur during the training process, causing the model to suppress high-frequency information in the generated image. In a single-control scenario, this suppression is beneficial for strict generation, but there will be certain problems in a multi-Control Net combination scenario. For example, when two initial control conditions coexist in an area, one has high-frequency information and the other is a silent control condition, a conflict will occur. Under the influence of the silent control condition, the model may unnecessarily suppress the high-frequency information in the image to be generated. In order to solve the above problem, in this embodiment, a data enhancement method is used to balance the distribution of areas lacking control conditions through data enhancement to help the model generate high-frequency information in the silent control condition area, thereby rebalancing the distribution of areas lacking control conditions.
[0191] Optionally, the implementation can apply segmentation mask data of the image samples to be generated on the initial control conditions to increase the diversity of the image regions corresponding to the silent control conditions. Through the above method, the image regions corresponding to the silent control conditions can be patched to help the model generate high-frequency information in the silent control condition regions, thereby reducing data bias during the training process.
[0192] As an optional embodiment, balanced feature combination and feature injection are proposed so that the synthesized feature vector forms an acute angle with each feature vector participating in the combination, so as to achieve the purpose of reducing feature vector conflicts.
[0193] In this embodiment, after obtaining multiple initial control conditions, the combination ratio of the multiple initial control conditions that meet the requirements can be determined, that is, the combination ratio of multiple Control Nets (that is, control condition injection models) can be determined. After obtaining the initial control conditions, the matching initial control conditions are processed using Control Net to obtain feature vectors of the multiple initial control conditions. The multiple feature vectors of the multiple initial control conditions can be combined according to the combination ratio to obtain the final target control conditions.
[0194] When combining multiple ControlNets, it is necessary to determine the ratio of the various initial control signals to be combined. In the sampling stage, if the initial control signals are combined as plug-ins based on user experiments, the image generation model may not be able to effectively balance the different initial control signals, which may lead to the technical problem of poor image generation. Therefore, in this implementation, in order to solve the above problem, by determining the combination ratio of multiple ControlNets, balanced feature combination and feature injection are achieved, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0195] In this embodiment, for the feature combination, in order to balance different initial control signals in the sampling stage, the characteristic vector of the initial control signal can be determined, and by controlling the sum of multiple characteristic vectors to form an acute angle with each characteristic vector, the purpose of balancing multiple initial control signals is achieved before injecting the target characteristic vector into the image generation model.
[0196] Figure 6 is a schematic diagram of an initial control condition combination according to an embodiment of the present application, such as Figure 6 As shown, the second weight data of the feature vector 601 can be determined, and by using the second weight data, the sum of the feature vector 601 and the feature vector 602 is controlled to form an acute angle with the feature vector 601 and the feature vector 602 respectively.
[0197] In this embodiment, after the target control vector is acquired, the target control vector may be converted into an injection parameter and injected into the original feature vector.
[0198] For feature injection, the final sum vector obtained by feature injection and feature combination can be controlled to form an acute angle with each vector that died in addition. Figure 7 is a schematic diagram of a feature injection according to an embodiment of the present application, such as Figure 7 As shown, the feature vectors of multiple initial control conditions are combined to obtain the target feature vector 701 (Combined Feature). The original feature vector 702 (Feature of U-Net) is obtained, the first weight data is determined, and the injection parameters are obtained for the target feature vector 701 based on the first weight data. The injection parameters are injected into the original feature vector 702 to obtain the injected original feature vector 703 (Feasible Domain). Among them, the angles between the injected original feature vector and the original feature vector 702, the target feature vector 703, and the multiple initial feature vectors are acute angles.
[0199] Optionally, by controlling the acute angle between the injected original feature vector and each feature vector involved in the addition, the control signal is injected into each layer of the neural network, and the coefficient of the original feature vector is maintained. First, the first weight data of the injected target feature vector can be calculated by the following formula:
[0200]
[0201]
[0202] Define a dynamic addition function (add function) to get the original feature vector after injection:
[0203]
[0204] In this embodiment, the key constraint imposed is to keep the coefficient of the original feature vector to 1 to ensure that the original U-Net data flow architecture can be maintained. In actual operation, the range of the first weight data can be limited to between 0 and 20 to avoid the input target feature vector being too strong and suppressing the original feature vector.
[0205] Optionally, since in the U-Net architecture, a sum operation is required when connecting the original feature vector with the upsampled feature vector, if the coefficient of the original feature vector is 1, then the sum operation is equivalent to simply concatenating the original feature vector with the upsampled feature vector without performing any weighting operation, thereby maintaining the original information flow mode in the U-Net architecture, ensuring that the network can maintain the importance and effectiveness of the original feature vector during the learning process, thereby better retaining the details and features of the image.
[0206] Optionally, in order to balance different initial control signals during the sampling phase, the target eigenvector may be controlled to form an acute angle with each eigenvector, thereby achieving the purpose of balancing various initial control signals before injection.
[0207] For example, the initial feature vectors (e.g., feature maps) associated with different initial control conditions (which may be initial control signals) can be expressed as , , here is just an example, there is no specific restriction on the form of the initial eigenvector, and the combination strategy defined by the following equation can be used:
[0208]
[0209] in, The function is explicitly defined as:
[0210]
[0211] As an optional embodiment, a loss function is also involved and added to optimize the conservatism of the control condition injection model.
[0212] Since Control Net is a control network adjusted with much less data than the original diffusion model, when combining multiple Control Nets, the conservatism of the loss function needs to be considered to ensure the stability and performance of the trained control condition injection model. The above loss function can be an enhanced score function, for example, a conditional score function. It should be noted that there is no specific restriction on the type of loss function here.
[0213] In this embodiment, Figure 8 is a schematic diagram of a model training according to an embodiment of the present application, such as Figure 8 As shown, obtain the image sample 801 to be generated, add noise information to the image sample 801 to be generated, and obtain the noise sample 802. Set the prompt word and generation time. Obtain the control condition sample, such as the edge map 805, add the segmentation mask data to the control condition sample, and obtain the enhanced control condition sample 806. Call the encoder 803 and decoder 804 in the sub-image generation model to process the noise sample 802, the prompt word, the generation time, and the data output by the sub-control condition injection model 807 to obtain the output data 808 of the sub-image generation model. The output data 808 can be the generated image sample.
[0214] Optionally, after determining the output data 808, the Jacobian matrix corresponding to the model is calculated, and then an estimation method (eg, Hutchinson) is used to estimate the parameters in the Jacobian matrix, thereby completing an unbiased estimation of the loss function.
[0215] Optionally, a conditional score function corresponding to the sub-image generation model is pre-constructed ( ):
[0216]
[0217] right An unbiased estimate can be made to obtain the estimated conservative loss:
[0218]
[0219] Among them, x t Can be used to characterize noise samples, S t Can be used to characterize output data. It can be used to characterize the Jacobian matrix associated with noise samples and output data.
[0220] In this embodiment, the Jacobian matrix of the model ( ) is decomposed into two parts: the Jacobian matrix of the original U-Net (which can include the encoder and decoder) ) and the additional Jacobian matrix introduced by ControlNet ( ):
[0221]
[0222] Since there is a huge gap between the amount of training data of the original U-Net and Control Net, this embodiment converts the Jacobian matrix ( ) and the additional Jacobian matrix introduced by Control Net ( ).
[0223] Optionally, based on the responsibility of conservatism, in the process of equipping the U-Net model with Control Net, the conservatism of the original Jacobian matrix is controlled by the original model parameters of U-Net. At the same time, the original model parameters of Control Net are mainly responsible for managing the additional conservatism introduced by ControlNet, which can be described as:
[0224]
[0225] Among them, ϕ can be used to characterize the initial model parameters of the sub-control condition injection model.
[0226] Furthermore, we can ignore Not included , and obtain the second loss function for optimizing the sub-control condition injection model:
[0227]
[0228] The gradient of the second loss function can be set equal to the gradient of the estimated conservative loss:
[0229]
[0230] Due to computational limitations, it is still desirable to completely remove , therefore, the simplified loss for Control Net optimization can be obtained:
[0231]
[0232] Optionally, a simplified loss function is used to process the generated image samples, noise samples, generated image samples and enhanced control condition samples to determine the target model parameters of the sub-control condition injection model, thereby obtaining a trained control condition injection model.
[0233] Optionally, based on the above, it can be determined The norm of is uniformly constrained by M, so that if the simplified loss is zero, the original loss is also 0.
[0234]
[0235] Optionally, during the model training process, the model parameters of the codec remain unchanged, and the original model parameters of the sub-control condition injection model are continuously updated.
[0236] Fig. 9 is a schematic diagram of generating an image to be generated according to an embodiment of the present application, such as Fig. 9 As shown, Fig. 9 The first column is the image to be generated after combining multiple image control signals, the second column is the result of the image to be generated directly using the controlNet combination, and the third column is the image to be generated generated by this embodiment. By comparing with other solutions, it can be seen that this embodiment can better ensure that the control signal is effective, and the overall picture is harmonious and natural.
[0237] In an embodiment of the present application, multiple initial control conditions are obtained, combined to obtain target control conditions, and injection parameters are determined based on the target control conditions. By utilizing the injection parameters, the multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected. The injection parameters and initial prompt information are used to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0238] The method embodiments provided in the embodiments of the present application may also be executed in a mobile terminal, a computer terminal or a similar computing device. Fig.10 is a hardware structure block diagram of a computer terminal (or mobile device) according to an image generation method of an embodiment of the present application, such as Fig.10As shown, the computer terminal 100 (or mobile device) may include one or more (1002a, 1002b, ..., 1002n are used to illustrate in the figure) processors 1002 (the processor 1002 may include but is not limited to a processing device such as a microprocessor (Microcontroller Unit, referred to as MCU) or a programmable logic device (Field Programmable Gate Array, referred to as FPGA)), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Fig.10 The structure shown is only for illustration and does not limit the structure of the above electronic device. Fig.10 More or fewer components as shown, or with Fig.10 Different configurations shown.
[0239] Fig.10 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the above-mentioned computer terminal 100 (or mobile device), but also as an exemplary block diagram of the above-mentioned server.
[0240] The memory 1004 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data processing method in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, realizing the above-mentioned data processing method. The memory 1004 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 1004 may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal 100 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0241] The transmission device 1006 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 100. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 1006 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0242] The display may be, for example, a touch screen liquid crystal display (LCD), which may enable a user to interact with a user interface of the computer terminal 100 (or mobile device).
[0243] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0244] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0245] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0246] According to an embodiment of the present application, there is also provided a method for implementing the above Figure 2An image generating device according to the image generating method shown.
[0247] Fig.11 is a schematic diagram of an image generating device according to an embodiment of the present application, such as Fig.11 As shown, the image generating device 1100 may include: a first acquiring unit 1102 , a first combining unit 1104 , a first determining unit 1106 and a first guiding unit 1108 .
[0248] The first acquisition unit 1102 is used to acquire a plurality of initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent the attribute information of the image to be generated and to guide the image generation model to generate the image to be generated.
[0249] The first combining unit 1104 is used to combine multiple initial control conditions to obtain target control conditions of the image generation model.
[0250] The first determination unit 1106 is used to determine the injection parameters of the image generation model based on the target control condition, wherein the injection parameters are used to characterize the influence of the target control condition on the attribute information.
[0251] The first guiding unit 1108 is used to guide the image generation model to generate the image to be generated by using the injection parameters and the initial prompt information of the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0252] It should be noted that the first acquisition unit 1102, the first combination unit 1104, the first determination unit 1106 and the first guidance unit 1108 correspond to the above steps S202 to S208, and the four units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiments. It should be noted that the above units can be hardware components or software components stored in a memory and processed by one or more processors, and the above units can also be run in the computer terminal provided in the above embodiments as part of the device.
[0253] According to an embodiment of the present application, another method for implementing the above Figure 3 An image generating device according to the image generating method shown.
[0254] Fig.12 is a schematic diagram of another image generating device according to an embodiment of the present application, such as Fig.12 As shown, the image generating device 1200 may include: a second acquiring unit 1202 , a converting unit 1204 , a second combining unit 1206 , a second determining unit 1208 and a second guiding unit 1210 .
[0255] The second acquiring unit 1202 is configured to acquire a control image from a client.
[0256] The conversion unit 1204 is used to convert the control image into an initial control condition corresponding to the image to be generated, wherein the initial control condition is used to represent the attribute information of the image to be generated and is used to guide the image generation model to generate the image to be generated.
[0257] The second combining unit 1206 is used to combine multiple initial control conditions to obtain target control conditions of the image generation model.
[0258] The second determining unit 1208 is used to determine the injection parameters of the image generation model based on the target control condition, wherein the injection parameters are used to characterize the influence of the target control condition on the attribute information.
[0259] The second guiding unit 1210 is used to guide the image generation model to generate the image to be generated by using the injection parameters and the initial prompt information of the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0260] It should be noted that the second acquisition unit 1202, the conversion unit 1204, the second combination unit 1206, the second determination unit 1208 and the second guidance unit 1210 correspond to the above steps S302 to S310, and the five units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiments. It should be noted that the above units can be hardware components or software components stored in a memory and processed by one or more processors, and the above units can also be run in the computer terminal provided in the above embodiments as part of the device.
[0261] According to an embodiment of the present application, another method for implementing the above Figure 4 An image generating device according to the image generating method shown.
[0262] Fig.13 is a schematic diagram of another image generating device according to an embodiment of the present application, such as Fig.13 As shown, the image generating device 1300 may include: a display unit 1302 and a processing unit 1304 .
[0263] The display unit 1302 is used to display the display objects in the virtual store on the e-commerce platform on the presentation screen of the operation interface.
[0264] The processing unit 1304 is used to generate and display the image to be generated including the display object in response to the generation instruction acting on the operation interface.
[0265] It should be noted that the above-mentioned display unit 1302 and processing unit 1304 correspond to the above-mentioned steps S402 to S404, and the examples and application scenarios implemented by the two units and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiments. It should be noted that the above-mentioned units can be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned units can also be run in the computer terminal provided in the above-mentioned embodiments as part of the device.
[0266] In the image generation device of this embodiment, in the embodiment of the present application, multiple initial control conditions are obtained, the multiple initial control conditions are combined to obtain target control conditions, and injection parameters are determined based on the target control conditions. By utilizing the injection parameters, the multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected. The injection parameters and initial prompt information are used to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0267] The embodiment of the present application may provide an electronic device, which may be any electronic device in a group of electronic devices. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.
[0268] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0269] In this embodiment, the electronic device can execute the program code in the image generation method.
[0270] Optionally, Fig.14 is a structural block diagram of an electronic device according to an embodiment of the present application. Fig.14 As shown, the electronic device A may include: one or more (only one is shown in the figure) processors 1402, a memory 1404, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0271] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0272] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain multiple initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent the attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; combine the multiple initial control conditions to obtain the target control conditions of the image generation model; based on the target control conditions, determine the injection parameters of the image generation model, wherein the injection parameters are used to characterize the degree of influence of the target control conditions on the attribute information; use the injection parameters and the initial prompt information of the image to be generated to guide the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0273] Optionally, the processor may also execute program codes of the following steps: obtaining an original feature vector of the initial prompt information; and determining injection parameters using the target feature vector and the original feature vector of the target control condition.
[0274] Optionally, the processor may also execute the following program code: determine first weight data based on the target feature vector and the original feature vector, wherein the first weight data is used to represent the degree of influence of the target feature vector on the original feature vector; and convert the target feature vector into an injection parameter using the first weight data.
[0275] Optionally, the processor may also execute program code of the following steps: injecting injection parameters into initial prompt information to obtain target prompt information; and using the target prompt information to guide the image generation model to generate the image to be generated.
[0276] Optionally, the processor may also execute the program code of the following steps: injecting the injection parameter into the original feature vector of the initial prompt information to obtain the injected original feature vector, wherein the injected original feature vector is used to represent the target prompt information.
[0277] Optionally, the processor may also execute program codes of the following steps: determining a characteristic vector of an initial control condition; combining characteristic vectors of multiple initial control conditions to obtain a target characteristic vector; and determining a target control condition corresponding to the target characteristic vector.
[0278] Optionally, the processor may also execute the program code of the following steps: determining second weight data corresponding to the feature vector, wherein the second weight data is used to indicate the importance of the feature vector to the determined target control condition; and combining the feature vectors of multiple initial control conditions according to the second weight data to obtain the target feature vector.
[0279] Optionally, the processor may also execute program code of the following steps: selecting a target number of feature vectors from multiple feature vectors; calling a combination strategy to determine second weight data corresponding to the selected feature vectors, wherein the combination strategy is used to represent the association relationship between multiple feature vectors.
[0280] Optionally, the processor may also execute the following program code steps: combining the target number of feature vectors according to the second weight data corresponding to the target number of feature vectors to obtain a first target feature vector; deleting the target number of feature vectors from multiple feature vectors, and selecting any feature vector from the multiple feature vectors after deletion; determining the second weight data corresponding to any selected feature vector using the first target feature vector and any selected feature vector; combining the first target feature vector and any selected feature vector according to the second weight data corresponding to any selected feature vector to obtain a target feature vector.
[0281] Optionally, the processor may also execute the program code of the following steps: a combining step, combining the first target feature vector and any selected feature vector according to the second weight data corresponding to any selected feature vector to obtain a second target feature vector, wherein the second target feature vector and the target feature vector are different; a deleting step, deleting any selected feature vector from a plurality of feature vectors, and selecting any feature vector again from at least one feature vector after deletion; a looping step, determining the second weight data corresponding to any feature vector selected again using the second target feature vector and any feature vector selected again, and combining the second target feature vector and any selected feature vector according to the second weight data corresponding to any feature vector selected again to obtain a third target feature vector; repeating the deleting step and the looping step until all the plurality of feature vectors are combined, and determining the third target feature vector determined based on the feature vector selected last time as the target feature vector.
[0282] Optionally, the processor may also execute the following program code steps: determining the control type of the initial control condition, wherein the control type is used to represent the control content of the initial control condition; calling a control condition injection model that matches the control type, wherein the control condition injection model is trained using enhanced control condition samples, and the control condition samples are used to adjust the attribute information samples of the image samples to be generated; inputting the initial control condition into the corresponding control condition injection model to obtain a feature vector.
[0283] Optionally, the processor may also execute the program code of the following steps: obtaining an image sample to be generated, wherein the image sample to be generated includes multiple regions; calling a preprocessing module of a sub-control condition injection model to preprocess the image sample to be generated to obtain a control condition sample, and determining segmentation mask data of the image sample to be generated, wherein the segmentation mask data is used to characterize the edge contours of different regions; using the segmentation mask data, enhancing the control condition sample to obtain an enhanced control condition sample.
[0284] Optionally, the processor may also execute the following program code: adding noise information to the image sample to be generated to obtain a noise sample; calling the sub-image generation model corresponding to the image generation model to convert the noise sample and the enhanced control condition sample to obtain a generated image sample, wherein the sub-image generation model includes at least a sub-control condition injection model; determining the target model parameters of the sub-control condition injection model based on the image sample to be generated, the noise sample, the generated image sample, and the enhanced control condition sample; and using the target model parameters to adjust the initial model parameters of the sub-control condition injection model to obtain a control condition injection model.
[0285] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a control image from the client; convert the control image into an initial control condition corresponding to the image to be generated, wherein the initial control condition is used to represent the attribute information of the image to be generated and is used to guide the image generation model to generate the image to be generated; combine multiple initial control conditions to obtain the target control condition of the image generation model; based on the target control condition, determine the injection parameter of the image generation model, wherein the injection parameter is used to characterize the degree of influence of the target control condition on the attribute information; use the injection parameter and the initial prompt information of the image to be generated to guide the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information.
[0286] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: display the display objects in the virtual store on the e-commerce platform on the presentation screen of the operation interface; generate and display the image to be generated including the display objects in response to the generation instruction acting on the operation interface; wherein the image to be generated is generated by guiding the image generation model using the injection parameters and the initial prompt information of the image to be generated, the initial prompt information is used to prompt the image content of the image to be generated under the attribute information of the image to be generated, the injection parameters are determined based on the target control conditions, and are used to characterize the degree of influence of the target control conditions on the attribute information, the target control conditions are obtained by combining multiple initial control conditions, the initial control conditions are used to represent the attribute information, and are used to guide the image generation model to generate the image to be generated.
[0287] In an embodiment of the present application, multiple initial control conditions are obtained, combined to obtain target control conditions, and injection parameters are determined based on the target control conditions. By utilizing the injection parameters, the multiple initial control conditions are injected into the image generation model, thereby achieving the purpose of balancing the multiple initial control conditions injected. The injection parameters and initial prompt information are used to guide the image generation model to generate the image to be generated, thereby achieving the technical effect of improving the image generation effect and solving the technical problem of poor image generation effect.
[0288] Those skilled in the art will understand that Fig.14 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, and a mobile Internet device (MID for short). The figure does not limit the structure of the above electronic device. For example, the electronic device A may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a configuration different from that shown in the figure.
[0289] A person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0290] The embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the storage medium is located is controlled to execute the program code executed by the method provided in the first embodiment, and the specific execution process is as shown above, which will not be repeated here.
[0291] Optionally, in this embodiment, the computer-readable storage medium may be located in any electronic device in a group of electronic devices in a computer network, or in any mobile terminal in a group of mobile terminals.
[0292] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and the computer program implements the method provided in the embodiment when executed by a processor.
[0293] The embodiments of the present application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program, and when the computer program is executed by a processor, the method provided in the embodiments is implemented.
[0294] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0295] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0296] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0297] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0298] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0299] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0300] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk, etc., and other media that can store program codes.
[0301] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, it does not depart from the principles of the present application.
Claims
1. A method for generating an image, characterized in that: include: Acquire a plurality of initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent attribute information of the image to be generated and are used to guide an image generation model to generate the image to be generated; Combining a plurality of the initial control conditions to obtain a target control condition of the image generation model; Based on the target control condition, determining an injection parameter of the image generation model, wherein the injection parameter is used to characterize the degree of influence of the target control condition on the attribute information; Using the injection parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information; Wherein, determining the injection parameters of the image generation model based on the target control condition includes: determining first weight data based on the target feature vector of the target control condition and the original feature vector of the initial prompt information, wherein the first weight data is used to represent the degree of influence of the target feature vector on the original feature vector; using the first weight data, converting the target feature vector into the injection parameter; Wherein, combining multiple initial control conditions to obtain the target control condition of the image generation model includes: combining multiple initial control conditions according to second weight data to obtain the target control condition, wherein the second weight data is used to indicate the importance of the initial control condition to the determined target control condition.
2. The method according to claim 1, characterized in that The method further comprises: Obtain an original feature vector of the initial prompt information.
3. The method according to claim 2, characterized in that: The step of converting the target feature vector into the injection parameter by using the first weight data includes: The product of the target feature vector and the first weight data is determined as the injection parameter.
4. The method according to claim 1, characterized in that: The step of guiding the image generation model to generate the image to be generated by using the injection parameters and the initial prompt information of the image to be generated includes: Injecting the injection parameter into the initial prompt information to obtain target prompt information; The target prompt information is used to guide the image generation model to generate the image to be generated.
5. The method according to claim 4, characterized in that The step of injecting the injection parameter into the initial prompt information to obtain the target prompt information includes: The injection parameter is injected into the original feature vector of the initial prompt information to obtain the injected original feature vector, wherein the injected original feature vector is used to represent the target prompt information.
6. The method according to claim 1, characterized in that The step of combining the plurality of initial control conditions according to the second weight data to obtain the target control condition includes: determining a characteristic vector of the initial control condition; According to the second weight data, combining the feature vectors of the plurality of initial control conditions to obtain a target feature vector; The target control condition corresponding to the target feature vector is determined.
7. The method according to claim 6, characterized in that The step of combining the feature vectors of the plurality of initial control conditions according to the second weight data to obtain a target feature vector comprises: Determining the second weight data corresponding to the feature vector; According to the second weight data, the feature vectors of a plurality of the initial control conditions are combined to obtain the target feature vector.
8. The method according to claim 7, characterized in that The determining the second weight data corresponding to the feature vector includes: Selecting a target number of the feature vectors from a plurality of the feature vectors; A combination strategy is called to determine the second weight data corresponding to the selected feature vector, wherein the combination strategy is used to represent the association relationship between multiple feature vectors.
9. The method according to claim 8, characterized in that Combining the feature vectors of the plurality of initial control conditions according to the second weight data to obtain the target feature vector comprises: Combining the target number of feature vectors according to the second weight data corresponding to the target number of feature vectors to obtain a first target feature vector; Deleting the target number of feature vectors from the plurality of feature vectors, and selecting any one of the feature vectors from the plurality of feature vectors after deletion; Determine second weight data corresponding to any of the selected feature vectors by using the first target feature vector and any of the selected feature vectors; According to the second weight data corresponding to any one of the selected feature vectors, the first target feature vector and any one of the selected feature vectors are combined to obtain the target feature vector.
10. The method according to claim 9, characterized in that The step of combining the first target feature vector and any selected feature vector according to the second weight data corresponding to any selected feature vector to obtain the target feature vector comprises: a combining step, combining the first target feature vector and any selected feature vector according to second weight data corresponding to any selected feature vector, to obtain a second target feature vector, wherein the second target feature vector is different from the target feature vector; A deleting step, deleting any one of the selected feature vectors from the plurality of feature vectors, and selecting any one of the feature vectors again from at least one of the deleted feature vectors; A loop step, using the second target feature vector and any of the feature vectors selected again, determining second weight data corresponding to any of the feature vectors selected again, and combining the second target feature vector and any of the feature vectors selected again according to the second weight data corresponding to any of the feature vectors selected again, to obtain a third target feature vector; The deleting step and the looping step are repeated until all the multiple feature vectors are combined, and a third target feature vector determined based on the feature vector selected last time is determined as the target feature vector.
11. The method according to claim 6, characterized in that The determining of the characteristic vector of the initial control condition comprises: Determining a control type of the initial control condition, wherein the control type is used to indicate a control content of the initial control condition; Retrieving a control condition injection model that matches the control type, wherein the control condition injection model is obtained by training using enhanced control condition samples, and the control condition samples are used to adjust the attribute information samples of the image samples to be generated; The initial control condition is input into the corresponding control condition injection model to obtain the feature vector.
12. The method according to claim 11, characterized in that The method further comprises: Acquire the image sample to be generated, wherein the image sample to be generated includes multiple regions; Recalling the preprocessing module of the sub-control condition injection model, preprocessing the image sample to be generated, obtaining the control condition sample, and determining the segmentation mask data of the image sample to be generated, wherein the segmentation mask data is used to characterize the edge contours of different regions; The control condition samples are enhanced using the segmentation mask data to obtain enhanced control condition samples.
13. The method according to claim 12, characterized in that The method further comprises: Adding noise information to the image sample to be generated to obtain a noise sample; Calling a sub-image generation model corresponding to the image generation model to transform the noise sample and the enhanced control condition sample to obtain a generated image sample, wherein the sub-image generation model at least includes a sub-control condition injection model; Determining target model parameters of the sub-control condition injection model based on the to-be-generated image sample, the noise sample, the generated image sample, and the enhanced control condition sample; The target model parameters are used to adjust the initial model parameters of the sub-control condition injection model to obtain the control condition injection model.
14. A method for generating an image, characterized in that: include: Get the control image from the client; Converting the control image into initial control conditions corresponding to the image to be generated, wherein the initial control conditions are used to represent attribute information of the image to be generated and are used to guide the image generation model to generate the image to be generated; Combining a plurality of the initial control conditions to obtain a target control condition of the image generation model; Based on the target control condition, determining an injection parameter of the image generation model, wherein the injection parameter is used to characterize the degree of influence of the target control condition on the attribute information; Using the injection parameters and the initial prompt information of the image to be generated, guiding the image generation model to generate the image to be generated, wherein the initial prompt information is used to prompt the image content of the image to be generated under the attribute information; Wherein, determining the injection parameters of the image generation model based on the target control condition includes: determining first weight data based on the target feature vector of the target control condition and the original feature vector of the initial prompt information, wherein the first weight data is used to represent the degree of influence of the target feature vector on the original feature vector; using the first weight data, converting the target feature vector into the injection parameter; Wherein, combining multiple initial control conditions to obtain the target control condition of the image generation model includes: combining multiple initial control conditions according to second weight data to obtain the target control condition, wherein the second weight data is used to indicate the importance of the initial control condition to the determined target control condition.
15. A method for generating an image, characterized in that: include: Displaying display objects in a virtual store on the e-commerce platform on a presentation screen of the operation interface; In response to a generation instruction applied to the operation interface, generating and displaying an image to be generated including the display object; The image to be generated is generated by guiding the image generation model using the injection parameters and the initial prompt information of the image to be generated, the initial prompt information is used to prompt the image content of the image to be generated under the attribute information of the image to be generated, the injection parameters are determined based on the target control conditions, and are used to characterize the degree of influence of the target control conditions on the attribute information, the target control conditions are obtained by combining multiple initial control conditions based on the combination ratio of multiple initial control conditions, the initial control conditions are used to represent the attribute information, and are used to guide the image generation model to generate the image to be generated; Among them, the injection parameter is obtained by converting the target feature vector using the first weight data, and the first weight data is determined based on the target feature vector of the target control condition and the original feature vector of the initial prompt information, and the first weight data is used to indicate the degree of influence of the target feature vector on the original feature vector.
16. A computer program product, characterized in that It comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Figure generation method and device, equipment and storage medium
CN118365730A