Model determination method and device, image generation method and device, equipment and medium
By determining and adjusting the target objects in the image generation model and introducing the characterization characteristics of the reference image, the problem of difficulty in balancing the performance between information maintenance and target dimensions in the prior art is solved, and a better image generation effect is achieved.
Patent Information
- Application Number
- CN202311568533.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
While maintaining certain information in the reference image, existing image generation models are difficult to achieve a performance balance between the information maintenance dimension and the target dimension (such as image editing, aesthetics, etc.), resulting in poor image generation results.
By obtaining the pending model, maintain the dimension-affected representation data and the target dimension-affected representation data according to the information of the candidate object, determine the target object from it, and adjust the pending model to obtain the image generation model. When generating an image, the image generation model introduces characterization features of the information to be retained in the reference image to be maintained in the data processing of the target object.
It realizes the performance of model under the target dimension while maintaining certain information in the reference image, improves the image generation effect, and takes into account the performance of multiple dimensions (such as information maintenance, image editing, aesthetics, etc.).
Smart Images

Figure CN120031990A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method for determining a model, an image generation method, an apparatus, a device, and a medium. Background Art
[0002] In some application scenarios, an image generation model pre-constructed (such as a diffusion model, etc.) may be used to perform an image generation task (such as a task of generating an image based on text, etc.). For example, in some application scenarios, the image generation model may perform image generation processing according to a prompt text provided by a user (such as the text "a puppy wearing sunglasses", etc.), so that the generated image conforms to the semantic content represented by the prompt text.
[0003] However, for some application scenarios, there may be a further requirement that the above-mentioned image generation model has the ability to preserve certain information in a reference image (such as the ability to preserve the face, the ability to preserve the hairstyle, etc.). For example, when the reference image is used to describe a target 1 (such as a person, an animal, a part of a body, etc.), and the image generation model is used to perform image generation processing according to a prompt text provided by a user and the reference image, it is not only required that the generated image conforms to the semantic content represented by the prompt text, but also required that the object described in the generated image is the target 1, so as to perform image processing (such as image style change processing, etc.) on the premise of preserving the description information of the target 1 in the reference image. Summary of the Invention
[0004] This application provides a method for determining a model, an image generation method, an apparatus, a device, and a medium, so as to achieve the purpose of performing image processing while maintaining certain information in a reference image.
[0005] To achieve the above object, the technical solutions provided in this application are as follows:
[0006] This application provides a method for determining a model, the method comprising:
[0007] Obtain a model to be processed, the model to be processed including at least one candidate object, and the at least one candidate object including at least one network layer;
[0008] Determine a target object from the at least one candidate object according to the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined according to a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object;
[0009] According to the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0010] In a possible implementation manner, different candidate objects are different network layers;
[0011] or,
[0012] Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
[0013] In one possible implementation, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
[0014] In a possible implementation manner, if different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object under the image editing dimension;
[0015] If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact characterization data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or the image generation quality dimension.
[0016] In a possible implementation manner, for any candidate object, a process of determining the impact characterization data of the candidate object includes:
[0017] According to the candidate object, the model to be processed is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is different from the state of the candidate object in the model to be processed;
[0018] Acquire a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by the model to be processed performing image generation processing on the prompt text, and the second generated image is determined by the adjusted model performing image generation processing based on the reference image and the prompt text;
[0019] Determining information retention dimension impact characterization data of the candidate object based on gap characterization data between information retention dimension performance characterization data of the first generated image and information retention dimension performance characterization data of the second generated image; the information retention dimension performance characterization data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance characterization data of the second generated image is determined based on the second generated image and the reference image;
[0020] The target dimension impact characterization data of the candidate object is determined based on the gap characterization data between the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image; the target dimension performance characterization data of the first generated image is determined based on the first generated image; the target dimension performance characterization data of the second generated image is determined based on the second generated image.
[0021] In a possible implementation manner, when generating the first generated image using the model to be processed, it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0022] When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0023] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention;
[0024] The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes:
[0025] The candidate object in the to-be-processed model is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
[0026] In a possible implementation manner, when generating the first generated image using the model to be processed, it is not necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0027] When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0028] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0029] The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes:
[0030] The candidate object in the model to be processed is adjusted to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
[0031] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by the model to be processed performing image generation processing according to the reference image and the prompt text;
[0032] or,
[0033] The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
[0034] In a possible implementation manner, if different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined according to the similarity between the first generated image and the prompt text;
[0035] The target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
[0036] In a possible implementation manner, if different candidate objects are network layer sets corresponding to different time steps, the target dimension impact representation data of the candidate object includes image editing dimension impact representation data and / or image generation quality dimension impact representation data;
[0037] The image editing dimension impact characterization data is determined based on the gap characterization data between the image editing dimension performance characterization data of the first generated image and the image editing dimension performance characterization data of the second generated image; the image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text;
[0038] The image generation quality dimension impact characterization data is determined based on the gap characterization data between the image generation quality dimension performance characterization data of the first generated image and the image generation quality dimension performance characterization data of the second generated image; the image generation quality dimension performance characterization data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance characterization data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
[0039] In a possible implementation manner, when the image generation model performs image generation processing based on the reference image and the prompt text, it is not necessary to introduce the characterization features of the information to be retained in the reference image into the data processing processes corresponding to other objects except the target object.
[0040] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-constructed basic model, the basic model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention;
[0041] Adjusting the model to be processed according to the target object to obtain an image generation model includes:
[0042] Performing adjustment processing on other objects except the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the state of other objects except the target object in the image generation model is consistent with the state of the other objects in the basic model.
[0043] In a possible implementation manner, the model to be processed is a pre-constructed basic model, and the basic model is used to perform image generation processing based on the prompt text;
[0044] Adjusting the model to be processed according to the target object to obtain an image generation model includes:
[0045] Performing adjustment processing on the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the state of other objects except the target object in the image generation model is consistent with the state of the other objects in the model to be processed.
[0046] In a possible implementation manner, if different candidate objects are different network layers, the method further includes:
[0047] The output weights corresponding to the target object in the image generation model are adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0048] The present application provides an image generation method, the method comprising:
[0049] Obtain the information description image and image constraint description text to be retained;
[0050] The image describing the information to be retained and the image constraint description text are input into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method provided in the present application.
[0051] The present application provides a model determination device, comprising:
[0052] A first acquisition unit, configured to acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer;
[0053] A first determining unit is used to determine a target object from the at least one candidate object according to the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined according to a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object;
[0054] A first adjustment unit is used to adjust the model to be processed according to the target object to obtain an image generation model, wherein the image generation model includes the target object, and when the image generation model performs image generation processing based on the reference image and the prompt text, characterization features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0055] The present application provides an image generation device, comprising:
[0056] A second acquisition unit is used to acquire the information description image to be retained and the image constraint description text;
[0057] The second determination unit is used to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method provided in the present application.
[0058] The present application provides an electronic device, the device comprising: a processor and a memory;
[0059] The memory is used to store instructions or computer programs;
[0060] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the model determination method or image generation method provided in the present application.
[0061] The present application provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes the model determination method or image generation method provided in the present application.
[0062] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the model determination method or image generation method provided by the present application.
[0063] Compared with the related art, this application has at least the following advantages:
[0064] In the technical solution provided by the present application, a model to be processed (such as a diffusion model, etc.) is first obtained, and the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; then, based on the information retention dimension impact characterization data and the target dimension impact characterization data of each candidate object, a target object is determined from the at least one candidate object, so that the information retention dimension impact characterization data of the target object is higher than the target dimension impact characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, the model to be processed is adjusted to obtain an image generation model, so that the image generation model includes The target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, so that the image generation model can achieve performance balance in the two dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.) as much as possible. In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, and aesthetics), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, which is beneficial to improving the image generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0066] Figure 1 A structural schematic diagram of a diffusion model provided in an embodiment of the present application;
[0067] Figure 2 A schematic diagram of the structure of another diffusion model provided in an embodiment of the present application;
[0068] Figure 3 A flow chart of a model determination method provided in an embodiment of the present application;
[0069] Figure 4 A flowchart of an image generation method provided in an embodiment of the present application;
[0070] Figure 5 A schematic diagram of the structure of a model determination device provided in an embodiment of the present application;
[0071] Figure 6A schematic diagram of the structure of an image generating device provided in an embodiment of the present application;
[0072] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] Through research, it is found that for some diffusion models, the diffusion model has strong generation ability, but the diffusion model does not have the ability to maintain certain information in the reference image. For example, given a picture of a cat, these diffusion models cannot guarantee that the cat in the generated image is still the cat in the reference image. Based on this, in order to make these diffusion models have the ability to maintain certain information in the reference image (for example, Figure 1 The ability of the diffusion model shown in the figure, etc.), can be achieved by the following method: first train an information extraction module that can extract and process this information (for example, Figure 2 The information extraction module shown in the figure is used to guide the diffusion models to generate images that can retain the information using the data extracted by the information extraction module. For ease of understanding, the following is an explanation with examples.
[0074] As an example, in some application scenarios (e.g., facial information preservation scenarios, etc.), for a diffusion model capable of preserving facial information in a reference image, the diffusion model may at least include an information extraction module (e.g., Figure 2 The information extraction module shown) and the denoising network (for example, Figure 2 The denoising network shown in the figure). The information extraction module is used to obtain the representation features of facial information from the reference image, so that the representation features can represent the facial information carried by the reference image. The denoising network is used to realize image generation processing by denoising; and the present application does not limit the denoising network, for example, it can be implemented using a Unet network. In addition, each network layer in the denoising network can perform corresponding data processing according to the representation features, so as to achieve the purpose of preserving facial information.
[0075] The study also found that for the diffusion model with information retention ability described in the above two paragraphs, the generated image of the diffusion model is limited by the reference image. Specifically, if the information retention ability of the diffusion model is stronger, the similarity between the image generated by the diffusion model and the reference image will be greater, resulting in the diffusion model's performance in other aspects (for example, image editing based on prompt text, aesthetics, color matching, etc.) being worse, which in turn leads to poor image generation effect.
[0076] Based on the above research, in order to better improve the image generation effect, the present application provides a model determination method, which includes: first obtaining a model to be processed (such as a diffusion model, etc.), the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; then, based on the information retention dimension influence characterization data of each candidate object and the target dimension influence characterization data, determine the target object from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, adjust the model to be processed to obtain the image Generate a model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, so that the image generation model can achieve performance balance in the two dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.) as much as possible. In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, which is beneficial to improving the image generation effect.
[0077] In addition, the present application does not limit the execution subject of the model determination method provided in the embodiment of the present application. For example, the model determination method provided in the embodiment of the present application can be applied to a terminal device or a server. For another example, the model determination method provided in the embodiment of the present application can also be implemented with the help of the data interaction process between the terminal device and the server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.
[0078] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0079] In order to better understand the technical solution provided by the present application, the model determination method provided by the present application is described below with reference to some drawings. Figure 3As shown, the model determination method provided in the embodiment of the present application includes the following S301-S303. Figure 3 A flowchart of a model determination method provided in an embodiment of the present application.
[0080] S301: Acquire a model to be processed, where the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer.
[0081] Among them, the model to be processed refers to the model needed to be used when constructing an image generation model that can take into account multi-dimensional performance (for example, information retention performance, image editing performance, image generation quality performance, etc.).
[0082] In addition, for the above-mentioned model to be processed, the model to be processed may include at least one candidate object. The candidate object refers to an object that exists in the model to be processed and can participate in the image generation process; and the present application does not limit the implementation method of the candidate object. For example, when the model to be processed is implemented using a diffusion model, the candidate object may be a network layer, or a part of the network layer corresponding to the time step, or all the network layers corresponding to the time step. It can be seen that in one possible implementation, the at least one candidate object may include at least one network layer, and different candidate objects are different network layers, so that an image generation model that can take into account multi-dimensional performance can be constructed by adjusting the relevant parameters in these network layers. In another possible implementation, the at least one candidate object may include at least one set of network layers corresponding to a time step, and different candidate objects are sets of network layers corresponding to different time steps, so that an image generation model that can take into account multi-dimensional performance can be constructed by adjusting the data processing process corresponding to these time steps (for example, the denoising process corresponding to these time steps). For any time step, the set of network layers corresponding to the time step includes at least one network layer.
[0083] In addition, the present application does not limit the implementation method of the network layer set corresponding to the above time step. For example, in some application scenarios, the network layer set corresponding to the time step can be used to record all network layers that exist in the model to be processed and have a corresponding relationship with the time step. For another example, in some application scenarios, in order to maximize the multi-dimensional performance balancing effect of the image generation model, the present application also provides a possible implementation method of the network layer set corresponding to the time step. Under this implementation method, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step, so that the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer. This is conducive to screening out which time step needs to introduce the information to be retained and which network layers corresponding to this time step need to enter the information to be retained, thereby helping to better improve the multi-dimensional performance balancing effect of the finally generated image generation model. Among them, the at least one network layer corresponding to the time step refers to the network layer that exists in the model to be processed and has a corresponding relationship with the time step; and the present application does not limit the implementation method of the at least one network layer corresponding to the time step. For example, the at least one network layer corresponding to the time step may include all network layers that exist in the model to be processed and have a corresponding relationship with the time step. It should be noted that the relevant content of the information retention dimension affecting the characterization data and the target dimension affecting the characterization data is shown below; and the implementation method of the determination process of "the network layer is determined from the at least one network layer corresponding to the time step based on the influence characterization data of the at least one network layer corresponding to the time step" is similar to the determination process of the target object when the target object is a network layer below. For the sake of brevity, it will not be repeated here.
[0084] In addition, the present application does not limit the implementation method of the above-mentioned model to be processed. For ease of understanding, two situations are described below.
[0085] Case 1: In some application scenarios, the present application can construct an image generation model that can take into account multi-dimensional performance by adding some information retention related parameters to an already constructed diffusion model that does not have information retention capability (for example, a diffusion model that can perform image generation processing based on prompt text, etc.).
[0086] Based on the above situation 1, it can be known that in one possible implementation, the above model to be processed can be a pre-constructed basic model, and the basic model is used to perform image generation processing based on the prompt text, that is, the basic model has the ability to generate text images. Among them, the basic model refers to a pre-constructed diffusion model that can perform image generation processing based on the prompt text; and the basic model does not have the ability to retain information. In addition, the basic model can refer to a text image generation model that is pre-trained based on a large amount of sample text, and the present application does not limit the construction method of the basic model. For example, in some application scenarios, the basic model can be implemented using any existing or future method that can be used to construct a text image generation model.
[0087] Based on the above content, it can be known that in some application scenarios, when the above model to be processed is a pre-built basic model, the model to be processed may at least include a denoising network. Among them, the denoising network is used to implement denoising processing; and the present application does not limit the denoising network. For example, the denoising network uses an upsampling module, a downsampling module and a time step sequence to implement denoising processing. Among them, the upsampling module and the downsampling model are components of the denoising network. The upsampling module is used to perform data compression processing. The downsampling module is used to perform data decompression processing (that is, data recovery processing). The time step sequence is a parameter of the denoising network, and the time step sequence is used to describe which time steps need to be executed in the denoising network. The corresponding data processing flow (that is, which time steps need to be executed. The denoising process).
[0088] Case 2: In some application scenarios, it is possible to construct a diffusion model with information retention capability (e.g. Figure 1 and Figure 2 By deleting some information and keeping relevant parameters from the diffusion model shown in the figure, an image generation model that can take into account multi-dimensional performance is constructed.
[0089] Based on the above situation 2, it can be known that in a possible implementation, the above model to be processed may refer to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used to learn information retention so that the model to be processed has information retention capabilities. That is, the model to be processed can perform image generation processing based on a reference image and a prompt text so that part of the information in the final generated image (for example, facial information, etc.) is consistent with part of the information in the reference image, so that the model to be processed has information retention capabilities. Among them, the reference image is used to provide the information that needs to be maintained during the image generation process (for example, facial information, etc.). The prompt text is used to describe in text some constraint information that is required to be based on the image generation process (for example, constraints such as the style of the generated image).
[0090] Based on the above content, we can know that in some application scenarios, when the above-mentioned model to be processed refers to a model obtained by fine-tuning the pre-built basic model (for example, Figure 1 or Figure 2 The diffusion model shown in the figure), the model to be processed may include an information extraction module and a denoising network; and the working principle of the model to be processed is as follows: the information extraction module first obtains the characterization features of the information to be retained from the reference image, so that the characterization features can characterize the information to be retained; and then the denoising network performs denoising processing based on the characterization features of the information to be retained to obtain the final generated image, so that the final generated image and the reference image maintain a high consistency in the information to be retained. Among them, the information extraction module is used to obtain the characterization features of the information to be retained from the reference image. The information to be retained refers to the information that exists in the reference image and needs to be retained during the image generation process; and the present application does not limit the information to be retained. For example, the information to be retained may refer to part of the information existing in the reference image (for example, information used to describe the facial area, information used to describe animals, information used to describe objects, etc.).
[0091] In addition, the present application does not limit the implementation method of the information extraction module in the above paragraph. For ease of understanding, two examples are used for illustration below.
[0092] Example 1, in some application scenarios, in order to reduce the workload of fine-tuning the basic model as much as possible when building a model with information retention capability based on the basic model, the existing text feature extraction module (e.g., Text Encoder) in the basic model can be used to complete the step of obtaining the representation features of the information to be retained from the reference image. Based on this, it can be seen that in a possible implementation, the above information extraction module may include a text-image mapping module and the text feature extraction module. Among them, the text-image mapping module is used to map the reference image into a text representation vector, so that the text representation vector can represent what the information to be retained in the reference image is, so that the text feature extraction module can subsequently encode the text representation vector to obtain the text expression features of the text representation vector, and regard the text expression features as the representation features of the information to be retained, so that the representation features of the information to be retained in the reference image can be obtained with the help of the existing text feature extraction module in the basic model.
[0093] It can be seen that in a possible implementation mode, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the working principle of the model to be processed can specifically include: first mapping the reference image to a text representation vector, so that the text representation vector can represent what the information to be retained in the reference image is; then encoding the text representation vector to obtain the representation features of the information to be retained in the reference image; then, denoising is performed based on the representation features to obtain a generated image, so that the generated image is consistent with the reference image in the information to be retained.
[0094] Example 2: In some application scenarios, in order to better improve the extraction effect of the characterization features of the information to be retained in the reference image, a new module with image feature extraction function can be added to the basic model, so that the module can be used to complete the step of obtaining the characterization features of the information to be retained from the reference image. Based on this, it can be seen that the above information extraction module can be implemented using an image feature extraction module, so that the information extraction module is used to perform image feature extraction processing on the reference image, and the extracted image features are used as the characterization features of the information to be retained in the reference image.
[0095] It can be seen that in a possible implementation, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the working principle of the model to be processed may specifically include: first performing image feature extraction processing on the reference image to obtain the characterization features of the information to be retained in the reference image; then, performing denoising processing based on the characterization features to obtain a generated image, so that the generated image is consistent with the reference image in the information to be retained.
[0096] Based on the above content, it can be known that for the model to be processed, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the model to be processed may include an information extraction module and a denoising network. Among them, the denoising network is used to perform denoising processing based on the output data of the information extraction module; and the present application does not limit the implementation method of the denoising network. For example, in some application scenarios, each attention layer in the denoising network (for example, cross-attention layer, self-attention layer, etc.) introduces the representation features of the information to be retained in the reference image, so that the attention layer can achieve the purpose of information retention by interacting the representation features with the latent features of the generated image; and the denoising process corresponding to each time step involved in the denoising network introduces the representation features of the information to be retained in the reference image, so that each round of denoising process in the denoising network needs to perform denoising processing based on the representation features of the information to be retained in the reference image, so that the denoising network exhibits a stronger information retention performance, so that the model to be processed including the denoising network has a stronger information retention performance, which causes the model to be processed to exhibit poor performance in some other dimensions (for example, image editing dimension, image generation quality dimension, etc.).
[0097] Based on the relevant content of S301 above, it can be known that in some application scenarios, if you want to build an image generation model that can take into account multi-dimensional performance (for example, information retention performance, image editing performance, image generation quality performance, etc.), you can first obtain a model to be processed with certain image generation capabilities, so that you can subsequently build this image generation model that can take into account multi-dimensional performance by making some adjustments to the data processing flow corresponding to some network layers and / or some time steps in the model to be processed.
[0098] S302: Determine a target object from at least one candidate object based on the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined based on a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object.
[0099] The influence characterization data of the i-th candidate object is used to characterize the influence presented by the i-th candidate object in at least one dimension. i is a positive integer, i≤I, I is a positive integer, and I represents the number of objects in the above at least one candidate object.
[0100] In addition, the present application does not limit at least one dimension in the above paragraph. For example, the at least one dimension may include an information retention dimension and a target dimension, and there is a preset opposition relationship between the target dimension and the information retention dimension. Among them, the information retention dimension is used to represent the performance of maintaining the information to be maintained in the reference image that is of concern in the image generation task. The target dimension is used to represent one or more aspects of performance (for example, image editing, aesthetics, etc.) other than the performance of maintaining the information to be maintained in the reference image that is of concern in the image generation task; and the same object will present an opposing influence relationship between the target dimension and the information retention dimension.
[0101] In addition, the present application does not limit the target dimension in the above paragraph. For example, the target dimension may include an image editing dimension and an image generation quality dimension. Among them, the image editing dimension is used to represent the performance of image editing based on prompt text (for example, editing a real image into a two-dimensional cartoon image, adding freckles, beards or wrinkles to a facial image, etc.) that is concerned in the image generation task. The image generation quality dimension is used to represent the performance of one or more quality states of the generated image (for example, aesthetics, accurate color matching, etc.) that is concerned in the image generation task.
[0102] Based on the above two paragraphs, it can be known that for the i-th candidate object above, the impact characterization data of the i-th candidate object can include the information retention dimension impact characterization data of the i-th candidate object and the target dimension impact characterization data of the i-th candidate object. Among them, the information retention dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the information retention dimension. The target dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the target dimension; and the present application does not limit the implementation method of the target dimension impact characterization data of the i-th candidate object. For example, if the i-th candidate object is a network layer, the target dimension impact characterization data of the i-th candidate object can be used to represent the impact presented by the i-th candidate object under the image editing dimension; if the i-th candidate object is a set of network layers corresponding to a time step, the target dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image editing dimension and / or the image generation quality dimension. It can be seen that in a possible implementation, when the i-th candidate object is a set of network layers corresponding to a time step, the target dimension impact characterization data of the i-th candidate object may include the image editing dimension impact characterization data of the i-th candidate object and / or the image generation quality dimension impact characterization data of the i-th candidate object. Among them, the image editing dimension impact characterization data of the i-th candidate object is used to represent the impact of the i-th candidate object under the image editing dimension. The image generation quality dimension impact characterization data of the i-th candidate object is used to represent the impact of the i-th candidate object under the image generation quality dimension (for example, dimensions such as aesthetics, color matching, and light distribution).
[0103] In addition, for the i-th candidate object mentioned above, the influence representation data of the i-th candidate object may be determined based on the reference image and the prompt text; and the determination process may specifically include the following steps 11 to 14.
[0104] Step 11: According to the ith candidate object, the model to be processed is adjusted to obtain an adjusted model, so that the state of the ith candidate object in the adjusted model is different from the state of the ith candidate object in the model to be processed.
[0105] Among them, the adjusted model refers to the model obtained by adjusting the i-th candidate object in the model to be processed, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed, so that the model to be processed and the adjusted model can be used as a control group in the subsequent process to determine the influence characterization data of the i-th candidate object.
[0106] In addition, the present application does not limit the implementation method of the above step 11. For ease of understanding, two situations are combined for explanation below.
[0107] Case 1: When the above-mentioned model to be processed is a pre-built basic model, the above-mentioned step 11 may specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the i-th candidate object.
[0108] It can be seen that for the above-mentioned model to be processed, if the model to be processed refers to a pre-built basic model, then the model to be processed may refer to a model that does not have the ability to retain information. Therefore, when the model to be processed is used to perform an impact assessment process on the i-th candidate object, the i-th candidate object in the model to be processed can be fine-tuned according to the fine-tuning learning goal of information retention to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the i-th candidate object, so that the i-th candidate object in the adjusted model has the ability to retain information, and then the adjusted model also has the ability to retain information. It should be noted that the present application does not limit the implementation method of the fine-tuning process in this paragraph. For example, it can be implemented by any fine-tuning method (for example, the lora fine-tuning method or the method of directly fine-tuning the network layer parameters, etc.).
[0109] Based on the content of the above paragraph, it can be known that in a possible implementation mode, when the above-mentioned model to be processed is a pre-built basic model, the above-mentioned step 11 can specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of other objects in the model to be processed, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the i-th candidate object, so that the adjusted model and the model to be processed can be used as a control group to determine the impact representation data of the i-th candidate object.
[0110] It should be noted that the present application does not limit the implementation method of the step of "adjusting the ith candidate object in the model to be processed" in the above paragraph. For example, if the ith candidate object is a network layer, the step may specifically include: adding the input weight of the characterization features of the information to be retained in the above reference image to the ith candidate object, and adjusting other existing parameters in the ith candidate object. For another example, if the ith candidate object is a set of network layers corresponding to a time step, the step may specifically include: adding the input weight of the characterization features of the information to be retained in the above reference image to all network layers in the set of network layers corresponding to the time step, and adjusting other existing parameters in these network layers, so as to achieve the purpose of introducing the characterization features of the information to be retained in the reference image into the set of network layers corresponding to the time step.
[0111] Case 2: When the model to be processed above refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used to retain learning information, the above step 11 can specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is consistent with the state of the i-th candidate object in the basic model.
[0112] It can be seen that for the above-mentioned model to be processed, if the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used to learn information retention, then the model to be processed may refer to a model with information retention capability. Therefore, when using the model to be processed to perform an impact assessment process on the i-th candidate object, the i-th candidate object in the model to be processed may be reverted to the state of the i-th candidate object in the basic model to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is consistent with the state of the i-th candidate object in the basic model, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of the other objects in the model to be processed, so that the adjusted model and the model to be processed can be used as a control group to determine the impact characterization data of the i-th candidate object.
[0113] It should be noted that the present application does not limit the implementation method of rollback. For example, the rollback may specifically refer to completely rolling back the parameter information of the ith candidate object to the state of the ith candidate object in the basic model, so that the parameter information of the ith candidate object after rollback is consistent with the parameter information of the ith candidate object in the basic model. Among them, the parameter information of the ith candidate object refers to the parameters involved in the data processing process corresponding to the ith candidate object; and the present application does not limit the parameter information of the ith candidate object. For example, when the ith candidate object is a network layer, the parameter information of the ith candidate object may refer to all parameters involved in the network layer. When the ith candidate object is a network layer set corresponding to a time step, the parameter information of the ith candidate object may refer to all parameters involved in all network layers in the network layer set corresponding to the time step. It can be seen that if the i-th candidate object above is a network layer, the fallback may specifically include deleting the introduction weight of the i-th candidate object for the representation features of the information to be retained in the reference image, and adjusting all parameters involved in the i-th candidate object except the introduction weight to the parameters of the i-th candidate object in the basic model, etc.; if the i-th candidate object is a set of network layers corresponding to a time step, the fallback may specifically include: for any network layer in the i-th candidate object, deleting the introduction weight of the network layer for the representation features of the information to be retained in the reference image, and adjusting all parameters involved in the network layer except the introduction weight to the parameters of the network layer in the basic model, etc.
[0114] Based on the relevant content of step 11 above, it can be known that for a model to be processed that includes at least one candidate object, the i-th candidate object in the model to be processed can be adjusted to obtain an adjusted model corresponding to the i-th candidate object, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of the other objects in the model to be processed, so that the adjusted model and the model to be processed can be used as a control group in the future to determine the impact characterization data of the i-th candidate object. Wherein, i is a positive integer, i≤I, I is a positive integer, and I represents the number of objects in the at least one candidate object above.
[0115] Step 12: Obtain a first generated image and a second generated image corresponding to the i-th candidate object, wherein the first generated image is determined by the image generation processing performed by the model to be processed on the prompt text, and the second generated image is determined by the image generation processing performed by the adjusted model based on the reference image and the prompt text.
[0116] The first generated image corresponding to the i-th candidate object refers to the image generated by the above-mentioned model to be processed for the prompt text.
[0117] In addition, the present application does not limit the determination process of the first generated image corresponding to the above-mentioned i-th candidate object. For example, when the model to be processed is a pre-built basic model, the first generated image can be determined by the model to be processed performing image generation processing according to the prompt text. It can be seen that in one possible implementation, when the model to be processed is a pre-built basic model, the determination process of the first generated image corresponding to the i-th candidate object is: input the prompt text into the model to be processed, so that the model to be processed performs image generation processing according to the prompt text, and obtains and outputs the first generated image.
[0118] For another example, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used for learning information retention, the first generated image corresponding to the above-mentioned i-th candidate object may be determined by the model to be processed performing image generation processing based on the reference image and the prompt text. It can be seen that in one possible implementation, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used for learning information retention, the process of determining the first generated image corresponding to the i-th candidate object is: inputting the prompt text and the reference image into the model to be processed, so that the model to be processed performs image generation processing based on the prompt text and the reference image, and obtains and outputs the first generated image.
[0119] The second generated image corresponding to the i-th candidate object refers to the image generated by the adjusted model corresponding to the i-th candidate object based on the reference image and the prompt text. The present application does not limit the determination process of the second generated image. For example, the determination process of the second generated image is: inputting the reference image and the prompt text into the adjusted model corresponding to the i-th candidate object, so that the adjusted model corresponding to the i-th candidate object performs image generation processing based on the reference image and the prompt text, and obtains and outputs the second generated image.
[0120] Based on the relevant contents of the first generated image and the second generated image corresponding to the i-th candidate object in the above text, it can be known that there is a certain difference between the generation process of the first generated image and the generation process of the second generated image, and the difference can be specifically: if the above-mentioned model to be processed is a pre-constructed basic model, then when the model to be processed is used to generate the first generated image, it is not necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object; however, when the adjusted model corresponding to the i-th candidate object is used to generate the second generated image, it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object. However, if the model to be processed refers to a model obtained by fine-tuning the pre-constructed basic model, and the fine-tuning is used to learn information retention, then when the model to be processed is used to generate the first generated image, it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the candidate object; however, when the adjusted model corresponding to the i-th candidate object is used to generate the second generated image, it is not necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object.
[0121] Based on the content of the previous paragraph, it can be known that for the first generated image and the second generated image corresponding to the i-th candidate object above, the data processing flow corresponding to the i-th candidate object used in generating the first generated image is different from the data processing flow corresponding to the i-th candidate object used in generating the second generated image, and the difference is: one data processing flow needs to introduce the representation features of the information to be maintained in the reference image, but the other data processing flow does not need to introduce the representation features of the information to be maintained in the reference image.
[0122] Based on the relevant content of step 12 above, it can be known that for the i-th candidate object above, the first generated image corresponding to the i-th candidate object can be determined using the above-mentioned model to be processed, and the second generated image corresponding to the i-th candidate object can be determined using the adjusted model corresponding to the i-th candidate object, so that the impact characterization data of the i-th candidate object can be determined subsequently based on the difference between the first generated image and the second generated image.
[0123] Step 13: For the first generated image and the second generated image corresponding to the i-th candidate object, determine the information retention dimension impact characterization data of the i-th candidate object based on the gap characterization data between the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image; the information retention dimension performance characterization data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance characterization data of the second generated image is determined based on the second generated image and the reference image.
[0124] In the present application, for the above i-th candidate object, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image can be calculated. Among them, the information retention dimension performance characterization data of the first generated image is used to describe the state of the first generated image in the information retention dimension; and the information retention dimension performance characterization data of the first generated image can be determined based on the first generated image and the reference image, so that the information retention dimension performance characterization data of the first generated image can represent the state of the first generated image in terms of retaining the information to be used in the reference image. The information retention dimension performance characterization data of the second generated image is used to describe the state of the second generated image in the information retention dimension; and the information retention dimension performance characterization data of the second generated image can be determined based on the second generated image and the reference image, so that the information retention dimension performance characterization data of the second generated image can represent the state of the second generated image in terms of retaining the information to be used in the reference image.
[0125] It should be noted that the present application does not limit the determination method of the information retention dimension performance characterization data of the first generated image above, for example, it can be determined by feature similarity. It can be seen that in a possible implementation, the determination process of the information retention dimension performance characterization data of the first generated image can be specifically: feature extraction processing is performed on the first generated image to obtain the image features of the first generated image, and feature extraction processing is performed on the reference image to obtain the image features of the reference image; then, the similarity between the image features of the first generated image and the image features of the reference image is determined as the information retention dimension performance characterization data of the first generated image. For example, in some application scenarios, when the information to be used above is facial description information, the information retention dimension performance characterization data of the first generated image can be judged with the help of a facial recognition model, so that the information retention dimension performance characterization data of the first generated image can be determined by judging whether the judgment result output by the facial recognition model for the first generated image is the same as the judgment result output by the facial recognition model for the reference image.
[0126] It should also be noted that the method for determining the information retention dimension performance characterization data of the second generated image above is similar to the "method for determining the information retention dimension performance characterization data of the first generated image" shown in the previous paragraph. For the sake of brevity, it will not be repeated here.
[0127] In addition, for the first generated image and the second generated image corresponding to the i-th candidate object above, after obtaining the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image, the gap characterization data between the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image can be calculated first, so that the gap characterization data can represent the model performance difference presented by the above-mentioned model to be processed and the adjusted model corresponding to the i-th candidate object in the information retention dimension, so that the gap characterization data can represent the influence of the i-th candidate object on the information retention dimension; then, based on the gap characterization data, the information retention dimension influence characterization data of the i-th candidate object is determined (for example, the gap characterization data is directly determined as the information retention dimension influence characterization data of the i-th candidate object, etc.), so that the information retention dimension influence characterization data can represent the influence of the i-th candidate object on the information retention dimension.
[0128] Based on the relevant content of step 13 above, it can be known that for the i-th candidate object above, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the information retention dimension impact representation data of the i-th candidate object can be calculated based on the first generated image, the second generated image and the reference image, so that the information retention dimension impact representation data can represent the impact of the i-th candidate object on the information retention dimension.
[0129] Step 14: For the first generated image and the second generated image corresponding to the i-th candidate object, determine the target dimension impact characterization data of the i-th candidate object based on the gap characterization data between the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image; the target dimension performance characterization data of the first generated image is determined based on the first generated image; and the target dimension performance characterization data of the second generated image is determined based on the second generated image.
[0130] In the present application, for the above i-th candidate object, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image can be calculated. Among them, the target dimension performance characterization data of the first generated image is used to describe the state of the first generated image in the target dimension; and the target dimension performance characterization data of the first generated image can be determined based on the first generated image, so that the target dimension performance characterization data of the first generated image can represent the state of the first generated image in other aspects (such as image editing, image quality, etc.) in addition to maintaining the information to be used in the reference image. The target dimension performance characterization data of the second generated image is used to describe the state of the second generated image in the target dimension; and the target dimension performance characterization data of the second generated image can be determined based on the second generated image, so that the target dimension performance characterization data of the second generated image can represent the state of the second generated image in other aspects (such as image editing, image quality, etc.) in addition to maintaining the information to be used in the reference image.
[0131] In addition, the present application does not limit the determination process of the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image. For example, when the i-th candidate object above is a network layer, the target dimension performance characterization data of the first generated image corresponding to the i-th candidate object can be determined based on the similarity between the first generated image and the prompt text (for example, the similarity between the first generated image and the prompt text can be directly determined as the target dimension performance characterization data of the first generated image), so that the target dimension performance characterization data can represent the performance of the first generated image in image editing; and the target dimension performance characterization data of the second generated image corresponding to the i-th candidate object can be determined based on the similarity between the second generated image and the prompt text (for example, the similarity between the second generated image and the prompt text can be directly determined as the target dimension performance characterization data of the second generated image), so that the target dimension performance characterization data can represent the performance of the second generated image in image editing.
[0132] It should be noted that the present application does not limit the calculation method of the "similarity between the first generated image and the prompt text" in the previous paragraph. For example, it can be implemented by using any existing or future method that can calculate the similarity between an image and a text, such as a method for calculating the similarity between an image and a text with the help of a Contrastive Language-Image Pre-Training (CLIP) model.
[0133] For another example, when the i-th candidate object above is a set of network layers corresponding to a time step, the target dimension impact representation data of the i-th candidate object may include the image editing dimension impact representation data of the i-th candidate object and / or the image generation quality dimension impact representation data of the i-th candidate object, so that the target dimension impact representation data of the i-th candidate object can represent the impact of the i-th candidate object on the image editing dimension and / or the image generation quality dimension.
[0134] The image editing dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image editing dimension; and the image editing dimension impact characterization data of the i-th candidate object is determined based on the difference characterization data between the image editing dimension performance characterization data of the first generated image corresponding to the i-th candidate object and the image editing dimension performance characterization data of the second generated image corresponding to the i-th candidate object. The image editing dimension performance characterization data of the first generated image is used to represent the performance presented by the first generated image in image editing; and the image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text (for example, the similarity between the first generated image and the prompt text can be directly determined as the image editing dimension performance characterization data of the first generated image). The image editing dimension performance characterization data of the second generated image is used to represent the performance presented by the second generated image in image editing; and the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text (for example, the similarity between the second generated image and the prompt text can be directly determined as the image editing dimension performance characterization data of the second generated image).
[0135] The image generation quality dimension impact characterization data of the i-th candidate object is used to represent the impact of the i-th candidate object in the image generation quality dimension (for example, quality dimensions such as aesthetics and color matching); and the image generation quality dimension impact characterization data of the i-th candidate object is determined based on the difference characterization data between the image generation quality dimension performance characterization data of the first generated image corresponding to the i-th candidate object and the image generation quality dimension performance characterization data of the second generated image corresponding to the i-th candidate object. Among them, the image generation quality dimension performance characterization data of the first generated image is used to represent the performance of the first generated image in terms of image generation quality; and the image generation quality dimension performance characterization data of the first generated image is determined by performing a quality assessment process on the first generated image under at least one quality dimension (for example, aesthetics, color matching, etc.). The image generation quality dimension performance characterization data of the second generated image is used to represent the performance of the second generated image in terms of image generation quality; and the image generation quality dimension performance characterization data of the second generated image is determined by performing a quality assessment process on the second generated image under at least one quality assessment dimension. Among them, the at least one quality assessment dimension refers to the dimension required to be based on when performing quality assessment processing on an image; and the present application does not limit the at least one quality assessment dimension, for example, it may include dimensions such as aesthetics, reasonableness of color matching, and reasonableness of light distribution.
[0136] It should be noted that the present application does not limit the implementation method of the quality assessment processing in the above paragraph. For example, it can be implemented with the help of any existing or future method for assessing the quality status of an image (for example, with the help of a pre-built machine learning model with image quality assessment performance, etc.).
[0137] Based on the relevant contents of steps 11 to 14 above, it can be known that for the i-th candidate object above, an adjustment method can be performed on the i-th candidate object in the model to be processed to obtain an adjusted model that has a control relationship with the model to be processed under the i-th candidate object, so that the impact characterization data of the i-th candidate object can be determined by comparing the generated images of the two models in some dimensions (for example, information retention dimension, image editing dimension, image generation quality dimension, etc.), so that the impact characterization data can characterize the impact presented by the i-th candidate object in at least one dimension (for example, information retention dimension, image editing dimension, image generation quality dimension, etc.), so that it can be determined based on the impact characterization data whether it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object.
[0138] The target object refers to an object that exists in at least one of the candidate objects above and satisfies the preset retention information introduction condition, so that the target object is used to represent the object whose characterization features of the information to be retained in the reference image need to be introduced into its corresponding data processing process. Among them, the preset retention information introduction condition can be based on the actual application scenario in advance; and the present application does not limit the preset retention information introduction condition. For example, the preset retention information introduction condition can specifically be: the information retention dimension of the target object affects the characterization data higher than the target dimension of the target object affects the characterization data (for example, the information retention dimension of the target object affects the characterization data much higher than the target dimension of the target object affects the characterization data, etc.).
[0139] Based on the relevant content of S302 above, it can be known that for at least one candidate object in the above model to be processed, the influence characterization data of each candidate object is first obtained (for example, information retention dimension influence characterization data and target dimension influence characterization data, etc.); then, based on the influence characterization data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object existing in the at least one candidate object, whose influence in the information retention dimension is far higher than the influence in other dimensions. In this way, when the characterization features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, it is possible to reduce the performance in other dimensions as little as possible under the premise of greatly improving the information retention performance, thereby effectively avoiding the adverse effects caused by a significant reduction in performance in other dimensions due to the improvement of the information retention performance.
[0140] S303: According to the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0141] In the present application, after determining the target object, the model to be processed can be adjusted according to the target object to obtain an image generation model, so that the image generation model includes the target object, and when the image generation model performs image generation processing according to the reference image and the prompt text, the characterization features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object. Among them, the data processing process corresponding to the target object refers to the data processing flow pre-configured for the target object; and the present application does not limit the data processing process corresponding to the target object. For example, when the target object is a network layer, the data processing process corresponding to the target object refers to the data processing process implemented by the network layer. For another example, when the target object is a set of network layers corresponding to a time step, the data processing process corresponding to the target object refers to the data processing process implemented by each network layer in the network layer set corresponding to the time step.
[0142] Based on the above content, it can be known that in some application scenarios, for the above image generation model, the characterization features of the information to be retained in the reference image need to be introduced into the data processing process corresponding to some objects in the image generation model, but the characterization features of the information to be retained in the reference image do not need to be introduced into the data processing process corresponding to another part of the objects. Based on this, it can be known that in one possible implementation, the image generation model has the following characteristics: when the image generation model performs image generation processing based on the reference image and the prompt text, the characterization features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
[0143] In addition, the present application does not limit the implementation method of the above S303. For ease of understanding, two situations are combined for explanation below.
[0144] Case 1, when the above-mentioned model to be processed is a pre-built basic model, the above S303 can specifically be: adjust the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of other objects in the model to be processed, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
[0145] It can be seen that in some application scenarios, when the above-mentioned model to be processed is a pre-built basic model, after determining the target object, the image generation model can be obtained by adjusting the parameters involved in the target object in the model to be processed, so that the target object in the image generation model has the ability to retain information. It should be noted that the present application does not limit the adjustment method. For example, if the target object is a network layer, the adjustment method may include adding the input weight of the characterization feature of the information to be retained to the network layer, and adjusting other existing parameters in the network layer, so that the network layer is used to perform corresponding data processing based on the characterization feature of the information to be retained, so that the characterization feature of the information to be retained can be introduced into the network layer; if the target object is a set of network layers corresponding to a time step, the adjustment method includes adding the input weight of the characterization feature of the information to be retained to all network layers in the network layer set corresponding to the time step, and adjusting other existing parameters in these network layers, so that part or all of the network layers involved in the denoising process need to perform corresponding data processing based on the characterization feature of the information to be retained, so that the characterization feature of the information to be retained can be introduced into the network layer set corresponding to the time step.
[0146] Case 2: When the model to be processed above refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used for learning information retention, S303 above can specifically be: adjusting the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the objects other than the target object.
[0147] It can be seen that in some application scenarios, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used to learn information retention, after determining the target object, the image generation model can be obtained by adjusting the parameters of other objects in the model to be processed except the target object, so that the target object in the image generation model retains the ability to retain information, but other objects do not retain the ability to retain information. It should be noted that the present application does not limit the adjustment method. For example, if the other object is a network layer, the adjustment method may include deleting the input weight of the characterization feature of the information to be maintained from the network layer, and returning other existing parameters in the network layer to the state of the corresponding parameters in the basic model; if the other object is a set of network layers corresponding to a time step, the adjustment method includes returning the data processing process described by the network layer set corresponding to the time step to the data processing process described by the network layer set corresponding to the corresponding time step in the basic model (for example, deleting the input weight of the characterization feature of the information to be maintained from all network layers in the network layer set corresponding to the time step, and returning other existing parameters in the network layers to the state of the corresponding parameters in the basic model, etc.), so that all network layers in the network layer set corresponding to the time step do not need to perform corresponding data processing based on the characterization feature of the information to be maintained, thereby achieving the goal of not introducing the characterization feature of the information to be maintained into the network layer set corresponding to the time step.
[0148] Based on the relevant contents of S301 to S303 above, it can be known that for the model determination method provided in the embodiment of the present application, the model to be processed (for example, a diffusion model, etc.) is first obtained, and the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer and / or at least one time step; then, based on the information retention dimension influence characterization data and the target dimension influence characterization data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, the model to be processed is adjusted to obtain to the image generation model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, so that the image generation model can achieve performance balance as much as possible in the two opposing dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.). In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, which is beneficial to improving the image generation effect.
[0149] In fact, in order to further improve the image generation effect, the present application also provides an implementation of the above model determination method. In this implementation, the model determination method may include the following step 21 in addition to the above S301-S303.
[0150] Step 21: When the target object above is a network layer, the output weight corresponding to the target object in the image generation model is adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0151] The output weight corresponding to the target object refers to the weighted weight that needs to be used for the output data when the output data of the target object is input into the next network layer, so that the input data of the next network layer includes the result obtained by weighted processing of the output data according to the weighted weight. The output data of the target object refers to the result output by the target object.
[0152] In addition, for the image generation model adjusted above, the performance improvement value presented by the adjusted image generation model in the target dimension is determined based on the difference between the target dimension impact characterization data of the adjusted image generation model and the target dimension impact characterization data of the image generation model before adjustment, so that the performance improvement value can represent the degree of performance improvement presented in the target dimension after adjusting the output weights corresponding to the target object in the image generation model before adjustment.
[0153] In addition, for the image generation model adjusted above, the performance degradation value presented by the adjusted image generation model on the information retention dimension is determined based on the difference between the information retention dimension impact characterization data of the adjusted image generation model and the information retention dimension impact characterization data of the image generation model before adjustment, so that the performance degradation value can represent the degree of performance degradation presented on the information retention dimension after adjusting the output weights corresponding to the target object in the image generation model before adjustment.
[0154] Based on the relevant content of step 21 above, it can be known that for the target object above, when the target object is a network layer, not only the parameters involved in the data processing process corresponding to the target object will affect the model performance of the image generation model, but also the output weights corresponding to the target object will affect the model performance of the image generation model. Therefore, in order to further balance the multi-dimensional model performance, the output weights corresponding to the target object in the image generation model can be fine-tuned, so that the performance improvement value of the fine-tuned image generation model in the target dimension is much higher than the performance degradation value of the fine-tuned image generation model in the information retention dimension, thereby further improving the image generation effect.
[0155] In order to better understand the model determination method provided in this application, two situations are described below.
[0156] Case 1: In some application scenarios, an image generation model that can take into account multi-dimensional performance can be constructed by adding some information retention related parameters to an already constructed diffusion model that does not have information retention capabilities (for example, a diffusion model that can perform image generation processing based on prompt text, etc.).
[0157] Based on the above situation 1, it can be known that the model determination method provided in the present application may include the following steps 31 to 37.
[0158] Step 31: Use some sample texts to train the diffusion model to obtain a basic model so that the basic model has a text image generation function, thereby enabling the basic model to perform image generation processing based on any prompt text.
[0159] The sample text refers to the training data required to be used when the diffusion model learns the image and text generation capability; and this application is not limited to the sample text.
[0160] Step 32: Determine the information retention dimension of some or all network layers in the above basic model that affects the characterization data.
[0161] In the present application, for the above basic model, one or more network layers in the basic model can be fine-tuned to obtain information retention dimension impact characterization data of the one or more network layers, so that the information retention dimension impact characterization data can represent the impact of the one or more network layers under the information retention dimension.
[0162] It should be noted that the present application does not limit the determination process of "the information retention dimension of one or more network layers affects the representation data" in the above paragraph. For example, it can be specifically as follows: first fine-tune one or more network layers in the basic model (for example, add the input weights of the representation features of the information to be retained to the one or more network layers, and adjust other existing parameters in the one or more network layers, etc.), and obtain an adjusted model so that the one or more network layers in the adjusted model have information retention capabilities; then determine the information retention dimension of the one or more network layers that affects the representation data based on the generated image of the basic model, the generated image of the adjusted model, and the reference image. Among them, the generated image of the basic model is obtained by the basic model performing image generation processing based on the prompt text; the generated image of the adjusted model is obtained by the adjusted model performing image generation processing based on the prompt text and the reference image.
[0163] It should also be noted that in some application scenarios (for example, face preservation scenarios), the information retention dimension of some or all network layers in the above basic model affects the characterization data, which can show that the network layers that can strongly affect the information retention effect are concentrated in the upsampling module in the denoising network of the diffusion model. Therefore, some attention layers (for example, cross-attention layers and self-attention layers) in the upsampling module can be fine-tuned (for example, lora fine-tuning or direct fine-tuning of the network layer, etc.) to obtain the fine-tuned model.
[0164] It should be noted that, for the fine-tuning part (e.g., one or more network layers) in the above basic model, the present application can adjust the influence of the fine-tuning part in the information preservation dimension by adjusting the weight of the output value of the parameter component of the fine-tuning part. For the model that directly adds a new attention layer to implement the information introduction processing to be used, the influence of the fine-tuning part in the information preservation dimension can be adjusted by directly adjusting the weight of the output value of the newly added attention layer.
[0165] Step 33: Determine how the image editing dimensions of some or all of the network layers in the above basic model affect the representation data.
[0166] In the present application, for the above basic model, one or more network layers in the basic model can be fine-tuned to obtain image editing dimension impact characterization data of the one or more network layers, so that the image editing dimension impact characterization data can represent the impact of the one or more network layers under the image editing dimension. Among them, the image editing dimension refers to the aspect that the model edits and modifies its generated content according to the prompt text (for example, adding items, changing styles, modifying part of the content, etc.). In addition, there is an antagonistic relationship between the performance of the model under the image editing dimension and the performance of the model under the information retention dimension. The specific antagonistic relationship is: if the performance of the model under the image editing dimension is stronger, then the performance of the model under the information retention dimension is poorer; if the performance of the model under the information retention dimension is stronger, then the performance of the model under the image editing dimension is poorer.
[0167] It should be noted that the determination process of "the image editing dimension of one or more network layers affects the representation data" in the above paragraph is similar to the determination process of "the information retention dimension of one or more network layers affects the representation data" above. For the sake of brevity, it will not be repeated here.
[0168] In addition, for the above step 33, in some application scenarios, in order to better improve the effect of determining the impact under the image editing dimension, the editing content that is difficult to implement in the above basic model can be used as a prompt text, so that the prompt text can be used during the execution of step 33 to determine the image editing dimension impact representation data of some or all network layers in the basic model. Among them, the "difficult to implement editing content" can be obtained by a certain method (for example, a big data analysis method, or a manually specified method, etc.).
[0169] It can be seen that, under a possible implementation mode, the acquisition process of the above-mentioned "difficult to implement editing content" can be specifically as follows: after obtaining the constructed basic model, the basic model can be used to perform image generation processing on some prompt texts to obtain the generated images corresponding to these prompt texts; then, for any prompt text, the editing performance corresponding to the prompt text is determined based on the similarity between the generated image corresponding to the prompt text and the prompt text, so that the editing performance can indicate whether the prompt text belongs to the "difficult to implement editing content", so that when it is determined that the editing performance meets the preset difficult editing conditions (for example, conditions below the preset performance threshold, etc.), the prompt text is used as the "difficult to implement editing content", so that the prompt text can be used in the subsequent step 33 to determine the image editing dimension impact characterization data of some or all network layers in the basic model.
[0170] Step 34: Based on the information retention dimension affecting the characterization data of some or all network layers in the above basic model and the image editing dimension affecting the characterization data, determine the target network layer that needs to introduce the information to be retained from the basic model, and fine-tune the target network layer in the basic model to obtain the image generation model, so that the data processing process corresponding to the target network layer in the image generation model needs to introduce the characterization features of the information to be retained in the reference image, and the data processing processes corresponding to other network layers in the image generation model except the target network layer do not need to introduce the characterization features of the information to be retained in the reference image, so that the target network layer in the image generation model has the ability to retain information, and the other network layers in the image generation model except the target network layer do not have the ability to retain information.
[0171] In the present application, for the above basic model, after determining that the information retention dimension of some or all network layers in the basic model affects the representation data and the image editing dimension affects the representation data, a network layer that has a greater impact on information retention but a smaller impact (or even no obvious impact) on editing ability can be selected from these network layers as the target network layer, and the target network layer in the basic model is fine-tuned (for example, adding the input weight of the representation feature of the information to be retained to the target network layer, and adjusting other existing parameters in the target network layer, etc.) to obtain an image generation model, so that the data processing process corresponding to the target network layer in the image generation model needs to introduce the representation features of the information to be retained in the reference image, so that the target network layer in the image generation model has information retention capability.
[0172] Step 35: Adjust the output weights of the target network layer in the above image generation model so that the performance improvement value presented by the adjusted image generation model in the image editing dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0173] It should be noted that the embodiments of the above Step 35 are not limited in this application. For ease of understanding, the following will be described with reference to three examples.
[0174] Example 1, when the representation features of the information to be used in the above reference image are determined based on the existing text feature extraction module, and the above Steps 31 - 34 are implemented by means of the LoRA fine-tuning strategy, if the above target network layer includes the cross-attention layer in the upsampling module, then Step 35 above can be specifically: Adjust the output weights of all fine-tuning parameters (for example, all fine-tuning parameters involved in the LoRA fine-tuning strategy) of all cross-attention layers in the upsampling module, and adjust the output weight from 1 to a smaller or even negative coefficient, which can significantly improve the editing ability while only slightly weakening the information retention ability.
[0175] Example 2, when the representation features of the information to be used in the above reference image are determined based on the newly added image feature extraction module, if the above target network layer includes the cross-attention layer in the upsampling module, then Step 35 above can be specifically: Adjust the output weight of the first cross-attention layer in the upsampling module, and adjust the output weight from 1 to a smaller or even negative coefficient, which can significantly improve the editing ability while only slightly weakening the information retention ability.
[0176] Example 3, if the above target network layer includes all network layers of the denoising network in the image generation model, then Step 35 above can be specifically: Reduce the output weights of all fine-tuning parameter biases to half of the original (a convenient implementation is to take the average of the fine-tuned model parameters and the pre-fine-tuned model parameters). In particular, the fine-tuning parameter components of all attention layer modules in the upsampling module can further reduce the weight to 0 or 0.1 to significantly improve the editing ability.
[0177] Step 36: Determine the information retention dimension impact representation data, image editing dimension impact representation data, and image generation quality dimension impact representation data for some or all time steps in the denoising network of the above adjusted image generation model.
[0178] In the present application, for the image generation model adjusted above, the information to be retained can be introduced and adjusted for one or more time steps in the image generation model to obtain characterization data affecting the information retention dimension, characterization data affecting the image editing dimension, and characterization data affecting the image generation quality dimension for the one or more time steps, so that a decision can be made based on these data whether to introduce the characterization features of the information to be retained in the reference image in the denoising process corresponding to the time step.
[0179] It should be noted that the present application does not limit the implementation method of the above step 36.
[0180] Step 37: Based on the characterization data affecting the information retention dimension of some or all time steps in the denoising network in the adjusted image generation model above, the characterization data affecting the image editing dimension and the characterization data affecting the image generation quality dimension, determine the target time step where the information to be retained needs to be introduced from the image generation model, and regress the denoising process corresponding to other time steps in the image generation model except the target time step to the denoising process corresponding to the corresponding time step in the basic model above, so that when the image generation model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to the target time step, but it is not necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to other time steps except the target time step, so that the performance improvement value of the image generation model after regressing in the image editing dimension and the image generation quality dimension is higher than the performance degradation value of the image generation model after regressing in the information retention dimension.
[0181] It should be noted that, for the diffusion model, since the generation process of the diffusion model is multi-step iterative, so that the early time step (that is, the time step that is executed earlier in the denoising process) mainly involves the generation of composition, light perception, etc., and the later time step (that is, the time step that is executed later in the denoising process) is related to details, so in this application, it is only necessary to introduce the characterization features of the information to be retained in the reference image in the denoising process corresponding to the intermediate time step. Based on this, it can be seen that there is no need to introduce the characterization features of the information to be retained in the reference image in the denoising process corresponding to the early time step and the later time step, which can better reduce the impact of the image generation model on the light perception aesthetics and editing ability of the above basic model. In addition, the intermediate time step can be adjusted according to the actual effect. For example, in some application scenarios, it can be that after the denoising process corresponding to 30% to 50% of the time steps is iteratively executed, the characterization features of the information to be retained in the reference image are introduced in the denoising process corresponding to each time step.
[0182] Based on the relevant contents of steps 31 to 37 above, it can be known that in some application scenarios, a basic model with image editing capabilities can be trained first; then, based on the information retention dimension impact representation data and the image editing dimension impact representation data of some network layers in the basic model, determine which network layers in the basic model need to be adjusted, so that these adjusted network layers perform corresponding data processing based on the representation features of the information to be retained in the reference image, so that the image generation model obtained by adjusting these network layers in the basic model can achieve the premise of greatly improving the information retention performance while minimizing the reduction of image editing performance; then, based on the information retention dimension impact representation data, the image editing dimension impact representation data and the image generation quality dimension impact representation data of some time steps in the denoising network of the image generation model, determine the image generation model. Which time steps need to introduce the representation features of the information to be kept, and which time steps do not need to introduce the representation features of the information to be kept, so that the image generation model can be adjusted according to these conclusions in the future, so that some time steps in the adjusted image generation model need to introduce the representation features of the information to be kept, and another part of the time steps do not need to introduce the representation features of the information to be kept, so that the adjusted image generation model can better balance the information retention performance and other performances (for example, image editing performance, image generation quality performance, etc.); finally, fine-tune some weight parameters in the adjusted image generation model, so that the fine-tuned image generation model can better balance the information retention performance and other performances (for example, image editing performance, image generation quality performance, etc.), so that the image generation model finally obtained has a better image generation effect.
[0183] Case 2: In some application scenarios, it is possible to construct a diffusion model with information retention capability (e.g. Figure 1 and Figure 2 By deleting some information and keeping relevant parameters from the diffusion model shown in the figure, an image generation model that can take into account multi-dimensional performance is constructed.
[0184] Based on the above situation 2, it can be known that the model determination method provided in the present application may include the following steps 41 to 49.
[0185] Step 41: Use some sample texts to train the diffusion model to obtain a basic model, so that the basic model has a text image generation function, so that the basic model can perform image generation processing based on any prompt text.
[0186] It should be noted that the relevant content of step 41 is similar to the relevant content of step 31 above, and for the sake of brevity, it will not be repeated here.
[0187] Step 42: Based on the basic model, construct a first model (e.g., Figure 1 or Figure 2 The diffusion model shown in FIG. 1 ) is configured so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0188] It should be noted that the present application does not limit the implementation method of the above step 42. For example, in some application scenarios (for example, when the representation features of the information to be retained in the reference image are determined with the help of the original text feature extraction module in the basic model), the step 42 can specifically be: adding a graphic-text mapping module to the basic module, and fine-tuning some or all of the network layers (for example, the attention layer) in the denoising network of the basic module to obtain a first model, so that the first model can at least include the graphic-text mapping module, the text feature extraction module and the denoising network, and the input data of some or all of the network layers (for example, the attention layer) in the denoising network include the output data of the text feature extraction module, so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0189] For example, in some application scenarios (for example, when the representation features of the information to be retained in the reference image are determined with the help of an additional image feature extraction module), step 42 can specifically be: adding an image feature extraction module to the basic module, and fine-tuning some or all of the network layers (for example, the attention layer) in the denoising network of the basic module to obtain a first model, so that the first model can at least include the image feature extraction module and the denoising network, and the input data of some or all of the network layers (for example, the attention layer) in the denoising network include the output data of the image feature extraction module, so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0190] Step 43: Using some sample images and the prompt texts corresponding to these sample images, fine-tune the first model to obtain a second model, so that the second model has better information retention capability. The fine-tuning process is used to learn information retention.
[0191] The sample image refers to the training data required to be used when the first model learns the information retention capability; and the present application does not limit the sample image.
[0192] In addition, the present application does not limit the implementation method of the fine-tuning process in step 43 above. For example, during the fine-tuning process, the sample image is used as a reference image so that the sample image is used to provide the information to be retained that needs to be retained during the image generation process. In this way, the second model obtained through step 43 can learn how to retain the information to be retained in the sample image, thereby enabling the second model to have a relatively strong information retention capability.
[0193] Step 44: Determine the information retention dimension of some or all network layers in the second model that affects the characterization data.
[0194] In the present application, for the second model above, one or more network layers in the second model can be fine-tuned to obtain information retention dimension impact characterization data of the one or more network layers, so that the information retention dimension impact characterization data can represent the impact of the one or more network layers under the information retention dimension.
[0195] It should be noted that the present application does not limit the determination process of "the information retention dimension of one or more network layers affects the characterization data" in the above paragraph. For example, it can be specifically as follows: first fine-tune one or more network layers in the second model (for example, revert one or more network layers in the second model to the state of the corresponding network layers in the above basic model, etc.), and obtain an adjusted model so that the one or more network layers in the adjusted model do not have the ability to retain information; then determine the information retention dimension of the one or more network layers that affects the characterization data based on the generated image of the second model, the generated image of the adjusted model, and the reference image. Among them, the generated image of the second model is obtained by the second model performing image generation processing based on the prompt text and the reference image; the generated image of the adjusted model is obtained by the adjusted model performing image generation processing based on the prompt text and the reference image.
[0196] It should also be noted that, for the fine-tuning part (e.g., one or more network layers) in the second model above, the present application can adjust the influence of the fine-tuning part in the information preservation dimension by adjusting the weight of the output value of the parameter component of the fine-tuning part. For the model that implements the introduction of the information to be used by directly adding a new attention layer, the influence of the fine-tuning part in the information preservation dimension can be adjusted by directly adjusting the weight of the output value of the newly added attention layer.
[0197] Step 45: Determine the image editing dimension impact of some or all network layers in the second model above on the representation data.
[0198] In the present application, for the second model above, one or more network layers in the second model can be fine-tuned to obtain image editing dimension impact characterization data of the one or more network layers, so that the image editing dimension impact characterization data can represent the impact of the one or more network layers under the image editing dimension.
[0199] It should be noted that the determination process of "the image editing dimension of one or more network layers affects the representation data" in the above paragraph is similar to the determination process of "the information retention dimension of one or more network layers affects the representation data" above. For the sake of brevity, it will not be repeated here.
[0200] In addition, for step 45 above, in some application scenarios, in order to better improve the effect of determining the impact under the image editing dimension, the editing content that is difficult to implement in the second model above can be used as a prompt text, so that the prompt text can be used during the execution of step 33 to determine the image editing dimension impact representation data of some or all network layers in the second model. Among them, the "difficult to implement editing content" can be obtained by a certain method (for example, a big data analysis method, or a manually specified method, etc.).
[0201] It can be seen that, under a possible implementation mode, the acquisition process of the above-mentioned "difficult to implement editing content" can be specifically as follows: after obtaining the constructed second model, the second model can be used to perform image generation processing on some prompt texts to obtain the generated images corresponding to these prompt texts; then, for any prompt text, the editing performance corresponding to the prompt text is determined based on the similarity between the generated image corresponding to the prompt text and the prompt text, so that the editing performance can indicate whether the prompt text belongs to the "difficult to implement editing content", so that when it is determined that the editing performance meets the preset difficult editing conditions (for example, conditions below the preset performance threshold, etc.), the prompt text is used as the "difficult to implement editing content", so that the prompt text can be used in the subsequent step 45 to determine the image editing dimension impact characterization data of some or all network layers in the second model.
[0202] Step 46: Based on the information retention dimension impact characterization data of some or all of the network layers in the second model above and the image editing dimension impact characterization data, determine the network layers to be adjusted that do not need to introduce the information to be retained (that is, other network layers except the target network layer that needs to introduce the information to be retained as mentioned above) from the second model, and roll back the network layers to be adjusted in the second model to the state of the corresponding network layers in the basic model above, to obtain the image generation model, so that the data processing process corresponding to the network layers to be adjusted in the image generation model does not need to introduce the characterization features of the information to be retained in the reference image.
[0203] In the present application, for the second model above, after determining that the information retention dimension of some or all network layers in the second model affects the characterization data and the image editing dimension affects the characterization data, a network layer that has a relatively large improvement in editing capability but a relatively small (even unnoticeable) decrease in information retention can be selected from these network layers as the network layer to be adjusted, and the network layer to be adjusted in the second model is rolled back to the state of the corresponding network layer in the basic model above to obtain an image generation model, so that the data processing process corresponding to the network layer to be adjusted in the image generation model does not need to introduce the characterization features of the information to be retained in the reference image.
[0204] Step 47: Adjust the output weights of the target network layer in the above image generation model so that the performance improvement value presented by the adjusted image generation model in the image editing dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension. The target network layer refers to the network layer in the image generation model that needs to perform corresponding data processing based on the representation features of the information to be retained in the reference image.
[0205] It should be noted that the relevant content of step 47 is similar to the relevant content of step 35 above, and for the sake of brevity, it will not be repeated here.
[0206] Step 48: Determine the characterization data affected by the information preservation dimension of some or all time steps in the denoising network in the image generation model adjusted above, the characterization data affected by the image editing dimension, and the characterization data affected by the image generation quality dimension.
[0207] It should be noted that the relevant content of step 48 is similar to the relevant content of the above step 36, and for the sake of brevity, it will not be repeated here.
[0208] Step 49: Based on the characterization data affecting the information retention dimension of some or all time steps in the denoising network in the adjusted image generation model above, the characterization data affecting the image editing dimension and the characterization data affecting the image generation quality dimension, determine the target time step where the information to be retained needs to be introduced from the image generation model, and regress the denoising process corresponding to other time steps in the image generation model except the target time step to the denoising process corresponding to the corresponding time step in the basic model above, so that when the image generation model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to the target time step, but it is not necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to other time steps except the target time step, so that the performance improvement value of the image generation model after regressing in the image editing dimension and the image generation quality dimension is higher than the performance degradation value of the image generation model after regressing in the information retention dimension.
[0209] It should be noted that the relevant content of step 49 is similar to the relevant content of step 37 above, and for the sake of brevity, it will not be repeated here.
[0210] Based on the relevant contents of steps 41 to 49 above, it can be known that in some application scenarios, a basic model with image editing capabilities can be trained first; then, based on the basic model, a second model with information retention capabilities can be constructed; secondly, based on the information retention dimension impact characterization data and the image editing dimension impact characterization data of some network layers in the second model, it is determined that those network layers in the second model need to be rolled back to the states of the corresponding network layers in the basic model, so that the image generation model obtained by rolling back these network layers in the second model can achieve a significant improvement in information retention performance while minimizing the reduction in image editing performance; then, based on the information retention dimension impact characterization data, the image editing dimension impact characterization data and the image generation quality dimension impact characterization data of some time steps in the denoising network of the image generation model, the image generation model is determined. Which time steps need to introduce the representation features of the information to be kept, and which time steps do not need to introduce the representation features of the information to be kept, so that the image generation model can be adjusted according to these conclusions in the future, so that some time steps in the adjusted image generation model need to introduce the representation features of the information to be kept, and another part of the time steps do not need to introduce the representation features of the information to be kept, so that the adjusted image generation model can better balance the information retention performance and other performance (for example, image editing performance, image generation quality performance, etc.); finally, fine-tune some weight parameters in the adjusted image generation model, so that the fine-tuned image generation model can better balance the information retention performance and other performance (for example, image editing performance, image generation quality performance, etc.), so that the image generation model finally obtained has a better image generation effect.
[0211] In fact, after determining the above image generation model, the image generation model can be used to perform certain image generation tasks. Based on this, the present application also provides an image generation method, such as Figure 4 As shown, the image generation method may include the following S401-S402. Figure 4 A flowchart of an image generation method provided in an embodiment of the present application.
[0212] S401: Acquire information description images and image constraint description text to be retained.
[0213] Among them, the image describing the information to be retained is used to limit the information that needs to be retained during the image generation process, so that the image describing the information to be retained can be used as a reference image during the image generation process; and the present application does not limit the method of obtaining the image describing the information to be retained. For example, in some application scenarios, the image describing the information to be retained can be an image specified by the user with the help of certain operations (such as image input operations or image selection operations, etc.).
[0214] The image constraint description text is used to limit the constraints required in the image generation process so that the image constraint description text can be used as prompt text in the image generation process; and the present application does not limit the method of obtaining the image constraint description text. For example, in some application scenarios, the image constraint description text can be text content specified by the user through certain operations (such as text input operations or text selection operations, etc.).
[0215] S402: Input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined by any implementation of the model determination method provided in the embodiments of the present application.
[0216] For more information about the image generation model, please refer to the above.
[0217] The target image refers to the image obtained by the image generation model performing image generation processing based on the above-mentioned image description of the information to be retained and the image constraint description text, so that the target image not only satisfies the constraints described by the image constraint description text, but also makes the target image consistent with the image description of the information to be retained in terms of the information to be retained.
[0218] Based on the relevant contents of S401 to S402 above, it can be known that in some application scenarios, a pre-built image generation model can be used to perform image generation processing based on the image describing the information to be maintained (for example, a reference image specified by the user) and the image constraint description text (for example, a prompt text specified by the user) to obtain a target image, so that the target image not only satisfies the constraints described by the image constraint description text, but also makes the target image consistent with the image describing the information to be maintained in terms of the information to be maintained, which is conducive to improving the image generation effect.
[0219] Based on the model determination method provided in the embodiment of the present application, the embodiment of the present application also provides a model determination device. Figure 5 Explain and illustrate. Figure 5 This is a schematic diagram of the structure of a model determination device provided in an embodiment of the present application. It should be noted that for the technical details of the model determination device provided in an embodiment of the present application, please refer to the relevant content of the model determination method above.
[0220] like Figure 5 As shown, the model determination device 500 provided in the embodiment of the present application includes:
[0221] A first acquisition unit 501 is used to acquire a model to be processed, where the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer;
[0222] A first determining unit 502 is configured to determine a target object from the at least one candidate object according to the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined according to a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object;
[0223] The first adjustment unit 503 is used to adjust the model to be processed according to the target object to obtain an image generation model, wherein the image generation model includes the target object, and when the image generation model performs image generation processing according to the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0224] In a possible implementation manner, different candidate objects are different network layers;
[0225] or,
[0226] Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
[0227] In one possible implementation, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
[0228] In a possible implementation manner, if different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object under the image editing dimension;
[0229] If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact characterization data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or the image generation quality dimension.
[0230] In a possible implementation manner, the model determination device 500 further includes:
[0231] A second adjustment unit is used for adjusting the model to be processed according to any candidate object to obtain an adjusted model, so that the state of the candidate object in the adjusted model is different from the state of the candidate object in the model to be processed;
[0232] a third acquisition unit, configured to acquire a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by the model to be processed performing image generation processing on the prompt text, and the second generated image is determined by the adjusted model performing image generation processing based on the reference image and the prompt text;
[0233] a third determining unit, configured to determine the information retention dimension impact characterization data of the candidate object based on the difference characterization data between the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image; the information retention dimension performance characterization data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance characterization data of the second generated image is determined based on the second generated image and the reference image;
[0234] The fourth determination unit is used to determine the target dimension impact characterization data of the candidate object based on the gap characterization data between the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image; the target dimension performance characterization data of the first generated image is determined based on the first generated image; the target dimension performance characterization data of the second generated image is determined based on the second generated image.
[0235] In a possible implementation manner, when generating the first generated image using the model to be processed, it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0236] When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0237] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention;
[0238] The second adjustment unit is specifically used to: adjust the candidate object in the model to be processed to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
[0239] In a possible implementation manner, when generating the first generated image using the model to be processed, it is not necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0240] When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0241] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0242] The second adjustment unit is specifically used to: adjust the candidate object in the model to be processed to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
[0243] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by the model to be processed performing image generation processing according to the reference image and the prompt text;
[0244] or,
[0245] The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
[0246] In one possible implementation, if different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
[0247] In a possible implementation manner, if different candidate objects are network layer sets corresponding to different time steps, the target dimension impact characterization data of the candidate object includes image editing dimension impact characterization data and / or image generation quality dimension impact characterization data; the image editing dimension impact characterization data is determined based on the gap characterization data between the image editing dimension performance characterization data of the first generated image and the image editing dimension performance characterization data of the second generated image; the image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text; the image generation quality dimension impact characterization data is determined based on the gap characterization data between the image generation quality dimension performance characterization data of the first generated image and the image generation quality dimension performance characterization data of the second generated image; the image generation quality dimension performance characterization data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance characterization data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
[0248] In one possible implementation, when the image generation model performs image generation processing based on the reference image and the prompt text, there is no need to introduce representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
[0249] In a possible implementation manner, the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention;
[0250] The first adjustment unit 503 is specifically used to: perform adjustment processing on other objects in the model to be processed except the target object to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the basic model.
[0251] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0252] The first adjustment unit 503 is specifically used to: adjust the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the model to be processed.
[0253] In a possible implementation manner, if different candidate objects are different network layers, the model determination device 500 further includes:
[0254] The third adjustment unit is used to adjust the output weights corresponding to the target object in the image generation model so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0255] Based on the relevant contents of the above-mentioned model determination device 500, it can be known that for the model determination device 500 provided in the embodiment of the present application, the model to be processed (for example, a diffusion model, etc.) is first obtained, and the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; then, based on the information retention dimension influence characterization data and the target dimension influence characterization data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, the model to be processed is adjusted to obtain an image Generate a model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, so that the image generation model can achieve performance balance in the two dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.) as much as possible. In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, which is beneficial to improving the image generation effect.
[0256] Based on the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device. Figure 6 Explain and illustrate. Figure 6This is a schematic diagram of the structure of an image generating device provided in an embodiment of the present application. It should be noted that for the technical details of the image generating device provided in an embodiment of the present application, please refer to the relevant contents of the image generating method above.
[0257] like Figure 6 As shown, the image generating device 600 provided in the embodiment of the present application includes:
[0258] The second acquisition unit 601 is used to acquire the information description image to be retained and the image constraint description text;
[0259] The second determination unit 602 is used to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using any implementation of the model determination method provided in the present application.
[0260] Based on the relevant content of the above-mentioned image generating device 600, it can be known that for the image generating device 600 provided in the embodiment of the present application, a pre-constructed image generation model can be used to perform image generation processing based on an image describing the information to be retained (for example, a reference image specified by a user) and an image constraint description text (for example, a prompt text specified by a user) to obtain a target image, so that the target image not only satisfies the constraints described by the image constraint description text, but also makes the target image consistent with the image describing the information to be retained in terms of the information to be retained, which is conducive to improving the image generation effect.
[0261] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the model determination method or image generation method provided in the embodiment of the present application.
[0262] See also Figure 7 , which shows a schematic diagram of the structure of an electronic device 700 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0263] like Figure 7As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0264] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0265] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0266] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept, and the technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0267] An embodiment of the present application also provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the model determination method or image generation method provided in the embodiment of the present application.
[0268] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0269] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0270] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0271] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.
[0272] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0273] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0274] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit / module does not, in some cases, constitute a limitation on the unit itself.
[0275] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0276] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0277] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0278] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0279] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0280] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0281] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model determination method, It is characterized in that The method comprises: Acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; Determine a target object from the at least one candidate object based on the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined based on a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object; According to the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
2. The method according to claim 1, It is characterized in that Different candidate objects are different network layers; or, Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
3. The method according to claim 2, It is characterized in that For any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
4. The method according to claim 1, It is characterized in that If different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object under the image editing dimension; If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact characterization data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or the image generation quality dimension.
5. The method according to claim 1, It is characterized in that For any candidate object, the process of determining the impact characterization data of the candidate object includes: According to the candidate object, the model to be processed is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is different from the state of the candidate object in the model to be processed; Acquire a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by the model to be processed performing image generation processing on the prompt text, and the second generated image is determined by the adjusted model performing image generation processing based on the reference image and the prompt text; Determining information retention dimension impact characterization data of the candidate object based on gap characterization data between information retention dimension performance characterization data of the first generated image and information retention dimension performance characterization data of the second generated image; the information retention dimension performance characterization data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance characterization data of the second generated image is determined based on the second generated image and the reference image; The target dimension impact characterization data of the candidate object is determined based on the gap characterization data between the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image; the target dimension performance characterization data of the first generated image is determined based on the first generated image; the target dimension performance characterization data of the second generated image is determined based on the second generated image.
6. The method according to claim 5, It is characterized in that When generating the first generated image using the model to be processed, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object; When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
7. The method according to claim 6, It is characterized in that The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, wherein the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes: The candidate object in the to-be-processed model is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
8. The method according to claim 5, It is characterized in that When generating the first generated image using the model to be processed, it is not necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object; When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
9. The method according to claim 8, It is characterized in that The model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text; The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes: The candidate object in the model to be processed is adjusted to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
10. The method according to claim 5, It is characterized in that The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by the image generation processing performed by the model to be processed according to the reference image and the prompt text; or, The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
11. The method according to claim 5, It is characterized in that If different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined according to the similarity between the first generated image and the prompt text; The target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
12. The method according to claim 5, It is characterized in that If different candidate objects are network layer sets corresponding to different time steps, the target dimension impact representation data of the candidate object includes image editing dimension impact representation data and / or image generation quality dimension impact representation data; The image editing dimension impact characterization data is determined based on the difference characterization data between the image editing dimension performance characterization data of the first generated image and the image editing dimension performance characterization data of the second generated image; The image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text; The image generation quality dimension impact characterization data is determined based on the gap characterization data between the image generation quality dimension performance characterization data of the first generated image and the image generation quality dimension performance characterization data of the second generated image; The image generation quality dimension performance characterization data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance characterization data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
13. The method according to claim 1, It is characterized in that When the image generation model performs image generation processing based on the reference image and the prompt text, there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
14. The method according to claim 13, It is characterized in that The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, wherein the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; The step of adjusting the model to be processed according to the target object to obtain an image generation model includes: Adjust the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model.
15. The method according to claim 13, It is characterized in that The model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text; The step of adjusting the model to be processed according to the target object to obtain an image generation model includes: The target object in the model to be processed is adjusted to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the model to be processed.
16. The method according to claim 1, It is characterized in that If different candidate objects are different network layers, the method further includes: The output weights corresponding to the target object in the image generation model are adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
17. A method for generating an image, It is characterized in that The method comprises: Obtain the information description image and image constraint description text to be retained; The image describing the information to be retained and the image constraint description text are input into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method described in any one of claims 1-16.
18. A model determination device, It is characterized in that include: A first acquisition unit, configured to acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; A first determining unit is used to determine a target object from the at least one candidate object according to the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined according to a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object; A first adjustment unit is used to adjust the model to be processed according to the target object to obtain an image generation model, wherein the image generation model includes the target object, and when the image generation model performs image generation processing based on the reference image and the prompt text, characterization features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
19. An image generating device, It is characterized in that include: A second acquisition unit is used to acquire the information description image to be retained and the image constraint description text; The second determination unit is used to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method described in any one of claims 1-16.
20. An electronic device, It is characterized in that The device comprises: a processor and a memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1 to 17.
21. A computer readable medium, It is characterized in that The computer-readable medium stores instructions or computer programs, and when the instructions or computer programs are executed on a device, the device executes the method according to any one of claims 1 to 17.