Model determination method, image generation method, apparatus, device and medium
By determining and adjusting the target objects in the image generation model and introducing the representation features of the reference image, the shortcomings of the existing model in retaining image information are solved, and the balance between information preservation and image generation effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/133793
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing image generation models have shortcomings in maintaining certain information in the reference image and cannot effectively retain specific information in the image, such as face or hairstyle.
By obtaining the pending model, the target object is determined based on the influence characterization data of the candidate object, and the pending model is adjusted to generate an image generation model. The image generation model introduces characterization features of information to be maintained in the reference image when generating the image to be generated to ensure that the performance of the information holding dimension is preferred.
It realizes the performance of the image generation model in the target dimension (such as image editing and aesthetics) while maintaining certain information in the reference image, and achieves a balance between information maintenance and image generation effect.
Smart Images

Figure CN2024133793_30052025_PF_FP_ABST
Abstract
Description
Model determination method, image generation method, device, equipment, and medium
[0001] This application claims priority to Chinese Patent Application No. 202311568533.5 filed on November 22, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field
[0002] The present application relates to a model determination method, image generation method, device, equipment, and medium. Background Art
[0003] In some application scenarios, a pre-built image generation model (e.g., a diffusion model) may be used to perform image generation tasks (e.g., tasks of generating images based on text). For example, in some application scenarios, the image generation model may perform image generation processing based on user-provided prompt text (e.g., the text "a puppy wearing sunglasses") to ensure that the generated image conforms to the semantic content represented by the prompt text.
[0004] However, for some application scenarios, the above image generation model may be further required to have the ability to preserve certain information in the reference image (e.g., facial preservation, hairstyle preservation, etc.). For example, when the reference image is used to describe target 1 (e.g., a person, animal, or body part), and the image generation model is used to perform image generation processing based on the prompt text provided by the user and the reference image, not only is the generated image required to conform to the semantic content represented by the prompt text, but the object described by the generated image is also required to be target 1, so as to achieve image processing (e.g., image style change processing, etc.) while preserving the description information of target 1 in the reference image. Summary of the Invention
[0005] The present application provides a model determination method, image generation method, device, equipment, and medium to achieve the purpose of image processing while maintaining certain information in a reference image.
[0006] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0007] The present application provides a model determination method, the method comprising:
[0008] Acquire a model to be processed, where the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer;
[0009] Determining a target object from the at least one candidate object based on the influence representation data of each candidate object; the influence representation data includes information-retention dimension influence representation data and target dimension influence representation data; the influence representation data is determined based on a reference image and a prompt text; the information-retention dimension influence representation data of the target object is higher than the target dimension influence representation data of the target object;
[0010] According to the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0011] In a possible implementation manner, different candidate objects are different network layers;
[0012] or,
[0013] Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
[0014] In one possible implementation, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
[0015] In a possible implementation manner, if different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object in the image editing dimension;
[0016] If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or image generation quality dimension.
[0017] In one possible implementation, for any candidate object, the process of determining the impact characterization data of the candidate object includes:
[0018] Adjusting the model to be processed according to the candidate object to obtain an adjusted model, so that the state of the candidate object in the adjusted model is different from the state of the candidate object in the model to be processed;
[0019] Obtaining a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by performing image generation processing by the to-be-processed model on the prompt text, and the second generated image is determined by performing image generation processing by the adjusted model based on the reference image and the prompt text;
[0020] determining information retention dimension impact representation data of the candidate object based on gap representation data between information retention dimension performance representation data of the first generated image and information retention dimension performance representation data of the second generated image; the information retention dimension performance representation data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance representation data of the second generated image is determined based on the second generated image and the reference image;
[0021] The target dimension impact representation data of the candidate object is determined based on the gap representation data between the target dimension performance representation data of the first generated image and the target dimension performance representation data of the second generated image; the target dimension performance representation data of the first generated image is determined based on the first generated image; the target dimension performance representation data of the second generated image is determined based on the second generated image.
[0022] In a possible implementation manner, when generating the first generated image using the model to be processed, it is necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0023] When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0024] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention;
[0025] The step of adjusting the model to be processed based on the candidate object to obtain an adjusted model includes:
[0026] An adjustment process is performed on the candidate object in the to-be-processed model to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
[0027] In one possible implementation, when generating the first generated image using the model to be processed, it is not necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0028] When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0029] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0030] The step of adjusting the model to be processed based on the candidate object to obtain an adjusted model includes:
[0031] The candidate object in the model to be processed is adjusted to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representational features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
[0032] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built base model, the base model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by performing image generation processing by the model to be processed based on the reference image and the prompt text;
[0033] or,
[0034] The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
[0035] In a possible implementation manner, if different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined based on a similarity between the first generated image and the prompt text;
[0036] The target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
[0037] In one possible implementation, if different candidate objects are network layer sets corresponding to different time steps, the target dimension impact representation data of the candidate object includes image editing dimension impact representation data and / or image generation quality dimension impact representation data;
[0038] The image editing dimension impact characterization data is determined based on the difference characterization data between the image editing dimension performance characterization data of the first generated image and the image editing dimension performance characterization data of the second generated image; the image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text;
[0039] The image generation quality dimension impact characterization data is determined based on the gap characterization data between the image generation quality dimension performance characterization data of the first generated image and the image generation quality dimension performance characterization data of the second generated image; the image generation quality dimension performance characterization data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance characterization data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
[0040] In one possible implementation, when the image generation model performs image generation processing based on the reference image and the prompt text, there is no need to introduce representational features of the information to be retained in the reference image into the data processing corresponding to other objects except the target object.
[0041] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention;
[0042] The step of adjusting the model to be processed according to the target object to obtain an image generation model includes:
[0043] Adjustments are performed on the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model.
[0044] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0045] The step of adjusting the model to be processed according to the target object to obtain an image generation model includes:
[0046] The target object in the model to be processed is adjusted to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the model to be processed.
[0047] In a possible implementation manner, if different candidate objects are different network layers, the method further includes:
[0048] The output weights corresponding to the target object in the image generation model are adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0049] The present application provides an image generation method, the method comprising:
[0050] Obtain the image describing the information to be retained and the image constraint description text;
[0051] The image describing the information to be retained and the image constraint description text are input into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method provided in this application.
[0052] The present application provides a model determination device, comprising:
[0053] A first acquiring unit is configured to acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer;
[0054] a first determining unit, configured to determine a target object from the at least one candidate object based on the influence representation data of each candidate object; the influence representation data comprising information-retention dimension influence representation data and target dimension influence representation data; the influence representation data being determined based on a reference image and a prompt text; the information-retention dimension influence representation data of the target object being higher than the target dimension influence representation data of the target object;
[0055] A first adjustment unit is used to adjust the model to be processed according to the target object to obtain an image generation model, where the image generation model includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representational features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0056] The present application provides an image generation device, comprising:
[0057] A second acquiring unit is used to acquire the image describing the information to be retained and the image constraint description text;
[0058] The second determination unit is used to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain the target image output by the image generation model; the image generation model is determined using the model determination method provided in this application.
[0059] The present application provides an electronic device, the device comprising: a processor and a memory;
[0060] The memory is used to store instructions or computer programs;
[0061] The processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes the model determination method or image generation method provided in this application.
[0062] The present application provides a computer-readable medium having instructions or computer programs stored therein. When the instructions or computer programs are executed on a device, the device executes the model determination method or image generation method provided in the present application.
[0063] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the model determination method or image generation method provided by the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0065] FIG1 is a schematic structural diagram of a diffusion model provided in an embodiment of the present application;
[0066] FIG2 is a schematic diagram of the structure of another diffusion model provided in an embodiment of the present application;
[0067] FIG3 is a flow chart of a model determination method provided in an embodiment of the present application;
[0068] FIG4 is a flow chart of an image generation method provided in an embodiment of the present application;
[0069] FIG5 is a schematic structural diagram of a model determination device provided in an embodiment of the present application;
[0070] FIG6 is a schematic structural diagram of an image generating device provided in an embodiment of the present application;
[0071] FIG7 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0072] Research has found that for some diffusion models, the diffusion model has a strong generation capability, but the diffusion model does not have the ability to maintain certain information in the reference image. For example, given a picture of a cat, these diffusion models cannot guarantee that the cat in the generated image is still the cat in the reference image. Based on this, in order to enable these diffusion models to have the ability to maintain certain information in the reference image (for example, the ability of the diffusion model shown in Figure 1, etc.), the following method can be used: first train an information extraction module that can extract and process this information (for example, the information extraction module shown in Figure 2), so that the data extracted by the information extraction module can be used to guide these diffusion models to generate images that can maintain this information. For ease of understanding, the following example is used to illustrate.
[0073] As an example, in some application scenarios (e.g., facial information preservation scenarios, etc.), for a diffusion model capable of preserving facial information in a reference image, the diffusion model may include at least an information extraction module (e.g., the information extraction module shown in FIG2 ) and a denoising network (e.g., the denoising network shown in FIG2 ). The information extraction module is used to obtain the representational features of facial information from the reference image, so that the representational features can represent the facial information carried by the reference image. The denoising network is used to implement image generation processing through denoising; and the present application does not limit the denoising network, for example, it can be implemented using a Unet network. In addition, each network layer in the denoising network can perform corresponding data processing based on the representational features, so as to achieve the purpose of preserving facial information.
[0074] Research has also found that for the diffusion model with information retention capabilities described in the above two paragraphs, the image generated by the diffusion model is limited by the reference image. Specifically, if the information retention capability of the diffusion model is stronger, the similarity between the image generated by the diffusion model and the reference image will be greater, resulting in the diffusion model's performance in other aspects (for example, image editing based on prompt text, aesthetics, color matching, etc.) being worse, which in turn leads to poor image generation effects.
[0075] Based on the above research, it can be known that in order to better improve the image generation effect, the present application provides a model determination method, which includes: first obtaining a model to be processed (such as a diffusion model, etc.), the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; then, based on the information retention dimension influence representation data and the target dimension influence representation data of each candidate object, determine the target object from the at least one candidate object, so that the information retention dimension influence representation data of the target object is higher than the target dimension influence representation data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, adjust the model to be processed to obtain the image Generate a model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, thereby enabling the image generation model to achieve performance balance in the two dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.) as much as possible. In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, thereby helping to improve the image generation effect.
[0076] In addition, the present application does not limit the execution subject of the model determination method provided in the embodiment of the present application. For example, the model determination method provided in the embodiment of the present application can be applied to a terminal device or a server. For another example, the model determination method provided in the embodiment of the present application can also be implemented with the help of a data interaction process between a terminal device and a server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.
[0077] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0078] To better understand the technical solutions provided by this application, the model determination method provided by this application is described below with reference to some figures. As shown in Figure 3, the model determination method provided by an embodiment of this application includes the following steps S301-S303. Figure 3 is a flow chart of a model determination method provided by an embodiment of this application.
[0079] S301: Acquire a model to be processed, where the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer.
[0080] Among them, the model to be processed refers to the model required to be used when constructing an image generation model that can take into account multi-dimensional performance (for example, information retention performance, image editing performance, image generation quality performance, etc.).
[0081] In addition, for the above-mentioned model to be processed, the model to be processed may include at least one candidate object. The candidate object refers to an object that exists in the model to be processed and can participate in the image generation process. Furthermore, this application does not limit the implementation method of the candidate object. For example, when the model to be processed is implemented using a diffusion model, the candidate object may be a network layer, or part of the network layer corresponding to the time step, or all the network layers corresponding to the time step. It can be seen that, in one possible implementation, the at least one candidate object may include at least one network layer, and different candidate objects are different network layers, so that an image generation model that can take into account multi-dimensional performance can be constructed by adjusting the relevant parameters in these network layers. In another possible implementation, the at least one candidate object may include a set of network layers corresponding to at least one time step, and different candidate objects are sets of network layers corresponding to different time steps, so that an image generation model that can take into account multi-dimensional performance can be constructed by adjusting the data processing processes corresponding to these time steps (for example, the denoising processes corresponding to these time steps). For any time step, the set of network layers corresponding to the time step includes at least one network layer.
[0082] In addition, the present application does not limit the implementation method of the network layer set corresponding to the above time step. For example, in some application scenarios, the network layer set corresponding to the time step can be used to record all network layers that exist in the model to be processed and have a corresponding relationship with the time step. For another example, in some application scenarios, in order to improve the multi-dimensional performance and balance effect of the image generation model as much as possible, the present application also provides a possible implementation method of the network layer set corresponding to the time step. Under this implementation method, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step, so that the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer. This is conducive to screening out which time step needs to introduce the information to be retained and which network layers corresponding to this time step need to enter the information to be retained, thereby helping to better improve the multi-dimensional performance and balance effect of the final generated image generation model. Among them, the at least one network layer corresponding to the time step refers to the network layer that exists in the model to be processed and has a corresponding relationship with the time step; and the present application does not limit the implementation method of the at least one network layer corresponding to the time step. For example, the at least one network layer corresponding to the time step may include all network layers that exist in the model to be processed and have a corresponding relationship with the time step. It should be noted that the relevant content of the information retention dimension impact characterization data and the target dimension impact characterization data is shown below; and the implementation method of the determination process of "the network layer is determined from the at least one network layer corresponding to the time step based on the impact characterization data of the at least one network layer corresponding to the time step" is similar to the determination process of the target object when the target object is a network layer below. For the sake of brevity, it will not be repeated here.
[0083] In addition, this application does not limit the implementation method of the above-mentioned model to be processed. For ease of understanding, the following description is combined with two situations.
[0084] Case 1: In some application scenarios, the present application can construct an image generation model that can take into account multi-dimensional performance by adding some information retention-related parameters to an already constructed diffusion model that does not have information retention capabilities (for example, a diffusion model that can perform image generation processing based on prompt text, etc.).
[0085] Based on the above situation 1, it can be known that in one possible implementation, the above model to be processed can be a pre-built basic model, and the basic model is used to perform image generation processing based on the prompt text, that is, the basic model has the ability to generate text images. Among them, the basic model refers to a pre-built diffusion model that can perform image generation processing based on the prompt text; and the basic model does not have the ability to retain information. In addition, the basic model can refer to a text image generation model that has been pre-trained based on a large amount of sample text, and this application does not limit the construction method of the basic model. For example, in some application scenarios, the basic model can be implemented using any existing or future method that can be used to construct a text image generation model.
[0086] Based on the content of the above paragraph, it can be seen that in some application scenarios, when the above-mentioned model to be processed is a pre-built basic model, the model to be processed may at least include a denoising network. Among them, the denoising network is used to implement denoising processing; and this application does not limit the denoising network. For example, the denoising network uses an upsampling module, a downsampling module and a time step sequence to implement denoising processing. Among them, the upsampling module and the downsampling model are components of the denoising network. The upsampling module is used to perform data compression processing. The downsampling module is used to perform data decompression processing (that is, data recovery processing). The time step sequence is a parameter of the denoising network, and the time step sequence is used to describe which time steps need to be executed in the denoising network. The data processing process corresponding to the corresponding time steps (that is, the denoising process corresponding to the corresponding time steps needs to be executed).
[0087] Case 2: In some application scenarios, an image generation model that can take into account multi-dimensional performance can be constructed by deleting some information-preserving related parameters from an already constructed diffusion model with information preservation capabilities (for example, the diffusion models shown in Figures 1 and 2).
[0088] Based on the above situation 2, it can be known that in one possible implementation, the above-mentioned model to be processed may refer to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used to learn information retention so that the model to be processed has information retention capabilities. That is, the model to be processed can perform image generation processing based on a reference image and a prompt text so that part of the information in the final generated image (for example, facial information, etc.) is consistent with part of the information in the reference image, so that the model to be processed has information retention capabilities. Among them, the reference image is used to provide the information required to be maintained during the image generation process (for example, facial information, etc.). The prompt text is used to describe in text some constraint information required to be based on the image generation process (for example, constraints such as the style of the generated image).
[0089] Based on the content of the above paragraph, it can be seen that in some application scenarios, when the model to be processed in the above text refers to a model obtained by fine-tuning a pre-built basic model (for example, the diffusion model shown in Figure 1 or Figure 2), the model to be processed may include an information extraction module and a denoising network; and the working principle of the model to be processed is as follows: the information extraction module first obtains the characterization features of the information to be retained from the reference image, so that the characterization features can represent the information to be retained; and then the denoising network performs denoising based on the characterization features of the information to be retained to obtain the final generated image, so that the final generated image maintains a high degree of consistency with the reference image in the information to be retained. Among them, the information extraction module is used to obtain the characterization features of the information to be retained from the reference image. The information to be retained refers to the information that exists in the reference image and needs to be retained during the image generation process; and this application does not limit the information to be retained. For example, the information to be retained may refer to part of the information existing in the reference image (for example, information used to describe the facial area, information used to describe animals, information used to describe objects, etc.).
[0090] In addition, this application does not limit the implementation method of the information extraction module in the above paragraph. For ease of understanding, two examples are used below for illustration.
[0091] Example 1, in some application scenarios, in order to reduce the workload of fine-tuning the basic model as much as possible when building a model with information retention capability based on the basic model, the text feature extraction module (for example, Text Encoder) already in the basic model can be used to complete the step of obtaining the representation features of the information to be retained from the reference image. Based on this, it can be seen that in a possible implementation, the above information extraction module may include a graphic-text mapping module and the text feature extraction module. Among them, the graphic-text mapping module is used to map the reference image into a text representation vector, so that the text representation vector can represent what the information to be retained in the reference image is, so that the text feature extraction module can subsequently encode the text representation vector to obtain the text expression features of the text representation vector, and regard the text expression features as the representation features of the information to be retained. In this way, it is possible to obtain the representation features of the information to be retained in the reference image with the help of the text feature extraction module already in the basic model.
[0092] It can be seen that in one possible implementation, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the working principle of the model to be processed can specifically include: first mapping the reference image to a text representation vector, so that the text representation vector can represent what the specific information to be retained in the reference image is; then encoding the text representation vector to obtain the representation features of the information to be retained in the reference image; then, denoising is performed based on the representation features to obtain a generated image, so that the generated image and the reference image are consistent in the information to be retained.
[0093] Example 2: In some application scenarios, to further improve the extraction of features representing the information to be retained in the reference image, a new module with image feature extraction functionality can be added to the base model so that this module can subsequently be used to complete the step of obtaining features representing the information to be retained from the reference image. Based on this, it can be seen that the above information extraction module can be implemented using an image feature extraction module, so that the information extraction module is used to extract image features from the reference image and use the extracted image features as the features representing the information to be retained in the reference image.
[0094] It can be seen that in one possible implementation, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the working principle of the model to be processed can specifically include: first performing image feature extraction processing on the reference image to obtain the representation features of the information to be retained in the reference image; then, performing denoising processing based on the representation features to obtain a generated image, so that the generated image is consistent with the reference image in the information to be retained.
[0095] Based on the above content, it can be seen that for the model to be processed, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the model to be processed may include an information extraction module and a denoising network. In which, the denoising network is used to perform denoising processing based on the output data of the information extraction module; and the present application does not limit the implementation method of the denoising network. For example, in some application scenarios, each attention layer in the denoising network (for example, cross-attention layer, self-attention layer, etc.) introduces the representation features of the information to be retained in the reference image, so that the attention layer can achieve the purpose of information retention by interacting the representation features with the latent features of the generated image; and the denoising process corresponding to each time step involved in the denoising network introduces the representation features of the information to be retained in the reference image, so that each round of denoising process in the denoising network needs to be denoised based on the representation features of the information to be retained in the reference image, so that the denoising network exhibits stronger information retention performance, thereby making the model to be processed including the denoising network have stronger information retention performance, which causes the model to be processed to exhibit poor performance in some other dimensions (for example, image editing dimension, image generation quality dimension, etc.).
[0096] Based on the relevant content of S301 above, it can be seen that in some application scenarios, if you want to build an image generation model that can take into account multi-dimensional performance (for example, information retention performance, image editing performance, image generation quality performance, etc.), you can first obtain a model to be processed with certain image generation capabilities, so that you can subsequently construct this image generation model that can take into account multi-dimensional performance by making some adjustments to the data processing flow corresponding to some network layers and / or some time steps in the model to be processed.
[0097] S302: Determine a target object from at least one candidate object based on the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined based on a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object.
[0098] The influence representation data of the i-th candidate object is used to represent the influence of the i-th candidate object in at least one dimension. i is a positive integer, i≤I, I is a positive integer, and I represents the number of objects in the at least one candidate object.
[0099] In addition, the present application does not limit at least one dimension in the above paragraph. For example, the at least one dimension may include an information retention dimension and a target dimension, and there is a preset opposition relationship between the target dimension and the information retention dimension. The information retention dimension is used to represent the performance of maintaining the information to be maintained in the reference image that is of concern in the image generation task. The target dimension is used to represent one or more performance aspects (for example, image editing, aesthetics, etc.) other than the performance of maintaining the information to be maintained in the reference image that is of concern in the image generation task; and the same object will present an opposing influence relationship between the target dimension and the information retention dimension.
[0100] In addition, this application does not limit the target dimension in the above paragraph. For example, the target dimension may include an image editing dimension and an image generation quality dimension. Among them, the image editing dimension is used to represent the performance of image editing based on prompt text (for example, editing a real image into a two-dimensional cartoon image, adding freckles, beards or wrinkles to a facial image, etc.) that is of concern in the image generation task. The image generation quality dimension is used to represent the performance of one or more quality states of the generated image (for example, aesthetics, accurate color matching, etc.) that is of concern in the image generation task.
[0101] Based on the above two paragraphs, it can be known that for the i-th candidate object above, the impact characterization data of the i-th candidate object can include the information retention dimension impact characterization data of the i-th candidate object and the target dimension impact characterization data of the i-th candidate object. Among them, the information retention dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the information retention dimension. The target dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the target dimension; and the present application does not limit the implementation method of the target dimension impact characterization data of the i-th candidate object. For example, if the i-th candidate object is a network layer, the target dimension impact characterization data of the i-th candidate object can be used to represent the impact presented by the i-th candidate object under the image editing dimension; if the i-th candidate object is a set of network layers corresponding to a time step, the target dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image editing dimension and / or the image generation quality dimension. It can be seen that in one possible implementation, when the i-th candidate object is a set of network layers corresponding to a time step, the target dimension impact characterization data of the i-th candidate object may include the image editing dimension impact characterization data of the i-th candidate object and / or the image generation quality dimension impact characterization data of the i-th candidate object. Among them, the image editing dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image editing dimension. The image generation quality dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image generation quality dimension (for example, dimensions such as aesthetics, color matching, and light distribution).
[0102] In addition, for the above-mentioned i-th candidate object, the influence representation data of the i-th candidate object may be determined based on the reference image and the prompt text; and the determination process may specifically include the following steps 11 to 14.
[0103] Step 11: Based on the i-th candidate object, adjust the model to be processed to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed.
[0104] Among them, the adjusted model refers to the model obtained by adjusting the i-th candidate object in the to-be-processed model, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the to-be-processed model, so that the to-be-processed model and the adjusted model can be used as a control group in the future to determine the influence characterization data of the i-th candidate object.
[0105] In addition, this application does not limit the implementation method of the above step 11. For ease of understanding, the following description is combined with two cases.
[0106] Case 1: When the above-mentioned model to be processed is a pre-built basic model, the above-mentioned step 11 can specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the i-th candidate object.
[0107] It can be seen that for the above-mentioned model to be processed, if the model to be processed refers to a pre-built basic model, then the model to be processed may refer to a model that does not have the ability to retain information. Therefore, when the model to be processed is used to perform an impact assessment process on the i-th candidate object, the i-th candidate object in the model to be processed can be fine-tuned according to the fine-tuning learning goal of information retention to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the i-th candidate object, thereby making the i-th candidate object in the adjusted model have the ability to retain information, and thus making the adjusted model also have the ability to retain information. It should be noted that this application does not limit the implementation method of the fine-tuning process in this paragraph. For example, it can be implemented with the help of any fine-tuning method (for example, the lora fine-tuning method or the method of directly fine-tuning the network layer parameters, etc.).
[0108] Based on the content of the above paragraph, it can be seen that in one possible implementation mode, when the above-mentioned model to be processed is a pre-built basic model, the above-mentioned step 11 can specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of other objects in the model to be processed, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the i-th candidate object, so that the adjusted model and the model to be processed can be used as a control group to determine the influence representation data of the i-th candidate object.
[0109] It should be noted that the present application does not limit the implementation method of the step of "adjusting the i-th candidate object in the model to be processed" in the above paragraph. For example, if the i-th candidate object is a network layer, the step may specifically include: adding the input weight of the representation feature of the information to be retained in the above reference image to the i-th candidate object, and adjusting other existing parameters in the i-th candidate object. For another example, if the i-th candidate object is a network layer set corresponding to a time step, the step may specifically include: adding the input weight of the representation feature of the information to be retained in the above reference image to all network layers in the network layer set corresponding to the time step, and adjusting other existing parameters in these network layers, so as to achieve the purpose of introducing the representation feature of the information to be retained in the reference image to the network layer set corresponding to the time step.
[0110] Case 2: When the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used to retain learning information, the above-mentioned step 11 may specifically be: adjusting the i-th candidate object in the model to be processed to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is consistent with the state of the i-th candidate object in the basic model.
[0111] It can be seen that for the above-mentioned model to be processed, if the model to be processed refers to a model obtained by fine-tuning the pre-built basic model, and the fine-tuning is used to learn information retention, then the model to be processed may refer to a model with information retention capability. Therefore, when the model to be processed is used to perform an impact assessment process on the i-th candidate object, the i-th candidate object in the model to be processed may be reverted to the state of the i-th candidate object in the basic model to obtain an adjusted model, so that the state of the i-th candidate object in the adjusted model is consistent with the state of the i-th candidate object in the basic model, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of the other objects in the model to be processed, so that the adjusted model and the model to be processed can be used as a control group to determine the impact characterization data of the i-th candidate object.
[0112] It should be noted that the present application does not limit the implementation method of the rollback. For example, the rollback may specifically refer to completely rolling back the parameter information of the i-th candidate object to the state of the i-th candidate object in the basic model, so that the parameter information of the i-th candidate object after the rollback is consistent with the parameter information of the i-th candidate object in the basic model. Among them, the parameter information of the i-th candidate object refers to the parameters involved in the data processing process corresponding to the i-th candidate object; and the present application does not limit the parameter information of the i-th candidate object. For example, when the i-th candidate object is a network layer, the parameter information of the i-th candidate object may refer to all parameters involved in the network layer. When the i-th candidate object is a network layer set corresponding to a time step, the parameter information of the i-th candidate object may refer to all parameters involved in all network layers in the network layer set corresponding to the time step. It can be seen that if the i-th candidate object above is a network layer, the fallback may specifically include deleting the introduction weight of the i-th candidate object for the representation features of the information to be retained in the reference image, and adjusting all parameters involved in the i-th candidate object except the introduction weight to the parameters of the i-th candidate object in the basic model, etc.; if the i-th candidate object is a set of network layers corresponding to a time step, the fallback may specifically include: for any network layer in the i-th candidate object, deleting the introduction weight of the network layer for the representation features of the information to be retained in the reference image, and adjusting all parameters involved in the network layer except the introduction weight to the parameters of the network layer in the basic model, etc.
[0113] Based on the relevant content of step 11 above, it can be known that for a model to be processed that includes at least one candidate object, the i-th candidate object in the model to be processed can be adjusted to obtain an adjusted model corresponding to the i-th candidate object, so that the state of the i-th candidate object in the adjusted model is different from the state of the i-th candidate object in the model to be processed, and the states of other objects except the i-th candidate object in the adjusted model are consistent with the states of the other objects in the model to be processed, so that the adjusted model and the model to be processed can be used as a control group to determine the influence characterization data of the i-th candidate object. Wherein, i is a positive integer, i≤I, I is a positive integer, and I represents the number of objects in the at least one candidate object above.
[0114] Step 12: Obtain the first generated image and the second generated image corresponding to the i-th candidate object, wherein the first generated image is determined by the image generation processing performed by the model to be processed on the prompt text, and the second generated image is determined by the image generation processing performed by the adjusted model based on the reference image and the prompt text.
[0115] The first generated image corresponding to the i-th candidate object refers to the image generated by the above-mentioned model to be processed for the prompt text.
[0116] In addition, this application does not limit the process of determining the first generated image corresponding to the i-th candidate object described above. For example, when the model to be processed is a pre-built basic model, the first generated image can be determined by the model to be processed performing image generation processing based on the prompt text. It can be seen that in one possible implementation, when the model to be processed is a pre-built basic model, the process of determining the first generated image corresponding to the i-th candidate object is: inputting the prompt text into the model to be processed, so that the model to be processed performs image generation processing based on the prompt text, and obtains and outputs the first generated image.
[0117] For another example, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used for learning information retention, the first generated image corresponding to the above-mentioned i-th candidate object can be determined by the model to be processed performing image generation processing based on the reference image and the prompt text. It can be seen that under one possible embodiment, when the model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning process is used for learning information retention, the process of determining the first generated image corresponding to the i-th candidate object is: inputting the prompt text and the reference image into the model to be processed, so that the model to be processed performs image generation processing based on the prompt text and the reference image, and obtains and outputs the first generated image.
[0118] The second generated image corresponding to the i-th candidate object refers to the image generated by the adjusted model corresponding to the i-th candidate object based on the reference image and the prompt text. Moreover, the present application does not limit the determination process of the second generated image. For example, the determination process of the second generated image is: the reference image and the prompt text are input into the adjusted model corresponding to the i-th candidate object, so that the adjusted model corresponding to the i-th candidate object performs image generation processing based on the reference image and the prompt text to obtain and output the second generated image.
[0119] Based on the above description of the first generated image and the second generated image corresponding to the i-th candidate object, it can be seen that there are certain differences between the generation process of the first generated image and the generation process of the second generated image, and the differences can be specifically as follows: if the above-mentioned model to be processed is a pre-constructed basic model, then when the model to be processed is used to generate the first generated image, it is not necessary to introduce the representation features of the information to be preserved in the reference image into the data processing process corresponding to the i-th candidate object; however, when the adjusted model corresponding to the i-th candidate object is used to generate the second generated image, it is necessary to introduce the representation features of the information to be preserved in the reference image into the data processing process corresponding to the i-th candidate object. However, if the model to be processed is a model obtained by fine-tuning the pre-constructed basic model, and the fine-tuning is used to learn information preservation, then when the model to be processed is used to generate the first generated image, it is necessary to introduce the representation features of the information to be preserved in the reference image into the data processing process corresponding to the candidate object; however, when the adjusted model corresponding to the i-th candidate object is used to generate the second generated image, it is not necessary to introduce the representation features of the information to be preserved in the reference image into the data processing process corresponding to the i-th candidate object.
[0120] Based on the content of the previous paragraph, it can be seen that for the first generated image and the second generated image corresponding to the i-th candidate object above, the data processing process corresponding to the i-th candidate object used when generating the first generated image is different from the data processing process corresponding to the i-th candidate object used when generating the second generated image, and the difference is: one data processing process needs to introduce the representation features of the information to be retained in the reference image, but the other data processing process does not need to introduce the representation features of the information to be retained in the reference image.
[0121] Based on the relevant content of step 12 above, it can be known that for the i-th candidate object above, the first generated image corresponding to the i-th candidate object can be determined using the above-mentioned model to be processed, and the second generated image corresponding to the i-th candidate object can be determined using the adjusted model corresponding to the i-th candidate object, so that the influence characterization data of the i-th candidate object can be determined subsequently based on the difference between the first generated image and the second generated image.
[0122] Step 13: For the first generated image and the second generated image corresponding to the i-th candidate object, determine the information retention dimension impact representation data of the i-th candidate object based on the gap representation data between the information retention dimension performance representation data of the first generated image and the information retention dimension performance representation data of the second generated image; the information retention dimension performance representation data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance representation data of the second generated image is determined based on the second generated image and the reference image.
[0123] In the present application, for the i-th candidate object above, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image can be calculated. The information retention dimension performance characterization data of the first generated image is used to describe the state of the first generated image in the information retention dimension; and the information retention dimension performance characterization data of the first generated image can be determined based on the first generated image and the reference image, so that the information retention dimension performance characterization data of the first generated image can represent the state of the first generated image in terms of retaining the information to be used in the reference image. The information retention dimension performance characterization data of the second generated image is used to describe the state of the second generated image in the information retention dimension; and the information retention dimension performance characterization data of the second generated image can be determined based on the second generated image and the reference image, so that the information retention dimension performance characterization data of the second generated image can represent the state of the second generated image in terms of retaining the information to be used in the reference image.
[0124] It should be noted that the present application does not limit the method for determining the information retention dimension performance characterization data of the first generated image mentioned above. For example, it can be determined by using a feature similarity method. It can be seen that under a possible implementation method, the process of determining the information retention dimension performance characterization data of the first generated image can be specifically as follows: performing feature extraction processing on the first generated image to obtain the image features of the first generated image, and performing feature extraction processing on the reference image to obtain the image features of the reference image; then, the similarity between the image features of the first generated image and the image features of the reference image is determined as the information retention dimension performance characterization data of the first generated image. For example, in some application scenarios, when the information to be used above is facial description information, the information retention dimension performance characterization data of the first generated image can be judged with the help of a facial recognition model, so that the information retention dimension performance characterization data of the first generated image can be determined subsequently by judging whether the judgment result output by the facial recognition model for the first generated image is the same as the judgment result output by the facial recognition model for the reference image.
[0125] It should also be noted that the method for determining the information retention dimension performance characterization data of the second generated image above is similar to the "method for determining the information retention dimension performance characterization data of the first generated image" shown in the previous paragraph. For the sake of brevity, it will not be repeated here.
[0126] In addition, for the first generated image and the second generated image corresponding to the i-th candidate object above, after obtaining the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image, the gap characterization data between the information retention dimension performance characterization data of the first generated image and the information retention dimension performance characterization data of the second generated image can be calculated first, so that the gap characterization data can represent the model performance difference presented by the above-mentioned model to be processed and the adjusted model corresponding to the i-th candidate object in the information retention dimension, so that the gap characterization data can represent the influence presented by the i-th candidate object in the information retention dimension; then, based on the gap characterization data, the information retention dimension influence characterization data of the i-th candidate object is determined (for example, the gap characterization data is directly determined as the information retention dimension influence characterization data of the i-th candidate object, etc.), so that the information retention dimension influence characterization data can represent the influence presented by the i-th candidate object in the information retention dimension.
[0127] Based on the relevant content of step 13 above, it can be known that for the i-th candidate object above, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the information retention dimension impact representation data of the i-th candidate object can be calculated based on the first generated image, the second generated image and the reference image, so that the information retention dimension impact representation data can represent the impact of the i-th candidate object on the information retention dimension.
[0128] Step 14: For the first generated image and the second generated image corresponding to the i-th candidate object, determine the target dimension impact representation data of the i-th candidate object based on the gap representation data between the target dimension performance representation data of the first generated image and the target dimension performance representation data of the second generated image; the target dimension performance representation data of the first generated image is determined based on the first generated image; the target dimension performance representation data of the second generated image is determined based on the second generated image.
[0129] In the present application, for the i-th candidate object above, after obtaining the first generated image and the second generated image corresponding to the i-th candidate object, the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image can be calculated. The target dimension performance characterization data of the first generated image is used to describe the state of the first generated image in the target dimension; and the target dimension performance characterization data of the first generated image can be determined based on the first generated image so that the target dimension performance characterization data of the first generated image can represent the state of the first generated image in certain aspects (such as image editing, image quality, etc.) other than the aspect of maintaining the information to be used in the reference image. The target dimension performance characterization data of the second generated image is used to describe the state of the second generated image in the target dimension; and the target dimension performance characterization data of the second generated image can be determined based on the second generated image so that the target dimension performance characterization data of the second generated image can represent the state of the second generated image in certain aspects (such as image editing, image quality, etc.) other than the aspect of maintaining the information to be used in the reference image.
[0130] In addition, the present application does not limit the determination process of the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image. For example, when the i-th candidate object above is a network layer, the target dimension performance characterization data of the first generated image corresponding to the i-th candidate object can be determined based on the similarity between the first generated image and the prompt text (for example, the similarity between the first generated image and the prompt text can be directly determined as the target dimension performance characterization data of the first generated image), so that the target dimension performance characterization data can represent the performance of the first generated image in image editing; and the target dimension performance characterization data of the second generated image corresponding to the i-th candidate object can be determined based on the similarity between the second generated image and the prompt text (for example, the similarity between the second generated image and the prompt text can be directly determined as the target dimension performance characterization data of the second generated image), so that the target dimension performance characterization data can represent the performance of the second generated image in image editing.
[0131] It should be noted that this application does not limit the calculation method of the "similarity between the first generated image and the prompt text" in the previous paragraph. For example, it can be implemented by using any existing or future method that can calculate the similarity between an image and a text, such as a method for calculating the similarity between an image and a text with the help of a Contrastive Language-Image Pre-Training (CLIP) model.
[0132] For another example, when the i-th candidate object above is a set of network layers corresponding to a time step, the target dimension impact representation data of the i-th candidate object may include the image editing dimension impact representation data of the i-th candidate object and / or the image generation quality dimension impact representation data of the i-th candidate object, so that the target dimension impact representation data of the i-th candidate object can represent the impact of the i-th candidate object on the image editing dimension and / or the image generation quality dimension.
[0133] The image editing dimension impact characterization data of the i-th candidate object is used to represent the impact presented by the i-th candidate object under the image editing dimension; and the image editing dimension impact characterization data of the i-th candidate object is determined based on the difference characterization data between the image editing dimension performance characterization data of the first generated image corresponding to the i-th candidate object and the image editing dimension performance characterization data of the second generated image corresponding to the i-th candidate object. The image editing dimension performance characterization data of the first generated image is used to represent the performance presented by the first generated image in terms of image editing; and the image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text (for example, the similarity between the first generated image and the prompt text can be directly determined as the image editing dimension performance characterization data of the first generated image). The image editing dimension performance characterization data of the second generated image is used to represent the performance presented by the second generated image in terms of image editing; and the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text (for example, the similarity between the second generated image and the prompt text can be directly determined as the image editing dimension performance characterization data of the second generated image).
[0134] The image generation quality dimension impact representation data of the i-th candidate object is used to represent the impact of the i-th candidate object in an image generation quality dimension (e.g., quality dimensions such as aesthetics and color matching). The image generation quality dimension impact representation data of the i-th candidate object is determined based on the difference representation data between the image generation quality dimension performance representation data of a first generated image corresponding to the i-th candidate object and the image generation quality dimension performance representation data of a second generated image corresponding to the i-th candidate object. The image generation quality dimension performance representation data of the first generated image is used to represent the performance of the first generated image in terms of image generation quality. The image generation quality dimension performance representation data of the first generated image is determined by performing a quality assessment process on the first generated image in terms of at least one quality dimension (e.g., aesthetics and color matching). The image generation quality dimension performance representation data of the second generated image is used to represent the performance of the second generated image in terms of image generation quality. The image generation quality dimension performance representation data of the second generated image is determined by performing a quality assessment process on the second generated image in terms of at least one quality assessment dimension. Among them, the at least one quality assessment dimension refers to the dimension required to be based on when performing quality assessment processing on an image; and this application does not limit the at least one quality assessment dimension. For example, it may include dimensions such as aesthetics, reasonable color matching, and reasonable light distribution.
[0135] It should be noted that this application does not limit the implementation method of the quality assessment processing in the previous paragraph. For example, it can be implemented with the help of any existing or future method for evaluating the quality status of an image (for example, with the help of a pre-built machine learning model with image quality assessment performance, etc.).
[0136] Based on the relevant content of steps 11 to 14 above, it can be known that for the i-th candidate object above, an adjustment method can be performed on the i-th candidate object in the model to be processed to obtain an adjusted model that has a control relationship with the model to be processed under the i-th candidate object, so that the impact characterization data of the i-th candidate object can be determined by comparing the generated images of the two models in some dimensions (for example, information retention dimension, image editing dimension, image generation quality dimension, etc.), so that the impact characterization data can characterize the impact presented by the i-th candidate object in at least one dimension (for example, information retention dimension, image editing dimension, image generation quality dimension, etc.), so that it can be determined based on the impact characterization data whether it is necessary to introduce the characterization features of the information to be retained in the reference image into the data processing process corresponding to the i-th candidate object.
[0137] The target object refers to an object that exists in at least one of the candidate objects mentioned above and meets the preset retention information introduction conditions, so that the target object is used to represent the object whose representation features of the information to be retained in the reference image need to be introduced into its corresponding data processing process. The preset retention information introduction conditions can be based on the actual application scenario in advance; and the present application does not limit the preset retention information introduction conditions. For example, the preset retention information introduction conditions can specifically be: the information retention dimension of the target object affects the representation data higher than the target dimension of the target object (for example, the information retention dimension of the target object affects the representation data much higher than the target dimension of the target object, etc.).
[0138] Based on the relevant content of S302 above, it can be known that for at least one candidate object in the above model to be processed, the influence representation data of each candidate object is first obtained (for example, information retention dimension influence representation data and target dimension influence representation data, etc.); then, based on the influence representation data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence representation data of the target object is higher than the target dimension influence representation data of the target object, so that the target object can represent the object existing in the at least one candidate object, which has an influence in the information retention dimension far higher than the influence in other dimensions. In this way, when the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, it is possible to achieve as little reduction in performance in other dimensions as possible under the premise of greatly improving the information retention performance, thereby effectively avoiding the adverse effects caused by a significant reduction in performance in other dimensions due to improving the information retention performance.
[0139] S303: Based on the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0140] In the present application, after determining the target object, the model to be processed can be adjusted according to the target object to obtain an image generation model, so that the image generation model includes the target object, and when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object. The data processing process corresponding to the target object refers to the data processing flow pre-configured for the target object; and the present application does not limit the data processing process corresponding to the target object. For example, when the target object is a network layer, the data processing process corresponding to the target object refers to the data processing process implemented by the network layer. For another example, when the target object is a set of network layers corresponding to a time step, the data processing process corresponding to the target object refers to the data processing process implemented by each network layer in the network layer set corresponding to the time step.
[0141] Based on the content in the previous paragraph, it can be seen that in some application scenarios, for the above-mentioned image generation model, the representation features of the information to be retained in the reference image need to be introduced into the data processing process corresponding to some objects within the image generation model, but the representation features of the information to be retained in the reference image do not need to be introduced into the data processing process corresponding to another part of the objects. Based on this, it can be seen that in one possible implementation, the image generation model has the following characteristics: when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects other than the target object.
[0142] In addition, this application does not limit the implementation method of the above S303. For ease of understanding, the following description is combined with two cases.
[0143] Case 1: When the above-mentioned model to be processed is a pre-built basic model, the above S303 can specifically be: adjusting the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of other objects in the model to be processed, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
[0144] It can be seen that in some application scenarios, when the above-mentioned model to be processed is a pre-built basic model, after determining the target object, the image generation model can be obtained by adjusting the parameters involved in the target object in the model to be processed, so that the target object in the image generation model has the ability to retain information. It should be noted that this application does not limit the adjustment method. For example, if the target object is a network layer, the adjustment method may include adding the input weight of the characterization feature of the information to be retained to the network layer, and adjusting other existing parameters in the network layer, so that the network layer is used to perform corresponding data processing based on the characterization feature of the information to be retained, so as to achieve the introduction of the characterization feature of the information to be retained into the network layer; if the target object is a set of network layers corresponding to a time step, the adjustment method includes adding the input weight of the characterization feature of the information to be retained to all network layers in the network layer set corresponding to the time step, and adjusting other existing parameters in these network layers, so that some or all network layers involved in the denoising process need to perform corresponding data processing based on the characterization feature of the information to be retained, so as to achieve the introduction of the characterization feature of the information to be retained into the network layer set corresponding to the time step.
[0145] Case 2: When the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used for learning information retention, the above S303 can specifically be: adjusting the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representational features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, and there is no need to introduce the representational features of the information to be retained in the reference image into the data processing process corresponding to the objects other than the target object.
[0146] It can be seen that in some application scenarios, when the above-mentioned model to be processed refers to a model obtained by fine-tuning a pre-built basic model, and the fine-tuning is used to learn information retention, after determining the target object, the image generation model can be obtained by adjusting the parameters of other objects in the model to be processed except the target object, so that the target object in the image generation model retains the ability to retain information, but other objects do not retain the ability to retain information. It should be noted that the present application does not limit the adjustment method. For example, if the other object is a network layer, the adjustment method may include deleting the input weight of the representation feature of the information to be retained from the network layer, and returning other existing parameters in the network layer to the state of the corresponding parameters in the basic model; if the other object is a network layer set corresponding to a time step, the adjustment method includes returning the data processing process described by the network layer set corresponding to the time step to the data processing process described by the network layer set corresponding to the corresponding time step in the basic model (for example, deleting the input weight of the representation feature of the information to be retained from all network layers in the network layer set corresponding to the time step, and returning other existing parameters in these network layers to the state of the corresponding parameters in the basic model, etc.), so that all network layers in the network layer set corresponding to the time step do not need to perform corresponding data processing based on the representation feature of the information to be retained, thereby achieving the goal of not introducing the representation feature of the information to be retained into the network layer set corresponding to the time step.
[0147] Based on the relevant contents of S301 to S303 above, it can be known that for the model determination method provided in the embodiment of the present application, the model to be processed (for example, a diffusion model, etc.) is first obtained, and the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer and / or at least one time step; then, based on the information retention dimension influence characterization data and the target dimension influence characterization data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, the model to be processed is adjusted to obtain to the image generation model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, thereby enabling the image generation model to achieve performance balance as much as possible in the two opposing dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.). In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, which is conducive to improving the image generation effect.
[0148] In fact, in order to further improve the image generation effect, the present application also provides an implementation of the above model determination method. Under this implementation, in addition to including the above S301-S303, the model determination method may also include the following step 21.
[0149] Step 21: When the target object above is a network layer, the output weight corresponding to the target object in the image generation model is adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0150] The output weight corresponding to the target object refers to the weighting weight that needs to be used for the output data of the target object when inputting the output data of the target object into the next network layer, so that the input data of the next network layer includes the result obtained by weighted processing of the output data according to the weighting weight. The output data of the target object refers to the result output by the target object.
[0151] In addition, for the image generation model adjusted above, the performance improvement value presented by the adjusted image generation model in the target dimension is determined based on the difference between the target dimension impact characterization data of the adjusted image generation model and the target dimension impact characterization data of the image generation model before adjustment, so that the performance improvement value can represent the degree of performance improvement presented in the target dimension after adjusting the output weights corresponding to the target object in the image generation model before adjustment.
[0152] In addition, for the image generation model adjusted above, the performance degradation value presented by the adjusted image generation model in the information retention dimension is determined based on the difference between the information retention dimension impact characterization data of the adjusted image generation model and the information retention dimension impact characterization data of the image generation model before adjustment, so that the performance degradation value can represent the degree of performance degradation presented in the information retention dimension after adjusting the output weights corresponding to the target object in the image generation model before adjustment.
[0153] Based on the relevant content of step 21 above, it can be seen that for the target object above, when the target object is a network layer, not only the parameters involved in the data processing process corresponding to the target object will affect the model performance of the image generation model, but also the output weights corresponding to the target object will affect the model performance of the image generation model. Therefore, in order to further balance the performance of the multi-dimensional model, the output weights corresponding to the target object in the image generation model can be fine-tuned so that the performance improvement value of the fine-tuned image generation model in the target dimension is much higher than the performance degradation value of the fine-tuned image generation model in the information retention dimension, thereby further improving the image generation effect.
[0154] In order to better understand the model determination method provided in this application, two situations are described below.
[0155] Case 1: In some application scenarios, an image generation model that can take into account multi-dimensional performance can be constructed by adding some information retention-related parameters to an already constructed diffusion model that does not have information retention capabilities (for example, a diffusion model that can generate images based on prompt text, etc.).
[0156] Based on the above situation 1, it can be seen that the model determination method provided in this application may include the following steps 31 to 37.
[0157] Step 31: Use some sample texts to train the diffusion model to obtain a basic model, so that the basic model has a text image generation function, thereby enabling the basic model to perform image generation processing based on any prompt text.
[0158] The sample text refers to the training data required for the diffusion model to learn the image and text generation capability; and this application is not limited to the sample text.
[0159] Step 32: Determine the information retention dimension of some or all network layers in the above basic model that affects the characterization data.
[0160] In this application, for the above basic model, one or more network layers in the basic model can be fine-tuned to obtain information retention dimension impact characterization data of the one or more network layers, so that the information retention dimension impact characterization data can represent the impact of the one or more network layers under the information retention dimension.
[0161] It should be noted that the present application does not limit the determination process of "the information retention dimension of one or more network layers affects the representation data" in the above paragraph. For example, it can be specifically as follows: first, fine-tune one or more network layers in the basic model (for example, add the input weights of the representation features of the information to be retained to the one or more network layers, and adjust other existing parameters in the one or more network layers, etc.) to obtain an adjusted model so that the one or more network layers in the adjusted model have information retention capabilities; then, based on the generated image of the basic model, the generated image of the adjusted model, and the reference image, determine the information retention dimension of the one or more network layers that affects the representation data. Among them, the generated image of the basic model is obtained by the basic model performing image generation processing based on the prompt text; the generated image of the adjusted model is obtained by the adjusted model performing image generation processing based on the prompt text and the reference image.
[0162] It should also be noted that in some application scenarios (for example, face preservation scenarios), the information retention dimension impact characterization data of some or all network layers in the above basic model can show that the network layers that can strongly affect the information retention effect are concentrated in the upsampling module in the denoising network of the diffusion model. Therefore, some attention layers (for example, cross-attention layers and self-attention layers) in the upsampling module can be fine-tuned (for example, lora fine-tuning or direct fine-tuning of network layers, etc.) to obtain the fine-tuned model.
[0163] It should also be noted that for the fine-tuning portion (e.g., one or more network layers) in the above basic model, the present application can adjust the influence of the fine-tuning portion in the information preservation dimension by adjusting the weight of the output value of the parameter component of the fine-tuning portion. For models that implement the processing of introducing the information to be used by directly adding a new attention layer, the influence of the fine-tuning portion in the information preservation dimension can be adjusted by directly adjusting the weight of the output value of the newly added attention layer.
[0164] Step 33: Determine how the image editing dimensions of some or all of the network layers in the above base model affect the representation data.
[0165] In the present application, for the above basic model, one or more network layers in the basic model can be fine-tuned to obtain image editing dimension impact characterization data of the one or more network layers, so that the image editing dimension impact characterization data can represent the impact of the one or more network layers under the image editing dimension. Wherein, the image editing dimension refers to the aspect that the model edits and modifies its generated content (for example, adding items, changing styles, modifying part of the content, etc.) according to the prompt text. In addition, there is an oppositional relationship between the performance of the model under the image editing dimension and the performance of the model under the information retention dimension. The specific oppositional relationship is: if the performance of the model under the image editing dimension is strong, the performance of the model under the information retention dimension is poor; if the performance of the model under the information retention dimension is strong, the performance of the model under the image editing dimension is poor.
[0166] It should be noted that the process of determining "the image editing dimension of one or more network layers affects the representation data" in the previous paragraph is similar to the process of determining "the information retention dimension of one or more network layers affects the representation data" above. For the sake of brevity, it will not be repeated here.
[0167] In addition, for step 33 above, in some application scenarios, to further improve the impact determination effect under the image editing dimension, the editing content that is difficult to implement in the above basic model can be used as prompt text, so that during the execution of step 33, this prompt text can be used to determine the image editing dimension impact representation data of some or all network layers in the basic model. The "difficult-to-implement editing content" can be obtained through certain methods (e.g., big data analysis methods, or manually specified methods, etc.).
[0168] It can be seen that, under a possible implementation method, the acquisition process of the above "difficult to implement editing content" can be specifically as follows: after obtaining the constructed basic model, the basic model can be used to perform image generation processing on some prompt texts to obtain the generated images corresponding to these prompt texts; then, for any prompt text, the editing performance corresponding to the prompt text is determined based on the similarity between the generated image corresponding to the prompt text and the prompt text, so that the editing performance can indicate whether the prompt text belongs to the "difficult to implement editing content", so that when it is determined that the editing performance meets the preset difficult editing conditions (for example, conditions lower than the preset performance threshold, etc.), the prompt text is used as the "difficult to implement editing content", so that the prompt text can be used in the subsequent step 33 to determine the image editing dimension impact characterization data of some or all network layers in the basic model.
[0169] Step 34: Based on the information retention dimension impact characterization data of some or all network layers in the above basic model and the image editing dimension impact characterization data, determine the target network layer that needs to introduce the information to be retained from the basic model, and fine-tune the target network layer in the basic model to obtain an image generation model, so that the data processing process corresponding to the target network layer in the image generation model needs to introduce the characterization features of the information to be retained in the reference image, and the data processing processes corresponding to other network layers in the image generation model other than the target network layer do not need to introduce the characterization features of the information to be retained in the reference image, so that the target network layer in the image generation model has the ability to retain information, and the other network layers in the image generation model other than the target network layer do not have the ability to retain information.
[0170] In the present application, for the above basic model, after determining the information retention dimension impact characterization data and the image editing dimension impact characterization data of some or all network layers in the basic model, a network layer that has a greater impact on information retention but a smaller impact (or even no obvious impact) on editing ability can be selected from these network layers as the target network layer, and the target network layer in the basic model is fine-tuned (for example, adding the input weight of the characterization feature of the information to be retained to the target network layer, and adjusting other existing parameters in the target network layer, etc.) to obtain an image generation model, so that the data processing process corresponding to the target network layer in the image generation model needs to introduce the characterization feature of the information to be retained in the reference image, so that the target network layer in the image generation model has information retention capability.
[0171] Step 35: Adjust the output weights of the target network layer in the above image generation model so that the performance improvement value presented by the adjusted image generation model in the image editing dimension is higher than the performance degradation value presented by the adjusted image generation model in the information preservation dimension.
[0172] It should be noted that this application does not limit the implementation method of the above step 35. For ease of understanding, it is explained below with reference to three examples.
[0173] Example 1: When the representational features of the information to be used in the above reference image are determined based on the existing text feature extraction module, and the above steps 31 to 34 are implemented with the help of the Lora fine-tuning strategy, if the above target network layer includes the cross-attention layer in the upsampling module, then the above step 35 can be specifically as follows: adjust the output weights of all fine-tuning parameters of all cross-attention layers in the upsampling module (for example, all fine-tuning parameters involved in the Lora fine-tuning strategy), and adjust the output weight from 1 to a smaller or even negative coefficient, which can significantly improve the editing ability while only slightly weakening the information retention ability.
[0174] Example 2: When the representational features of the information to be used in the above reference image are determined based on the newly added image feature extraction module, if the above target network layer includes the cross-attention layer in the upsampling module, then the above step 35 can be specifically as follows: adjusting the output weight of the first cross-attention layer in the upsampling module, and adjusting the output weight from 1 to a smaller or even negative coefficient, which can significantly improve the editing ability while only slightly weakening the information retention ability.
[0175] Example 3: If the target network layer above includes all network layers of the denoising network in the image generation model, then the above step 35 can be specifically as follows: the output weights of all fine-tuning parameter biases are reduced to half of the original (a convenient implementation method is to average the fine-tuning model parameters with the model parameters before fine-tuning). In particular, the fine-tuning parameter components of all attention layer modules of the upsampling module can be further reduced to 0 or 0.1 to significantly improve the editing ability.
[0176] Step 36: Determine the characterization data affected by the information preservation dimension of some or all time steps in the denoising network in the adjusted image generation model, the characterization data affected by the image editing dimension, and the characterization data affected by the image generation quality dimension.
[0177] In the present application, for the image generation model adjusted above, the information to be retained can be introduced and adjusted for one or more time steps in the image generation model to obtain the information retention dimension impact characterization data, image editing dimension impact characterization data and image generation quality dimension impact characterization data of the one or more time steps, so that it can be decided based on these data whether to introduce the characterization features of the information to be retained in the reference image in the denoising process corresponding to the time step.
[0178] It should be noted that this application does not limit the implementation method of the above step 36.
[0179] Step 37: Based on the information retention dimension impact characterization data of some or all time steps in the denoising network in the adjusted image generation model above, the image editing dimension impact characterization data and the image generation quality dimension impact characterization data, determine the target time step where the information to be retained needs to be introduced from the image generation model, and revert the denoising process corresponding to other time steps in the image generation model except the target time step to the denoising process corresponding to the corresponding time step in the basic model above, so that when the image generation model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to the target time step, but it is not necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to other time steps except the target time step, so that the performance improvement value of the image generation model after regressing in the image editing dimension and the image generation quality dimension is higher than the performance degradation value of the image generation model after regressing in the information retention dimension.
[0180] It should be noted that for the diffusion model, since the generation process of the diffusion model is multi-step iterative, so that the early time steps (that is, the time steps that are executed relatively early in the denoising process) are mainly related to the generation of composition, light perception, etc., and the later time steps (that is, the time steps that are executed relatively late in the denoising process) are related to details, so in this application, it is only necessary to introduce the representation features of the information to be retained in the reference image in the denoising process corresponding to the intermediate time steps. Based on this, it can be seen that there is no need to introduce the representation features of the information to be retained in the reference image in the denoising process corresponding to the early time steps and the later time steps, which can better reduce the impact of the image generation model on the light perception aesthetics and editing ability of the above basic model. In addition, the intermediate time steps can be adjusted according to the actual effect. For example, in some application scenarios, it can be after iteratively executing the denoising process corresponding to 30% to 50% of the time steps, to start introducing the representation features of the information to be retained in the reference image in the denoising process corresponding to each time step.
[0181] Based on the relevant contents of steps 31 to 37 above, it can be known that in some application scenarios, a basic model with image editing capabilities can be trained first; then, based on the information retention dimension influence representation data and the image editing dimension influence representation data of some network layers in the basic model, it is determined which network layers in the basic model need to be adjusted, so that these adjusted network layers perform corresponding data processing based on the representation features of the information to be retained in the reference image, so that the image generation model obtained by adjusting these network layers in the basic model can achieve a significant improvement in information retention performance with the least possible reduction in image editing performance; then, based on the information retention dimension influence representation data, the image editing dimension influence representation data and the image generation quality dimension influence representation data of some time steps in the denoising network of the image generation model, the image generation model is determined. Which time steps need to introduce the representation features of the information to be kept, and which time steps do not need to introduce the representation features of the information to be kept, so that the image generation model can be adjusted according to these conclusions in the future, so that some time steps in the adjusted image generation model need to introduce the representation features of the information to be kept, and another part of the time steps do not need to introduce the representation features of the information to be kept, so that the adjusted image generation model can better balance the information retention performance and other performance (such as image editing performance, image generation quality performance, etc.); finally, fine-tune certain weight parameters in the adjusted image generation model, so that the fine-tuned image generation model can better balance the information retention performance and other performance (such as image editing performance, image generation quality performance, etc.), so that the image generation model finally obtained has a better image generation effect.
[0182] Case 2: In some application scenarios, an image generation model that can take into account multi-dimensional performance can be constructed by deleting some information-preserving related parameters from an already constructed diffusion model with information preservation capabilities (for example, the diffusion models shown in Figures 1 and 2).
[0183] Based on the above situation 2, it can be seen that the model determination method provided in this application may include the following steps 41 to 49.
[0184] Step 41: Use some sample texts to train the diffusion model to obtain a basic model, so that the basic model has a text image generation function, thereby enabling the basic model to perform image generation processing based on any prompt text.
[0185] It should be noted that the relevant content of step 41 is similar to the relevant content of step 31 above, and for the sake of brevity, it will not be repeated here.
[0186] Step 42: Based on the basic model, a first model (eg, the diffusion model shown in FIG. 1 or FIG. 2 ) is constructed so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0187] It should be noted that the present application does not limit the implementation method of the above step 42. For example, in some application scenarios (for example, when the representation features of the information to be retained in the reference image are determined with the help of the original text feature extraction module in the basic model), the step 42 can be specifically: adding a graphic-text mapping module to the basic module, and fine-tuning some or all of the network layers (for example, the attention layer) in the denoising network of the basic module to obtain a first model, so that the first model can at least include the graphic-text mapping module, the text feature extraction module and the denoising network, and the input data of some or all of the network layers (for example, the attention layer) in the denoising network include the output data of the text feature extraction module, so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0188] For example, in some application scenarios (for example, when the representational features of the information to be retained in the reference image are determined with the help of an additional image feature extraction module), step 42 can specifically be: adding an image feature extraction module to the basic module, and fine-tuning some or all of the network layers (for example, the attention layer) in the denoising network of the basic module to obtain a first model, so that the first model can at least include the image feature extraction module and the denoising network, and the input data of some or all of the network layers (for example, the attention layer) in the denoising network include the output data of the image feature extraction module, so that the first model has the function of performing image generation processing based on the reference image and the prompt text.
[0189] Step 43: Using some sample images and the prompt texts corresponding to these sample images, fine-tune the first model to obtain a second model, so that the second model has better information retention capability. The fine-tuning process is used to learn information retention.
[0190] The sample image refers to the training data required to be used when the first model learns the information retention capability; and this application does not limit the sample image.
[0191] In addition, the present application does not limit the implementation method of the fine-tuning process in step 43 above. For example, during the fine-tuning process, the sample image is used as a reference image so that the sample image is used to provide the information to be retained that needs to be retained during the image generation process. In this way, the second model obtained after step 43 can learn how to retain the information to be retained in the sample image, thereby enabling the second model to have a relatively strong information retention capability.
[0192] Step 44: Determine the information retention dimension impact characterization data of some or all network layers in the second model.
[0193] In the present application, for the second model above, fine-tuning processing can be performed on one or more network layers in the second model to obtain information retention dimension impact characterization data of the one or more network layers, so that the information retention dimension impact characterization data can represent the impact of the one or more network layers under the information retention dimension.
[0194] It should be noted that the present application does not limit the determination process of "the information retention dimension of one or more network layers affecting the characterization data" in the above paragraph. For example, it can be specifically as follows: first, fine-tune one or more network layers in the second model (for example, revert one or more network layers in the second model to the state of the corresponding network layer in the above basic model, etc.), and obtain an adjusted model so that the one or more network layers in the adjusted model do not have the ability to retain information; then, based on the generated image of the second model, the generated image of the adjusted model, and the reference image, determine the information retention dimension affecting the characterization data of the one or more network layers. The generated image of the second model is obtained by the second model performing image generation processing based on the prompt text and the reference image; the generated image of the adjusted model is obtained by the adjusted model performing image generation processing based on the prompt text and the reference image.
[0195] It should also be noted that for the fine-tuning portion (e.g., one or more network layers) in the second model above, the present application can adjust the influence of the fine-tuning portion in the information preservation dimension by adjusting the weight of the output value of the parameter component of the fine-tuning portion. For models that implement the processing of introducing information to be used by directly adding a new attention layer, the influence of the fine-tuning portion in the information preservation dimension can be adjusted by directly adjusting the weight of the output value of the newly added attention layer.
[0196] Step 45: Determine the image editing dimension impact characterization data of some or all network layers in the second model above.
[0197] In the present application, for the second model above, one or more network layers in the second model can be fine-tuned to obtain image editing dimension impact characterization data of the one or more network layers, so that the image editing dimension impact characterization data can represent the impact of the one or more network layers under the image editing dimension.
[0198] It should be noted that the process of determining "the image editing dimension of one or more network layers affects the representation data" in the previous paragraph is similar to the process of determining "the information retention dimension of one or more network layers affects the representation data" above. For the sake of brevity, it will not be repeated here.
[0199] In addition, regarding step 45 above, in some application scenarios, to further improve the effect of determining the impact of image editing dimensions, the editing content that is difficult to implement in the second model above can be used as prompt text, so that during the execution of step 33, this prompt text is used to determine the image editing dimension impact representation data of some or all network layers in the second model. The "difficult-to-implement editing content" can be obtained through certain methods (e.g., big data analysis methods, or manually specified methods, etc.).
[0200] It can be seen that, under a possible implementation method, the acquisition process of the above-mentioned "difficult-to-edit content" can be specifically as follows: after obtaining the constructed second model, the second model can be used to perform image generation processing on some prompt texts to obtain the generated images corresponding to these prompt texts; then, for any prompt text, the editing performance corresponding to the prompt text is determined based on the similarity between the generated image corresponding to the prompt text and the prompt text, so that the editing performance can indicate whether the prompt text belongs to the "difficult-to-edit content", so that when it is determined that the editing performance meets the preset difficult editing conditions (for example, conditions lower than the preset performance threshold, etc.), the prompt text is used as the "difficult-to-edit content", so that the prompt text can be used in the subsequent step 45 to determine the image editing dimension impact characterization data of some or all network layers in the second model.
[0201] Step 46: Based on the information retention dimension impact characterization data of some or all network layers in the second model above and the image editing dimension impact characterization data, determine the network layers to be adjusted that do not need to introduce the information to be retained (that is, other network layers except the target network layer that needs to introduce the information to be retained above) from the second model, and roll back the network layers to be adjusted in the second model to the state of the corresponding network layers in the basic model above to obtain the image generation model, so that the data processing process corresponding to the network layers to be adjusted in the image generation model does not need to introduce the characterization features of the information to be retained in the reference image.
[0202] In the present application, for the second model above, after determining the information retention dimension impact characterization data and the image editing dimension impact characterization data of some or all network layers in the second model, a network layer that has a relatively large improvement in editing capability but a relatively small (or even insignificant) decrease in information retention can be selected from these network layers as the network layer to be adjusted, and the network layer to be adjusted in the second model is rolled back to the state of the corresponding network layer in the basic model above to obtain an image generation model, so that the data processing process corresponding to the network layer to be adjusted in the image generation model does not need to introduce the characterization features of the information to be retained in the reference image.
[0203] Step 47: Adjust the output weights of the target network layer in the above image generation model so that the performance improvement of the adjusted image generation model in the image editing dimension is greater than the performance degradation of the adjusted image generation model in the information preservation dimension. The target network layer is the network layer in the image generation model that needs to perform data processing based on the characteristics representing the information to be preserved in the reference image.
[0204] It should be noted that the relevant content of step 47 is similar to the relevant content of step 35 above, and for the sake of brevity, it will not be repeated here.
[0205] Step 48: Determine the characterization data affected by the information preservation dimension of some or all time steps in the denoising network in the adjusted image generation model, the characterization data affected by the image editing dimension, and the characterization data affected by the image generation quality dimension.
[0206] It should be noted that the relevant content of step 48 is similar to the relevant content of step 36 above, and for the sake of brevity, it will not be repeated here.
[0207] Step 49: Based on the information retention dimension impact characterization data of some or all time steps in the denoising network in the adjusted image generation model above, the image editing dimension impact characterization data and the image generation quality dimension impact characterization data, determine the target time step where the information to be retained needs to be introduced from the image generation model, and revert the denoising process corresponding to other time steps in the image generation model except the target time step to the denoising process corresponding to the corresponding time step in the basic model above, so that when the image generation model performs image generation processing based on the reference image and the prompt text, it is necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to the target time step, but it is not necessary to introduce the characterization features of the information to be retained in the reference image into the denoising process corresponding to other time steps except the target time step, so that the performance improvement value of the image generation model after regressing in the image editing dimension and the image generation quality dimension is higher than the performance degradation value of the image generation model after regressing in the information retention dimension.
[0208] It should be noted that the relevant content of step 49 is similar to the relevant content of step 37 above, and for the sake of brevity, it will not be repeated here.
[0209] Based on the relevant contents of steps 41 to 49 above, it can be known that in some application scenarios, a basic model with image editing capabilities can be trained first; then, based on the basic model, a second model with information retention capabilities can be constructed; secondly, based on the information retention dimension impact representation data and the image editing dimension impact representation data of some network layers in the second model, it is determined that those network layers in the second model need to be rolled back to the states of the corresponding network layers in the basic model, so that the image generation model obtained by rolling back these network layers in the second model can achieve a significant improvement in information retention performance with the least possible reduction in image editing performance; then, based on the information retention dimension impact representation data, the image editing dimension impact representation data and the image generation quality dimension impact representation data of some time steps in the denoising network of the image generation model, the image generation model is determined. Which time steps need to introduce the representation features of the information to be retained, and which time steps do not need to introduce the representation features of the information to be retained, so that the image generation model can be adjusted based on these conclusions in the future, so that some time steps in the adjusted image generation model need to introduce the representation features of the information to be retained, and another part of the time steps do not need to introduce the representation features of the information to be retained, so that the adjusted image generation model can better balance the information retention performance and other performance (for example, image editing performance, image generation quality performance, etc.); finally, fine-tune certain weight parameters in the adjusted image generation model, so that the fine-tuned image generation model can better balance the information retention performance and other performance (for example, image editing performance, image generation quality performance, etc.), so that the image generation model finally obtained has a better image generation effect.
[0210] In fact, after determining the above image generation model, the image generation model can be used to perform certain image generation tasks. Based on this, the present application also provides an image generation method, as shown in Figure 4, which may include the following S401-S402. Figure 4 is a flowchart of an image generation method provided in an embodiment of the present application.
[0211] S401: Acquire information description images and image constraint description text to be retained.
[0212] Among them, the image describing the information to be retained is used to limit the information that needs to be retained during the image generation process, so that the image describing the information to be retained can be used as a reference image during the image generation process; and this application does not limit the method of obtaining the image describing the information to be retained. For example, in some application scenarios, the image describing the information to be retained can be an image specified by the user through certain operations (such as image input operations or image selection operations, etc.).
[0213] The image constraint description text is used to limit the constraints required in the image generation process, so that the image constraint description text can be used as prompt text in the image generation process; and this application does not limit the method of obtaining the image constraint description text. For example, in some application scenarios, the image constraint description text can be text content specified by the user through certain operations (such as text input operations or text selection operations, etc.).
[0214] S402: Input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using any implementation of the model determination method provided in the embodiments of the present application.
[0215] For more information about the image generation model, please refer to the above text.
[0216] The target image refers to the image obtained by the image generation model based on the above-mentioned image description of the information to be retained and the image constraint description text to perform image generation processing, so that the target image not only satisfies the constraints described in the image constraint description text, but also makes the target image consistent with the image description of the information to be retained in terms of the information to be retained.
[0217] Based on the relevant content of S401 to S402 above, it can be seen that in some application scenarios, a pre-built image generation model can be used to perform image generation processing based on the image describing the information to be retained (for example, a reference image specified by the user) and the image constraint description text (for example, a prompt text specified by the user) to obtain a target image, so that the target image not only satisfies the constraints described by the image constraint description text, but also makes the target image consistent with the image describing the information to be retained in terms of the information to be retained, which is conducive to improving the image generation effect.
[0218] Based on the model determination method provided in the embodiments of the present application, the embodiments of the present application also provide a model determination device, which is explained and illustrated below in conjunction with Figure 5. Figure 5 is a schematic diagram of the structure of a model determination device provided in the embodiments of the present application. It should be noted that for the technical details of the model determination device provided in the embodiments of the present application, please refer to the relevant content of the model determination method above.
[0219] As shown in FIG5 , the model determination device 500 provided in an embodiment of the present application includes:
[0220] A first acquiring unit 501 is configured to acquire a model to be processed, where the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer;
[0221] A first determining unit 502 is configured to determine a target object from the at least one candidate object based on the influence representation data of each candidate object; the influence representation data includes information-preserving dimension influence representation data and target dimension influence representation data; the influence representation data is determined based on a reference image and a prompt text; the information-preserving dimension influence representation data of the target object is higher than the target dimension influence representation data of the target object;
[0222] The first adjustment unit 503 is used to adjust the model to be processed according to the target object to obtain an image generation model, where the image generation model includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
[0223] In a possible implementation manner, different candidate objects are different network layers;
[0224] or,
[0225] Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
[0226] In one possible implementation, for any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
[0227] In a possible implementation manner, if different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object in the image editing dimension;
[0228] If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or image generation quality dimension.
[0229] In a possible implementation manner, the model determination device 500 further includes:
[0230] a second adjusting unit, configured to adjust the model to be processed according to any candidate object to obtain an adjusted model, so that a state of the candidate object in the adjusted model is different from a state of the candidate object in the model to be processed;
[0231] a third acquiring unit, configured to acquire a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by performing image generation processing by the to-be-processed model on the prompt text, and the second generated image is determined by performing image generation processing by the adjusted model based on the reference image and the prompt text;
[0232] a third determining unit, configured to determine information retention dimension impact representation data of the candidate object based on gap representation data between the information retention dimension performance representation data of the first generated image and the information retention dimension performance representation data of the second generated image; the information retention dimension performance representation data of the first generated image is determined based on the first generated image and the reference image; and the information retention dimension performance representation data of the second generated image is determined based on the second generated image and the reference image;
[0233] The fourth determination unit is used to determine the target dimension impact representation data of the candidate object based on the gap representation data between the target dimension performance representation data of the first generated image and the target dimension performance representation data of the second generated image; the target dimension performance representation data of the first generated image is determined based on the first generated image; the target dimension performance representation data of the second generated image is determined based on the second generated image.
[0234] In a possible implementation manner, when generating the first generated image using the model to be processed, it is necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0235] When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0236] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention;
[0237] The second adjustment unit is specifically used to: adjust the candidate object in the to-be-processed model to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
[0238] In one possible implementation, when generating the first generated image using the model to be processed, it is not necessary to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to the candidate object;
[0239] When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
[0240] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0241] The second adjustment unit is specifically used to: adjust the candidate object in the model to be processed to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
[0242] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built base model, the base model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by performing image generation processing by the model to be processed based on the reference image and the prompt text;
[0243] or,
[0244] The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
[0245] In one possible implementation, if different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
[0246] In one possible implementation, if different candidate objects are network layer sets corresponding to different time steps, the target dimension impact representation data of the candidate object includes image editing dimension impact representation data and / or image generation quality dimension impact representation data; the image editing dimension impact representation data is determined based on the gap representation data between the image editing dimension performance representation data of the first generated image and the image editing dimension performance representation data of the second generated image; the image editing dimension performance representation data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance representation data of the second generated image is determined based on the similarity between the second generated image and the prompt text; the image generation quality dimension impact representation data is determined based on the gap representation data between the image generation quality dimension performance representation data of the first generated image and the image generation quality dimension performance representation data of the second generated image; the image generation quality dimension performance representation data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance representation data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
[0247] In one possible implementation, when the image generation model performs image generation processing based on the reference image and the prompt text, there is no need to introduce representational features of the information to be retained in the reference image into the data processing corresponding to other objects except the target object.
[0248] In one possible implementation, the model to be processed is a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing based on the prompt text, and the fine-tuning processing is used to learn information retention;
[0249] The first adjustment unit 503 is specifically used to: adjust the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model.
[0250] In a possible implementation manner, the model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text;
[0251] The first adjustment unit 503 is specifically used to: adjust the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the model to be processed.
[0252] In a possible implementation manner, if different candidate objects are different network layers, the model determination device 500 further includes:
[0253] The third adjustment unit is used to adjust the output weights corresponding to the target object in the image generation model so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
[0254] Based on the relevant content of the above-mentioned model determination device 500, it can be known that for the model determination device 500 provided in the embodiment of the present application, the model to be processed (for example, a diffusion model, etc.) is first obtained, and the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; then, based on the information retention dimension influence characterization data and the target dimension influence characterization data of each candidate object, the target object is determined from the at least one candidate object, so that the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object, so that the target object can represent the object that can mainly affect the information retention performance in the model to be processed; then, based on the target object, the model to be processed is adjusted to obtain an image Generate a model so that the image generation model includes the target object, so that when the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object, thereby enabling the image generation model to achieve performance balance in the two dimensions of information retention dimension and target dimension (for example, image editing, aesthetics, etc.) as much as possible. In this way, the image generation model can take into account the performance in multiple dimensions (for example, information retention, image editing, aesthetics, etc.), so as to achieve the model performance presented in the target dimension while maintaining certain information in the reference image as well as possible, thereby helping to improve the image generation effect.
[0255] Based on the image generation method provided in the embodiments of this application, the embodiments of this application also provide an image generation device, which will be explained and illustrated below in conjunction with Figure 6. Figure 6 is a schematic diagram of the structure of an image generation device provided in the embodiments of this application. It should be noted that for the technical details of the image generation device provided in the embodiments of this application, please refer to the relevant content of the image generation method above.
[0256] As shown in FIG6 , an image generation device 600 provided in an embodiment of the present application includes:
[0257] The second acquisition unit 601 is used to acquire the image describing the information to be retained and the image constraint description text;
[0258] The second determination unit 602 is used to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain the target image output by the image generation model; the image generation model is determined using any implementation of the model determination method provided in this application.
[0259] Based on the relevant content of the above-mentioned image generating device 600, it can be known that for the image generating device 600 provided in the embodiment of the present application, a pre-built image generation model can be used to perform image generation processing based on the image describing the information to be retained (for example, a reference image specified by the user) and the image constraint description text (for example, a prompt text specified by the user) to obtain a target image, so that the target image not only satisfies the constraints described by the image constraint description text, but also makes the target image consistent with the image describing the information to be retained in terms of the information to be retained, which is conducive to improving the image generation effect.
[0260] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the model determination method or image generation method provided in the embodiment of the present application.
[0261] Referring to FIG7 , a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG7 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0262] As shown in Figure 7, electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of electronic device 700 are also stored in RAM 703. Processing device 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0263] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.
[0264] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0265] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0266] An embodiment of the present application also provides a computer-readable medium, which stores instructions or computer programs. When the instructions or computer programs are executed on a device, the device executes any implementation of the model determination method or image generation method provided in the embodiment of the present application.
[0267] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0268] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0269] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0270] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.
[0271] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0272] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0273] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.
[0274] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0275] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0276] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.
[0277] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0278] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0279] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0280] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model determination method, comprising: Acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; Determine a target object from the at least one candidate object based on the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined based on a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object; According to the target object, the model to be processed is adjusted to obtain an image generation model, which includes the target object. When the image generation model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
2. The method according to claim 1, wherein: Different candidate objects are different network layers; or, Different candidate objects are network layer sets corresponding to different time steps, and the network layer set includes at least one network layer.
3. The method according to claim 2, wherein: For any network layer in the network layer set corresponding to the time step, the network layer is determined from at least one network layer corresponding to the time step based on the influence characterization data of at least one network layer corresponding to the time step; the information retention dimension influence characterization data of the network layer is higher than the target dimension influence characterization data of the network layer.
4. The method according to any one of claims 1 to 3, wherein: If different candidate objects are different network layers, then for any candidate object, the target dimension impact representation data of the candidate object is used to represent the impact presented by the candidate object under the image editing dimension; If different candidate objects are sets of network layers corresponding to different time steps, then for any candidate object, the target dimension impact characterization data of the candidate object is used to represent the impact of the candidate object in the image editing dimension and / or the image generation quality dimension.
5. The method according to any one of claims 1 to 4, wherein: For any candidate object, the process of determining the impact characterization data of the candidate object includes: According to the candidate object, the model to be processed is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is different from the state of the candidate object in the model to be processed; Acquire a first generated image and a second generated image corresponding to the candidate object, wherein the first generated image is determined by the model to be processed performing image generation processing on the prompt text, and the second generated image is determined by the adjusted model performing image generation processing based on the reference image and the prompt text; Determining information retention dimension impact characterization data of the candidate object based on gap characterization data between information retention dimension performance characterization data of the first generated image and information retention dimension performance characterization data of the second generated image; the information retention dimension performance characterization data of the first generated image is determined based on the first generated image and the reference image; the information retention dimension performance characterization data of the second generated image is determined based on the second generated image and the reference image; The target dimension impact characterization data of the candidate object is determined based on the gap characterization data between the target dimension performance characterization data of the first generated image and the target dimension performance characterization data of the second generated image; the target dimension performance characterization data of the first generated image is determined based on the first generated image; the target dimension performance characterization data of the second generated image is determined based on the second generated image.
6. The method according to claim 5, wherein: When generating the first generated image using the model to be processed, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object; When the adjusted model is used to generate the second generated image, it is not necessary to introduce the representative features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
7. The method according to claim 6, wherein: The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, wherein the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes: The candidate object in the to-be-processed model is adjusted to obtain an adjusted model, so that the state of the candidate object in the adjusted model is consistent with the state of the candidate object in the basic model.
8. The method according to any one of claims 5 to 7, wherein: When generating the first generated image using the model to be processed, it is not necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object; When the adjusted model is used to generate the second generated image, it is necessary to introduce the characterizing features of the information to be retained in the reference image into the data processing process corresponding to the candidate object.
9. The method according to claim 8, wherein: The model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text; The step of adjusting the model to be processed according to the candidate object to obtain an adjusted model includes: The candidate object in the model to be processed is adjusted to obtain an adjusted model, so that when the adjusted model performs image generation processing based on the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the candidate object.
10. The method according to claim 5, wherein: The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; the first generated image is determined by the image generation processing performed by the model to be processed according to the reference image and the prompt text; or, The model to be processed is a pre-built basic model; the first generated image is determined by the model to be processed performing image generation processing based on the prompt text.
11. The method according to any one of claims 5 to 10, wherein: If different candidate objects are different network layers, the target dimension performance characterization data of the first generated image is determined according to the similarity between the first generated image and the prompt text; The target dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text.
12. The method according to any one of claims 5 to 11, wherein: If different candidate objects are network layer sets corresponding to different time steps, the target dimension impact representation data of the candidate object includes image editing dimension impact representation data and / or image generation quality dimension impact representation data; The image editing dimension impact characterization data is determined based on the difference characterization data between the image editing dimension performance characterization data of the first generated image and the image editing dimension performance characterization data of the second generated image; The image editing dimension performance characterization data of the first generated image is determined based on the similarity between the first generated image and the prompt text; the image editing dimension performance characterization data of the second generated image is determined based on the similarity between the second generated image and the prompt text; The image generation quality dimension impact characterization data is determined based on the gap characterization data between the image generation quality dimension performance characterization data of the first generated image and the image generation quality dimension performance characterization data of the second generated image; The image generation quality dimension performance characterization data of the first generated image is determined by performing quality assessment processing on the first generated image under at least one quality assessment dimension; the image generation quality dimension performance characterization data of the second generated image is determined by performing quality assessment processing on the second generated image under the at least one quality assessment dimension.
13. The method according to claim 1, wherein: When the image generation model performs image generation processing based on the reference image and the prompt text, there is no need to introduce the representation features of the information to be retained in the reference image into the data processing process corresponding to other objects except the target object.
14. The method according to claim 13, wherein: The model to be processed refers to a model obtained by fine-tuning a pre-built basic model, wherein the basic model is used to perform image generation processing according to the prompt text, and the fine-tuning processing is used to learn information retention; The step of adjusting the model to be processed according to the target object to obtain an image generation model includes: Adjust the objects other than the target object in the model to be processed to obtain an image generation model, so that the state of the target object in the image generation model is consistent with the state of the target object in the model to be processed, and the states of the objects other than the target object in the image generation model are consistent with the states of the other objects in the basic model.
15. The method according to claim 13, wherein: The model to be processed is a pre-built basic model, and the basic model is used to perform image generation processing according to the prompt text; The step of adjusting the model to be processed according to the target object to obtain an image generation model includes: The target object in the model to be processed is adjusted to obtain an image generation model, so that the state of the target object in the image generation model is different from the state of the target object in the model to be processed, and the states of other objects except the target object in the image generation model are consistent with the states of the other objects in the model to be processed.
16. The method according to any one of claims 1 to 15, wherein: If different candidate objects are different network layers, the method further includes: The output weights corresponding to the target object in the image generation model are adjusted so that the performance improvement value presented by the adjusted image generation model in the target dimension is higher than the performance degradation value presented by the adjusted image generation model in the information retention dimension.
17. A method for generating an image, comprising: Obtain the information description image and image constraint description text to be retained; The image describing the information to be retained and the image constraint description text are input into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method described in any one of claims 1-16.
18. A model determination device, comprising: A first acquisition unit is configured to acquire a model to be processed, wherein the model to be processed includes at least one candidate object, and the at least one candidate object includes at least one network layer; A first determining unit is configured to determine a target object from the at least one candidate object according to the influence characterization data of each candidate object; the influence characterization data includes information retention dimension influence characterization data and target dimension influence characterization data; the influence characterization data is determined according to a reference image and a prompt text; the information retention dimension influence characterization data of the target object is higher than the target dimension influence characterization data of the target object; The first adjustment unit is configured to adjust the model to be processed according to the target object to obtain an image generation model, wherein the image generation model includes the target object, and when the image generation model performs image generation processing according to the reference image and the prompt text, the representation features of the information to be retained in the reference image are introduced into the data processing process corresponding to the target object.
19. An image generating device, comprising: A second acquisition unit is configured to acquire an image describing information to be retained and an image constraint description text; The second determination unit is configured to input the image describing the information to be retained and the image constraint description text into a predetermined image generation model to obtain a target image output by the image generation model; the image generation model is determined using the model determination method described in any one of claims 1-16.
20. An electronic device, comprising: Processor and memory; The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory so that the electronic device performs the method according to any one of claims 1 to 17.
21. A computer-readable medium storing instructions or computer programs, which, when executed on a device, enable the device to execute the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image generation method and device and electronic equipment
CN115965791A
Image generation method and device, electronic equipment and storage medium
CN116757923A
Image generation method and device
CN116883528A
Image editing method and device, electronic equipment and storage medium
CN116958326A
Generating artistic content from a text prompt or a style image utilizing a neural network model
US20230267652A1