Data processing method, device and equipment and readable storage medium
By combining the image generation network and the reverse network, the backward diffusion and forward diffusion methods are used to solve the problem of low image inversion accuracy in the prior art, and higher image generation accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510228172.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
Existing image inversion methods rely on multi-step sampling, resulting in amplification of approximation errors that will affect the accuracy of the iteration process, resulting in artifacts or details lost in the recovered images, and low accuracy.
Through the combination of the image generation network and the reverse network, the backward diffusion and forward diffusion methods are used to generate noise loss values and image loss values, adjust the model parameters of the reverse network, and improve the accuracy of image inversion.
Improves the accuracy of image generation, reduces artifacts and details loss, and enhances the stability and consistency of the image inversion process.
Smart Images

Figure CN120147478A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method, apparatus, device, and readable storage medium. Background Art
[0002] The goal of image inversion is to recover the original image from a degraded observed (such as blurred, noisy, low-resolution) image. Due to the non-uniqueness and instability of the target solution, most existing image inversion methods rely on approximate calculations through multi-step sampling. Multi-step sampling requires multiple iterations of denoising. The approximation error in the previous steps of iteration is relatively large. The large approximation error will affect the accuracy of the subsequent iteration process. The approximation error will continuously amplify and accumulate during the iteration process, resulting in inaccurate final results, artifacts or detail loss in the recovered image, and low accuracy. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, device, and readable storage medium, which can improve the accuracy of image generation.
[0004] On the one hand, an embodiment of this application provides a data processing method, including:
[0005] Obtain a text prompt and an initial noise vector, and perform backward diffusion on the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector;
[0006] Input the initial latent vector and the text prompt into an initial reverse network, perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector;
[0007] Generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0008] Generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is used to convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and the service noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0009] Among them, the time step corresponding to the initial reverse network is 1. Performing forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and performing backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector includes:
[0010] Input the initial latent vector and the text prompt into the initial reverse network. Perform single-step forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector. Perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector.
[0011] Among them, the time step corresponding to the initial reverse network is M, where M is a positive integer greater than 1. Input the initial latent vector and the text prompt into the initial reverse network. Perform single-step forward diffusion on the initial latent vector through the initial reverse network to obtain a reconstructed noise vector. Perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector, including:
[0012] Obtain the unit latent vector S at the i-th step of the time step i , and input the unit latent vector S i and the text prompt into the initial reverse network. Perform the i-th step of forward diffusion on the unit latent vector S i through the initial reverse network and the text prompt to obtain the unit noise vector P output by the initial reverse network i ; if i is 1, the unit latent vector S i is the initial latent vector; i is a positive integer less than or equal to M;
[0013] According to the initial latent vector, the time step number i corresponding to the unit noise vector P i , the unit noise vector P i and the unit noise vector P i-1 , generate a sampled noise vector Q i , and perform backward diffusion on the sampled noise vector Q i through the image generation network and the text prompt to obtain the unit latent vector S i+1 ; if i is less than M, the unit latent vector S i+1 is the unit latent vector at the (i + 1)-th step of the time step; if i is 1, the unit noise vector P i-1 is an empty vector; if i is greater than 1, the unit noise vector P i-1 is obtained through the initial reverse network after the (i - 1)-th step of forward diffusion;
[0014] If i = M, determine the unit latent vector S i+1 as the reconstructed latent vector.
[0015] Among them, based on the image loss value and the noise loss value, generate a reverse loss value. Adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network, including:
[0016] Input the initial latent vector and the reconstructed latent vector into the discriminator network. Through the discriminator network, the initial latent vector, and the reconstructed latent vector, generate the first adversarial loss value corresponding to the inverse network and the second adversarial loss value corresponding to the discriminator network;
[0017] Based on the image loss value, the noise loss value, and the first adversarial loss value, generate the inverse loss value. Adjust the model parameters of the initial inverse network through the inverse loss value, and adjust the model parameters of the discriminator network through the second adversarial loss value. When both the initial inverse network and the discriminator network meet the model convergence conditions, determine the initial inverse network that meets the model convergence conditions as the inverse network.
[0018] Among them, generating the first adversarial loss value corresponding to the inverse network and the second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector includes:
[0019] Generate the first discriminant parameter for the reconstructed latent vector through the discriminator network. Based on the first discriminant parameter, generate the first expected mean for the reconstructed latent vector. Generate the first adversarial loss value corresponding to the inverse network based on the first expected mean;
[0020] Generate the second discriminant parameter for the initial latent vector through the discriminator network. Based on the second discriminant parameter, generate the second expected mean for the initial latent vector. Generate the second adversarial loss value corresponding to the discriminator network based on the first expected mean and the second expected mean.
[0021] Among them, the noise loss value includes the noise reconstruction loss value and the regularization loss value; generating the noise loss value based on the initial noise vector and the reconstructed noise vector includes:
[0022] Generate the regularization loss value based on the noise probability distribution of the initial noise vector and the noise probability distribution of the reconstructed noise vector;
[0023] Determine the mean square error between the initial noise vector and the reconstructed noise vector as the noise reconstruction loss value, and determine the sum of the regularization loss value and the noise reconstruction loss value as the noise loss value.
[0024] Among them, generating the image loss value based on the reconstructed latent vector and the initial latent vector includes:
[0025] Determine the mean square error between the initial latent vector and the reconstructed latent vector as the image reconstruction loss value. Obtain the service hyperparameter for balancing the noise reconstruction loss value and the image reconstruction loss value, and perform a dot product operation on the service hyperparameter and the image reconstruction loss value to obtain the image loss value.
[0026] Among them, the image generation network includes a basic generation network and a refinement generation network; obtaining a text prompt and an initial noise vector, and performing backward diffusion on the initial noise vector through the image generation network and the text prompt to obtain an initial latent vector, including:
[0027] Obtain a text prompt and an initial noise image, extract features from the text prompt to obtain a text prompt vector, and perform image encoding on the initial noise image to obtain an initial noise vector;
[0028] Input the initial noise vector into the basic generation network. In the basic generation network, perform backward diffusion on the initial noise vector through the text prompt vector to obtain a basic latent vector;
[0029] Input the basic latent vector into the refinement generation network. In the refinement generation network, perform backward diffusion on the basic latent vector through the text prompt vector to obtain an initial latent vector.
[0030] Among them, the initial inverse network includes a downsampling network and an upsampling network, and the input of the upsampling network is the output of the downsampling network layer; perform forward diffusion on the initial latent vector through the initial inverse network and the text prompt to obtain a reconstructed noise vector, including:
[0031] In the downsampling network of the initial inverse network, perform self-attention processing on the initial latent vector to obtain a downsampling attention vector, and perform cross-attention processing on the downsampling attention vector and the text prompt to obtain a downsampling feature vector; the feature dimension of the downsampling feature vector is smaller than that of the initial latent vector;
[0032] In the upsampling network of the initial inverse network, perform self-attention processing on the downsampling feature vector to obtain an upsampling attention vector, and perform cross-attention processing on the upsampling attention vector and the text prompt to obtain a reconstructed noise vector; the feature dimension of the reconstructed noise vector is equal to that of the initial latent vector.
[0033] Among them, performing cross-attention processing on the downsampling attention vector and the text prompt to obtain a downsampling feature vector includes:
[0034] Perform a dot product operation on the downsampling attention vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the downsampling attention vector and the key parameter matrix to obtain a key vector, and perform a dot product operation on the text prompt vector corresponding to the text prompt and the value parameter matrix to obtain a value vector;
[0035] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the dimension of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain a normalized vector, and perform a dot product operation on the normalized vector and the value vector to obtain a downsampling feature vector.
[0036] Another aspect of the embodiments of the present application provides a data processing method, including:
[0037] Obtain a first image and a target text prompt input by an object, perform image recognition on the first image to obtain an image recognition result, generate a business text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method provided in the embodiments of the present application;
[0038] Convert the first latent vector into a business noise vector through the reverse network and the business text prompt;
[0039] Input the business noise vector and the target text prompt into an image generation network. In the image generation network, perform backward diffusion on the business noise vector through the target text prompt to obtain a second latent vector;
[0040] Perform image decoding on the second latent vector to obtain a second image that matches the target text prompt.
[0041] One aspect of the embodiments of the present application provides a data processing device, including:
[0042] A backward diffusion module, configured to obtain a text prompt and an initial noise vector, and perform backward diffusion on the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector;
[0043] An image inversion module, configured to input the initial latent vector and the text prompt into an initial reverse network, perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector;
[0044] A loss calculation module, configured to generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0045] A model adjustment module, configured to generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is configured to convert the latent vector representing the first image into a business noise vector through the business text prompt corresponding to the first image, and the business noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0046] In a possible implementation, the time step corresponding to the initial reverse network is 1; when the image inversion module is used to perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector, it is specifically used to perform the following operations:
[0047] Input the initial latent vector and the text prompt into the initial reverse network, perform single-step forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector.
[0048] In a possible implementation, the time step corresponding to the initial reverse network is M, where M is a positive integer greater than 1; when the image inversion module is used to input the initial latent vector and the text prompt into the initial reverse network, perform single-step forward diffusion on the initial latent vector through the initial reverse network to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector, it is specifically used to perform the following operations:
[0049] Obtain the unit latent vector S in the i-th step at the time step i , input the unit latent vector S i and the text prompt into the initial reverse network, and perform the i-th step of forward diffusion on the unit latent vector S i through the initial reverse network and the text prompt to obtain the unit noise vector P output by the initial reverse network i ; if i is 1, the unit latent vector S i is the initial latent vector; i is a positive integer less than or equal to M;
[0050] According to the initial latent vector, the time step i corresponding to the unit noise vector P i , the unit noise vector P i and the unit noise vector P i-1 , generate a sampled noise vector Q i , perform backward diffusion on the sampled noise vector Q i through the image generation network and the text prompt to obtain the unit latent vector S i+1 ; if i is less than M, the unit latent vector S i+1 is the unit latent vector in the (i + 1)-th step at the time step; if i is 1, the unit noise vector P i-1 is an empty vector; if i is greater than 1, the unit noise vector P i-1 is obtained through the initial reverse network after the (i - 1)-th step of forward diffusion;
[0051] If i = M, the unit latent vector S i+1 is determined as the reconstructed latent vector.
[0052] In a possible implementation, the model adjustment module is used to generate an inverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial inverse network through the inverse loss value to obtain the inverse network. Specifically, it is used to perform the following operations:
[0053] Input the initial latent vector and the reconstructed latent vector into the discriminator network, and generate the first adversarial loss value corresponding to the inverse network and the second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector;
[0054] Generate an inverse loss value based on the image loss value, the noise loss value, and the first adversarial loss value, adjust the model parameters of the initial inverse network through the inverse loss value, and adjust the model parameters of the discriminator network through the second adversarial loss value. When both the initial inverse network and the discriminator network meet the model convergence condition, the initial inverse network that meets the model convergence condition is determined as the inverse network.
[0055] In a possible implementation, when the model adjustment module is used to generate the first adversarial loss value corresponding to the inverse network and the second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector, it is specifically used to perform the following operations:
[0056] Generate the first discriminant parameter for the reconstructed latent vector through the discriminator network, generate the first expected mean for the reconstructed latent vector based on the first discriminant parameter, and generate the first adversarial loss value corresponding to the inverse network based on the first expected mean;
[0057] Generate the second discriminant parameter for the initial latent vector through the discriminator network, generate the second expected mean for the initial latent vector based on the second discriminant parameter, and generate the second adversarial loss value corresponding to the discriminator network based on the first expected mean and the second expected mean.
[0058] In a possible implementation, the noise loss value includes a noise reconstruction loss value and a regularization loss value; when the loss calculation module is used to generate the noise loss value based on the initial noise vector and the reconstructed noise vector, it is specifically used to perform the following operations:
[0059] Generate the regularization loss value based on the noise probability distribution of the initial noise vector and the noise probability distribution of the reconstructed noise vector;
[0060] Determine the mean square error between the initial noise vector and the reconstructed noise vector as the noise reconstruction loss value, and determine the sum of the regularization loss value and the noise reconstruction loss value as the noise loss value.
[0061] In a possible implementation, when the loss calculation module is used to generate an image loss value based on the reconstructed latent vector and the initial latent vector, it is specifically used to perform the following operations:
[0062] Determine the mean square error between the initial latent vector and the reconstructed latent vector as the image reconstruction loss value, obtain a service hyperparameter for balancing the noise reconstruction loss value and the image reconstruction loss value, and perform a dot product operation on the service hyperparameter and the image reconstruction loss value to obtain the image loss value.
[0063] In a possible implementation, the image generation network includes a basic generation network and a refinement generation network; when the backward diffusion module is used to obtain a text prompt and an initial noise vector and perform backward diffusion on the initial noise vector through the image generation network and the text prompt to obtain the initial latent vector, it is specifically used to perform the following operations:
[0064] Obtain a text prompt and an initial noise image, perform feature extraction on the text prompt to obtain a text prompt vector, and perform image encoding on the initial noise image to obtain an initial noise vector;
[0065] Input the initial noise vector into the basic generation network, and in the basic generation network, perform backward diffusion on the initial noise vector through the text prompt vector to obtain a basic latent vector;
[0066] Input the basic latent vector into the refinement generation network, and in the refinement generation network, perform backward diffusion on the basic latent vector through the text prompt vector to obtain the initial latent vector.
[0067] In a possible implementation, the initial inverse network includes a downsampling network and an upsampling network, and the input of the upsampling network is the output of the downsampling network layer; when the image inversion module is used to perform forward diffusion on the initial latent vector through the initial inverse network and the text prompt to obtain a reconstructed noise vector, it is specifically used to perform the following operations:
[0068] In the downsampling network of the initial inverse network, perform self-attention processing on the initial latent vector to obtain a downsampling attention vector, and perform cross-attention processing on the downsampling attention vector and the text prompt to obtain a downsampling feature vector; the feature dimension of the downsampling feature vector is smaller than the feature dimension of the initial latent vector;
[0069] In the upsampling network of the initial inverse network, perform self-attention processing on the downsampling feature vector to obtain an upsampling attention vector, and perform cross-attention processing on the upsampling attention vector and the text prompt to obtain a reconstructed noise vector; the feature dimension of the reconstructed noise vector is equal to the feature dimension of the initial latent vector.
[0070] In a possible implementation, when the image inversion module is used to perform cross-attention processing on the downsampled attention vector and the text prompt to obtain the downsampled feature vector, it is specifically used to perform the following operations:
[0071] Perform a dot product operation on the downsampled attention vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the downsampled attention vector and the key parameter matrix to obtain a key vector, and perform a dot product operation on the text prompt vector corresponding to the text prompt and the value parameter matrix to obtain a value vector;
[0072] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the dimension of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain a normalized vector, and perform a dot product operation on the normalized vector and the value vector to obtain the downsampled feature vector.
[0073] On the one hand, an embodiment of the present application provides another data processing device, including:
[0074] An input processing module, configured to obtain a first image and a target text prompt input by an object, perform image recognition on the first image to obtain an image recognition result, generate a service text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method provided in the embodiment of the present application;
[0075] A conversion processing module, configured to convert the first latent vector into a service noise vector through the reverse network and the service text prompt;
[0076] A noise prediction module, configured to input the service noise vector and the target text prompt into an image generation network, and in the image generation network, perform backward diffusion on the service noise vector through the target text prompt to obtain a second latent vector;
[0077] An image generation module, configured to perform image decoding on the second latent vector to obtain a second image that matches the target text prompt.
[0078] On the one hand, an embodiment of the present application provides a computer device, including: a processor, a memory, and a network interface;
[0079] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, and the memory is used to store a computer program. When the computer program is executed by the processor, the computer device executes the method provided in the embodiment of the present application.
[0080] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiments of the present application.
[0081] One aspect of the embodiments of the present application provides a computer program product including a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program so that the computer device executes the method provided by the embodiments of the present application.
[0082] In the embodiments of the present application, an initial inverse network parameterizes the image inversion process of converting a latent vector into a noise vector. By directly capturing the mapping relationship between the latent vector and the noise vector, the mapping of data from one domain (latent vector) to another domain (noise vector) is learned. By calculating the noise loss value between the initial noise vector and the reconstructed noise vector, and the image loss value between the reconstructed latent vector and the initial latent vector, an explicit training objective is defined for training the initial inverse network, enabling the initial inverse network to learn to match the reconstructed noise vector with the initial noise vector and the reconstructed latent vector with the initial latent vector, ensuring the consistency of the image inversion process of converting the latent vector into the noise vector, that is, the initial latent vector can be inversely transformed into a reconstructed noise vector through image inversion, and this reconstructed noise still has consistency with the initial noise vector, thereby improving the accuracy of image inversion by the trained inverse network. Therefore, the noise vector generated by the trained inverse network has higher precision and can more accurately reconstruct the features of the input image. The inverse network can be used to more accurately convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and then the image generation network can generate a second image associated with the first image through the service noise vector, improving the accuracy of image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0084] Figure 1 It is a schematic diagram of a network architecture provided by the embodiments of the present application;
[0085] Figure 2 It is a schematic diagram of a data processing scenario provided by the embodiments of the present application;
[0086] Figure 3 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;
[0087] Figure 4 It is a schematic diagram of a model structure provided by an embodiment of the present application Figure 1 ;
[0088] Figure 5 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;
[0089] Figure 6 It is a schematic diagram of a model structure provided by an embodiment of the present application Figure 2 ;
[0090] Figure 7 It is a schematic diagram of a model structure provided by an embodiment of the present application Figure 3 ;
[0091] Figure 8 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application Figure 1 ;
[0092] Figure 9 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application Figure 2 ;
[0093] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application Detailed implementation manners
[0094] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application
[0095] Please refer to Figure 1 , Figure 1 It is a schematic diagram of a network architecture provided by an embodiment of the present application. As Figure 1As shown in the figure, the network architecture may include a business server 100 and a cluster of terminal devices. The cluster of terminal devices may include terminal devices 10a, 10b, …, 10n. Among them, any terminal device in the cluster of terminal devices may have a communication connection with the business server 100. For example, there is a communication connection between terminal device 10a and business server 100, and there is a communication connection between terminal device 10b and business server 100. Among them, the above communication connection does not limit the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be through other methods, which are not limited in this application.
[0096] Among them, each terminal device in the cluster of terminal devices may include: smart phones, tablets, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances (such as smart TVs), wearable devices, vehicle-mounted terminals, aircraft and other intelligent terminals with data processing functions. Among them, the vehicle-mounted terminal can be a terminal device in the intelligent transportation scenario and the assisted driving scenario. It should be understood that, as Figure 1 shown, each terminal device in the cluster of terminal devices may be installed with an application client with data processing functions. When the application client runs on each terminal device, it can perform data interaction with the above Figure 1 shown business server 100 respectively.
[0097] Among them, the application client may specifically include: vehicle-mounted client, smart home client, entertainment client (such as game client), multimedia client (such as video client), social client, and information client (such as news client), etc. Among them, the application client in the embodiments of this application may be integrated in a certain client (such as a social client), or may be an independent client (such as a news client). The embodiments of this application do not limit the type of the application client.
[0098] The application client may be an image editing application in the multimedia client. Any terminal device in the cluster of terminal devices may perform graphic editing through the application client. The process is shown in Figure 2 , Figure 2 which is a schematic diagram of a data processing scenario provided by the embodiments of this application. As Figure 2 shown, the terminal device may be deployed with a reverse network. The content of the reverse network can be seen in the specific description of the corresponding embodiment below Figure 3 . The business server may be deployed with an image generation network. The image generation network may be a backward diffusion structure of a diffusion model, which is used to diffuse a noise vector representing noise into a latent vector representing an image. Among them, the terminal device may be the above Figure 1For any one of the terminal devices in the corresponding embodiments, for example, it may be terminal device 10a, and the service server may be the above Figure 1 service server 100 in the corresponding embodiments. Both the reverse network and the image generation network may include a U-net network structure implemented based on the attention mechanism (QKV), and the embodiments of the present application do not limit this here.
[0099] The object (user) can input image A and a target text prompt in the application client of the terminal device. Among them, the target text prompt may be a text prompt for describing the image editing performed on the image content of image A, or may be a text prompt for describing the image content obtained after image editing of image A. For example, the target text prompt may be expressed as "a girl standing in the park". Or it may be expressed as "replace the background of the picture with a park", and the embodiments of the present application do not limit this here.
[0100] The terminal device can perform image recognition on image A to obtain an image recognition result, and the terminal device can generate a service text prompt for describing the image content of image A based on the image recognition result. For example, the service text prompt may be expressed as "a girl standing indoors".
[0101] The terminal device can perform image encoding on image A to obtain a first latent vector corresponding to image A, and through the reverse network and the service text prompt, convert the first latent vector corresponding to image A into a service noise vector. The terminal device can send the service noise vector and the target text prompt to the service server. Among them, the first latent vector can be used to characterize the features of the image content of the first image. The service noise vector can characterize the noise image corresponding to the first image and can be used to restore the first image after gradually denoising.
[0102] The service server can input the service noise vector and the target text prompt into the image generation network. In the image generation network, the service noise vector is backward diffused through the target text prompt to obtain a second latent vector, and the second latent vector is decoded into an image B that matches the target text prompt. The service server can send image B to the terminal device.
[0103] In the embodiments of the present application, the inverse network can perform image inversion on the latent vector representing the image to obtain the noise vector representing the noise. The image inversion can replace the forward diffusion process in the diffusion model, and the image generation network can represent the backward diffusion process in the diffusion model. Through the distributed deployment of the inverse network and the image generation network, the object can send the locally processed service noise vector and the target text prompt to the server for image generation, thereby protecting the privacy and security of the original image data. Optionally, the inverse network and the image generation network can also be both deployed in the service server to reduce the computing pressure on the terminal device. The inverse network in the embodiments of the present application can be applied to image generation scenarios that rely on image inversion, such as the image-to-image scenario, the image editing scenario, etc. In the image editing scenario, the target text prompt can represent the image editing process required for the first image, so as to conveniently and easily implement advanced image editing, such as changing the subject in the image content, adjusting the image style, etc.
[0104] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data processing method provided by the embodiments of the present application Figure 1 , and this data processing method can be executed by a computer device, and the computer device can be any one of the service server 100 or the terminal device cluster as shown in Figure 1 , for example, it can be the terminal device 10a. The following will take this data processing method being executed by a computer device as an example for description. Among them, this data processing method can at least include the following steps S101 - step S104:
[0105] Step S101, obtain the text prompt and the initial noise vector, and perform backward diffusion on the initial noise vector through the image generation network and the text prompt to obtain the initial latent vector;
[0106] Specifically, please also refer to Figure 4 , Figure 4 which is a schematic diagram of a model structure proposed by the embodiments of the present application Figure 1 , as shown in Figure 4 , the text prompt can be represented as c, and the initial noise vector can be represented as ∈. The computer device can obtain the text prompt c and the initial noise vector ∈. Among them, the text prompt c can be randomly generated and is used to describe the text of the image content. The initial noise vector ∈ can be a noise vector that conforms to the Gaussian distribution (normal distribution), that is, ∈ ~ N(0,1), and for example, it can be sampled and generated from a noise image that conforms to the Gaussian distribution.
[0107] As shown in Figure 4As shown, the image generation network can be denoted as G. The computer device can input the text prompt c and the initial noise vector ∈ into the image generation network G. The image generation network G can be the backward diffusion structure of a diffusion model, which is used to diffuse the noise vector representing noise into a latent vector representing an image. The image generation network G can be the backward diffusion structure of a few-step diffusion model (a model that reduces the number of iteration steps in the diffusion model to achieve fast image generation), and the image generation network G can also include a network structure implemented based on the attention mechanism (QKV). Among them, the few-step diffusion model can be SDXL-Turbo (Stable Diffusion eXtra Large-Turbo, a distilled model of SDXL) and Flux.1-Schnell (a model of the Flux.1 series), and the embodiments of the present application do not limit this here.
[0108] The computer device can perform backward diffusion (Backward Process) on the initial noise vector ∈ through the image generation network G and the text prompt c to obtain the initial latent vector z. 0 The initial latent vector can also be denoted as G(∈, c, T), where T is the number of time steps used by the image generation model G for backward diffusion. Among them, backward diffusion refers to predicting noise or removing noise in the noise vector based on the information indicated by the text prompt through a series of steps, so as to recover or generate a latent vector representing an image.
[0109] Step S102: Input the initial latent vector and the text prompt into the initial inverse network, perform forward diffusion on the initial latent vector through the initial inverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector.
[0110] Specifically, please also refer to Figure 4 As Figure 4 shown, the computer device can input the initial latent vector z 0 and the text prompt c into the initial inverse network F. The initial inverse network F can be a neural network for learning image inversion, and the initial inverse network F can also include a network structure implemented based on the attention mechanism (QKV). For example, it can be an MMDIT network (Multi-Modal Diffusion Transformer, a multi-modal DiT network), a U-net network (a U-shaped noise prediction network), and the embodiments of the present application do not limit this here. Image inversion means that given an original image (or a latent vector representing an image) and a descriptive prompt (a text prompt describing the image content of the original image), image inversion aims to find an initial noise corresponding to the original image, and this initial noise can accurately reconstruct the original image after gradually denoising.
[0111] The computer device can perform a single-step forward diffusion (Forward Process) on the initial latent vector z through the initial reverse network F and the text prompt c 0 to obtain a reconstructed noise vector The reconstructed noise vector can also be expressed as F(z 0 , c, T). Among them, forward diffusion refers to adding Gaussian-distributed noise to the latent vector based on the information indicated by the text prompt through a series of steps, so as to generate a noise vector associated with the latent vector. The computer device can perform backward diffusion on the reconstructed noise vector through the image generation network G and the text prompt c to obtain a reconstructed latent vector
[0112] It can be understood that since the initial reverse network F can learn the direct mapping relationship between the latent vector and the noise vector, the initial reverse network F can convert the latent vector into a noise vector through a single-step forward diffusion. Among them, the single-step forward diffusion means that the number of time steps of forward diffusion is 1.
[0113] Step S103, generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0114] Specifically, please also refer to Figure 4 , such as Figure 4 shown, the computer device can generate a noise loss value L based on the initial noise vector ∈ and the reconstructed noise vector Noise . Among them, the noise loss value L Noise can include a regularization loss value L reg (Regularization Loss) and a noise reconstruction loss value (Noise Reconstruction Loss). The regularization loss value can be a loss value used to prevent the model from overfitting. For example, it can refer to the KL divergence (Kullback-Leibler Divergence, relative entropy) between the noise probability distribution of the initial noise vector and the noise probability distribution of the reconstructed noise vector. Its calculation process can be as shown in formula (1):
[0115] L reg = KL((z 0 , c, T), (N(0, 1))) Formula (1)
[0116] Among them, the noise probability distribution of the initial noise vector is the Gaussian distribution, expressed as N(0, 1).
[0117] Noise reconstruction loss value It can be a loss used to measure the difference between the output noise and the input noise. The noise reconstruction loss value can also be a Noise Cycle Loss, which is used to encourage the reconstructed noise vector obtained after image inversion to match the initial noise vector ∈. For example, the noise reconstruction loss value can be the Mean Squared Error (MSE) between the initial noise vector and the reconstructed noise vector, and its calculation process can be as shown in formula (2):
[0118]
[0119] The computer device can use the regularization loss value L reg and the noise reconstruction loss value The sum of them is determined as the noise loss value L Noise . Based on the reconstructed latent vector and the initial latent vector, an image loss value L Image is generated. The image loss value L Image can be generated by the image reconstruction loss value . The image reconstruction loss value can be a loss used to measure the difference between the output image and the input image. The image reconstruction loss value can also be an ImageCycle Loss, which is used to encourage the reconstructed latent vector obtained after image inversion to match the initial latent vector z 0 . For example, the image reconstruction loss value can be the mean square error between the initial image vector and the reconstructed image vector, and its calculation process can be as shown in formula (3):
[0120]
[0121] The computer device can also introduce a business hyperparameter λ for balancing the noise reconstruction loss value and the image reconstruction loss value . The dot product operation result of the business hyperparameter λ and the image reconstruction loss value is determined as the image loss value L Image .
[0122] Step S104: Based on the image loss value and the noise loss value, generate a reverse loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is used to convert the latent vector representing the first image into a business noise vector through the business text prompt corresponding to the first image, and the business noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0123] Specifically, the computer device can generate an inverse loss value L Image based on the image loss value L Noise and the noise loss value L full For example, the sum of the image loss value L Image and the noise loss value L Noise can be determined as the inverse loss value L full The model parameters of the initial inverse network are adjusted through the inverse loss value. When the model parameters of the initial inverse network meet the model convergence condition, the inverse network is obtained. Among them, the model convergence condition can be that during the training process, if the loss value of the loss function does not decrease significantly after multiple training epochs or the training epochs of the initial inverse network reach a preset maximum value, etc. The embodiments of the present application do not limit the model convergence condition here. The inverse network is used to convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and the service noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0124] In the embodiments of the present application, an initial inverse network is used to parameterize the image inversion process of converting a latent vector into a noise vector. By directly capturing the mapping relationship between the latent vector and the noise vector, the mapping of data from one domain (latent vector) to another domain (noise vector) is learned. By calculating the noise loss value between the initial noise vector and the reconstructed noise vector, and the image loss value between the reconstructed latent vector and the initial latent vector, an explicit training objective is defined for training the initial inverse network, enabling the initial inverse network to learn to match the reconstructed noise vector with the initial noise vector and the reconstructed latent vector with the initial latent vector, ensuring the consistency of the image inversion process of converting the latent vector into the noise vector, that is, the initial latent vector can be inversely transformed through image inversion to obtain a reconstructed noise vector, and this reconstructed noise still has consistency with the initial noise vector, thereby improving the accuracy of the image inversion performed by the trained inverse network. Therefore, the noise vector generated by the trained inverse network has higher precision and can more accurately reconstruct the features of the input image. The inverse network can be used to more accurately convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and then the image generation network can generate a second image associated with the first image through the service noise vector, improving the accuracy of image generation.
[0125] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a data processing method provided by the embodiments of the present application Figure 2 This data processing method can be executed by a computer device, and the computer device can be, for example, Figure 1The business server 100 shown or any one of the terminal device clusters, for example, can be the terminal device 10a. Hereinafter, an example will be given in which this data processing method is executed by a computer device. Among them, the data processing method can at least include the following steps S201-S205:
[0126] Step S201, obtain a text prompt and an initial noise image, extract features from the text prompt to obtain a text prompt vector, perform image encoding on the initial noise image to obtain an initial noise vector; input the initial noise vector into a basic generation network, and in the basic generation network, perform backward diffusion on the initial noise vector through the text prompt vector to obtain a basic latent vector; input the basic latent vector into a refined generation network, and in the refined generation network, perform backward diffusion on the basic latent vector through the text prompt vector to obtain an initial latent vector.
[0127] Specifically, the computer device can obtain a text prompt and an initial noise image. Among them, the text prompt can be randomly generated text for describing the image content. The initial noise image can be sampled and generated from a noise image conforming to a Gaussian distribution. The computer device can perform text splitting (tokenization) on the text prompt to obtain a token sequence corresponding to the text prompt. For example, if the text prompt is "anime characters in an outdoor scene", the token sequence corresponding to the object text information after removing stop words can be expressed as [anime characters, outdoor scene]. Among them, [anime characters] and [outdoor scene] are both tokens. A token in the token sequence refers to the smallest basic unit obtained after splitting the main text information. The method of text splitting can be word-based, character-based, subword-based, and the embodiments of the present application do not limit this here.
[0128] The computer device can perform text encoding on the text prompt through a text encoder to obtain a text prompt vector. The text encoder can be the encoder part in models such as a pre-trained image text encoding model (Contrastive Language-Image Pre-Training, CLIP) and a Bayesian bidirectional encoder (Bidirectional Encoder Representations from Transformers, BERT).
[0129] The computer device can perform image encoding on the initial noise image through an image encoder to obtain an initial latent vector. The image encoder can be the encoder in a variational auto encoder (VAE), and the embodiments of the present application do not limit this here.
[0130] For ease of understanding, the image generation network SDXL-Turbo is taken as an example for illustration. The processing process of SDXL-Turbo is a two-stage cascaded diffusion model. SDXL-Turbo may include a base generation network (Base Models) and a refinement generation network (Refiner Models). The base generation network and the refinement generation network can be connected in series. Both the base generation network and the refinement generation network can have a U-net structure, but the number of model parameters of the base generation network is different from that of the refinement generation network.
[0131] The computer device can input an initial noise vector and a text prompt vector into the base generation network. In the base generation network of the image generation network, the initial noise vector is backward diffused through the text prompt vector to obtain a base latent vector. Among them, backward diffusion refers to predicting noise or removing noise in the noise vector based on the information indicated by the text prompt through a series of steps, so as to recover or generate a latent vector representing the image.
[0132] The computer device can input the base latent vector into the refinement generation network. The refinement generation network is used to further perform more detailed denoising steps on the base latent vector generated by the base generation network. In the refinement generation network of the image generation network, the computer device can continue to perform backward diffusion on the base latent vector through the text prompt vector to obtain an initial latent vector.
[0133] Step S202: In the downsampling network of the initial inverse network, perform self-attention processing on the initial latent vector to obtain a downsampled attention vector, and perform cross-attention processing on the downsampled attention vector and the text prompt to obtain a downsampled feature vector; the feature dimension of the downsampled feature vector is smaller than that of the initial latent vector; in the upsampling network of the initial inverse network, perform self-attention processing on the downsampled feature vector to obtain an upsampled attention vector, and perform cross-attention processing on the upsampled attention vector and the text prompt to obtain a reconstructed noise vector; the feature dimension of the reconstructed noise vector is equal to that of the initial latent vector.
[0134] Specifically, for ease of understanding, taking the initial inverse network as a U-net network as an example, the initial inverse network may include a downsampling network (Down Block) and an upsampling network (Up Block). Please also refer to Figure 6 , Figure 6 which is a schematic diagram of a model structure provided by an embodiment of the present application Figure 2 , such as Figure 6As shown, both the downsampling network and the upsampling network of the initial inverse network can include several network units (Blocks), and each network unit can include a skip connection layer, a self-attention layer, and a cross-attention layer. Among them, the skip connection layer is used to send the received vector to the connected network layer. The modules composed of each network unit are the same, but the corresponding feature dimension directions (i.e., the parameter dimensions in the network layer) are different, so as to realize the sampling process from down-mid-up.
[0135] The initial inverse network can inject the text prompt vector corresponding to the text prompt c into the cross-attention layer in each network unit, and perform feature fusion and sampling (reducing the feature dimension or restoring the feature dimension) through the cross-attention processing of the cross-attention layer in each network unit.
[0136] Taking a network unit A in the downsampling network of the initial inverse network as an example, this network unit A includes a first connection layer (the skip connection layer corresponding to network unit A), a first self-attention layer (the self-attention layer corresponding to network unit A), and a first cross-attention layer (the cross-attention layer corresponding to network unit A). The first connection layer can send the vector received by network unit A (which can be a forward noise vector) to the skip connection layer corresponding to the connected network unit, for example, it can be the second connection layer (the skip connection layer corresponding to a network unit B in the downsampling network layer).
[0137] The input of network unit A can be an initial noise vector. In the first self-attention layer, the computer device can perform self-attention processing on the initial noise vector to obtain a first self-attention result vector (downsampling attention vector). In the first cross-attention layer, cross-attention processing is performed on the first self-attention result vector (downsampling attention vector) and the injected text prompt vector, and feature fusion and downsampling (reducing the feature dimension) are performed through the cross-attention processing to obtain a first cross-attention result vector. Network unit A can input the first cross-attention result vector into the next network unit connected in series, and continue to perform the same processing through the next network unit to obtain the downsampling feature vector output in the downsampling network. Among them, the feature dimension of the downsampling feature vector is smaller than the feature dimension of the initial latent vector.
[0138] Taking the example of performing cross-attention processing on the downsampled attention vector and the text prompt to obtain the downsampled feature vector, the processing process of the cross-attention processing can be as follows: perform a dot product operation on the downsampled attention vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the downsampled attention vector and the key parameter matrix to obtain a key vector, and perform a dot product operation on the text prompt vector corresponding to the text prompt and the value parameter matrix to obtain a value vector; generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the dimension of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain a normalized vector, and perform a dot product operation on the normalized vector and the value vector to obtain the downsampled feature vector.
[0139] Specifically, the computer device can perform a dot product operation on the downsampled attention vector and the query parameter matrix W q to obtain a query vector Q, perform a dot product operation on the downsampled attention vector and the key parameter matrix W K to obtain a key vector K, perform a dot product operation on the text prompt vector and the value parameter matrix W V to obtain a value vector V. The computer device can perform a dot product operation on the query vector Q and the transposed operation result K T of the key vector K to obtain an attention score vector. Based on the dimension value d K of the key vector K, take the square root to perform dimensionality reduction processing on the attention score matrix, perform normalization processing on the dimensionally reduced attention score vector to obtain a normalized vector, and perform a dot product operation on the normalized vector and the value vector V to obtain the downsampled feature vector. The downsampled feature vector can be represented as Z att , and its process can be as shown in formula (4):
[0140]
[0141] where softmax is the normalization processing.
[0142] The computer device can input the downsampled feature vector into the upsampling network. In the upsampling network of the initial reverse network, taking a certain network unit B in the upsampling network of the initial reverse network as an example, this network unit B includes a second connection layer (the skip connection layer corresponding to the network unit B), a second self-attention layer (the second self-attention layer corresponding to the network unit B), and a second cross-attention layer (the cross-attention layer corresponding to the network unit B). The first connection layer of the network unit A in the downsampling network is connected to the second connection layer of the network unit B in the upsampling network, and the second connection layer can obtain the initial noise vector sent by the network unit A through the first connection layer.
[0143] In network unit B, the computer device can perform self-attention processing on the downsampled feature vector to obtain a second self-attention result vector (upsampled attention vector), perform cross-attention processing on the second self-attention result vector (upsampled attention vector) and the text prompt vector, and perform feature fusion and upsampling (restoring the feature dimension) through the cross-attention processing to obtain a second cross-attention result vector. Network unit B can input the second cross-attention result vector into the next network unit connected in series, and continue to perform the same processing through the next network unit to obtain the upsampled feature vector output by the last network unit in the upsampling network. The computer device can determine the upsampled feature vector as the reconstructed noise vector. Among them, the feature dimension of the reconstructed noise vector is equal to the feature dimension of the initial latent vector.
[0144] It can be understood that since the initial reverse network can learn the direct mapping relationship between the latent vector and the noise vector, the initial reverse network can convert the latent vector into the noise vector through single-step forward diffusion or few-step forward diffusion. Among them, single-step forward diffusion means that the number of time steps of forward diffusion is 1, and few-step forward diffusion means forward diffusion with the number of time steps less than the number of time steps required by the diffusion model. For example, the number of time steps of few-step forward diffusion can be 2 to 5.
[0145] The reverse network of the embodiments of the present application can be applied to single-step or few-step, because the reverse network proposed in the embodiments of the present application can reduce the error generated during the diffusion process. Therefore, when the reverse network is applied to few-step, the accuracy of the diffusion process can be further improved by increasing the sampling process during the diffusion process.
[0146] For ease of understanding, taking the time step corresponding to the initial reverse network as M as an example for illustration, M is a positive integer greater than 1. When the time step is M, the processing process of the initial reverse network can be: obtaining the unit latent vector S in the i-th step at the time step i , inputting the unit latent vector S i and the text prompt into the initial reverse network, and performing the i-th step of forward diffusion on the unit latent vector S i through the initial reverse network and the text prompt to obtain the unit noise vector P output by the initial reverse network i ; if i is 1, the unit latent vector S i is the initial latent vector; i is a positive integer less than or equal to M; according to the initial latent vector, the time step i corresponding to the unit noise vector P i , the unit noise vector P i and the unit noise vector P i-1 , generating the sampled noise vector Q i , and performing processing on the sampled noise vector Q i through the image generation network and the text promptPerform backward diffusion to obtain the unit latent vector S i+1 ; if i is less than M, the unit latent vector S i+1 is the unit latent vector in the (i + 1)-th step of the time step; if i is 1, the unit noise vector P i-1 is an empty vector; if i is greater than 1, the unit noise vector P i-1 is obtained through the initial backward network after the (i - 1)-th step of forward diffusion; if i = M, the unit latent vector S i+1 is determined as the reconstruction latent vector.
[0147] Specifically, please also participate Figure 7 , Figure 7 is a schematic diagram of a model structure provided by an embodiment of the present application Figure 3 , as Figure 7 shown, the initial backward network can be represented as F, and the initial backward network F can also include a network structure implemented based on the attention mechanism (QKV), for example, it can be U-net or DiT. The image generation network can be represented as G, and the image generation network G can also include a network structure implemented based on the attention mechanism (QKV), for example, it can be networks such as U-net, SegNet (a deep fully convolutional neural network with an encoder-decoder architecture), or DeepLab (a dilated convolutional network for image semantic segmentation based on deep learning), etc. The text prompt can be represented as c. The computer device can obtain the unit latent vector S i in the i-th step of the time step. If i is 1, the unit latent vector S i is the above-mentioned initial latent vector z 0 . The computer device can perform the i-th step of forward diffusion on the unit latent vector S i through the initial backward network F and the text prompt c to obtain the unit noise vector P i output by the initial backward network.
[0148] The computer device can perform noise processing on the initial latent vector z 0 , the time step i corresponding to the unit noise vector P i , the unit noise vector P i and the unit noise vector P i-1 to generate the sampled noise vector Q i . The process of noise processing can be as shown in formula (5):
[0149] Q i = concat(AddNoise(z 0 , P i-1 , i), P i ) Formula (5)
[0150] Among them, the concat function represents vector concatenation along the channel dimension. The AddNoise function is a noise processing function, which can be determined by the type of the model. For example, when the image generation network is SDXL, a series of predetermined beta values can be used to add noise to the unit noise vector P based on the time step i. i Among them, the beta value can be determined by the data distribution of the initial latent vector z. 0 When the image generation network is the flux model based on rectified flow, the computer device can generate a linear function through the time step i and the data distribution of the initial latent vector z. The linear function can be expressed as a straight-line path between the data distribution and the standard normal distribution. The noise processing method in the embodiments of the present application is not limited herein. Optionally, the computer device can also directly determine the unit noise vector P as the sampled noise vector Q. 0 When i = 1, the unit noise vector P is an empty vector, for example, it can be a vector of all zeros. When i is greater than 1, the unit noise vector P is obtained through the initial reverse network after the (i - 1)-step forward diffusion. i It can be understood that adding additional noise to the unit noise vector P through noise processing can allow the reverse network to more finely control the noise level during the generation process, so that the image generation network introduces different degrees of details and variations at different stages of image generation, improving the image generation quality. This further enhances the accuracy of inversion (image inversion), encourages the initial reverse network to cooperate between inversions at different time steps, learn and use the unit noise vector obtained in the previous time step, and ensure that each stage contributes to the overall inversion quality through the cooperation method, thereby improving the accuracy and detail retention. i When i = 1, the unit noise vector P i-1 is an empty vector, for example, it can be a vector of all zeros. When i is greater than 1, the unit noise vector P i-1 is obtained through the initial reverse network after the (i - 1)-step forward diffusion.
[0151] It can be understood that adding additional noise to the unit noise vector P through noise processing can allow the reverse network to more finely control the noise level during the generation process, so that the image generation network introduces different degrees of details and variations at different stages of image generation, improving the image generation quality. This further enhances the accuracy of inversion (image inversion), encourages the initial reverse network to cooperate between inversions at different time steps, learn and use the unit noise vector obtained in the previous time step, and ensure that each stage contributes to the overall inversion quality through the cooperation method, thereby improving the accuracy and detail retention. i It can be understood that adding additional noise to the unit noise vector P through noise processing can allow the reverse network to more finely control the noise level during the generation process, so that the image generation network introduces different degrees of details and variations at different stages of image generation, improving the image generation quality. This further enhances the accuracy of inversion (image inversion), encourages the initial reverse network to cooperate between inversions at different time steps, learn and use the unit noise vector obtained in the previous time step, and ensure that each stage contributes to the overall inversion quality through the cooperation method, thereby improving the accuracy and detail retention.
[0152] The computer device can input the text prompt c and the sampled noise vector Q i into the image generation network G. Through the image generation network and the text prompt c, the sampled noise vector Q i is backward diffused to obtain the unit latent vector S i+1 . If i is less than M, the computer device can determine the unit latent vector S i+1 as the unit latent vector in the (i + 1)-th step of the time step. In the (i + 1)-th step of the time step, the unit latent vector S i+1 and the time step i + 1 are continuously input into the initial reverse network F, so as to extend the single-step structure of the initial reverse network F to multiple steps. Among them, the time step i + 1 is used to indicate the diffusion stage currently executed by the initial reverse network F.
[0153] If i = M, it represents the unit latent vector S i+1 is generated in the M-th step of the time step. The computer device can, in the M-th step of the time step, input the unit latent vector S M and the time step number M into the initial reverse network F to obtain the unit noise vector P M , and generate a sampled noise vector Q through noise processing M . Input the sampled noise vector Q M into the image generation network G, and perform backward diffusion on the sampled noise vector Q through the image generation network and the text prompt c M to obtain the reconstructed latent vector That is, when i = M, the unit latent vector S M+1 is determined as the reconstructed latent vector
[0154] For example, in the 1st step of the time step, the computer device can input the text prompt c and the unit latent vector S 1 (i.e., the initial latent vector z 0 ) into the reverse network, and perform forward diffusion on the unit latent vector S through the reverse network F and the text prompt c 1 to obtain the unit noise vector P 1 . The computer device can, based on the time step number 1, the initial latent vector z 0 and the unit noise vector P 0 (a vector of all zeros), perform noise processing on the noise vector P 1 to obtain the sampled noise vector Q 1 . The computer device can input the text prompt c and the sampled noise vector Q 1 into the graphics generation network G, and perform backward diffusion on the sampled noise vector Q through the graphics generation network G and the text prompt c 1 to obtain the unit latent vector S 2 .
[0155] In the 2nd step of the time step, the computer device can input the text prompt c and the unit latent vector S 2 into the reverse network, and perform forward diffusion on the unit latent vector S through the reverse network F and the text prompt c 2 to obtain the unit noise vector P 2 . The computer device can, based on the time step number 2, the initial latent vector z 0 and the unit noise vector P 1 (the unit noise vector output in the previous step of the current time step, the unit noise vector P i-1 ), perform noise processing on the unit noise vector P 2 to obtain the sampled noise vector Q 2。The computer device can input the text prompt c and the sampled noise vector Q 2 into the graphics generation network G, and perform backward diffusion on the sampled noise vector Q through the graphics generation network G and the text prompt c 2 to obtain the unit latent vector S 3 。
[0156] In the third step of the time step, the computer device can continue to input the unit latent vector S 3 into the inverse network until the obtained unit latent vector is the unit latent vector S M+1 At this time, the computer device can determine the unit latent vector S M+1 as the reconstructed latent vector
[0157] Step S203: Generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0158] Specifically, reference can be made to the specific content of step S103 in the corresponding embodiment above Figure 3 and details are not described herein again.
[0159] Step S204: Input the initial latent vector and the reconstructed latent vector into the discriminator network, and generate a first adversarial loss value corresponding to the inverse network and a second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector;
[0160] Specifically, the training sample data used in the embodiments of the present application can be generated by inputting the text prompt and the initial noise vector into the image generation model, thereby reducing the labor cost of data annotation. At the same time, the most accurate initial noise vector and text prompt of the generated image can be determined. Since using synthetic data for training may cause distribution shift, it may produce overly smooth results and be accompanied by loss of details. To solve this problem, the embodiments of the present application can also introduce a discriminator network (Discriminator) and propose to use adversarial learning to enhance the accuracy of image inversion.
[0161] The computer device can input the initial latent vector and the reconstructed latent vector into the discriminator network D. The discriminator network D can be a generative adversarial network (GAN). The discriminator network D can be used to distinguish the initial latent vector from the reconstructed latent vector and generate discrimination parameters for the input data. The discrimination parameters can belong to [0, 1]. The closer the discrimination parameter is to 1, it means that the discriminator network D believes that the input data is the initial latent vector. The closer the discrimination parameter is to 0, it means that the discriminator network D believes that the input data is the reconstructed latent vector. The computer device can generate the first adversarial loss value corresponding to the inverse network F through the discriminator network, the initial latent vector, and the reconstructed latent vector. and the second adversarial loss value corresponding to the discriminator network D The first adversarial loss value and the second adversarial loss value can both be adversarial losses.
[0162] The computer device can generate the first discrimination parameter for the reconstructed latent vector through the discriminator network D. Based on the first discrimination parameter, generate the first expected mean for the reconstructed latent vector. Based on the first expected mean, generate the first adversarial loss value corresponding to the inverse network. The calculation process of the first adversarial loss value can be shown as in formula (6):
[0163]
[0164] where represents the first discrimination parameter of the discriminator network D for the reconstructed latent vector , represents the data distribution that the reconstructed latent vector in a batch of input data conforms to. represents the first expected mean for the input reconstructed latent vector .
[0165] Generate the second discrimination parameter for the initial latent vector through the discriminator network. Based on the second discrimination parameter, generate the second expected mean for the initial latent vector. Based on the first expected mean and the second expected mean, generate the second adversarial loss value corresponding to the discriminator network. The calculation process of the second adversarial loss value can be shown as in formula (7):
[0166]
[0167] where [D(z 0 )] represents the second discrimination parameter of the discriminator network D for the initial latent vector z 0 , represents the initial latent vector z 0In terms of the data distribution, for the initial latent vector z 0 The second expected mean value.
[0168] It can be understood that in the embodiment of the present application, by introducing a discriminator network, the initial inverse network can learn the adversarial loss, so that it can better learn the feature representation associated with the reconstructed latent vector and the initial latent vector, thereby improving the resistance of the trained inverse network to input perturbations and enhancing the robustness.
[0169] Step S205: Generate an inverse loss value based on the image loss value, the noise loss value, and the first adversarial loss value, adjust the model parameters of the initial inverse network through the inverse loss value, and adjust the model parameters of the discriminator network through the second adversarial loss value. When both the initial inverse network and the discriminator network meet the model convergence condition, determine the initial inverse network that meets the model convergence condition as the inverse network.
[0170] Specifically, the computer device can add the image loss value L Image , the noise loss value L Noise and the first adversarial loss value to obtain the inverse loss value L full , adjust the model parameters of the initial inverse network through the inverse loss value, and adjust the model parameters of the discriminator network through the second adversarial loss value . When both the initial inverse network and the discriminator network meet the model convergence condition, determine the initial inverse network that meets the model convergence condition as the inverse network. The inverse network can also be called the Flash In model. Among them, the model convergence condition can be that during the training process, if the loss value of the loss function does not significantly decrease after multiple training batches or the training batches of the initial inverse network reach a preset maximum value, etc. The embodiment of the present application does not limit the model convergence condition here.
[0171] In the embodiments of the present application, an initial reverse network is used to parameterize the image inversion process of converting a latent vector into a noise vector. By directly capturing the mapping relationship between the latent vector and the noise vector, the mapping of data from one domain (latent vector) to another domain (noise vector) is learned, so that the number of time steps used in the forward diffusion process of the initial reverse network can be single-step or few steps, reducing the number of time steps used in forward diffusion and improving the efficiency of image generation. By calculating the noise loss value between the initial noise vector and the reconstructed noise vector, and the image loss value between the reconstructed latent vector and the initial latent vector, an explicit training objective is defined for training the initial reverse network, enabling the initial reverse network to learn to match the reconstructed noise vector with the initial noise vector and the reconstructed latent vector with the initial latent vector, ensuring the consistency of the image inversion process of converting the latent vector into the noise vector, that is, the initial latent vector can be inversely transformed into the reconstructed noise vector through image inversion, and the reconstructed noise still has consistency with the initial noise vector, thereby improving the accuracy of image inversion by the trained reverse network. Therefore, the noise vector generated by the trained reverse network has higher precision and can more accurately reconstruct the features of the input image. The reverse network can be used to more accurately convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and then the image generation network can generate a second image associated with the first image through the service noise vector, improving the efficiency and accuracy of image generation.
[0172] On the other hand, the reverse network of the embodiments of the present application can be applied to single-step or few steps. Since the reverse network proposed in the embodiments of the present application can reduce the error generated in the diffusion process, when the reverse network is applied to at least a few steps, the accuracy of the diffusion process can be further improved by increasing the sampling process in the diffusion process. By introducing a discriminator network, the initial reverse network can learn the adversarial loss, so that it can better learn the feature representation associated with the reconstructed latent vector and the initial latent vector, thereby improving the resistance of the trained reverse network to input perturbations and enhancing the robustness.
[0173] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data processing device provided by the embodiments of the present application. Figure 1 As Figure 8 shown, the data processing device 1 includes a backward diffusion module 810, an image inversion module 820, a loss calculation module 830, and a model adjustment module 840.
[0174] The backward diffusion module 810 is configured to obtain a text prompt and an initial noise vector, and perform backward diffusion on the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector;
[0175] The image inversion module 820 is used to input the initial latent vector and the text prompt into the initial reverse network, perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector;
[0176] The loss calculation module 830 is used to generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0177] The model adjustment module 840 is used to generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is used to convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and the service noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0178] In a possible implementation manner, the time step corresponding to the initial reverse network is 1; when the image inversion module 820 is used to perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector, it is specifically used to perform the following operations:
[0179] Input the initial latent vector and the text prompt into the initial reverse network, perform single-step forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector.
[0180] In a possible implementation manner, the time step corresponding to the initial reverse network is M, where M is a positive integer greater than 1; when the image inversion module 820 is used to input the initial latent vector and the text prompt into the initial reverse network, perform single-step forward diffusion on the initial latent vector through the initial reverse network to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector, it is specifically used to perform the following operations:
[0181] Obtain the unit latent vector S in the i-th step at the time step i , and input the unit latent vector S i and the text prompt into the initial reverse network, and perform the i-th step forward diffusion on the unit latent vector S i through the initial reverse network and the text prompt to obtain the unit noise vector P output by the initial reverse network i ; if i is 1, then the unit latent vector Si is the initial latent vector; i is a positive integer less than or equal to M;
[0182] According to the initial latent vector, the unit noise vector P i corresponding time step i, the unit noise vector P i and the unit noise vector P i-1 , generate the sampling noise vector Q i , perform backward diffusion on the sampling noise vector Q through the image generation network and the text prompt i to obtain the unit latent vector S i+1 ; if i is less than M, the unit latent vector S i+1 is the unit latent vector in the (i + 1)-th step of the time step; if i is 1, the unit noise vector P i-1 is an empty vector; if i is greater than 1, the unit noise vector P i-1 is obtained through the initial reverse network after forward diffusion in the (i - 1)-th step;
[0183] If i = M, determine the unit latent vector S i+1 as the reconstructed latent vector.
[0184] In a possible implementation, the model adjustment module 840 is used to generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value. When obtaining the reverse network, it is specifically used to perform the following operations:
[0185] Input the initial latent vector and the reconstructed latent vector into the discriminator network, and generate the first adversarial loss value corresponding to the reverse network and the second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector;
[0186] Generate a reverse loss value based on the image loss value, the noise loss value, and the first adversarial loss value, adjust the model parameters of the initial reverse network through the reverse loss value, and adjust the model parameters of the discriminator network through the second adversarial loss value. When both the initial reverse network and the discriminator network meet the model convergence condition, determine the initial reverse network that meets the model convergence condition as the reverse network.
[0187] In a possible implementation, when the model adjustment module 840 is used to generate the first adversarial loss value corresponding to the reverse network and the second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector, and the reconstructed latent vector, it is specifically used to perform the following operations:
[0188] Generate the first discrimination parameter for the reconstructed latent vector through the discriminator network. Based on the first discrimination parameter, generate the first expected mean for the reconstructed latent vector, and generate the first adversarial loss value corresponding to the inverse network based on the first expected mean;
[0189] Generate the second discrimination parameter for the initial latent vector through the discriminator network. Based on the second discrimination parameter, generate the second expected mean for the initial latent vector, and generate the second adversarial loss value corresponding to the discriminator network based on the first expected mean and the second expected mean.
[0190] In a possible implementation, the noise loss value includes a noise reconstruction loss value and a regularization loss value; when the loss calculation module 830 is used to generate the noise loss value based on the initial noise vector and the reconstructed noise vector, it is specifically used to perform the following operations:
[0191] Generate the regularization loss value based on the noise probability distribution of the initial noise vector and the noise probability distribution of the reconstructed noise vector;
[0192] Determine the mean square error between the initial noise vector and the reconstructed noise vector as the noise reconstruction loss value, and determine the sum of the regularization loss value and the noise reconstruction loss value as the noise loss value.
[0193] In a possible implementation, when the loss calculation module 830 is used to generate the image loss value based on the reconstructed latent vector and the initial latent vector, it is specifically used to perform the following operations:
[0194] Determine the mean square error between the initial latent vector and the reconstructed latent vector as the image reconstruction loss value, obtain the service hyperparameter for balancing the noise reconstruction loss value and the image reconstruction loss value, and perform a dot product operation on the service hyperparameter and the image reconstruction loss value to obtain the image loss value.
[0195] In a possible implementation, the image generation network includes a basic generation network and a refinement generation network; when the backward diffusion module 810 is used to obtain the text prompt and the initial noise vector and perform backward diffusion on the initial noise vector through the image generation network and the text prompt to obtain the initial latent vector, it is specifically used to perform the following operations:
[0196] Obtain the text prompt and the initial noise image, perform feature extraction on the text prompt to obtain the text prompt vector, and perform image encoding on the initial noise image to obtain the initial noise vector;
[0197] Input the initial noise vector into the basic generation network. In the basic generation network, perform backward diffusion on the initial noise vector through the text prompt vector to obtain the basic latent vector;
[0198] Input the base latent vector into the refinement generation network. In the refinement generation network, perform backward diffusion on the base latent vector through the text prompt vector to obtain the initial latent vector.
[0199] In a possible implementation, the initial inverse network includes a downsampling network and an upsampling network. The input of the upsampling network is the output of the downsampling network layer. When the image inversion module 820 is used to perform forward diffusion on the initial latent vector through the initial inverse network and the text prompt to obtain the reconstructed noise vector, it is specifically used to perform the following operations:
[0200] In the downsampling network of the initial inverse network, perform self-attention processing on the initial latent vector to obtain the downsampling attention vector, and perform cross-attention processing on the downsampling attention vector and the text prompt to obtain the downsampling feature vector; the feature dimension of the downsampling feature vector is smaller than the feature dimension of the initial latent vector.
[0201] In the upsampling network of the initial inverse network, perform self-attention processing on the downsampling feature vector to obtain the upsampling attention vector, and perform cross-attention processing on the upsampling attention vector and the text prompt to obtain the reconstructed noise vector; the feature dimension of the reconstructed noise vector is equal to the feature dimension of the initial latent vector.
[0202] In a possible implementation, when the image inversion module 820 is used to perform cross-attention processing on the downsampling attention vector and the text prompt to obtain the downsampling feature vector, it is specifically used to perform the following operations:
[0203] Perform a dot product operation on the downsampling attention vector and the query parameter matrix to obtain the query vector, perform a dot product operation on the downsampling attention vector and the key parameter matrix to obtain the key vector, and perform a dot product operation on the text prompt vector corresponding to the text prompt and the value parameter matrix to obtain the value vector;
[0204] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the dimension of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain the normalized vector, and perform a dot product operation on the normalized vector and the value vector to obtain the downsampling feature vector.
[0205] In the embodiments of the present application, an initial inverse network is used to parameterize the image inversion process of converting a latent vector into a noise vector. By directly capturing the mapping relationship between the latent vector and the noise vector, the mapping of data from one domain (latent vector) to another domain (noise vector) is learned, so that the number of time steps used in the forward diffusion process of the initial inverse network can be a single step or a small number of steps, reducing the number of time steps used in forward diffusion and improving the efficiency of image generation. By calculating the noise loss value between the initial noise vector and the reconstructed noise vector, and the image loss value between the reconstructed latent vector and the initial latent vector, an explicit training objective is defined for training the initial inverse network, enabling the initial inverse network to learn to match the reconstructed noise vector with the initial noise vector and the reconstructed latent vector with the initial latent vector, ensuring the consistency of the image inversion process of converting the latent vector into the noise vector, that is, the initial latent vector can be inversely transformed into a reconstructed noise vector through image inversion, and this reconstructed noise still has consistency with the initial noise vector, thereby improving the accuracy of image inversion by the trained inverse network. Therefore, the noise vector generated by the trained inverse network has higher precision and can more accurately reconstruct the features of the input image. The inverse network can be used to more accurately convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and then the image generation network can generate a second image associated with the first image through the service noise vector, improving the efficiency and accuracy of image generation.
[0206] On the other hand, the inverse network of the embodiments of the present application can be applied to a single step or a small number of steps because the inverse network proposed in the embodiments of the present application can reduce the error generated in the diffusion process. Therefore, when the inverse network is applied to at least a small number of steps, the accuracy of the diffusion process can be further improved by increasing the sampling process in the diffusion process. By introducing a discriminator network, the initial inverse network can learn the adversarial loss, so that it can better learn the feature representation associated with the reconstructed latent vector and the initial latent vector, thereby improving the resistance of the trained inverse network to input perturbations and enhancing the robustness.
[0207] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0208] Please refer to Figure 9 , Figure 9It is a structural schematic diagram of a data processing device provided by an embodiment of the present application Figure 2 . As Figure 9 shown, the data processing device 2 includes an input processing module 910, a conversion processing module 920, a noise prediction module 930, and an image generation module 940.
[0209] The input processing module 910 is configured to obtain a first image and a target text prompt input by an object, perform image recognition on the first image to obtain an image recognition result, generate a service text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method provided by the embodiment of the present application;
[0210] The conversion processing module 920 is configured to convert the first latent vector into a service noise vector through the reverse network and the service text prompt;
[0211] The noise prediction module 930 is configured to input the service noise vector and the target text prompt into an image generation network, and in the image generation network, perform backward diffusion on the service noise vector through the target text prompt to obtain a second latent vector;
[0212] The image generation module 940 is configured to perform image decoding on the second latent vector to obtain a second image that matches the target text prompt.
[0213] The reverse network in the embodiment of the present application can be applied to image generation scenarios that rely on image inversion, such as image-to-image scenarios, image editing scenarios, etc. In the image editing scenario, the target text prompt can represent the image editing process required for the first image, so as to conveniently and easily implement advanced image editing, such as changing the main body in the image content, adjusting the image style, etc. At the same time, the reverse network can be obtained by learning the mapping relationship between the latent vector and the noise vector, so that the number of time steps used in the forward diffusion process can be single-step or few-step, reducing the number of time steps used in the forward diffusion and improving the efficiency of image generation or graphic editing.
[0214] On the other hand, in the process of obtaining the inverse network by training the initial inverse network, by calculating the noise loss value between the initial noise vector and the reconstructed noise vector, and the image loss value between the reconstructed latent vector and the initial latent vector, an explicit training objective is defined for training the initial inverse network, enabling the initial inverse network to learn to match the reconstructed noise vector with the initial noise vector and the reconstructed latent vector with the initial latent vector, ensuring the consistency of the image inversion process of converting the latent vector into the noise vector, that is, the initial latent vector can be inversely transformed into the reconstructed noise vector through image inversion, and this reconstructed noise still has consistency with the initial noise vector, thereby improving the accuracy of the image inversion performed by the trained inverse network. The inverse network can be used to more accurately convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and then enable the image generation network to generate a second image associated with the first image through the service noise vector, improving the efficiency and accuracy of image generation.
[0215] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0216] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 10 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, for example, at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As Figure 10As shown in the figure, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0217] In the computer device 1000 as shown in Figure 10 the figure, the network interface 1004 can provide network communication network elements; the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005.
[0218] When the computer device 1000 is used to execute the data processing device 1, it is to achieve:
[0219] Obtain a text prompt and an initial noise vector, perform backward diffusion on the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector;
[0220] Input the initial latent vector and the text prompt into an initial reverse network, perform forward diffusion on the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and perform backward diffusion on the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector;
[0221] Generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and generate an image loss value based on the reconstructed latent vector and the initial latent vector;
[0222] Generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is used to convert the latent vector representing the first image into a service noise vector through the service text prompt corresponding to the first image, and the service noise vector is used to instruct the image generation network to generate a second image associated with the first image.
[0223] When the computer device 1000 is used to execute the data processing device 1, it is to achieve:
[0224] Obtain the first image input by the object and the target text prompt, perform image recognition on the first image to obtain an image recognition result, generate a service text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method provided in the embodiments of the present application;
[0225] Convert the first latent vector into a service noise vector through the reverse network and the service text prompt;
[0226] Input the service noise vector and the target text prompt into the image generation network, and in the image generation network, perform backward diffusion on the service noise vector through the target text prompt to obtain a second latent vector;
[0227] Perform image decoding on the second latent vector to obtain a second image that matches the target text prompt.
[0228] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the descriptions of the data processing method in any of the foregoing Figure 3 and Figure 5 corresponding embodiments, which will not be elaborated herein. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0229] In addition, it should be noted here that: the embodiments of the present application also provide a computer-readable storage medium, and a computer program is stored in the above computer-readable storage medium. When the above processor executes the above computer program, it can execute the descriptions of the above data processing method in any of the foregoing Figure 3 and Figure 5 corresponding embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0230] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been displayed or will be displayed.
[0231] In addition, it should be noted here that: the embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in any of the foregoing Figure 3 and Figure 5 corresponding embodiments.
[0232] In the description, claims, and drawings of the embodiments of this application, the terms "first", "second", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products, or equipment.
[0233] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of network elements in the above description. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described network elements for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0234] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0235] The steps in the method of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0236] The modules in the device of the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0237] The foregoing disclosure is only for the preferred embodiments of the present application, and of course cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that: include: Obtaining a text prompt and an initial noise vector, and back-diffusing the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector; Inputting the initial latent vector and the text prompt into an initial reverse network, forward diffusing the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and backward diffusing the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector; generating a noise loss value based on the initial noise vector and the reconstructed noise vector, and generating an image loss value based on the reconstructed latent vector and the initial latent vector; Based on the image loss value and the noise loss value, a reverse loss value is generated, and the model parameters of the initial reverse network are adjusted by the reverse loss value to obtain a reverse network; the reverse network is used to convert the potential vector representing the first image into a business noise vector through the business text prompt corresponding to the first image, and the business noise vector is used to instruct the image generation network to generate a second image associated with the first image.
2. The method according to claim 1, characterized in that The time step corresponding to the initial reverse network is 1, the initial potential vector is forward diffused through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and the reconstructed noise vector is backward diffused through the image generation network and the text prompt to obtain a reconstructed potential vector, including: The initial latent vector and the text prompt are input into the initial reverse network, the initial latent vector is forward diffused in a single step through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and the reconstructed noise vector is backward diffused through the image generation network and the text prompt to obtain a reconstructed latent vector.
3. The method according to claim 1, characterized in that The time step corresponding to the initial reverse network is M, where M is a positive integer greater than 1; the initial potential vector and the text prompt are input into the initial reverse network, the initial potential vector is forward diffused in a single step through the initial reverse network to obtain a reconstructed noise vector, and the reconstructed noise vector is backward diffused through the image generation network and the text prompt to obtain a reconstructed potential vector, including: Get the unit latent vector S at the i-th time step i , the unit potential vector S i The text prompt is input into the initial reverse network, and the unit potential vector S is converted by the initial reverse network and the text prompt. i Perform the i-th step forward diffusion to obtain the unit noise vector P output by the initial reverse network i ; If i is 1, then the unit potential vector S i is the initial potential vector; i is a positive integer less than or equal to M; According to the initial potential vector, the unit noise vector P i The corresponding time step number i, the unit noise vector P i and the unit noise vector P i-1 , generating the sampling noise vector Q i , the sampled noise vector Q is generated by the image generation network and the text prompt i Perform back diffusion to obtain the unit potential vector S i+1 ; If the i is less than M, then the unit potential vector S i+1 is the unit potential vector in the i+1th step of the time step; if i is 1, then the unit noise vector P i-1 is an empty vector; if i is greater than 1, the unit noise vector P i-1 is obtained after the initial reverse network undergoes the i-1th step of forward diffusion; If i=M, the unit latent vector S i+1 Determined as the reconstructed latent vector.
4. The method according to claim 1, characterized in that The method of generating a reverse loss value based on the image loss value and the noise loss value, and adjusting the model parameters of the initial reverse network by using the reverse loss value to obtain a reverse network includes: Inputting the initial latent vector and the reconstructed latent vector into a discriminator network, and generating a first adversarial loss value corresponding to the reverse network and a second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial latent vector and the reconstructed latent vector; Based on the image loss value, the noise loss value and the first adversarial loss value, a reverse loss value is generated, the model parameters of the initial reverse network are adjusted by the reverse loss value, and the model parameters of the discriminator network are adjusted by the second adversarial loss value. When the initial reverse network and the discriminator network both meet the model convergence conditions, the initial reverse network that meets the model convergence conditions is determined as the reverse network.
5. The method according to claim 4, characterized in that The step of generating a first adversarial loss value corresponding to the reverse network and a second adversarial loss value corresponding to the discriminator network through the discriminator network, the initial potential vector and the reconstructed potential vector comprises: Generate a first discriminant parameter for the reconstructed latent vector through a discriminator network, generate a first expected mean for the reconstructed latent vector based on the first discriminant parameter, and generate a first adversarial loss value corresponding to the reverse network based on the first expected mean; A second discriminant parameter for the initial latent vector is generated through the discriminator network, a second expected mean for the initial latent vector is generated based on the second discriminant parameter, and a second adversarial loss value corresponding to the discriminator network is generated based on the first expected mean and the second expected mean.
6. The method according to claim 1, characterized in that The noise loss value includes a noise reconstruction loss value and a regularization loss value; the generating the noise loss value based on the initial noise vector and the reconstructed noise vector includes: Generating the regularization loss value based on the noise probability distribution of the initial noise vector and the noise probability distribution of the reconstructed noise vector; The mean square error between the initial noise vector and the reconstructed noise vector is determined as the noise reconstruction loss value, and the sum of the regularization loss value and the noise reconstruction loss value is determined as the noise loss value.
7. The method according to claim 6, characterized in that The generating an image loss value based on the reconstructed latent vector and the initial latent vector comprises: The mean square error between the initial latent vector and the reconstructed latent vector is determined as the image reconstruction loss value, a business hyperparameter for balancing the noise reconstruction loss value and the image reconstruction loss value is obtained, and a dot multiplication operation is performed on the business hyperparameter and the image reconstruction loss value to obtain an image loss value.
8. The method according to claim 1, characterized in that The image generation network includes a basic generation network and a refined generation network; the acquiring of text prompts and an initial noise vector, and back-diffusion of the initial noise vector through the image generation network and the text prompt to obtain an initial potential vector include: Acquire a text prompt and an initial noise image, perform feature extraction on the text prompt to obtain a text prompt vector, perform image encoding on the initial noise image to obtain an initial noise vector; Inputting the initial noise vector into the base generation network, in which the initial noise vector is back-diffused through the text prompt vector to obtain a base latent vector; The basic latent vector is input into the refined generation network, and in the refined generation network, the basic latent vector is back-diffused through the text prompt vector to obtain an initial latent vector.
9. The method according to claim 8, characterized in that The initial reverse network includes a downsampling network and an upsampling network, and the input of the upsampling network is the output of the downsampling network layer; the forward diffusion of the initial potential vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector includes: In the downsampling network of the initial reverse network, the initial latent vector is subjected to self-attention processing to obtain a downsampled attention vector, and the downsampled attention vector and the text prompt are subjected to cross-attention processing to obtain a downsampled feature vector; the feature dimension of the downsampled feature vector is smaller than the feature dimension of the initial latent vector; In the upsampling network of the initial reverse network, the downsampled feature vector is self-attention processed to obtain an upsampled attention vector, and the upsampled attention vector and the text prompt are cross-attention processed to obtain a reconstructed noise vector; the feature dimension of the reconstructed noise vector is equal to the feature dimension of the initial potential vector.
10. The method according to claim 9, characterized in that The cross-attention processing is performed on the down-sampled attention vector and the text prompt to obtain a down-sampled feature vector, including: Performing a dot multiplication operation on the downsampled attention vector and the query parameter matrix to obtain a query vector, performing a dot multiplication operation on the downsampled attention vector and the key parameter matrix to obtain a key vector, and performing a dot multiplication operation on the text prompt vector corresponding to the text prompt and the value parameter matrix to obtain a value vector; An attention score vector is generated based on the query vector and the key vector, the attention score vector is dimensionally reduced based on the dimension of the key vector, the reduced-dimensional attention score vector is normalized to obtain a normalized vector, and a point multiplication operation is performed on the normalized vector and the value vector to obtain a downsampled feature vector.
11. A data processing method, characterized in that: include: Acquire a first image and a target text prompt input by the object, perform image recognition on the first image to obtain an image recognition result, generate a business text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method according to any one of claims 1 to 10; Converting the first potential vector into a business noise vector through the reverse network and the business text prompt; Inputting the business noise vector and the target text prompt into the image generation network, and in the image generation network, back-diffusion is performed on the business noise vector through the target text prompt to obtain a second latent vector; The second latent vector is image decoded to obtain a second image matching the target text prompt.
12. A data processing device, characterized in that: include: A back diffusion module, used to obtain a text prompt and an initial noise vector, and back diffuse the initial noise vector through an image generation network and the text prompt to obtain an initial latent vector; An image inversion module, configured to input the initial latent vector and the text prompt into an initial reverse network, forward diffuse the initial latent vector through the initial reverse network and the text prompt to obtain a reconstructed noise vector, and backward diffuse the reconstructed noise vector through the image generation network and the text prompt to obtain a reconstructed latent vector; A loss calculation module, configured to generate a noise loss value based on the initial noise vector and the reconstructed noise vector, and to generate an image loss value based on the reconstructed latent vector and the initial latent vector; A model adjustment module is used to generate a reverse loss value based on the image loss value and the noise loss value, and adjust the model parameters of the initial reverse network through the reverse loss value to obtain a reverse network; the reverse network is used to convert the potential vector representing the first image into a business noise vector through the business text prompt corresponding to the first image, and the business noise vector is used to instruct the image generation network to generate a second image associated with the first image.
13. A data processing device, characterized in that: include: An input processing module, used to obtain a first image and a target text prompt of an object input, perform image recognition on the first image to obtain an image recognition result, generate a business text prompt based on the image recognition result, perform image encoding on the first image to obtain a first latent vector; the reverse network is trained by the method according to any one of claims 1 to 10; A conversion processing module, used for converting the first potential vector into a business noise vector through the reverse network and the business text prompt; A noise prediction module, configured to input the business noise vector and the target text prompt into the image generation network, and in the image generation network, back-diffuse the business noise vector through the target text prompt to obtain a second latent vector; An image generation module is used to perform image decoding on the second latent vector to obtain a second image matching the target text prompt.
14. A computer device, characterized in that: include: Processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method described in any one of claims 1-11.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 11.
16. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer-readable storage medium and is suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 11.