Image processing method and device and electronic equipment
By directly using text description information in the image processing model for image iterative processing, the image detail loss problem caused by encoding and decoding is solved, and more efficient image processing effect and cross-model applicability are achieved.
Patent Information
- Application Number
- CN202510192979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the image processing model causes image details loss through encoding and decoding operations, affecting the processing effect.
The pre-trained image processing model is used to determine the image change trend parameters based on the initial and target text description information, and iteratively process iteratively in the image space to generate the target image, avoiding the encoding and decoding process.
It reduces image detail loss, improves image processing effect, and has good cross-model applicability and efficient transformation process.
Smart Images

Figure CN120355818A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly, to an image processing method, apparatus, and electronic device. Background Art
[0002] In the related art, an image can be modified according to a text description by an image processing model established based on artificial intelligence. The image processing model needs to input an image encoding into an intermediate feature space such as a latent space, and then generate new image content according to the target text description. However, the operation mode of first encoding and then decoding the image in this way usually causes loss of image details, resulting in poor image processing effects. Summary of the Invention
[0003] In view of this, an object of the present disclosure is to provide an image processing method, apparatus, and electronic device to reduce the loss of image details and improve the image processing effect.
[0004] In a first aspect, an embodiment of the present disclosure provides an image processing method, the method including: obtaining an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; performing a first image processing on the initial image based on the initial text description information by a pre-trained image processing model to determine a first image change trend parameter, and performing a second image processing on the initial image based on the target text description information by the image processing model to determine a second image change trend parameter; wherein, the first image change trend parameter is used to indicate: in the process of the first image processing, the change trend of the initial image; the second image change trend parameter is used to indicate: in the process of the second image processing, the change trend of the initial image; processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0005] Second aspect, embodiments of the present disclosure provide an image processing apparatus, the apparatus comprising: an initial image acquisition module, configured to acquire an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; a parameter determination module, configured to perform a first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform a second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; wherein, the first image change trend parameter is used to indicate: the change trend of the initial image during the first image processing; the second image change trend parameter is used to indicate: the change trend of the initial image during the second image processing; an image generation module, configured to process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0006] Third aspect, embodiments of the present invention provide an electronic device, comprising a processor and a memory, the memory storing machine-executable instructions capable of being executed by the processor, and the processor executing the machine-executable instructions to implement the above-mentioned image processing method.
[0007] Fourth aspect, embodiments of the present invention provide a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above-mentioned image processing method.
[0008] Embodiments of the present invention bring the following beneficial effects:
[0009] The above-mentioned image processing method, apparatus and electronic device acquire an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; perform a first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform a second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information. This method does not require encoding and decoding processing of the image, reduces the loss of image details, and improves the image processing effect.
[0010] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0011] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure;
[0014] Figure 2 A schematic diagram of the structure of an image processing device provided by an embodiment of the present disclosure;
[0015] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0017] With the development of artificial intelligence (AI) technology, the demand for "letting computers generate or modify images based on text descriptions" is growing, especially in industries such as marketing, design, and entertainment. For example, a company may want to use a sentence to describe "change the product image from a white background to a blue background, and add some decorative elements" to automatically generate or modify the corresponding product poster; or use a sentence to describe "change a photo of a kitten into a photo that looks like a dog" and preserve the clarity and details of the original image as much as possible in the process.
[0018] Techniques for manipulating image content using natural language descriptions are becoming increasingly mature. Common methods rely on encoding the input image into an intermediate feature space (such as a latent space), and then generating new image content based on the target text description. It can also be described as: first converting the image into a certain "intermediate representation" inside the model, simply understood as the "hidden image" that the model sees for itself, and then converting it back to the final image. The above "encoding first and then decoding" process is also called the "retrospective" editing method. This approach usually has the following problems:
[0019] 1. Detail loss: Each time encoding and decoding are performed, details of the original image may be lost or distorted.
[0020] 2. Dependence on the structure of a specific model. If a new image generation model is used, the previous "intermediate representation" may not be directly usable.
[0021] 3. High computational cost: The processes of encoding and decoding require a large amount of computational resources and time.
[0022] Based on this, an image processing method, apparatus, and electronic device provided in the embodiments of the present disclosure can be applied to scenarios for image processing.
[0023] See Figure 1 , first, an image processing method provided in the embodiments of the present invention will be introduced. The method includes the following steps:
[0024] Step S102, obtain an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image.
[0025] The above initial image can be captured by a camera, generated by a drawing software, or obtained by other means, which is not limited herein. The above initial text description information can be generated manually or by artificial intelligence, such as generated by an image recognition model. The initial text description information needs to accurately describe the display content in the initial image. Specifically, it can describe the types of objects displayed in the initial image, such as people, certain animals, buildings, etc., and can also describe the color of the object, the atmosphere generated by the object, etc., which is not limited herein.
[0026] The above target text description information is usually used to describe the display content of the target image to be generated based on the initial image. Generally speaking, the target text description information is different from the initial text description information. The target text description information needs to highlight the content that is different between the target image and the initial image. For example, if a ship is shown in the initial image, and the target image to be generated needs to show a panda instead of a ship, then the target text description information can be "a panda", or "a panda in a bamboo forest", etc., without limitation here.
[0027] Step S104, perform first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine the first image change trend parameter, and perform second image processing on the initial image based on the target text description information through the image processing model to determine the second image change trend parameter; wherein, the first image change trend parameter is used to indicate: during the first image processing, the change trend of the initial image; the second image change trend parameter is used to indicate: during the second image processing, the change trend of the initial image.
[0028] The above image processing model is usually a pre-trained diffusion model or flow model. The image processing model can perform image processing on the input image based on the input text description information, so that the processed image conforms to the input text description information.
[0029] After the image processing model performs first image processing on the initial image based on the initial text description information, the type of the object shown in the processed image usually does not change compared with the initial image. The specific display content may or may not change, which is related to the specific model used. When the image processing model performs first image processing on the initial image, it is necessary to maintain the content such as the type of the object shown in the initial image described in the initial text description information. For example, when the initial description text information describes the object in the initial image as a hot air balloon, when the image processing model performs first image processing on the initial image, it is necessary to change the content of the initial image from the currently shown hot air balloon to a hot air balloon. The hot air balloons before and after the change may have different displayed details.
[0030] After the image processing model performs second image processing on the initial image based on the target text description information, the details such as the type of the object shown in the processed image need to conform to the target text description information. For example, if the initial image shows a hot air balloon, and the target text description information describes an airplane, the image processing model needs to perform second image processing on the initial image to change the hot air balloon shown in the initial image to an airplane.
[0031] The process of the image processing model performing the first image processing on the initial image based on the initial text description information is similar to the process of performing the second image processing on the initial image based on the target text description information. Usually, it is necessary to perform image processing on the initial image iteratively, that is, perform multiple image processings. The direction of each image processing can usually be determined from the data generated during the processing by the image processing model. The direction of image processing can also be called the trend of image processing, and can be specifically represented by an image change trend parameter.
[0032] Different image processing models will generate different data during the image processing, and the corresponding determined image change trend parameters are also different. For example, the negative gradient of the noise residual generated during the image processing by the diffusion model can be used as the image change trend parameter; for the flow model, the parameters output by the forward network in the model can be directly obtained, and the change data of the parameters at different iteration times can be determined, so that these changes are used as the image change trend parameters, and the image change trend parameters can also be determined by calculating the gradient of the log-likelihood.
[0033] For the convenience of writing, the image change trend parameter determined during the process of the image processing model performing image processing on the initial image based on the initial text description information is called the first image change trend parameter, and the image change trend parameter determined during the process of the image processing model performing image processing on the initial image based on the target text description information is called the second image change trend parameter.
[0034] Step S106, process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0035] The first image change trend parameter represents the image change trend when the initial image maintains the current display content, and the second image change trend parameter represents the image change trend when the initial image changes to the display content indicated by the target text description information. The difference between the two can be regarded as the image change trend when the initial image changes from the current display content to the display content indicated by the target text description information.
[0036] Since the image processing process is usually an iterative process, the first image change trend parameter and the first image change trend parameter corresponding to each time step in the iterative process can be subtracted from each other, and then the obtained result is superimposed on the initial image, so that an iterative process can be completed without relying on the image processing model. When the iteration is completed, the target image corresponding to the target text description information can be obtained.
[0037] The above-mentioned image processing method includes obtaining an initial image and target text description information; the initial image has corresponding initial text description information, which is used to describe the display content in the initial image; performing a first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and performing a second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information. This method does not require encoding and decoding processing of the image, reduces the loss of image details, and improves the image processing effect.
[0038] The following embodiments provide a specific method for performing a first image processing on an initial image based on initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and performing a second image processing on the initial image based on target text description information through the image processing model to determine a second image change trend parameter.
[0039] When the image processing model is a diffusion model, the way the image processing model processes images is usually as follows: First, noise is gradually added to the initial image until it finally becomes pure Gaussian noise. This process is called the forward diffusion process. The model learns a reverse denoising process, that is, starting from pure noise, gradually removing the noise, and finally generating a high-quality image. This denoising process is iterative, and each step depends on a noise prediction network (ε-predictor). At each step of the denoising process, the noise prediction network inputs the current noisy image and text description, and predicts the noise residual contained in the image. This noise residual can be understood as the "extra" noise component that the model believes is in the current image and needs to be removed to be closer to the true image distribution. Further, the negative value of the noise residual can be regarded as the image change trend parameter.
[0040] When determining the first image change trend parameter, the following operations can be performed through the diffusion model: generating a first noise map corresponding to the initial image, then determining a first predicted noise residual based on the initial text description information and the first noise map, and performing a first denoising process on the first noise map based on the first predicted noise residual; and then determining the first image change trend parameter based on the first predicted noise residual.
[0041] Denoising processing is usually an iterative process. Specifically, the first denoising processing includes first sub-operations performed at multiple time steps. Each first sub-operation has a corresponding first predicted noise residual. For each first predicted noise residual, it is necessary to determine the first negative gradient corresponding to the first predicted noise residual, which can specifically be the negative value of the first predicted noise residual, or the output value of the scoring function when determining the first predicted noise; then arrange the multiple first negative gradients in the execution order of the corresponding first denoising processing to obtain the first image change trend parameter;
[0042] When determining the second image change trend parameter, the following operations are performed through the diffusion model: generating a second noise map corresponding to the initial image, determining the second predicted noise residual based on the target text description information and the first noise map, and performing second denoising processing on the second noise map based on the second predicted noise residual; determining the second image change trend parameter based on the second predicted noise residual.
[0043] Similarly, the second denoising processing also includes second sub-operations performed at multiple time steps. Each second sub-operation has a corresponding second predicted noise residual. For each second predicted noise residual, it is necessary to determine the second negative gradient corresponding to the second predicted noise residual, which can specifically be the negative value of the second predicted noise parameter, or the output value of the scoring function when determining the second predicted noise; arrange the multiple second negative gradients in the execution order of the corresponding second denoising processing to obtain the second image change trend parameter.
[0044] When the image processing model is a flow model, since the flow model includes a forward network, the first output parameter of the forward network can be obtained during the first image processing of the initial image by the flow model based on the initial text description information. When the forward network includes multiple layers, the output parameters of each layer can be obtained and accumulated to get the current first output parameter; then based on the first output parameter, the first image change trend parameter is determined. Generally speaking, the first output parameter of the forward network can be obtained in each iteration process (corresponding to multiple time steps) during multiple iteration processes, so the first image change trend parameter includes the first output parameters arranged according to the time steps. Similarly, the second output parameter of the forward network can be obtained during the second image processing of the initial image by the flow model based on the target text description information; and based on the second output parameter, the second image change trend parameter is determined.
[0045] The following embodiments provide a specific method for processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate the target image corresponding to the target text description information.
[0046] In practical applications, it is necessary to determine the target trend change parameter based on the first image change trend parameter and the second image change trend parameter; usually, it is obtained by taking the difference between the second image change trend parameter and the first image change trend parameter. Then, the initial image is processed based on the target trend change parameter to generate the target image corresponding to the target text description information.
[0047] In specific implementation, the first image change trend parameter includes a plurality of first sub-parameters arranged in sequence according to the corresponding time steps. Suppose 10 time steps are used in the first image processing process, then the first image change trend parameter has a first sub-parameter corresponding to the first time step, a second sub-parameter corresponding to the second time step,..., and a first sub-parameter corresponding to the tenth time step. Similarly, the second image change trend parameter usually includes a plurality of second sub-parameters arranged in sequence according to the corresponding time steps. The number of second sub-parameters in the second image change trend parameter may be the same as or different from the number of first sub-parameters in the first image change trend parameter. Correspondingly, the number of time steps used in the second image processing process is the same as or different from the number of time steps used in the first image processing process.
[0048] When the number of second sub-parameters is the same as the number of first sub-parameters, for each first sub-parameter, the second sub-parameter corresponding to the first sub-parameter can be determined; among them, the time step corresponding to the first sub-parameter and the corresponding second sub-parameter is the same; then, the difference between the second sub-parameter and the first sub-parameter is determined as the target sub-parameter corresponding to the time step corresponding to the first sub-parameter; furthermore, based on the time steps corresponding to the multiple target sub-parameters, the multiple target sub-parameters are arranged to obtain the target trend change parameter.
[0049] When the number of second sub-parameters is different from the number of first sub-parameters, the corresponding curves can be respectively fitted based on the first sub-parameters and the second sub-parameters. Suppose the curve corresponding to the first sub-parameters is the first curve, and the curve corresponding to the second sub-parameters is the second curve. N points can be sampled from the first curve, N points can be sampled from the second curve, and the corresponding relationship of the sampling points can be determined. The sampling value of the point belonging to the second curve in the corresponding sampling points is subtracted from the sampling value of the point belonging to the first curve, and the obtained result can be used as the target sub-parameter. The target sub-parameters are arranged according to the corresponding sampling order to obtain the target trend change parameter.
[0050] After determining the target trend change parameter, the first target sub-parameter among multiple target sub-parameters can be determined as the current sub-parameter; and based on the product of the current sub-parameter and the evolution step size, the image state increment is determined; where the evolution step size is the time interval between adjacent time steps. Then, the image state increment needs to be superimposed on the initial image to obtain the current image. Next, the target sub-parameters after the current sub-parameter need to be updated to the current sub-parameter; and based on the product of the current sub-parameter and the evolution step size, the image state increment is updated; then the updated image state increment is superimposed on the current image, and the superimposed image is determined as the current image. Continue to execute the step of updating the target sub-parameters after the current sub-parameter to the current sub-parameter until the current sub-parameter is the last target sub-parameter among multiple target sub-parameters, and finally the current image is determined as the target image corresponding to the target text description information.
[0051] In the above process, perturbations can also be added to each or some of the current images. First, the perturbation parameter needs to be determined, and this perturbation parameter can be determined by random noise. Then, the current image is updated based on the perturbation parameter. Generally speaking, the random noise can be directly superimposed on the current image, or a perturbation intensity parameter can be set, and based on the perturbation intensity parameter, the weights of the random noise and the current image are determined, and the current image is updated based on these weights. In specific implementation, multiple perturbation intensity parameters can be set to update the current image respectively to obtain multiple updated current images, and then the process of updating the current image through the target sub-parameters continues using the multiple updated current images.
[0052] The embodiment of the present invention also provides another image processing method, and this method is implemented based on the method Figure 1 shown. This method can directly use text information to guide the evolution of image content, without reconstructing intermediate representations, and at the same time has good cross-model applicability.
[0053] In this method, a non-backtracking text-driven image transformation can be realized. Simply put, "non-backtracking" is to avoid the process of "encoding first and then decoding", and directly modify or transform the original image according to the text description.
[0054] The core goal of this method is: according to a text description, directly "transform" the image into the required one, which can not only retain the high-quality details of the original image but also flexibly modify the overall content.
[0055] The core concept of this method is to utilize the image evolution law contained in the pre-trained generative flow model to construct a dynamic model that directly guides the transfer of the source image state to the target image state. This model is driven based on the difference in the potential evolution directions corresponding to the source text description and the target text description.
[0056] The key idea of this method is as follows: Prepare a "pre-trained image generation or transformation model", which can specifically be a "generative flow model", such as a "diffusion model", etc. This model can be understood as an "expert that has learned how to develop an image in a certain direction according to text descriptions". When it is necessary to transform a "cat into a dog", let the model show "if I want to turn a cat into a dog, in which direction should I take steps one by one". The biggest difference between this method and the traditional method is that instead of first compressing the cat picture into the "latent space" inside the model and then decoding it back, it directly makes gradual adjustments "on the original image". This reduces information loss and does not have to rely too much on the internal structure of a specific model.
[0057] This method is implemented in the following way:
[0058] 1. Define the direction of image "transformation": According to the model's understanding of the "source text" (equivalent to the above "initial text description information", taking "cat" as an example) and the "target text" (equivalent to the above "target text description information", taking "dog" as an example), respectively show how to change "if we want to keep the cat" and how to change "if we want to turn it into a dog". The difference between the two is the adjustment direction for "turning a cat into a dog".
[0059] In specific implementation, it is necessary to input a source image (such as a photo of a "cat"), a source text (corresponding to the original image, such as "This is a cat"), and a target text (corresponding to the target description, such as "This is a dog") into the model.
[0060] Specifically, a general, pre-trained "generative flow model" can be adopted. It has learned the correspondence between "text and image content" and can also predict to a certain extent "how an image will evolve under different text guidance".
[0061] Then, calculate the "transformation vector". The model needs to show: if I want to keep the "cat" as a "cat", in which direction should it change (because even a cat can have various detailed adjustments); if I want to turn this picture into a "dog", in which direction should it change.
[0062] Then, by taking the difference between the two, we get the direction of "how to turn a cat into a dog", and this direction can be understood as the compass for image deformation.
[0063] 2. Iterate step by step: Divide this "adjustment direction" into many small steps. At each step, let the image change a little bit, and finally reach the target image we want. Random perturbation and multiple attempts: Add a little randomness at each step to prevent getting stuck in an unsatisfactory result. Multiple random transformations can be tried in parallel, and then the results can be averaged or the best one can be selected to ensure more stable effects.
[0064] After calculating the transformation vector, this "compass" can be applied to the image to gradually change the image in this direction, with each change being very small, so as to ensure that the quality and details of the original image can be retained as much as possible.
[0065] At each step, a little random variation can be added, and after trying multiple versions, the average or the best one can be taken to make the result more rich and stable.
[0066] 3 Output: After multiple steps of iteration, the output result is an image that "looks like a dog", and details such as its background, style, and resolution also retain the information of the "source image" as much as possible.
[0067] There is no need to go back and reconstruct during this process. There is no need to "encode first and then decode" to trace back the intermediate representation of the image throughout the process, and information loss is minimized as much as possible.
[0068] The "transformation vector" in the above text can also be a state transition vector field that can guide the evolution of the content of the source image to the content of the target image. This vector field is composed of the differences in the content evolution trends of the pre-trained generative flow model under different text descriptions.
[0069] Explaining the meaning of this vector field with a real-life phenomenon is as follows: Suppose walking from point A to point B on a map, different routes correspond to different speeds and directions; similarly, the goal of this method is to make the image "walk" from the source image (point A) to the target image (point B). The "state transition vector field" is equivalent to the route and speed that guide the image to "walk".
[0070] Suppose the content evolution trend of the pre-trained generative flow model at a given text description P and time step τ is U(X, τ, P), where X represents the image. The "content evolution trend" here can be understood as the change direction and amplitude of the image content per unit time under the guidance of a specific text description.
[0071] Among them, the "content evolution trend" U(X, τ, P) can be understood as the "change direction" of the image X at time τ under the guidance of the text description P in the model. The text description will affect how the image changes. For example, the text prompt "a cat" will guide the image to generate the appearance of a cat, and "a dog" will guide the image to generate the appearance of a dog.
[0072] In specific implementation, the pre-trained generative flow model used can be an architecture based on a diffusion model or a flow model. For different models, the way to obtain the content evolution trend U(X, τ, P) is different:
[0073] Diffusion models (such as Stable Diffusion): In diffusion models, the image generation process can be regarded as a process of gradually denoising from noise. At each step of denoising, the model predicts a noise residual ε_predicted. Specifically, the negative gradient of this residual (i.e., -ε_predicted or the "Score function") can be used as the content evolution trend of the image under this text description and time step. Specifically, given an image X, a text description P, and a time step τ, call the noise prediction network of the diffusion model to obtain the noise residual ε_predicted(X,τ,P), and then calculate -ε_predicted(X,τ,P) as U(X,τ,P).
[0074] U(X,τ,P) = -ε_predicted(X,τ,P)
[0075] GAN-based flow models or other flow models: For such models, the content evolution trend U(X,τ,P) can be obtained in the following ways: 1. Directly call the forward network in the model to obtain the changes in the output of each layer, and accumulate these changes as the content evolution trend; 2. Obtain it by calculating the gradient of the log-likelihood. The specific implementation depends on the specific flow model architecture. For example, for an ODE-based flow model, the reverse flow direction of the current state can be obtained by solving the reverse differential equation and used as the content evolution trend U(X,τ,P).
[0076] Regardless of the model used, it is necessary to utilize the generation or transformation capabilities inherent in the model, rather than constructing an additional or intermediate representation space, and directly use the output or gradient-related results of the model as the guidance for the evolution of the image content.
[0077] For the input image X_input and its corresponding descriptive text P_input, the content evolution trend of the model in this state (equivalent to the above "first image change trend parameter") is:
[0078] D_input(X,τ) = U(X,τ,P_input)
[0079] For the target image content and its descriptive text P_target, the content evolution trend of the model in this state (equivalent to the above "second image change trend parameter") is:
[0080] D_target(X,τ) = U(X,τ,P_target)
[0081] Then a transformation vector field (equivalent to the above "target image change trend parameter") can be defined to guide the transfer of the input image content to the target image content:
[0082] V_transform(S(τ), τ) = D_target(S_target(τ), τ) - D_input(S_input(τ), τ)
[0083] Where S(τ) is the image state during the transformation process, and S_target(τ) and S_input(τ) represent the states of the image content starting from the initial state under the guidance of the target text description and the source text description at the evolution step τ, respectively.
[0084] Since the image needs to be transformed from the source image to the target image, and the "transformation vector field" V_transform(S(τ), τ) tells the image in which direction it should "go" to reach the target image. It is calculated by subtracting "how the source image should change" from "how the target image should change".
[0085] Time step (τ): In the diffusion model, the time step τ usually corresponds to the discrete time points during the diffusion process, representing the denoising process from pure noise to a complete image. In the ODE-based flow model, τ can correspond to the discrete time of the continuous flow process. Smaller τ values correspond to the early stage of the denoising process, when the overall structure of the image is mainly captured, and larger τ values correspond to the later stage of the denoising process, when the detailed information of the image is mainly captured.
[0086] Evolution step size (δτ): The choice of δτ affects the smoothness of the transformation trajectory and the computational cost. A smaller δτ can make the transformation process smoother but requires more computational resources; a larger δτ can speed up the transformation but may lead to a decrease in image quality. In the present invention, the range of δτ is usually selected between 0.001 and 0.01, and the specific value will be adjusted according to the scale of the model and the resolution of the image. For cases with sufficient computational resources, it is usually recommended to select a smaller δτ value to ensure the quality of the transformation result.
[0087] To achieve a smooth transition from the input image content to the target image content, a kinetic equation can be constructed to describe the transformation trajectory:
[0088] dS(τ) / dτ = V_transform(S(τ), τ)
[0089] The initial condition is S(0) = X_input. By numerically solving this kinetic equation, a transformation trajectory starting from the input image content and finally approaching the target image content can be obtained.
[0090] This kinetic equation is like a navigation system, telling the image how to adjust itself at each step and finally reach the state of the target image.
[0091] In practical engineering applications, it is necessary to discretize and solve the above continuous dynamic equations. Specifically, a discretization scheme based on the forward difference method can be adopted:
[0092] Let the discrete evolution step size be δτ, then the iteration formula is:
[0093] S_{k + 1}=S_k+δτ*V_transform(S_k,τ_k)
[0094] Where S_k represents the image state at the k-th evolution step, and τ_k is the corresponding evolution time.
[0095] Since the dynamic equation is continuous and cannot be directly solved by a computer, it is necessary to "discretize" it, that is, to divide the continuous evolution process into many very small steps. Just like dividing a journey into small steps to walk.
[0096] The following shows the comparison experiment results of some key hyperparameters:
[0097] Perturbation intensity (βk):
[0098] βk is too large: There will be too much random noise in the image transformation result, leading to a decrease in image quality.
[0099] βk is too small: The image transformation is likely to fall into a local optimum and it is difficult to break out of the appearance of the source image.
[0100] Evolution step size (δτ):
[0101] δτ is larger: The image transformation speed is faster, but the detailed information of the image may be lost.
[0102] δτ is smaller: The image transformation is smoother, but the computational cost is larger.
[0103] Number of iterations:
[0104] Too few iterations: The image transformation is insufficient and may not achieve the expected effect.
[0105] Too many iterations: The computational cost is large and may lead to excessive image transformation.
[0106] To enhance the robustness of the transformation process, a transformation strategy with perturbation injection is further proposed. At each evolution step, a small random perturbation (noise) is added to the input image state, and the transformation vector field is calculated based on the perturbed input image state and the current transformation state. The specific iteration formula is as follows:
[0107] Let the injected random perturbation be η_k, and its statistical characteristics conform to a preset distribution, such as a Gaussian distribution with a mean of zero, then:
[0108] S_input_perturbed = (1 - β_k) * X_input + β_k * η_k
[0109] Among them, β_k is the perturbation intensity parameter.
[0110] The purpose of adding noise is to make the image transformation process more "flexible" and avoid falling into local optimal solutions. Just like adding a little left - right sway to a walking person to make it easier for him to avoid obstacles on the road.
[0111] The calculation and update of the transformation vector field are as follows:
[0112] V_transform_k = D_target(S_target_k, τ_k) - D_input(S_input_perturbed, τ_k)
[0113] The update of the image state is as follows:
[0114] S_{k + 1}=S_k+δτ*V_transform_k
[0115] To further improve the quality of the transformation results, an optional ensemble averaging strategy is also adopted. At each evolution step, N transformation operations with perturbation injection can be performed, and the multiple transformation results are ensemble - averaged as the final state of the current step. In practice, the value of N is usually taken as 1 - 5.
[0116] The random perturbation η_k follows a Gaussian distribution with zero mean, and its variance can be adjusted as needed. Usually, a relatively large variance is used in the initial stage of iteration to increase the exploration space, and the variance is gradually reduced in the later stage of iteration (similar to the annealing process) to ensure the stability of convergence in the later stage.
[0117] In practice, for each iteration step, N copies of candidate images with random perturbations can be generated in parallel, and then these candidate images are averaged or the best candidate image is selected according to a certain evaluation criterion, and then this result is used as the output of this iteration.
[0118] The method has the following details:
[0119] The random perturbation η_k conforms to a Gaussian distribution with zero mean, and its variance gradually decreases according to the number of iterations (simulating the annealing process) to ensure sufficient exploration in the initial stage of iteration and the stability of convergence in the later stage.
[0120] In practice, for each iteration step, N (for example, N = 3) copies of candidate images with random perturbations can be generated in parallel, and then these candidate images are averaged, or a weighted average is performed to obtain the final output.
[0121] Model Selection: Use pre-trained generative flow models such as Stable Diffusion or flow-based models.
[0122] Text Encoding: Use a pre-trained text encoder (such as the text encoder of CLIP) to convert the text description into a numerical vector.
[0123] Parameter Settings:
[0124] δτ (discrete evolution step): Usually selected between 0.001 and 0.01.
[0125] β_k (perturbation intensity parameter): Usually selected between 0.01 and 0.05 and gradually decreased.
[0126] Number of Iterations: Usually between 20 and 50.
[0127] Number of Ensemble Averagings: Usually selected between 1 and 5.
[0128] The following are some example scenarios to illustrate the applicability of the method in different style transformations and attribute transformations:
[0129] Style Transformation: Transform a "beach photo on a sunny day" into a "beach photo on a rainy day" or into a "beach photo at sunset".
[0130] Attribute Transformation: Transform a "daytime city street scene" into a "nighttime city street scene", or transform a "summer green grassland" into a "winter snowfield".
[0131] Object Replacement: Replace a picture of an "apple" with an "orange", or replace a picture of a "blue car" with a "red car".
[0132] The main difference between this method and the "backtracking" editing method is:
[0133] "Backtracking" Editing: Taking the diffusion model as an example, the conventional method first maps the source image to the latent space (such as Latent code), then generates a new latent vector according to the target text prompt, and finally decodes the new latent vector back to the image. This process will lose some detailed information of the image, resulting in a mismatch in style between the edited image and the original image, and requires two encoding / decoding processes, with a large computational overhead.
[0134] The "non-backtracking" editing of this method directly performs iterative transformations in the image space, utilizing the content evolution ability of the generative flow model itself without mapping the image to the latent space. This method avoids the information loss and computational overhead caused by encoding / decoding, and better preserves the details and overall structure of the image.
[0135] The above method can achieve the following technical effects:
[0136] High-quality content transformation: Since the reconstruction of the intermediate representation is avoided, the detailed information and overall structure of the original image can be better retained.
[0137] Good cross-model applicability: The core of this method lies in utilizing the content evolution law of the generative flow model, and it can be applied to flow models with different structures.
[0138] Efficient transformation process: Without complex optimization calculations, the transformation efficiency is high.
[0139] Flexible control of the transformation degree: The degree of transformation can be controlled by adjusting parameters such as the evolution step size and perturbation intensity.
[0140] For the above method embodiments, refer to Figure 2 An image processing apparatus as shown, the apparatus includes:
[0141] An initial image acquisition module 202, configured to acquire an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image;
[0142] A parameter determination module 204, configured to perform a first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform a second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; wherein, the first image change trend parameter is used to indicate: during the first image processing, the change trend of the initial image; the second image change trend parameter is used to indicate: during the second image processing, the change trend of the initial image;
[0143] An image generation module 206, configured to process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0144] The above-mentioned image processing device obtains an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; the initial image is subjected to first image processing based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and the initial image is subjected to second image processing based on the target text description information through the image processing model to determine a second image change trend parameter; the initial image is processed based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information. This method does not require encoding and decoding processing of the image, reduces the loss of image details, and improves the image processing effect.
[0145] The above-mentioned image processing model includes a diffusion model; the parameter determination module is further configured to: perform the following operations through a pre-trained diffusion model: generate a first noise map corresponding to the initial image, determine a first predicted noise residual based on the initial text description information and the first noise map, and perform a first denoising process on the first noise map based on the first predicted noise residual; determine a first image change trend parameter based on the first predicted noise residual; the parameter determination module is further configured to: perform the following operations through the diffusion model: generate a second noise map corresponding to the initial image, determine a second predicted noise residual based on the target text description information and the first noise map, and perform a second denoising process on the second noise map based on the second predicted noise residual; determine a second image change trend parameter based on the second predicted noise residual.
[0146] The above-mentioned image processing model includes a diffusion model; the first denoising process includes first sub-operations performed at multiple time steps; each first sub-operation has a corresponding first predicted noise residual; the parameter determination module is further configured to: for each first predicted noise residual, determine a first negative gradient corresponding to the first predicted noise residual; arrange the multiple first negative gradients in the execution order of the corresponding first denoising process to determine a first image change trend parameter; the second denoising process includes second sub-operations performed at multiple time steps; each second sub-operation has a corresponding second predicted noise residual; the parameter determination module is further configured to: for each second predicted noise residual, determine a second negative gradient corresponding to the second predicted noise residual; arrange the multiple second negative gradients in the execution order of the corresponding second denoising process to obtain a second image change trend parameter.
[0147] The above image processing model includes a flow model; the flow model includes a forward network; the parameter determination module is further configured to: obtain the first output parameters of the forward network during the first image processing of the initial image by the flow model based on the initial text description information; determine the first image change trend parameters based on the first output parameters; the step of determining the second image change trend parameters includes: obtaining the second output parameters of the forward network during the second image processing of the initial image by the flow model based on the target text description information; determining the second image change trend parameters based on the second output parameters.
[0148] The above image generation module is further configured to: determine the target trend change parameters based on the first image change trend parameters and the second image change trend parameters; process the initial image based on the target trend change parameters to generate the target image corresponding to the target text description information.
[0149] The above first image change trend parameters include a plurality of first sub-parameters arranged in sequence according to the corresponding time steps; the second image change trend parameters include a plurality of second sub-parameters arranged in sequence according to the corresponding time steps; the above image generation module is further configured to: for each first sub-parameter, determine the second sub-parameter corresponding to the first sub-parameter; the time step corresponding to the first sub-parameter and the corresponding second sub-parameter is the same; determine the difference between the second sub-parameter and the first sub-parameter as the target sub-parameter corresponding to the time step of the first sub-parameter; arrange the plurality of target sub-parameters based on the time steps corresponding to the plurality of target sub-parameters to obtain the target trend change parameters.
[0150] The above target trend change parameters include a plurality of target sub-parameters arranged in sequence according to the corresponding time steps; the above image generation module is further configured to: determine the first target sub-parameter among the plurality of target sub-parameters as the current sub-parameter; determine the image state increment based on the product of the current sub-parameter and the evolution step size; superimpose the image state increment on the initial image to obtain the current image; update the target sub-parameter after the current sub-parameter as the current sub-parameter; update the image state increment based on the product of the current sub-parameter and the evolution step size; superimpose the updated image state increment on the current image, and determine the superimposed image as the current image; continue to execute the step of updating the target sub-parameter after the current sub-parameter as the current sub-parameter until the current sub-parameter is the last target sub-parameter among the plurality of target sub-parameters, and determine the current image as the target image corresponding to the target text description information.
[0151] The above device further includes: a perturbation parameter determination module for determining perturbation parameters; an image update module for updating the current image based on the perturbation parameters.
[0152] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the above image processing method. For example:
[0153] Obtain an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; perform first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; wherein, the first image change trend parameter is used to indicate: the change trend of the initial image during the first image processing; the second image change trend parameter is used to indicate: the change trend of the initial image during the second image processing; process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0154] The above method does not require encoding and decoding processing of the image, reduces the loss of image details, and improves the image processing effect.
[0155] Optionally, the above image processing model includes a diffusion model; the step of performing first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter includes: performing the following operations through the pre-trained diffusion model: generating a first noise map corresponding to the initial image, determining a first predicted noise residual based on the initial text description information and the first noise map, and performing first denoising processing on the first noise map based on the first predicted noise residual; determining the first image change trend parameter based on the first predicted noise residual; the step of performing second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter includes: performing the following operations through the diffusion model: generating a second noise map corresponding to the initial image, determining a second predicted noise residual based on the target text description information and the first noise map, and performing second denoising processing on the second noise map based on the second predicted noise residual; determining the second image change trend parameter based on the second predicted noise residual.
[0156] Optionally, the above image processing model includes a diffusion model; the first denoising process includes first sub-operations performed at multiple time steps; each first sub-operation has a corresponding first predicted noise residual; the step of determining the first image change trend parameter based on the first predicted noise residual includes: for each first predicted noise residual, determining a first negative gradient corresponding to the first predicted noise residual; arranging the multiple first negative gradients in the execution order of the corresponding first denoising process to determine the first image change trend parameter; the second denoising process includes second sub-operations performed at multiple time steps; each second sub-operation has a corresponding second predicted noise residual; the step of determining the second image change trend parameter based on the second predicted noise residual includes: for each second predicted noise residual, determining a second negative gradient corresponding to the second predicted noise residual; arranging the multiple second negative gradients in the execution order of the corresponding second denoising process to obtain the second image change trend parameter.
[0157] Optionally, the above image processing model includes a flow model; the flow model includes a forward network; the step of determining the first image change trend parameter includes: obtaining a first output parameter of the forward network during the first image processing of the initial image by the flow model based on the initial text description information; determining the first image change trend parameter based on the first output parameter; the step of determining the second image change trend parameter includes: obtaining a second output parameter of the forward network during the second image processing of the initial image by the flow model based on the target text description information; determining the second image change trend parameter based on the second output parameter.
[0158] Optionally, the step of processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information includes: determining a target trend change parameter based on the first image change trend parameter and the second image change trend parameter; processing the initial image based on the target trend change parameter to generate a target image corresponding to the target text description information.
[0159] Optionally, the first image change trend parameter includes multiple first sub-parameters arranged in sequence according to the corresponding time steps; the second image change trend parameter includes multiple second sub-parameters arranged in sequence according to the corresponding time steps; the step of determining the target trend change parameter based on the first image change trend parameter and the second image change trend parameter includes: for each first sub-parameter, determining a second sub-parameter corresponding to the first sub-parameter; the time step corresponding to the first sub-parameter and the corresponding second sub-parameter is the same; determining the difference between the second sub-parameter and the first sub-parameter as the target sub-parameter corresponding to the time step corresponding to the first sub-parameter; arranging the multiple target sub-parameters based on the time steps corresponding to the multiple target sub-parameters to obtain the target trend change parameter.
[0160] Optionally, the above-mentioned target trend change parameter includes a plurality of target sub-parameters arranged according to corresponding time steps; the step of generating a target image corresponding to the target text description information by processing the initial image based on the target trend change parameter includes: determining the first target sub-parameter among the plurality of target sub-parameters as the current sub-parameter; determining an image state increment based on the product of the current sub-parameter and the evolution step size; superimposing the image state increment on the initial image to obtain the current image; updating the target sub-parameter after the current sub-parameter as the current sub-parameter; updating the image state increment based on the product of the current sub-parameter and the evolution step size; superimposing the updated image state increment on the current image, and determining the superimposed image as the current image; continuing to execute the step of updating the target sub-parameter after the current sub-parameter as the current sub-parameter until the current sub-parameter is the last target sub-parameter among the plurality of target sub-parameters, and determining the current image as the target image corresponding to the target text description information.
[0161] Optionally, the above method further includes: determining a perturbation parameter; updating the current image based on the perturbation parameter.
[0162] See Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above image processing method.
[0163] Furthermore, Figure 3 The electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0164] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only one bidirectional arrow is used in
[0165] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the method of the foregoing embodiments.
[0166] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above image processing method.
[0167] An image processing method, apparatus, and electronic device provided by an embodiment of the present disclosure include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments, for example:
[0168] Obtain an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; perform first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; wherein, the first image change trend parameter is used to indicate: the change trend of the initial image during the first image processing; the second image change trend parameter is used to indicate: the change trend of the initial image during the second image processing; process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
[0169] The above method does not require encoding and decoding processing of the image, reduces the loss of image details, and improves the image processing effect.
[0170] Optionally, the above image processing model includes a diffusion model; the step of performing first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter includes: performing the following operations through a pre-trained diffusion model: generating a first noise map corresponding to the initial image, determining a first predicted noise residual based on the initial text description information and the first noise map, and performing first denoising processing on the first noise map based on the first predicted noise residual; determining the first image change trend parameter based on the first predicted noise residual; the step of performing second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter includes: performing the following operations through the diffusion model: generating a second noise map corresponding to the initial image, determining a second predicted noise residual based on the target text description information and the first noise map, and performing second denoising processing on the second noise map based on the second predicted noise residual; determining the second image change trend parameter based on the second predicted noise residual.
[0171] Optionally, the above image processing model includes a diffusion model; the first denoising process includes first sub-operations performed at multiple time steps; each first sub-operation has a corresponding first predicted noise residual; the step of determining the first image change trend parameter based on the first predicted noise residual includes: for each first predicted noise residual, determining the first negative gradient corresponding to the first predicted noise residual; arranging the multiple first negative gradients in the execution order of the corresponding first denoising process to determine the first image change trend parameter; the second denoising process includes second sub-operations performed at multiple time steps; each second sub-operation has a corresponding second predicted noise residual; the step of determining the second image change trend parameter based on the second predicted noise residual includes: for each second predicted noise residual, determining the second negative gradient corresponding to the second predicted noise residual; arranging the multiple second negative gradients in the execution order of the corresponding second denoising process to obtain the second image change trend parameter.
[0172] Optionally, the above image processing model includes a flow model; the flow model includes a forward network; the step of determining the first image change trend parameter includes: obtaining the first output parameter of the forward network during the first image processing of the initial image by the flow model based on the initial text description information; determining the first image change trend parameter based on the first output parameter; the step of determining the second image change trend parameter includes: obtaining the second output parameter of the forward network during the second image processing of the initial image by the flow model based on the target text description information; determining the second image change trend parameter based on the second output parameter.
[0173] Optionally, the step of processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate the target image corresponding to the target text description information includes: determining the target trend change parameter based on the first image change trend parameter and the second image change trend parameter; processing the initial image based on the target trend change parameter to generate the target image corresponding to the target text description information.
[0174] Optionally, the first image change trend parameter includes multiple first sub-parameters arranged in sequence according to the corresponding time steps; the second image change trend parameter includes multiple second sub-parameters arranged in sequence according to the corresponding time steps; the step of determining the target trend change parameter based on the first image change trend parameter and the second image change trend parameter includes: for each first sub-parameter, determining the second sub-parameter corresponding to the first sub-parameter; the time step corresponding to the first sub-parameter and the corresponding second sub-parameter is the same; determining the difference between the second sub-parameter and the first sub-parameter as the target sub-parameter corresponding to the time step corresponding to the first sub-parameter; arranging the multiple target sub-parameters based on the time steps corresponding to the multiple target sub-parameters to obtain the target trend change parameter.
[0175] Optionally, the above-mentioned target trend change parameter includes multiple target sub-parameters arranged according to corresponding time steps; the step of generating a target image corresponding to the target text description information by processing the initial image based on the target trend change parameter includes: determining the first target sub-parameter among the multiple target sub-parameters as the current sub-parameter; determining an image state increment based on the product of the current sub-parameter and the evolution step size; superimposing the image state increment on the initial image to obtain the current image; updating the target sub-parameter after the current sub-parameter as the current sub-parameter; updating the image state increment based on the product of the current sub-parameter and the evolution step size; superimposing the updated image state increment on the current image, and determining the superimposed image as the current image; continuing to execute the step of updating the target sub-parameter after the current sub-parameter as the current sub-parameter until the current sub-parameter is the last target sub-parameter among the multiple target sub-parameters, and determining the current image as the target image corresponding to the target text description information.
[0176] Optionally, the above method further includes: determining a perturbation parameter; updating the current image based on the perturbation parameter.
[0177] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0178] In addition, in the description of the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.
[0179] If the above function is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of the present disclosure, or the part that contributes to the prior art, or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0180] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present disclosure. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0181] Finally, it should be noted that the above embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; Performing first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and performing second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; Wherein, the first image change trend parameter is used to indicate the change trend of the initial image during the first image processing; the second image change trend parameter is used to indicate the change trend of the initial image during the second image processing; Processing the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
2. The method according to claim 1, characterized in that, The image processing model includes a diffusion model; The step of performing first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter includes: Performing the following operations through a pre-trained diffusion model: generating a first noise map corresponding to the initial image, determining a first predicted noise residual based on the initial text description information and the first noise map, and performing first denoising processing on the first noise map based on the first predicted noise residual; Determining a first image change trend parameter based on the first predicted noise residual; The step of performing second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter includes: Performing the following operations through the diffusion model: Generating a second noise map corresponding to the initial image, determining a second predicted noise residual based on the target text description information and the first noise map, and performing second denoising processing on the second noise map based on the second predicted noise residual; Determining a second image change trend parameter based on the second predicted noise residual.
3. The method according to claim 2, wherein The image processing model includes a diffusion model; the first denoising processing includes first sub-operations performed at multiple time steps; each of the first sub-operations has a corresponding first predicted noise residual; The step of determining a first image change trend parameter based on the first predicted noise residual includes: For each of the first predicted noise residuals, determining a first negative gradient corresponding to the first predicted noise residual; Arranging the multiple first negative gradients in the execution order of the corresponding first denoising processing to determine a first image change trend parameter; The second denoising processing includes second sub-operations performed at multiple time steps; each of the second sub-operations has a corresponding second predicted noise residual; The step of determining a second image change trend parameter based on the second predicted noise residual includes: For each of the second predicted noise residuals, determining a second negative gradient corresponding to the second predicted noise residual; Arrange the multiple second negative gradients in the execution order corresponding to the second denoising process to obtain the second image change trend parameter.
4. The method according to claim 1, wherein The image processing model includes a flow model; the flow model includes a forward network. The step of determining the first image change trend parameter includes: Obtain the first output parameter of the forward network during the first image processing of the initial image by the flow model based on the initial text description information. Based on the first output parameter, determine the first image change trend parameter. The step of determining the second image change trend parameter includes: Obtain the second output parameter of the forward network during the second image processing of the initial image by the flow model based on the target text description information. Based on the second output parameter, determine the second image change trend parameter.
5. The method according to claim 1, characterized in that, The step of generating the target image corresponding to the target text description information by processing the initial image based on the first image change trend parameter and the second image change trend parameter includes: Based on the first image change trend parameter and the second image change trend parameter, determine the target trend change parameter. Process the initial image based on the target trend change parameter to generate the target image corresponding to the target text description information.
6. The method according to claim 5, wherein The first image change trend parameter includes multiple first sub-parameters arranged in sequence according to the corresponding time steps; the second image change trend parameter includes multiple second sub-parameters arranged in sequence according to the corresponding time steps. The step of determining the target trend change parameter based on the first image change trend parameter and the second image change trend parameter includes: For each first sub-parameter, determine the second sub-parameter corresponding to the first sub-parameter; the time step corresponding to the first sub-parameter is the same as that of the corresponding second sub-parameter. Determine the difference between the second sub-parameter and the first sub-parameter as the target sub-parameter corresponding to the time step of the first sub-parameter. Based on the time steps corresponding to the multiple target sub-parameters, arrange the multiple target sub-parameters to obtain the target trend change parameter.
7. The method according to claim 5, wherein The target trend change parameter includes multiple target sub-parameters arranged according to the corresponding time steps; the interval between adjacent time steps is the evolution step size. The step of generating the target image corresponding to the target text description information by processing the initial image based on the target trend change parameter includes: Determine the first target sub-parameter among the multiple target sub-parameters as the current sub-parameter. Based on the product of the current sub-parameter and the evolution step size, determine the image state increment. Superimpose the image state increment on the initial image to obtain the current image. Update the target sub-parameter after the current sub-parameter as the current sub-parameter. Based on the product of the current sub-parameter and the evolution step size, update the image state increment. Superimpose the updated image state increment on the current image, and determine the superimposed image as the current image. Continue to execute the step of updating the target sub-parameter after the current sub-parameter to the current sub-parameter until the current sub-parameter is the last target sub-parameter among the multiple target sub-parameters, and determine the current image as the target image corresponding to the target text description information.
8. The method according to claim 7, characterized in that, The method further includes: Determine the perturbation parameter; Update the current image based on the perturbation parameter.
9. An image processing apparatus, characterized in that, The apparatus includes: An initial image acquisition module, configured to acquire an initial image and target text description information; the initial image has corresponding initial text description information; the initial text description information is used to describe the display content in the initial image; A parameter determination module, configured to perform first image processing on the initial image based on the initial text description information through a pre-trained image processing model to determine a first image change trend parameter, and perform second image processing on the initial image based on the target text description information through the image processing model to determine a second image change trend parameter; Wherein, the first image change trend parameter is used to indicate the change trend of the initial image during the first image processing; the second image change trend parameter is used to indicate the change trend of the initial image during the second image processing; An image generation module, configured to process the initial image based on the first image change trend parameter and the second image change trend parameter to generate a target image corresponding to the target text description information.
10. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the image processing method according to any one of claims 1-8.
11. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the image processing method according to any one of claims 1-8.