Image processing method and device, electronic equipment, storage medium and program product
By determining the correlation and feature differences between the current control point area and the source control point, and updating the control point and hidden vector, the problem of inaccurate control point adjustment in image editing is solved, and image quality and efficiency are improved.
Patent Information
- Application Number
- CN202410067502.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, how to accurately adjust the control points to improve the final image quality during image editing processing is an urgent problem.
By obtaining the regional characteristics of the current control point area and the distribution parameters of the source control point, the correlation and characteristic differences between the control point area and the source control point are determined, and the updated control point is determined in the current control point area based on the correlation and characteristic differences, and the hidden vector is updated to generate the processed image.
While retaining image features, precisely control the degree of image deformation, improving the quality and efficiency of generated images.
Smart Images

Figure CN120339143A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, device, electronic device, storage medium and program product. Background Art
[0002] In recent years, with the development of computer technology, image editing has been widely used in computer animation, computer games and computer vision. In the prior art, image editing technology is often used to drag the control points of the original image to perform deformation operations to achieve services such as animation character posture editing, scene editing, character face editing, etc.
[0003] However, how to accurately adjust the control points during the image editing process to improve the quality of the final image is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present application provide an image processing method, device, electronic device, storage medium and program product, which can improve the quality of the generated processed image.
[0005] An embodiment of the present application provides an image processing method, comprising: obtaining regional features of a current control point region from current image features of an image to be processed, the current image features being obtained by a current latent vector of the image to be processed, the image to be processed comprising a source control point and a current control point corresponding to the source control point, and the current control point region being a region related to the current control point; determining the correlation between the current control point region and the source control point according to the regional features of the current control point region and a source control point distribution parameter, the source control point distribution parameter being used to represent the Gaussian distribution of the source control point in the image to be processed; determining a feature difference between the current control point region and the source control point according to the regional features of the current control point region and the control point features of the source control point; determining an updated current control point in the current control point region in combination with the correlation and the feature difference; updating the current latent vector according to the updated current control point to obtain an updated current latent vector, so as to generate a processed image according to the updated current control point and the updated current latent vector.
[0006] An embodiment of the present application further provides an image processing apparatus, including: an acquisition unit, configured to acquire regional features of a current control point region from current image features of an image to be processed, where the current image features are obtained from a current latent vector of the image to be processed, the image to be processed includes source control points and current control points corresponding to the source control points, and the current control point region is a region related to the current control points; a correlation determination unit, configured to determine the correlation between the current control point region and the source control points according to the regional features of the current control point region and source control point distribution parameters, where the source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed; a difference determination unit, configured to determine the feature difference between the current control point region and the source control points according to the regional features of the current control point region and the control point features of the source control points; a control point determination unit, configured to determine updated current control points in the current control point region by combining the correlation and the feature difference; and an update unit, configured to update the current latent vector according to the updated current control points to obtain an updated current latent vector, so as to generate a processed image according to the updated current control points and the updated current latent vector.
[0007] In some embodiments, the image processing apparatus further includes a feature learning unit, and the feature learning unit is configured to: acquire an image generation model; perform feature learning on the current latent vector of the image to be processed through the image generation model, and use the feature vector output by an intermediate layer of the image generation model as the current image features of the image to be processed.
[0008] In some embodiments, the source control point distribution parameters include a pre-trained convolution matrix, the pre-trained convolution matrix is trained based on Gaussian distribution information of the source control points in the image to be processed, and the correlation determination unit is specifically configured to: acquire the pre-trained convolution matrix; perform convolution calculation on the regional features of the current control point region through the pre-trained convolution matrix to determine the correlation between the current control point region and the source control points.
[0009] In some embodiments, the image processing apparatus further includes a pre-training unit, and the pre-training unit is configured to: acquire an initial convolution matrix and Gaussian distribution information of the source control points in the image to be processed; use the Gaussian distribution information of the source control points in the image to be processed as a target value; perform convolution calculation on the regional features of an initial control point region through the initial convolution matrix to determine the initial correlation between points in the initial control point region and the source control points, where the initial control point region is a region related to the source control points; and train the initial convolution matrix according to the difference between the initial correlation and the target value to obtain the pre-trained convolution matrix.
[0010] In some embodiments, the difference determination unit is specifically configured to: obtain the control point features of the source control point from the initial latent vector of the image to be processed; use the feature distance between the region features of the current control point region and the control point features of the source control point as the feature difference between the current control point region and the source control point.
[0011] In some embodiments, the control point determination unit is specifically configured to: calculate the confidence of the points within the current control point region by combining the correlation and the feature difference; determine the updated current control point within the current control point region according to the confidence.
[0012] In some embodiments, the update unit is specifically configured to: update the current latent vector according to the updated current control point to obtain an updated current latent vector; use the updated current control point as the current control point, and use the updated current latent vector as the current latent vector, and return to execute the steps of obtaining the region features of the current control point region from the image features of the image to be processed and subsequent steps until the updated current control point matches the specified target control point, and generate a processed image from the updated current latent vector.
[0013] In some embodiments, the updating the current latent vector according to the updated current control point to obtain an updated current latent vector includes: obtaining the region features of the reference control point region corresponding to the confidence of the updated current control point according to the confidence of the updated current control point, where the confidence is obtained from the correlation and the feature difference, and the reference control point region includes one of the current control point region and the initial control point region, and the initial control point region is a region related to the source control point; determining a loss value according to the region features of the reference control point region and the region features corresponding to the updated current control point; updating the current latent vector according to the loss value to obtain an updated current latent vector.
[0014] In some embodiments, the obtaining the region features of the reference control point region corresponding to the confidence of the updated current control point includes: if the confidence of the updated current control point is greater than or equal to a preset threshold, obtaining the region features of the current control point region; if the confidence of the updated current control point is less than the preset threshold, obtaining the region features of the initial control point region.
[0015] In some embodiments, the image to be processed further includes a preset deformation region, the source control points and the specified target control points are located within the preset deformation region, and determining the loss value according to the regional features of the reference control point region and the regional features corresponding to the updated current control points includes: determining the difference in regional features of the control point regions according to the regional features of the reference control point region and the regional features corresponding to the updated current control points; determining the difference in deformation region features according to the regional features of the preset deformation region in the initial image features and the regional features of the preset deformation region in the current image features, where the initial image features are obtained from the initial hidden vector of the image to be processed; and determining the loss value by combining the difference in regional features of the control point regions and the difference in deformation region features.
[0016] In some embodiments, the regional features corresponding to the updated current control points are obtained through the following steps: determining the offset of the updated current control points relative to the specified target control points according to the feature distance between the specified target control points and the updated current control points; determining the offset control points corresponding to the updated current control points according to the updated current control points and the offset; and obtaining the regional features of the offset control point region from the current image features of the image to be processed, so as to use the regional features of the offset control point region as the regional features corresponding to the updated current control points, where the offset control point region is the region related to the offset control points.
[0017] An embodiment of the present application further provides an electronic device, including a processor and a memory, where the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in any one of the image processing methods provided by the embodiments of the present application.
[0018] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiments of the present application.
[0019] An embodiment of the present application further provides a computer program product, including a computer program or instructions, where when the computer program or instructions are executed by a processor, the steps in any one of the image processing methods provided by the embodiments of the present application are implemented.
[0020] In an embodiment of the present application, the regional feature of the current control point area can be obtained from the current image features of the image to be processed. The current image features are obtained from the current latent vector of the image to be processed. The image to be processed includes source control points and current control points corresponding to the source control points. The current control point area is an area related to the current control points; according to the regional feature of the current control point area and the source control point distribution parameter, the correlation between the current control point area and the source control points is determined. The source control point distribution parameter is used to represent the Gaussian distribution of the source control points in the image to be processed; according to the regional feature of the current control point area and the control point features of the source control points, the feature difference between the current control point area and the source control points is determined; combining the correlation and the feature difference, an updated current control point is determined within the current control point area; according to the updated current control point, the current latent vector is updated to obtain an updated current latent vector, so as to generate a processed image according to the updated current control point and the updated current latent vector.
[0021] In the present application, during the process of performing deformation processing on the image to be processed, the next current control point (i.e., the updated current control point) in the process is determined through the correlation and feature difference between the current control point area and the source control points, so as to gradually drag the source control points based on the updated current control points to generate a processed image. In this process, through the correlation between the current control point area and the source control points, it can be ensured that the updated current control points have a high similarity with the source control points, so that the generated processed image can better retain the features of the source control points. At the same time, the degree of feature change is measured through the feature difference to control the degree of image deformation. In this way, by combining the correlation and feature difference between the current control point area and the source control points, the degree of image deformation can be accurately controlled while retaining the features of the image to be processed, improving the quality of the generated processed image. In addition, the embodiment of the present application also determines the correlation between the current control point area and the source control points based on the source control point distribution parameter representing the Gaussian distribution of the source control points in the image to be processed, so as to utilize the Gaussian distribution information of the source control points to increase the influence of the points related to the source control points in the current control point area, improve the accuracy of determining the updated current control points within the current control point area, and improve the efficiency and quality of the generated processed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1a It is a schematic diagram of the scenario of the image processing method provided by an embodiment of the present application;
[0024] Figure 1b It is a schematic flowchart of the image processing method provided by an embodiment of the present application;
[0025] Figure 1c It is a comparison schematic diagram of the image to be processed and the processed image provided by an embodiment of the present application;
[0026] Figure 1d It is a schematic diagram of the training process of the convolution matrix provided by an embodiment of the present application;
[0027] Figure 2a It is a schematic diagram of the image processing system provided by an embodiment of the present application;
[0028] Figure 2b It is a schematic flowchart of the image processing method provided by another embodiment of the present application;
[0029] Figure 2c It is the image setting page provided by an embodiment of the present application;
[0030] Figure 2d It is a schematic flowchart of the processing process of dragging any control point provided by an embodiment of the present application;
[0031] Figure 2e It is a comparison schematic diagram of the present application and the prior art for the image processing results of different types provided by an embodiment of the present application;
[0032] Figure 3 It is a schematic structural diagram of the image processing device provided by an embodiment of the present application;
[0033] Figure 4 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0035] An embodiment of the present application provides an image processing method, device, electronic device, storage medium, and program product.
[0036] Among them, the image processing device can be specifically integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.; the server can be a single server or a server cluster composed of multiple servers. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0037] In some embodiments, the image processing device can also be integrated in multiple electronic devices. For example, the image processing device can be integrated in multiple servers, and the image processing method of the present application can be implemented by multiple servers.
[0038] In some embodiments, the server can also be implemented in the form of a terminal.
[0039] For example, referring to Figure 1a , the image processing method can be implemented by a server. The server can obtain the image to be processed from the terminal through a network, and obtain the regional feature of the current control point area from the current image features of the image to be processed. The current image features are obtained from the current hidden vector of the image to be processed. The image to be processed includes source control points and current control points corresponding to the source control points. The current control point area is an area related to the current control point; according to the regional feature of the current control point area and the source control point distribution parameter, determine the correlation between the current control point area and the source control point. The source control point distribution parameter is related to the Gaussian distribution of the source control point; according to the regional feature of the current control point area and the control point feature of the source control point, determine the feature difference between the current control point area and the source control point; combine the correlation and the feature difference to determine the updated current control point within the current control point area; according to the updated current control point, update the current hidden vector to obtain the updated current hidden vector, so as to generate the processed image according to the updated current control point and the updated current hidden vector, and the server returns the processed image to the terminal through the network.
[0040] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0041] The following will be described in detail respectively. It should be noted that the order of the following embodiments does not limit the preferred order of the embodiments. It can be understood that in the specific implementation manner of this application, when it comes to data related to the image to be processed and user operations, etc., when the embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0042] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0043] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the field of vision such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to specific downstream tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0045] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pretrained models are the latest development results of deep learning, integrating the above technologies.
[0046] A pre-training model (PTM), also known as a foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a vast amount of unlabeled data, and by leveraging the function approximation ability of the large-parameter DNN, the PTM extracts common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0047] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0048] In this embodiment, an image processing method related to artificial intelligence is provided, as Figure 1b shown. The specific process of this image processing method can be as follows:
[0049] 110. Obtain the regional features of the current control point area from the current image features of the image to be processed.
[0050] Among them, the image to be processed refers to the original image to be processed. The current image feature is obtained from the current latent vector of the image to be processed. The image to be processed includes source control points and current control points corresponding to the source control points. In the embodiments of the present application, by dragging the control points of the image to be processed, the processed image can be obtained. This dragging process can be understood as a process of updating the positions of the control points. The image processing method provided by the embodiments of the present application can be applied in various scenarios. For example, the image to be processed can include, but is not limited to, one or a combination of a human image, an animal image, a building image, a food image, etc. Taking the image to be processed as a human image as an example, with the user's individual permission or consent, the image of the user can be obtained as the image to be processed for subsequent image processing.
[0051] Among them, the latent vector refers to the potential feature vector used to represent the image to be processed. The latent vector is the initial vector used to generate the image. For an image, the vector obtained after encoding the image is the latent vector. It should be noted that before dragging the control points of the image to be processed (i.e., updating the control points), the latent vector of the image to be processed refers to the original latent vector of the image to be processed (hereinafter referred to as the initial latent vector). As the control points of the image to be processed are continuously dragged, the latent vector can be updated synchronously. In the embodiments of the present application, the initial latent vector can be obtained by encoding the input image to be processed through an encoder.
[0052] Among them, the control point refers to the point used to control the deformation of the image. By dragging the control points of the image (i.e., updating the positions of the control points), the deformation operation of the image is performed. The source control point refers to the control point before updating the position of the control point. The source control point can be set according to the application scenario or actual requirements. In some embodiments, the source control point can be determined by the user's operation of setting the control points of the image to be processed. For example, the operation of setting the control points can include, but is not limited to, clicking one or more positions in the image to be processed and taking them as the source control points, or dragging the mouse or finger to draw one or more box selection areas in the image to be processed and taking specific points (such as the center point) therein as the source control points, or inputting the coordinate information of the source control points to set the position of the source control points in the image to be processed.
[0053] Wherein, the current control point refers to the control point at the current moment, and the current control point is obtained by updating the position of the source control point (i.e., dragging). It can be understood that the current control point corresponding to the moment before updating the position of the control point is the source control point. It should be noted that in the embodiments of the present application, the process of updating the position of the control point (i.e., one-step or multi-step dragging operations) can be performed on the source control point of the image to be processed. Before the first update of the control point position, the source control point 1 can be used as the current control point. For example, during the process of updating the position of the source control point n times, the control point 2 (obtained by updating the position of the source control point 1), the control point 3 (obtained by updating the position of the control point 2),..., the control point n + 1 (obtained by updating the position of the control point n) can be obtained in sequence. That is, the control point 2, the control point 2,..., the control point n + 1 are the current control points obtained after the first drag, the second drag,..., the nth drag, respectively.
[0054] Wherein, the control point area refers to the area related to the control point, and the current control point area is the area related to the current control point. The area feature refers to the feature representation of the area related to the control point in the image. In some embodiments, the control point area is an area with a preset shape centered on the control point. For example, the current control point area can be a circular area with a preset radius or a rectangular area with a preset side length centered on the current control point.
[0055] For example, it can be understood that the feature representation of an image is usually in the form of a feature map. Therefore, according to the position and size of the current control point area, the features at the corresponding position and size can be extracted from the current image features as the area features of the current control point area.
[0056] In some embodiments, useful feature representations of the image to be processed can be extracted by performing feature learning on the latent vector of the image to be processed. Specifically, before obtaining the area features of the current control point area from the current image features of the image to be processed, it further includes:
[0057] Performing feature learning on the current latent vector of the image to be processed to obtain the current image features.
[0058] Wherein, feature learning refers to the process of obtaining new feature representations (i.e., current image features) by learning the distribution of the original data (i.e., the latent vector of the image to be processed). It can be understood that feature learning is also a process of extracting useful feature representations from the original data. During the feature learning process, relevant features can be automatically learned according to the input original data to obtain feature representations that can be used to generate new images. The feature representations can effectively represent the original data and are more easily generalized, that is, the feature representations have better performance.
[0059] For example, in the embodiments of the present application, a neural network model can be used to perform feature learning on the current latent vector of the image to be processed. For example, the neural network model can include, but is not limited to, one or a combination of more of a convolutional neural network (CNN), a recurrent neural network (RNN), a variational autoencoder (VAE), a generative adversarial network (GAN), a diffusion model (Diffusion), or a manifold generation model (MGM), etc.
[0060] In some embodiments, feature learning can be performed on the current latent vector of the image to be processed through an image generation model, and the feature vector output by the intermediate layer of the image generation model can be used as the current image feature. Since the feature vector output by the intermediate layer of the image generation model not only contains a high-level abstract representation of the image to be processed but also has high interpretability, when used in the subsequent current control point update process, it can better understand the relationship between the features of the image to be processed, improve the accuracy and interpretability of updating the control points, and thus improve the quality of the generated processed image. Specifically, performing feature learning on the current latent vector of the image to be processed to obtain the current image feature includes:
[0061] Obtain an image generation model;
[0062] Through the image generation model, perform feature learning on the current latent vector of the image to be processed, and use the feature vector output by the intermediate layer of the image generation model as the current image feature of the image to be processed.
[0063] Among them, the image generation model refers to a model that generates new images by learning the distribution of raw data. For example, the image generation model can include, but is not limited to, one or a combination of more of an autoencoder (VAE), a generative adversarial network (GAN), a manifold generation model (MGM), and a diffusion model (Diffusion), etc.
[0064] Among them, the intermediate layer refers to the layer between the input layer and the output layer of the neural network model. When the neural network model processes the input data, it extracts features layer by layer and performs abstract representation. Each intermediate layer can generate an intermediate layer output, also called a feature vector.
[0065] In the embodiments of the present application, for different image generation models, the intermediate layer output used as the current image feature can be different intermediate layer output feature vectors, specifically depending on the structure and algorithm of the image generation model.
[0066] In some embodiments, the image generation model includes a generative adversarial network, and the current image features are the feature vectors output by the generator of the generative adversarial network. For example, the generative adversarial network may include a generator and a discriminator. The generator is used to generate data, and the discriminator is used to determine whether the generated data is real or machine-generated. Thus, in this application, the current latent vector of the image to be processed can be input into the generator of the generative adversarial network to generate a feature representation of an image similar to the image to be processed through the generator, and use it as the current image features.
[0067] In some embodiments, the image generation model includes a diffusion model, and the current image features are the feature vectors output by the diffusion process of the diffusion model. For example, the current latent vector of the image to be processed can be input into the diffusion model to generate a feature representation of an image similar to the image to be processed through the diffusion process of the diffusion model, and use it as the current image features. Specifically, the diffusion process is an iterative process that transforms the current image state into the next image state at each step. In each step, the image state gradually evolves, becoming more real or vivid from the initial state. Through continuous iteration, a series of image states can be obtained, and finally, the complete diffusion process can be obtained. The diffusion latent code (i.e., the feature representation) of the input image to be processed can be recovered from the image state output by the diffusion process through an inversion technique, and this diffusion latent code is used as the current image features.
[0068] In some embodiments, a LoRA (Low-Rank Adaptation) model can be introduced into the diffusion model to fine-tune the diffusion model using the LoRA algorithm to retain the content and style of the input model image. Through the fine-tuning of LoRA, the diffusion model can find a balance between high fidelity and efficiency.
[0069] In some embodiments, the image generation model is a pre-trained model, which can provide better performance for the feature learning process and improve the accuracy of the image features obtained by feature learning. In addition, in the embodiments of this application, the model parameters of the pre-trained image generation model can be fixed, that is, in the image processing process of the embodiments of this application, the model parameters of the image generation model are no longer updated, so as not to increase the time of image processing and improve the processing efficiency.
[0070] 120. Determine the correlation between the current control point region and the source control points according to the regional features of the current control point region and the source control point distribution parameters.
[0071] Among them, the source control point distribution parameter is used to represent the Gaussian distribution of the source control points in the image to be processed. The source control point distribution parameter can be a parameter that directly or indirectly represents the Gaussian distribution information of the source control points in the image to be processed, and this Gaussian distribution information can describe the distribution of the source control points in the image to be processed. For example, the source control point distribution parameter can be a parameter that directly represents the Gaussian distribution information of the source control points in the image to be processed. For example, the source control point distribution parameter can be the Gaussian distribution information centered on the source control points.
[0072] Among them, the correlation is an index used to measure the degree of correlation between the points in the current control point area and the source control points, and it can be expressed as similarity, feature distance, or other forms.
[0073] In the embodiments of the present application, through the Gaussian distribution of the source control points, the weights of the points related to the source control points in the current control point area can be increased, and the weights of the surrounding points of the points related to the source control points can be reduced, so as to suppress the scores of the surrounding points (i.e., background points) of the points related to the source control points, reduce the influence of the surrounding points of the points related to the source control points in the current control point area, increase the influence of the points related to the source control points in the current control point area, improve the accuracy of determining the updated current control points in the current control point area, and improve the quality of the generated processed image.
[0074] In some embodiments, the source control point distribution parameter can be a parameter that indirectly represents the Gaussian distribution characteristics of the source control points, such as a pre-trained convolution matrix. Specifically, the source control point distribution parameter includes a pre-trained convolution matrix, and the pre-trained convolution matrix is trained based on the Gaussian distribution information of the source control points in the image to be processed. According to the regional characteristics of the current control point area and the source control point distribution parameter, determining the correlation between the current control point area and the source control points includes:
[0075] Obtain the pre-trained convolution matrix;
[0076] Through the pre-trained convolution matrix, perform convolution calculation on the regional characteristics of the current control point area to determine the correlation between the current control point area and the source control points.
[0077] Among them, the convolution matrix refers to the weight matrix in the convolutional neural network, also known as the convolution kernel or filter. The convolution matrix can be used to perform convolution operations on the input, so as to extract the features of the input data. The pre-trained convolution matrix refers to the convolution matrix trained on a pre-defined task.
[0078] For example, the convolution calculation of the regional features can be implemented through the function G(F(Θ), z), where G represents a convolution function, Θ represents the surrounding area of the current control point p (i.e., the current control point area), F represents the current image feature, F(Θ) represents the regional feature of the current control point area, and z represents the pre-trained convolution matrix. The pre-trained convolution matrix can be a convolution filter with a size of 1×C×1×1, where C represents the number of feature convolution channels of the input. The result calculated by the function G(F(Θ), z) contains the correlation between each point in the current control point area and the source control point. In this way, by performing convolution calculation on the regional features through the pre-trained convolution matrix, the Gaussian distribution features of the source control points implicit in the pre-trained convolution matrix can be utilized to reduce the influence of the surrounding points of the points related to the source control point in the current control point area, so as to improve the accuracy of determining the updated current control point in the current control point area. In addition, compared with directly using the Gaussian distribution features of the source control points as the source control point distribution parameters, performing convolution calculation using the pre-trained convolution matrix has better transferability and generalization ability, and can be applied to image processing tasks under different types, different sizes, different shapes, and different lighting conditions.
[0079] In some embodiments, before obtaining the pre-trained convolution matrix, the Gaussian distribution information of the source control points can be used as the target value to train the pre-trained convolution matrix, so that the pre-trained convolution matrix learns the Gaussian distribution features of the source control points, enabling the pre-trained convolution matrix to better learn the features of the source control points in the image, so as to improve the accuracy of determining the updated current control point in the current control point area and improve the quality of the generated processed image. Specifically, the pre-trained convolution matrix is obtained through the following steps:
[0080] Obtain the initial convolution matrix and the Gaussian distribution information of the source control points in the image to be processed;
[0081] Use the Gaussian distribution information of the source control points in the image to be processed as the target value;
[0082] Through the initial convolution matrix, perform convolution calculation on the regional features of the initial control point area to determine the initial correlation between the points in the initial control point area and the source control point. The initial control point area is the area related to the source control point;
[0083] According to the difference between the initial correlation and the target value, train the initial convolution matrix to obtain the pre-trained convolution matrix.
[0084] Among them, the initial convolution matrix refers to the convolution matrix used to train the pre-trained convolution matrix. The initial convolution matrix can be a matrix set according to the application scenario or actual requirements, or a randomly initialized convolution matrix, which is not limited here.
[0085] Among them, the Gaussian distribution information refers to the information used to represent the Gaussian distribution characteristics of the source control points in the image to be processed. For example, the Gaussian distribution information of the source control points can be a Gaussian distribution graph centered on the control source control points. This Gaussian distribution graph can be expressed as a two-dimensional or multi-dimensional probability density function, where the probability density values present the shape of a Gaussian distribution and are centered on the control source control points. Specifically, the value of the Gaussian distribution function can be calculated according to the positions of the source control points in the image to be processed to generate the Gaussian distribution graph of the source control points. The shape and parameters of the Gaussian distribution graph can be set according to actual requirements or application scenarios. The position of the source control points corresponds to the center of the Gaussian distribution graph. The closer a point in the graph is to the source control points, the higher its probability density value, and the farther a point is from the source control points, the lower its probability density value.
[0086] Among them, the target value refers to the true output to be learned during the process of training the initial convolution matrix.
[0087] Among them, the initial correlation refers to the correlation between the points in the initial control point area and the source control points.
[0088] For example, the feature distance between the initial correlation and the target value can be calculated as the difference between the initial correlation and the target value, and this difference can be used as the loss to train the initial convolution matrix. Specifically, the following loss function can be used to calculate the difference between the initial correlation and the target value:
[0089] L = ∥G(F(Θ), z) - y∥ 2 ;
[0090] Among them, y represents the Gaussian distribution graph of the source control points, F(Θ) represents the regional features of the initial control point area, and z represents the initial convolution matrix. Thus, as Figure 1d shown in the training process of the convolution matrix, the initial correlation G(F(Θ), z) between the points in the initial control point area and the source control points is obtained by performing convolution calculation on the regional features of the initial control point area through the initial convolution matrix. The Euclidean distance (also known as the L2 distance) ∥G(F(Θ), z) - y∥ 2 between the initial correlation G(F(Θ), z) and the target value y is calculated to obtain the value of the loss function L. The initial convolution matrix is updated backward according to the value of the loss function L to make the loss function L as small as possible (i.e., the loss function L converges), and the finally updated convolution matrix is used as the pre-trained convolution matrix. Specifically, the initial convolution matrix can be updated backward through optimization algorithms such as backpropagation algorithm and gradient descent to obtain the pre-trained convolution matrix.
[0091] In some embodiments, after obtaining the pre-trained convolutional matrix corresponding to the image to be processed during training, the convolutional matrix is fixed during the subsequent processing of the image to be processed. That is, when updating the position at the current control point, the pre-trained convolutional matrix is no longer updated, so as not to increase the time for image processing and improve the processing efficiency.
[0092] 130. Determine the feature difference between the current control point region and the source control point according to the region feature of the current control point region and the control point feature of the source control point.
[0093] Among them, the control point feature refers to the feature representation of the control point in the image. In the embodiments of the present application, the control point feature of the source control point can be the feature representation of the source control point in the original feature representation of the image to be processed. For example, in some embodiments, the control point feature of the source control point can be obtained from the initial latent vector of the image to be processed. It can be understood that the initial latent vector of the image to be processed is in the form of a feature map. Therefore, according to the position and size of the source control point, the feature corresponding to the position and size of the source control point can be extracted from the initial latent vector of the image to be processed as the control point feature of the source control point.
[0094] Among them, the feature difference is an index used to measure the difference between features. For example, the feature difference can be expressed as a feature distance, a feature difference, an information entropy, an average absolute difference, or other forms.
[0095] In some embodiments, the feature difference degree between the region feature of the current control point region and the control point feature of the source control point can be measured by the feature distance. Specifically, determining the feature difference between the current control point region and the source control point according to the region feature of the current control point region and the control point feature of the source control point includes:
[0096] Obtain the control point feature of the source control point from the initial latent vector of the image to be processed;
[0097] Take the feature distance between the region feature of the current control point region and the control point feature of the source control point as the feature difference between the current control point region and the source control point.
[0098] For example, the feature corresponding to the position and size of the source control point can be extracted from the initial latent vector of the image to be processed as the control point feature of the source control point. Then, by calculating the Manhattan distance (also known as the L1 distance) between the region feature of the current control point region and the control point feature of the source control point, the feature difference between the current control point region and the source control point can be obtained. Specifically, the Manhattan distance between the region feature of the current control point region and the control point feature of the source control point can be calculated using the formula ∥F(Θ)-f∥1, where F(Θ) represents the region feature of the current control point region and f represents the control point feature of the source control point.
[0099] 140. Determine the updated current control point within the current control point area by combining the correlation and feature differences.
[0100] Among them, the updated current control point refers to the control point obtained by updating the position of the current control point. This updated current control point can be used as the current control point for the next position update.
[0101] For example, the correlation and feature differences can be combined to find a point in the current control point area that is most relevant to the source control point as the updated current control point. In this way, it can be considered that the position of the current control point corresponding to the source control point is updated to the updated current control point, thereby determining the updated current control point.
[0102] It should be noted that when there are multiple source control points, for any source control point, the correlation between the points in the current control point area corresponding to the source control point and the source control point, as well as the feature difference between the current control point area corresponding to the source control point and the source control point, can be determined respectively. Then, by combining the correlation and feature difference corresponding to the source control point, the updated current control point corresponding to the source control point is determined within the current control point area corresponding to the source control point. By analogy, the updated current control points corresponding to each source control point are determined.
[0103] In the embodiments of the present application, multiple rounds of position update processes of the current control point can be executed until the position of the updated current control point matches the specified target control point position, and then the process ends to generate the processed image. The specified target control point is a specific position point that the source control point needs to reach, and the specified target control point can be set according to the application scenario or actual requirements. For example, the specified target control point can be determined by the user's operation on the control points of the image to be processed. For example, Figure 1c As shown in the comparison schematic diagram of the image to be processed and the processed image, the user can set the source control point and the corresponding specified target control point on the face of the person in the image to be processed through the control point setting operation. In this way, through multiple rounds of position updates, the source control point can be dragged gradually closer until it reaches the specified target control point. During the dragging process, multiple current control points are passed through, and thus the image content at the source control point is gradually deformed to the specified target control point through image deformation, so as to realize operations such as anime character pose editing, scene editing, and character face editing in the image.
[0104] Obviously, in the embodiments of the present application, during the process of performing deformation processing on the image to be processed, the next current control point (i.e., the updated current control point) in the deformation processing is determined based on the correlation and feature difference between the current control point area and the source control points, so as to gradually drag the source control points based on the updated current control points to generate the processed image. In this process, through the correlation between the current control point area and the source control points, it can be ensured that the updated current control points have a high similarity with the source control points, so that the generated processed image can better retain the features of the source control points. At the same time, the degree of feature change is measured by the feature difference to control the degree of image deformation. In this way, by combining the correlation and feature difference between the current control point area and the source control points, it is possible to accurately control the degree of image deformation while retaining the features of the image to be processed, and improve the quality of the generated processed image.
[0105] In some embodiments, the updated current control point can be determined according to the confidence level, so as to preferentially select the point with a higher confidence level as the updated current control point, and improve the reliability of determining the updated current control point. Specifically, in combination with the correlation and the feature difference, determining the updated current control point within the current control point area includes:
[0106] Combining the correlation and the feature difference, calculate the confidence level of the points within the current control point area;
[0107] According to the confidence level, determine the updated current control point within the current control point area.
[0108] Among them, the confidence level refers to the degree of belief or credibility in a certain result or hypothesis, and the confidence level can be used to represent the evaluation of the possibility of a certain event occurring. In the embodiments of the present application, the confidence level can be used to represent the reliability of the points within the current control point area as the updated current control point.
[0109] For example, perform a weighted sum on the correlation between the current control point area and the source control points and the feature difference between the current control point area and the source control points to calculate the confidence level of each point within the current control point area. Among them, an exponential operation can be performed on the negative value of the feature difference, so that the greater the feature difference, the smaller the confidence level. Specifically, the confidence level of the points within the current control point area can be calculated by the following formula:
[0110]
[0111] Among them, S(Θ) represents the confidence of all points within the current control point area, λ is a balance coefficient, which is used to balance the weight relationship between various terms in multi-objective optimization or polynomials. λ can be set according to actual requirements or application scenarios. G(F(Θ), z) represents the correlation between the current control point area and the source control point, and ∥F(Θ) - f∥1 represents the feature difference between the current control point area and the source control point. Thus, the point with the highest confidence can be found from S(Θ) as the updated current control point.
[0112] 150. According to the updated current control point, update the current latent vector to obtain the updated current latent vector, so as to generate a processed image based on the updated current control point and the updated current latent vector.
[0113] For example, in the multi-round position update process of the current control point in the embodiments of the present application, the current latent vector of the image to be processed can be synchronously updated, so that each time feature learning is performed on the updated current latent vector of the image to be processed, making the processed image more natural and improving the quality of the generated processed image.
[0114] In some embodiments, during the multi-round position update process of the current control point, the current latent vector can be continuously updated to use the updated current latent vector as the latent vector for the next position update. Specifically, according to the updated current control point, update the current latent vector to obtain the updated current latent vector, so as to generate a processed image based on the updated current control point and the updated current latent vector, including:
[0115] According to the updated current control point, update the current latent vector to obtain the updated current latent vector;
[0116] Use the updated current control point as the current control point, and the updated current latent vector as the current latent vector, and return to execute the steps of obtaining the regional features of the current control point area from the image features of the image to be processed and subsequent steps until the updated current control point matches the specified target control point, and generate a processed image from the updated current latent vector.
[0117] Among them, each element in the latent vector of the image to be processed represents a certain feature of the image. Therefore, by modifying different elements of the latent vector, various aspects of the generated image can be controlled, such as color, texture, and shape, etc., to update the latent vector. For example, in the embodiments of the present application, the current latent vector can be updated by dragging or modifying the elements corresponding to the image pixel points.
[0118] For example, during the position update process of multiple rounds of current control points, the current control points and the latent vectors can be continuously updated. The updated current control points and the updated current latent vectors are respectively used as the current control points and the latent vectors for the next position update. During the next position update, feature learning and subsequent processing are performed on the latent vector to obtain the updated current control points for the next position update. The process ends until the current control points gradually approach and reach the specified target control points through multiple position updates. The current latent vector at the last position update (i.e., the current latent vector used to obtain the current image features at the last time) is decoded to generate the processed image.
[0119] In some embodiments, the regional features of different control point regions can be selected according to the confidence of the updated current control points to supervise the position update process of the current control points, so as to better adapt to the learning process under different confidences and improve the reliability of the generated processed image. Specifically, according to the updated current control points, the current latent vector is updated to obtain the updated current latent vector, including:
[0120] According to the confidence of the updated current control points, the regional features of the reference control point region corresponding to the confidence of the updated current control points are obtained. The confidence is obtained from the correlation and the feature difference. The reference control point region includes one of the current control point region and the initial control point region. The initial control point region is the region related to the source control point;
[0121] According to the regional features of the reference control point region and the regional features corresponding to the updated current control points, the loss value is determined;
[0122] According to the loss value, the current latent vector is updated to obtain the updated current latent vector.
[0123] Among them, the regional features corresponding to the current control points refer to the regional features of the control point region corresponding to the current control points. For example, the control point region can be the current control point region with a preset shape centered on the current control point or other control point regions related to the current control points.
[0124] For example, according to the confidence of the updated current control points, the control point region corresponding to the current control point region or the initial control point region can be selected as the reference control point region to calculate the loss value, and the current latent vector is updated in reverse according to the loss value to obtain the updated current latent vector. Specifically, the current latent vector can be updated by optimization algorithms such as backpropagation algorithm and gradient descent to obtain the updated current latent vector.
[0125] In some embodiments, when the confidence of the updated current control point is relatively high, the regional feature corresponding to the current control point area can be selected to supervise the position update process of the current control point. When the confidence of the updated current control point is relatively low, the regional feature of the initial control point area can be selected to supervise the position update process of the current control point. In this way, when the confidence is low, the original features of the image to be processed are used to supervise the position update process of the current control point, so as to prevent the control point features of the current control point from deviating from the original features of the image to be processed, retain the features of the image to be processed, and increase the consistency between the generated processed image and the image to be processed. Specifically, according to the confidence of the updated current control point, the regional feature of the reference control point area corresponding to the confidence of the updated current control point is obtained, including:
[0126] If the confidence of the updated current control point is greater than or equal to the preset threshold, obtain the regional feature of the current control point area;
[0127] If the confidence of the updated current control point is less than the preset threshold, obtain the regional feature of the initial control point area.
[0128] Wherein, the preset threshold is a value used to judge the magnitude of the confidence of the updated current control point, and this value can be set according to actual requirements or application scenarios.
[0129] For example, when the confidence is greater than or equal to the preset threshold, the feature distance between the regional feature corresponding to the current control point area and the regional feature corresponding to the updated control point can be calculated to obtain a loss value. When the confidence is less than the preset threshold, the feature distance between the regional feature of the initial control point area and the regional feature corresponding to the updated control point can be calculated to obtain a loss value. In this way, in the multi-round position update process of the current control point, in each position update process, due to the different confidences of the current control point after the current update, different regional features can be selected to calculate the loss value, so as to dynamically select the corresponding supervision method according to the confidence during the process of dragging the source control point to the target control point, so as to prevent the control point features of the current control point from deviating from the original features of the image to be processed.
[0130] In some embodiments, the preset threshold can be determined according to the confidence of the updated current control point obtained in the first update. For example, during the multi-round position update process of the current control point, the confidence of the updated current control point determined during the first position update (i.e., when the current control point is the source control point) can be used as the preset threshold, or the updated current control point can be multiplied by a preset weight parameter to obtain the preset threshold, and the preset weight parameter is used to control the intensity of the preset threshold.
[0131] In some embodiments, before determining the loss value based on the regional features of the reference control point region and the regional features corresponding to the updated current control point, the regional features corresponding to the updated current control point used to calculate the loss value may be determined based on the offset of the updated current control point relative to the specified target control point, so that when the current control point is iteratively updated based on the loss value, the current control point can converge to the specified target control point until the current control point finally reaches the target control point, thereby improving the efficiency and accuracy of processing the processed image. Specifically, the regional features corresponding to the updated current control point are obtained through the following steps:
[0132] Determine the offset of the updated current control point relative to the specified target control point according to the feature distance between the specified target control point and the updated current control point;
[0133] Determine the offset control point corresponding to the updated current control point according to the updated current control point and the offset;
[0134] Obtain the regional features of the offset control point region from the current image features of the image to be processed, and use the regional features of the offset control point region as the regional features corresponding to the updated current control point. The offset control point region is the region related to the offset control point.
[0135] Wherein, the offset control point refers to the point obtained by offsetting the updated current control point according to the offset.
[0136] For example, it can be achieved through Calculate the offset of the updated current control point relative to the specified target control point. Here, d is the offset, t represents the specified target control point, and p represents the updated current control point. Thus, calculate the Euclidean distance (i.e., L2 distance) between the updated current control point and the specified target control point, and use it as the offset of the updated current control point relative to the specified target control point. Then, based on the position of the updated current control point p plus the offset d, determine the offset control point p + d after the updated current control point is offset according to the offset. Then, according to the position of the offset control point p + d, obtain the regional feature of the area with a preset shape centered on the offset control point p + d (i.e., the offset control point area) from the current image features, and use the obtained regional feature as the regional feature corresponding to the updated current control point. It should be noted that in the process of updating the position of the current control point in multiple rounds in the embodiments of the present application, the current hidden vector can be reversely updated through the loss value until the loss value converges or the updated control point reaches the specified target control point. It can be understood that by calculating the loss value using the regional feature of the reference control point area and the regional feature of the offset control point area, when the loss value reversely updates the current hidden vector, the updated current control point can gradually approach and reach the specified target control point along the direction pointing to the target control point (i.e., the direction corresponding to the offset), so that the current control point can converge to the specified target control point until the current control point finally reaches the target control point, thereby improving the efficiency and accuracy of processing the processed image.
[0137] In some embodiments, when setting the source control point, a deformation area can also be set in the image to be processed to constrain the image to be processed to deform only within the deformation area, so that the area outside the deformation area remains unchanged, thereby improving the accuracy and controllability of generating the processed image. Specifically, the image to be processed further includes a preset deformation area, and the source control point and the specified target control point are located within the preset deformation area. Determining the loss value according to the regional feature of the reference control point area and the regional feature of the updated control point area includes:
[0138] Determine the difference in the regional feature of the control point area according to the regional feature of the reference control point area and the regional feature corresponding to the updated current control point;
[0139] Determine the difference in the regional feature of the deformation area according to the regional feature of the preset deformation area in the initial image features and the regional feature of the preset deformation area in the current image features. The initial image features are obtained from the initial hidden vector of the image to be processed;
[0140] Combine the difference in the regional feature of the control point area and the difference in the regional feature of the deformation area to determine the loss value.
[0141] Among them, the initial image feature refers to the image feature obtained from the initial latent vector of the image to be processed. For example, during the process of updating the position of the current control point in multiple rounds, the image feature (i.e., the initial image feature) obtained by performing feature learning on the initial latent vector during the first position update can be directly obtained for use in the loss value calculation process.
[0142] Among them, the preset deformation area refers to a specific area in the image to be processed that is preset and expected to be deformed. In some embodiments, when the user performs the control point setting operation, the deformation area setting operation can also be performed to set a deformation area including the source control point and the specified target control point in the image to be processed. For example, the deformation area setting operation may include, but is not limited to, dragging the mouse or finger to draw one or more box selection areas in the image to be processed as the preset deformation area, or inputting the coordinate information of the deformation control point area to set the position of the preset deformation area in the image to be processed.
[0143] Among them, the control point area feature difference is an index for measuring the difference between the reference control point area and the control point area corresponding to the updated current control point. The deformation area feature difference is an index for measuring the difference between the preset deformation area in the initial image feature and the preset deformation area in the current image feature. For example, the control point area feature difference and the deformation area feature difference can be expressed as feature distance, feature difference, information entropy, mean absolute difference or other forms.
[0144] For example, the control point area feature difference can be obtained by calculating the feature distance between the area feature of the reference control point area and the area feature of the control point area corresponding to the updated current control point. And the deformation area feature difference can be obtained by calculating the feature distance between the area feature of the deformation area in the initial image feature and the area feature of the deformation area in the current image feature. The control point area feature difference and the deformation area feature difference can be weighted and summed to obtain the loss value. Therefore, the area feature of the deformation area is incorporated into the loss value calculation process to constrain the image to be processed to deform only within the deformation area. As Figure 1c shown in the comparison schematic diagram of the image to be processed and the processed image, the head of the person in the left image to be processed is set as the deformation area. Therefore, in the generated processed image on the right, as the source control point is dragged to the target control point, the head of the person in the deformation area deflects to the right, while the other non-deformation areas remain unchanged.
[0145] Specifically, if the confidence level is greater than the preset threshold, the loss value can be calculated through the following loss function:
[0146]
[0147] Wherein, n represents the number of source control points (i.e., there are n current control points and n updated current control points), F0 represents the initial image feature, F represents the current image feature obtained from the current hidden vector in the current position update, p i represents the i-th current control point, Θ(p i ) represents the surrounding area of the i-th current control point (i.e., the control point area corresponding to the i-th current control point), F(Θ(p i )) represents the area feature of the control point area corresponding to the i-th current control point (i.e., the area feature of the current control point area), p i represents the updated current control point, d i is the offset, p i represents the t-th specified target control point, p i +d i represents the point determined by adding the offset to the i-th updated current control point (i.e., the i-th offset control point), Θ(p i +d i ) represents the surrounding area of the i-th updated current control point (i.e., the control point area corresponding to the i-th updated current control point), F(Θ(p i +d i )) represents the area feature of the control point area corresponding to the i-th offset control point (i.e., the area feature of the offset control point area), M is a predefined mask, M is used to control the deformation area in the image to be processed, ε is a weight coefficient, and ε can be set according to actual requirements or application scenarios.
[0148] Therefore, in the above formula, the Manhattan distance (also known as the L1 distance) between the control point area corresponding to any updated current control point in the current control point area and the control point area corresponding to the corresponding updated current control point's offset control point is calculated by ||F(Θ(p i )) - F(Θ(p i +d i ))||1, and then the sum of the Manhattan distances corresponding to all current control points is obtained to get the control point area feature difference. And, the feature difference between the initial image feature and the current image feature can be calculated by (F - F0), and then the inverse mask (1 - M) of the mask M corresponding to the preset deformation area is multiplied element by element with this feature difference (F - F0), and the L1 norm (i.e., the Manhattan distance) of the multiplication result is calculated to obtain the deformation area feature difference. Then, the control point area feature difference and the deformation area feature difference are weighted and summed through the preset weight coefficient ε to obtain the loss value. Among them, mask processing refers to the method of using a mask to selectively operate on specific areas of an image. In the embodiments of the present application, the deformation area in the image to be processed can be defined by setting a mask, so as to obtain the features of the deformation area from the feature map through mask processing.
[0149] Specifically, if the confidence level is less than or equal to a preset threshold, the loss value can be calculated through the following loss function:
[0150] L2 = ||F0(Θ(p0)) - F(Θ(p + d))||1 + ε‖(F - F0)·(1 - M)‖1;
[0151] Wherein, F0 represents the initial image feature, F represents the current image feature obtained from the current hidden vector in the current position update, p0 represents the source control point, Θ(p0) represents the surrounding area of the source control point Θ(p0) (i.e., the initial control point area), F0(Θ(p0)) represents the area feature of the control point area corresponding to the source control point (i.e., the area feature of the initial control point area), p represents the updated current control point, p + d represents the point determined by adding the offset to the updated current control point (i.e., the offset control point), d is the offset, t represents the specified target control point, Θ(p + d) represents the surrounding area of the updated current control point (i.e., the control point area corresponding to the updated current control point), F(Θ(p + d)) represents the area feature of the control point area corresponding to the offset control point (i.e., the area feature of the offset control point area), that is, M is a predefined mask, M is used to control the deformation area in the image to be processed, ε is a weight coefficient, and ε can be set according to actual requirements or application scenarios.
[0152] Therefore, in the above formula, the Manhattan distance (also known as the L1 distance) between the initial control point area corresponding to any source control point in the initial control point area and the offset control point area corresponding to the corresponding updated current control point is calculated through ||F0(Θ(p0)) - F(Θ(p + d))||1, and then the Manhattan distances corresponding to all current control points are summed to obtain the control point area feature difference. And, the feature difference between the initial image feature and the current image feature can be calculated through (F - F0), and then the inverse mask (1 - M) of the mask M corresponding to the preset deformation area is multiplied element by element with the feature difference (F - F0), and the L1 norm (i.e., the Manhattan distance) of the multiplication result is calculated to obtain the deformation area feature difference. Then, the loss value is obtained by weighted summing the control point area feature difference and the deformation area feature difference through a preset weight coefficient ε.
[0153] The image processing solution provided by the embodiments of the present application can be applied to various image processing scenarios. For example, taking the processing of human images as an example, from the current image features of the image to be processed, the regional features of the current control point area are obtained. The current image features are obtained from the current latent vector of the image to be processed. The image to be processed includes source control points and corresponding current control points. The current control point area is the area related to the current control point; according to the regional features of the current control point area and the source control point distribution parameters, the correlation between the current control point area and the source control points is determined. The source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed; according to the regional features of the current control point area and the control point features of the source control points, the feature difference between the current control point area and the source control points is determined; combining the correlation and the feature difference, the updated current control point is determined within the current control point area; according to the updated current control point, the current latent vector is updated to obtain the updated current latent vector, so as to generate the processed image according to the updated current control point and the updated current latent vector.
[0154] As can be seen from the above, in the embodiments of the present application, during the process of deforming the image to be processed, the next current control point (i.e., the updated current control point) in the process is determined through the correlation and feature difference between the current control point area and the source control points, so as to gradually drag the source control points based on the updated current control points to generate the processed image. In this process, through the correlation between the current control point area and the source control points, it can be ensured that the updated current control points have a high similarity with the source control points, so that the generated processed image can better retain the features of the source control points. At the same time, the degree of feature change is measured through the feature difference to control the degree of image deformation. In this way, by combining the correlation and feature difference between the current control point area and the source control points, the degree of image deformation can be accurately controlled while retaining the features of the image to be processed, improving the quality of the generated processed image. In addition, the embodiments of the present application also determine the correlation between the current control point area and the source control points based on the source control point distribution parameters representing the Gaussian distribution of the source control points in the image to be processed, so as to increase the influence of the points related to the source control points in the current control point area by using the Gaussian distribution information of the source control points, improve the accuracy of determining the updated current control point within the current control point area, and improve the efficiency and quality of the generated processed image.
[0155] According to the method described in the above embodiments, further detailed description will be made below.
[0156] In this embodiment, taking the image processing process based on an image generation model as an example, the method of the embodiments of the present application will be described in detail.
[0157] The image processing method of the embodiments of the present application can be implemented by an image processing system, such as Figure 2aSchematic diagram of the image processing system shown, which may include a front end and a back end. The back end includes an image generation module, a point following module, and a motion supervision module.
[0158] As Figure 2b shown, the specific process of an image processing method is as follows:
[0159] 210. Set source control points, target control points, and a deformation area in the image to be processed.
[0160] For example, for an image in an initial state given by a creator of an image, video, animation, or other service (i.e., the image to be processed), the creator can upload the image to be processed on an image setting page as Figure 2c shown. The image setting page is displayed at the front end of the image processing system. The creator can also perform pre - image - processing settings on the image to be processed displayed in the input box on the left side of the image setting page according to their specific needs. As shown in the figure, set source control points and target control points on the face of the image character, and set the character's head as the deformation area, and can input parameters for the image processing process in the setting box below the image setting page, such as image generation model parameters, control point area size, etc., so that the finally generated image meets the creation requirements. With the separate permission or consent of the image right holder, by clicking the run button on the image setting page, the processing of the image to be processed can start in the back end of the image processing system.
[0161] 220. Through the image generation model, perform feature learning on the current latent vector of the image to be processed, and use the feature vector output by the middle layer of the image generation model as the current image feature of the image to be processed.
[0162] For example, the back end of the image processing system can perform multi - step dragging (position updating) on the source control points set by the user until they reach the target control points to achieve drag - and - drop image editing processing. The processing flow of dragging control points at any step (such as the t - th step) by the back end of the image processing system is as Figure 2d shown. Among them, drag - and - drop image editing refers to the process of dragging the image content of the source control point to the position of the target point to generate a logically reasonable and high - quality image for the image provided by the user and the defined initial source control points and target control points.
[0163] Specifically, as Figure 2dFor the processing flow of dragging any control point shown, for any t-step drag, the image generation module can perform feature learning on the current hidden vector of the t-th step of the image to be processed through the image generation model, and use the feature vector output by the intermediate layer of the image generation model as the current image feature of the image to be processed. Among them, the generation model generates new samples by learning the distribution of data. A typical application of the generation model is image generation, where the model is trained to generate new images similar to the images in the training set. Generation models usually use deep neural networks to implement and can be trained using different loss functions and optimization algorithms. Some commonly used generation models include variational autoencoders (VAEs), generative adversarial networks (GANs), and manifold generation models (MGMs), etc.
[0164] The image generation model adopted in the embodiments of this application can be a generative adversarial network (GAN), a diffusion model (Diffusion). When the image generation model is a generative adversarial network, the input image to be processed is encoded and the generator of the generative adversarial network generates intermediate features (i.e., the current image features). When the image generation model is a diffusion model, first LoRA is trained according to the input image to be processed, and then the intermediate features (i.e., the current image features) generated by passing the image to be processed through the encoder and the diffusion process are obtained.
[0165] 230. Obtain the regional feature of the current control point region from the current image feature of the image to be processed, where the current control point region is the region related to the current control point.
[0166] 240. Determine the correlation between the current control point region and the source control point according to the regional feature of the current control point region and the source control point distribution parameters.
[0167] 250. Determine the feature difference between the current control point region and the source control point according to the regional feature of the current control point region and the control point feature of the source control point.
[0168] 260. Combine the correlation and the feature difference to calculate the confidence of the points in the current control point region.
[0169] 270. Determine the updated current control point within the current control point region according to the confidence.
[0170] For example, in the embodiments of the present application, for the dragging process of control points, a new point following direction based on discriminative learning is designed (for following the updated control points for each step of dragging). By combining the original feature difference with the following score generated by the discriminative following model, the confidence is determined, thereby improving the point following accuracy and dragging accuracy. In the process of drag-and-drop image editing, due to the uncertainty of optimization, it is impossible to directly judge the position reached by the control source control point. Therefore, a following module is required to locate and follow the position of the current control source control point, that is, a point follower. The point follower in the embodiments of the present application is a point following module. Discriminative learning refers to a learning method that uses a discriminative approach to distinguish the target to be discriminated from similar targets, aiming to increase the distance between the target and the interference items.
[0171] Specifically, as Figure 2d shown in the processing flow of any step of dragging a control point, for any t-step dragging, the point following module can determine the confidence of points within the current control point area through the point following model, and use the point with the highest confidence within the current control point area as the updated current control point. In the embodiments of the present application, the specific form of the point following model is the convolution kernel of a convolutional layer, and its output serves as the confidence score for point following (that is, the confidence of points within the current control point area). Specifically, the confidence score S of the point following model is obtained according to the following formula:
[0172]
[0173] where λ is a balance coefficient, f represents the feature of the control source control point in the initial state (the control point feature of the source control point), and ∥F(Θ)-f∥1 represents the difference between the feature of the current control point area at the t-th step and the feature of the source control point in the initial state (that is, the feature difference between the current control point area and the source control point). The point with the highest confidence score obtained through S is used as the position reached by the current control point at the t-th step, that is, the updated current control point.
[0174] In the above formula, the point following model learns a function G(F(Θ),z), where G represents a convolutional function, Θ is the local area around the current control point p (that is, the current control point area), F represents the feature of the intermediate layer (that is, the image feature of the image to be processed), and z is the learned following model (that is, the pre-trained convolutional matrix). If the following model z matches the content at a certain position and identifies it as the updated current control point p, a higher confidence score is returned; otherwise, a lower confidence score is returned. In particular, since the following model z is learned before the latent vector optimization and remains unchanged in all dragging steps, this method hardly adds extra running time required for image editing. Finally, the classification score generated by the following model is combined with the original feature difference score to obtain the confidence score, so as to achieve discriminative and accurate point positioning.
[0175] In the embodiment of the present application, the learning process of the following model (i.e., the convolutional matrix) z is carried out before the operation process. This model is a convolutional filter of size 1×C×1×1, where C represents the number of input feature convolutional channels. The learning process of the following model z, in this learning process, f in the above is used to initialize z, and the weights are updated under the guidance of the following loss function:
[0176] L = ∥G(F(Θ), z) - t∥ 2 ;
[0177] where y is a Gaussian distribution map centered on the initial source control point (i.e., the Gaussian distribution map of the source control point), which is used as the target value for training. The learning process of the following model z only needs to optimize the following model z through this loss function. Through optimization, the confidence performance of controlling the source control point is enhanced, and at the same time, the confidence scores of interference points such as background points are suppressed. In the subsequent operation steps, the following model z remains unchanged.
[0178] 280. Update the current latent vector according to the updated current control point to obtain the updated current latent vector.
[0179] For example, in the embodiment of the present application for the dragging process of the control point, based on the following score, a confidence-based latent vector enhancement strategy is further explored to achieve more complete motion supervision. Motion supervision refers to an online loss function used to control the movement of the control point content, and the single-step dragging operation is achieved by optimizing this loss function. This loss function can realize the motion supervision of the control point by supervising the update of the current latent vector. As can be seen from the schematic diagrams of the current image features corresponding to the images at the t-th step and the schematic diagrams of the current image features corresponding to the images at the (t + 1)-th step shown in the processing flow of any step of dragging the control point as shown in Figure 2d each time the current latent vector is updated, the corresponding image changes, so the step-by-step dragging of the control point can be realized by updating the latent vector.
[0180] Specifically, as shown in the processing flow of any step of dragging the control point as shown in Figure 2d for any t-step dragging, the motion supervision module can update the current latent vector by using a confidence-based latent vector enhancement strategy to make each step of dragging (i.e., position update) more stable. The motion supervision module introduces the maximum value of the following confidence score to represent the supervision quality of the current dragging step, and uses the maximum confidence score of the first-step dragging to generate a threshold to determine whether to use latent vector enhancement. When the current state editing quality is high enough (i.e., the confidence is greater than or equal to the preset threshold), the loss value can be calculated through the following loss function to update the current latent vector:
[0181]
[0182] When the editing quality in the current state is low (i.e., the confidence is less than the preset threshold), the loss value can be calculated through the following loss function to update the current latent vector:
[0183] L2 = ||F0(Θ(p0)) - F(Θ(p + d))||1 + ε‖(F - F0) • (1 - M)‖1;
[0184] Where F0 represents the intermediate feature in the initial state (i.e., the initial image feature), Θ(p0) represents the surrounding area of the initial control point p0 (i.e., the initial control point area), p is the control point in the current dragging step (i.e., the updated current control point), d is the set offset for this step, ε is the weight coefficient, and M is the deformation area input by the user (i.e., the deformation area). In this way, it is possible to prevent the current content of the control point from deviating significantly from the template in the initial state, thus achieving high-quality motion supervision.
[0185] 290. Take the updated current control point as the current control point, and take the updated current latent vector as the current latent vector, and return to execute steps 220 to 290 until the updated current control point matches the specified target control point, and generate the processed image from the updated current latent vector.
[0186] For example, as Figure 2d shown in the processing flow of dragging the control point at any step. For any t-step dragging, after determining the updated current control point and the updated current latent vector, the updated current control point at the t-th step can be taken as the current control point at the (t + 1)-th step, and the updated current latent vector at the t-th step can be taken as the current latent vector at the (t + 1)-th step, and execute the next dragging until the current control point reaches the specified target control point. At this time, the image generation module can decode the current latent vector in the last dragging through the decoder in the image generation model to generate the processed image. It should be noted that in multi-step dragging, when the updated current control point determined in step 270 matches the position of the target control point, it is regarded as the current control point reaching the target control point, and thus the process can be ended (i.e., no longer execute the subsequent steps 280 and 290), and the updated latent vector obtained in the last update is decoded to generate the processed image.
[0187] Specifically, when the image generation model is a generative adversarial network, the generative adversarial network locates the position of the current control source control point using the discriminative following method and generates a confidence score. It performs complete motion supervision based on the confidence score, and finally passes the edited latent vector (i.e., the updated current latent vector obtained in the last drag) through the decoder to generate the processed image. When the image generation model is a diffusion model, the above-mentioned discriminative point following and complete motion supervision are performed multiple times on the generated intermediate features, and finally the modified latent vector (i.e., the updated current latent vector obtained in the last drag) is subjected to inverse diffusion denoising and decoded to generate the processed image.
[0188] It should be noted that when the image generation model is a generative adversarial network, the image processing method of the embodiment of the present application is more proficient in dealing with larger deformations and more creative content generation, such as turning a lion with its mouth closed into a roaring state. When the image generation model is a generative adversarial network, the image processing method of the embodiment of the present application is more proficient in generating high-quality and high-fidelity editing results. By adopting different image generation models, the image processing method of the embodiment of the present application can be adapted to various scenarios, including but not limited to image, video, animation or other business creation tasks, and including but not limited to anime character pose editing, scene editing, face swapping or facial feature replacement in them. To solve the problem of point-to-point fine image processing in practical applications, thereby assisting image creation work in various application scenarios, effectively improving the efficiency and diversity of creation. As Figure 2e As can be seen from the comparison schematic diagram of the image processing results of the present application and the prior art for different types of images provided by the embodiment of the present application shown in the figure, for various different types of images such as landscape images, animal images, home images, and human images, the distance between the currently updated control point and the specified target control point in the processed image generated by the prior art for processing the image to be processed is relatively far, while the distance between the currently updated control point and the specified target control point in the processed image generated by the image method provided by the embodiment of the present application for processing the image to be processed is closer, and basically the two points can completely overlap, which shows that the image method provided by the embodiment of the present application can more precisely control the degree of image deformation and improve the quality of the generated processed image. Especially in the processing result of the home image in the figure, the processed image obtained by the prior art does not achieve the expected effect of deforming the lamp in the figure, while the position where the lamp finally deforms in the processed image obtained by the present application completely fits the specified target control area.
[0189] In addition, when the image processing system processes in the background to obtain the processed image, it can be Figure 2c displayed in the output box on the right side of the image setting page as shown in the figure, and the processed image can be modified. The creator can perform subsequent creation tasks based on the modified image.
[0190] As can be seen from the above, the embodiment of the present application can suppress the following confidence scores of interference points (background points) and improve the following confidence scores of source control points through a simple but more robust following model. At the beginning of dragging, the weights of the following model are updated under the supervision of a customized similarity learning function. Then, the following model is combined with the original feature difference method to achieve robust and accurate point following. On the other hand, through a confidence-based latent vector enhancement strategy, motion supervision is made more complete at each dragging step, thereby achieving a stable image processing effect.
[0191] To better implement the above method, the embodiment of the present application also provides an image processing device, which can be specifically integrated in an electronic device. The electronic device can be a terminal, a server, or other devices. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or other devices; the server can be a single server or a server cluster composed of multiple servers.
[0192] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the image processing device specifically integrated in the server as an example.
[0193] For example, as Figure 3 shown, the image processing device may include an acquisition unit 310, a correlation determination unit 320, a difference determination unit 330, a control point determination unit 340, and an update unit 350, as follows:
[0194] (1) Acquisition unit 310
[0195] It is used to obtain the regional features of the current control point area from the current image features of the image to be processed. The current image features are obtained from the current latent vector of the image to be processed. The image to be processed includes source control points and current control points corresponding to the source control points. The current control point area is the area related to the current control point.
[0196] In some embodiments, the image processing device further includes a feature learning unit, and the feature learning unit is used to: obtain an image generation model; through the image generation model, perform feature learning on the current latent vector of the image to be processed, and use the feature vector output by the middle layer of the image generation model as the current image features of the image to be processed.
[0197] (2) Correlation determination unit 320
[0198] It is used to determine the correlation between the current control point area and the source control points according to the regional features of the current control point area and the source control point distribution parameters. The source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed.
[0199] In some embodiments, the source control point distribution parameter includes a pre-trained convolutional matrix, which is trained based on the Gaussian distribution information of the source control points in the image to be processed. The correlation determination unit is specifically configured to: obtain the pre-trained convolutional matrix; perform convolutional calculation on the regional features of the current control point region through the pre-trained convolutional matrix to determine the correlation between the current control point region and the source control points.
[0200] In some embodiments, the image processing device further includes a pre-training unit, which is configured to: obtain an initial convolutional matrix and the Gaussian distribution information of the source control points in the image to be processed; use the Gaussian distribution information of the source control points in the image to be processed as the target value; perform convolutional calculation on the regional features of the initial control point region through the initial convolutional matrix to determine the initial correlation between the points in the initial control point region and the source control points, where the initial control point region is the region related to the source control points; train the initial convolutional matrix according to the difference between the initial correlation and the target value to obtain the pre-trained convolutional matrix.
[0201] (III) Difference determination unit 330
[0202] It is configured to determine the feature difference between the current control point region and the source control points according to the regional features of the current control point region and the control point features of the source control points.
[0203] In some embodiments, the difference determination unit is specifically configured to: obtain the control point features of the source control points from the initial hidden vector of the image to be processed; use the feature distance between the regional features of the current control point region and the control point features of the source control points as the feature difference between the current control point region and the source control points.
[0204] (IV) Control point determination unit 340
[0205] It is configured to combine the correlation and the feature difference to determine the updated current control point within the current control point region.
[0206] In some embodiments, the control point determination unit is specifically configured to: combine the correlation and the feature difference to calculate the confidence of the points within the current control point region; determine the updated current control point within the current control point region according to the confidence.
[0207] (V) Update unit 350
[0208] It is configured to update the current hidden vector according to the updated current control point to obtain the updated current hidden vector, so as to generate the processed image according to the updated current control point and the updated current hidden vector.
[0209] In some embodiments, the updating unit is specifically configured to: update the current latent vector according to the updated current control point to obtain an updated current latent vector; use the updated current control point as the current control point and the updated current latent vector as the current latent vector, and return to execute the subsequent steps of obtaining the regional feature of the current control point area from the image features of the image to be processed until the updated current control point matches the specified target control point, and generate a processed image from the updated current latent vector.
[0210] In some embodiments, updating the current latent vector according to the updated current control point to obtain an updated current latent vector includes: obtaining the regional feature of the reference control point area corresponding to the confidence level of the updated current control point according to the confidence level of the updated current control point, where the confidence level is obtained from the correlation and the feature difference, and the reference control point area includes one of the current control point area and the initial control point area, and the initial control point area is the area related to the source control point; determining a loss value according to the regional feature of the reference control point area and the regional feature corresponding to the updated current control point; and updating the current latent vector according to the loss value to obtain an updated current latent vector.
[0211] In some embodiments, obtaining the regional feature of the reference control point area corresponding to the confidence level of the updated current control point includes: if the confidence level of the updated current control point is greater than or equal to a preset threshold, obtaining the regional feature of the current control point area; if the confidence level of the updated current control point is less than the preset threshold, obtaining the regional feature of the initial control point area.
[0212] In some embodiments, the image to be processed further includes a preset deformation area, and the source control point and the specified target control point are located within the preset deformation area. Determining a loss value according to the regional feature of the reference control point area and the regional feature corresponding to the updated current control point includes: determining the difference in control point area features according to the regional feature of the reference control point area and the regional feature corresponding to the updated current control point; determining the difference in deformation area features according to the regional feature of the preset deformation area in the initial image features and the regional feature of the preset deformation area in the current image features, where the initial image features are obtained from the initial latent vector of the image to be processed; and determining a loss value by combining the difference in control point area features and the difference in deformation area features.
[0213] In some embodiments, the regional feature corresponding to the updated current control point is obtained through the following steps: determining the offset of the updated current control point relative to the specified target control point according to the feature distance between the specified target control point and the updated current control point; determining the offset control point corresponding to the updated current control point according to the updated current control point and the offset; and obtaining the regional feature of the offset control point area from the current image features of the image to be processed, so as to use the regional feature of the offset control point area as the regional feature corresponding to the updated current control point, where the offset control point area is the area related to the offset control point.
[0214] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments and will not be elaborated herein.
[0215] Thus, in the process of deforming the image to be processed in the embodiments of the present application, the next current control point (i.e., the updated current control point) in the process is determined through the correlation and feature difference between the current control point area and the source control point, so as to gradually drag the source control point based on the updated current control point to generate the processed image. In this process, through the correlation between the current control point area and the source control point, it can be ensured that the updated current control point has a high similarity with the source control point, so that the generated processed image can better retain the features of the source control point. At the same time, the feature difference is used to measure the degree of feature change to control the degree of image deformation. In this way, by combining the correlation and feature difference between the current control point area and the source control point, it is possible to accurately control the degree of image deformation while retaining the features of the image to be processed, and improve the quality of the generated processed image. In addition, the embodiments of the present application also determine the correlation between the current control point area and the source control point based on the Gaussian distribution source control point distribution parameter representing the source control point in the image to be processed, so as to use the Gaussian distribution information of the source control point to increase the influence of the points related to the source control point in the current control point area, improve the accuracy of determining the updated current control point in the current control point area, and improve the efficiency and quality of the generated processed image.
[0216] The embodiments of the present application also provide an electronic device, which can be a device such as a terminal or a server. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0217] In some embodiments, the image processing device can also be integrated in multiple electronic devices. For example, the image processing device can be integrated in multiple servers, and the image processing method of the present application is implemented by multiple servers.
[0218] In this embodiment, the electronic device in this embodiment will be described in detail by taking the server as an example. For example, as Figure 4 shown, it shows a schematic structural diagram of the server involved in the embodiments of the present application. Specifically:
[0219] The server may include components such as a processor 410 with one or more processing cores, a memory 420 with one or more computer-readable storage media, a power supply 430, an input module 440, and a communication module 450. Those skilled in the art can understand that Figure 4 the server structure shown in does not constitute a limitation on the server, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0220] The processor 410 is the control center of the server, connecting various parts of the entire server through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 420, and by calling data stored in the memory 420, it executes various functions of the server and processes data. In some embodiments, the processor 410 may include one or more processing cores; in some embodiments, the processor 410 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 410.
[0221] The memory 420 can be used to store software programs and modules. The processor 410 executes various functional applications and data processing by running the software programs and modules stored in the memory 420. The memory 420 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the server. In addition, the memory 420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 420 may also include a memory controller to provide the processor 410 with access to the memory 420.
[0222] The server further includes a power supply 430 for powering each component. In some embodiments, the power supply 430 may be logically connected to the processor 410 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 430 may further include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.
[0223] The server may further include an input module 440, which may be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0224] The server may further include a communication module 450. In some embodiments, the communication module 450 may include a wireless module, and the server may perform short-range wireless transmission through the wireless module of the communication module 450, thereby providing users with wireless broadband Internet access. For example, the communication module 450 may be used to help users send and receive emails, browse web pages, and access streaming media, etc.
[0225] Although not shown, the server may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 410 in the server will load the executable files corresponding to the processes of one or more application programs into the memory 420 according to the following instructions, and the processor 410 will run the application programs stored in the memory 420 to implement various functions as follows:
[0226] Obtain the regional features of the current control point area from the current image features of the image to be processed. The current image features are obtained from the current latent vector of the image to be processed. The image to be processed includes source control points and current control points corresponding to the source control points. The current control point area is an area related to the current control point; determine the correlation between the current control point area and the source control points according to the regional features of the current control point area and the source control point distribution parameters. The source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed; determine the feature difference between the current control point area and the source control points according to the regional features of the current control point area and the control point features of the source control points; combine the correlation and the feature difference to determine the updated current control points within the current control point area; update the current latent vector according to the updated current control points to obtain an updated current latent vector, so as to generate a processed image according to the updated current control points and the updated current latent vector.
[0227] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.
[0228] As can be seen from the above, in the process of deforming the image to be processed in the embodiment of the present application, the next current control point (i.e., the updated current control point) in the process is determined through the correlation and feature difference between the current control point area and the source control point, so as to gradually drag the source control point based on the updated current control point to generate the processed image. In this process, through the correlation between the current control point area and the source control point, it can be ensured that the updated current control point has a high similarity with the source control point, so that the generated processed image can better retain the features of the source control point. At the same time, the degree of feature change is measured through the feature difference to control the degree of image deformation. In this way, by combining the correlation and feature difference between the current control point area and the source control point, the degree of image deformation can be accurately controlled while retaining the features of the image to be processed, improving the quality of the generated processed image. In addition, the embodiment of the present application also determines the correlation between the current control point area and the source control point based on the Gaussian distribution source control point distribution parameter representing the source control point in the image to be processed, so as to use the Gaussian distribution information of the source control point to increase the influence of the points related to the source control point in the current control point area, improve the accuracy of determining the updated current control point in the current control point area, and improve the efficiency and quality of the generated processed image.
[0229] Those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods of the embodiments can be completed by instructions, or by instructions controlling relevant hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0230] For this reason, the embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiment of the present application. For example, the instructions can execute the following steps:
[0231] Obtain the regional features of the current control point area from the current image features of the image to be processed. The current image features are obtained from the current hidden vector of the image to be processed. The image to be processed includes source control points and the corresponding current control points. The current control point area is the area related to the current control point; determine the correlation between the current control point area and the source control point according to the regional features of the current control point area and the source control point distribution parameter, and the source control point distribution parameter is used to represent the Gaussian distribution of the source control point in the image to be processed; determine the feature difference between the current control point area and the source control point according to the regional features of the current control point area and the control point features of the source control point; combine the correlation and the feature difference to determine the updated current control point in the current control point area; update the current hidden vector according to the updated current control point to obtain the updated current hidden vector, so as to generate the processed image according to the updated current control point and the updated current hidden vector.
[0232] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0233] According to one aspect of the present application, a computer program product or a computer program is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps in the methods provided in the various alternative implementations in the above embodiments are implemented. The computer program / instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program / instruction from the computer-readable storage medium, and the processor executes the computer program / instruction, so that the electronic device executes the methods provided in the various alternative implementations in the above embodiments.
[0234] Since the instructions stored in the storage medium can execute the steps in any one of the image processing methods provided in the embodiments of the present application, the beneficial effects achievable by any one of the image processing methods provided in the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be elaborated here.
[0235] The above has introduced in detail an image processing method, device, electronic device, storage medium and program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain the regional feature of the current control point area from the current image features of the image to be processed, where the current image features are obtained from the current latent vector of the image to be processed, the image to be processed includes source control points and corresponding current control points, and the current control point area is an area related to the current control point; Determine the correlation between the current control point area and the source control points according to the regional feature of the current control point area and the source control point distribution parameters, where the source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed; Determine the feature difference between the current control point area and the source control points according to the regional feature of the current control point area and the control point features of the source control points; Combine the correlation and the feature difference to determine the updated current control points within the current control point area; Update the current latent vector according to the updated current control points to obtain an updated current latent vector, so as to generate a processed image according to the updated current control points and the updated current latent vector.
2. The image processing method according to claim 1, wherein The source control point distribution parameters include a pre-trained convolution matrix, which is trained based on the Gaussian distribution information of the source control points in the image to be processed. The determining the correlation between the current control point area and the source control points according to the regional feature of the current control point area and the source control point distribution parameters includes: Obtain the pre-trained convolution matrix; Perform convolution calculation on the regional feature of the current control point area through the pre-trained convolution matrix to determine the correlation between the current control point area and the source control points.
3. The image processing method according to claim 2, characterized in that The pre-trained convolution matrix is obtained through the following steps: Obtain an initial convolution matrix and the Gaussian distribution information of the source control points in the image to be processed; Use the Gaussian distribution information of the source control points in the image to be processed as the target value; Perform convolution calculation on the regional feature of the initial control point area through the initial convolution matrix to determine the initial correlation between the points in the initial control point area and the source control points, where the initial control point area is an area related to the source control points; Train the initial convolution matrix according to the difference between the initial correlation and the target value to obtain the pre-trained convolution matrix.
4. The image processing method according to claim 1, wherein The determining the feature difference between the current control point area and the source control points according to the regional feature of the current control point area and the control point features of the source control points includes: Obtain the control point features of the source control points from the initial latent vector of the image to be processed; Use the feature distance between the regional feature of the current control point area and the control point features of the source control points as the feature difference between the current control point area and the source control points.
5. The image processing method according to claim 1, characterized in that The combining the correlation and the feature difference to determine the updated current control points within the current control point area includes: Combine the correlation and the feature difference to calculate the confidence of the points within the current control point area; Determine the updated current control point within the current control point region according to the confidence level.
6. The image processing method according to claim 1, wherein Updating the current latent vector according to the updated current control point to obtain an updated current latent vector, and generating a processed image according to the updated current control point and the updated current latent vector, includes: Update the current latent vector according to the updated current control point to obtain an updated current latent vector; Take the updated current control point as the current control point, and take the updated current latent vector as the current latent vector, and return to execute the steps of obtaining the region feature of the current control point region from the image features of the image to be processed and subsequent steps until the updated current control point matches the specified target control point, and generate a processed image from the updated current latent vector.
7. The image processing method according to claim 6, wherein Updating the current latent vector according to the updated current control point to obtain an updated current latent vector, includes: Obtain the region feature of the reference control point region corresponding to the confidence level of the updated current control point according to the confidence level of the updated current control point, where the confidence level is obtained from the correlation and the feature difference, and the reference control point region includes one of the current control point region and the initial control point region, and the initial control point region is the region related to the source control point; Determine the loss value according to the region feature of the reference control point region and the region feature corresponding to the updated current control point; Update the current latent vector according to the loss value to obtain an updated current latent vector.
8. The image processing method according to claim 7, wherein Obtaining the region feature of the reference control point region corresponding to the confidence level of the updated current control point according to the confidence level of the updated current control point, includes: If the confidence level of the updated current control point is greater than or equal to the preset threshold, obtain the region feature of the current control point region; If the confidence level of the updated current control point is less than the preset threshold, obtain the region feature of the initial control point region.
9. The image processing method according to claim 7, wherein The image to be processed further includes a preset deformation region, and the source control point and the specified target control point are located within the preset deformation region. Determining the loss value according to the region feature of the reference control point region and the region feature corresponding to the updated current control point, includes: Determine the control point region feature difference according to the region feature of the reference control point region and the region feature corresponding to the updated current control point; Determine the deformation region feature difference according to the region feature of the preset deformation region in the initial image feature and the region feature of the preset deformation region in the current image feature, where the initial image feature is obtained from the initial latent vector of the image to be processed; Combine the control point region feature difference and the deformation region feature difference to determine the loss value.
10. The image processing method according to claim 7, wherein The region feature corresponding to the updated current control point is obtained through the following steps: Determine the offset of the updated current control point relative to the specified target control point according to the feature distance between the specified target control point and the updated current control point; Determine the offset control point corresponding to the updated current control point according to the updated current control point and the offset; Obtain the regional feature of the offset control point area from the current image features of the image to be processed, and use the regional feature of the offset control point area as the regional feature corresponding to the updated current control point, where the offset control point area is the area related to the offset control point.
11. The image processing method according to any one of claims 1 to 10, characterized in that, Before obtaining the regional feature of the current control point area from the current image features of the image to be processed, it further includes: Obtain an image generation model; Through the image generation model, perform feature learning on the current latent vector of the image to be processed, and use the feature vector output by the middle layer of the image generation model as the current image features of the image to be processed.
12. An image processing apparatus, characterized in that, It includes: An acquisition unit for obtaining the regional feature of the current control point area from the current image features of the image to be processed, where the current image features are obtained from the current latent vector of the image to be processed, the image to be processed includes source control points and the current control points corresponding to the source control points, and the current control point area is the area related to the current control point; A correlation determination unit for determining the correlation between the current control point area and the source control points according to the regional feature of the current control point area and the source control point distribution parameters, where the source control point distribution parameters are used to represent the Gaussian distribution of the source control points in the image to be processed; A difference determination unit for determining the feature difference between the current control point area and the source control points according to the regional feature of the current control point area and the control point features of the source control points; A control point determination unit for determining the updated current control point within the current control point area by combining the correlation and the feature difference; An update unit for updating the current latent vector according to the updated current control point to obtain an updated current latent vector, so as to generate a processed image according to the updated current control point and the updated current latent vector.
13. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in the image processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the image processing method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, the steps in the image processing method according to any one of claims 1 to 11 are implemented.