An image processing method, system, storage medium and terminal device

By acquiring the three-dimensional coefficients of the face image and combining the combined data processing of the pre-trained image processing model, the problem of poor face image processing effect in the prior art is solved, and image processing effect with higher resolution and clarity is achieved.

CN115131196BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210348927.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-06-27
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

The prior art is not effective when processing face images, especially face images with low definition and low quality, and it is difficult to effectively improve the clarity and quality of the image.

Method used

By acquiring the target face image and its three-dimensional coefficients, combining the data into the pre-trained image processing model, and generating the processed image, the resolution is higher than that of the input image.

Benefits of technology

The processing effect of low-definition face images is significantly improved, and the image resolution is higher after the generated processing is also significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131196B_ABST
    Figure CN115131196B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose an image processing method, system, storage medium, and terminal device, which are applied in the field of artificial intelligence. The image processing system can obtain a three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image, and call a pre-trained image processing model. The image processing model can obtain a processed image of the target face image according to the merged data of the target face image and its three-dimensional image, and the resolution of the processed image is higher than that of the target face image. This application combines the corresponding three-dimensional image when processing a target face image with low clarity (i.e., low resolution), so as to improve the effect of the processed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an image processing method, system, storage medium and terminal device. Background Art

[0002] During the process of image shooting or network transmission, the captured image may be unclear and of low quality due to factors such as inaccurate focusing, excessive noise, and too high picture compression ratio. Therefore, it is necessary to process the image to improve its clarity and quality.

[0003] In the existing image processing, the image can be processed based on the machine learning model of artificial intelligence to obtain a high-definition image. However, when using the existing image processing method to process some images (such as face images), due to the complexity of the objects in these images, the effect of the processed image is not very good. Summary of the Invention

[0004] The embodiments of the present invention provide an image processing method, system, storage medium and terminal device, which improve the processing effect of face images.

[0005] On the one hand, an embodiment of the present invention provides an image processing method, including:

[0006] Obtain a target face image and the three-dimensional coefficients of the face image included in the target face image;

[0007] Obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image;

[0008] Call a pre-trained image processing model;

[0009] The image processing model obtains the processed image of the target face image according to the merged data of the target face image and the three-dimensional image, outputs the processed image, and the resolution of the processed image is higher than that of the target face image.

[0010] On the other hand, an embodiment of the present invention provides an image processing system, including:

[0011] A coefficient acquisition unit for obtaining a target face image and the three-dimensional coefficients of the face image included in the target face image;

[0012] A three-dimensional image unit for obtaining the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image;

[0013] A model call unit for calling a pre-trained image processing model;

[0014] A processing unit, which is configured to obtain a processed image of the target face image according to the merged data of the target face image and the three-dimensional image by the image processing model, and output the processed image, where the resolution of the processed image is higher than that of the target face image.

[0015] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium storing a plurality of computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the image processing method as described in one aspect of the embodiments of the present invention.

[0016] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory;

[0017] The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by a processor to perform the image processing method as described in one aspect of the embodiments of the present invention; the processor is used to implement each of the computer programs in the plurality of computer programs.

[0018] It can be seen that in the method of this embodiment, the image processing system can obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image, and call the pre-trained image processing model. The image processing model obtains the processed image of the target face image according to the merged data of the target face image and its three-dimensional image, and the resolution of the processed image is higher than that of the target face image. In this way, when processing a target face image with low clarity (i.e., low resolution), the corresponding three-dimensional image is combined, and the data of the three-dimensional image can describe the face image in more detail, so that the effect of the processed image obtained by processing is well improved. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a schematic diagram of an image processing method provided by an embodiment of the present invention;

[0021] Figure 2 is a flowchart of an image processing method provided by an embodiment of the present invention;

[0022] Figure 3a is a schematic structural diagram of an image processing model called in an embodiment of the present invention;

[0023] Figure 3bIt is a schematic structural diagram of another image processing model called in an embodiment of the present invention;

[0024] Figure 4 It is a flowchart of a method for training an image processing model in an embodiment of the present invention;

[0025] Figure 5 It is a schematic diagram of an initial image processing model determined when training an image processing model in an application embodiment of the present invention;

[0026] Figure 6 It is a schematic diagram of a second decoding module in an application embodiment of the present invention;

[0027] Figure 7 It is a schematic diagram of an image processing method provided in an application embodiment of the present invention;

[0028] Figure 8 It is a schematic diagram of a distributed system to which the image processing method is applied in another application embodiment of the present invention;

[0029] Figure 9 It is a schematic diagram of a block structure in another application embodiment of the present invention;

[0030] Figure 10 It is a schematic logical structure diagram of an image processing system provided in an embodiment of the present invention;

[0031] Figure 11 It is a schematic logical structure diagram of a terminal device provided in an embodiment of the present invention. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] An embodiment of the present invention provides an image processing method, which mainly processes any face image with low clarity and low quality (i.e., the target face image), such as Figure 1 As shown, the image processing system of this embodiment can process the target face image according to the following method:

[0035] Obtain the target face image and the three-dimensional coefficients of the face image included in the target face image; obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image; call the pre-trained image processing model; the image processing model obtains the processed image of the target face image according to the merged data of the target face image and the three-dimensional image, and outputs the processed image, and the resolution of the processed image is higher than the resolution of the target face image.

[0036] In practical applications, the above-mentioned image processing system can be mainly applied to the following application terminals: mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0037] The pre-trained image processing model here is a machine learning model based on artificial intelligence. Among them, artificial intelligence (AI) uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theory, methods, technologies and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.

[0038] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0039] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0040] In this way, when processing a target face image with low clarity (i.e., low resolution), the corresponding three-dimensional image is combined. The data of the three-dimensional image can describe the face image in more detail, so that the effect of the processed image obtained is well improved.

[0041] An embodiment of the present invention provides an image processing method, mainly the method executed by an image processing system. The flowchart is as Figure 2 shown, including:

[0042] Step 101, obtain a target face image and the three-dimensional coefficients of the face image included in the target face image.

[0043] It can be understood that in one case, the image processing system can provide an interaction interface with the user. In this way, the user can specify a certain face image as the target face image through the interaction interface and initiate the image processing process of this embodiment. Or, in another case, when certain specific events occur, the image processing system can be triggered to use a certain face image as the target face image and initiate the image processing process of this embodiment.

[0044] Generally, the target face image is a two-dimensional image with low clarity and low quality, and the obtained three-dimensional coefficients of the face image refer to the information required to perform three-dimensional reconstruction on the face image included in the two-dimensional target face image to obtain the corresponding three-dimensional image, such as the shape and posture of the face included in the target face image.

[0045] Specifically, the image processing system can use a pre-set coefficient prediction model to obtain the three-dimensional coefficients of any face image. Here, the coefficient prediction model is a machine learning model based on artificial intelligence, which can be trained through a certain training method, and the operation logic of the trained coefficient prediction model is stored in the system in advance. When the process of this embodiment is initiated, the pre-set coefficient prediction module in the system can be directly called to obtain the three-dimensional coefficients of the target face image.

[0046] Among them, the coefficient prediction model can adopt a network with any structure, such as Residual Neural Network (ResNet) 50, etc.

[0047] Step 102: Obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image.

[0048] Specifically, the image processing system can first use the method of three-dimensional reconstruction to reconstruct the three-dimensional (3D) mesh information of the target face image according to the three-dimensional coefficients of the target face image. The three-dimensional mesh information can include the shape (S) and texture (T) of multiple faces that make up the three-dimensional face. Here, the three-dimensional face is the three-dimensional face corresponding to the face included in the target face image. Then, the three-dimensional mesh information is projected onto a two-dimensional plane through rendering to obtain the three-dimensional image of the target face image.

[0049] Among them, the three-dimensional network information refers to the information of all the faces (such as triangular faces, etc.) composed of three-dimensional objects. In this embodiment, the image processing system can call a three-dimensional deformable face model (3DMM), and the 3DMM can reconstruct the three-dimensional network information according to the three-dimensional coefficients obtained in the above steps.

[0050] It should be noted that when reconstructing the three-dimensional mesh information according to the three-dimensional coefficients, the three-dimensional mesh information of a small part of the face area can be reconstructed, such as the three-dimensional network information of the area except for the hair and ears; or the three-dimensional mesh information of a larger area of the face can be reconstructed, so that the effect of the restored three-dimensional image is better.

[0051] Step 103: Call a pre-trained image processing model.

[0052] Step 104: The image processing model obtains the processed image of the target face image according to the combined data of the target face image and the three-dimensional image, and outputs the processed image. The resolution of the processed image is higher than that of the target face image.

[0053] Here, the combined data of the target face image and the three-dimensional image refers to the data obtained by concatenating the data of the target face image and the data of the three-dimensional image.

[0054] Specifically, the image processing system can call an image processing model with the following structure, such as Figure 3a and Figure 3b As shown, the image processing model may include: a mapping module 11, a plurality of encoding modules 10, a plurality of decoding modules (including a first decoding module 12 and a plurality of second decoding modules 13), and an output module 14, where: the output of one encoding module 10 is connected to one input of one decoding module, and the output of the mapping module 11 is respectively connected to another input of the plurality of decoding modules; the plurality of encoding modules 10 are connected in series, the plurality of decoding modules are also connected in series, and the output of the last decoding module among the plurality of decoding modules is connected to the output module 14. Among them, the number of encoding modules 10 is equal to the number of decoding modules.

[0055] The relationship between the encoding module 10 and the mapping module can be as Figure 3a shown. There is no direct connection between the mapping module 11 and the encoding module 10, and their inputs are the same, both being the target face image and its three-dimensional image. The relationship between the encoding module 10 and the mapping module 11 can also be as Figure 3b shown. The output of the last encoding module 10 among the plurality of encoding modules 10 is connected to the mapping module 11.

[0056] In this way, when obtaining the processed image, the mapping module 11 can jointly determine the latent variable based on the data of the target face image and the three-dimensional image, and the latent variable is used to adjust the convolution weights of the convolution calculations involved in each decoding module; the encoding module 10 extracts the spatial features of a resolution of the combined data of the target face image and the three-dimensional image, that is, one encoding module 10 corresponds to one resolution, and the resolutions corresponding to different encoding modules 10 are different; the second decoding module 13 among the plurality of decoding modules performs convolution calculations based on the latent variable, the spatial features of a resolution extracted by one encoding module, and the feature information of a resolution obtained by the previous decoding module of the second decoding module 13 to obtain the feature information of another resolution; further, the output module 14 obtains the processed image of the target face image based on the feature information obtained by the last second decoding module 13.

[0057] Specifically, when the second decoding module 13 obtains the feature information of another resolution, the second decoding module 13 can first adjust the convolution weights involved in the convolution calculation according to the latent variable to obtain the adjusted convolution weights, and then perform convolution calculations on the feature information of a resolution obtained by the previous decoding module according to the adjusted convolution weights to obtain the convolved features; finally, according to the spatial features of a resolution extracted by one encoding module and the convolved features, the feature information of another resolution is obtained.

[0058] In this process, the first decoding module 12 among multiple decoding modules performs convolution calculations based on the latent variable and the spatial features of a resolution obtained by an encoding module to obtain feature information of another resolution. Specifically, the first decoding module 12 can first adjust the convolution weights involved in the convolution calculation according to the latent variable to obtain the adjusted convolution weights, and then perform convolution calculation on the preset initial feature information according to the adjusted convolution weights to obtain the convolved features; finally, according to the spatial features of a resolution extracted by an encoding module and the convolved features, feature information of another resolution is obtained. And in this process, after the first encoding module 10 among multiple encoding modules 10 directly extracts the spatial features based on the data of the target face image and the three-dimensional image, spatial features of a resolution are obtained; other encoding modules 10 (encoding modules 10 other than the first encoding module) among multiple encoding modules 10 obtain spatial features of another resolution according to the spatial features of a resolution obtained by the previous encoding module 10.

[0059] In addition, it should be noted that in the process of obtaining the processed image above, when the mapping module 11 determines the latent variable, in the case as Figure 3a shown, the mapping module 11 directly determines the latent variable according to the combined data of the target face image and the three-dimensional image; in the case as Figure 3b shown, after the combined data of the target face image and the three-dimensional image passes through the feature extraction of multiple encoding modules 10, spatial features of a certain resolution are obtained, and the mapping module 11 determines the latent variable according to the spatial features of a certain resolution output by the last encoding module 10 among multiple encoding modules 10.

[0060] It can be seen that in the method of this embodiment, the image processing system can obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image, and call the pre-trained image processing model. The image processing model obtains the processed image of the target face image according to the combined data of the target face image and its three-dimensional image, and its resolution is higher than that of the target face image. In this way, when processing a target face image with low clarity (i.e., low resolution), the corresponding three-dimensional image is combined, and the data of the three-dimensional image can describe the face image in more detail, so that the effect of the processed image obtained is well improved.

[0061] In a specific embodiment, the image processing model called in step 103 above can be pre-trained according to the following steps. The flowchart is as Figure 4 shown and includes:

[0062] Step 201, determine the initial image processing model, and the initial image processing model includes: a processing sub-module and a discriminant sub-module.

[0063] It can be understood that when determining the initial image processing model, the image processing system will determine the multi-layer structure included in the initial image processing model and the initial values of the parameters in each layer mechanism. Among them, the parameters of the initial image processing model refer to the fixed parameters that are used in the calculation process of each layer structure in the initial image processing model and do not need to be assigned values at any time, such as parameter scale, number of network layers, weight values, and other parameters.

[0064] Specifically, as Figure 5 shown, the image processing system can determine that the initial image processing model includes: a processing sub-module 20 and a discrimination sub-module 21. The processing sub-module 20 is used to obtain a processed image of the first face sample image according to the first face sample image and its corresponding three-dimensional face image, and the discrimination sub-module 21 is used to discriminate whether the processed image obtained by the processing sub-module 20 is real. In this embodiment, the discrimination sub-module 21 also needs to discriminate whether the second face sample image is real.

[0065] Specifically, the processing sub-module 20 may include the following structure: a mapping module 210, a plurality of encoding modules 220, a plurality of decoding modules (including a first decoding module 230 and a second decoding module 240), and an output module 250, where:

[0066] The mapping module 210 is used to combine the data of the first face sample image and its three-dimensional face image, that is, to determine a latent variable according to the combined data of the first face sample image and its three-dimensional face image. The latent variable is used to adjust the convolution weights involved in the convolution calculations of each decoding module; the encoding module 220 is used to extract the spatial features of a resolution of the combined data of the first face sample image and its three-dimensional face image. The first encoding module 220 directly extracts the spatial features of a resolution of the first face sample image and its three-dimensional face image, and the other encoding modules 220 can obtain the spatial features of another resolution according to the spatial features of a resolution obtained by the previous encoding module 220; the second decoding module 240 in the plurality of decoding modules is used to perform convolution calculations according to the latent variable, the spatial features of a resolution extracted by an encoding module 220, and the feature information of a resolution obtained by the previous decoding module of the second decoding module 240 to obtain the feature information of another resolution, and the first decoding module 230 in the plurality of decoding modules will perform convolution calculations according to the latent variable and the spatial features of a resolution obtained by an encoding module to obtain the feature information of another resolution; the output module 250 obtains the processed image of the first face sample image according to the feature information obtained by the last second decoding module 240.

[0067] It should be noted that the processing sub-module 20 and the discrimination sub-module 21 determined here constitute a Generative Adversarial Network (GAN). Using a GAN can achieve unsupervised machine learning. A GAN mainly consists of a generator network and a discriminator network. In this embodiment, the generator network is specifically the processing sub-module 20, and the discriminator network is specifically the discrimination sub-module 21. Practice has proved that for any low-resolution face image, the quality of the high-resolution image obtained by the image processing model trained through the GAN is improved.

[0068] Step 202: Determine training samples. The training samples include multiple sample groups, and each sample group includes a low-resolution first face sample image, its corresponding high-resolution second face sample image, and a three-dimensional face image.

[0069] Specifically, when obtaining the first face sample image, the second face sample image can be downsampled to obtain the first face sample image. Among them, the downsampling process can include, but is not limited to, any of the following processing methods: adding blur, downsampling, noise, and Joint Photographic Experts Group (JPEG) compression, etc. Among them:

[0070] Adding blur processing randomly adds Gaussian blur, motion blur, etc. The standard deviation of Gaussian blur is randomly selected within a certain range. Motion blur includes a variety of (38 kinds) custom blur kernels; downsampling processing is to reduce the image resolution, and the sampling method is randomly selected from biliner, bicubic, area, etc.; noise processing includes Gaussian noise, Poisson noise, etc., and the noise intensity is randomly selected within a certain range; JPEG compression processing is to simulate the decline in image quality during the image saving process. The compression ratio is randomly selected between 5% and 50%.

[0071] Step 203: The processing sub-module 20 obtains the processed image of the first face sample image according to the merged data of the first face sample image and its corresponding three-dimensional face image, and the discrimination sub-module 21 discriminates whether the processed image obtained by the processing sub-module is real.

[0072] Specifically, the encoding module 220 of the processing sub-module 20 in the initial image processing model extracts the spatial features of the merged data of the first face sample image and its 3D face image; at the same time, the mapping module 210 obtains the latent variable according to the merged data of the first face sample image and its 3D face image to adjust the convolutional weights in the decoding module; then, based on the convolutional weights adjusted by the latent variable, the decoding module performs convolutional calculation on the spatial features extracted by the encoding module 220 to obtain the post-convolution features, and finally the output module 250 outputs the processed image of the first face sample image according to the post-convolution features.

[0073] The discrimination sub-module 21 then discriminates the processed image obtained by the processing sub-module 20. In this embodiment, the discrimination sub-module 21 needs to discriminate whether the processed image obtained by the processing sub-module 20 is true, and also needs to discriminate whether the second face sample image is true.

[0074] Step 204: Adjust the initial image processing model according to the result obtained by the discrimination sub-module 21 and the second face sample image in the training sample. The processing sub-module 20 in the adjusted initial image processing model is the above-mentioned pre-trained image processing model. This image processing model is mainly used to obtain the corresponding high-resolution image according to the merged data of any low-resolution image and its 3D image.

[0075] Specifically, the image processing system will first calculate the loss function related to the initial image processing model according to the result obtained by the discrimination sub-module 21 in the above step 203 and the corresponding second face sample image. This loss function is used to indicate the error between the processed image of the first face sample image obtained by the processing sub-module 20 and the actual high-resolution image of the first face sample image (i.e., the second face sample image), such as the cross-entropy loss function, etc. This error can be obtained through the result obtained by the discrimination sub-module 21 and the second face sample image; then, according to the loss function, adjust the parameter values of the parameters in the initial image processing model.

[0076] In the embodiment of the present invention, when calculating the loss function, the image processing system can calculate but is not limited to the following loss sub-functions: feature loss sub-function, adversarial loss sub-function, and reconstruction loss sub-function, and use the function calculation values of the feature loss sub-function, adversarial loss sub-function, and reconstruction loss sub-function as the loss function related to the initial image processing model, such as using the sum value of these sub-functions as this loss function. Among them:

[0077] Adversarial loss sub-function L GAN, which is used to represent the result D(G(input)) indicating whether the processed image of the first face sample image obtained by the above discriminant sub-module 21 and discriminant processing sub-module 20 is true, and the expected result D(GT) of whether the actually high-resolution image of the first face sample image (i.e., the second face sample image) is true by the discriminant sub-module 21. Specifically, it can be represented by the following formula 1, where D represents the discriminant sub-module 21, G represents the processing sub-module 20, GT represents the high-resolution second face sample image in the training samples, and input represents the low-resolution first face sample image in the training samples and its corresponding 3D face image:

[0078]

[0079] Reconstruction loss sub-function L rec , which is used to represent the difference between the processed image G(input) of the first face sample image obtained by the above processing sub-module 20 and the actually high-resolution image of the first face sample image (i.e., the second face sample image GT). Specifically, it can be represented by the following formula 2. The reconstruction loss sub-function can include the difference between the processed image G(input) and the second face sample image GT, and the difference between the feature information LPIPS(G(input)) of the processed image and the feature information LPIPS(GT) of the second face sample image:

[0080] L rec =|G(input)-GT| l -|LPIPS(G(input))-LPIPS(GT)| l (2)

[0081] Feature loss sub-function L rec , which is used to represent the difference between the feature information of the processed image G(input) of the first face sample image obtained by the processing sub-module and the feature information of the actually high-resolution image of the first face sample image (i.e., the second face sample image GT). Specifically, it can be represented by the following formula 3. The feature loss sub-function can include the distance between the feature information of the processed image G(input) and the feature information of the second face sample image GT:

[0082] L ID =1 - F cos (F ArcFace (G(input)), F ArcFace (GT)) (3)

[0083] Therefore, the loss function L related to the initial image processing model calculated by the image processing system can be represented by the following formula 4:

[0084] L = LGAN +L rec +L ID (4)

[0085] The training process of the image processing model is to minimize the value of the above loss function. This training process continuously optimizes the parameter values of the initial image processing model determined in step 201 through a series of mathematical optimization means such as backpropagation derivative calculation and gradient descent, and makes the calculated value of the above loss function drop to the lowest. Specifically, when the function value of the calculated loss function is relatively large, such as greater than a preset value, the parameter values need to be changed. For example, the weight value of a certain neuron connection is reduced, etc., so that the function value of the loss function calculated according to the adjusted parameter values is reduced.

[0086] It should be noted that in the actual process of adjusting the initial image processing model according to the calculated loss function, in one case, the parameter values of the processing sub-module 20 and the discriminant sub-module 21 can be adjusted simultaneously according to the loss function obtained from the above formula 4. In another case, the parameter values of the processing sub-module 20 can be fixed first, and the parameter values of the discriminant sub-module 21 are adjusted according to the adversarial loss sub-function obtained from the above formula 1; then the parameter values of the adjusted discriminant sub-module 21 are fixed, and the parameter values of the processing sub-module 20 are adjusted according to the loss function obtained from the above formulas 1 to 4.

[0087] Among them, in the process of adjusting the parameter values of the discriminant sub-module 21, it is to minimize the adversarial loss sub-function obtained from the above formula 1, that is, it is necessary to make the discriminant sub-module 21 determine that the processed image of the first face sample image obtained by the processing sub-module 20 is false, while determining that the second face sample image is true.

[0088] In addition, it should be noted that the above steps 203 to 204 are an adjustment of the parameter values in the initial image processing model based on the results obtained from the processing sub-module 20 and the discriminant sub-module 21 in the initial image processing model. In actual applications, it is necessary to continuously loop through the above steps 203 to 204 until the adjustment of the parameter values meets a certain stop condition.

[0089] Therefore, after the image processing system executes steps 201 to 204 of the above embodiment, it is also necessary to determine whether the current adjustment of the parameter value meets a preset stop condition. When it is met, the process ends, and the processing sub-module in the initial image processing model obtained through the adjustment in step 204 is used as the preset image processing model. When it is not met, for the initial image processing model after adjusting the parameter value, return to execute the above steps 203 to 204. The preset stop condition includes, but is not limited to, any one of the following conditions: the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, that is, the adjusted parameter value reaches convergence; and the number of times of adjusting the parameter value is equal to the preset number of times, etc.

[0090] It can be seen that in the process of pre-training the image processing model in this embodiment, the adversarial network formed by the processing sub-module 20 and the discriminant sub-module 21 can avoid manual annotation of any sample image in the determined training samples, realizing unsupervised training of the image processing model and saving the time spent on manual annotation.

[0091] The following uses a specific application example to illustrate the image processing method of the present invention. The method of this embodiment may include the following two parts:

[0092] (1) Pre-training the image processing model

[0093] Specifically, the method adopted by the image processing system when training the image processing model is similar to the training method as shown above Figure 4 , except that in this embodiment:

[0094] (1) In this embodiment, when determining the structure of the initial image processing model, specifically, it can be determined that the processing sub-module 20 includes 7 decoding modules (including a first decoding module 230 and 6 second decoding modules 240) and 7 encoding modules 220, and the resolutions corresponding to the encoding modules 220 and the decoding modules can both increase from 4*4 to 512*512. The structure of the second decoding module 240 is as shown in Figure 6 , where:

[0095] Assume that a spatial feature of a resolution of the merged data of the first face sample image and its three-dimensional face image extracted by an encoding module 220 connected to the second decoding module 240 is And the feature information of a resolution obtained by the previous decoding module of the second decoding module 240 is where is the same as the resolution corresponding to , so that:

[0096] After the latent variable w of the mapping module 210 is input into the second decoding module 240, it is processed by A241, Mod242, and Demod243 and then input into a convolutional layer (Conv) 244 to adjust the convolutional weights of the convolutional layer, obtaining the adjusted convolutional weights; and the feature information of one resolution can pass through upsampling 245 and a convolutional layer 244, and perform convolutional calculations in the convolutional layer 244 according to the adjusted convolutional weights to obtain the post-convolution features Among them, when the mapping module 210 obtains the latent variable w, it is mainly based on the first low-resolution face sample image I lq and its corresponding 3D face image I 3d The combined data of is obtained, which is specifically represented by the following formula 5:

[0097] w = MLP(I lq , I 3d ) (5)

[0098] The spatial features of one resolution After passing through another convolutional layer 246, it is input into a concatenation layer C247. The concatenation layer C247 will be based on the spatial features The obtained post-convolution features And the post-convolution features obtained based on the feature information The obtained post-convolution features Are concatenated to obtain the feature information of another resolution Specifically, it can be represented by the following formula 6:

[0099]

[0100] (2) In this embodiment, the learning rate ratios of the respective encoding modules 220, output module 250, and discrimination sub-module 21 in the processing sub-module 20 can be 100:10:1.

[0101] It should be noted that after the image processing model is trained, the running logic of the image processing model can be preset in the system.

[0102] (2) As Figure 7 shown, the image processing system can process any low-resolution target face image:

[0103] Step 301, the user can, through the user interface provided by the image processing system, specify a face image with a relatively low resolution (i.e., low clarity) as the target face image and trigger the processing of the target face image.

[0104] In step 302, the image processing system can predict the three-dimensional coefficients (coeff) of the face image contained in the target face image I according to a pre-trained 3D coefficient prediction network, such as ResNet50, etc., which can be specifically expressed by the following formula 7: lq coeff = F

[0105] (I res50 ) (7) lq

[0106] In step 303, the image processing system can reconstruct the three-dimensional mesh information (3D mesh) according to the obtained three-dimensional coefficients by calling a three-dimensional deformable face model (3DMM). The three-dimensional network information can include the shape (S) and texture (T) of multiple faces that make up the three-dimensional face (the face image in the target face image). Then, the three-dimensional mesh information can be projected onto a two-dimensional plane through a renderer to obtain the three-dimensional image I of the target face image 3d . The data of the three-dimensional image can include information such as face shape and lighting. Specifically, when obtaining the three-dimensional image, it can be expressed by the following formulas 8 and 9:

[0107] S, T = F 3dmm (coeff) (8)

[0108] I 3d = F render (S,T) (9)

[0109] In step 304, the image processing system calls a pre-set image processing model, and takes the merged data obtained by merging the target face image and the data of its three-dimensional image, that is, the 6-channel concatenated data, as the input of the image processing model.

[0110] In step 305, the image processing model outputs the processed image of the target face image according to the merged data of the target face image and its three-dimensional image.

[0111] Specifically, the encoding module in the image processing model extracts the spatial features of the merged data of the target face image and its three-dimensional image. The mapping module obtains the latent variable according to the merged data of the target face image and its three-dimensional image to adjust the convolution weights in the decoding module. Then, the decoding module performs convolution calculation on the spatial features extracted by the encoding module based on the convolution weights adjusted by the latent variable, so as to obtain the convolved features. Finally, the output module outputs the processed image of the target face image according to the convolved features.

[0112] It can be seen that through the pre-trained image processing model in this embodiment, by combining the two-dimensional target face image and its three-dimensional image, the resolution of the obtained processed image is high and the effect is improved.

[0113] ​The following uses another specific application example to illustrate the image processing method in the present invention. The image processing system in the embodiments of the present invention is mainly a distributed system 100, which may include a client 300 and multiple nodes 200 (any form of computing device accessing the network, such as a server or a user terminal). The client 300 and the nodes 200 are connected through network communication.

[0114] Taking the distributed system as a blockchain system as an example, refer to Figure 8 It is an optional structural schematic diagram of the distributed system 100 provided by the embodiments of the present invention applied to the blockchain system, formed by multiple nodes 200 (any form of computing device accessing the network, such as a server or a user terminal) and a client 300. A peer-to-peer (P2P) network is formed between the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In the distributed system, any machine such as a server or a terminal can join and become a node. The node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.

[0115] Refer to Figure 8 The functions of each node in the shown blockchain system are involved, including:

[0116] 1) Routing, the basic function of the node, used to support communication between nodes.

[0117] In addition to the routing function, the node may also have the following functions:

[0118] 2) Application, used to be deployed in the blockchain, implement specific services according to actual business needs, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to the temporary block.

[0119] For example, the service implemented by the application includes the code of the image processing function, and the image processing function mainly includes:

[0120] Obtain the target face image and the three-dimensional coefficients of the face image included in the target face image; obtain the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image; call the pre-trained image processing model; the image processing model obtains the processed image of the target face image according to the merged data of the target face image and the three-dimensional image, and outputs the processed image, and the resolution of the processed image is higher than that of the target face image.

[0121] 3) A blockchain, including a series of blocks that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by nodes in the blockchain system.

[0122] See Figure 9 This is an optional schematic diagram of the block structure provided by the embodiments of the present invention. Each block includes the hash value of the transaction records stored in this block (the hash value of this block), and the hash value of the previous block. The blocks are connected to form a blockchain through the hash values. In addition, the block may also include information such as the timestamp when the block is generated. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.

[0123] The embodiments of the present invention also provide an image processing system, and its structural schematic diagram is as Figure 10 shown, and specifically may include:

[0124] A coefficient acquisition unit 30, configured to acquire a target face image and acquire the three-dimensional coefficients of the target face image.

[0125] A three-dimensional image unit 31, configured to acquire the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image acquired by the coefficient acquisition unit 30.

[0126] The three-dimensional image unit 31 is specifically configured to reconstruct the three-dimensional network information of the target face image according to the three-dimensional coefficients of the target face image; the three-dimensional network information includes the shapes and textures of multiple faces that make up the three-dimensional face, and the three-dimensional face is the three-dimensional face corresponding to the face included in the target face image; project the three-dimensional network information onto a two-dimensional plane by means of rendering to obtain the three-dimensional image of the target face image.

[0127] A model calling unit 32, configured to call a pre-trained image processing model.

[0128] The model calling unit 32 is specifically configured to call an image processing model with the following structure: the image processing model includes: a mapping module, a plurality of encoding modules, a plurality of decoding modules, and an output module, where: the output of one encoding module is connected to one input of one decoding module, and the output of the mapping module is respectively connected to the other inputs of the plurality of decoding modules; the plurality of encoding modules are connected in series, the plurality of decoding modules are connected in series, and the output of the last decoding module among the plurality of decoding modules is connected to the output module.

[0129] The processing unit 33 is configured to obtain a processed image of the target face image according to the merged data of the target face image and the three-dimensional image obtained by the three-dimensional image unit 31 by using the image processing model called by the model calling unit 32, and output the processed image, where the resolution of the processed image is higher than that of the target face image.

[0130] Specifically, the processing unit 33 is configured to determine a latent variable by combining the data of the target face image and the three-dimensional image through the mapping module, where the latent variable is used to adjust the convolution weights of the convolution calculations involved in the decoding module; extract the spatial features of a resolution of the merged data of the target face image and the three-dimensional image through the encoding module; the second decoding module among the multiple decoding modules performs convolution calculations according to the latent variable, the spatial features of a resolution extracted by one encoding module, and the feature information of the one resolution obtained by the previous decoding module of the second decoding module to obtain the feature information of another resolution; and the output module obtains the processed image of the target face image according to the feature information obtained by the last second decoding module.

[0131] Among them, when the second decoding module in the processing unit 33 performs convolution calculations according to the latent variable, the spatial features of a resolution extracted by one encoding module, and the feature information of the one resolution obtained by the previous decoding module of the second decoding module to obtain the feature information of another resolution, specifically, the second encoding module adjusts the convolution weights involved in the convolution calculations according to the latent variable to obtain adjusted convolution weights, and performs convolution calculations on the feature information of the one resolution according to the adjusted convolution weights to obtain convolved features; and obtains the feature information of another resolution according to the spatial features of the one resolution and the convolved features. And in the process that the image processing model obtains the processed image of the target face image according to the merged data of the target face image and the three-dimensional image, the first decoding module among the multiple decoding modules performs convolution calculations according to the latent variable and the spatial features obtained by one encoding module to obtain the feature information of one resolution.

[0132] Further, the image processing system in this embodiment may further include:

[0133] A training unit 34 for determining an initial image processing model; the initial image processing model includes: a processing sub-module and a discrimination sub-module; determining training samples, where the training samples include a plurality of sample groups, and each sample group includes a low-resolution first face sample image, its corresponding high-resolution second face sample image, and a three-dimensional face image; the processing sub-module obtains a processed image of the first face sample image according to the merged data of the first face sample image and its corresponding three-dimensional face image, and the discrimination sub-module discriminates whether the processed image obtained by the processing sub-module is genuine; according to the result obtained by the discrimination sub-module and the second face sample image in the training samples, adjusting the initial image processing model, and the processing sub-module in the adjusted initial image processing model is a pre-trained image processing model called by the model calling unit 32.

[0134] Wherein, the first face sample image is obtained by performing a resolution reduction process on the second face sample image, and the resolution reduction process includes any one of the following processing methods: adding blur, downsampling, noise, and Joint Photographic Experts Group (JPEG) compression processing.

[0135] When adjusting the initial image processing model according to the result obtained by the discrimination sub-module and the second face sample image in the training samples, the training unit 34 is specifically configured to calculate a loss function related to the initial image processing model according to the result obtained by the discrimination sub-module and the second face sample image in the training samples, and adjust the parameter values of the parameters in the initial image processing model according to the loss function. Among them, when calculating the loss function related to the initial image processing model according to the result obtained by the discrimination sub-module, the training unit 34 is specifically configured to calculate a feature loss sub-function, an adversarial loss sub-function, and a reconstruction loss sub-function; wherein, the feature loss sub-function is used to represent the difference between the feature information of the processed image of the first face sample image obtained by the processing sub-module and the feature information of the second face sample image; the adversarial loss sub-function is used to represent the result of the discrimination sub-module discriminating whether the processed image of the first face sample image obtained by the processing sub-module is genuine, and the expectation of the result of the discrimination sub-module discriminating whether the second face sample image is genuine; the reconstruction loss sub-function is used to represent the difference between the processed image of the first face sample image obtained by the processing sub-module and the second face sample image; the function calculation values of the feature loss sub-function, the adversarial loss sub-function, and the reconstruction loss sub-function are used as the loss function.

[0136] Further, the training unit 34 is further configured to stop adjusting the parameter values when the number of times of adjusting the parameter values is equal to a preset number of times, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold value.

[0137] It can be seen that in the image processing system of this embodiment, the 3D image unit 31 can obtain the 3D image of the target face image according to the 3D coefficients of the target face image, and the model calling unit 32 calls the pre-trained image processing model. In the processing unit 33, the image processing model obtains the processed image of the target face image according to the merged data of the target face image and its 3D image, and the resolution of the processed image is higher than that of the target face image. In this way, when processing a target face image with low clarity (i.e., low resolution), the corresponding 3D image is combined. The data of the 3D image can describe the face image in more detail, so that the effect of the processed image obtained by processing is well improved.

[0138] An embodiment of the present invention further provides a terminal device, and its structural schematic diagram is as Figure 11 shown. The terminal device may vary greatly due to configuration or performance, and may include one or more central processing units (CPUs) 40 (for example, one or more processors) and a memory 41, and one or more storage media 42 for storing application programs 421 or data 422 (for example, one or more mass storage devices). Among them, the memory 41 and the storage media 42 may be transient storage or persistent storage. The program stored in the storage media 42 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the terminal device. Further, the central processing unit 40 may be configured to communicate with the storage media 42 and execute a series of instruction operations in the storage media 42 on the terminal device.

[0139] Specifically, the application program 421 stored in the storage media 42 includes an application program for image processing, and this program may include the coefficient acquisition unit 30, the 3D image unit 31, the model calling unit 32, the processing unit 33, and the training unit 34 in the above image processing system, which will not be elaborated here. Further, the central processing unit 40 may be configured to communicate with the storage media 42 and execute a series of operations corresponding to the image processing application program stored in the storage media 42 on the terminal device.

[0140] The terminal device may further include one or more power supplies 43, one or more wired or wireless network interfaces 44, one or more input / output interfaces 45, and / or one or more operating systems 423, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0141] The steps performed by the image processing system in the above method embodiment may be based on the Figure 11 structure of the shown terminal device.

[0142] Further, on the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing a plurality of computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the image processing method executed by the above image processing system.

[0143] On the other hand, an embodiment of the present invention further provides a terminal device, including a processor and a memory;

[0144] The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the image processing method executed by the above image processing system; the processor is used to implement each of the computer programs in the plurality of computer programs.

[0145] In addition, according to one aspect of the present application, there is provided a computer program product or a computer program, and the computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image processing method provided in the above various alternative implementation manners.

[0146] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0147] The above has introduced in detail an image processing method, system, storage medium and terminal device provided by an embodiment of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An image processing method, characterized in that, Including: Obtain a target face image and three-dimensional coefficients of the face image included in the target face image; Obtain a three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image; Call a pre-trained image processing model, where the image processing model includes: a mapping module, an encoding module, a decoding module, and an output module; The image processing model obtains a processed image of the target face image according to the combined data of the target face image and the three-dimensional image, and outputs the processed image. The resolution of the processed image is higher than that of the target face image, including: the encoding module in the image processing model extracts spatial features of the combined data of the target face image and the three-dimensional image, the mapping module obtains a latent variable according to the combined data of the target face image and the three-dimensional image to adjust the convolutional weights in the decoding module, and then the decoding module performs convolutional calculation on the spatial features extracted by the encoding module based on the convolutional weights adjusted by the latent variable, so as to obtain post-convolution features. Finally, the output module outputs a processed image of the target face image according to the post-convolution features.

2. The method according to claim 1, wherein The obtaining of the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image specifically includes: Reconstruct three-dimensional network information of the target face image according to the three-dimensional coefficients of the target face image; the three-dimensional network information includes the shapes and textures of multiple faces constituting the three-dimensional face, and the three-dimensional face is the three-dimensional face corresponding to the face included in the target face image; Project the three-dimensional network information onto a two-dimensional plane by means of rendering to obtain a three-dimensional image of the target face image.

3. The method according to claim 1, wherein The encoding module includes multiple encoding modules, and the decoding module includes multiple decoding modules, where: The output of one encoding module is connected to one input of one decoding module, and the output of the mapping module is respectively connected to the other inputs of the multiple decoding modules; the multiple encoding modules are connected in series, the multiple decoding modules are connected in series, and the output of the last decoding module among the multiple decoding modules is connected to the output module.

4. The method according to claim 3, wherein The encoding module in the image processing model extracts spatial features of the combined data of the target face image and the three-dimensional image, the mapping module obtains a latent variable according to the combined data of the target face image and the three-dimensional image to adjust the convolutional weights in the decoding module, and then the decoding module performs convolutional calculation on the spatial features extracted by the encoding module based on the convolutional weights adjusted by the latent variable, so as to obtain post-convolution features. Finally, the output module outputs a processed image of the target face image according to the post-convolution features, specifically including: Determine a latent variable through the mapping module by combining the data of the target face image and the three-dimensional image, and the latent variable is used to adjust the convolutional weights of the convolutional calculation involved in the decoding module; Extract spatial features of a resolution of the combined data of the target face image and the three-dimensional image through the encoding module; The second decoding module among the multiple decoding modules performs convolution calculation based on the latent variable, the spatial features of one resolution extracted by one encoding module, and the feature information of the one resolution obtained by the previous decoding module of the second decoding module, to obtain the feature information of another resolution; The output module obtains the processed image of the target face image according to the feature information obtained by the last second decoding module.

5. The method according to claim 4, wherein The second decoding module performs convolution calculation based on the latent variable, the spatial features of one resolution extracted by one encoding module, and the feature information of the one resolution obtained by the previous decoding module of the second decoding module, to obtain the feature information of another resolution, specifically including: The second decoding module adjusts the convolution weights involved in the convolution calculation according to the latent variable, to obtain the adjusted convolution weights, and performs convolution calculation on the feature information of the one resolution according to the adjusted convolution weights to obtain the convolved features; According to the spatial features of the one resolution and the convolved features, the feature information of another resolution is obtained.

6. The method according to claim 4, wherein In the process that the image processing model obtains the processed image of the target face image according to the merged data of the target face image and the three-dimensional image, the first decoding module among the multiple decoding modules performs convolution calculation based on the latent variable and the spatial features of one resolution obtained by one encoding module, to obtain the feature information of another resolution.

7. The method according to any one of claims 1 to 6, wherein Determine an initial image processing model; the initial image processing model includes: a processing sub-module and a discrimination sub-module; Determine training samples, where the training samples include multiple sample groups, and each sample group includes a low-resolution first face sample image, its corresponding high-resolution second face sample image, and a three-dimensional face image; The processing sub-module obtains the processed image of the first face sample image according to the merged data of the first face sample image and its corresponding three-dimensional face image, and the discrimination sub-module discriminates whether the processed image obtained by the processing sub-module is real; According to the result obtained by the discrimination sub-module and the second face sample image in the training samples, adjust the initial image processing model, and the processing sub-module in the adjusted initial image processing model is the pre-trained image processing model.

8. The method according to claim 7, characterized in that The first face sample image is obtained by performing downsampling on the second face sample image, and the downsampling includes any one of the following processing methods: adding blur, downsampling, noise, and Joint Photographic Experts Group compression processing.

9. The method according to claim 7, wherein The adjusting the initial image processing model according to the result obtained by the discrimination sub-module and the second face sample image in the training samples specifically includes: Calculating a loss function related to the initial image processing model according to the result obtained by the discrimination sub-module and the second face sample image in the training samples, and adjusting the parameter values of the parameters in the initial image processing model according to the loss function.

10. The method according to claim 9, wherein Calculating a loss function related to the initial image processing model based on the result obtained from the discrimination sub-module and the second face sample image in the training sample, specifically including: Calculating a feature loss sub-function, an adversarial loss sub-function, and a reconstruction loss sub-function; Among them, the feature loss sub-function is used to represent the difference between the feature information of the processed image of the first face sample image obtained by the processing sub-module and the feature information of the second face sample image; the adversarial loss sub-function is used to represent the result of whether the processed image of the first face sample image obtained by the discrimination sub-module for discriminating the processing sub-module is true, and the expectation of the result of whether the discrimination sub-module discriminates the second face sample image is true; the reconstruction loss sub-function is used to represent the difference between the processed image of the first face sample image obtained by the processing sub-module and the second face sample image; Taking the function calculation values of the feature loss sub-function, the adversarial loss sub-function, and the reconstruction loss sub-function as the loss function.

11. The method according to claim 9, wherein When the number of adjustments to the parameter value is equal to the preset number, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, then stop adjusting the parameter value.

12. An image processing system, characterized in that, Including: A coefficient acquisition unit for acquiring a target face image and the three-dimensional coefficients of the target face image; A three-dimensional image unit for acquiring the three-dimensional image of the target face image according to the three-dimensional coefficients of the target face image; A model call unit for calling a pre-trained image processing model, where the image processing model includes: a mapping module, an encoding module, a decoding module, and an output module; A processing unit for the image processing model to obtain a processed image of the target face image according to the combined data of the target face image and the three-dimensional image, and output the processed image, and the resolution of the processed image is higher than the resolution of the target face image, including: the encoding module in the image processing model extracts the spatial features of the combined data of the target face image and the three-dimensional image, the mapping module obtains a latent variable according to the combined data of the target face image and the three-dimensional image to adjust the convolutional weights in the decoding module, and then the decoding module performs convolutional calculation on the spatial features extracted by the encoding module based on the convolutional weights adjusted by the latent variable, so as to obtain the post-convolution features, and finally the output module outputs the processed image of the target face image according to the post-convolution features.

13. An image processing method, characterized in that, Including: Determining an initial image processing model; The initial image processing model includes: a processing sub-module and a discrimination sub-module; Determining a training sample, where the training sample includes multiple sample groups, and each sample group includes a low-resolution first face sample image and its corresponding high-resolution second face sample image and three-dimensional face image; The processing sub-module obtains a processed image of the first face sample image according to the combined data of the first face sample image and its corresponding three-dimensional face image, and the discrimination sub-module discriminates whether the processed image obtained by the processing sub-module is true; According to the result obtained by the discrimination sub-module and the second face sample image in the training sample, adjust the initial image processing model. The processing sub-module in the adjusted initial image processing model is an image processing model, which is used to obtain a corresponding high-resolution image according to the combined data of any low-resolution image and its three-dimensional image. Wherein, the image processing model includes: a mapping module, an encoding module, a decoding module, and an output module; including: the encoding module in the image processing model extracts the spatial features of the combined data of the low-resolution image and the three-dimensional image, and the mapping module obtains a latent variable according to the combined data of the low-resolution image and the three-dimensional image to adjust the convolution weights in the decoding module. Then, based on the convolution weights adjusted by the latent variable, the decoding module performs convolution calculation on the spatial features extracted by the encoding module to obtain the convolved features. Finally, the output module outputs the processed image of the low-resolution image according to the convolved features.

14. The method according to claim 13, wherein The determination of the initial image processing model specifically includes: Determine that the processing sub-module in the initial image processing model includes the following structure: The processing sub-module includes: a mapping module, a plurality of encoding modules, a plurality of decoding modules, and an output module, wherein: The output of one encoding module is connected to one input of one decoding module, and the output of the mapping module is respectively connected to the other inputs of the plurality of decoding modules; the plurality of encoding modules are connected in series, the plurality of decoding modules are connected in series, and the output of the last decoding module among the plurality of decoding modules is connected to the output module.

15. The method according to claim 13 or 14, characterized in that, The adjustment of the initial image processing model according to the result obtained by the discrimination sub-module and the second face sample image in the training sample specifically includes: Calculate the feature loss sub-function, the adversarial loss sub-function, and the reconstruction loss sub-function; Wherein, the feature loss sub-function is used to represent the difference between the feature information of the processed image of the first face sample image obtained by the processing sub-module and the feature information of the second face sample image; the adversarial loss sub-function is used to represent the result of whether the discrimination sub-module discriminates that the processed image of the first face sample image obtained by the processing sub-module is true, and the expectation of the result of whether the discrimination sub-module discriminates that the second face sample image is true; the reconstruction loss sub-function is used to represent the difference between the processed image of the first face sample image obtained by the processing sub-module and the second face sample image; Take the function calculation values of the feature loss sub-function, the adversarial loss sub-function, and the reconstruction loss sub-function as the loss function; Adjust the parameter values of the parameters in the initial image processing model according to the loss function.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of computer programs, and the computer programs are suitable for being loaded and executed by a processor to execute the image processing method according to any one of claims 1 to 11, or to execute the image processing method according to any one of claims 13 to 15.

17. A terminal device, characterized in that, Including a processor and a memory; The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by a processor to perform the image processing method according to any one of claims 1 to 11, or to perform the image processing method according to any one of claims 13 to 15; the processor is used to implement each of the plurality of computer programs.

18. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the image processing method according to any one of claims 1 to 11, or performs the image processing method according to any one of claims 13 to 15.

Citation Information

Patent Citations

  • Face super-resolution method combined with 3D face structure prior

    CN113034370A