Image processing method, model training method, device and equipment

By acquiring the multi-dimensional features of facial images and processing them using the first machine learning model, the problem of insufficient generalization ability of machine learning models in the existing technology is solved, and high-quality image processing is achieved in various application scenarios.

CN113628122BActive Publication Date: 2025-09-16ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010388839.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-09
Publication Date
2025-09-16
Estimated Expiration
2040-05-09

AI Technical Summary

Technical Problem

Existing technologies use machine learning models trained with artificial synthetic datasets, but they lack generalization capabilities when processing images and cannot adapt to various application scenarios, resulting in poor image processing quality and efficiency.

Method used

By obtaining multi-dimensional features of the facial image to be processed, including key point features, contour features, texture features and color features, a first machine learning model is used to perform image processing based on these features to obtain a target image.

Benefits of technology

It realizes the effective processing of face images in any application scenario, improves the quality and effect of image processing, and expands the scope of application and practicality of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113628122B_ABST
    Figure CN113628122B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention provide an image processing method, a model training method, an apparatus and a device, the method comprising: obtaining a facial image to be processed; determining multidimensional features corresponding to the facial image, the multidimensional features comprising at least two different image features corresponding to the facial image; inputting the multidimensional features and the facial image into a first machine learning model, so that the first machine learning model processes the facial image based on the multidimensional features to obtain a target image corresponding to the facial image; the first machine learning model is trained to determine a target image corresponding to the facial image based on the multidimensional features, the clarity of the target image being different from the clarity of the facial image. The technical solution provided in this embodiment can realize processing operations on facial images in any application scenario through multidimensional features, and also ensures the quality and effect of image processing, so that the method can be widely applied to various application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method, a model training method, a device and equipment. Background Art

[0002] In the field of image processing technology, enhancing and clarifying blurred facial images in images or videos has a wide range of application scenarios. For example, in surveillance and security, enhancing low-definition facial images can help determine the identity of the person being monitored, or repairing facial images in old photos, old movies and TV dramas can not only improve the image quality, but also enhance the audience's viewing experience.

[0003] Currently, when performing image enhancement processing, a machine learning model can be generated by training artificially synthesized data sets, and then image enhancement processing can be performed on blurred face images based on the above machine learning model.

[0004] However, although the use of artificially synthesized datasets to learn and train machine learning models has good performance, artificially synthesized datasets cannot cover all application scenarios included in actual scenarios. Therefore, when using machine learning models for image processing, they do not have generalization capabilities and cannot guarantee the quality and efficiency of image processing in various application scenarios. Summary of the Invention

[0005] The embodiments of the present invention provide an image processing method, a model training method, an apparatus and equipment, which can realize processing operations on facial images in any application scenario through multi-dimensional features, and also ensure the quality and effect of image processing, so that the method can be widely applied to various application scenarios.

[0006] In a first aspect, an embodiment of the present invention provides an image processing method, comprising:

[0007] Obtain the face image to be processed;

[0008] Determining a multi-dimensional feature corresponding to the facial image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image;

[0009] Inputting the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image;

[0010] In which, the first machine learning model is trained to determine a target image corresponding to the facial image based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the facial image.

[0011] In a second aspect, an embodiment of the present invention provides an image processing device, including:

[0012] A first acquisition module is used to acquire a face image to be processed;

[0013] a first determining module, configured to determine a multi-dimensional feature corresponding to the facial image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image;

[0014] a first processing module, configured to input the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image;

[0015] In which, the first machine learning model is trained to determine a target image corresponding to the facial image based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the facial image.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the image processing method in the above-mentioned first aspect.

[0017] In a fourth aspect, an embodiment of the present invention provides a computer storage medium for storing a computer program, wherein the computer program enables a computer to implement the image processing method in the first aspect when executed.

[0018] In a fifth aspect, an embodiment of the present invention provides an image processing method, including:

[0019] Get the image to be processed;

[0020] Determining a multi-dimensional feature corresponding to the image to be processed, wherein the multi-dimensional feature includes at least two different image features corresponding to the image to be processed;

[0021] Inputting the multidimensional features and the image to be processed into a first machine learning model, so that the first machine learning model processes the image to be processed based on the multidimensional features to obtain a target image corresponding to the image to be processed;

[0022] In which, the first machine learning model is trained to determine a target image corresponding to the image to be processed based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the image to be processed.

[0023] In a sixth aspect, an embodiment of the present invention provides an image processing device, including:

[0024] A second acquisition module is used to acquire the image to be processed;

[0025] a second determining module, configured to determine a multi-dimensional feature corresponding to the image to be processed, wherein the multi-dimensional feature includes at least two different image features corresponding to the image to be processed;

[0026] a second processing module, configured to input the multi-dimensional features and the image to be processed into a first machine learning model, so that the first machine learning model processes the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed;

[0027] In which, the first machine learning model is trained to determine a target image corresponding to the image to be processed based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the image to be processed.

[0028] In the seventh aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the image processing method in the above-mentioned fifth aspect.

[0029] In an eighth aspect, an embodiment of the present invention provides a computer storage medium for storing a computer program, wherein the computer program enables a computer to implement the image processing method in the fifth aspect when executed.

[0030] In a ninth aspect, an embodiment of the present invention provides a model training method, comprising:

[0031] Acquire a first image and a reference image corresponding to the first image, wherein the definition of the reference image is different from the definition of the first image;

[0032] determining a multi-dimensional feature corresponding to the first image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image;

[0033] Learning and training are performed based on the first image, the reference image and the multidimensional features to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multidimensional features, and the clarity of the target image is different from the clarity of the first image.

[0034] In a tenth aspect, an embodiment of the present invention provides a model training device, comprising:

[0035] a third acquisition module, configured to acquire a first image and a reference image corresponding to the first image, wherein the clarity of the reference image is different from that of the first image;

[0036] a third determining module, configured to determine a multi-dimensional feature corresponding to the first image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image;

[0037] The third processing module is used to perform learning and training based on the first image, the reference image and the multidimensional features to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multidimensional features, and the clarity of the target image is different from the clarity of the first image.

[0038] In the eleventh aspect, an embodiment of the present invention provides an electronic device comprising: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the model training method in the above-mentioned ninth aspect.

[0039] In the twelfth aspect, an embodiment of the present invention provides a computer storage medium for storing a computer program, which enables the computer to implement the model training method in the above-mentioned ninth aspect when executed.

[0040] The image processing method, model training method, device and equipment provided in this embodiment obtain a facial image to be processed; determine the multi-dimensional features corresponding to the facial image, and then the first machine learning model can perform image processing on the facial image based on the multi-dimensional features, thereby realizing the processing operation of the facial image in any application scenario through the multi-dimensional features, making the method widely applicable to various application scenarios, thereby effectively improving the practicality of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A schematic flow chart of an image processing method provided by an embodiment of the present invention;

[0043] Figure 2 An application scenario diagram of an image processing method provided by an embodiment of the present invention Figure 1 ;

[0044] Figure 3 An application scenario diagram of an image processing method provided by an embodiment of the present invention Figure 2 ;

[0045] Figure 4 A schematic diagram of a process for analyzing and processing the facial image using a second machine learning model to determine multi-dimensional features corresponding to the facial image, provided in an embodiment of the present invention;

[0046] Figure 5 A schematic diagram of an embodiment of the present invention providing a method for analyzing and processing the facial image using a second machine learning model to determine multi-dimensional features corresponding to the facial image;

[0047] Figure 6 A schematic diagram of inputting the multi-dimensional features and the facial image into a first machine learning model provided by an embodiment of the present invention;

[0048] Figure 7 A schematic flow chart of another image processing method provided by an embodiment of the present invention;

[0049] Figure 8 A schematic diagram of a process for determining multi-dimensional features corresponding to the facial image provided by an embodiment of the present invention;

[0050] Figure 9 A schematic diagram of a process for obtaining a modulation function corresponding to the convolution kernel provided in an embodiment of the present invention;

[0051] Figure 10 A schematic diagram of a process for determining a modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector provided in an embodiment of the present invention;

[0052] Figure 11 A flowchart of an image processing method provided by an application embodiment of the present invention;

[0053] Figure 12 A schematic flow chart of another image processing method provided by an embodiment of the present invention;

[0054] Figure 13 A flowchart of a model training method provided by an embodiment of the present invention;

[0055] Figure 14 A schematic structural diagram of an image processing device provided by an embodiment of the present invention;

[0056] Figure 15 For Figure 14 A schematic structural diagram of an electronic device corresponding to the image processing device provided by the illustrated embodiment;

[0057] Figure 16 A schematic structural diagram of another image processing device provided by an embodiment of the present invention;

[0058] Figure 17 For Figure 16 A schematic structural diagram of an electronic device corresponding to the image processing device provided by the illustrated embodiment;

[0059] Figure 18 A schematic structural diagram of a model training device provided by an embodiment of the present invention;

[0060] Figure 19 For Figure 18 A schematic structural diagram of an electronic device corresponding to the model training device provided in the illustrated embodiment. DETAILED DESCRIPTION

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0062] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. "A plurality" generally includes at least two, but does not exclude the inclusion of at least one.

[0063] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0064] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0065] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or system. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or system comprising the element.

[0066] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0067] In order to facilitate understanding of the technical solution of this application, the following is a brief description of the prior art:

[0068] With the rapid development of science and technology, users have increasingly higher expectations for video and photo quality. However, some older photos and classic films are often blurry, resulting in a poor viewing experience. Facial images play a crucial role in film and television productions. Whether in film or photographs, these images often contain people, and users are more sensitive to the clarity of these images.

[0069] Common deep learning image restoration algorithms based on convolutional neural networks (CNN) all use image pairs (low-quality images, high-quality images) to train the network. Among them, deep learning image restoration algorithms can include at least one of the following: Super-Resolution Generative Adversarial Networks (SRGAN), Deep Residual Channel Attention Networks RCAN, Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN), etc.

[0070] However, when using the above-mentioned trained convolutional neural network for image processing, there are the following disadvantages: the low-quality images used in the above-mentioned network training are generally obtained through artificial downsampling, which easily makes the trained network unsuitable for real low-quality facial images; in addition, the prior knowledge of facial structure is not fully utilized to accurately and reliably process the image.

[0071] In order to solve the above technical problems, this embodiment provides an image processing method, a model training method, an apparatus and a device, which obtains a facial image to be processed and determines a multidimensional feature corresponding to the facial image. The above-mentioned multidimensional feature includes at least two different image features corresponding to the facial image, for example: the multidimensional feature may include at least two of the following: key point features, contour features, texture features, and color features; after obtaining the multidimensional feature, the multidimensional feature and the facial image can be input into a first machine learning model, so that a target image corresponding to the facial image can be obtained, thereby realizing that the facial image in any application scenario (real scenario) can be processed based on the multidimensional feature corresponding to the facial image, thereby ensuring the quality and effect of image processing, reducing the difficulty of image processing, and making the method widely applicable to various application scenarios, further improving the scope of application and practicality of the method.

[0072] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflicts between the embodiments, the following embodiments and features therein may be combined with each other.

[0073] Figure 1 A flowchart of an image processing method provided by an embodiment of the present invention; Figure 1 As shown, this embodiment provides an image processing method. The execution subject of the method may be an image processing device. It is understood that the image processing device can be implemented as software or a combination of software and hardware. Specifically, the processing method may include:

[0074] Step S101: Acquire a face image to be processed.

[0075] Step S102: Determine a multi-dimensional feature corresponding to the facial image, where the multi-dimensional feature includes at least two different image features corresponding to the facial image.

[0076] Step S103: Input the multi-dimensional features and the facial image into the first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image.

[0077] The following is a detailed explanation of each of the above steps:

[0078] Step S101: Acquire a face image to be processed.

[0079] Among them, the facial image to be processed refers to a facial image that needs to be processed. It can be understood that the above-mentioned image processing may include at least one of the following: image enhancement processing, image blurring processing, image rendering processing, image editing processing, etc. Specifically, the above-mentioned image enhancement processing can increase the clarity, local details, etc. of the facial image display, image blurring processing can reduce the clarity, local details, etc. of the facial image display, image rendering processing can perform whitening, beautification and other rendering processing on the facial subject in the facial image, and image editing processing can perform various types of editing operations on the facial image, such as image filtering processing, image texture processing, image cropping processing, etc.

[0080] In addition, the facial image to be processed may include at least one of the following: image information captured by a camera, image information in video information, a composite image, and the like. It is understood that the number of images to be processed may be one or more. When there are multiple images to be processed, the multiple images to be processed may constitute an image sequence, thereby enabling image processing operations to be performed on the image sequence. Furthermore, the image type of the images to be processed may be static images or dynamic images, thereby enabling image processing operations to be performed on static images or dynamic images.

[0081] In addition, this embodiment does not limit the specific implementation method for the image processing device to obtain the facial image to be processed. Those skilled in the art can make settings based on specific application requirements and design requirements. For example, the shooting device can be connected to the enhancement device for communication. After the shooting device captures the facial image to be processed, the image processing device can obtain the facial image to be processed through the shooting device. Specifically, the image processing device can actively obtain the facial image to be processed obtained by the shooting device, or the shooting device can actively send the facial image to be processed to the enhancement device, so that the image processing device can obtain the facial image to be processed. Alternatively, the facial image to be processed can be stored in a preset area, and the image processing device can obtain the facial image to be processed by accessing the preset area.

[0082] Step S102: Determine a multi-dimensional feature corresponding to the facial image, where the multi-dimensional feature includes at least two different image features corresponding to the facial image.

[0083] After acquiring the facial image, the facial image can be analyzed and processed to determine the multidimensional features corresponding to the facial image. The above-mentioned multidimensional features can include at least two different image features corresponding to the facial image. For example, the multidimensional features can include at least two of the following: key point features, contour features, texture features, and color features.

[0084] In addition, this embodiment does not limit the specific implementation method for determining the multi-dimensional features corresponding to the facial image. Those skilled in the art can set it according to specific application requirements and design requirements. For example, one achievable method can determine the multi-dimensional features corresponding to the facial image through a preset machine learning model. Specifically, determining the multi-dimensional features corresponding to the facial image may include:

[0085] Step S1021: Analyze and process the facial image using the second machine learning model to determine the multidimensional features corresponding to the facial image. The second machine learning model is trained to determine the multidimensional features corresponding to the facial image.

[0086] Among them, the second machine learning model can be pre-trained to determine the multi-dimensional features corresponding to the facial image. It can be understood that in different application scenarios, the number of multi-dimensional features corresponding to the facial image determined may be the same or different.

[0087] In addition, a second machine learning model can be generated by training a convolutional neural network, that is, the convolutional neural network is trained using a preset reference image and multi-dimensional features corresponding to the reference image, thereby obtaining the second machine learning model. After the second machine learning model is generated, the second machine learning model can be used to analyze and process the facial image, thereby obtaining multi-dimensional features corresponding to the facial image.

[0088] In this embodiment, the facial image is analyzed and processed by the trained second machine learning model to obtain multi-dimensional features corresponding to the facial image. This not only effectively ensures the accuracy and reliability of obtaining the multi-dimensional features, but also ensures the quality and efficiency of obtaining the target image based on the multi-dimensional features, further improving the stability and reliability of the use of this method.

[0089] Of course, those skilled in the art may also adopt other methods to determine the multi-dimensional features corresponding to the facial image, as long as they can ensure that the multi-dimensional features corresponding to the facial image are accurately obtained, which will not be elaborated here.

[0090] Step S103: Input the multi-dimensional features and the facial image into the first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image.

[0091] Among them, after obtaining the multi-dimensional features, the multi-dimensional features and the facial image can be input into the first machine learning model, so that the first machine learning model can analyze and process the facial image based on the multi-dimensional features, thereby realizing the analysis and processing of the facial image using the multi-dimensional features as the guiding features of image processing, thereby ensuring the quality and efficiency of processing the facial image, that is, obtaining the target image corresponding to the facial image. The above-mentioned first machine learning model is trained to determine the target image corresponding to the facial image based on the multi-dimensional features. It should be noted that the above-mentioned second machine learning model and the first machine learning model can be different machine learning models, or the second machine learning model and the first machine learning model can be the same machine learning model.

[0092] In addition, the clarity of the target image obtained is different from the clarity of the facial image, that is, the relationship between the clarity of the target image and the clarity of the facial image can include: the clarity of the target image is higher than the clarity of the image to be processed; or, the clarity of the target image is lower than the clarity of the facial image. It can be understood that when the clarity of the target image is higher than the clarity of the facial image, the first machine learning model is trained to determine the target image for enhancing the facial image based on multi-dimensional features. When the clarity of the target image is lower than the clarity of the facial image, the second machine learning model is trained to determine the target image for blurring the facial image based on multi-dimensional features.

[0093] In addition, when a target image corresponding to a facial image is obtained, the number of target images may be at least one. When the number of target images is multiple, a final target image may be determined based on the similarity between the multiple target images and the facial image. Specifically, at least one similarity may correspond to at least one target image and the facial image, and the similarity between the target image and the facial image may include: the similarity between the structure and appearance of the face in the target image and the structure and appearance of the face in the facial image. The above-mentioned structure of the face includes at least one of the following: facial orientation (front, left, right, etc.), posture (head up, head down, etc.), and position information of the face relative to the image (center position, left position, right position, etc.); the appearance of the face includes at least one of the following: hair features, skin color features, brightness features, and color features.

[0094] It is understood that the similarities between different target images and facial images may be the same or different. After obtaining the similarities between the facial image and the different target images, at least one target image may be sorted based on the similarity, thereby obtaining a sorted queue of at least one target image based on different similarities. Based on the sorted queue, a target image with the highest similarity may be obtained, and this selected target image may be determined as the final target reference image. This effectively ensures the quality and effectiveness of image processing.

[0095] For example 1, refer to the attached Figure 2 As shown, an image processing method capable of implementing an image enhancement operation is used as an example for explanation. In this case, the execution subject of the image processing method is an image processing device, which is communicatively connected to a client. When a user has an image enhancement requirement, an image processing request corresponding to the image enhancement requirement can be generated on the client. The image processing request corresponds to a face image, and then the client can transmit the generated image processing request and the face image to the image processing device. After receiving the image processing request and the face image, the image processing device can process the face image based on the image processing request, specifically including:

[0096] Step 1: Receive image processing request and face image.

[0097] Step 2: Process the face image to obtain multi-dimensional features corresponding to the face image.

[0098] Step 3: Input the facial image and the multi-dimensional features into a preset first machine learning model to obtain a target image corresponding to the facial image, where the clarity of the target image is higher than that of the facial image.

[0099] Step 4: The target image is transmitted to the client, so that the client can display the target image through a preset display area, so that the user can view the target image after image enhancement processing.

[0100] For example 2, refer to the attached Figure 3 As shown, an image processing method capable of implementing an image blur operation is used as an example for explanation. In this case, the execution subject of the image processing method is an image processing device, which is communicatively connected to a client. When a user has an image blur requirement, an image processing request corresponding to the image blur requirement can be generated on the client. The image processing request corresponds to a face image, and then the client can transmit the generated image processing request and the face image to the image processing device. After receiving the image processing request and the face image, the image processing device can process the face image based on the image processing request, specifically including:

[0101] Step 1: Receive image processing request and face image.

[0102] Step 2: Process the face image to obtain multi-dimensional features corresponding to the face image.

[0103] Step 3: Input the facial image and the multi-dimensional features into a preset first machine learning model to obtain a target image corresponding to the facial image, where the clarity of the target image is lower than that of the facial image.

[0104] Step 4: The target image is transmitted to the client, so that the client can display the target image through a preset display area, so that the user can view the target image after the image blurring process.

[0105] The image processing method provided in this embodiment obtains a facial image to be processed, determines the multi-dimensional features corresponding to the facial image to be processed, and inputs the multi-dimensional features and the facial image into a first machine learning model, thereby enabling the first machine learning model to use the multi-dimensional features as guiding information for analyzing and processing the facial image, and then obtaining a target image corresponding to the facial image; this effectively achieves the need to obtain a high-definition facial image, that is, it is possible to perform image processing operations in any application scenario (real scenario), and also ensures the quality and effect of image processing, reduces the difficulty of image processing, and enables the image processing method to be widely applicable to various application scenarios, further improving the scope of application and practicality of the method.

[0106] In some instances, when analyzing and processing a facial image using a second machine learning model to determine multi-dimensional features corresponding to the facial image, the second machine learning model includes: one or more second network units, the plurality of second network units being connected in series, the second network units being configured to analyze and process received second input information to determine second output information corresponding to the second input information. The second input information may include any one of the following: a facial image, or second output information output by a second network unit at a previous level.

[0107] For details, please refer to the attached Figure 4-Figure 5 As shown, in this embodiment, analyzing and processing the facial image using the second machine learning model to determine the multi-dimensional features corresponding to the facial image may include:

[0108] Step S401: When analyzing and processing a facial image using a second machine learning model, obtain one or more second output information output by one or more second network units.

[0109] Step S402: Determine one or more second output information as multi-dimensional features corresponding to the facial image.

[0110] Among them, the second machine learning model may include one or more second network units. When the second machine learning model is used to analyze and process the facial image, that is, one or more second network units are used to analyze and process the facial image. Since multiple second network units are connected in series in sequence, the second network unit at the next level can obtain the analysis and processing results (second output information) of the second network unit at the previous level, and analyze and process the analysis and processing results of the second network unit at the previous level to determine the multi-dimensional features corresponding to the facial image.

[0111] For example, if Figure 5 As shown, the second machine learning model may include: a second network unit A1, a second network unit A2...a second network unit An and a second network unit An+1; wherein, the output port of the A1 unit is communicatively connected to the input port of the A2 unit, the output port of the An-1 unit is communicatively connected to the input port of the An unit, and the output port of the An unit is communicatively connected to the input port of the An+1 unit, thereby realizing a plurality of second network units connected in series in sequence.

[0112] After acquiring a facial image, the facial image can be input into unit A1. Unit A1 can analyze and process the facial image to obtain second output information B1 corresponding to the facial image. After acquiring second output information B1, unit B1 can be sent to unit A2. Unit A2 can then analyze and process B1 to obtain second output information B2 corresponding to B1. Similarly, when unit An-1 generates second output information Bn-1, it can send Bn-1 to the second network unit An. After unit An acquires Bn-1, it can analyze and process Bn-1 to obtain second output information Bn. Bn can then be sent to unit An+1, thereby acquiring one or more second output information output by one or more second network units.

[0113] After obtaining one or more second output information, one or more second output information can be determined as multi-dimensional features corresponding to the facial image. At this time, the multi-dimensional features may include at least two different image features corresponding to the facial image (the second output information B1 output by the A1 unit, the second output information B2 output by the A2 unit...the second output information Bn output by the An unit, the second output information Bn+1 output by the An+1 unit), thereby effectively ensuring the accuracy and reliability of the acquisition of the multi-dimensional features corresponding to the facial image.

[0114] In some instances, the first machine learning model may include: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine first output information corresponding to the first input information.

[0115] The first input information may include any one of the following: multi-dimensional features corresponding to the facial image, guiding feature information, and first output information output by the first network unit at the previous level. The guiding feature information may include at least one of the following: a facial semantic map, a key point location map, and a heat map. It is understood that the guiding feature information may refer to feature information input into the first machine learning model based on application requirements and design requirements.

[0116] In addition, when the first input information of the first network unit includes multidimensional features corresponding to the facial image, that is, the second output information of the second network unit in the second machine learning model can be input into the first network unit, so that the first network unit can analyze and process the facial image based on the second output information. The number of the above-mentioned first network units and second network units can be the same or different. When the number of first network units is greater than the number of second network units, the multiple second output information output by the multiple second network units can be input into part of the first network units. When the number of first network units is less than the number of second network units, the multiple second output information output by the multiple second network units can be input into part or all of the first network units. When the number of first network units is equal to the number of second network units, the multiple second output information output by the multiple second network units can be input into part or all of the first network units.

[0117] In this embodiment, the first machine learning model may include one or more first network units, and the multiple first network units are connected in series in sequence. In addition, the first input information of the first network unit may include any one of the following: multi-dimensional features corresponding to the facial image, guiding feature information, and the first output information output by the first network unit at the previous level, thereby effectively ensuring that the first machine learning model can stably and reliably process the image to be processed based on the multi-dimensional features as guiding information, further ensuring the quality and efficiency of image processing.

[0118] Figure 7 A flow chart of another image processing method provided by an embodiment of the present invention; based on the above embodiment, continue to refer to the attached Figure 7 As shown, the method in this embodiment may further include:

[0119] Step S701: Acquire guidance feature information.

[0120] Step S702: Input the guiding feature information into the first network unit included in the first machine learning model, so that the first network unit processes the face image based on the guiding feature information and the multi-dimensional features to obtain a target image corresponding to the face image.

[0121] To further improve the quality and efficiency of image processing, guidance feature information for analyzing and processing facial images can be obtained. It is understood that the guidance feature information can be input by a user into the image processing device, or the guidance feature information can be sent to the image processing device by another device, or the guidance feature information can be stored in a preset area of ​​the image processing device and can be obtained by accessing the preset area. Of course, those skilled in the art can also use other methods to obtain the guidance feature information, as long as the accuracy and reliability of the guidance feature information can be guaranteed, and these will not be elaborated here.

[0122] After obtaining the guiding feature information, the guiding feature information can be input into the first network unit included in the first machine learning model, so that the first network unit can process the facial image based on the guiding feature information and the multi-dimensional features, thereby obtaining a target image corresponding to the facial image. Specifically, since the first machine learning model includes one or more first network units, when obtaining the target image corresponding to the facial image, it can include: in one or more first network units, the first output information output by the last-level first network unit can be determined as the target image corresponding to the facial image, thereby effectively ensuring the quality and efficiency of the analysis and processing of the facial image.

[0123] It should be noted that the number of first network units included in the first machine learning model may vary due to different application scenarios and application requirements. That is, when learning and training the first machine learning model, a first machine learning model including different numbers of first network units can be trained based on different application scenarios and application requirements. The image processing effect of the first machine learning model can be suitable for different application scenarios and can meet different image processing requirements.

[0124] In this embodiment, by obtaining guiding feature information and then inputting the guiding feature information into the first network unit included in the first machine learning model, the first network unit can process the facial image based on the guiding feature information and multi-dimensional features, thereby further ensuring the quality and efficiency of the analysis and processing of the facial image, and improving the stability and reliability of the method.

[0125] Figure 8A schematic diagram of a flow chart for determining multi-dimensional features corresponding to a face image according to an embodiment of the present invention; based on the above embodiment, further reference is made to the attached Figure 8 As shown, this embodiment provides another implementation method for determining multi-dimensional features corresponding to a facial image. Specifically, in this embodiment, determining multi-dimensional features corresponding to a facial image may include:

[0126] Step S801: Obtain a convolution kernel for processing a facial image and a modulation function corresponding to the convolution kernel.

[0127] Step S802: Process the facial image based on the convolution kernel and the modulation function to obtain multi-dimensional features corresponding to the facial image.

[0128] Among them, the convolution kernel is used to analyze and process the facial image. It can be understood that different application scenarios or different application requirements may correspond to different convolution kernels. After obtaining the convolution kernel, the modulation function corresponding to the convolution kernel can be obtained. Specifically, this embodiment does not limit the specific implementation method of obtaining the modulation function corresponding to the convolution kernel. Those skilled in the art can set it according to specific application requirements and design requirements. For example: the correspondence between the convolution kernel and the modulation function is pre-configured, and the modulation function corresponding to the convolution kernel can be determined based on the above correspondence, etc. As long as the accuracy and reliability of the modulation function can be guaranteed, it will not be repeated here. After obtaining the convolution kernel and the modulation function, the facial image can be processed based on the convolution kernel and the modulation function, so that the multi-dimensional features corresponding to the facial image can be obtained.

[0129] In this embodiment, a convolution kernel for processing a facial image and a modulation function corresponding to the convolution kernel are obtained, and then the facial image is processed based on the convolution kernel and the modulation function to obtain multi-dimensional features corresponding to the facial image. This effectively ensures the accuracy and reliability of the acquisition of multi-dimensional features.

[0130] Figure 9 The flow chart of obtaining the modulation function corresponding to the convolution kernel provided by the embodiment of the present invention; Based on the above embodiment, continue to refer to the attached Figure 9 As shown, this embodiment provides another method for obtaining a modulation function corresponding to a convolution kernel. Specifically, obtaining a modulation function corresponding to a convolution kernel in this embodiment may include:

[0131] Step S901: Determine a first original input vector of second input information on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis.

[0132] Step S902: Determine a modulation function corresponding to a convolution kernel based on the first original input vector and the second original input vector.

[0133] The second input information may include any one of the following: a facial image, or the second output information output by the second network unit at the previous level. After obtaining the second input information, the first original input vector of the second input information on the first spatial coordinate axis and the second original input vector on the second spatial coordinate axis may be determined. Specifically, a preset coordinate system corresponding to the facial image may be determined first. The preset coordinate system may include the first spatial coordinate axis and the second spatial coordinate axis, and the first spatial coordinate axis and the second spatial coordinate axis are perpendicular to each other. After obtaining the second input information, the second input information in the preset coordinate system may be analyzed and processed, thereby obtaining the first original input vector of the second input information on the first spatial coordinate axis and the second original input vector on the second spatial coordinate axis.

[0134] The first original input vector and the second original input vector are used to identify the information characteristics of the second input information in the preset coordinate system; after obtaining the first original input vector and the second original input vector, the first original input vector and the second original input vector can be analyzed and processed to determine the modulation function corresponding to the convolution kernel. Figure 10 As shown, in this embodiment, determining the modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector may include:

[0135] Step S9021: Determine a first mapping function for mapping the first original input vector to a preset spatial coordinate axis, and a second mapping function for mapping the second original input vector to a preset spatial coordinate axis.

[0136] Step S9022: Based on the first mapping function and the second mapping function, determine a modulation function corresponding to the convolution kernel.

[0137] Specifically, after obtaining the first original input vector, the first original input vector can be mapped to a preset spatial coordinate axis, thereby obtaining a first mapping function corresponding to the first original input vector; similarly, after obtaining the second original input vector, the second original input vector can be mapped to a preset spatial coordinate axis, thereby obtaining a second mapping function corresponding to the second original input vector. After obtaining the first mapping function and the second mapping function, the modulation function corresponding to the convolution kernel can be determined based on the first mapping function and the second mapping function, thereby effectively ensuring the accuracy and reliability of obtaining the modulation function corresponding to the convolution kernel, further improving the practicality of the method.

[0138] For specific applications, refer to the attached Figure 11 As shown, this application embodiment provides an image processing method that can perform face repair processing on a facial image to be processed. Face repair (Face Renovation) refers to reconstructing a low-quality facial image (or video frame) containing complex degradation for real application scenarios, thereby obtaining a corresponding high-definition, realistic, and natural target facial image. Compared with the facial image to be processed, the target facial image has more realistic facial details, making the facial texture details (such as wrinkles, hair, etc.) more vivid and lifelike. Specifically, the method may include:

[0139] Step 1: Get the face image to be processed;

[0140] Step 2: Input the facial image into a second machine learning model to determine the multi-dimensional features corresponding to the facial image, where the multi-dimensional features may include at least two different image features corresponding to the facial image.

[0141] like Figure 11 As shown, the second machine learning model can include one or more second network units, and the one or more second network units can analyze and process the facial image, thereby determining the multi-dimensional features corresponding to the facial image (for example, key points, facial contours, texture, color). Specifically, since the number of second network units is one or more, and there are multiple second network units, and when multiple second network units are used to analyze and process the facial image, the second network unit at the first level can analyze and process the facial image, thereby obtaining the first-level output result, and then the first-level output result can be input to the second network unit at the second level, and the second network unit at the second level can analyze and process the first-level output result, thereby obtaining the second-level output result. By analogy, the second network unit at each level can analyze and process the received input information and output the corresponding output result. After the above process, the output result output by the second network unit at each level can be obtained, and then the output result output by the second network unit at each level can be determined as the multi-dimensional feature corresponding to the facial image.

[0142] Specifically, when the second network unit processes the input information (a facial image or the output result of the second network unit at the previous level), it can include: obtaining a convolution kernel and an adaptive weight modulation function, wherein the convolution kernel can be a four-dimensional floating-point matrix of a fixed size C*C`*S*S, where the above C refers to the input channel width, C` refers to the output channel width, and S is used to limit the operation range of the convolution processing.

[0143] In addition, the adaptive weight modulation function can be obtained through neural network training, and the modulation function is used to perform a nonlinear transformation on the input features (the face image, or the output result output by the second network unit of the previous level), so as to determine the multi-dimensional features corresponding to the face image. Specifically, when analyzing and processing the face image using the convolution kernel and the adaptive weight modulation function, the analysis and processing can be performed according to the following formula:

[0144]

[0145] Among them, DRAFT(F;W) i It is the result output by the second network unit at each level. It is an adaptive weight modulation function. It is understandable that different application scenarios can correspond to different fj refers to the second original input vector of the input information on the second spatial coordinate axis j, fi refers to the first original input vector of the input information on the first spatial coordinate axis i; W∈R C×C×S×S , R is the convolution kernel, F is the input information, Ω(i) is the sliding window centered on the i coordinate axis, i and j are the preset 2D spatial coordinate axes, w is the preset coefficient, Δji is the offset between coordinates i and j, which is used to index elements in w, and b is the bias vector corresponding to the convolution kernel.

[0146] It should be noted that the above It can be obtained by the following formula:

[0147]

[0148] in, refers to the adaptive weight modulation function, exp is the exponential function, It refers to the first mapping function that maps the first original input vector to the preset spatial coordinate axis. It refers to the second mapping function that maps the second original input vector to the preset spatial coordinate axis.

[0149] In some instances, before the second network unit inputs the corresponding second output result to the second network unit of the next level, the second output result can be downsampled to achieve feature screening processing on the second output result, and then the processed second output result can be input to the second network unit of the next level. This can effectively reduce the memory space occupied by the second output result and further improve the quality and efficiency of data processing by the second network unit.

[0150] Step 3: Input the multidimensional features into the first machine learning model so that the first machine learning model can analyze and process the facial image based on the multidimensional features, and determine a target image corresponding to the facial image, where the clarity of the target image is higher than that of the facial image.

[0151] Among them, for the first machine learning model, the first machine learning model can analyze and process the received facial image. Specifically, the first machine learning model can use multi-dimensional features as guiding feature information to repair the facial image, for example, it can add details to the facial image. It should be noted that the first machine learning model can include one or more first network units, each of which can process the currently received input information and input the obtained first output result to the first network unit of the next level, and iterate until the target image corresponding to the facial image is determined.

[0152] It should be noted that the image processing method provided in this application embodiment is not limited to being used for image repair processing of facial images. For example, it can be used for image repair processing of facial images with complex backgrounds, or animal portraits, etc.

[0153] In addition, when the first machine learning model includes multiple first network units and the second machine learning model includes multiple second network units, the cascade mode of the multiple first network units and the multiple second network units can be a nested structure, or the multiple first network units and the multiple second network units can adopt a series structure, a parallel structure or a series and parallel combination structure, and the above-mentioned first machine learning model and the second machine learning model can be obtained through learning and training iterations using a recurrent neural network (RNN, LSTM).

[0154] In addition, the number of first network units can be the same as or different from the number of second network units. For example, the first machine learning model is composed of five cascaded first network units, and the second machine learning model is composed of three cascaded second network units. In this case, the results output by each second network unit in the second machine learning model can be shared with the five first network units in the first machine learning model, for example: second network unit D1->first network unit S5; second network unit D2->first network unit S4; second network unit D3->first network unit S1, first network unit S2, first network unit S3. As can be seen from the above, the mapping relationship between the first network unit and the second network unit can be a one-to-one mapping relationship or a one-to-many mapping relationship, etc.

[0155] Step 3`: Obtain the guiding feature information input by the user, and input the multi-dimensional features into the first machine learning model, so that the first machine learning model can analyze and process the facial image based on the guiding feature information and the multi-dimensional features, and determine the target image corresponding to the facial image, and the clarity of the target image is higher than the clarity of the facial image.

[0156] Among them, the guiding feature information may include at least one of the following: a facial semantic map, a key point positioning map, and a heat map. It can be understood that the guiding feature information is not limited to the information exemplified above. Those skilled in the art may also include other types of feature information, which will not be repeated here.

[0157] The image processing method provided in this embodiment can adapt to the processing of any complex noise and degraded images in various real scenes. Through the cascaded first machine learning model and the second machine learning model, multi-dimensional features can be adaptively screened out, and then the facial image is processed based on the multi-dimensional feature information. This effectively ensures the quality and efficiency of acquiring high-quality target images, and reduces the difficulty of image processing, so that the image processing method can be widely applied to various application scenarios, further improving the practicality of the method.

[0158] Figure 12 A flowchart of another image processing method provided by an embodiment of the present invention; Figure 12 As shown, this embodiment provides another image processing method. The execution subject of this method can be an image processing device. It is understandable that the image processing device can be implemented as software or a combination of software and hardware. Specifically, the processing method may include:

[0159] Step S1201: Acquire an image to be processed.

[0160] Step S1202: Determine a multi-dimensional feature corresponding to the image to be processed, where the multi-dimensional feature includes at least two different image features corresponding to the image to be processed.

[0161] Step S1203: Input the multi-dimensional features and the image to be processed into the first machine learning model, so that the first machine learning model processes the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed.

[0162] In which, the first machine learning model is trained to determine a target image corresponding to the image to be processed based on multi-dimensional features, and the clarity of the target image is different from the clarity of the image to be processed.

[0163] The following is a detailed explanation of each of the above steps:

[0164] Step S1201: Acquire an image to be processed.

[0165] Among them, the image to be processed is a biological facial image that needs to be processed. It can be understood that the above-mentioned image processing may include at least one of the following: image enhancement processing, image blurring processing, image rendering processing, image editing processing, etc. Specifically, the above-mentioned image enhancement processing can increase the clarity of the display of the image to be processed, image blurring processing can reduce the clarity of the display of the image to be processed, image rendering processing can perform whitening, beautification and other rendering processing on the target in the image to be processed, and image editing processing can perform various types of editing operations on the image to be processed, such as image filtering processing, image texture processing, image cropping processing, etc.

[0166] Additionally, a biological facial image may refer to: a human face image, a cat face image, a dog face image, or a facial portrait of another biological creature, etc. The image to be processed may include at least one of the following: image information obtained by a camera, image information in video information, a composite image, etc. It is understood that the number of images to be processed may be one or more. When there are multiple images to be processed, the multiple images to be processed may constitute an image sequence, thereby enabling image processing operations to be performed on the image sequence. Furthermore, the image type of the image to be processed may be a static image or a dynamic image, thereby enabling image processing operations to be performed on static images or dynamic images.

[0167] In addition, this embodiment does not limit the specific implementation method for the image processing device to obtain the image to be processed. Those skilled in the art can configure it according to specific application requirements and design requirements. For example, the shooting device can be in communication with the enhancement device. After the shooting device captures the image to be processed, the image processing device can obtain the image to be processed through the shooting device. Specifically, the image processing device can actively obtain the image to be processed obtained by the shooting device, or the shooting device can actively send the image to be processed to the enhancement device, so that the image processing device can obtain the image to be processed. Alternatively, the image to be processed can be stored in a preset area, and the image processing device can obtain the image to be processed by accessing the preset area.

[0168] Step S1202: Determine a multi-dimensional feature corresponding to the image to be processed, where the multi-dimensional feature includes at least two different image features corresponding to the image to be processed.

[0169] Step S1203: Input the multi-dimensional features and the image to be processed into the first machine learning model, so that the first machine learning model processes the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed.

[0170] The specific implementation methods and effects of the above steps in this embodiment are the same as those in the above Figure 1 The specific implementation methods and effects of steps S102-S103 in the embodiment are similar, and the details can be referred to the above descriptions, which will not be repeated here. Figure 1 The difference between the embodiments is that, in this embodiment, a face image to be processed is used as an example of an image to be processed to implement the image processing method in this embodiment.

[0171] In some instances, determining the multidimensional features corresponding to the image to be processed may include: using a second machine learning model to analyze and process the image to be processed to determine the multidimensional features corresponding to the image to be processed, and the second machine learning model is trained to determine the multidimensional features corresponding to the image to be processed.

[0172] In some instances, the second machine learning model includes: one or more second network units, multiple second network units are connected in series in sequence, and the second network units are used to analyze and process the received second input information to determine second output information corresponding to the second input information.

[0173] In some examples, the second input information includes any one of the following: an image to be processed, or second output information output by a second network unit at an upper level.

[0174] In some instances, analyzing and processing the image to be processed using a second machine learning model to determine the multidimensional features corresponding to the image to be processed may include: obtaining one or more second output information output by one or more second network units when analyzing and processing the image to be processed using the second machine learning model; and determining the one or more second output information as the multidimensional features corresponding to the image to be processed.

[0175] In some instances, the first machine learning model includes: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine first output information corresponding to the first input information.

[0176] In some instances, the first input information includes any one of the following: multi-dimensional features corresponding to the image to be processed, guiding feature information, and first output information output by the first network unit at the previous level.

[0177] In some examples, the guidance feature information includes at least one of the following: a semantic map, a key point location map, and a heat map.

[0178] In some examples, the number of the first network unit and the number of the second network unit are the same or different.

[0179] In some instances, the method in this embodiment may further include: obtaining guiding feature information; inputting the guiding feature information into a first network unit included in a first machine learning model, so that the first network unit processes the image to be processed based on the guiding feature information and multi-dimensional features to obtain a target image corresponding to the image to be processed.

[0180] In some examples, obtaining a target image corresponding to the image to be processed may include: in one or more first network units, determining first output information output by a last-level first network unit as the target image corresponding to the image to be processed.

[0181] In some examples, the multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

[0182] In some instances, determining the multidimensional features corresponding to the image to be processed may include: obtaining a convolution kernel for processing the image to be processed and a modulation function corresponding to the convolution kernel; processing the image to be processed based on the convolution kernel and the modulation function to obtain the multidimensional features corresponding to the image to be processed.

[0183] In some instances, obtaining a modulation function corresponding to a convolution kernel may include: determining a first original input vector of the image to be processed on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis; and determining a modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector.

[0184] In some instances, determining a modulation function corresponding to a convolution kernel based on a first original input vector and a second original input vector may include: determining a first mapping function that maps the first original input vector to a preset spatial coordinate axis, and a second mapping function that maps the second original input vector to a preset spatial coordinate axis; and determining a modulation function corresponding to the convolution kernel based on the first mapping function and the second mapping function.

[0185] In some instances, obtaining a target image corresponding to an image to be processed may include: obtaining an area to be processed corresponding to the image to be processed, covering the area to be processed with a preset mosaic configuration, generating a mosaic image, and determining the mosaic image as the target image corresponding to the image to be processed.

[0186] Among them, different application scenarios may correspond to different images to be processed. Specifically, the images to be processed may be game interface images, face images, text images to be reviewed, etc.; in order to ensure the security and reliability of data display and avoid data leakage, the relevant parts of the image to be processed can be mosaic processed, that is, a target image with a mosaic effect is generated.

[0187] For example: when the image to be processed is a face image, in order to avoid the leakage of face information, the face display area corresponding to the image to be processed can be determined, and then the face display area can be covered with a preset mosaic, thereby generating a target image with a mosaic effect. Alternatively, when the image to be processed is a game interface image, in order to avoid the leakage of game-related information (account information, password information, etc.) and to ensure the security and reliability of game-related information, the game-related information area corresponding to the image to be processed can be determined, and then the game-related information area can be covered with a preset mosaic, thereby generating a target image with a mosaic effect. Alternatively, when the image to be processed is a text image to be reviewed, in order to avoid the leakage of text information, the text display area corresponding to the text image to be reviewed can be determined, and then the whole or part of the text display area can be covered with a preset mosaic, thereby generating a target image with a mosaic effect.

[0188] In this embodiment, by obtaining the area to be processed corresponding to the image to be processed, the area to be processed is covered with a preset mosaic to generate a target image with a mosaic effect, thereby effectively ensuring the flexibility and reliability of processing the target image and further improving the stability and reliability of the method.

[0189] In some instances, when determining the multidimensional features corresponding to the image to be processed, it can include: obtaining configuration rules corresponding to the image to be processed, and determining the multidimensional features corresponding to the image to be processed based on the configuration rules, so as to input the multidimensional features and the facial image into the first machine learning model, so that the first machine learning model processes the facial image based on the multidimensional features to obtain a target image corresponding to the facial image.

[0190] Specifically, different application scenarios may correspond to different multi-dimensional features. Therefore, after obtaining the image to be processed, in order to improve the quality and efficiency of analyzing and processing the image to be processed, the configuration rules corresponding to the image to be processed can be obtained (used to determine the multi-dimensional features corresponding to the image to be processed). Specifically, multiple configuration rules are pre-configured, and then the mapping relationship between the image to be processed and the configuration rules can be obtained, and the configuration rules corresponding to the image to be processed can be determined based on the above mapping relationship; or, the image to be processed can be analyzed and processed to determine the image category corresponding to the image to be processed (person image, data type image, etc.), and the configuration rules corresponding to the image to be processed are determined based on the image category.

[0191] After obtaining the configuration rules, the multidimensional features corresponding to the facial image can be determined based on the determined configuration rules, and then the obtained multidimensional features and the image to be processed are input into the first machine learning model, so that the first machine learning model processes the image to be processed based on the multidimensional features and obtains the target image corresponding to the image to be processed. This effectively ensures the quality and efficiency of processing the image to be processed and further improves the stability and reliability of image processing.

[0192] The execution process and technical effects of the above method in this embodiment are similar to those in Figures 1-11 The execution process and technical effects of the method in the illustrated embodiment are similar, and for details, please refer to the above statements, which will not be repeated here.

[0193] Figure 13 A flow chart of a model training method provided by an embodiment of the present invention; see the attached Figure 13 As shown, this embodiment provides a model training method. The execution subject of this method can be a model training device. It is understandable that the model training device can be implemented as software or a combination of software and hardware. Specifically, the method may include:

[0194] Step S1301: Acquire a first image and a reference image corresponding to the first image, wherein the definition of the reference image is different from that of the first image.

[0195] Step S1302: Determine a multi-dimensional feature corresponding to the first image, where the multi-dimensional feature includes at least two different image features corresponding to the facial image.

[0196] Step S1303: Perform learning and training based on the first image, the reference image and the multi-dimensional features to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multi-dimensional features. The clarity of the target image is different from that of the first image.

[0197] The first image and the reference image are the same image with different sharpness. In a specific implementation, the sharpness of the reference image can be higher than that of the first image, or lower than that of the first image. The first image and the reference image can be stored in a preset area, and the first image and the reference image can be obtained by accessing the preset area. In a specific application, the multiple first images can be multiple preset blurred images, and the above-mentioned first images can include at least one of the following: image information obtained by a camera, image information in video information, a composite image, etc. This embodiment does not limit the specific implementation method for the training device to obtain the first image. Those skilled in the art can configure it according to specific application and design requirements. For example, the camera can be in communication with the training device. After the camera captures the first image, the training device can obtain the first image through the camera. Specifically, the training device can actively obtain the first image obtained by the camera, or the camera can actively send the first image to the training device, so that the training device obtains the first image. Alternatively, the first image can be stored in a preset area, and the training device can obtain the first image by accessing the preset area.

[0198] After acquiring the first image, the first image can be analyzed and processed to obtain a multidimensional feature corresponding to the first image. The multidimensional feature can include at least two different image features corresponding to the facial image. After acquiring the multidimensional feature, learning and training can be performed based on the multidimensional feature, the reference image, and the first image. Specifically, a spatially adaptive convolutional residual network can be learned and trained based on the multidimensional feature, the reference image, and the first image to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image, where the clarity of the target image is different from that of the first image.

[0199] The model training method provided in this embodiment obtains a first image and a reference image corresponding to the first image; determines the multidimensional features corresponding to the first image, and performs learning and training based on the multidimensional features, the reference image and the first image, so as to obtain a first machine learning model suitable for processing images in all application scenarios. The first machine learning model can determine the target image corresponding to the first image, and realize image analysis and processing based on the generated first machine learning model, thereby effectively ensuring the scope of application of the first machine learning model and improving the practicality of the model training method.

[0200] In some examples, the multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

[0201] In some instances, the method in this embodiment may further include: obtaining guiding feature information corresponding to the first image; performing learning and training based on the first image, the reference image, the multidimensional features and the guiding feature information to obtain a first machine learning model, the first machine learning model being used to determine a target image corresponding to the first image based on the multidimensional features, the clarity of the target image being different from the clarity of the first image.

[0202] In some instances, the guided feature information includes at least one of the following: a face semantic map, a key point location map, a heat map, and output feature information obtained by processing the first image.

[0203] The specific execution process and technical effects of the above steps in this embodiment are similar to the specific execution process and technical effects of obtaining the first machine learning model by learning and training based on the first image, the reference image and the multi-dimensional features in the above embodiment. Please refer to the above statements for details and will not be repeated here.

[0204] Figure 14 A schematic diagram of the structure of an image processing device provided by an embodiment of the present invention; Figure 14 As shown, this embodiment provides an image processing device that can perform the above Figure 1 The corresponding image processing method, the image processing device may include a first acquisition module 11, a first determination module 12 and a first processing module 13; specifically,

[0205] A first acquisition module 11 is used to acquire a face image to be processed;

[0206] A first determining module 12 is configured to determine a multi-dimensional feature corresponding to the facial image, where the multi-dimensional feature includes at least two different image features corresponding to the facial image;

[0207] A first processing module 13 is configured to input the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image;

[0208] In which, the first machine learning model is trained to determine a target image corresponding to a facial image based on multi-dimensional features, and the clarity of the target image is different from the clarity of the facial image.

[0209] In some instances, when the first determination module 12 determines the multidimensional features corresponding to the facial image, the first determination module 12 can be used to perform: analyzing and processing the facial image using a second machine learning model to determine the multidimensional features corresponding to the facial image, and the second machine learning model is trained to determine the multidimensional features corresponding to the facial image.

[0210] In some instances, the second machine learning model includes: one or more second network units, multiple second network units are connected in series in sequence, and the second network units are used to analyze and process the received second input information to determine second output information corresponding to the second input information.

[0211] In some instances, the second input information includes any one of the following: a face image, or second output information output by a second network unit at an upper level.

[0212] In some instances, when the first determination module 12 uses the second machine learning model to analyze and process the facial image to determine the multidimensional features corresponding to the facial image, the first determination module 12 can be used to perform: when using the second machine learning model to analyze and process the facial image, obtain one or more second output information output by one or more second network units; determine the one or more second output information as multidimensional features corresponding to the facial image.

[0213] In some instances, the first machine learning model includes: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine first output information corresponding to the first input information.

[0214] In some instances, the first input information includes any one of the following: multi-dimensional features corresponding to the facial image, guiding feature information, and first output information output by the first network unit at the previous level.

[0215] In some instances, the guiding feature information includes at least one of the following: a face semantic map, a key point location map, and a heat map.

[0216] In some examples, the number of the first network unit and the number of the second network unit are the same or different.

[0217] In some examples, the first acquisition module 11 and the first processing module 13 in this embodiment can be used to perform the following steps:

[0218] A first acquisition module 11 is used to acquire guidance feature information;

[0219] The first processing module 13 is used to input the guiding feature information into the first network unit included in the first machine learning model, so that the first network unit processes the facial image based on the guiding feature information and the multi-dimensional features to obtain a target image corresponding to the facial image.

[0220] In some instances, when the first processing module 13 obtains a target image corresponding to a facial image, the first processing module 13 can be used to execute: in one or more first network units, determining the first output information output by the last-level first network unit as the target image corresponding to the facial image.

[0221] In some examples, the multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

[0222] In some instances, when the first determination module 12 determines the multidimensional features corresponding to the facial image, the first determination module 12 can be used to execute: obtaining a convolution kernel for processing the facial image and a modulation function corresponding to the convolution kernel; processing the facial image based on the convolution kernel and the modulation function to obtain the multidimensional features corresponding to the facial image.

[0223] In some instances, when the first determination module 12 obtains the modulation function corresponding to the convolution kernel, the first determination module 12 can be used to perform: determining a first original input vector of the second input information on the first spatial coordinate axis and a second original input vector on the second spatial coordinate axis; based on the first original input vector and the second original input vector, determining the modulation function corresponding to the convolution kernel.

[0224] In some instances, when the first determination module 12 determines the modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector, the first determination module 12 can be used to perform: determining a first mapping function that maps the first original input vector to a preset spatial coordinate axis, and a second mapping function that maps the second original input vector to a preset spatial coordinate axis; based on the first mapping function and the second mapping function, determining the modulation function corresponding to the convolution kernel.

[0225] Figure 14 The device shown can perform Figures 1-11 For the method of the embodiment shown in FIG. 1 , reference may be made to the description of the part not described in detail in the embodiment. Figures 1-11 The implementation process and technical effects of this technical solution can be found in Figures 1-11 The description in the illustrated embodiment will not be repeated here.

[0226] In one possible design, Figure 14 The structure of the image processing device shown can be implemented as an electronic device, which can be a mobile phone, tablet computer, server and other devices. Figure 15 As shown, the electronic device may include: a first processor 21 and a first memory 22. The first memory 22 is used to store the corresponding electronic device to perform the above Figures 1-11The program of the image processing method provided in the illustrated embodiment is configured so that the first processor 21 is configured to execute the program stored in the first memory 22 .

[0227] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the first processor 21, the following steps can be implemented:

[0228] Obtain the face image to be processed;

[0229] Determining a multi-dimensional feature corresponding to the facial image, the multi-dimensional feature including at least two different image features corresponding to the facial image;

[0230] Inputting the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model processes the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image;

[0231] In which, the first machine learning model is trained to determine a target image corresponding to a facial image based on multi-dimensional features, and the clarity of the target image is different from the clarity of the facial image.

[0232] Furthermore, the first processor 21 is also used to execute the aforementioned Figures 1-11 All or part of the steps in the illustrated embodiments.

[0233] The structure of the electronic device may further include a first communication interface 23 for the electronic device to communicate with other devices or a communication network.

[0234] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by electronic devices, which includes instructions for executing the above Figures 1-11 The procedures involved in the image processing method in the method embodiment shown.

[0235] Figure 16 A schematic diagram of another image processing device according to an embodiment of the present invention; Figure 16 As shown, this embodiment provides another image processing device, which can perform the above Figure 12 The corresponding image processing method, the image processing device may include a second acquisition module 31, a second determination module 32 and a second processing module 33; specifically,

[0236] A second acquisition module 31 is used to acquire an image to be processed;

[0237] A second determining module 32 is configured to determine a multi-dimensional feature corresponding to the image to be processed, where the multi-dimensional feature includes at least two different image features corresponding to the image to be processed;

[0238] A second processing module 33 is configured to input the multi-dimensional features and the image to be processed into the first machine learning model, so that the first machine learning model processes the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed;

[0239] In which, the first machine learning model is trained to determine a target image corresponding to the image to be processed based on multi-dimensional features, and the clarity of the target image is different from the clarity of the image to be processed.

[0240] In some instances, when the second determination module 32 determines the multidimensional features corresponding to the image to be processed, the second determination module 32 can be used to perform: using a second machine learning model to analyze and process the image to be processed to determine the multidimensional features corresponding to the image to be processed, and the second machine learning model is trained to determine the multidimensional features corresponding to the image to be processed.

[0241] In some instances, the second machine learning model includes: one or more second network units, multiple second network units are connected in series in sequence, and the second network units are used to analyze and process the received second input information to determine second output information corresponding to the second input information.

[0242] In some examples, the second input information includes any one of the following: an image to be processed, or second output information output by a second network unit at an upper level.

[0243] In some instances, when the second determination module 32 uses the second machine learning model to analyze and process the image to be processed to determine the multidimensional features corresponding to the image to be processed, the second determination module 32 can be used to perform: when using the second machine learning model to analyze and process the image to be processed, obtain one or more second output information output by one or more second network units; determine the one or more second output information as the multidimensional features corresponding to the image to be processed.

[0244] In some instances, the first machine learning model includes: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine first output information corresponding to the first input information.

[0245] In some instances, the first input information includes any one of the following: multi-dimensional features corresponding to the image to be processed, guiding feature information, and first output information output by the first network unit at the previous level.

[0246] In some examples, the guidance feature information includes at least one of the following: a semantic map, a key point location map, and a heat map.

[0247] In some examples, the number of the first network unit and the number of the second network unit are the same or different.

[0248] In some examples, the second acquisition module 31 and the second processing module 33 in this embodiment can be used to perform the following steps:

[0249] The second acquisition module 31 is used to obtain guidance feature information;

[0250] The second processing module 33 is used to input the guiding feature information into the first network unit included in the first machine learning model, so that the first network unit processes the image to be processed based on the guiding feature information and the multi-dimensional features to obtain a target image corresponding to the image to be processed.

[0251] In some instances, when the second processing module 33 obtains a target image corresponding to the image to be processed, the second processing module 33 can be used to execute: in one or more first network units, determining the first output information output by the last-level first network unit as the target image corresponding to the image to be processed.

[0252] In some examples, the multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

[0253] In some instances, when the second determination module 32 determines the multidimensional features corresponding to the image to be processed, the second determination module 32 can be used to perform: obtaining a convolution kernel for processing the image to be processed and a modulation function corresponding to the convolution kernel; processing the image to be processed based on the convolution kernel and the modulation function to obtain the multidimensional features corresponding to the image to be processed.

[0254] In some instances, when the second determination module 32 obtains the modulation function corresponding to the convolution kernel, the second determination module 32 can be used to perform: determining a first original input vector of the image to be processed on the first spatial coordinate axis and a second original input vector on the second spatial coordinate axis; based on the first original input vector and the second original input vector, determining the modulation function corresponding to the convolution kernel.

[0255] In some instances, when the second determination module 32 determines the modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector, the second determination module 32 can be used to perform: determining a first mapping function that maps the first original input vector to a preset spatial coordinate axis, and a second mapping function that maps the second original input vector to a preset spatial coordinate axis; based on the first mapping function and the second mapping function, determining the modulation function corresponding to the convolution kernel.

[0256] Figure 16 The device shown can perform Figure 12 For the method of the embodiment shown in FIG. 1 , reference may be made to the description of the part not described in detail in the embodiment. Figure 12 The implementation process and technical effects of this technical solution can be found in Figure 12 The description in the illustrated embodiment will not be repeated here.

[0257] In one possible design, Figure 16 The structure of the image processing device shown can be implemented as an electronic device, which can be a mobile phone, tablet computer, server and other devices. Figure 17 As shown, the electronic device may include: a second processor 41 and a second memory 42. The second memory 42 is used to store the corresponding electronic device to perform the above Figure 12 The program of the image processing method provided in the illustrated embodiment is configured so that the second processor 41 is configured to execute the program stored in the second memory 42 .

[0258] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the second processor 41, the following steps can be implemented:

[0259] Get the image to be processed;

[0260] Determining a multi-dimensional feature corresponding to the image to be processed, where the multi-dimensional feature includes at least two different image features corresponding to the image to be processed;

[0261] Inputting the multi-dimensional features and the image to be processed into a first machine learning model, so that the first machine learning model processes the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed;

[0262] In which, the first machine learning model is trained to determine a target image corresponding to the image to be processed based on multi-dimensional features, and the clarity of the target image is different from the clarity of the image to be processed.

[0263] Furthermore, the second processor 41 is also used to execute the aforementioned Figure 12 All or part of the steps in the illustrated embodiments.

[0264] The structure of the electronic device may further include a second communication interface 43 for the electronic device to communicate with other devices or a communication network.

[0265] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by electronic devices, which includes instructions for executing the above Figure 12 The procedures involved in the image processing method in the method embodiment shown.

[0266] Figure 18A schematic diagram of the structure of a model training device provided by an embodiment of the present invention; Figure 18 As shown, this embodiment provides a model training device, which can perform the above Figure 13 The corresponding model training method, the model training device may include a third acquisition module 51, a third determination module 52 and a third training module 53; specifically,

[0267] a third acquisition module 51, configured to acquire a first image and a reference image corresponding to the first image, wherein the clarity of the reference image is different from that of the first image;

[0268] A third determining module 52 is configured to determine a multi-dimensional feature corresponding to the first image, where the multi-dimensional feature includes at least two different image features corresponding to the facial image;

[0269] The third processing module 53 is used to perform learning and training based on the first image, the reference image and the multi-dimensional features to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the first image.

[0270] In some examples, the multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

[0271] In some examples, the third acquisition module 51 and the third processing module 53 in this embodiment can be used to perform the following steps:

[0272] A third acquisition module 51 is used to acquire guidance feature information corresponding to the first image;

[0273] The third processing module 53 is used to perform learning and training based on the first image, the reference image, the multi-dimensional features and the guiding feature information to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multi-dimensional features. The clarity of the target image is different from the clarity of the first image.

[0274] In some instances, the guided feature information includes at least one of the following: a face semantic map, a key point location map, a heat map, and output feature information obtained by processing the first image.

[0275] Figure 18 The device shown can perform Figure 13 For the method of the embodiment shown in FIG. 1 , reference may be made to the description of the part not described in detail in the embodiment. Figure 13 The implementation process and technical effects of this technical solution can be found in Figure 13 The description in the illustrated embodiment will not be repeated here.

[0276] In one possible design, Figure 18 The structure of the model training device shown can be implemented as an electronic device, which can be a mobile phone, tablet computer, server and other devices. Figure 15 As shown, the electronic device may include: a third processor 61 and a third memory 62. The third memory 62 is used to store the corresponding electronic device to perform the above Figure 13 The program of the model training method provided in the illustrated embodiment, the third processor 61 is configured to execute the program stored in the third memory 62.

[0277] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the third processor 61, the following steps can be implemented:

[0278] Acquire a first image and a reference image corresponding to the first image, wherein the definition of the reference image is different from the definition of the first image;

[0279] determining a multi-dimensional feature corresponding to the first image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image;

[0280] Learning and training are performed based on the first image, the reference image and the multi-dimensional features to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multi-dimensional features, and the clarity of the target image is different from the clarity of the first image.

[0281] Furthermore, the third processor 61 is also used to execute the aforementioned Figure 13 All or part of the steps in the illustrated embodiments.

[0282] The structure of the electronic device may further include a third communication interface 63 for the electronic device to communicate with other devices or a communication network.

[0283] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by electronic devices, which includes instructions for executing the above Figure 13 The procedures involved in the model training method in the method embodiment shown.

[0284] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0285] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by a combination of hardware and software. Based on this understanding, the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0286] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0287] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0288] These computer program instructions can also be loaded onto a computer or other programmable device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0289] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0290] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0291] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0292] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image processing method, characterized in that: include: Obtain the face image to be processed; Determining a multi-dimensional feature corresponding to the facial image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image; Inputting the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model performs image blurring processing on the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image; The first machine learning model is trained to determine a target image corresponding to the facial image based on the multi-dimensional features, and the clarity of the target image is lower than that of the facial image; Determining the multi-dimensional features corresponding to the facial image includes: Obtaining a convolution kernel for processing a facial image and a modulation function corresponding to the convolution kernel, wherein the convolution kernel is a four-dimensional floating-point matrix having a fixed size of input channel width * output channel width * a computing range for limiting convolution processing * a computing range for limiting convolution processing; the modulation function is determined based on a first original input vector on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis of second input information, wherein the second input information includes the facial image; The facial image is processed based on the convolution kernel and the modulation function to obtain multi-dimensional features corresponding to the facial image.

2. The method according to claim 1, characterized in that Determining a multi-dimensional feature corresponding to the facial image includes: The facial image is analyzed and processed using a second machine learning model to determine multidimensional features corresponding to the facial image, where the second machine learning model is trained to determine multidimensional features corresponding to the facial image.

3. The method according to claim 2, characterized in that The second machine learning model includes: one or more second network units, multiple second network units are connected in series in sequence, and the second network unit is used to analyze and process the received second input information to determine second output information corresponding to the second input information.

4. The method according to claim 3, characterized in that The second input information includes: second output information output by the second network unit at an upper level.

5. The method according to claim 3, characterized in that Analyzing and processing the facial image using a second machine learning model to determine multi-dimensional features corresponding to the facial image includes: When analyzing and processing the facial image using the second machine learning model, obtaining one or more second output information output by the one or more second network units; The one or more second output information are determined as multi-dimensional features corresponding to the facial image.

6. The method according to any one of claims 3 to 5, characterized in that The first machine learning model includes: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine the first output information corresponding to the first input information.

7. The method according to claim 6, characterized in that The first input information includes any one of the following: multi-dimensional features corresponding to the facial image, guiding feature information, and first output information output by the first network unit at the previous level.

8. The method according to claim 7, characterized in that The guiding feature information includes at least one of the following: a face semantic map, a key point positioning map, and a heat map.

9. The method according to claim 6, characterized in that The number of the first network units and the number of the second network units are the same or different.

10. The method according to claim 7, characterized in that The method further comprises: obtaining the guidance feature information; The guiding feature information is input into the first network unit included in the first machine learning model, so that the first network unit processes the facial image based on the guiding feature information and multi-dimensional features to obtain a target image corresponding to the facial image.

11. The method according to claim 7, characterized in that Obtaining a target image corresponding to the face image, comprising: In one or more first network units, the first output information output by the last-level first network unit is determined as a target image corresponding to the facial image.

12. The method according to any one of claims 1 to 5, characterized in that The multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

13. The method according to claim 1, wherein Obtaining a modulation function corresponding to the convolution kernel, including: Determine a first original input vector of the second input information on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis; A modulation function corresponding to the convolution kernel is determined based on the first original input vector and the second original input vector.

14. The method according to claim 13, characterized in that Determining a modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector includes: Determining a first mapping function for mapping the first original input vector to a preset spatial coordinate axis, and a second mapping function for mapping the second original input vector to a preset spatial coordinate axis; Based on the first mapping function and the second mapping function, a modulation function corresponding to the convolution kernel is determined.

15. An image processing method, characterized in that: include: Get the image to be processed; Determining a multi-dimensional feature corresponding to the image to be processed, wherein the multi-dimensional feature includes at least two different image features corresponding to the image to be processed; Inputting the multidimensional features and the image to be processed into a first machine learning model, so that the first machine learning model performs image blurring processing on the image to be processed based on the multidimensional features to obtain a target image corresponding to the image to be processed; The first machine learning model is trained to determine a target image corresponding to the image to be processed based on the multi-dimensional features, and the clarity of the target image is lower than that of the image to be processed; Determining the multi-dimensional features corresponding to the image to be processed includes: Obtaining a convolution kernel for processing the image to be processed and a modulation function corresponding to the convolution kernel, wherein the convolution kernel is a four-dimensional floating-point matrix with a fixed size of input channel width * output channel width * a computing range for limiting convolution processing * a computing range for limiting convolution processing, and the modulation function is determined based on a first original input vector on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis of second input information, wherein the second input information includes the image to be processed; The image to be processed is processed based on the convolution kernel and the modulation function to obtain multi-dimensional features corresponding to the image to be processed.

16. The method according to claim 15, characterized in that Determining a multi-dimensional feature corresponding to the image to be processed includes: The image to be processed is analyzed and processed using a second machine learning model to determine multidimensional features corresponding to the image to be processed, and the second machine learning model is trained to determine multidimensional features corresponding to the image to be processed.

17. The method according to claim 16, characterized in that The second machine learning model includes: one or more second network units, multiple second network units are connected in series in sequence, and the second network unit is used to analyze and process the received second input information to determine second output information corresponding to the second input information.

18. The method according to claim 17, characterized in that The second input information includes: second output information output by the second network unit at an upper level.

19. The method according to claim 17, wherein Analyzing and processing the image to be processed using a second machine learning model to determine multi-dimensional features corresponding to the image to be processed, including: When analyzing and processing the image to be processed using the second machine learning model, obtaining one or more second output information output by the one or more second network units; The one or more second output information are determined as multi-dimensional features corresponding to the image to be processed.

20. The method according to any one of claims 17 to 19, characterized in that The first machine learning model includes: one or more first network units, multiple first network units are connected in series in sequence, and the first network unit is used to analyze and process the received first input information to determine the first output information corresponding to the first input information.

21. The method according to claim 20, characterized in that The first input information includes any one of the following: multi-dimensional features corresponding to the image to be processed, guiding feature information, and first output information output by the first network unit at the previous level.

22. The method according to claim 21, characterized in that The guiding feature information includes at least one of the following: a semantic map, a key point positioning map, and a heat map.

23. The method according to claim 20, characterized in that The number of the first network units and the number of the second network units are the same or different.

24. The method according to claim 21, characterized in that The method further comprises: obtaining the guidance feature information; The guiding feature information is input into the first network unit included in the first machine learning model, so that the first network unit processes the image to be processed based on the guiding feature information and multi-dimensional features to obtain a target image corresponding to the image to be processed.

25. The method according to claim 21, characterized in that Obtaining a target image corresponding to the image to be processed, comprising: In one or more first network units, the first output information output by the last-level first network unit is determined as a target image corresponding to the image to be processed.

26. The method according to any one of claims 15 to 19, characterized in that The multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

27. The method according to claim 15, wherein Obtaining a modulation function corresponding to the convolution kernel, including: Determine a first original input vector of the image to be processed on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis; A modulation function corresponding to the convolution kernel is determined based on the first original input vector and the second original input vector.

28. The method according to claim 27, characterized in that Determining a modulation function corresponding to the convolution kernel based on the first original input vector and the second original input vector includes: Determining a first mapping function for mapping the first original input vector to a preset spatial coordinate axis, and a second mapping function for mapping the second original input vector to a preset spatial coordinate axis; Based on the first mapping function and the second mapping function, a modulation function corresponding to the convolution kernel is determined.

29. A model training method, characterized in that: include: Acquire a first image and a reference image corresponding to the first image, wherein the definition of the reference image is different from the definition of the first image; determining a multi-dimensional feature corresponding to the first image, the multi-dimensional feature comprising at least two different image features corresponding to the first image; Performing learning and training based on the first image, a reference image, and the multi-dimensional features to obtain a first machine learning model, wherein the first machine learning model is used to determine a target image corresponding to the first image based on the multi-dimensional features, where the clarity of the target image is lower than that of the first image; Determining the multi-dimensional features corresponding to the first image includes: Obtaining a convolution kernel for processing a first image and a modulation function corresponding to the convolution kernel, wherein the convolution kernel is a four-dimensional floating-point matrix having a fixed size of input channel width * output channel width * a computing range for limiting convolution processing * a computing range for limiting convolution processing; the modulation function is determined based on a first original input vector on a first spatial coordinate axis and a second original input vector on a second spatial coordinate axis of second input information, wherein the second input information includes the first image; The first image is processed based on the convolution kernel and the modulation function to obtain multi-dimensional features corresponding to the first image.

30. The method according to claim 29, wherein The multi-dimensional features include at least two of the following: key point features, contour features, texture features, and color features.

31. The method according to claim 29, wherein The method further comprises: Acquiring guidance feature information corresponding to the first image; Learning and training are performed based on the first image, the reference image, the multidimensional features and the guiding feature information to obtain a first machine learning model. The first machine learning model is used to determine a target image corresponding to the first image based on the multidimensional features, and the clarity of the target image is lower than the clarity of the first image.

32. The method according to claim 31, characterized in that The guiding feature information includes at least one of the following: a face semantic map, a key point positioning map, a heat map, and output feature information obtained by processing the first image.

33. An image processing device, characterized in that include: A first acquisition module is used to acquire a face image to be processed; a first determining module, configured to determine a multi-dimensional feature corresponding to the facial image, the multi-dimensional feature comprising at least two different image features corresponding to the facial image; a first processing module, configured to input the multi-dimensional features and the facial image into a first machine learning model, so that the first machine learning model performs image blurring processing on the facial image based on the multi-dimensional features to obtain a target image corresponding to the facial image; The first machine learning model is trained to determine a target image corresponding to the facial image based on the multi-dimensional features, and the clarity of the target image is lower than that of the facial image; The first determining module is further configured to: obtain a convolution kernel for processing a facial image and a modulation function corresponding to the convolution kernel, and process the facial image based on the convolution kernel and the modulation function to obtain a multi-dimensional feature corresponding to the facial image; In which, the convolution kernel is a four-dimensional floating-point matrix with a fixed size of input channel width * output channel width * used to limit the operation range of convolution processing * used to limit the operation range of convolution processing, and the modulation function is determined based on the first original input vector of the second input information on the first spatial coordinate axis and the second original input vector on the second spatial coordinate axis, and the second input information includes the facial image.

34. An electronic device, characterized in that: include: A memory, a processor; wherein the memory is used to store one or more computer instructions, wherein when the one or more computer instructions are executed by the processor, the image processing method according to any one of claims 1 to 14 is implemented.

35. An image processing device, characterized in that include: A second acquisition module is used to acquire the image to be processed; a second determining module, configured to determine a multi-dimensional feature corresponding to the image to be processed, wherein the multi-dimensional feature includes at least two different image features corresponding to the image to be processed; a second processing module, configured to input the multi-dimensional features and the image to be processed into a first machine learning model, so that the first machine learning model performs image blurring processing on the image to be processed based on the multi-dimensional features to obtain a target image corresponding to the image to be processed; The first machine learning model is trained to determine a target image corresponding to the image to be processed based on the multi-dimensional features, and the clarity of the target image is lower than that of the image to be processed; The second determining module is further configured to: obtain a convolution kernel for processing the image to be processed and a modulation function corresponding to the convolution kernel, and process the image to be processed based on the convolution kernel and the modulation function to obtain a multi-dimensional feature corresponding to the image to be processed; In which, the convolution kernel is a four-dimensional floating-point matrix with a fixed size of input channel width * output channel width * used to limit the operation range of convolution processing * used to limit the operation range of convolution processing, and the modulation function is determined based on the first original input vector of the second input information on the first spatial coordinate axis and the second original input vector on the second spatial coordinate axis, and the second input information includes the image to be processed.

36. An electronic device, characterized in that: include: A memory, a processor; wherein the memory is used to store one or more computer instructions, wherein when the one or more computer instructions are executed by the processor, the image processing method as described in any one of claims 15 to 28 is implemented.

37. A model training device, characterized in that: include: a third acquisition module, configured to acquire a first image and a reference image corresponding to the first image, wherein the clarity of the reference image is different from that of the first image; a third determining module, configured to determine a multi-dimensional feature corresponding to the first image, the multi-dimensional feature comprising at least two different image features corresponding to the first image; a third processing module, configured to perform learning and training based on the first image, a reference image, and the multi-dimensional features to obtain a first machine learning model, wherein the first machine learning model is configured to determine a target image corresponding to the first image based on the multi-dimensional features, where the clarity of the target image is lower than that of the first image; The third determining module is further configured to: obtain a convolution kernel for processing the first image and a modulation function corresponding to the convolution kernel, and process the first image based on the convolution kernel and the modulation function to obtain a multi-dimensional feature corresponding to the first image; In which, the convolution kernel is a four-dimensional floating-point matrix with a fixed size of input channel width * output channel width * used to limit the operation range of convolution processing * used to limit the operation range of convolution processing, and the modulation function is determined based on the first original input vector of the second input information on the first spatial coordinate axis and the second original input vector on the second spatial coordinate axis, and the second input information includes the first image.

38. An electronic device, characterized in that: include: A memory and a processor; wherein the memory is used to store one or more computer instructions, wherein when the one or more computer instructions are executed by the processor, the model training method as described in any one of claims 29 to 32 is implemented.

Citation Information

Patent Citations

  • Image processing method and device and storage medium

    CN108921782A

  • A method and apparatus for generating a human face key point detection model

    CN109214343A