Image processing method and apparatus, computing device, storage medium, and program product

By extracting the set of image feature points and determining the location mapping to update pixel values, the problem of large computational load and time consumption in existing technologies for face replacement is solved, achieving efficient and natural face replacement effects and meeting real-time requirements.

CN117152805BActive Publication Date: 2026-05-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-05-20
Publication Date
2026-05-26

Smart Images

  • Figure CN117152805B_ABST
    Figure CN117152805B_ABST
Patent Text Reader

Abstract

This disclosure relates to an image processing method, comprising: extracting a first set of feature points representing the first facial region from a first image containing a first facial region; extracting a second set of feature points representing the second facial region from a second image containing a second facial region, wherein the second set of feature points corresponds to the first set of feature points; determining a positional mapping from a set of first pixels in the first facial region to a set of second pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points; determining a target region in the second facial region based on the positional mapping and the positions of at least a portion of the first pixels in the first facial region; and updating the pixel values ​​of at least a portion of the second pixels in the target region based on the pixel values ​​of at least a portion of the first pixels. This method facilitates convenient and efficient face replacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to an image processing method, an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] The development of computer technology has brought numerous changes to people's work, study, and life, thereby giving rise to many new demands. Currently, in many application scenarios, there is a need to replace the faces of people, animals, cartoon characters, etc., that is, to replace at least a portion of the facial region in one image with at least a portion of the facial region in another image. For example, in fields such as live video streaming or online chat, face replacement between images or videos of people or people and cartoon characters can enhance entertainment; in fields such as online shopping, face replacement between images or videos of people or people and virtual characters can achieve functions such as virtual dress-up; in fields such as film and television production, video editing, or image editing, face replacement can meet certain special effects requirements; and so on. However, current methods for implementing face replacement often rely on deep learning network structures such as encoders and decoders. These structures require training on a large amount of image data before being used for face replacement processing. The training process is computationally intensive, time-consuming, and labor-intensive, making it difficult to meet the needs of many online or real-time processing business scenarios. Summary of the Invention

[0003] In view of the above, this disclosure provides a method, apparatus, computing device, computer-readable storage medium, and computer program product for image processing, which can alleviate, reduce, or even eliminate the above-mentioned problems.

[0004] According to one aspect of this disclosure, an image processing method is provided, comprising: extracting a first set of feature points representing the first facial region from a first image containing a first facial region; extracting a second set of feature points representing the second facial region from a second image containing a second facial region, wherein the second set of feature points has a correspondence with the first set of feature points, the correspondence indicating that at least a portion of the second feature points in the second set correspond one-to-one with at least a portion of the first feature points in the first set of feature points; determining a positional mapping from a set of first pixels in the first facial region to a set of second pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points; determining a target region in the second facial region based on the positional mapping and the positions of at least a portion of the first pixels in the first facial region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first facial region; and updating the pixel values ​​of at least a portion of the second pixels in the target region based at least on the pixel values ​​of at least a portion of the first pixels.

[0005] In some embodiments, updating the pixel values ​​of at least a portion of the second pixels in the target region based on the pixel values ​​of at least a portion of the first pixels includes: for each second pixel in the target region, updating the pixel value of the second pixel based on the pixel value of the first pixel corresponding to the second pixel in the at least a portion of the first pixels.

[0006] In some embodiments, updating the pixel values ​​of at least a portion of the second pixels in the target region based on at least a portion of the pixel values ​​of the first pixels includes: updating the pixel values ​​of at least a portion of the second pixels in the target region through image fusion processing from the first image to the second image based on the pixel values ​​of the first pixels.

[0007] In some embodiments, updating the pixel values ​​of at least a portion of the second pixels in the target region through image fusion processing from the first image to the second image based on the pixel values ​​of at least a portion of the first pixels includes: determining the gradient of the first image at at least a portion of the first pixels based on the pixel values ​​of at least a portion of the first pixels; and fusing the first facial region of the first image into the target region of the second image based on the gradient of the first image at at least a portion of the first pixels, such that the pixel values ​​of each second pixel at the edge of the target region in the fused second image remain unchanged, and the gradient loss of at least a portion of the second pixels in the target region of the fused second image relative to the gradient loss of at least a portion of the first pixels in the first image is minimized.

[0008] In some embodiments, determining the position mapping from the first set of pixels in the first facial region to the second set of pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points includes: determining at least one of a translation mapping vector, a scaling mapping factor, and a rotation mapping matrix from the first set of pixels in the first facial region to the second set of pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points.

[0009] In some embodiments, determining the position mapping from the first set of pixels in the first facial region to the second set of pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points includes: substituting the coordinates of at least a portion of the first feature points into a predetermined position mapping function to obtain the mapped coordinates of at least a portion of the first feature points, wherein the predetermined position mapping function includes at least one undetermined mapping parameter; determining the position mapping loss based on the mapped coordinates of at least a portion of the first feature points and the coordinates of at least a portion of the second feature points; and determining the at least one undetermined mapping parameter based on the position mapping loss, such that the position mapping loss is minimized.

[0010] In some embodiments, the first feature point set includes at least one first external feature point characterizing a first facial contour in a first facial region and at least one first internal feature point characterizing facial features in the first facial region, and wherein the second feature point set includes at least one second external feature point characterizing a second facial contour in a second facial region and at least one second internal feature point characterizing facial features in the second facial region.

[0011] In some embodiments, extracting a first set of feature points characterizing the first facial region from a first image containing the first facial region includes: determining a first outer bounding region surrounding the first facial region based on the first image; determining at least one first outer feature point based on the first outer bounding region; determining a first inner bounding region surrounding facial features in the first facial region based on the first image; determining at least one first inner feature point based on the first inner bounding region; and determining a first set of feature points based on the at least one first outer feature point and the at least one first inner feature point.

[0012] In some embodiments, the first outer surrounding region includes a minimum rectangular region that surrounds the first facial contour in the first facial region, or the first inner surrounding region includes a minimum rectangular region that surrounds the facial features in the first facial region.

[0013] In some embodiments, extracting a second set of feature points characterizing the second facial region from a second image containing the second facial region includes: determining a second outer bounding region surrounding the second facial region based on the second image; determining at least one second outer feature point based on the second outer bounding region; determining a second inner bounding region surrounding facial features in the second facial region based on the second image; determining at least one second inner feature point based on the second inner bounding region; and determining a second set of feature points based on the at least one second outer feature point and the at least one second inner feature point.

[0014] In some embodiments, the second outer surrounding region includes a minimum rectangular region that surrounds the second facial contour in the second facial region, or the second inner surrounding region includes a minimum rectangular region that surrounds the facial features in the second facial region.

[0015] In some embodiments, determining at least one first internal feature point based on a first enclosing region includes: dividing the first enclosing region into a plurality of first sub-regions, such that each first sub-region contains a different facial sensory region; and determining at least one first internal feature point based on the plurality of first sub-regions. Alternatively, determining at least one second internal feature point based on a second enclosing region includes: dividing the second enclosing region into a plurality of second sub-regions, such that each second sub-region contains a different facial sensory region; and determining at least one second internal feature point based on the plurality of second sub-regions.

[0016] In some embodiments, extracting a first set of feature points representing the first facial region from a first image containing a first facial region includes: scaling the first image to a first scaled image with a predetermined size; determining the first set of feature points based on the first scaled image, and wherein extracting a second set of feature points representing the second facial region from a second image containing a second facial region includes: scaling the second image to a second scaled image with a predetermined size; determining the second set of feature points based on the second scaled image.

[0017] In some embodiments, determining a first set of feature points based on a first scaled image includes: inputting the first scaled image to a feature point extraction model to determine the first set of feature points, wherein the feature point extraction model is configured to extract a set of feature points representing a facial region from an input image having a predetermined size, and wherein determining a second set of feature points based on a second scaled image includes: inputting the second scaled image to the feature point extraction model to determine the second set of feature points.

[0018] In some embodiments, at least one of the first facial region and the second facial region includes a human facial region, an animal facial region, or a cartoon character facial region.

[0019] According to another aspect of this disclosure, an image processing apparatus is provided, comprising: a first extraction module configured to extract a first set of feature points representing the first facial region from a first image containing a first facial region; a second extraction module configured to extract a second set of feature points representing the second facial region from a second image containing a second facial region, the second set of feature points having a correspondence with the first set of feature points, the correspondence indicating that at least a portion of the second feature points in the second set correspond one-to-one with at least a portion of the first feature points in the first set of feature points; a first determination module configured to determine a position mapping from a set of first pixels in the first facial region to a set of second pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points; a second determination module configured to determine a target region in the second facial region based on the position mapping and the positions of at least a portion of the first pixels in the first facial region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first facial region; and an update module configured to update the pixel values ​​of at least a portion of the second pixels in the target region based at least on the pixel values ​​of at least a portion of the first pixels.

[0020] According to another aspect of this disclosure, a computing device is provided, comprising: a memory configured to store computer-executable instructions; and a processor configured to perform the method described in the foregoing aspect when the computer-executable instructions are executed by the processor.

[0021] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed, perform the method described in the foregoing aspect.

[0022] According to another aspect of this disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the method described in the foregoing aspect.

[0023] In the image processing methods and apparatus provided according to some embodiments of this disclosure, the replacement of the second facial region in the second image (i.e., replacing the second facial region in the second image with the first facial region in the first image) can be achieved by image processing of a first image containing a first facial region and a second image containing a second facial region. This image processing process includes extracting or detecting corresponding feature points of the two images, determining the positional mapping between pixels in the two images based on the correspondence of feature points, determining the target region of the second image based on the positional mapping, and updating pixels in the target region (i.e., facial replacement). In other words, compared with related technologies, the image processing methods and apparatus provided according to some embodiments of this disclosure only involve simple and efficient image processing techniques such as image detection, image mapping, and pixel value replacement or updating, thereby avoiding the inefficiency and time-consuming drawbacks caused by large amounts of training data and complex training processes in related technologies, and achieving a convenient and efficient facial replacement process. Simultaneously, due to the use of relatively refined detection technology for key facial feature points (such as facial features, facial contours, etc.), a more accurate positional mapping can be determined, thereby achieving a more natural and realistic facial replacement effect. This can better meet the real-time facial replacement needs in more application scenarios and significantly improve the user experience.

[0024] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description

[0025] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0026] Figure 1A and 1B The illustrations illustrate example application scenarios where image processing solutions provided by some embodiments of this disclosure can be applied;

[0027] Figure 2 An exemplary schematic diagram of a face replacement scheme in the related art is shown;

[0028] Figure 3An example flowchart illustrating an image processing method according to some embodiments of the present disclosure is shown schematically;

[0029] Figure 4 Exemplary facial feature points are illustrated schematically according to some embodiments of the present disclosure;

[0030] Figure 5 A schematic diagram of a scheme for extracting a set of feature points characterizing a facial region, according to some embodiments of the present disclosure, is shown as an example.

[0031] Figure 6 An exemplary schematic diagram illustrating the principle of image fusion processing according to some embodiments of the present disclosure is shown;

[0032] Figure 7 The illustrations show example effect diagrams of image processing methods according to some embodiments of the present disclosure;

[0033] Figure 8 An example block diagram of an image processing apparatus according to some embodiments of the present disclosure is illustrated schematically; and

[0034] Figure 9 An example block diagram of a computing device according to some embodiments of the present disclosure is illustrated schematically. Detailed Implementation

[0035] Figure 1A and 1B Example application scenarios 100A and 100B are illustrated, in which image processing solutions provided by some embodiments of the present disclosure can be applied.

[0036] like Figure 1A As shown, application scenario 100A includes a computing device 110. Exemplarily, a user can provide a first image 120A and a second image 120B through the input interface of the computing device 110; alternatively, the first image 120A and the second image 120B can be different parts of the same image, and the user can specify one part of the image as the first image 120A and another part as the second image 120B through the input interface of the computing device 110. An application can be deployed on the computing device 110, which can be used to enable the computing device 110 to execute image processing methods provided according to some embodiments of this disclosure to perform facial replacement on facial regions in the first image 120A and the second image 120B, for example, replacing at least a portion of the facial region in the second image 120B with at least a portion of the facial region in the first image 120A, thereby obtaining a composite image 130. The above process will be described in detail below. Exemplarily, the computing device 110 can use its output interface to present the obtained composite image 130 to the user.

[0037] like Figure 1B As shown, application scenario 100B includes computing device 110'. Similarly, a user can provide or specify a first image 120A and a second image 120B through the input interface of computing device 110'. Furthermore, application scenario 100B may also include server 140. An application can be deployed on server 140, which can be used to enable computing device 110 to perform image processing methods provided according to some embodiments of this disclosure. Exemplarily, computing device 110' can send the received first image 120A and second image 120B to server 140 via network 150. Server 140 can perform facial replacement on facial regions in the first image 120A and second image 120B to obtain a composite image 130, and send the composite image 130 to computing device 110' via network 150. Computing device 110' can present the composite image 130 to the user using its output interface.

[0038] In addition to being deployed separately on a computing device or server, the image processing methods provided according to some embodiments of this disclosure may optionally be deployed partly on a computing device and partly on a server.

[0039] Optionally, Figure 1A and 1B Although the computing devices 110 and 110' shown are respectively a desktop computer and a smartphone, they can also be other types of computing devices. For example, the computing device can be one or more combinations of desktop computers, laptop computers, tablet computers, smartphones, smartwatches, smart glasses, in-vehicle devices, etc.

[0040] Optionally, server 140 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Furthermore, it should be understood that server 140 is shown only as an example; in practice, other devices or combinations of devices with computing and storage capabilities can be used alternatively or additionally to provide the corresponding services.

[0041] Alternatively, network 150 can be a wired network connected via a cable, fiber optic cable, or a wireless network such as 2G, 3G, 4G, 5G, Wi-Fi, Bluetooth, ZigBee, Li-Fi, etc.

[0042] Figure 2 An exemplary schematic diagram of a face replacement scheme 200 in the related art is shown.

[0043] Face replacement scheme 200 is an autoencoder-decoder pairing structure scheme. In this scheme, the encoder can extract implicit features from the facial image, for example, by using a neural network structure to reduce and compress the input facial image data and output the corresponding implicit features. The decoder can then reconstruct the facial image based on these implicit features. Figure 2 As shown, to swap facial regions in original facial image A and original facial image B, two encoder-decoder pairs can be trained using the two sets of facial images. The encoder parameters in the two encoder-decoder pairs can be shared, while the decoder parameters are not. This parameter-sharing strategy between the encoders ensures that the two encoders remain identical, allowing them to learn commonalities between the two sets of facial regions, such as the common features of facial features like the mouth and nose. After training, the trained encoder can extract implicit features from the original facial image A, i.e., the features of face A, and then the trained decoder B can reconstruct the facial image based on the features of face A, thus obtaining facial image B reconstructed from the original facial image A. This achieves the effect of replacing facial regions in the original facial image B with facial regions from the original facial image A. Alternatively, the effect of replacing facial regions in the original facial image A with facial regions from the original facial image B can be achieved in a similar manner. Furthermore, in related technologies, additional mechanisms can be added to the above encoder-decoder structure to further improve the quality of face replacement.

[0044] Figure 2 The face replacement scheme 200 and similar schemes shown can achieve good replacement results. However, these schemes require multiple facial images during the training phase to train a complete encoder and decoder, thereby ensuring good face replacement results. For example, the above training process typically requires more than 50 facial images. In addition, the encoder and decoder in the above schemes often involve deep neural networks, and their training process is computationally intensive and time-consuming, making it difficult to apply to many application scenarios, such as some scenarios that require real-time face replacement or scenarios with limited computing resources.

[0045] Based on the above analysis, the applicant proposes a new image processing method that does not require a complex training process and can maintain good face replacement results while ensuring processing efficiency, thereby helping to solve or alleviate the above problems.

[0046] Figure 3An example flowchart of an image processing method 300 according to some embodiments of the present disclosure is illustrated schematically. Exemplarily, the image processing method 300 can be deployed in... Figure 1A The computing device 110 in scenario 100A shown, or it can be deployed on... Figure 1B On server 140 or a combination of computing device 110' and server 140 in scenario 100B shown.

[0047] In step 310, a first set of feature points characterizing the first facial region can be extracted from the first image containing the first facial region. Exemplarily, when the image processing method 300 runs on a computing device or server, the first image can be provided by a user through the input interface of the computing device or server, or it can be read by the computing device or server from internal or external storage, or it can be received by the computing device or server from another computing device or server via a network. Furthermore, the first image can also be a portion of an image cropped from a larger image; this cropping process can be performed manually by the user or automatically by the computing device or server. The first image can be a standalone static image or a frame image from a video. The first image can contain the first facial region and optionally may contain other regions besides the first facial region. The first facial region can refer to an image region containing the first face, which may contain only the first face or additionally include other regions not related to the first face. The first set of feature points can refer to a set consisting of at least one feature point characterizing the first facial region, wherein the feature point can be a point used to characterize the location of the facial contour, facial features, or other facial features of the first face in the first facial region. The number of feature points in the first set of feature points can be set according to the needs of the specific application scenario.

[0048] In some embodiments, the user may specify the first set of feature points. For example, the user may manually mark the feature points included in the first set of feature points in the first image through an input interface of a computing device or server, so that the computing device or server can extract the first set of feature points accordingly; or, the user may input a file listing the locations of the feature points included in the first set of feature points through an input interface of a computing device or server, so that the computing device or server can read the corresponding first set of feature points based on the file; or, before providing the first image, the feature points included in the first set of feature points may be identified with specific markers in the first image, so that the computing device or server can extract the first set of feature points by recognizing the specific markers; and so on.

[0049] In some embodiments, the first feature point set can be automatically determined by a computing device or server. For example, key parts such as facial contours and facial features can be located by analyzing grayscale changes, brightness variations, etc., in the first image, thereby extracting the first feature point set representing these key parts; alternatively, the brightness, grayscale, and other information of pixels in the first image can be analyzed by combining information such as facial symmetry and prior rules of facial feature distribution to extract the aforementioned first feature point set; alternatively, the color information of pixels in the first image can be analyzed by combining different color features of different facial features, such as the color features of lips, eyes, eyebrows, etc., to extract the aforementioned first feature point set; alternatively, a machine learning model can be used to extract the aforementioned first feature point set in the first image; and so on.

[0050] In step 320, a second set of feature points representing the second facial region can be extracted from the second image containing the second facial region. The second set of feature points may correspond to the first set of feature points, indicating a one-to-one correspondence between at least a portion of the second feature points in the second set and at least a portion of the first feature points in the first set. The second image may be an image containing the second facial region similar to the first image, and its acquisition process can be performed according to the various embodiments described in step 310 for acquiring the first image. Optionally, the second image and the first image may be two independent images or different portions of the same image. The second set of feature points may be a set similar to the first set of feature points, where some or all of the second feature points may have a one-to-one correspondence with some or all of the first feature points in the first set. The extraction process of the second set of feature points can be performed according to the various embodiments described in step 310 for extracting the first set of feature points. Optionally, depending on specific application requirements, the extraction method of the second set of feature points may be the same as or different from the extraction method of the first set of feature points.

[0051] In step 330, a positional mapping from a first set of pixels in a first facial region to a second set of pixels in a second facial region can be determined based on the correspondence between the first set of feature points and the second set of feature points. In some embodiments, a positional mapping can be determined based on a one-to-one correspondence between at least a portion of the first feature points in the first set and at least a portion of the second feature points in the second set of feature points, and this positional mapping can be used as the positional mapping from the first set of pixels in the first facial region to the second set of pixels in the second facial region. Exemplarily, the positional mapping can be expressed in the form of a function, a transformation matrix, etc. In some embodiments, the positional mapping function from the first set of pixels in the first facial region to the second set of pixels in the second facial region can be determined using various data fitting methods. For example, for each coordinate dimension, the coordinate values ​​of at least a portion of the first feature points in the first feature point set under that coordinate dimension can be regarded as independent variables, and the coordinate values ​​of at least a portion of the second feature points in the second feature point set under that coordinate dimension can be regarded as dependent variables. A linear fitting function, a polynomial fitting function, or other nonlinear fitting function can be used, employing methods such as least squares or Newton's iteration method, to fit the functional relationship between the independent and dependent variables, thereby obtaining the fitted position mapping function, which serves as the position mapping from the first set of first pixels in the first facial region to the second set of second pixels in the second facial region. In some embodiments, the position mapping function, position mapping matrix, and position mapping vector can be solved by constructing a position loss function and minimizing that position loss function. The position loss function can represent the position loss of at least a portion of the first feature points in the first feature point set during the mapping process. This can be measured, for example, by the sum, sum of squares, or average of the distances between the mapped positions of at least a portion of the first feature points in the first feature point set and the corresponding positions of the second feature points in the second feature point set. These distances can be Euclidean distance, cosine distance, Mahalanobis distance, Manhattan distance, etc. Since the first feature point set and the second feature point set are used to characterize the corresponding features of the first facial region and the second facial region, such as the position and shape of facial contours, facial features or other facial features, the position mapping determined based on the correspondence between the first feature point set and the second feature point set can be used to represent the position mapping relationship from the first pixel point set in the first facial region to the second pixel point set in the second facial region.

[0052] In step 340, a target region can be determined in the second face region based on the position mapping and the positions of at least a portion of the first pixels in the first face region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first face region. For example, the position mapping determined in step 330 can be applied to the positions of at least a portion of the first pixels in the first face region to obtain at least one target position in the second face region. The set of second pixels at the at least one target position can constitute the target region, or the smallest region covering the second pixels at the at least one target position can be determined as the target region. Optionally, during the mapping process, for each first pixel, its mapped position coordinates are rounded up, rounded down, or rounded to obtain the corresponding target position. In some examples, such as in an example where the first face region in the first image is larger than the second face region in the second image, the positions of multiple first pixels in the first face region may be mapped to the same target position; or, in some examples, such as in an example where the first face region in the first image is smaller than the second face region in the second image, the positions of adjacent first pixels in the first face region may be mapped to non-adjacent target positions; and so on. These will not affect the implementation of this method. Optionally, at least a portion of the first pixels in the first facial region may include all or part of the pixels constituting the first face. For example, the smallest closed curve enclosing all feature points in the first feature point set can be determined, and some or all of the first pixels within the closed curve can be used as at least a portion of the first pixels in the first facial region; or, a triangle, quadrilateral, pentagon, hexagon, etc., covering all the first feature points in the first feature point set can be determined, and some or all of the first pixels within the determined shape region can be used as at least a portion of the first pixels in the first facial region; or, the user can pre-specify the facial region to be mapped, and some or all of the first pixels within the specified region can be used as at least a portion of the first pixels in the first facial region; and so on.

[0053] In step 350, the pixel values ​​of at least a portion of the second pixels in the target region can be updated based on at least a portion of the pixel values ​​of the first pixels. For example, for each second pixel in the target region, its pixel value can be updated based on the pixel values ​​of one or more corresponding first pixels in the first face region. Optionally, for a second pixel in the target region, if it has multiple corresponding first pixels in the first face region, its pixel value can be updated based on the pixel value of one of the multiple first pixels, or it can be updated based on the mean or median of the pixel values ​​of the multiple first pixels, etc. Optionally, for a second pixel in the target region, if it does not have a corresponding first pixel in the first face region, its pixel value can remain unchanged, or it can be updated using the updated pixel values ​​of one or more neighboring second pixels, or it can be updated based on the pixel values ​​of the corresponding first pixels in the first face region of one or more neighboring second pixels, etc. Alternatively, by way of example, gradient information of the first facial region can be determined based on the pixel values ​​of at least a portion of the first pixels in the first facial region, and the pixel values ​​of the second pixels in the target region can be updated based on the determined gradient information and the pixel values ​​of a portion of the pixels in the second image, and so on.

[0054] The image processing method 300 described above enables face replacement without using an encoder-decoder structure. During face replacement, only a first image and a second image are required; no large amount of training data needs to be collected beforehand, nor is a time-consuming and labor-intensive training process necessary. Therefore, this scheme can achieve face replacement conveniently and efficiently. Furthermore, due to the use of refined facial key feature point detection technology (such as facial features and contours), more accurate positional mappings can be determined, resulting in a more natural and realistic face replacement effect. This better meets the real-time face replacement needs in more application scenarios and significantly improves the user experience.

[0055] In some embodiments, at least one of the first facial region and the second facial region includes a human facial region, an animal facial region, or a cartoon character facial region. This allows for the replacement of human faces, animal faces, or cartoon character faces, or even the replacement of human faces with animal faces or cartoon character faces.

[0056] In some embodiments, the first set of feature points determined in step 310 may include at least one first external feature point characterizing a first facial contour in the first facial region and at least one first internal feature point characterizing facial features in the first facial region. Similarly, the second set of feature points determined in step 320 includes at least one second external feature point characterizing a second facial contour in the second facial region and at least one second internal feature point characterizing facial features in the second facial region. (Illustratively,) Figure 4 Example external feature points and example internal feature points characterizing a facial region are shown. Figure 4 As shown, several external feature points 410 are marked on the facial contour in the left facial image. These external feature points can be used to characterize the features of the facial contour. Several internal feature points 420 are marked on the eyebrows, eyes, and mouth in the right facial image. These internal feature points can be used to characterize facial features within the facial contour, such as the position or shape of facial features or other facial features. For example, internal feature points can be feature points on the facial contour. External feature points can better describe the position and shape of the facial contours of the first and second faces, thus helping to more accurately replace the area of ​​the second face with the first face. Internal feature points can better describe the position and shape of facial features in the first and second faces. This helps to appropriately adjust the position and shape of facial features in the first face during the face replacement process, so that the distribution of the replaced facial features is closer to the distribution of facial features in the second face, and can retain the expression features of the second face to a certain extent, thus providing a replacement effect that is closer to the distribution of facial features and emotional state of the second face.

[0057] Furthermore, dividing the first feature point into a first external feature point and a first internal feature point, and dividing the second feature point into a second external feature point and a second internal feature point, has the following advantages: Compared to the internal facial features, the facial contour usually has fewer local texture or color features, and there is more background interference. Moreover, compared to the facial features, the shape transitions of the facial contour are often more ambiguous and difficult to accurately determine. Therefore, the position of external feature points is more difficult to determine than that of internal feature points. Determining both external and internal feature points simultaneously may result in excessive computational effort being spent on determining external feature points, which is detrimental to the accurate positioning of internal feature points. Therefore, dividing feature points into external and internal feature points and determining them separately helps improve the accuracy of the determined internal feature points and also improves the efficiency of feature point determination.

[0058] In some embodiments, the extraction process of the first feature point set in step 310 and / or the second feature point set in step 320 can be based on... Figure 5 The proposed solution 500 will be implemented.

[0059] Specifically, at least one first external feature point in the first feature point set in step 310 can be extracted through the following process: First, a first outer bounding region surrounding the first facial region can be determined based on the first image; then, at least one first external feature point can be determined based on the first outer bounding region. (Illustratively,) Figure 5 Image 510 in the image can be considered as the first image. First, a first enclosing region 520, as indicated by the dashed line, can be initially determined in image 510. This first enclosing region 520 is shown as a rectangular region enclosing at least a portion of the facial contour of the first facial region. Exemplarily, the first enclosing region can be the smallest rectangular region enclosing the first facial contour of the first facial region. Alternatively, the first enclosing region 520 can also be a region of other shapes, such as an elliptical or circular region that better conforms to the facial contour, such as the smallest elliptical or circular region enclosing the first facial contour of the first facial region. Exemplarily, the first enclosing region 520 can be determined by manual annotation or automatically by facial recognition. Subsequently, at least one first external feature point 522 can be determined in the first enclosing region 520. These first external feature points 522 are points on the facial contour and can participate in forming the first feature point set 540. Exemplarily, the first external feature points 522 can be determined by the aforementioned inflection points based on grayscale, brightness, or color changes, or they can be determined by a machine learning model, etc. By first roughly determining the first enclosing region, and then refining the determination of the first external feature points based on the first enclosing region, it is helpful to improve the extraction efficiency and accuracy of the first external feature point set.

[0060] Similarly, at least one second external feature point in the second feature point set in step 320 can be extracted through the following process: First, a second outer bounding region surrounding the second facial region can be determined based on the second image; then, at least one second external feature point can be determined based on the second outer bounding region. Exemplarily, the second outer bounding region may include the smallest rectangular region surrounding the second facial contour within the second facial region. This process can be performed similarly to the extraction process of the first feature point set described above and can offer similar advantages, which will not be repeated here. Furthermore, similarly, these second external feature points can participate in constituting the second feature point set.

[0061] Optionally, in step 340, when determining the target region in the second face region based on the position mapping and the positions of at least a portion of the first pixels in the first face region, the pixels in the first enclosing region described in the foregoing embodiment can be used as at least a portion of the pixels in the first face region to determine the target region in the second face region. This simplifies the determination of at least a portion of the first pixels in the first face region that should be mapped. Furthermore, when the first enclosing region is a minimum rectangular, elliptical, or circular region that encloses at least a portion of the facial contour surrounding the first face region, the pixels in the determined first face region can have fewer redundant pixels, thereby simplifying the face replacement operation while maintaining good replacement results.

[0062] At least one first internal feature point in the first feature point set in step 310 can be extracted through the following process: First, based on the first image, a first inner bounding region surrounding the facial features in the first facial region can be determined; then, based on the first inner bounding region, at least one first internal feature point can be determined. (Illustratively,) Figure 5 Image 510 in the image can be considered as the first image. First, a first enclosing region 530, as indicated by the dashed line, can be initially determined in image 510. This first enclosing region 530 is shown as a rectangular region enclosing the facial features. Exemplarily, the first enclosing region may include the smallest rectangular region enclosing the facial features in the first facial region. Optionally, the first enclosing region 530 may also be a region of other shapes. Exemplarily, the first enclosing region 530 can be determined by manual annotation or automatically by facial recognition. Subsequently, at least one first internal feature point 532 can be determined in the first enclosing region 530. These first internal feature points 532 can be points within the facial contour, such as points on the facial feature contour, which can participate in forming the first feature point set 540. That is, the first feature point set 540 can be determined based on at least one first external feature point 522 and at least one first internal feature point 532. Exemplarily, the first internal feature point 532 can be determined by the aforementioned inflection points based on grayscale, brightness, or color changes, or it can be determined by a machine learning model, etc. By first roughly determining the first inner enclosing region, and then refining the determination of the first internal feature points based on the first inner enclosing region, it is helpful to improve the extraction efficiency and accuracy of the first internal feature point set.

[0063] Furthermore, in some embodiments, at least one first internal feature point can be determined based on the first enclosing region through the following steps. Specifically, the first enclosing region can first be divided into multiple first sub-regions, such that each first sub-region contains different facial feature regions; then, at least one first internal feature point can be determined based on the multiple first sub-regions.

[0064] Indicatively, such as Figure 5 As shown, the first enclosing region 530 can be further divided into multiple first sub-regions 534. Each sub-region may include a facial feature region, such as eyes, eyebrows, mouth, nose, etc., and optionally may also include other feature regions, such as skin textures like nasolabial folds or other facial features. For example, the first enclosing region 530 can be divided into six first sub-regions 534, including two eyebrow regions, two eye regions, one mouth region, and one nose region. Optionally, after initially determining at least one first internal feature point based on the first enclosing region 530, the first enclosing region 530 can be further divided into multiple first sub-regions 534 based on the determined first internal feature points, such that each first sub-region 534 can be a minimum region including a set of initially determined first internal feature points, wherein the set of first internal feature points is used to characterize one of the eyes, eyebrows, mouth, nose, or other facial features.

[0065] Furthermore, at least one initially determined first internal feature point can be adjusted and / or at least one new first internal feature point can be determined more precisely based on multiple first sub-regions 534. For example, each first sub-region 534 can be considered an independent image region, and the first internal feature points within it can be independently refined. Optionally, multiple first sub-regions 534 can be processed by rotation or other operations to obtain processed first sub-regions 534, and at least one first internal feature point can be adjusted and / or determined based on the processed first sub-regions 534. For example, for each first sub-region 534, its rotation angle can be estimated, which can adjust the facial features in the first sub-region to be placed or extended along a preset direction or approximately along a preset direction. For example, for the first sub-region containing the eyes, the rotation angle can make the line connecting the inner and outer corners of the eyes horizontal or approximately horizontal; for the first sub-region containing the nose, the rotation angle can make the axis of symmetry of the nose vertical or approximately vertical; for the first sub-region containing the mouth, the rotation angle can make the line connecting the corners of the mouth horizontal or approximately horizontal; for the first sub-region containing the eyebrows, the rotation angle can make the line connecting the inner and outer corners of the eyebrows horizontal or approximately horizontal; and so on. Optionally, the facial features in each first sub-region can be adjusted to other preset angles. Optionally, the preset angles can be manually set or learned by the algorithm model during training. Then, based on the rotated first sub-regions 536, the position coordinates of each first internal feature point can be determined more precisely and accurately. For the position coordinates of each first internal feature point determined based on the rotated first sub-regions 536, a reverse rotation transformation can be performed using the rotation angle applied to the corresponding first sub-region 534, so that the position coordinates of each first internal feature point after the reverse rotation transformation can reflect its actual position in the facial region 510.

[0066] By dividing the region into sub-regions as described above, each sub-region can be analyzed and processed more precisely, which helps to more accurately determine at least one of the first internal feature points.

[0067] Similarly, at least one second internal feature point in the second feature point set in step 320 can be extracted through the following process: First, based on the second image, a second inner surrounding region enclosing the facial features in the second facial region can be determined; then, at least one second internal feature point can be determined based on the second inner surrounding region. In some embodiments, at least one first internal feature point can be determined based on the first inner surrounding region through the following steps: dividing the second inner surrounding region into multiple second sub-regions, each sub-region containing different facial feature regions; and determining at least one second internal feature point based on the multiple second sub-regions. Exemplarily, the second inner surrounding region may include the smallest rectangular region, or the smallest elliptical region, the smallest circular region, etc., surrounding the facial features in the second facial region. The above process can be performed similarly to the extraction process of the first feature point set described above and can bring similar advantages, which will not be repeated here. Furthermore, similarly, these second internal feature points can participate in constituting the second feature point set. That is, the second feature point set can be determined based on at least one second external feature point and at least one second internal feature point.

[0068] In some embodiments, to generate a composite image with a predetermined size and to simplify the feature point extraction process or subsequent mapping processing, the sizes of the first image and / or the second image can be pre-processed to unify their sizes to a predetermined size. Specifically, step 310 may include: scaling the first image to a first scaled image with a predetermined size, and determining a first set of feature points based on the first scaled image. Step 320 may include: scaling the second image to a second scaled image with a predetermined size, and determining a second set of feature points based on the second scaled image. Optionally, the predetermined size may be a pre-specified size, or it may be determined based on the sizes of the first image and the second image, such as the smaller, larger, or average size of the two. Optionally, if the predetermined size is equal to the size of the first image or the second image, scaling of the first image or the second image may be omitted, or it may be understood that the scaling ratio of the first image or the second image is 1. Alternatively, scaling of the first and / or second image can be achieved through various downsampling or upsampling methods. For example, the scaling process can be achieved by discarding some pixels according to a preset rule, adding some pixels, or determining the pixel value of the corresponding position after scaling based on the pixel values ​​of multiple adjacent pixels. Alternatively, when the non-facial area in the first or second image is large, the image size can also be reduced by cropping part of the background area.

[0069] In some embodiments, the feature point extraction process can be implemented using a machine learning model. Specifically, a first scaled image and a second scaled image can be input to a feature point extraction model to determine a first set of feature points and a second set of feature points, wherein the feature point extraction model is configured to extract a set of feature points representing facial regions from an input image of a predetermined size. Exemplarily, the feature point extraction model can be a cascaded CNN (Convolutional Neural Network) model. Specifically, the input to this model can be an image containing facial regions, and the output can be 68 feature points of the face. These feature points can be used to identify the facial contours and the position and shape information of facial features. Specifically, these feature points can be further divided into 51 inner points and 17 outer points, where the inner points are used to identify facial features and are more detailed, while the outer points are used to describe facial contours, shapes, and other information. To improve extraction accuracy, as described in the above embodiments, the model can employ different methods to extract internal and external feature points: For external feature points, the minimum bounding box (MBC), i.e., the smallest rectangular region that can encompass the external feature points, can be estimated first. Then, 17 external feature points can be directly predicted within this region. This process can be similarly referenced. Figure 5 The process of extracting external feature points 522 from facial image 510 is described. Similarly, for internal feature points, their minimum bounding box (i.e., the smallest rectangular region that can encompass the internal feature points) can be estimated first, and preliminary predictions can be made for at least a portion of the 51 internal feature points within this region. This process can be similarly referenced. Figure 5 The process of extracting internal feature points 532 from facial image 510 is described. However, to further improve the prediction accuracy of internal feature points, a more precise analysis of the facial region can be performed using segmentation and rotation techniques. This allows for fine-tuning and updating of the initially predicted internal feature points to obtain more accurate results. This can be similar to [the previous section, which is incomplete and requires further context]. Figure 5 The process of dividing and rotating the sub-regions is described.

[0070] A trained machine learning model can quickly and accurately extract the corresponding set of feature points from a first or second image. In addition to the cascaded CNN network model described above, other known models or self-designed models can also be used to implement this extraction process, such as classifier models, DPM (Deformable Part Model), Dlib models, RetinaFace models, libfacedetect models, seetaface models, etc.

[0071] In some embodiments, step 330 may include: determining at least one of a translation mapping vector, a scaling mapping factor, and a rotation mapping matrix from a first set of pixels in a first facial region to a second set of pixels in a second facial region, based on the correspondence between the first set of feature points and the second set of feature points. By using one or more of the translation mapping vector, scaling mapping factor, and rotation mapping matrix, the mapping relationship from the first set of pixels in the first facial region to the second set of pixels in the second facial region can be described simply and quantitatively, which facilitates faster implementation of subsequent mapping transformations. Specifically, the translation mapping vector can be used to describe the translation transformation from the position coordinates of a pixel in the first set of pixels in the first facial region to the position coordinates of the corresponding pixel in the second set of pixels in the second facial region, and it can be, for example, a 2D vector; the position scaling mapping factor can be used to describe the scaling transformation from the position coordinates of a pixel in the first set of pixels in the first facial region to the position coordinates of the corresponding pixel in the second set of pixels in the second facial region, and it can be an integer or a floating-point number; the position rotation mapping matrix can be used to describe the rotation transformation from the position coordinates of a pixel in the first set of pixels in the first facial region to the position coordinates of the corresponding pixel in the second set of pixels in the second facial region, and it can be, for example, a 2x2 matrix.

[0072] In some embodiments, the mapping in step 330 can be determined such that the position mapping loss of the feature points is minimized during the mapping process. The position mapping loss can represent the difference between the position of a first feature point in a first set of feature points after mapping and the position of a second feature point in a second set of feature points. For example, step 330 may include: substituting the coordinates of at least a portion of the first feature points into a predetermined position mapping function to obtain the mapped coordinates of at least a portion of the first feature points, wherein the predetermined position mapping function includes mapping parameters to be determined; determining the position mapping loss based on the mapped coordinates of at least a portion of the first feature points and the corresponding coordinates of at least a portion of the second feature points; and determining the mapping parameters to be determined based on the position mapping loss, such that the position mapping loss is minimized.

[0073] For example, the absolute values ​​of the coordinate differences between the mapped coordinates of at least a portion of the first feature points and the coordinates of the corresponding at least a portion of the second feature points can be statistically analyzed. These absolute values ​​of coordinate differences can then be summed, averaged, or summed of squares to obtain the aforementioned position mapping loss. When the position mapping loss is minimized, the resulting position mapping can more accurately represent the positional transformation from the coordinates of at least a portion of the first feature points to the coordinates of the corresponding at least a portion of the second feature points. Therefore, when using the obtained position mapping to transform the position of at least a portion of the first pixels in the first facial region to determine the target region in the second facial region, the corresponding target region can be determined more accurately, thereby achieving a more accurate face replacement effect.

[0074] Optionally, when determining the aforementioned location mapping loss, different weights can be assigned to different feature points according to specific application requirements to achieve different facial replacement effects. For example, when determining the location mapping loss, a first weight can be assigned to certain feature points, while a second weight can be assigned to the remaining feature points. The second weight can be less than the first weight. In this way, the determined location mapping can better map the feature points with the first weight to the appropriate positions. For example, when focusing more on facial contours, higher weights can be assigned to feature points on the facial contours; when focusing more on the contours of facial features, higher weights can be assigned to feature points on the contours of facial features; or, higher weights can be assigned to certain turning points of facial or facial feature contours (such as feature points at the corners of the eyes, eyebrows, and corners of the mouth); and so on.

[0075] The following example, describing position mapping using translation mapping vectors, scaling mapping factors, and rotation mapping matrices, illustrates the process of determining position mapping. For example, assume... Let i be the first feature point of the first facial region. It is the i-th second feature point in the second facial region, where the value of i ranges, for example, from... Let R be a 2x2 rotation mapping matrix, s be a scaling mapping factor, and T be a two-dimensional translation mapping vector. Then, a set of R, s, and T can be determined such that the following position mapping loss function is minimized:

[0076] .

[0077] Alternatively, in examples where different weights are set for different feature points, the following location mapping loss function can be used:

[0078] ,

[0079] Where, k i This represents the weight value corresponding to the feature point.

[0080] After determining the translation mapping vector, scaling mapping factor, and rotation mapping matrix according to the above example, in step 340, the determined translation mapping vector, scaling mapping factor, and rotation mapping matrix can be used to transform the position coordinates of at least a portion of the first pixels in the first facial region, thereby determining the corresponding target region in the second image. Compared to more complex mapping functions or other mapping mechanisms, the position transformation achieved using the translation mapping vector, scaling mapping factor, and rotation mapping matrix is ​​simpler and faster, while maintaining high accuracy.

[0081] As mentioned earlier, in step 350, the pixel values ​​of the pixels in the target region can be updated in various ways. In some embodiments, for each second pixel in at least a portion of the second pixels in the target region, the pixel value of the second pixel can be updated based on the pixel value of the first pixel corresponding to the second pixel in at least a portion of the first pixels. For example, for each second pixel in at least a portion of the second pixels in the target region, the pixel value of the first pixel corresponding to the second pixel in the first face region can be used as the pixel value of the second pixel; or, the pixel value of the first pixel corresponding to the second pixel in the first face region can be pre-processed, such as multiplying by a pre-set coefficient or adding to a pre-set coefficient, to determine the pixel value of the second pixel. In this way, the pixel values ​​of at least a portion of the second pixels in the target region in the second image can be quickly determined based on the pixel values ​​of at least a portion of the first pixels in the first face region, thereby quickly completing face replacement.

[0082] In some embodiments, to achieve better face replacement effects, the pixel values ​​of at least a portion of the second pixels in the target region can be updated through image fusion processing from the first image to the second image, based on the pixel values ​​of at least a portion of the first pixels. In this disclosure, fusion processing can refer to processing the pixel values ​​of at least the pixels near the fusion boundary when merging two images into one, so that there is no obvious boundary in the merged image.

[0083] Figure 6 A schematic diagram 600 illustrating the principle of image fusion processing according to some embodiments of the present disclosure is shown. Indicatively, as... Figure 6As shown, the fusion process can be understood as follows: when fusing the first image and the second image 620, for example, fusing a portion of the first image into the target region 610 of the second image 620, there may be no obvious boundary at the boundary 630 of the target region 610. In the aforementioned embodiment, when directly updating the pixel value of the second pixel in the target region based on the pixel value of at least a portion of the first pixel in the first facial region, if there is a significant difference between the first facial region in the first image and the second facial region in the second image, such as differences in skin texture between the first and second faces, the resulting composite image after face replacement using the above scheme may have a clear boundary at the boundary of the target region. To avoid this boundary, a fusion process can be used to generate the final composite image.

[0084] Optionally, the pixel values ​​of each pixel in the target region of the second image can be determined based on the pixel values ​​of the corresponding pixels in the first facial region, and then the pixel values ​​of each pixel in the target region can be updated through fusion processing; or, a target image with the same shape and size as the target region can be generated based on the pixel values ​​of at least some pixels in the first facial region, and then the target image and the second image can be fused to determine the pixel values ​​of each pixel in the target region; or, image fusion can be performed directly based on at least some pixels in the first facial region of the first image and the second image; and so on.

[0085] In some embodiments, the above-described fusion process can be implemented through the following steps: First, the gradient of the first image at at least a portion of the first pixels can be determined based on the pixel values ​​of at least a portion of the first pixels; then, based on the gradient of the first image at at least a portion of the pixels, the first facial region of the first image can be fused into the target region of the second image, such that the pixel value of each second pixel at the edge of the target region in the fused second image remains unchanged, and the gradient loss of at least a portion of the second pixels in the target region of the fused second image relative to the gradient of at least a portion of the first pixels in the first image is minimized. For example, for each pixel in at least a portion of the first pixels in the first facial region, the gradient of that pixel can be determined based on the pixel value of that first pixel and / or the pixel values ​​of its neighboring first pixels. Then, based on the pixel values ​​of the second pixels at the edge of the target region in the second image and the determined gradient of the first image at at least a portion of the pixels, the updated pixel values ​​of each second pixel in the target region can be determined, such that after fusion, the pixel value of each second pixel at the edge of the target region remains unchanged, and the difference between the gradient of at least a portion of the second pixels in the target region and the gradient of the first image at at least a portion of the first pixels is minimized, thereby minimizing the gradient loss. Optionally, the gradient loss can be represented by the sum, average, or sum of squares of the absolute values ​​of the differences between the gradients of at least a portion of the second pixels in the target region and the gradients of the first image at at least a portion of the first pixels. In this way, the image gradient within the target region can match the corresponding portion in the first facial region, while avoiding abrupt changes in pixel values ​​at the edges of the target region. This effectively avoids a sense of boundary at the edges of the target region and maintains the pixel value change trend within the target region consistent with the corresponding portion in the first facial region, achieving a more natural face replacement effect.

[0086] For example, the pixel value of a pixel in the target region can be determined by solving the following Poisson equation:

[0087] ,

[0088] Where f represents the synthesized image after face replacement, f* represents the second image, v represents the gradient of at least one pixel in the first facial region of the first image corresponding to the target region, ▽f represents the gradient of the synthesized image, Ω represents the target region, and ∂Ω represents the edge portion of the target region. When the first and second images are single-channel grayscale images, the above gradient can refer to the gradient of grayscale values; when the first and second images are multi-channel images, such as RGB images including R (red), G (green), and B (blue) channels, the above gradient can refer to the gradient of the color values ​​of each channel, that is, the color values ​​of each channel can be solved by solving three Poisson equations respectively. Through the above formula, while ensuring that the pixel values ​​of the pixels at the edge of the target region in the second image remain unchanged, the gradient of the synthesized image in the target region is made closest to the corresponding part in the first facial region, thereby avoiding abrupt changes at the edge of the target region and avoiding unsightly boundaries.

[0089] In addition, the above-described fusion process can also be implemented using other methods. For example, the pixel values ​​of the second image and the target region's pixels within a preset neighborhood that falls within the target region's boundary can be adjusted to create a smooth transition. For instance, the pixel values ​​of the innermost and outermost pixels in the preset neighborhood can be used as reference values, and the pixel values ​​of other pixels within the preset neighborhood can be determined by interpolation. Alternatively, the fusion process can be achieved by determining a mask that matches the target region and applying Gaussian filtering to the mask's boundaries. And so on.

[0090] Figure 7 An example effect diagram 700 schematically illustrates an image processing scheme according to some embodiments of the present disclosure. For example... Figure 7 As shown, the leftmost column displays the first image, the middle column displays the second image, and the rightmost column displays the composite image after face replacement according to some embodiments of this disclosure. It is evident that the facial regions in the first image can be well replaced with the corresponding facial regions in the second image, and during the synthesis process, the distribution and shape of the facial features in the first facial region can be adaptively adjusted according to the corresponding features of the second facial region. Furthermore, due to the use of fusion processing, there are no obvious boundaries in the composite image, but rather a natural replacement effect.

[0091] Figure 8 An example block diagram of an image processing apparatus 800 according to some embodiments of the present disclosure is shown schematically. Figure 8 As shown, the image processing apparatus 800 includes a first extraction module 810, a second extraction module 820, a first determination module 830, a second determination module 840, and an update module 850. Exemplarily, this image processing apparatus 800 can be deployed in... Figure 1AOn the computing device 110 shown, or can be deployed on Figure 1B On the computing device 110', server 140, or a combination of both shown.

[0092] Specifically, the first extraction module 810 can be configured to extract a first set of feature points representing the first facial region from a first image containing a first facial region; the second extraction module 820 can be configured to extract a second set of feature points representing the second facial region from a second image containing a second facial region, wherein the second set of feature points has a correspondence with the first set of feature points, and the correspondence indicates that at least a portion of the second feature points in the second set correspond one-to-one with at least a portion of the first feature points in the first set of feature points; the first determination module 830 can be configured to determine a positional mapping from a set of first pixels in the first facial region to a set of second pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points; the second determination module 840 can be configured to determine a target region in the second facial region based on the positional mapping and the position of at least a portion of the first pixels in the first facial region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first facial region; and the update module 850 can be configured to update the pixel values ​​of at least a portion of the second pixels in the target region based at least on the pixel values ​​of at least a portion of the first pixels.

[0093] It should be understood that the image processing device 800 can be implemented in software, hardware, or a combination of both. Multiple different modules can be implemented in the same software or hardware architecture, or a single module can be implemented by multiple different software or hardware architectures.

[0094] Furthermore, the image processing apparatus 800 can be used to implement the image processing method 300 described above, the details of which have been described in detail above and will not be repeated here for the sake of brevity. The image processing apparatus 800 can have the same features and advantages as described with respect to the aforementioned method.

[0095] Figure 9 An example block diagram of a computing device 900 according to some embodiments of the present disclosure is illustrated schematically. For example, it may represent... Figure 1A and 1B The computing device 110 or 110' shown, server 140, or other types of computing devices that can be used to deploy the image processing apparatus 800 provided in this disclosure.

[0096] As shown in the figure, the example computing device 900 includes a processing system 901, one or more computer-readable media 902, and one or more I / O interfaces 903 that are communicatively coupled to each other. Although not shown, the computing device 900 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus may include any or a combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of a variety of bus architectures, or may include control and data lines.

[0097] Processing system 901 represents the functionality of performing one or more operations using hardware. Therefore, processing system 901 is illustrated as including hardware elements 904 that can be configured as processors, function blocks, etc. This may include application-specific integrated circuits (ASICs) or other logic devices formed using one or more semiconductors implemented in the hardware. Hardware element 904 is not limited by its forming material or the processing mechanism employed therein. For example, a processor may consist of semiconductors and / or transistors (e.g., integrated circuits (ICs)). In such a context, processor-executable instructions may be electronically executable instructions.

[0098] Computer-readable medium 902 is illustrated as including memory / storage device 905. Memory / storage device 905 represents a memory / storage device associated with one or more computer-readable media. Memory / storage device 905 may include volatile storage media (such as random access memory (RAM)) and / or non-volatile storage media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). Memory / storage device 905 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). Exemplarily, memory / storage device 905 may be used to store the first image, second image, extracted first and second feature point sets, determined mapping relationships, etc., mentioned in the above embodiments. Computer-readable medium 902 may be configured in various other ways as further described below.

[0099] One or more input / output interfaces 903 represent functions that allow a user to type commands and information into a computing device 900 and also allow information to be presented to the user and / or sent to other components or devices using various input / output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch functionality (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., capable of detecting non-touch-related movements as gestures using visible or invisible wavelengths (such as infrared frequencies), network interface cards (NICs), receivers, and so on). Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, haptic-responsive devices, network interface cards (NICs), transmitters, and so on. Exemplarily, in the embodiments described above, an input device may allow a user to provide a first image and a second image, etc., and an output device may allow a user to view a composite image with a replaced face, etc.

[0100] The computing device 900 also includes an image processing application 906. The image processing application 906 can be stored as computing program instructions in a memory / storage device 905. The image processing application 906, together with the processing system 901, etc., implements... Figure 8 The image processing apparatus 800 described includes all the functions of each module.

[0101] This document describes various technologies in the general context of software, hardware, components, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc., that perform specific tasks or implement specific abstract data types. The terms "module," "function," etc., as used herein generally refer to software, firmware, hardware, or a combination thereof. The technologies described herein are platform-independent, meaning that these technologies can be implemented on a variety of computing platforms with various processors.

[0102] Implementations of the described modules and technologies may be stored on or transmitted across some form of computer-readable medium. Computer-readable medium may include a variety of media accessible by the computing device 900. By way of example and not limitation, computer-readable medium may include "computer-readable storage medium" and "computer-readable signal medium".

[0103] In contrast to simple signal transmission, carrier waves, or signals themselves, a "computer-readable storage medium" refers to a medium and / or device capable of persistently storing information, and / or a tangible storage device. Therefore, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented using methods or techniques suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, hard disks, cassette tapes, magnetic tapes, disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of art suitable for storing desired information and accessible by a computer.

[0104] "Computer-readable signal medium" refers to a signal-bearing medium configured to transmit instructions, such as via a network, to computing device 900. A signal medium typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. Signal media also includes any information transmission medium. By way of example and not limitation, signal media includes wired media such as wired networks or direct connections, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0105] As previously described, hardware element 904 and computer-readable medium 902 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements may include components of integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implementations or other hardware devices in silicon. In this context, hardware elements can serve as processing devices for executing program tasks defined by instructions, modules, and / or logic embodied by the hardware element, and as hardware devices for storing instructions for execution, such as the previously described computer-readable storage medium.

[0106] The foregoing combinations can also be used to implement the various techniques and modules described herein. Therefore, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage medium and / or by one or more hardware elements 904. The computing device 900 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, modules can be implemented at least partially in hardware as modules executable as software by the computing device 900, for example, by using the computer-readable storage medium and / or hardware elements 904 of a processing system. Instructions and / or functions can be executed / operated by, for example, one or more computing devices 900 and / or processing system 901 to implement the techniques, modules, and examples described herein.

[0107] The techniques described herein can be supported by these various configurations of computing device 900, and are not limited to specific examples of the techniques described herein.

[0108] It should be understood that, for clarity, embodiments of this disclosure have been described with reference to different functional units. However, it will be apparent that, without departing from this disclosure, the functionality of each functional unit may be implemented in a single unit, in multiple units, or as part of other functional units. For example, functionality described as being performed by a single unit may be performed by multiple different units. Therefore, references to a particular functional unit are considered merely as references to the appropriate unit used to provide the described functionality, and not as indicating a strict logical or physical structure or organization. Thus, this disclosure may be implemented in a single unit, or may be physically and functionally distributed among different units and circuits.

[0109] This disclosure provides a computer-readable storage medium storing computer-readable instructions thereon, which, when executed, implement the above-described image processing method.

[0110] This disclosure provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform the image processing methods provided in the various embodiments described above.

[0111] By studying the accompanying drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps, and "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not imply that a combination of these measures cannot be used for profit.

[0112] It is understood that the specific embodiments of this disclosure involve image data such as facial images. When the embodiments involving such data described in this disclosure are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

Claims

1. An image processing method, comprising: Extract a first set of feature points representing the first facial region from a first image containing the first facial region. The first set of feature points includes at least one first external feature point representing a first facial contour in the first facial region and at least one first internal feature point representing facial features in the first facial region. Extract a second set of feature points representing the second facial region from a second image containing the second facial region. The second set of feature points includes at least one second external feature point representing the second facial contour in the second facial region and at least one second internal feature point representing the facial features in the second facial region. The second set of feature points has a correspondence with the first set of feature points, and the correspondence indicates that at least a portion of the second feature points in the second set of feature points correspond one-to-one with at least a portion of the first feature points in the first set of feature points. Based on the correspondence between the first feature point set and the second feature point set, determine the positional mapping from the first pixel set in the first facial region to the second pixel set in the second facial region; Based on the location mapping and the positions of at least a portion of the first pixels in the first facial region, a target region is determined in the second facial region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first facial region; as well as The pixel values ​​of at least a portion of the second pixels in the target region are updated based on at least a portion of the pixel values ​​of the first pixels. The step of extracting a first set of feature points representing the first facial region from a first image containing the first facial region includes: Based on the first image, a first outer bounding region surrounding the first facial region is determined; Based on the first enclosing region, the at least one first external feature point is determined; Based on the first image, a first inner surrounding region is determined that surrounds the facial features in the first facial region. Based on the first enclosed region, the at least one first internal feature point is determined; A first feature point set is determined based on the at least one first external feature point and the at least one first internal feature point.

2. The method according to claim 1, wherein, Updating the pixel values ​​of at least a portion of the second pixels in the target region based at least a portion of the pixel values ​​of the first pixels includes: For each second pixel in the at least part of the second pixels in the target region, the pixel value of the second pixel is updated based on the pixel value of the first pixel corresponding to the second pixel in the at least part of the first pixels.

3. The method according to claim 1, wherein, Updating the pixel values ​​of at least a portion of the second pixels in the target region based at least a portion of the pixel values ​​of the first pixels includes: Based on the pixel values ​​of at least a portion of the first pixels, the pixel values ​​of at least a portion of the second pixels in the target region are updated through image fusion processing from the first image to the second image.

4. The method according to claim 3, wherein, The step of updating the pixel values ​​of at least a portion of the second pixels in the target region through image fusion processing from the first image to the second image based on the pixel values ​​of the at least a portion of the first pixels includes: Based on the pixel values ​​of the at least a portion of the first pixels, determine the gradient of the first image at the at least a portion of the first pixels; Based on the gradient of the first image at at least a portion of the first pixels, the first facial region of the first image is fused into the target region of the second image, such that the pixel value of each second pixel at the edge of the target region of the fused second image remains unchanged, and the gradient loss of the at least a portion of the second pixels of the fused second image in the target region relative to the gradient loss of the at least a portion of the first pixels of the first image is minimized.

5. The method according to claim 1, wherein, The step of determining the position mapping from the first set of pixels in the first facial region to the second set of pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points includes: Based on the correspondence between the first feature point set and the second feature point set, at least one of the following is determined: a translation mapping vector, a scaling mapping factor, and a rotation mapping matrix from the first pixel point set in the first facial region to the second pixel point set in the second facial region.

6. The method according to claim 1, wherein, The step of determining the position mapping from the first set of pixels in the first facial region to the second set of pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points includes: Substitute the coordinates of the at least part of the first feature points into a predetermined position mapping function to obtain the mapped coordinates of the at least part of the first feature points, wherein the predetermined position mapping function includes at least one undetermined mapping parameter; Based on the mapped coordinates of at least a portion of the first feature points and the coordinates of at least a portion of the second feature points, determine the position mapping loss; Based on the location mapping loss, at least one undetermined mapping parameter is determined such that the location mapping loss is minimized.

7. The method of claim 1, wherein the first outer surrounding region comprises a minimum rectangular region surrounding a first facial contour in the first facial region, or the first inner surrounding region comprises a minimum rectangular region surrounding facial features in the first facial region.

8. The method according to claim 1, wherein, The step of extracting a second set of feature points representing the second facial region from a second image containing the second facial region includes: Based on the second image, a second outer bounding region surrounding the second facial region is determined; Based on the second enclosing region, the at least one second external feature point is determined; Based on the second image, a second inner surrounding region is determined that surrounds the facial features in the second facial region; Based on the second enclosed region, the at least one second internal feature point is determined; The second feature point set is determined based on the at least one second external feature point and the at least one second internal feature point.

9. The method of claim 8, wherein the second outer surrounding region comprises a minimum rectangular region surrounding a second facial contour in the second facial region, or the second inner surrounding region comprises a minimum rectangular region surrounding facial features in the second facial region.

10. The method according to claim 8, wherein, Determining the at least one first internal feature point based on the first enclosing region includes: The first inner enclosing region is divided into multiple first sub-regions, such that each first sub-region contains different facial features. Based on the plurality of first sub-regions, determine the at least one first internal feature point, or Wherein, determining the at least one second internal feature point based on the second enclosing region includes: The second inner surrounding region is divided into multiple second sub-regions, such that each second sub-region contains different facial features. Based on the plurality of second sub-regions, at least one second internal feature point is determined.

11. The method according to claim 1, wherein, The step of extracting a first set of feature points representing the first facial region from a first image containing the first facial region includes: Scale the first image to a first scaled image with a predetermined size; Based on the first scaled image, the first set of feature points is determined, and The step of extracting a second set of feature points representing the second facial region from a second image containing the second facial region includes: The second image is scaled to a second scaled image having the predetermined size; Based on the second scaled image, the second set of feature points is determined.

12. The method according to claim 11, wherein, Determining the first set of feature points based on the first scaled image includes: The first scaled image is input to a feature point extraction model to determine the first set of feature points, wherein the feature point extraction model is configured to extract a set of feature points representing facial regions from an input image having the predetermined size. The step of determining the second set of feature points based on the second scaled image includes: The second scaled image is input into the feature point extraction model to determine the second set of feature points.

13. The method according to claim 1, wherein, At least one of the first facial region and the second facial region includes a human facial region, an animal facial region, or a cartoon character facial region.

14. An image processing apparatus, comprising: A first extraction module is configured to extract a first set of feature points representing the first facial region from a first image containing the first facial region. The first set of feature points includes at least one first external feature point representing a first facial contour in the first facial region and at least one first internal feature point representing facial features in the first facial region. The extraction of the first set of feature points representing the first facial region from the first image includes: determining a first outer region surrounding the first facial region based on the first image; determining the at least one first external feature point based on the first outer region; determining a first inner region surrounding facial features in the first facial region based on the first image; determining the at least one first internal feature point based on the first inner region; and determining a first set of feature points based on the at least one first external feature point and the at least one first internal feature point. The second extraction module is configured to extract a second set of feature points representing the second facial region from a second image containing the second facial region. The second set of feature points includes at least one second external feature point representing the second facial contour in the second facial region and at least one second internal feature point representing the facial features in the second facial region. The second set of feature points has a correspondence with the first set of feature points, and the correspondence indicates that at least a portion of the second feature points in the second set of feature points corresponds one-to-one with at least a portion of the first feature points in the first set of feature points. The first determining module is configured to determine the position mapping from the first set of first pixels in the first facial region to the second set of second pixels in the second facial region based on the correspondence between the first set of feature points and the second set of feature points. The second determining module is configured to determine a target region in the second facial region based on the location mapping and the positions of at least a portion of the first pixels in the first facial region, wherein at least a portion of the second pixels in the target region correspond one-to-one with at least a portion of the first pixels in the first facial region; and The update module is configured to update the pixel values ​​of at least a portion of the second pixels in the target region based at least a portion of the pixel values ​​of the first pixels.

15. A computing device, comprising: Memory, which is configured to store computer-executable instructions; A processor configured to perform the method according to any one of claims 1 to 13 when the computer-executable instructions are executed by the processor.

16. A computer-readable storage medium storing computer-executable instructions that, when executed, perform the method according to any one of claims 1 to 13.

17. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 13.