Model training method, facial data processing method, equipment, medium and product

By acquiring multiple sampled facial images and inputting facial data models, the problem that facial texture map data in the prior art cannot accurately reflect facial multi-angle textures, achieving higher accuracy and reliability.

CN114898027BActive Publication Date: 2025-05-16ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210289349.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-05-16
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

When obtaining facial texture map data in the prior art, the data obtained based on a single image cannot accurately reflect the texture of the unpresented part of the face, and in the presence of occlusion, it cannot accurately reflect the texture of the obstructed part, resulting in poor accuracy and reliability.

Method used

By acquiring multiple sampled facial images collected from different angles, the facial texture map data corresponding to each sampled facial image is obtained, and inputting it into the facial data model to obtain the target facial texture map data. Then, based on the target face texture map data, multiple target face images are obtained, and in response to the multiple target face images matching multiple sample face images, the facial data model is determined as the target face data model.

Benefits of technology

It improves the accuracy and reliability of facial texture map data, can accurately reflect the texture characteristics of the face at multiple angles, and reduces the amount of calculation required to obtain facial texture map data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898027B_ABST
    Figure CN114898027B_ABST
Patent Text Reader

Abstract

The disclosed embodiment discloses a model training method, a facial data processing method, a device, a medium and a product, the method comprising: obtaining a plurality of sampled facial images collected from different angles, and obtaining facial texture map data corresponding to each sampled facial image; obtaining a facial data model, taking the facial texture map data corresponding to the plurality of sampled facial images as input, to obtain target facial texture map data output by the facial data model; rendering based on the target facial texture map data, to obtain a plurality of target facial images; in response to the matching of the plurality of target facial images with the plurality of sampled facial images, determining the facial data model as a target facial data model. Based on the target facial data model obtained by the above scheme, facial texture map data that can accurately reflect the texture features of the face at multiple angles can be obtained, thereby improving the accuracy of the facial texture map data and the reliability of the facial texture map data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a model training method, a facial data processing method, a device, a medium and a product. Background Art

[0002] In recent years, with the development of virtual reality (VR) technology, many related applications such as virtual communication, virtual YouTuber (VTuber) live broadcast, immersive interactive video games, etc. have gradually become popular in people's lives. Among them, in the above applications, it is generally necessary to obtain facial texture map data of the face in order to display the corresponding virtual image based on the facial texture map data.

[0003] In the related art, when acquiring facial texture map data of a face, a single image of the face can be acquired first, and facial texture map data of the face in the single image can be generated based on the acquired single image. However, in the above scheme, although facial texture map data can be acquired relatively conveniently, since the facial texture map data in the above scheme is acquired based on a single image, the acquired facial texture map data cannot accurately reflect the facial texture corresponding to the facial part not included in the single image. For example, when the single image is a frontal view image of the face or a near-frontal view image of the face, the facial texture map data acquired based on the single image cannot reliably reflect the texture of the side of the face; in addition, if there is occlusion of the face in the single image, the facial texture map data acquired based on the single image cannot accurately reflect the texture of the occluded part of the face. Therefore, the accuracy and reliability of the facial texture map data acquired by the above scheme are relatively poor. Summary of the invention

[0004] In order to solve the problems in the related art, the embodiments of the present disclosure provide a model training method, a facial data processing method, a device, a medium and a product.

[0005] In a first aspect, an embodiment of the present disclosure provides a model training method, the method comprising:

[0006] Acquire multiple sampled facial images collected from different angles, and acquire facial texture map data corresponding to each sampled facial image;

[0007] Acquire a facial data model, and use facial texture map data corresponding to a plurality of sampled facial images as input to acquire target facial texture map data output by the facial data model;

[0008] Rendering is performed based on the target facial texture map data to obtain multiple target facial images;

[0009] In response to the plurality of target facial images matching the plurality of sample facial images, a facial data model is determined as a target facial data model.

[0010] In an implementation of the present disclosure, obtaining facial texture map data corresponding to each sampled facial image includes:

[0011] Performing facial reconstruction on each of the multiple sampled facial images to obtain facial geometry data corresponding to each sampled facial image;

[0012] According to each sampled facial image and the facial geometry data corresponding to each sampled facial image, the facial texture map data corresponding to each sampled facial image is obtained.

[0013] In an implementation of the present disclosure, facial texture map data corresponding to a plurality of sampled facial images are used as input to obtain target facial texture map data output by a facial data model, including:

[0014] Taking facial texture map data and facial geometry data corresponding to a plurality of sampled facial images as input, to obtain target facial texture map data and target facial geometry data output by a facial data model;

[0015] Rendering is performed based on the target facial texture map data to obtain multiple target facial images, including:

[0016] Rendering is performed based on the target facial texture map data and the target facial geometry data to obtain a plurality of target facial images.

[0017] In an implementation of the present disclosure, in response to matching a plurality of target facial images with a plurality of sampled facial images, determining a facial data model as a target facial data model includes:

[0018] In response to the plurality of target facial images matching the plurality of sampled facial images and the plurality of target facial images not including artifact regions, the facial data model is determined as a target facial data model.

[0019] In an implementation of the present disclosure, before taking the facial texture map data corresponding to the plurality of sampled facial images as input, the method further includes:

[0020] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0021] The facial texture map data corresponding to multiple sampled facial images are used as input, including:

[0022] The facial texture map data of the artifact-removed data corresponding to the multiple sampled facial images is taken as input.

[0023] In an implementation of the present disclosure, before acquiring a plurality of sampled facial images collected from different angles, the method further includes:

[0024] Get sampled facial videos;

[0025] Get multiple sample facial images collected from different angles, including:

[0026] The video frames of the sampled facial video are intercepted at intervals of a sampling time threshold to obtain a plurality of sampled facial images.

[0027] In a second aspect, an embodiment of the present disclosure provides a facial data processing method, the method comprising:

[0028] Acquire multiple detection facial images collected from different angles, and acquire facial texture map data corresponding to each detection facial image;

[0029] A target texture geometric model is obtained, and facial texture map data corresponding to a plurality of detected facial images are used as input to obtain target facial texture map data output by the target texture geometric model.

[0030] In an implementation of the present disclosure, obtaining facial texture map data corresponding to each detected facial image includes:

[0031] Performing facial reconstruction on each of the multiple facial images to obtain facial geometry data corresponding to each facial image;

[0032] According to each detected facial image and the facial geometry data corresponding to each detected facial image, the facial texture map data corresponding to each detected facial image is obtained.

[0033] In an implementation of the present disclosure, facial texture map data corresponding to a plurality of detected facial images are used as input to obtain target facial texture map data output by a target texture geometric model, including:

[0034] The facial texture map data and facial geometry data corresponding to the plurality of detected facial images are taken as input to obtain the target facial texture map data and target facial geometry data output by the target facial data model.

[0035] In an implementation of the present disclosure, before acquiring a plurality of detection facial images collected from different angles, the method further includes:

[0036] Get the detected face video;

[0037] Get multiple detected facial images collected from different angles, including:

[0038] The video frames of the detected face video are intercepted at intervals of the detection time threshold to obtain multiple detected face images.

[0039] In an implementation of the present disclosure, before taking the facial texture map data corresponding to the plurality of detected facial images as input, the method further includes:

[0040] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0041] The facial texture map data corresponding to multiple detected facial images are used as input, including:

[0042] The facial texture map data with artifact removal data corresponding to the plurality of detected facial images is taken as input.

[0043] In a third aspect, an embodiment of the present disclosure provides a model training device, the device comprising:

[0044] A first image acquisition module is configured to acquire a plurality of sampled facial images collected from different angles, and acquire facial texture map data corresponding to each sampled facial image;

[0045] A first target texture acquisition module is configured to acquire a facial data model, and take facial texture map data corresponding to a plurality of sampled facial images as input to acquire target facial texture map data output by the facial data model;

[0046] An image rendering module is configured to render based on the target facial texture map data to obtain a plurality of target facial images;

[0047] The target model acquisition module is configured to determine the facial data model as the target facial data model in response to matching the multiple target facial images with the multiple sampled facial images.

[0048] In a fourth aspect, an embodiment of the present disclosure provides a facial data processing device, the device comprising:

[0049] A second image acquisition module is configured to acquire a plurality of detection facial images collected from different angles, and acquire facial texture map data corresponding to each detection facial image;

[0050] The second target texture acquisition module is configured to acquire a target texture geometric model and take facial texture map data corresponding to a plurality of detected facial images as input to acquire target facial texture map data output by the target texture geometric model.

[0051] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising a memory and at least one processor; the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by at least one processor to implement the first aspect, any implementation of the first aspect, the second aspect, and any method steps in any implementation of the second aspect.

[0052] In the sixth aspect, a computer-readable storage medium is provided in an embodiment of the present disclosure, on which computer instructions are stored. When the computer instructions are executed by a processor, the method steps of any one of the first aspect, any one of the implementation methods of the first aspect, the second aspect, and any one of the implementation methods of the second aspect are implemented.

[0053] In the seventh aspect, a computer program product is provided in an embodiment of the present disclosure, comprising a computer program / instruction, which, when executed by a processor, implements the method steps of the first aspect, any one of the implementations of the first aspect, the second aspect, and any one of the implementations of the second aspect.

[0054] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0055] According to the technical solution provided by the embodiment of the present disclosure, a plurality of sampled facial images collected from different angles are obtained, and facial texture map data corresponding to each sampled facial image is obtained, wherein, since the facial texture map data corresponding to each sampled facial image is obtained based on a single sampled facial image, the facial texture map data can reflect the texture features of the face at a single angle; a facial data model is obtained, and the facial texture map data corresponding to the plurality of sampled facial images are used as input to obtain target facial texture map data output by the facial data model; rendering is performed based on the target facial texture map data to obtain a plurality of target facial images; and in response to the matching of the plurality of target facial images with the plurality of sampled facial images, the facial data model is determined as a target facial data model. Among them, when multiple target facial images are matched with multiple sampled facial images, it can be understood that the target facial texture map data can accurately reflect the texture features of the face at multiple angles. Therefore, the acquired target facial data model has learned the rules between the facial texture map data acquired based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles. That is, the target facial data model obtained based on the above scheme can obtain facial texture map data that can accurately reflect the texture features of the face at multiple angles, thereby improving the accuracy of the facial texture map data and the reliability of the facial texture map data.

[0056] According to the technical solution provided by the embodiments of the present disclosure, by performing facial reconstruction on each sampled facial image in a plurality of sampled facial images to obtain facial geometry data corresponding to each sampled facial image, and obtaining facial texture map data corresponding to each sampled facial image based on each sampled facial image and the facial geometry data corresponding to each sampled facial image, the accuracy of the acquired facial texture map data can be improved and the amount of calculation required to obtain the facial texture map data can be reduced.

[0057] According to the technical solution provided in the embodiments of the present disclosure, considering that the facial geometry data acquired based on a single image may not be able to reflect the geometric structure of the facial part that is not presented in the single image (for example, when the single image does not include some facial organs, such as ears and eyebrows on one side of the face, etc.), the facial geometry data may not be able to accurately reflect the geometric structure of the face at different angles, and the accuracy is poor, which in turn leads to poor accuracy of the multiple target facial images acquired. By taking the facial texture map data and facial geometry data corresponding to a plurality of sampled facial images as input to obtain the target facial texture map data and target facial geometry data output by the facial data model, and rendering based on the target facial texture map data and target facial geometry data to obtain a plurality of target facial images, it can be ensured that the facial geometry data output by the facial data model can accurately reflect the geometric structure of the face at different angles when the plurality of target facial images are matched with the plurality of sampled facial images. Therefore, the obtained target facial data model has learned the regularity between the facial texture map data obtained based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles, and the regularity between the facial geometry data obtained based on a single facial image and the facial geometry data that can accurately reflect the geometric structure of the face at multiple angles. That is, based on the target facial data model obtained by the above scheme, facial texture map data that can accurately reflect the texture features of the face at multiple angles and facial geometry data of the geometric structure of the face at multiple angles can be obtained, thereby improving the accuracy of the acquired facial texture map data and facial geometry data, and improving the reliability of the facial texture map data and facial geometry data.

[0058] According to the technical solution provided by the embodiment of the present disclosure, by responding to the matching of multiple target facial images with multiple sampled facial images, and the multiple target facial images do not include artifact areas, the facial data model is determined as the target facial data model, thereby ensuring that the target facial texture map data obtained through the target facial data model has a high accuracy.

[0059] According to the technical solution provided by the embodiments of the present disclosure, in response to the facial texture map data including artifact data, the artifact data is removed from the facial texture map data, and the facial texture map data with artifact data removed corresponding to multiple sampled facial images is used as input, so that the accuracy of the facial texture map data input into the facial data model can be improved, thereby improving the efficiency of training the facial data model.

[0060] According to the technical solution provided by the embodiment of the present disclosure, by acquiring a sampled facial video and intercepting the video frames of the sampled facial video at intervals of a sampling time threshold to acquire multiple sampled facial images, the difficulty of acquiring the sampled facial images can be reduced.

[0061] According to the technical solution provided by the embodiment of the present disclosure, multiple detection facial images collected from different angles are obtained, and facial texture map data corresponding to each detection facial image is obtained; and a target texture geometric model is obtained, and the facial texture map data corresponding to the multiple detection facial images are used as input to obtain target facial texture map data output by the target texture geometric model. Among them, since the target facial data model has learned the law between facial texture map data obtained based on a single facial image and facial texture map data that can accurately reflect the texture features of the face at multiple angles, that is, based on the target facial data model, facial texture map data that can accurately reflect the texture features of the face at multiple angles, namely, target facial texture map data, can be obtained, thereby improving the accuracy of the acquired target facial texture map data and improving the reliability of the target facial texture map data.

[0062] According to the technical solution provided by the embodiments of the present disclosure, by performing facial reconstruction on each of the multiple detected facial images to obtain facial geometry data corresponding to each detected facial image, and obtaining facial texture map data corresponding to each detected facial image based on each detected facial image and the facial geometry data corresponding to each detected facial image, the accuracy of the acquired facial texture map data can be improved and the amount of calculation required to obtain the facial texture map data can be reduced.

[0063] According to the technical solution provided by the embodiment of the present disclosure, considering that the facial geometry data obtained based on a single image may not reflect the geometric structure of the facial part that is not presented in the single image (for example, when the single image does not include some facial organs, such as ears and eyebrows on one side of the face, etc.), the facial geometry data may not accurately reflect the geometric structure of the face at different angles, and the accuracy is poor, which may further lead to the accuracy of the facial image obtained based on the target facial texture map data being poor. By taking the facial texture map data and facial geometry data corresponding to multiple detected facial images as input to obtain the target facial texture map data and target facial geometry data output by the facial data model, it is possible to obtain the target facial texture map data that can accurately reflect the texture features of the face at multiple angles and the target facial geometry data of the geometric structure of the face at multiple angles, thereby improving the accuracy of the acquired target facial texture map data and target facial geometry data.

[0064] According to the technical solution provided by the embodiment of the present disclosure, by acquiring a detection face video and intercepting the video frames of the detection face video at intervals of a detection time threshold to acquire multiple detection face images, the difficulty of acquiring the detection face images can be reduced.

[0065] According to the technical solution provided by the embodiments of the present disclosure, by removing artifact data from the facial texture map data in response to the facial texture map data including artifact data, and taking the facial texture map data with artifact data removed corresponding to multiple detected facial images as input, the accuracy of the facial texture map data input into the facial data model can be improved, thereby improving the efficiency of training the facial data model.

[0066] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings. In the accompanying drawings:

[0068] Figure 1 A flowchart of a model training method according to an embodiment of the present disclosure is shown.

[0069] Figure 2 A flowchart of a facial data processing method according to an embodiment of the present disclosure is shown.

[0070] Figure 3 A structural block diagram of a model training device according to an embodiment of the present disclosure is shown.

[0071] Figure 4A structural block diagram of a facial data processing device according to an embodiment of the present disclosure is shown.

[0072] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0073] Figure 6 It is a schematic diagram of the structure of a computer system suitable for implementing the method according to the embodiment of the present disclosure. DETAILED DESCRIPTION

[0074] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0075] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of labels, numbers, steps, behaviors, components, parts, or combinations thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other labels, numbers, steps, behaviors, components, parts, or combinations thereof exist or are added.

[0076] It should also be noted that, in the absence of conflict, the embodiments and labels in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0077] In order to obtain facial texture map data, the inventors of the present disclosure considered the following solutions:

[0078] In the related art, when acquiring facial texture map data of a face, a single image of the face may be acquired first, and then the facial texture map data of the face in the single image may be generated based on the acquired single image.

[0079] Disadvantages of this solution: In the above solution, although facial texture map data can be obtained more conveniently, since the facial texture map data in the above solution is obtained based on a single image, the acquired facial texture map data cannot accurately reflect the facial texture corresponding to the facial part not included in the single image. For example, when the single image is a frontal view image of the face or a near-frontal view image of the face, the facial texture map data obtained based on the single image cannot reliably reflect the texture of the side of the face; in addition, if there is occlusion of the face in the single image, the facial texture map data obtained based on the single image cannot accurately reflect the texture of the occluded part of the face. Therefore, the accuracy and reliability of the facial texture map data obtained by the above solution are poor. Exemplarily, when it is necessary to generate a corresponding virtual image based on an input video, corresponding facial texture map data and a facial three-dimensional model can be obtained based on the video frames in the video. However, since the accuracy and reliability of the facial texture map data obtained based on a single video frame in the video are relatively poor, it can only be applied to the model corresponding to the single video frame to generate the corresponding virtual image, and cannot be applied to the models corresponding to other video frames in the video to generate the corresponding virtual image. Therefore, it is necessary to obtain corresponding facial texture map data and a facial three-dimensional model based on each frame in the video, which results in a large amount of data to be processed when generating the virtual image, thereby increasing the cost.

[0080] Taking into account the shortcomings of the above schemes, the inventors of the present disclosure have proposed a new scheme: by acquiring multiple sampled facial images collected from different angles, and acquiring facial texture map data corresponding to each sampled facial image, wherein, since the facial texture map data corresponding to each sampled facial image is acquired based on a single sampled facial image, the facial texture map data can reflect the texture features of the face at a single angle; acquiring a facial data model, and taking the facial texture map data corresponding to the multiple sampled facial images as input to acquire target facial texture map data output by the facial data model; rendering based on the target facial texture map data to acquire multiple target facial images; and determining the facial data model as a target facial data model in response to matching the multiple target facial images with the multiple sampled facial images. Among them, when multiple target facial images are matched with multiple sampled facial images, it can be understood that the target facial texture map data can accurately reflect the texture features of the face at multiple angles. Therefore, the acquired target facial data model has learned the rules between the facial texture map data acquired based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles. That is, the target facial data model obtained based on the above scheme can obtain facial texture map data that can accurately reflect the texture features of the face at multiple angles, thereby improving the accuracy of the facial texture map data and the reliability of the facial texture map data.

[0081] In order to solve the above problems, the present disclosure proposes a model training method, a facial data processing method, a device, a medium and a product.

[0082] Figure 1 FIG. 1 is a flow chart of a model training method according to an embodiment of the present disclosure. Figure 1 As shown, the model training method includes steps S101-S104.

[0083] In step S101, a plurality of sampled facial images collected from different angles are obtained, and facial texture map data corresponding to each sampled facial image is obtained.

[0084] In step S102, a facial data model is obtained, and facial texture map data corresponding to a plurality of sampled facial images are used as input to obtain target facial texture map data output by the facial data model.

[0085] In step S103, rendering is performed based on the target facial texture map data to obtain a plurality of target facial images.

[0086] In step S104, in response to the matching of the plurality of target facial images with the plurality of sampled facial images, the facial data model is determined as the target facial data model.

[0087] In one embodiment of the present disclosure, the sampled facial image may be understood as an image including at least one face, and the face in the sampled facial image may be understood as the face of the target user. It should be noted that in order to improve the training efficiency of the facial data model, the multiple sampled facial images may only include the face of one target user.

[0088] In one embodiment of the present disclosure, the multiple sampled facial images collected from different angles can be understood as the multiple sampled facial images in which the angles between the direction facing the face and the image collection direction are different. Exemplarily, among the multiple sampled facial images collected from different angles, a certain sampled facial image may include an image corresponding to the left ear of the face, but not an image corresponding to the right ear of the face, and another sampled facial image other than the certain sampled facial image may include an image corresponding to the right ear of the face, but not an image corresponding to the left ear of the face.

[0089] In one embodiment of the present disclosure, obtaining a plurality of sampled facial images captured from different angles can be understood as performing image capture angle detection on the plurality of sampled facial images acquired in advance, and selecting a plurality of sampled facial images captured from different angles from the plurality of sampled facial images acquired in advance according to the detection result. It can also be intercepting a plurality of video frames from a corresponding video, performing image capture angle detection on the plurality of video frames, and selecting a plurality of video frames from the plurality of video frames as the plurality of sampled facial images captured from different angles according to the detection result.

[0090] In one embodiment of the present disclosure, facial texture map data may be understood as image parameters such as color and brightness indicating different positions on the facial surface in a corresponding image.

[0091] In one embodiment of the present disclosure, facial texture map data corresponding to each sampled facial image can be obtained by inputting each sampled facial image into a single facial texture model to obtain facial texture map data output by the single facial texture model; facial texture map data corresponding to each sampled facial image can also be obtained by parsing each sampled facial image based on a pre-acquired algorithm to obtain facial texture map data corresponding to each sampled facial image, etc.

[0092] In one embodiment of the present disclosure, the facial data model can be understood as being pre-acquired or acquired from other devices or systems. The facial data model can be an encoder-decoder model, a neural network (NN) model, a convolutional neural network (CNN) model, or a long short-term memory network (LSTM) model, etc., and the present disclosure does not specifically limit this.

[0093] In one embodiment of the present disclosure, rendering is performed based on the target facial texture map data to obtain multiple target facial images. This can be understood as determining the pixel values ​​of the pixel points at different positions on the surface of the facial model corresponding to the target facial texture map data, and then projecting them onto the imaging plane of the virtual camera according to the three-dimensional spatial positions of the different pixel points, that is, obtaining the correspondence between the pixel points at different positions on the surface of the facial model and the pixel points on the virtual imaging plane, and according to the correspondence, copying the pixel values ​​of the pixel points at different positions on the surface of the facial model as the pixel values ​​of the pixel points at the corresponding positions of the virtual imaging plane, thereby obtaining the corresponding target facial image.

[0094] In one embodiment of the present disclosure, a plurality of target facial images are matched with a plurality of sampled facial images, which can be understood as the similarity between the sampled facial images captured from the corresponding angle in the plurality of sampled facial images and the target facial images whose virtual imaging angle is the corresponding angle in the plurality of target facial images is greater than or equal to the similarity threshold, wherein the imaging angle can be understood as the angle between the face and the virtual image acquisition direction. It should be noted that when the plurality of target facial images do not match the plurality of sampled facial images, the parameters of the facial data model can be adjusted, and steps S101 to S104 are executed cyclically to achieve the purpose of training the facial data model. When the plurality of target facial images match the plurality of sampled facial images, it can be understood that the facial data model meets the training requirements.

[0095] According to the technical solution provided by the embodiment of the present disclosure, a plurality of sampled facial images collected from different angles are obtained, and facial texture map data corresponding to each sampled facial image is obtained, wherein, since the facial texture map data corresponding to each sampled facial image is obtained based on a single sampled facial image, the facial texture map data can reflect the texture features of the face at a single angle; a facial data model is obtained, and the facial texture map data corresponding to the plurality of sampled facial images are used as input to obtain target facial texture map data output by the facial data model; rendering is performed based on the target facial texture map data to obtain a plurality of target facial images; and in response to the matching of the plurality of target facial images with the plurality of sampled facial images, the facial data model is determined as a target facial data model. Among them, when multiple target facial images are matched with multiple sampled facial images, it can be understood that the target facial texture map data can accurately reflect the texture features of the face at multiple angles. Therefore, the acquired target facial data model has learned the rules between the facial texture map data acquired based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles. That is, the target facial data model obtained based on the above scheme can obtain facial texture map data that can accurately reflect the texture features of the face at multiple angles, thereby improving the accuracy of the facial texture map data and the reliability of the facial texture map data.

[0096] In an implementation of the present disclosure, in step S101, facial texture map data corresponding to each sampled facial image is obtained, which can be achieved by the following steps:

[0097] Performing facial reconstruction on each of the multiple sampled facial images to obtain facial geometry data corresponding to each sampled facial image;

[0098] According to each sampled facial image and the facial geometry data corresponding to each sampled facial image, the facial texture map data corresponding to each sampled facial image is obtained.

[0099] In one embodiment of the present disclosure, the facial geometry data corresponding to the sampled facial image can be understood as being used to indicate the three-dimensional structure of the face in the corresponding sampled facial image, and can also be understood as being used to indicate the spatial position coordinates of at least one sampling point on the facial surface in the corresponding sampled facial image.

[0100] In one embodiment of the present disclosure, facial reconstruction is performed on each of the multiple sampled facial images to obtain facial geometry data corresponding to each sampled facial image. This can be understood as parsing each sampled facial image based on a pre-acquired facial reconstruction algorithm to obtain facial geometry data corresponding to each sampled facial image; it can also be understood as inputting each sampled facial image into a pre-acquired facial geometry model to obtain facial geometry data corresponding to each sampled facial image. The facial geometry model can be a deep convolutional neural network model (DCNN), or a three-dimensional deformable model (3D Morphable Model, 3DMM), etc.

[0101] In one embodiment of the present disclosure, facial texture map data corresponding to each sampled facial image is obtained based on each sampled facial image and the facial geometry data corresponding to each sampled facial image. This can be understood as substituting each sampled facial image and the facial geometry data corresponding to each sampled facial image into a calculation based on a pre-acquired geometry texture algorithm to obtain facial texture map data corresponding to each sampled facial image. Exemplarily, multiple pixel points in the area where the face is located in the sampled facial image can be mapped to a three-dimensional space where the facial geometry is located determined according to the facial geometry data, that is, the correspondence between the multiple pixel points in the area where the face is located in the sampled facial image and the pixel points on the surface of the facial geometry determined according to the facial geometry data is determined, and the pixel values ​​of the multiple pixel points on the surface of the facial geometry can be determined based on the correspondence, and then the facial texture map data is obtained based on the pixel values ​​of the multiple pixel points on the surface of the facial geometry.

[0102] According to the technical solution provided by the embodiments of the present disclosure, by performing facial reconstruction on each sampled facial image in a plurality of sampled facial images to obtain facial geometry data corresponding to each sampled facial image, and obtaining facial texture map data corresponding to each sampled facial image based on each sampled facial image and the facial geometry data corresponding to each sampled facial image, the accuracy of the acquired facial texture map data can be improved and the amount of calculation required to obtain the facial texture map data can be reduced.

[0103] In an implementation of the present disclosure, in step S102, the facial texture map data corresponding to the plurality of sampled facial images are used as input to obtain the target facial texture map data output by the facial data model, which can be implemented by the following steps:

[0104] Taking facial texture map data and facial geometry data corresponding to a plurality of sampled facial images as input, to obtain target facial texture map data and target facial geometry data output by a facial data model;

[0105] In step S103, rendering is performed based on the target facial texture map data to obtain multiple target facial images, which can be achieved by the following steps:

[0106] Rendering is performed based on the target facial texture map data and the target facial geometry data to obtain a plurality of target facial images.

[0107] According to the technical solution provided in the embodiments of the present disclosure, considering that the facial geometry data acquired based on a single image may not be able to reflect the geometric structure of the facial part that is not presented in the single image (for example, when the single image does not include some facial organs, such as ears and eyebrows on one side of the face, etc.), the facial geometry data may not be able to accurately reflect the geometric structure of the face at different angles, and the accuracy is poor, which in turn leads to poor accuracy of the multiple target facial images acquired. By taking the facial texture map data and facial geometry data corresponding to a plurality of sampled facial images as input to obtain the target facial texture map data and target facial geometry data output by the facial data model, and rendering based on the target facial texture map data and target facial geometry data to obtain a plurality of target facial images, it can be ensured that the facial geometry data output by the facial data model can accurately reflect the geometric structure of the face at different angles when the plurality of target facial images are matched with the plurality of sampled facial images. Therefore, the obtained target facial data model has learned the regularity between the facial texture map data obtained based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles, and the regularity between the facial geometry data obtained based on a single facial image and the facial geometry data that can accurately reflect the geometric structure of the face at multiple angles. That is, based on the target facial data model obtained by the above scheme, facial texture map data that can accurately reflect the texture features of the face at multiple angles and facial geometry data of the geometric structure of the face at multiple angles can be obtained, thereby improving the accuracy of the acquired facial texture map data and facial geometry data, and improving the reliability of the facial texture map data and facial geometry data.

[0108] In an implementation of the present disclosure, in step S104, in response to matching the multiple target facial images with the multiple sampled facial images, determining the facial data model as the target facial data model can be implemented by the following steps:

[0109] In response to the plurality of target facial images matching the plurality of sampled facial images and the plurality of target facial images not including artifact regions, the facial data model is determined as a target facial data model.

[0110] In an implementation of the present disclosure, an artifact region may be understood as a region including repeated images; or may be understood as a region including multiple facial biometric image features whose distance is less than a threshold value of the artifact distance. Among them, facial biometric image features may be understood as organs, parts of organs, skin manifestations (e.g., acne, moles), etc. When a target facial image includes an artifact region, it may be understood that the target facial image has a mapping error due to the low accuracy of the acquired target facial texture mapping data.

[0111] According to the technical solution provided by the embodiment of the present disclosure, by responding to the matching of multiple target facial images with multiple sampled facial images, and the multiple target facial images do not include artifact areas, the facial data model is determined as the target facial data model, thereby ensuring that the target facial texture map data obtained through the target facial data model has a high accuracy.

[0112] In an implementation of the present disclosure, in step S102, before taking the facial texture map data corresponding to the plurality of sampled facial images as input, the method further includes the following steps:

[0113] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0114] In step S102, the facial texture map data corresponding to the plurality of sampled facial images are taken as input, which can be achieved by the following steps:

[0115] The facial texture map data of the artifact-removed data corresponding to the multiple sampled facial images is taken as input.

[0116] In an implementation of the present disclosure, artifact data can be understood as, among the image parameters of different positions on the facial surface obtained according to the facial texture map data, if the image parameters of two areas whose facial surface distance is less than the artifact distance threshold have a high degree of similarity, then the data in the facial texture map data used to obtain the image parameters of the two areas can be understood as artifact data. Among them, the high degree of similarity of the image parameters of the two areas can be understood as the high degree of similarity of the image parameters of the two areas themselves, and can also be understood as the image parameters of the two areas reflect similar or identical facial biological image features. When the facial texture map data includes artifact data, it can be understood that the accuracy of the facial texture map data is low and the reliability of the artifact data is poor.

[0117] According to the technical solution provided by the embodiments of the present disclosure, in response to the facial texture map data including artifact data, the artifact data is removed from the facial texture map data, and the facial texture map data with artifact data removed corresponding to multiple sampled facial images is used as input, so that the accuracy of the facial texture map data input into the facial data model can be improved, thereby improving the efficiency of training the facial data model.

[0118] In an implementation of the present disclosure, before acquiring a plurality of sampled facial images collected from different angles in step S101, the method further includes the following steps:

[0119] Get sampled facial videos;

[0120] In step S101, a plurality of sampled facial images collected from different angles are obtained, which can be achieved by the following steps:

[0121] The video frames of the sampled facial video are intercepted at intervals of a sampling time threshold to obtain a plurality of sampled facial images.

[0122] In an implementation of the present disclosure, the sampled facial video may be understood as a video including at least one face, and the face in the sampled facial video may be understood as the face of the target user. It should be noted that in order to improve the training efficiency of the facial data model, the sampled facial video may only include the face of one target user.

[0123] In an implementation of the present disclosure, obtaining a sampled facial video may be understood as reading and writing a sampled facial video obtained in advance, or may be understood as receiving a sampled facial video sent by an image acquisition device.

[0124] In an implementation of the present disclosure, the sampling time threshold may be understood as being obtained in advance or from other devices or systems. Intercepting video frames of the sampled facial video at intervals of the sampling time threshold may be understood as determining multiple interception moments within the acquisition time interval of the sampled facial video at intervals of the sampling time threshold, and intercepting video frames whose acquisition time is the interception moment in the sampled facial video.

[0125] In an implementation of the present disclosure, considering that it is difficult for a person to maintain the angle between the face and the video acquisition direction at the same angle for a long time during the process of collecting the sampled facial video, it can be considered that the video frames at different times in the sampled facial video are collected from different angles. It should be noted that when the angle difference between the image acquisition angles corresponding to the multiple sampled facial images is large, the multiple sampled facial images can reflect the texture features and facial geometry of the face in more directions. Therefore, during the process of collecting the sampled facial video, prompt information for prompting the user to actively turn the head can be displayed; or, the sampled facial video can be detected, and when it is determined according to the detection result that the rotation angle of the face in the sampled facial video is greater than or equal to the facial rotation angle threshold, the video frames of the sampled facial video are intercepted at intervals of the sampling time threshold.

[0126] According to the technical solution provided by the embodiment of the present disclosure, by acquiring a sampled facial video and intercepting the video frames of the sampled facial video at intervals of a sampling time threshold to acquire multiple sampled facial images, the difficulty of acquiring the sampled facial images can be reduced.

[0127] Figure 2 FIG. 1 is a flowchart of a facial data processing method according to an embodiment of the present disclosure. Figure 2 As shown, the facial data processing method includes steps S201-S202.

[0128] In step S201, a plurality of detection facial images collected from different angles are obtained, and facial texture map data corresponding to each detection facial image is obtained.

[0129] In step S202, a target texture geometric model is obtained, and facial texture map data corresponding to a plurality of detected facial images are taken as input to obtain target facial texture map data output by the target texture geometric model.

[0130] In one embodiment of the present disclosure, the detected facial image may be understood as an image including at least one face, and the face in the detected facial image may be understood as the face of the target user. It should be noted that in order to improve the training efficiency of the facial data model, the multiple detected facial images may only include the face of one target user.

[0131] In one embodiment of the present disclosure, the multiple facial images for detection collected from different angles can be understood as the multiple facial images for detection, in which the angles between the direction in which the face is facing and the image collection direction are different. Exemplarily, among the multiple facial images for detection collected from different angles, a certain facial image for detection may include an image corresponding to the left ear of the face, but not an image corresponding to the right ear of the face, and another facial image for detection other than the certain facial image for detection may include an image corresponding to the right ear of the face, but not an image corresponding to the left ear of the face.

[0132] In one embodiment of the present disclosure, obtaining a plurality of detection facial images captured from different angles can be understood as performing image capture angle detection on the plurality of detection facial images acquired in advance, and selecting a plurality of detection facial images captured from different angles from the plurality of detection facial images acquired in advance according to the detection results. It can also be capturing a plurality of video frames from a corresponding video, performing image capture angle detection on the plurality of video frames, and selecting a plurality of video frames from the plurality of video frames as the plurality of detection facial images captured from different angles according to the detection results.

[0133] In one embodiment of the present disclosure, facial texture map data may be understood as image parameters such as color and brightness indicating different positions on the facial surface in a corresponding image.

[0134] In one embodiment of the present disclosure, facial texture map data corresponding to each detected facial image is obtained by inputting each detected facial image into a single facial texture model to obtain facial texture map data output by the single facial texture model; facial texture map data corresponding to each detected facial image is obtained by parsing each detected facial image based on a pre-acquired algorithm to obtain facial texture map data corresponding to each detected facial image, etc.

[0135] In one embodiment of the present disclosure, the target texture geometric model can be understood as being obtained in advance or obtained from other devices or systems. The target texture geometric model can be understood as being obtained based on any of the above-mentioned model training methods.

[0136] According to the technical solution provided by the embodiment of the present disclosure, multiple detection facial images collected from different angles are obtained, and facial texture map data corresponding to each detection facial image is obtained; and a target texture geometric model is obtained, and the facial texture map data corresponding to the multiple detection facial images are used as input to obtain target facial texture map data output by the target texture geometric model. Among them, since the target facial data model has learned the law between facial texture map data obtained based on a single facial image and facial texture map data that can accurately reflect the texture features of the face at multiple angles, that is, based on the target facial data model, facial texture map data that can accurately reflect the texture features of the face at multiple angles, namely, target facial texture map data, can be obtained, thereby improving the accuracy of the acquired target facial texture map data and improving the reliability of the target facial texture map data.

[0137] In an implementation of the present disclosure, in step S201, obtaining facial texture map data corresponding to each detected facial image can be achieved by the following steps:

[0138] Performing facial reconstruction on each of the multiple facial images to obtain facial geometry data corresponding to each facial image;

[0139] According to each detected facial image and the facial geometry data corresponding to each detected facial image, the facial texture map data corresponding to each detected facial image is obtained.

[0140] In one embodiment of the present disclosure, the facial geometry data corresponding to the detected facial image can be understood as being used to indicate the three-dimensional structure of the face in the corresponding detected facial image, and can also be understood as being used to indicate the spatial position coordinates of at least one detection point on the facial surface in the corresponding detected facial image.

[0141] In one embodiment of the present disclosure, facial reconstruction is performed on each of the multiple facial images to obtain facial geometry data corresponding to each facial image. This can be understood as parsing each facial image based on a pre-acquired facial reconstruction algorithm to obtain facial geometry data corresponding to each facial image; it can also be understood as inputting each facial image into a pre-acquired facial geometry model to obtain facial geometry data corresponding to each facial image. The facial geometry model can be a deep convolutional neural network model (DCNN), or a three-dimensional deformable model (3D Morphable Model, 3DMM), etc.

[0142] In one embodiment of the present disclosure, facial texture map data corresponding to each detected facial image is obtained based on each detected facial image and the facial geometry data corresponding to each detected facial image. This can be understood as substituting each detected facial image and the facial geometry data corresponding to each detected facial image into a calculation based on a pre-acquired geometry texture algorithm to obtain facial texture map data corresponding to each detected facial image. Exemplarily, multiple pixel points in the area where the face is located in the detected facial image can be mapped to a three-dimensional space where the facial geometry is located determined according to the facial geometry data, that is, the correspondence between multiple pixel points in the area where the face is located in the detected facial image and the pixel points on the surface of the facial geometry determined according to the facial geometry data is determined, and the pixel values ​​of the multiple pixel points on the surface of the facial geometry can be determined based on the correspondence, and then the facial texture map data is obtained based on the pixel values ​​of the multiple pixel points on the surface of the facial geometry.

[0143] According to the technical solution provided by the embodiments of the present disclosure, by performing facial reconstruction on each of the multiple detected facial images to obtain facial geometry data corresponding to each detected facial image, and obtaining facial texture map data corresponding to each detected facial image based on each detected facial image and the facial geometry data corresponding to each detected facial image, the accuracy of the acquired facial texture map data can be improved and the amount of calculation required to obtain the facial texture map data can be reduced.

[0144] In an implementation of the present disclosure, in step S202, the facial texture map data corresponding to the plurality of detected facial images are used as input to obtain the target facial texture map data output by the target texture geometric model, which can be implemented by the following steps:

[0145] The facial texture map data and facial geometry data corresponding to the plurality of detected facial images are taken as input to obtain the target facial texture map data and target facial geometry data output by the target facial data model.

[0146] According to the technical solution provided by the embodiment of the present disclosure, considering that the facial geometry data obtained based on a single image may not reflect the geometric structure of the facial part that is not presented in the single image (for example, when the single image does not include some facial organs, such as ears and eyebrows on one side of the face, etc.), the facial geometry data may not accurately reflect the geometric structure of the face at different angles, and the accuracy is poor, which may further lead to the accuracy of the facial image obtained based on the target facial texture map data being poor. By taking the facial texture map data and facial geometry data corresponding to multiple detected facial images as input to obtain the target facial texture map data and target facial geometry data output by the facial data model, it is possible to obtain the target facial texture map data that can accurately reflect the texture features of the face at multiple angles and the target facial geometry data of the geometric structure of the face at multiple angles, thereby improving the accuracy of the acquired target facial texture map data and target facial geometry data.

[0147] In an implementation of the present disclosure, before acquiring a plurality of detection facial images collected from different angles in step S201, the method further includes the following steps:

[0148] Get the detected face video;

[0149] In step S201, a plurality of detected facial images collected from different angles are obtained, which can be achieved by the following steps:

[0150] The video frames of the detected face video are intercepted at intervals of the detection time threshold to obtain multiple detected face images.

[0151] In an implementation of the present disclosure, the face detection video can be understood as a video including at least one face, and the face in the face detection video can be understood as the face of the target user. It should be noted that in order to improve the training efficiency of the facial data model, the face detection video can only include the face of one target user.

[0152] In an implementation of the present disclosure, acquiring the detected facial video may be understood as reading and writing the detected facial video acquired in advance, or may be understood as receiving the detected facial video sent by the image acquisition device.

[0153] In an implementation of the present disclosure, the detection time threshold may be understood as being obtained in advance or from other devices or systems. Intercepting the video frames of the detected face video at intervals of the detection time threshold may be understood as determining multiple interception moments within the acquisition time interval of the detected face video at intervals of the detection time threshold, and intercepting the video frames of the detected face video whose acquisition time is the interception moment.

[0154] In an implementation of the present disclosure, considering that it is difficult for a person to maintain the angle between the face and the video acquisition direction at the same angle for a long time during the process of collecting the detection face video, it can be considered that the video frames at different times in the detection face video are collected from different angles. It should be noted that when the angle difference between the image acquisition angles corresponding to the multiple detection face images is large, the multiple detection face images can reflect the texture features and facial geometry of the face in more directions. Therefore, during the process of collecting the detection face video, prompt information for prompting the user to actively turn the head can be displayed; or, the detection face video can be detected, and when it is determined according to the detection result that the rotation angle of the face in the detection face video is greater than or equal to the face rotation angle threshold, the video frames of the detection face video are intercepted at intervals of the detection time threshold.

[0155] According to the technical solution provided by the embodiment of the present disclosure, by acquiring a detection face video and intercepting the video frames of the detection face video at intervals of a detection time threshold to acquire multiple detection face images, the difficulty of acquiring the detection face images can be reduced.

[0156] In an implementation of the present disclosure, in step S202, before taking the facial texture map data corresponding to the plurality of detected facial images as input, the method further includes the following steps:

[0157] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0158] In step S201, the facial texture map data corresponding to the plurality of detected facial images are taken as input, which can be achieved by the following steps:

[0159] The facial texture map data with artifact removal data corresponding to the plurality of detected facial images is taken as input.

[0160] In an implementation of the present disclosure, artifact data can be understood as, among the image parameters of different positions on the facial surface obtained according to the facial texture map data, if the image parameters of two areas whose facial surface distance is less than the artifact distance threshold have a high degree of similarity, then the data in the facial texture map data used to obtain the image parameters of the two areas can be understood as artifact data. Among them, the high degree of similarity of the image parameters of the two areas can be understood as the high degree of similarity of the image parameters of the two areas themselves, and can also be understood as the image parameters of the two areas reflect similar or identical facial biological image features. When the facial texture map data includes artifact data, it can be understood that the accuracy of the facial texture map data is low and the reliability of the artifact data is poor.

[0161] According to the technical solution provided by the embodiments of the present disclosure, by removing artifact data from the facial texture map data in response to the facial texture map data including artifact data, and taking the facial texture map data with artifact data removed corresponding to multiple detected facial images as input, the accuracy of the facial texture map data input into the facial data model can be improved, thereby improving the efficiency of training the facial data model.

[0162] The following reference Figure 3 A model training device according to an embodiment of the present disclosure is described. Figure 3 A structural block diagram of a model training device according to an embodiment of the present disclosure is shown.

[0163] like Figure 3 As shown, the model training device 100 includes:

[0164] The first image acquisition module 101 is configured to acquire multiple sampled facial images collected from different angles, and acquire facial texture map data corresponding to each sampled facial image;

[0165] The first target texture acquisition module 102 is configured to acquire a facial data model, and take facial texture map data corresponding to a plurality of sampled facial images as input to acquire target facial texture map data output by the facial data model;

[0166] The image rendering module 103 is configured to perform rendering based on the target facial texture map data to obtain a plurality of target facial images;

[0167] The target model acquisition module 104 is configured to determine the facial data model as the target facial data model in response to matching the multiple target facial images with the multiple sampled facial images.

[0168] According to the technical solution provided by the embodiment of the present disclosure, a plurality of sampled facial images collected from different angles are obtained, and facial texture map data corresponding to each sampled facial image is obtained, wherein, since the facial texture map data corresponding to each sampled facial image is obtained based on a single sampled facial image, the facial texture map data can reflect the texture features of the face at a single angle; a facial data model is obtained, and the facial texture map data corresponding to the plurality of sampled facial images are used as input to obtain target facial texture map data output by the facial data model; rendering is performed based on the target facial texture map data to obtain a plurality of target facial images; and in response to the matching of the plurality of target facial images with the plurality of sampled facial images, the facial data model is determined as a target facial data model. Among them, when multiple target facial images are matched with multiple sampled facial images, it can be understood that the target facial texture map data can accurately reflect the texture features of the face at multiple angles. Therefore, the acquired target facial data model has learned the rules between the facial texture map data acquired based on a single facial image and the facial texture map data that can accurately reflect the texture features of the face at multiple angles. That is, the target facial data model obtained based on the above scheme can obtain facial texture map data that can accurately reflect the texture features of the face at multiple angles, thereby improving the accuracy of the facial texture map data and the reliability of the facial texture map data.

[0169] Those skilled in the art will understand that referring to Figure 3 The technical solution described can be compared with the above Figure 1 The corresponding embodiments are combined to have Figure 1 The technical effects achieved by the corresponding embodiments. For details, please refer to Figure 1 The specific contents of the description of the corresponding embodiments will not be repeated here.

[0170] The following reference Figure 4 A facial data processing device according to an embodiment of the present disclosure is described. Figure 4 A structural block diagram of a facial data processing device according to an embodiment of the present disclosure is shown.

[0171] like Figure 5 As shown, the facial data processing device 200 includes:

[0172] The second image acquisition module 201 is configured to acquire multiple detection facial images collected from different angles, and acquire facial texture map data corresponding to each detection facial image;

[0173] The second target texture acquisition module 202 is configured to acquire a target texture geometric model and take facial texture map data corresponding to a plurality of detected facial images as input to acquire target facial texture map data output by the target texture geometric model.

[0174] According to the technical solution provided by the embodiment of the present disclosure, multiple detection facial images collected from different angles are obtained, and facial texture map data corresponding to each detection facial image is obtained; and a target texture geometric model is obtained, and the facial texture map data corresponding to the multiple detection facial images are used as input to obtain target facial texture map data output by the target texture geometric model. Among them, since the target facial data model has learned the law between facial texture map data obtained based on a single facial image and facial texture map data that can accurately reflect the texture features of the face at multiple angles, that is, based on the target facial data model, facial texture map data that can accurately reflect the texture features of the face at multiple angles, namely, target facial texture map data, can be obtained, thereby improving the accuracy of the acquired target facial texture map data and improving the reliability of the target facial texture map data.

[0175] Those skilled in the art will understand that referring to Figure 4 The technical solution described can be compared with the above Figure 2 The corresponding embodiments are combined to have Figure 2 The technical effects achieved by the corresponding embodiments. For details, please refer to Figure 2 The specific contents of the description of the corresponding embodiments will not be repeated here.

[0176] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0177] The present disclosure also provides an electronic device, such as Figure 5 As shown, the electronic device 300 includes at least one processor 301. And a memory 302 that is communicatively connected to the at least one processor 301. The memory 302 stores instructions that can be executed by the at least one processor 301, and the instructions can be executed by the at least one processor 301 to implement the following steps:

[0178] In a first aspect, an embodiment of the present disclosure provides a model training method, the method comprising:

[0179] Acquire multiple sampled facial images collected from different angles, and acquire facial texture map data corresponding to each sampled facial image;

[0180] Acquire a facial data model, and use facial texture map data corresponding to a plurality of sampled facial images as input to acquire target facial texture map data output by the facial data model;

[0181] Rendering is performed based on the target facial texture map data to obtain multiple target facial images;

[0182] In response to the plurality of target facial images matching the plurality of sample facial images, a facial data model is determined as a target facial data model.

[0183] In an implementation of the present disclosure, obtaining facial texture map data corresponding to each sampled facial image includes:

[0184] Performing facial reconstruction on each of the multiple sampled facial images to obtain facial geometry data corresponding to each sampled facial image;

[0185] According to each sampled facial image and the facial geometry data corresponding to each sampled facial image, the facial texture map data corresponding to each sampled facial image is obtained.

[0186] In an implementation of the present disclosure, facial texture map data corresponding to a plurality of sampled facial images are used as input to obtain target facial texture map data output by a facial data model, including:

[0187] Taking facial texture map data and facial geometry data corresponding to a plurality of sampled facial images as input, to obtain target facial texture map data and target facial geometry data output by a facial data model;

[0188] Rendering is performed based on the target facial texture map data to obtain multiple target facial images, including:

[0189] Rendering is performed based on the target facial texture map data and the target facial geometry data to obtain a plurality of target facial images.

[0190] In an implementation of the present disclosure, in response to matching a plurality of target facial images with a plurality of sampled facial images, determining a facial data model as a target facial data model includes:

[0191] In response to the plurality of target facial images matching the plurality of sampled facial images and the plurality of target facial images not including artifact regions, the facial data model is determined as a target facial data model.

[0192] In an implementation of the present disclosure, before taking the facial texture map data corresponding to the plurality of sampled facial images as input, the method further includes:

[0193] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0194] The facial texture map data corresponding to multiple sampled facial images are used as input, including:

[0195] The facial texture map data of the artifact-removed data corresponding to the multiple sampled facial images is taken as input.

[0196] In an implementation of the present disclosure, before acquiring a plurality of sampled facial images collected from different angles, the method further includes:

[0197] Get sampled facial videos;

[0198] Get multiple sample facial images collected from different angles, including:

[0199] The video frames of the sampled facial video are intercepted at intervals of a sampling time threshold to obtain a plurality of sampled facial images.

[0200] In a second aspect, an embodiment of the present disclosure provides a facial data processing method, the method comprising:

[0201] Acquire multiple detection facial images collected from different angles, and acquire facial texture map data corresponding to each detection facial image;

[0202] A target texture geometric model is obtained, and facial texture map data corresponding to a plurality of detected facial images are used as input to obtain target facial texture map data output by the target texture geometric model.

[0203] In an implementation of the present disclosure, obtaining facial texture map data corresponding to each detected facial image includes:

[0204] Performing facial reconstruction on each of the multiple facial images to obtain facial geometry data corresponding to each facial image;

[0205] According to each detected facial image and the facial geometry data corresponding to each detected facial image, the facial texture map data corresponding to each detected facial image is obtained.

[0206] In an implementation of the present disclosure, facial texture map data corresponding to a plurality of detected facial images are used as input to obtain target facial texture map data output by a target texture geometric model, including:

[0207] The facial texture map data and facial geometry data corresponding to the plurality of detected facial images are taken as input to obtain the target facial texture map data and target facial geometry data output by the target facial data model.

[0208] In an implementation of the present disclosure, before acquiring a plurality of detection facial images collected from different angles, the method further includes:

[0209] Get the detected face video;

[0210] Get multiple detected facial images collected from different angles, including:

[0211] The video frames of the detected face video are intercepted at intervals of the detection time threshold to obtain multiple detected face images.

[0212] In an implementation of the present disclosure, before taking the facial texture map data corresponding to the plurality of detected facial images as input, the method further includes:

[0213] In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data;

[0214] The facial texture map data corresponding to multiple detected facial images are used as input, including:

[0215] The facial texture map data with artifact removal data corresponding to the plurality of detected facial images is taken as input.

[0216] like Figure 6 As shown, the computer system 400 includes a processing unit 401, which can perform various processes in the embodiments shown in the above figures according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 to the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0217] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc. An output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc. A storage section 408 including a hard disk, etc. And a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that the computer program read therefrom is installed into the storage section 408 as needed. Among them, the processing unit 401 can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0218] In particular, according to an embodiment of the present disclosure, the method described above with reference to the accompanying drawings can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a readable medium thereof, and the computer program includes a program code for executing the method in the accompanying drawings. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 409, and / or installed from a removable medium 411. For example, an embodiment of the present disclosure includes a readable storage medium on which computer instructions are stored, and when the computer instructions are executed by a processor, the program code for executing the method in the accompanying drawings is implemented.

[0219] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the road map or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a different order from the order marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0220] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.

[0221] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the node in the above-mentioned embodiment. It may also be a computer-readable storage medium that exists independently and is not assembled into a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.

[0222] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other.

Claims

1. A model training method, wherein: The method comprises: Acquire multiple sampled facial images collected from different angles, and acquire facial texture map data corresponding to each sampled facial image; Acquire a facial data model, and use the facial texture map data corresponding to the plurality of sampled facial images as input to acquire target facial texture map data output by the facial data model, wherein the target facial texture map data is used to indicate texture features of the face at multiple angles; Rendering is performed based on the target facial texture map data to obtain a plurality of target facial images; In response to the plurality of target facial images matching the plurality of sampled facial images, determining the facial data model as a target facial data model, the target facial data model having learned a regularity between facial texture map data acquired based on a single facial image and facial texture map data reflecting texture features of a face at multiple angles; The matching of the plurality of target facial images with the plurality of sampled facial images comprises: A similarity between a sampled facial image acquired from a corresponding angle in the plurality of sampled facial images and a target facial image whose virtual imaging angle is the corresponding angle in the plurality of target facial images is greater than or equal to a similarity threshold.

2. The model training method according to claim 1, wherein: The step of obtaining facial texture map data corresponding to each sampled facial image includes: Performing facial reconstruction on each of the plurality of sampled facial images to obtain facial geometry data corresponding to each sampled facial image; According to each sampled facial image and the facial geometry data corresponding to each sampled facial image, the facial texture map data corresponding to each sampled facial image is obtained.

3. The model training method according to claim 2, wherein: The step of taking the facial texture map data corresponding to the plurality of sampled facial images as input to obtain the target facial texture map data output by the facial data model comprises: Taking the facial texture map data and facial geometry data corresponding to the plurality of sampled facial images as input, to obtain the target facial texture map data and target facial geometry data output by the facial data model; The step of performing rendering based on the target facial texture map data to obtain a plurality of target facial images includes: Rendering is performed based on the target facial texture map data and the target facial geometry data to obtain the multiple target facial images.

4. The model training method according to any one of claims 1 to 3, wherein: In response to the plurality of target facial images matching the plurality of sampled facial images, determining the facial data model as a target facial data model comprises: In response to the plurality of target facial images matching the plurality of sampled facial images and the plurality of target facial images not including artifact regions, the facial data model is determined as the target facial data model.

5. The model training method according to any one of claims 1 to 3, wherein: Before taking the facial texture map data corresponding to the plurality of sampled facial images as input, the method further comprises: In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data; The step of taking the facial texture map data corresponding to the plurality of sampled facial images as input comprises: The facial texture map data with artifacts removed corresponding to the plurality of sampled facial images is taken as input.

6. The model training method according to any one of claims 1 to 3, wherein: Before acquiring a plurality of sampled facial images collected from different angles, the method further includes: Get sampled facial videos; The step of acquiring a plurality of sampled facial images collected from different angles comprises: The video frames of the sampled facial video are intercepted at intervals of a sampling time threshold to obtain the multiple sampled facial images.

7. A facial data processing method, wherein: The method comprises: Acquire multiple detection facial images collected from different angles, and acquire facial texture map data corresponding to each detection facial image; Obtain a target facial data model trained according to claim 1, and use facial texture map data corresponding to the multiple detected facial images as input to obtain target facial texture map data output by the target facial data model, wherein the target facial texture map data is used to indicate texture features of the face at multiple angles.

8. The facial data processing method according to claim 7, wherein: The step of obtaining facial texture map data corresponding to each detected facial image includes: Performing facial reconstruction on each of the plurality of detected facial images to obtain facial geometry data corresponding to each detected facial image; According to each detected facial image and the facial geometry data corresponding to each detected facial image, the facial texture map data corresponding to each detected facial image is obtained.

9. The facial data processing method according to claim 8, wherein: The step of taking the facial texture map data corresponding to the plurality of detected facial images as input to obtain the target facial texture map data output by the target facial data model comprises: The facial texture map data and facial geometry data corresponding to the plurality of detected facial images are used as input to obtain target facial texture map data and target facial geometry data output by the target facial data model.

10. The facial data processing method according to any one of claims 7 to 9, wherein: Before acquiring a plurality of detection facial images collected from different angles, the method further includes: Get the detected face video; The step of acquiring a plurality of detected facial images collected from different angles comprises: The video frames of the detected face video are intercepted at intervals of a detection time threshold to obtain the multiple detected face images.

11. The facial data processing method according to any one of claims 7 to 9, wherein: Before taking the facial texture map data corresponding to the plurality of detected facial images as input, the method further comprises: In response to the facial texture map data including artifact data, removing the artifact data from the facial texture map data; The step of taking the facial texture map data corresponding to the plurality of detected facial images as input comprises: The facial texture map data with artifacts removed corresponding to the plurality of detected facial images is used as input.

12. An electronic device comprising a memory and at least one processor; wherein: The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the at least one processor to implement the method steps described in any one of claims 1-11.

13. A computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed by a processor, implement the method steps described in any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the method steps described in any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Texture generation model training method, and image processing method and device

    CN112802075A

  • Face texture feature extraction method and device, 3D face reconstruction method and device and storage medium

    CN113111861A