Face three-dimensional key point extraction method and device, model creation method and system

By directly mapping 2D facial key points to structured point cloud data, and combining PFLD networks and rigid body transformation matrices, the problem of complex 3D facial key point extraction in existing technologies is solved, achieving efficient and accurate 3D facial model creation.

CN115830663BActive Publication Date: 2025-12-19SHENZHEN ANHUA OPTOELECTRONICS TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210590814.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-12-19
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Existing technologies for 3D face reconstruction require the creation of a pixel lookup table between 3D depth images and 2D color images, resulting in complex and inefficient extraction of 3D facial key points.

Method used

By acquiring two-dimensional image data and point cloud data of the face, the image is stored in a structured manner. The two-dimensional key points of the face are directly mapped to the structured point cloud data, avoiding the need to build a pixel lookup table. The two-dimensional key points are extracted using a PFLD network and registered using a rigid body transformation matrix and an iterative nearest point method.

Benefits of technology

It simplifies the extraction process of 3D key points of human face, improves registration efficiency and accuracy, and enables rapid and accurate creation of 3D face models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830663B_ABST
    Figure CN115830663B_ABST
Patent Text Reader

Abstract

The application provides a face three-dimensional key point extraction method and device, a model creation method and system, the face three-dimensional key point extraction method comprising: obtaining first two-dimensional image data and corresponding point cloud data of a face; extracting face two-dimensional key points from the first two-dimensional image data; performing image structured storage on the point cloud data to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixel points at the same face position between the face point cloud structured storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have three-dimensional data channels; mapping the extracted face two-dimensional key points to the face point cloud structured storage data; and taking the three-dimensional points in the face point cloud structured storage data that are mapped by the face two-dimensional key points as face three-dimensional key points. The application can not need to establish a pixel lookup table, and is beneficial to simplifying the face three-dimensional key point extraction mode.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human face three-dimensional model, and particularly relates to a human face three-dimensional key point extraction method and device, a model creation method and system. BACKGROUND

[0002] A human face is a complex three-dimensional body, and how to quickly, completely and accurately realize human face three-dimensional digitization is a key research problem in the field of computer vision and image graphics.

[0003] In the current three-dimensional human face reconstruction technology, a three-dimensional human face reconstruction method based on structured light is not affected by the complexity of the surface texture of the measured object, and the reconstructed three-dimensional human face model is realistic and high in precision. The method mainly includes the following steps: human face point cloud acquisition, human face point cloud preprocessing, human face point cloud registration, human face three-dimensional reconstruction, etc. The human face point cloud registration is an important link in three-dimensional human face digitization, and the extraction of human face three-dimensional key points affects the efficiency of registration.

[0004] The patent document with the publication number CN113544744 A discloses a head posture measurement method. In the prior art, the following method is used to extract human face three-dimensional key points: based on the two-dimensional key points of the face of a target object, indexing is performed in the point cloud data of the face of the target object to obtain the point cloud data of the face key points corresponding to the two-dimensional key points of the face of the target object (the three-dimensional key points of the face of the target object). However, this method needs to establish a pixel lookup table between a three-dimensional depth image and a two-dimensional color image in the point cloud data generation process, and realizes the indexing process of two-dimensional mapping to three-dimensional through the established pixel lookup table. SUMMARY

[0005] Based on the above-mentioned status quo, the main purpose of the present application is to provide a human face three-dimensional key point extraction method and device, a model creation method and system, and a storage medium, which can not need to establish a pixel lookup table, and are beneficial to simplify the extraction method of human face three-dimensional key points.

[0006] To achieve the above-mentioned purpose, the technical scheme of the present application provides a human face three-dimensional key point extraction method, which comprises the following steps:

[0007] Step 100: acquiring first two-dimensional image data of a human face and corresponding point cloud data, wherein the first two-dimensional image data comprises a plurality of pixel points, and the point cloud data comprises a plurality of three-dimensional points;

[0008] Step 200: extracting human face two-dimensional key points from the first two-dimensional image data;

[0009] Step 300: image structure storage is performed on the point cloud data, and corresponding face point cloud structure storage data is obtained, wherein the same face position three-dimensional points and pixel points in the face point cloud structure storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structure storage data have three-dimensional data channels for storing their own three-dimensional coordinates.

[0010] Step 400: the extracted face two-dimensional key points are mapped to the face point cloud structure storage data, and the three-dimensional points in the face point cloud structure storage data mapped by the face two-dimensional key points are taken as face three-dimensional key points.

[0011] Further, the step 200 comprises:

[0012] Step 201: face detection is performed on the first two-dimensional image data, and position and size information of the face is obtained.

[0013] Step 202: image cropping is performed on the first two-dimensional image data according to the position and size information of the face.

[0014] Step 203: face two-dimensional key point extraction is performed on the cropped image data.

[0015] Further, in the step 200, the PFLD network is used to extract the face two-dimensional key points.

[0016] Further, in the step 100, the point cloud data is obtained by the following method:

[0017] Second two-dimensional image data containing structured light modulation information of the face is obtained, and the point cloud data is obtained according to the structured light modulation information in the second two-dimensional image data.

[0018] Further, the three-dimensional data channel comprises a first channel, a second channel, a third channel and a fourth channel, wherein the first channel is used to store the X coordinate in the three-dimensional coordinate, the second channel is used to store the Y coordinate in the three-dimensional coordinate, the third channel is used to store the Z coordinate in the three-dimensional coordinate, and the fourth channel is used to store the effective identification of the three-dimensional point, and the effective identification is used to indicate whether the three-dimensional coordinate stored by the three-dimensional point is effective.

[0019] The technical scheme of the present application also provides a face three-dimensional model creation method, comprising:

[0020] Step 10: The left face data, the front face data and the right face data are processed respectively by using any one of the face three-dimensional key point extraction methods described above, so as to obtain the left face three-dimensional key point, the front face three-dimensional key point and the right face three-dimensional key point, wherein the left face data includes the first two-dimensional image data and the corresponding left face point cloud data of the left face, the front face data includes the first two-dimensional image data and the corresponding front face point cloud data of the front face, and the right face data includes the first two-dimensional image data and the corresponding right face point cloud data of the right face;

[0021] Step 20: The left face point cloud data and the front face point cloud data are registered according to the left face three-dimensional key point and the front face three-dimensional key point, and the right face point cloud data and the front face point cloud data are registered according to the right face three-dimensional key point and the front face three-dimensional key point, so as to obtain the face three-dimensional model obtained by splicing the left face point cloud data, the front face point cloud data and the right face point cloud data.

[0022] Further, the step 20 includes:

[0023] Step 21: A first key point pair located in a set face region is obtained, a rigid transformation matrix parameter between the left face point cloud data and the front face point cloud data is calculated according to the first key point pair, and the left face point cloud data and the front face point cloud data are registered according to the calculated rigid transformation matrix parameter, wherein a left face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair;

[0024] Step 22: A second key point pair located in the set face region is obtained, a rigid transformation matrix parameter between the right face point cloud data and the front face point cloud data is calculated according to the second key point pair, and the right face point cloud data and the front face point cloud data are registered according to the calculated rigid transformation matrix parameter, wherein a right face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair;

[0025] The set face region includes at least one of a brow region, a nose region, a mouth region and a chin region.

[0026] Further, after the step 20, the face three-dimensional model creation method further includes:

[0027] Step 30: The face three-dimensional model is registered by using an iterative closest point method.

[0028] The technical scheme of the present application also provides a face three-dimensional key point extraction device, which includes:

[0029] An acquisition module is configured to acquire first two-dimensional image data and corresponding point cloud data of a face, the first two-dimensional image data including a plurality of pixel points, and the point cloud data including a plurality of three-dimensional points.

[0030] An extraction module is configured to extract face two-dimensional key points from the first two-dimensional image data.

[0031] A point cloud data storage module is configured to store the point cloud data in an image structure to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixel points at the same face position in the face point cloud structured storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have a three-dimensional data channel for storing their own three-dimensional coordinates.

[0032] A mapping processing module is configured to map the extracted face two-dimensional key points to the face point cloud structured storage data, and take the three-dimensional points in the face point cloud structured storage data that are mapped by the face two-dimensional key points as face three-dimensional key points.

[0033] Further, the extraction module includes:

[0034] A detection unit is configured to detect a face from the first two-dimensional image data to obtain position and size information of the face.

[0035] A cropping unit is configured to crop the first two-dimensional image data according to the position and size information of the face.

[0036] An extraction unit is configured to extract face two-dimensional key points from the cropped image data.

[0037] Further, the extraction module is configured to extract face two-dimensional key points by using a PFLD network.

[0038] Further, the acquisition module acquires the point cloud data by the following method:

[0039] Acquire second two-dimensional image data of the face containing structured light modulation information, and obtain the point cloud data according to the structured light modulation information in the second two-dimensional image data.

[0040] Further, the three-dimensional data channel includes a first channel, a second channel, a third channel and a fourth channel, wherein the first channel is configured to store the X coordinate in the three-dimensional coordinate, the second channel is configured to store the Y coordinate in the three-dimensional coordinate, the third channel is configured to store the Z coordinate in the three-dimensional coordinate, and the fourth channel is configured to store an effective identifier of the three-dimensional point, the effective identifier being configured to indicate whether the three-dimensional coordinate stored by the three-dimensional point is valid.

[0041] The technical scheme of the present application also provides a face three-dimensional key point extraction device, comprising a processor and a memory coupled with the processor, wherein the memory stores instructions for the processor to execute, and when the processor executes the instructions, any one of the face three-dimensional key point extraction methods described above can be implemented.

[0042] The technical scheme of the present application also provides a face three-dimensional model creation system, comprising:

[0043] Any one of the face three-dimensional key point extraction devices described above is used to process left face data, front face data and right face data respectively, so as to obtain left face three-dimensional key points, front face three-dimensional key points and right face three-dimensional key points, wherein the left face data comprises first two-dimensional image data and corresponding left face point cloud data of the left face, the front face data comprises first two-dimensional image data and corresponding front face point cloud data of the front face, and the right face data comprises first two-dimensional image data and corresponding right face point cloud data of the right face.

[0044] A first registration device is used to register the left face point cloud data and the front face point cloud data according to the left face three-dimensional key points and the front face three-dimensional key points, and to register the right face point cloud data and the front face point cloud data according to the right face three-dimensional key points and the front face three-dimensional key points, so as to obtain a face three-dimensional model obtained by splicing the left face point cloud data, the front face point cloud data and the right face point cloud data.

[0045] Further, the first registration device comprises:

[0046] A first registration processing module is used to obtain a first key point pair located in a set face area, to calculate rigid transformation matrix parameters between the left face point cloud data and the front face point cloud data according to the first key point pair, and to register the left face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameters, wherein a left face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair.

[0047] A second registration processing module is used to obtain a second key point pair located in the set face area, to calculate rigid transformation matrix parameters between the right face point cloud data and the front face point cloud data according to the second key point pair, and to register the right face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameters, wherein a right face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair.

[0048] The set face area comprises at least one of a brow area, a nose area, a mouth area and a chin area.

[0049] Further, the system further comprises:

[0050] A second registration device is configured to register the three-dimensional face model by using an iterative closest point method.

[0051] The technical scheme of the present application further provides a three-dimensional face model creation system, which comprises a processor and a memory coupled to the processor, wherein the memory stores instructions for the processor to execute, and when the processor executes the instructions, any one of the above-mentioned three-dimensional face model creation methods can be implemented.

[0052] The technical scheme of the present application further provides a computer readable storage medium, which stores a computer program, and when the program is executed by a processor, any one of the above-mentioned three-dimensional face key point extraction methods or any one of the above-mentioned three-dimensional face model creation methods can be implemented.

[0053] The three-dimensional face key point extraction method provided by the present application stores point cloud data according to the image structure of the two-dimensional image data of the face region, so that after obtaining the two-dimensional face key point pixels, the face three-dimensional key points can be directly mapped to the face point cloud structured storage data to obtain the face three-dimensional key points, without the need to establish a pixel lookup table, thereby simplifying the extraction method of the face three-dimensional key points and improving the registration efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0054] The above and other objects, features and advantages of the present application will become more apparent from the following description of embodiments of the present application taken with reference to the accompanying drawings, in which:

[0055] Figure 1 is a flowchart of a three-dimensional face key point extraction method provided by an embodiment of the present application;

[0056] Figure 2 is a schematic diagram of a three-dimensional scanning device provided by an embodiment of the present application;

[0057] Figure 3 is a schematic diagram of a face image acquisition provided by an embodiment of the present application;

[0058] Figure 4 is a schematic diagram of face point cloud structured storage data provided by an embodiment of the present application;

[0059] Figure 5 is a schematic diagram of a face registration result provided by an embodiment of the present application. DETAILED DESCRIPTION

[0060] The present application will be described below based on examples, but the present application is not limited to these examples only. In the following detailed description of the present application, some specific details are described in detail in order to avoid obscuring the essence of the present application, and well-known methods, processes, procedures, elements are not described in detail.

[0061] In addition, those skilled in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.

[0062] Unless the context clearly requires otherwise, throughout the description and the claims, the words "comprise", "comprising", and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of "including, but not limited to".

[0063] In the description of the present application, it should be understood that the terms "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "multiple" is two or more.

[0064] It should be noted that the step number (letter or number) is used in the present application to refer to certain specific method steps, which is only for the purpose of convenience and brevity of description, and in no way limits the order of these method steps by letters or numbers. Those skilled in the art can understand that the order of the related method steps should be determined by the technology itself, and should not be improperly limited by the existence of step numbers.

[0065] Referring to Figure 1 , Figure 1 is a flowchart of a face three-dimensional key point extraction method provided by an embodiment of the present application, and the face three-dimensional key point extraction method comprises:

[0066] Step 100: obtaining first two-dimensional image data of a face and corresponding point cloud data, the first two-dimensional image data comprising a plurality of pixel points, and the point cloud data comprising a plurality of three-dimensional points;

[0067] Among them, the first two-dimensional image data is a two-dimensional image data of a face collected by an image device (such as a camera) without structured light modulation information;

[0068] It is understood that the aforementioned point cloud data refers to the point cloud data of the same shooting object as the aforementioned first two-dimensional image data. The point cloud data of the face region can be obtained in the following way: obtain the second two-dimensional image data of the face containing structured light modulation information (the second two-dimensional image data and the first two-dimensional image data come from the same shooting object), and obtain the point cloud data based on the structured light modulation information in the second two-dimensional image data. Specifically, the corresponding point cloud data can be reconstructed by N-step phase shift and Gray code dephase wrapping technology.

[0069] For example, an image of a face can be captured first using an imaging device to obtain the first two-dimensional image data. Then, structured light can be projected onto the face, and multiple face images with structured light stripes can be captured at the same location to obtain the second two-dimensional image data. The point cloud data of the face can be demodulated through the deformed stripes in these face images.

[0070] For example, after acquiring second two-dimensional image data containing structured light modulation information of the left face, the front face, and the right face, the point cloud data of the left face can be obtained from the second two-dimensional image data containing structured light modulation information of the left face, the point cloud data of the front face can be obtained from the second two-dimensional image data containing structured light modulation information of the front face, and the point cloud data of the right face can be obtained from the second two-dimensional image data containing structured light modulation information of the right face.

[0071] For example, in this embodiment of the invention, a single-node 3D scanning device based on structured light can be used to acquire 2D facial images and point cloud data, such as... Figure 2 As shown, the 3D scanning device consists of two cameras 1 and 3 (one is a left camera and the other is a right camera; the cameras can be monochrome industrial cameras) and a digital micro-projector device 2. The point cloud data acquisition process is as follows: structured light is projected onto the object to be measured (such as a human head) through the digital micro-projector device 2, forming a special image on the surface of the object. The structured light information modulated by the human face surface is collected by the camera. By decoding the two-dimensional image collected by the camera, the absolute phase information of the object can be obtained. Combined with the principle of binocular stereo vision, the three-dimensional data (point cloud data) of the object can be obtained.

[0072] For example, such as Figure 3As shown, in the process of creating a three-dimensional face model, a front face is taken as a central axis, and the face is rotated within a preset angle (such as ±60°) to realize the collection of left face data, front face data and right face data, wherein the left face data includes first two-dimensional image data and corresponding left face point cloud data of the left face, the front face data includes first two-dimensional image data and corresponding front face point cloud data of the front face, and the right face data includes first two-dimensional image data and corresponding right face point cloud data of the right face;

[0073] Step 200: extracting face two-dimensional key points from the first two-dimensional image data to obtain a plurality of pixel points as the face two-dimensional key points;

[0074] Step 300: performing image structured storage on the point cloud data to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixel points at the same face position between the face point cloud structured storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have three-dimensional data channels for storing their own three-dimensional coordinates;

[0075] That is, each three-dimensional point in the face point cloud structured storage data has a three-dimensional data channel, and the value of the three-dimensional data channel represents the three-dimensional coordinates created for the face position of the corresponding pixel point in the face two-dimensional image data;

[0076] It can be understood that the storage position of the pixel point in the first two-dimensional image data is the pixel position of the pixel point;

[0077] Step 400: mapping the extracted face two-dimensional key points to the face point cloud structured storage data, and taking the three-dimensional points in the face point cloud structured storage data mapped by the face two-dimensional key points as face three-dimensional key points.

[0078] It can be understood that in the present application, the face three-dimensional key points should be valid three-dimensional points (i.e. the three-dimensional coordinates of the three-dimensional points are valid coordinates);

[0079] Since the face point cloud structured storage data is structured storage data established according to the corresponding two-dimensional image data, the three-dimensional points and the pixel points at the same face position have the same storage position, so after obtaining the storage position (such as pixel coordinates) of the face two-dimensional key points, the corresponding face three-dimensional key points can be obtained through the three-dimensional points at the corresponding storage position in the face point cloud structured storage data, so that after the extraction of the face two-dimensional key points is completed, the face three-dimensional key points can be obtained through simple mapping, without the need to establish a pixel lookup table between three-dimensional and two-dimensional, thereby the implementation mode can be simplified and the registration efficiency can be improved.

[0080] The face three-dimensional key point extraction method provided by the embodiment of the present application can store point cloud data in a corresponding image structure according to two-dimensional image data of a face region, so that after obtaining face two-dimensional key point pixels, the face two-dimensional key point pixels can be directly mapped to face point cloud structured storage data to obtain face three-dimensional key points, without the need to establish a pixel lookup table, thereby facilitating the simplification of the face three-dimensional key point extraction method and improving the registration efficiency.

[0081] Preferably, in an embodiment of the present application, in order to improve the detection speed and detection accuracy of the key points, image cropping can be performed on the first face two-dimensional image data to reduce the detection range, wherein the step 200 can include:

[0082] Step 201: performing face detection on the first two-dimensional image data to obtain position and size information of the face; wherein the first two-dimensional image data can be face two-dimensional image data collected by an image device (such as a camera); the first two-dimensional image data can be single-channel grayscale image data;

[0083] After collecting the first two-dimensional image data of the face, face image detection can be performed first to locate the position and size information of the face, for example, a residual SSD network in a DNN module based on OpenCV can be used for face prediction;

[0084] Step 202: performing image cropping on the first two-dimensional image data according to the position and size information of the face;

[0085] After locating the position and size information of the face, a rectangular frame can be used for cropping to reduce the area (such as the background area, hair area, etc.) in the image that is irrelevant to the key points, so that when feature extraction is performed on the cropped image, the detection range can be reduced, thereby improving the detection speed and detection accuracy of the key points;

[0086] Step 203: extracting face two-dimensional key points from the cropped image data.

[0087] In the step 200 of the embodiment of the present application, a neural network can be used to extract face two-dimensional key points from face two-dimensional image data, preferably, in the embodiment of the present application, PFLD (A Practical Facial Landmark Detector, a practical face feature point detection algorithm) can be used to extract face two-dimensional key points, and the PFLD network has the advantages of high precision, small model, and fast speed, and can achieve good key point extraction efficiency, in the present embodiment, the required PFLD network can be obtained in the following way:

[0088] (1): Obtain a data set and label the data set, the labeled data set containing a plurality of two-dimensional images of human faces, the key points in each two-dimensional image of the human face being labeled, wherein the labeled key points can include a plurality of pixel points in the edge region of the face, a plurality of pixel points in the eyebrow region, a plurality of pixel points in the nose region, a plurality of pixel points in the mouth region, and a plurality of pixel points in the eye region, for example, 98 key points can be labeled for each two-dimensional image of the human face;

[0089] After obtaining the labeled data set, the data set can be divided into a training set and a test set, the division ratio of the training set can be about 75%, for example, an open source data set such as the WFLW data set can be used, which can include 10,000 face images, 7,500 of which are divided into a training set and 2,500 of which are divided into a test set;

[0090] (2): Data preprocessing is performed on the training set and the test set, including data augmentation and obtaining Euler angles, wherein the data augmentation can include random cropping, vertical flipping, rotation, etc.;

[0091] (3): The PFLD network model is trained using the training set, wherein the Adam optimizer can be used during training, the learning rate is adjusted to 1e-4, the weight decay is 1e-6, the training batchsize is 256, and the iteration is performed for about 400 rounds;

[0092] (4): After the training is completed, the trained PFLD network is tested using the test set to detect the effect of key point extraction.

[0093] Through the above method, the neural network based on pose information supervision can be obtained.

[0094] It can be understood that since the key points in the data set during training are labeled in a certain order, the human face two-dimensional key points extracted from the first two-dimensional image data of the human face by the trained PFLD network in step 200 are also arranged in the order of labeling, and the position of the human face two-dimensional key points can be known through the order of the human face two-dimensional key points, for example, the first 33 human face two-dimensional key points are key points in the edge region of the face, the 34th-51st human face two-dimensional key points are key points in the eyebrow region, the 52nd-60th human face two-dimensional key points are key points in the nose region, the 61st-76th, 97th and 98th human face two-dimensional key points are key points in the eye region, and the 77th-96th human face two-dimensional key points are key points in the nose region.

[0095] In the present application, the face point cloud structured storage data contains three-dimensional coordinate information of three-dimensional points, so each three-dimensional point in the face point cloud structured storage data has multiple channels, specifically four channels, for example, see Figure 4 Each pixel point in the face point cloud structured storage data includes a first channel (X channel), a second channel (Y channel), a third channel (Z channel), and a fourth channel (P channel), wherein the first channel is used to store the X coordinate in the three-dimensional coordinate, the second channel is used to store the Y coordinate in the three-dimensional coordinate, the third channel is used to store the Z coordinate in the three-dimensional coordinate, and the fourth channel of the three-dimensional point is used to store the valid identification (valid bit) of the three-dimensional point, which is used to indicate whether the three-dimensional coordinate data stored by the three-dimensional point is valid, i.e., whether the corresponding pixel position is reconstructed into a three-dimensional point.

[0096] For example, for a pixel point with a pixel position of the third row and the first column in the face two-dimensional image data, the three-dimensional coordinate of the three-dimensional point created by the pixel position is (X1, Y1, Z1), and in the face point cloud structured storage data, the value of the first channel of the three-dimensional point with a storage position of the third row and the first column is X1, the value of the second channel is Y1, the value of the third channel is Z1, and the value of the fourth channel is 1, representing that the three-dimensional coordinate data stored by the three-dimensional point is valid; for a pixel point with a pixel position of the second row and the first column in the face two-dimensional image data, the pixel position does not create a three-dimensional point, and in the face point cloud structured storage data, the value of the fourth channel of the three-dimensional point with a storage position of the second row and the first column is 0, representing that the three-dimensional coordinate data stored by the three-dimensional point is invalid, and the values of the first channel, the second channel, and the third channel can be set to be preset 0.

[0097] The embodiment of the present application also provides a face three-dimensional model creation method, which comprises:

[0098] Step 10: adopting the face three-dimensional key point extraction method described above to process left face data, front face data, and right face data respectively, so as to obtain left face three-dimensional key points, front face three-dimensional key points, and right face three-dimensional key points, wherein the left face data includes first two-dimensional image data and corresponding left face point cloud data of the left face, the front face data includes first two-dimensional image data and corresponding front face point cloud data of the front face, and the right face data includes first two-dimensional image data and corresponding right face point cloud data of the right face;

[0099] Step 20: registering the left-side face point cloud data with the front face point cloud data according to the left-side face three-dimensional key points and the front face three-dimensional key points, and registering the right-side face point cloud data with the front face point cloud data according to the right-side face three-dimensional key points and the front face three-dimensional key points, so as to obtain a three-dimensional face model obtained by splicing the left-side face point cloud data, the front face point cloud data and the right-side face point cloud data, that is, by registering the side face data with the front face data, a complete face model can be obtained.

[0100] wherein the overlapping parts of the side face data and the front face data are different due to different shooting angles. Therefore, in order to make the algorithm more robust and reliable, preferably, only robust feature points (key points) can be selected for registration operation, wherein the step 20 can include:

[0101] Step 21: obtaining a first key point pair located in a set face region, calculating a rigid transformation matrix parameter between the left-side face point cloud data and the front face point cloud data according to the first key point pair, and registering the left-side face point cloud data with the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a left-side face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair;

[0102] wherein the set face region can include a brow region, a nose region, a mouth region and a chin region in a face edge region;

[0103] Since the face two-dimensional key points extracted from the first two-dimensional image data of the face by the PFLD network are arranged in the order of annotation, the face position of the face two-dimensional key points can be known by the order of the face two-dimensional key points, that is, the face position of the corresponding face three-dimensional key points, so that in all the left-side face three-dimensional key points and all the front face three-dimensional key points, the left-side face three-dimensional key points and the front face three-dimensional key points with the same order can be taken as a first key pair, so as to obtain several second key pairs;

[0104] After obtaining several first key pairs, the first key pairs located in the set face region are screened out, and then SVD decomposition is used to obtain the rigid transformation matrix parameter R, t between the left-side face point cloud data and the front face point cloud data, and the left-side face point cloud data and the front face point cloud data are registered and spliced by using the parameter;

[0105] Step 22: obtaining a second key point pair located in the above set face region, calculating the rigid transformation matrix parameter between the right face point cloud data and the front face point cloud data according to the second key point pair, and registering the right face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a right face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair;

[0106] Among all the extracted right face three-dimensional key points and all the front face three-dimensional key points, the right face three-dimensional key points and the front face three-dimensional key points with the same order can be taken as a second key pair, so that a plurality of second key pairs can be obtained;

[0107] After obtaining a plurality of second key pairs, the second key pairs located in the set face region are screened out, and then SVD decomposition is used to obtain the rigid transformation matrix parameter R, t between the right face point cloud data and the front face point cloud data, and the right face point cloud data and the front face point cloud data are registered and spliced by using the parameter;

[0108] After registering and splicing the left face point cloud data, the right face point cloud data and the front face point cloud data, a complete face point cloud model can be obtained.

[0109] For example, in the above embodiment using the PFLD network, among the 98 extracted face two-dimensional key points, more than 30 points such as eyebrow region points (37th, 38th, 39th, 40th, 43rd, 44th, 50th, 51st eyebrow feature points), nose region points (52nd, 53rd, 54th, 55th, 56th, 57th, 58th nose feature points), mouth region points (78th, 79th, 80th, 84th, 85th, 86th, 89th, 90th, 91st, 93rd, 94th, 95th mouth feature points) and chin region points (15th, 16th, 17th, 18th chin feature points) can be selected, and the region points in the above three postures are further screened and registered.

[0110] For example, the results of registering the left face data, the right face data and the front face data respectively by using the obtained three-dimensional key point pairs are as shown in Figure 5

[0111] The face three-dimensional model creation method provided by the embodiment fully utilizes the characteristics of the face image, reduces the three-dimensional point cloud registration problem to two-dimensional, simplifies the registration problem, improves the registration speed, and through the corresponding image structured storage of the face point cloud data, the face two-dimensional key points can be directly mapped to the face point cloud structured storage data to obtain the face three-dimensional key points, without the need to establish a pixel lookup table, which is conducive to further improving the registration speed; ​

[0112] In addition, the PFLD network is used for extraction and registration of two-dimensional key points, which has better robustness and stronger anti-interference ability, can fully utilize the texture information of the face image, and greatly reduces the occurrence of false registration. Meanwhile, the registration process can be unaffected by the number of point clouds, and can reduce the registration problem of hundreds of thousands of point clouds to tens of key points, greatly simplifying the registration process, improving the registration speed, and achieving a coarse registration method of 33 fps on the CPU, which is conducive to real-time operation.

[0113] The coarse registration of the face point cloud model can be achieved through the above steps 10 and 20. Preferably, in the embodiment of the application, after the step 20, the face three-dimensional model creation method further comprises:

[0114] Step 30: using the Iterative Closest Point method (ICP) to register the face three-dimensional model, thereby achieving fine registration.

[0115] The ICP registration process can include:

[0116] Step 1: selecting point cloud P as the source point cloud and point cloud Q as the target point cloud;

[0117] Step 2: searching in the target point cloud Q for the point q i with the closest Euclidean distance to each point p i in the source point cloud P, and constructing two sets of corresponding point sets P1={p1,p2,…,p k1} and Q1={q1,q2,…,q k1};

[0118] Step 3: using the Singular Value Decomposition (SVD) or quaternion method to calculate the transformation parameters R and t of the two point sets through the corresponding point sets;

[0119] Step 4: using the transformation parameters solved in the previous step to perform rigid transformation on the source point cloud, to obtain new corresponding point sets P2={p1,p2,…,p k2} and Q2={q1,q2,…,q k2}

[0120] Step 5: when the value of the objective function E(R,t) is greater than a given threshold, repeating step 3 until the iteration termination condition is met, for example, the objective function is less than a certain threshold or the number of iterations reaches the requirement.

[0121] The face three-dimensional model creation method provided by the embodiment of the present application can quickly provide a good initial posture through the rough registration of the face point cloud model by steps 10 and 20, and can make the convergence faster and avoid the problem that the ICP algorithm is prone to fall into local optimization through the subsequent fine registration algorithm of ICP.

[0122] The face three-dimensional model creation method based on the pose information supervision network provided by the present application uses a three-dimensional scanning device based on structured light to collect face image information and point cloud information from the left side, the right side and the front of the face, and uses a face detector based on an OpenCV DNN module to detect the face image, locate the position and size information of the face, and cut the face image with a rectangular frame, then sends the cut image into a pose information supervision network (PFLD network) to extract two-dimensional key points of the face, maps the two-dimensional key points back to a three-dimensional space by using camera calibration information to obtain three-dimensional key points, and then the three-dimensional key points of the left side and the front (front) face can form a key point pair, the three-dimensional key points of the right side and the front face can also form a key point pair, and the rotation matrix and the translation vector between the point clouds are obtained through SVD decomposition based on the obtained three-dimensional key point pairs, so that the left face point cloud and the right face point cloud are respectively quickly registered with the front face point cloud, and the initial registration matrix of the two frames of point clouds is automatically obtained, and finally the point cloud is fine registered by using an improved ICP iterative algorithm, so that the three-dimensional face point cloud registration is quickly and accurately realized.

[0123] The embodiment of the present application also provides a face three-dimensional key point extraction device, which comprises:

[0124] An acquisition module is configured to acquire first two-dimensional image data of a face and corresponding point cloud data, wherein the first two-dimensional image data comprises a plurality of pixel points, and the point cloud data comprises a plurality of three-dimensional points.

[0125] An extraction module is configured to extract two-dimensional key points of the face from the first two-dimensional image data.

[0126] A point cloud data storage module is configured to store the point cloud data in an image structure to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixel points at the same face position in the face point cloud structured storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have a three-dimensional data channel for storing their own three-dimensional coordinates.

[0127] A mapping processing module is configured to map the extracted two-dimensional key points of the face to the face point cloud structured storage data, and take the three-dimensional points in the face point cloud structured storage data that are mapped by the two-dimensional key points of the face as the three-dimensional key points of the face.

[0128] In an embodiment of the present application, the extraction module comprises:

[0129] a detection unit configured to perform face detection on the first two-dimensional image data to obtain position and size information of a face;

[0130] a cropping unit configured to perform image cropping on the first two-dimensional image data according to the position and size information of the face;

[0131] an extraction unit configured to perform face two-dimensional key point extraction on the cropped image data.

[0132] In an embodiment of the present application, the extraction module is configured to perform face two-dimensional key point extraction using a PFLD network.

[0133] In an embodiment of the present application, the acquisition module acquires the point cloud data in the following manner:

[0134] acquiring second two-dimensional image data of the face containing structured light modulation information, and obtaining the point cloud data according to the structured light modulation information in the second two-dimensional image data.

[0135] In an embodiment of the present application, the three-dimensional data channel comprises a first channel, a second channel, a third channel and a fourth channel, wherein the first channel is configured to store X coordinates in the three-dimensional coordinates, the second channel is configured to store Y coordinates in the three-dimensional coordinates, the third channel is configured to store Z coordinates in the three-dimensional coordinates, and the fourth channel is configured to store an effective identifier of a three-dimensional point, the effective identifier being configured to indicate whether the three-dimensional coordinates stored by the three-dimensional point are valid.

[0136] The embodiment of the present application further provides a face three-dimensional key point extraction device, comprising a processor and a memory coupled with the processor, wherein the memory stores instructions for the processor to execute, and when the processor executes the instructions, the face three-dimensional key point extraction method described above can be realized.

[0137] The embodiment of the present application further provides a face three-dimensional model creation system, comprising:

[0138] The face three-dimensional key point extraction device described above is configured to process left face data, front face data and right face data respectively, so as to obtain left face three-dimensional key points, front face three-dimensional key points and right face three-dimensional key points, wherein the left face data comprises first two-dimensional image data of a left face and corresponding left face point cloud data, the front face data comprises first two-dimensional image data of a front face and corresponding front face point cloud data, and the right face data comprises first two-dimensional image data of a right face and corresponding right face point cloud data.

[0139] The first registration device is configured to register the left-side face point cloud data and the front face point cloud data according to the left-side face three-dimensional key points and the front face three-dimensional key points, and to register the right-side face point cloud data and the front face point cloud data according to the right-side face three-dimensional key points and the front face three-dimensional key points, so as to obtain a face three-dimensional model in which the left-side face point cloud data, the front face point cloud data and the right-side face point cloud data are spliced together.

[0140] In an embodiment of the present application, the first registration device comprises:

[0141] The first registration processing module is configured to obtain a first key point pair located in a set face region, to calculate a rigid transformation matrix parameter between the left-side face point cloud data and the front face point cloud data according to the first key point pair, and to register the left-side face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a left-side face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair.

[0142] The second registration processing module is configured to obtain a second key point pair located in the set face region, to calculate a rigid transformation matrix parameter between the right-side face point cloud data and the front face point cloud data according to the second key point pair, and to register the right-side face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a right-side face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair.

[0143] The set face region comprises at least one of a brow region, a nose region, a mouth region and a chin region.

[0144] In an embodiment of the present application, the system further comprises:

[0145] The second registration device is configured to register the face three-dimensional model by using an iterative closest point method.

[0146] The embodiment of the present application further provides a face three-dimensional model creation system, comprising a processor and a memory coupled with the processor, wherein the memory stores instructions for the processor to execute, and when the processor executes the instructions, the face three-dimensional model creation method described above can be realized.

[0147] The embodiment of the present application further provides a computer readable storage medium storing a computer program, which is executed by a processor to realize the face three-dimensional key point extraction method or the face three-dimensional model creation method described above.

[0148] Those skilled in the art can understand that the above-mentioned preferred solutions can be freely combined and superimposed without conflict.

[0149] It should be understood that the above-described implementations are merely exemplary, and not limiting, and that various obvious or equivalent modifications or substitutions for the above-described details can be made by those skilled in the art without departing from the spirit of the present application, and all such modifications or substitutions are intended to be included within the scope of the claims of the present application.

Claims

1. A method for extracting three-dimensional key points of a human face, characterized in that, The method comprises the steps of: Step 100: acquiring first two-dimensional image data of a face and corresponding point cloud data, the first two-dimensional image data comprising a plurality of pixels, and the point cloud data comprising a plurality of three-dimensional points; Step 200: extracting face two-dimensional key points from the first two-dimensional image data; Step 300: performing image structured storage on the point cloud data to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixels at the same face position in the face point cloud structured storage data and the first two-dimensional image data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have a three-dimensional data channel for storing their own three-dimensional coordinates; Step 400: mapping the extracted face two-dimensional key points to the face point cloud structured storage data, and taking the three-dimensional points in the face point cloud structured storage data that are mapped by the face two-dimensional key points as face three-dimensional key points.

2. The method of claim 1, wherein, The step 200 comprises the steps of: Step 201: performing face detection on the first two-dimensional image data to obtain position and size information of the face; Step 202: performing image cropping on the first two-dimensional image data according to the position and size information of the face; Step 203: extracting face two-dimensional key points from the cropped image data.

3. The method according to claim 1 or 2, characterized in that, In the step 200, the PFLD network is used to extract the face two-dimensional key points.

4. The method according to any one of claims 1 to 3, characterized in that, In the step 100, the point cloud data is acquired in the following manner: Acquiring second two-dimensional image data of the face containing structured light modulation information, and obtaining the point cloud data according to the structured light modulation information in the second two-dimensional image data.

5. The method according to any one of claims 1 to 4, characterized in that, The three-dimensional data channel comprises a first channel, a second channel, a third channel and a fourth channel, wherein the first channel is used to store the X coordinate in the three-dimensional coordinate, the second channel is used to store the Y coordinate in the three-dimensional coordinate, the third channel is used to store the Z coordinate in the three-dimensional coordinate, and the fourth channel is used to store an effective identifier of the three-dimensional point, the effective identifier being used to indicate whether the three-dimensional coordinate stored by the three-dimensional point is valid.

6. A method of creating a three-dimensional model of a human face, characterized by, The method comprises the steps of: Step 10: processing left face data, front face data and right face data respectively by using the method according to any one of claims 1-5, to obtain left face three-dimensional key points, front face three-dimensional key points and right face three-dimensional key points, wherein the left face data comprises first two-dimensional image data of a left face and corresponding left face point cloud data, the front face data comprises first two-dimensional image data of a front face and corresponding front face point cloud data, and the right face data comprises first two-dimensional image data of a right face and corresponding right face point cloud data; Step 20: registering the left face point cloud data with the front face point cloud data according to the left face three-dimensional key points and the front face three-dimensional key points, and registering the right face point cloud data with the front face point cloud data according to the right face three-dimensional key points and the front face three-dimensional key points, to obtain a face three-dimensional model obtained by splicing the left face point cloud data, the front face point cloud data and the right face point cloud data.

7. The method of claim 6, wherein, The step 20 comprises: Step 21: obtaining a first key point pair located in a set face region, calculating a rigid transformation matrix parameter between the left face point cloud data and the front face point cloud data according to the first key point pair, and registering the left face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a left face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair; Step 22: obtaining a second key point pair located in the set face region, calculating a rigid transformation matrix parameter between the right face point cloud data and the front face point cloud data according to the second key point pair, and registering the right face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameter, wherein a right face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair; The set face region includes at least one of a brow region, a nose region, a mouth region, and a chin region.

8. The method according to claim 6 or 7, characterized in that, After the step 20, the face three-dimensional model creation method further includes: Step 30: registering the face three-dimensional model by using an iterative closest point method.

9. A device for extracting three-dimensional key points of a human face, characterized by comprising: Including: An acquisition module is configured to acquire first two-dimensional image data of a face and corresponding point cloud data, the first two-dimensional image data includes a plurality of pixel points, and the point cloud data includes a plurality of three-dimensional points; An extraction module is configured to extract face two-dimensional key points from the first two-dimensional image data; A point cloud data storage module is configured to store the point cloud data in an image structure to obtain corresponding face point cloud structured storage data, wherein the three-dimensional points and the pixel points at the same face position in the face point cloud structured storage data have the same storage position, and the three-dimensional points in the face point cloud structured storage data have a three-dimensional data channel for storing their own three-dimensional coordinates; A mapping processing module is configured to map the extracted face two-dimensional key points to the face point cloud structured storage data, and take the three-dimensional points in the face point cloud structured storage data that are mapped by the face two-dimensional key points as face three-dimensional key points.

10. The apparatus of claim 9, wherein, The extraction module includes: A detection unit is configured to detect a face from the first two-dimensional image data to obtain position and size information of the face; A cropping unit is configured to crop the first two-dimensional image data according to the position and size information of the face; An extraction unit is configured to extract face two-dimensional key points from the cropped image data.

11. The apparatus of claim 9 or 10, wherein, The extraction module is configured to extract face two-dimensional key points by using a PFLD network.

12. The device of any one of claims 9-11, wherein, The acquisition module acquires the point cloud data by: Acquiring second two-dimensional image data of the face containing structured light modulation information, and obtaining the point cloud data according to the structured light modulation information in the second two-dimensional image data.

13. The device of any of claims 9-12, wherein, The three-dimensional data channel comprises a first channel, a second channel, a third channel and a fourth channel, wherein the first channel is used to store the X coordinate in the three-dimensional coordinate, the second channel is used to store the Y coordinate in the three-dimensional coordinate, the third channel is used to store the Z coordinate in the three-dimensional coordinate, and the fourth channel is used to store an effective identifier of the three-dimensional point, and the effective identifier is used to indicate whether the three-dimensional coordinate stored by the three-dimensional point is valid.

14. A face three-dimensional key point extraction device, characterized by, Comprise: A processor and a memory coupled to the processor, wherein the memory stores instructions for execution by the processor, and when the processor executes the instructions, the method of any one of claims 1-5 can be implemented.

15. A face three-dimensional model creating system characterized by comprising: Comprise: The face three-dimensional key point extraction device of any one of claims 9-14 is used to process left face data, front face data and right face data respectively, so as to obtain left face three-dimensional key points, front face three-dimensional key points and right face three-dimensional key points, wherein the left face data comprises first two-dimensional image data and corresponding left face point cloud data of the left face, the front face data comprises first two-dimensional image data and corresponding front face point cloud data of the front face, and the right face data comprises first two-dimensional image data and corresponding right face point cloud data of the right face; A first registration device is used to register the left face point cloud data and the front face point cloud data according to the left face three-dimensional key points and the front face three-dimensional key points, and register the right face point cloud data and the front face point cloud data according to the right face three-dimensional key points and the front face three-dimensional key points, so as to obtain a face three-dimensional model obtained by splicing the left face point cloud data, the front face point cloud data and the right face point cloud data.

16. The system of claim 15, wherein, The first registration device comprises: A first registration processing module is used to obtain a first key point pair located in a set face region, calculate rigid transformation matrix parameters between the left face point cloud data and the front face point cloud data according to the first key point pair, and register the left face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameters, wherein a left face three-dimensional key point and a front face three-dimensional key point at the same face position form a first key point pair; A second registration processing module is used to obtain a second key point pair located in the set face region, calculate rigid transformation matrix parameters between the right face point cloud data and the front face point cloud data according to the second key point pair, and register the right face point cloud data and the front face point cloud data according to the calculated rigid transformation matrix parameters, wherein a right face three-dimensional key point and a front face three-dimensional key point at the same face position form a second key point pair; The set face region comprises at least one of a brow region, a nose region, a mouth region and a chin region.

17. The system of claim 15 or 16, wherein, The system further comprises: A second registration device is used to register the face three-dimensional model by using an iterative closest point method.

18. A face three-dimensional model creating system characterized by comprising: Comprise: A processor, a memory coupled with the processor, wherein the memory has stored instructions for execution by the processor, and when the processor executes the instructions, the method of any one of claims 6-8 is implemented.

19. A computer readable storage medium storing a computer program, wherein the computer program comprises instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 18. The program is executed by the processor to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • 3D human face quick identity authentication method and apparatus

    CN108427871A

  • Point cloud splicing method and device

    CN110120013A

  • Three-dimensional face reconstruction method, apparatus and device, and storage medium

    CN110956691A

  • Head posture measuring method and device

    CN113544744A