Face irregular shape recognition method and device, terminal and storage medium
By calculating the dispersion of facial geometric features through a 3D face reconstruction network, the problem of facial anomaly recognition relying on doctors' experience is solved, improving recognition accuracy and diagnostic efficiency, and supporting early disease intervention.
Patent Information
- Application Number
- CN202310722579.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-16
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-06-16
AI Technical Summary
In existing technologies, facial deformity recognition relies on doctors' extensive experience, resulting in low recognition efficiency and high costs, making it difficult to provide timely intervention in the early stages of the disease.
By acquiring facial geometric features through a 3D face reconstruction network, calculating the dispersion of facial key points in the geometric features of the target region, and then determining the category of facial deformity, the system can provide auxiliary information for clinical diagnosis.
It improves the accuracy of facial anomaly recognition, helps doctors diagnose diseases early, reduces reliance on doctors' experience, and enables timely intervention.
Smart Images

Figure CN116705303B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital healthcare, and in particular to a method, apparatus, terminal device, and storage medium for facial anomaly recognition. Background Technology
[0002] In Traditional Chinese Medicine (TCM) diagnosis, facial abnormalities can be categorized into four types: facial swelling, cheek swelling, thinning face with prominent cheekbones, and facial asymmetry (drooping mouth and eyes). The appearance of these symptoms leads to abnormal facial features and can even disrupt a patient's daily life. Therefore, early intervention in the initial stages of the disease can be very beneficial. Furthermore, the results of tests identifying facial abnormalities can serve as supplementary information for clinical diagnosis, thus better assisting doctors in making informed decisions.
[0003] Accurate identification of facial deformities in users requires extensive clinical experience, but the gap between the number of users and the number of doctors is enormous, and the development of an experienced doctor requires significant time investment. Therefore, achieving accurate identification of facial deformities is of great importance. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, terminal device, and storage medium for recognizing facial deformities, aiming to improve the accuracy of facial deformity recognition and thus better assist doctors in diagnosing diseases.
[0005] In a first aspect, embodiments of this application provide a facial anomaly recognition method, applied to a terminal device, comprising:
[0006] A facial image of the target object is acquired and input into a 3D face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple facial patches. Each facial patch is determined by the coordinate information of three adjacent facial key points, and the multiple facial patches constitute the reconstructed 3D face model of the target object.
[0007] The geometric features of the target region corresponding to the region to be analyzed in the face image of the target object are determined from the facial geometric features.
[0008] The target coordinate information corresponding to the target facial key point is determined from the coordinate information of the facial key points in the geometric features of the target region, and the neighborhood coordinate information corresponding to the neighboring facial key points is determined from the facial geometric features of the target object based on the target coordinate information. The neighboring facial key points are adjacent to the target facial key points.
[0009] Calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determine the dispersion of facial key points in the geometric features of the target region based on the distance information.
[0010] The facial deformity category of the target object is determined based on the degree of dispersion.
[0011] Secondly, embodiments of this application also provide a facial anomaly recognition device, comprising:
[0012] The data acquisition module is used to acquire the face image of the target object and input the face image into the three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple face patches. Each face patch is determined by the coordinate information of three adjacent facial key points, and the multiple face patches constitute the reconstructed three-dimensional face model of the target object.
[0013] The data processing module is used to determine the geometric features of the target region corresponding to the region to be analyzed in the face image of the target object from the facial geometric features.
[0014] The data collection module is used to determine the target coordinate information corresponding to the target facial key point from the coordinate information of the facial key points in the geometric features of the target region, and to determine the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, wherein the neighboring facial key points are adjacent to the target facial key point.
[0015] The data calculation module is used to calculate the distance information between the normal of the target coordinate information and the normal of each of the neighboring coordinate information, and to determine the dispersion of facial key points in the geometric features of the target region based on the distance information.
[0016] The data analysis module is used to determine the facial deformity category of the target object based on the degree of dispersion.
[0017] Thirdly, embodiments of this application also provide a terminal device, which includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for connecting and communicating between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of any of the facial anomaly recognition methods provided in this application specification.
[0018] Fourthly, embodiments of this application also provide a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the facial anomaly recognition method as provided in any of the present application specifications.
[0019] This application provides a method, apparatus, terminal device, and storage medium for facial anomaly recognition. The method obtains the facial geometric features of a target object by passing a facial image through a 3D reconstruction network. Then, based on these facial geometric features, it obtains the geometric features of the target region corresponding to the area to be analyzed in the facial image. From the geometric features of the target region corresponding to the area to be analyzed, it selects target coordinate information corresponding to multiple target facial key points and determines the neighboring coordinate information of neighboring facial key points. It calculates the distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determines the dispersion of facial key points in the target region geometric features based on the distance information. Finally, it determines the facial anomaly category of the target object based on the dispersion. This allows for the determination of facial anomaly categories using the target coordinate information and the distance information calculated from the neighboring coordinate information in the user's facial geometric features, improving the accuracy of facial anomaly identification. The determination results can then be used as auxiliary information for clinical diagnosis, better assisting doctors in diagnosing diseases and enabling timely early intervention for users in the early stages of illness, providing better treatment. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a facial anomaly recognition method provided in this application embodiment;
[0022] Figure 2 This is a schematic diagram of a three-dimensional face model reconstructed using the facial anomaly recognition method provided in the embodiments of this application;
[0023] Figure 3 A schematic diagram illustrating the process of using a preset model to reconstruct a three-dimensional face from a face image in the facial anomaly recognition method provided in this application embodiment;
[0024] Figure 4 A schematic diagram of facial key points and neighboring facial key points in a 3D face model;
[0025] Figure 5 This is a schematic diagram of the module structure of a facial anomaly recognition device provided in an embodiment of this application;
[0026] Figure 6 This is a schematic block diagram of a terminal device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0029] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0030] This application provides a method, apparatus, terminal device, and storage medium for facial anomaly recognition. The facial anomaly recognition method can be applied to a terminal device, which can be a mobile phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or server. The server can be a standalone server or a server cluster.
[0031] This application provides a method, apparatus, terminal device, and storage medium for facial anomaly recognition. The method involves obtaining the facial geometric features of a target object by passing a facial image through a 3D reconstruction network. Then, based on these features, it obtains the geometric features of the target region corresponding to the area to be analyzed in the facial image. From the geometric features of the target region corresponding to the area to be analyzed, it selects target coordinate information corresponding to multiple target facial key points and determines the neighboring coordinate information of neighboring facial key points. It calculates the distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determines the dispersion of facial key points in the target region's geometric features based on the distance information. Finally, it determines the facial anomaly category of the target object based on the dispersion. This allows for the determination of facial anomaly categories using the target coordinate information and the distance information calculated from the neighboring coordinate information in the user's facial geometric features, improving the accuracy of facial anomaly identification. The determination results can then be used as auxiliary information for clinical diagnosis, better assisting doctors in diagnosing diseases and enabling timely early intervention for users in the early stages of illness, leading to better treatment.
[0032] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0033] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a facial anomaly recognition method provided in an embodiment of this application.
[0034] like Figure 1 As shown, the facial anomaly recognition method includes steps S1 to S5.
[0035] Step S1: Obtain the face image of the target object and input the face image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple face patches. Each face patch is determined by the coordinate information of three adjacent facial key points, and the multiple face patches constitute the reconstructed three-dimensional face model of the target object.
[0036] For example, by acquiring a facial image of a target object and inputting it into a 3D face reconstruction network, the 3D facial geometric features of the target object corresponding to the 2D facial image are obtained. Compared with a 2D plane, 3D space can reflect more information and allows for a more intuitive, comprehensive, and convenient observation of the target object from any viewpoint, providing a stereoscopic visual effect. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple facial patches, with each facial patch determined by the coordinate information of three adjacent facial key points. These multiple facial patches constitute the reconstructed 3D facial model of the target object.
[0037] For example, facial photos can be taken with a mobile phone, without strict limitations on pose, expression, lighting, or background environment. Inputting the facial image into a 3D face reconstruction network yields a reconstructed 3D face model represented by several facial landmarks and several face patches in a 3D Euclidean coordinate system. Each face patch consists of three facial landmarks. The number of facial landmarks and face patches is determined by the 3D reconstruction algorithm. For example, the ARKitFace dataset uses an iPhone to collect 1220 vertices and 2403 triangles, resulting in a 3D face model like this... Figure 2 As shown, the specific 3D reconstruction algorithm, the number of facial landmarks after 3D reconstruction, and the number of facial patches are not limited. The facial landmarks are as follows: Figure 2 As shown in Figure 101, the face image is as follows: Figure 2 As shown in Figure 102.
[0038] In some implementations, acquiring a facial image of a target object and inputting the facial image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object includes: receiving a facial image of the target object sent by an image acquisition device, wherein the facial image is an image obtained by the image acquisition device by acquiring an object image of the target object and cropping the facial region of the target object from the object image; and inputting the facial image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object.
[0039] For example, the image acquisition device acquires an image of the target object, obtains the face location information in the object image through image processing, and then extracts the face region of the object image based on the face location information to obtain the face image of the target object. The face image of the target object is then sent to the terminal device. After receiving the face image of the target object sent by the image acquisition device, the terminal device inputs the face image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object.
[0040] For example, an image acquisition device includes an image acquisition module and an image processing module. The image acquisition module is mainly used to capture images of the target object, while the image processing module is mainly used to process the images acquired by the image acquisition module to obtain a facial image of the target object. The image acquisition module can be a mobile phone or a camera, which captures images of the target object and sends them to the image processing module. After receiving the image of the target object from the image acquisition module, the image processing module performs face detection processing on the image to obtain the location and size of the face in the image. Based on the location and size of the face in the image, the facial region is cropped from the image to obtain the facial image of the target object, and then the facial image is sent to the terminal device.
[0041] Optionally, face detection can scale the object image to different sizes and then scan the object image with a window of the same size. That is, select the area bounded by the window on the object image as an observation object, and then slide the window to update the area bounded by the window. Then, feature extraction is performed on the area bounded by the window to obtain the feature vector corresponding to the area bounded by the window. Then, based on the feature vector, it is determined whether the area bounded by the window contains a face. If the area bounded by the window contains a face, the information of the area bounded by the window is converted into the position and size of the face in the object image. Otherwise, the window is continued to slide and the area bounded by the window is updated.
[0042] Optionally, determining whether the region selected by the window contains exactly one face based on the feature vector can be seen as classifying the region selected by the window based on the feature vector. The classification categories include face windows and non-face windows.
[0043] The target object's face image is obtained by cropping the image based on the position and size of the face in the object image, and then sent to the terminal device. After receiving the target object's face image sent by the image acquisition device, the terminal device inputs the face image into the 3D face reconstruction network to obtain the target object's facial geometric features.
[0044] In some implementations, inputting the face image into a 3D face reconstruction network to obtain the facial geometric features of the target object includes: inputting the face image into the image feature extraction network of the 3D face reconstruction network to obtain two-dimensional feature information of the face image and initialize a 3D face model for the face image; and continuously adjusting the 3D face model using the two-dimensional feature information based on the graph convolutional neural network of the 3D face reconstruction network to obtain a target 3D face model, wherein the target 3D face model is composed of the facial geometric features of the target object.
[0045] For example, spatial three-dimensional image information is obtained from two-dimensional face images. Compared with a two-dimensional plane, three-dimensional space can reflect more information and is more intuitive and comprehensive for people to accept. The face image is processed by an image feature extraction network to extract features and initialize a three-dimensional face model. Using a graph convolutional neural network of a three-dimensional face reconstruction network, the initialized three-dimensional face model is continuously adjusted according to the two-dimensional feature information, adding details to the three-dimensional model from coarse to fine, thereby transforming the initial three-dimensional face model into the target three-dimensional face model.
[0046] For example, such as Figure 3 As shown, firstly, an ellipsoid of a fixed size (e.g., with radii of 0.2m, 0.2m, and 0.8m on the three axes) is initialized for any input face image as its initial 3D face model; then, the face image is processed through a multi-layer convolutional neural network for feature extraction; and a cascaded mesh deformation module is designed, wherein the mesh deformation module is composed of a graph convolutional network, and a perceptual feature pooling layer is used to concentrate the image features of the two-dimensional projection of each node in the mesh deformation module together, and the node state of the graph convolutional network in the 3D mesh is adjusted using the two-dimensional image features. Finally, the initial ellipsoid is continuously deformed according to these features until it approaches a realistic 3D face model.
[0047] In some implementations, acquiring a facial image of a target object and inputting the facial image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object further includes: acquiring an object image of the target object from an image acquisition device and sending the object image to the terminal device; after receiving the object image, determining the facial region of the target object from the object image and cropping the facial region to obtain a facial image of the target object; and inputting the facial image into the three-dimensional face reconstruction network to obtain the facial geometric features of the target object.
[0048] For example, an image acquisition device acquires an image of the target object, and the image acquisition device sends the acquired image to a terminal device. After receiving the image, the terminal device obtains the face location information in the image, and then extracts the face region of the image based on the face location information to obtain the face image of the target object. The face image is then input into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object.
[0049] For example, the image acquisition device can be a mobile phone or a camera, which captures an image of the target object and sends it to the terminal device. After receiving the image of the target object from the image acquisition module, the terminal device performs face detection processing on the image to obtain the location and size of faces within it. Based on the location and size of the faces in the image, the terminal device extracts the face region to obtain the face image of the target object. This face image is then input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
[0050] Optionally, face detection can scale the object image to different sizes and then scan the object image with a window of the same size. That is, select the area bounded by the window on the object image as an observation object, and then slide the window to update the area bounded by the window. Then, feature extraction is performed on the area bounded by the window to obtain the feature vector corresponding to the area bounded by the window. Then, based on the feature vector, it is determined whether the area bounded by the window contains a face. If the area bounded by the window contains a face, the information of the area bounded by the window is converted into the position and size of the face in the object image. Otherwise, the window is continued to slide and the area bounded by the window is updated.
[0051] Optionally, determining whether the region selected by the window contains exactly one face based on the feature vector can be seen as classifying the region selected by the window based on the feature vector. The classification categories include face windows and non-face windows.
[0052] Step S2: Determine the target region geometric features corresponding to the region to be analyzed in the face image of the target object from the facial geometric features.
[0053] For example, in traditional Chinese medicine diagnosis, facial abnormalities can be divided into four categories: facial swelling, cheek swelling, thin face and prominent cheekbones, and facial asymmetry. Different facial abnormalities correspond to different facial areas. Therefore, when determining whether a target person has a certain facial abnormality, it is necessary to analyze the corresponding area.
[0054] For example, when the facial abnormality is facial swelling, the geometric features of the target region corresponding to the area to be analyzed are the two cheek areas of the face; when the facial abnormality is asymmetry of the mouth and eyes, the geometric features of the target region corresponding to the area to be analyzed include the area around the mouth and the area around the eyes.
[0055] For example, the facial region can be divided according to the facial structure. The left and right sides of the face can be obtained by dividing the face along the vertical midline of the nose. Then, the eyebrows, eyes, nose, and mouth can be divided into different regions using horizontal lines. When performing facial anomaly recognition, the corresponding divided regions can be set as the regions to be analyzed, thereby obtaining the geometric features of the target region corresponding to the regions to be analyzed.
[0056] Step S3: Determine the target coordinate information corresponding to the target facial key point from the coordinate information of the facial key points in the geometric features of the target region, and determine the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information. The neighboring facial key points are adjacent to the target facial key points.
[0057] For example, taking facial swelling around the cheekbone as an example, assuming that the visual appearance of facial swelling is round and approximately spherical, it will cause the normals of key points in the geometric features of the corresponding target area to tend to be discrete. If a key point and several of its neighboring key points are selected, the degree of dispersion of facial key points in the geometric features of the target area can be obtained by analyzing the key point and its several neighboring key points.
[0058] For example, the area around the cheekbone can be selected as the region to be analyzed. Based on the geometric features of the target region corresponding to the region to be analyzed, one point can be selected as a keypoint, and the surrounding 26 points can be selected as the corresponding neighborhood facial keypoints. The number of neighborhood facial keypoints can be selected depending on the specific characteristics of the target region to be analyzed; it is not limited to 26. Figure 4 As shown, the first colored circle indicates that the selected point is the facial key point, and the second colored circle indicates that the selected point is the neighboring facial key point adjacent to the selected facial key point. The colors of the first and second colored circles can be set as needed, such as the first colored circle being a gray circle and the second colored circle being a white circle.
[0059] In some implementations, determining target coordinate information corresponding to a target facial key point from the coordinate information of facial key points in the geometric features of the target region, and determining neighborhood coordinate information corresponding to neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, includes: determining first target coordinate information corresponding to a first target facial key point from the coordinate information of facial key points in the geometric features of the target region; determining first neighborhood coordinate information corresponding to a first neighboring facial key point from the facial geometric features of the target object based on the first target coordinate information; determining second target coordinate information corresponding to a second target facial key point from the coordinate information of facial key points in the geometric features of the target region; determining second neighborhood coordinate information corresponding to a second neighboring facial key point from the facial geometric features of the target object based on the second target coordinate information; and constructing the target coordinate information corresponding to the target facial key point by combining the first target coordinate information corresponding to the first target facial key point and the second target coordinate information corresponding to the second target facial key point.
[0060] For example, taking facial swelling around the cheekbone as an example, if a key point and several neighboring key points are selected, the dispersion of facial key points in the geometric features of the target area can be obtained by analyzing the key point and its neighboring key points. The selection of key points is random. In order to reduce the misjudgment rate of random selection, it is necessary to select key points and their corresponding neighboring key points multiple times, and then comprehensively judge whether the target object has facial abnormalities based on the analysis results of multiple key point selections.
[0061] For example, if the analysis initially determines that the target object does not have surface anomalies when selecting key points and their corresponding neighborhood coordinates within the geometric features of the target area to be analyzed, but then determines that the target object does have surface anomalies when selecting key points and their corresponding neighborhood coordinates a second time, then analyzing only one key point would result in a random judgment, significantly reducing the accuracy of surface anomaly identification. Therefore, it is necessary to select key points and their corresponding neighborhood key points multiple times, and then comprehensively judge whether the target object has surface anomalies based on the analysis results of multiple key point selections, thereby improving the accuracy of surface anomaly identification.
[0062] Step S4: Calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determine the dispersion of facial key points in the geometric features of the target region based on the distance information.
[0063] For example, taking facial swelling around the cheekbone as an example, assuming that the visual appearance of facial puffiness is round and approximately spherical, this will cause the normals of the corresponding area to tend to be dispersed. Select a key point and obtain the target coordinate information corresponding to the key point, as well as its neighboring key points and their corresponding neighboring coordinate information. Calculate the distance information between the normal of the target coordinate information and the normals of all neighboring coordinate information. If this distance information is sufficiently dispersed but not excessively dispersed, it indicates that the area is approximately spherical, i.e., facial swelling. Then, based on the distance information, determine the degree of dispersion of facial key points in the geometric features of the target area.
[0064] In some implementations, calculating the distance information between the normal of the target coordinate information and the normal of each of the neighboring coordinate information, and determining the dispersion degree of facial key points in the geometric features of the target region based on the distance information, includes: calculating the cosine distance between the normal of the target coordinate information and the normal of each of the neighboring coordinate information, calculating the cosine mean and cosine variance corresponding to the cosine distance, and using the cosine mean and cosine variance as the distance information; and calculating the dispersion degree of facial key points in the geometric features of the target region based on the distance information and a preset value.
[0065] For example, taking facial swelling around the cheekbones as an example, assuming that the visual appearance of facial puffiness is round and approximately spherical, this will cause the normals of the corresponding area to tend to be dispersed. Select a key point and obtain the target coordinate information corresponding to the key point, as well as the coordinate information of its neighboring key points and their corresponding neighboring key points. Calculate the distance between the normal of the target coordinate information and the normals of all neighboring coordinate information. If this distance is sufficiently dispersed but not excessively dispersed, it indicates that the area is approximately spherical, i.e., facial swelling. Cosine distance can be used as a distance metric, and the cosine mean and cosine difference can be calculated as measures of dispersion. If the normal of a key point is basically in the same direction as the normals of neighboring key points, then the cosine distance is approximately 0, the variance is approximately 0, and the area is approximately planar. If the normal of a key point differs greatly from the normals of neighboring key points, then the cosine distance is relatively large, the variance may be large, and the area is approximately angular.
[0066] For example, such as Figure 4 As shown, the gray circles represent keypoints, and the white circles represent neighboring facial keypoints. A keypoint was selected for facial puffiness detection near the right cheekbone, along with its 26 neighboring keypoints. The keypoint and its neighbors form 26 pairs of normals. The cosine distance (angle) of each pair of normals is calculated, denoted as theta1, theta2, ..., theta26, where theta_i = ...<normal_center,normal_neighbor_i> The expression `<>` represents the inner product, `normal_center` represents the normal direction of the keypoint, and `normal_neighbor_i` represents the normal direction of neighboring keypoints. The keypoint normal is calculated by the vector sum of the surface normals, which in turn are obtained by the vector product of the edge vectors. The edge vectors are obtained by the coordinate differences between points, i.e., the coordinate differences between the keypoint and each neighboring keypoint. After obtaining `theta1`, `theta2`, ..., `theta26`, the mean and variance of these 26 `theta` values are calculated.
[0067] Optionally, to improve robustness to noise, when calculating the mean and variance of the 26 theta values, the maximum and minimum values can be removed before calculating the mean and variance.
[0068] Step S5: Determine the facial deformity category of the target object based on the degree of dispersion.
[0069] For example, in TCM diagnosis, facial abnormalities can be divided into four categories: facial swelling, cheek swelling, thin face and prominent cheekbones, and facial asymmetry. The geometric features of the target area selected for each type of facial abnormality are different, and the degree of dispersion calculated based on the geometric features of the target area is also different, thus the conditions for judging facial abnormalities are also different.
[0070] For example, when selecting the left and right sides of the mouth as the target area, the cosine mean and cosine variance of each area on the left and right sides of the mouth are calculated. When the difference between the cosine mean and cosine variance of the two sides is greater than a certain threshold, it can be determined that the left and right sides of the mouth are very symmetrical, and then facial abnormalities, including asymmetry of the mouth and eyes, can be identified.
[0071] In some implementations, the terminal device stores a correlation between facial deformity categories and dispersion levels. Determining the facial deformity category of the target object based on the dispersion level includes: determining the facial deformity category of the target object based on the dispersion level and the correlation level, and outputting the facial deformity category.
[0072] For example, based on images of facial deformities accumulated from clinical work, facial deformities are categorized according to actual needs based on images and experience. Then, a numerical range representing the dispersion degree corresponding to each facial deformity category is calculated from the images of facial deformity images within that category. This establishes a correlation between facial deformity categories and dispersion degrees, which is then stored in the terminal device. As clinical data increases, the correlation between facial deformity categories and dispersion degrees is continuously adjusted and optimized. When performing facial deformity detection on a target subject's facial image, it is necessary to detect all facial deformity categories in the image to determine which categories the target subject suffers from.
[0073] Please see Figure 5 , Figure 5 A facial anomaly recognition device 200 provided in this application embodiment is applied to a terminal device. The facial anomaly recognition device 200 includes:
[0074] The data acquisition module 201 is used to acquire the face image of the target object and input the face image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple face patches. Each face patch is determined by the coordinate information of three adjacent facial key points, and the multiple face patches constitute the three-dimensional face model of the target object after reconstruction.
[0075] The data processing module 202 is used to determine the target region geometric features corresponding to the region to be analyzed in the face image of the target object from the facial geometric features.
[0076] The data collection module 203 is used to determine the target coordinate information corresponding to the target facial key point from the coordinate information of the facial key points in the geometric features of the target region, and to determine the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object according to the target coordinate information, wherein the neighboring facial key points are adjacent to the target facial key points.
[0077] The data calculation module 204 is used to calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and to determine the dispersion of facial key points in the geometric features of the target region based on the distance information.
[0078] Data analysis module 205 is used to determine the facial deformity category of the target object based on the degree of dispersion.
[0079] In some implementations, the data acquisition module 201, during the process of acquiring a facial image of the target object and inputting the facial image into a 3D face reconstruction network to obtain the facial geometric features of the target object, performs the following:
[0080] Receive a face image of a target object sent by an image acquisition device, wherein the face image is an image obtained by the image acquisition device by acquiring an object image of the target object and cropping the face region of the target object from the object image;
[0081] The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
[0082] In some implementations, the data acquisition module 201 performs the following steps during the process of inputting the face image into a 3D face reconstruction network to obtain the facial geometric features of the target object:
[0083] The face image is input into the image feature extraction network of the three-dimensional face reconstruction network to obtain the two-dimensional feature information of the face image and initialize a three-dimensional face model for the face image;
[0084] A graph convolutional neural network based on a 3D face reconstruction network continuously adjusts the 3D face model using the 2D feature information to obtain a target 3D face model, which is composed of the facial geometric features of the target object.
[0085] In some implementations, the data acquisition module 201, after acquiring a facial image of the target object and inputting the facial image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object, also performs the following:
[0086] The image acquisition device acquires an image of the target object and sends the image to the terminal device.
[0087] After receiving the object image, the face region of the target object is determined from the object image, and the face region is cropped to obtain the face image of the target object;
[0088] The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
[0089] In some implementations, the data collection module 203, in the process of determining the target coordinate information corresponding to the target facial key points from the coordinate information of facial key points in the geometric features of the target region, and determining the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, performs the following:
[0090] The coordinate information of the first target corresponding to the first target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region;
[0091] Based on the first target coordinate information, the first neighborhood coordinate information corresponding to the first neighborhood facial key points is determined from the facial geometric features of the target object;
[0092] The coordinate information of the second target corresponding to the second target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region;
[0093] Based on the second target coordinate information, the second neighborhood coordinate information corresponding to the second neighborhood facial key points is determined from the facial geometric features of the target object;
[0094] The first target coordinate information corresponding to the first target facial key point and the second target coordinate information corresponding to the second target facial key point constitute the target coordinate information corresponding to the target facial key point.
[0095] In some implementations, the data calculation module 204 performs the following steps during the process of calculating the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determining the dispersion of facial key points in the geometric features of the target region based on the distance information:
[0096] Calculate the cosine distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and calculate the cosine mean and cosine variance corresponding to the cosine distance, and use the cosine mean and cosine variance as the distance information;
[0097] The dispersion of facial key points in the geometric features of the target region is calculated based on the distance information and preset values.
[0098] In some specific embodiments, the terminal device stores the correlation between facial deformity categories and the degree of dispersion. During the process of determining the facial deformity category of the target object based on the degree of dispersion, the data analysis module 205 executes the following:
[0099] Based on the degree of dispersion and the correlation, the facial deformity category of the target object is determined and the facial deformity category is output.
[0100] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned embodiments of the facial anomaly recognition method, and will not be repeated here.
[0101] Please see Figure 6 , Figure 6 A schematic block diagram of the structure of a terminal device provided in an embodiment of this application.
[0102] like Figure 6 As shown, the terminal device 300 includes a processor 301 and a memory 302, which are connected by a bus 303, such as an I2C (Inter-integrated Circuit) bus.
[0103] Specifically, processor 301 provides computing and control capabilities to support the operation of the entire server. Processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0104] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0105] Those skilled in the art will understand that Figure 6 The structure shown in the figure is merely a block diagram of a portion of the structure related to the embodiments of this application, and does not constitute a limitation on the terminal device to which the embodiments of this application are applied. A specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0106] The processor 301 is used to run a computer program stored in the memory, and implements the facial anomaly recognition method provided in any embodiment of this application when executing the computer program.
[0107] In some implementations, the processor 301 is used to run a computer program stored in memory, applied to a terminal device, and performs the following steps when executing the computer program:
[0108] A facial image of the target object is acquired and input into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple facial patches. Each facial patch is determined by the coordinate information of three adjacent facial key points, and the multiple facial patches constitute the reconstructed three-dimensional face model of the target object.
[0109] Determine the target region geometric features corresponding to the region to be analyzed in the face image of the target object from the facial geometric features;
[0110] The target coordinate information corresponding to the target facial key point is determined from the coordinate information of the facial key points in the geometric features of the target region, and the neighborhood coordinate information corresponding to the neighboring facial key points is determined from the facial geometric features of the target object based on the target coordinate information. The neighboring facial key points are adjacent to the target facial key points.
[0111] Calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determine the dispersion degree of facial key points in the geometric features of the target region based on the distance information;
[0112] The facial deformity category of the target object is determined based on the degree of dispersion.
[0113] In some implementations, during the process of acquiring a facial image of a target object and inputting the facial image into a 3D face reconstruction network to obtain the facial geometric features of the target object, the processor 301 performs the following:
[0114] Receive a face image of a target object sent by an image acquisition device, wherein the face image is an image obtained by the image acquisition device by acquiring an object image of the target object and cropping the face region of the target object from the object image;
[0115] The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
[0116] In some implementations, during the process of inputting the face image into a 3D face reconstruction network to obtain the facial geometric features of the target object, the processor 301 performs the following:
[0117] The face image is input into the image feature extraction network of the three-dimensional face reconstruction network to obtain the two-dimensional feature information of the face image and initialize a three-dimensional face model for the face image;
[0118] A graph convolutional neural network based on a 3D face reconstruction network continuously adjusts the 3D face model using the 2D feature information to obtain a target 3D face model, which is composed of the facial geometric features of the target object.
[0119] In some implementations, during the process of acquiring the face image of the target object and inputting the face image into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object, the processor 301 also performs the following:
[0120] The image acquisition device acquires an image of the target object and sends the image to the terminal device.
[0121] After receiving the object image, the face region of the target object is determined from the object image, and the face region is cropped to obtain the face image of the target object;
[0122] The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
[0123] In some implementations, during the process of determining the target coordinate information corresponding to the target facial key points from the coordinate information of facial key points in the geometric features of the target region, and determining the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, the processor 301 performs the following:
[0124] The coordinate information of the first target corresponding to the first target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region;
[0125] Based on the first target coordinate information, the first neighborhood coordinate information corresponding to the first neighborhood facial key points is determined from the facial geometric features of the target object;
[0126] The coordinate information of the second target corresponding to the second target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region;
[0127] Based on the second target coordinate information, the second neighborhood coordinate information corresponding to the second neighborhood facial key points is determined from the facial geometric features of the target object;
[0128] The first target coordinate information corresponding to the first target facial key point and the second target coordinate information corresponding to the second target facial key point constitute the target coordinate information corresponding to the target facial key point.
[0129] In some implementations, the processor 301 performs the following steps during the process of calculating the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determining the dispersion of facial key points in the geometric features of the target region based on the distance information:
[0130] Calculate the cosine distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and calculate the cosine mean and cosine variance corresponding to the cosine distance, and use the cosine mean and cosine variance as the distance information;
[0131] The dispersion of facial key points in the geometric features of the target region is calculated based on the distance information and preset values.
[0132] In some implementations, the terminal device stores a correlation between facial deformity categories and the degree of dispersion. During the process of determining the facial deformity category of the target object based on the degree of dispersion, the processor 301 executes:
[0133] Based on the degree of dispersion and the correlation, the facial deformity category of the target object is determined and the facial deformity category is output.
[0134] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the corresponding process in the aforementioned embodiments of the facial anomaly recognition method, and will not be repeated here.
[0135] This application also provides a storage medium for computer-readable storage, which stores one or more programs that can be executed by one or more processors to implement the steps of any of the facial anomaly recognition methods provided in the embodiments of this application.
[0136] The storage medium can be an internal storage unit of the terminal device described in the aforementioned embodiments, such as the terminal device's memory. Alternatively, the storage medium can be an external storage device of the terminal device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.
[0137] It will be understood by those skilled in the art that all or some of the steps in the methods disclosed above, and the functional modules / units in the apparatus, can be implemented as software, firmware, hardware, and suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0138] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0139] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for recognizing facial irregularities, applied to a terminal device, characterized in that, The method includes: A facial image of the target object is acquired and input into a three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple facial patches. Each facial patch is determined by the coordinate information of three adjacent facial key points, and the multiple facial patches constitute the reconstructed three-dimensional face model of the target object. Determine the target region geometric features corresponding to the region to be analyzed in the face image of the target object from the facial geometric features; The target coordinate information corresponding to the target facial key point is determined from the coordinate information of the facial key points in the geometric features of the target region, and the neighborhood coordinate information corresponding to the neighboring facial key points is determined from the facial geometric features of the target object based on the target coordinate information. The neighboring facial key points are adjacent to the target facial key points. Calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determine the dispersion degree of facial key points in the geometric features of the target region based on the distance information; The facial deformity category of the target object is determined based on the degree of dispersion. Specifically, determining the target coordinate information corresponding to the target facial key points from the coordinate information of the facial key points in the geometric features of the target region, and determining the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information is as follows: The key points of the target face and their corresponding neighboring face key points were determined multiple times. The step of calculating the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and determining the dispersion of facial key points in the geometric features of the target region based on the distance information, includes: Calculate the cosine distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and calculate the cosine mean and cosine variance corresponding to the cosine distance, and use the cosine mean and cosine variance as the distance information; The dispersion of facial key points in the geometric features of the target region is calculated based on the distance information and preset values.
2. The method according to claim 1, characterized in that, The process of acquiring a facial image of the target object and inputting the facial image into a 3D face reconstruction network to obtain the facial geometric features of the target object includes: Receive a face image of a target object sent by an image acquisition device, wherein the face image is an image obtained by the image acquisition device by acquiring an object image of the target object and cropping the face region of the target object from the object image; The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
3. The method according to claim 2, characterized in that, The step of inputting the face image into a 3D face reconstruction network to obtain the facial geometric features of the target object includes: The face image is input into the image feature extraction network of the three-dimensional face reconstruction network to obtain the two-dimensional feature information of the face image and initialize a three-dimensional face model for the face image; A graph convolutional neural network based on a 3D face reconstruction network continuously adjusts the 3D face model using the 2D feature information to obtain a target 3D face model, which is composed of the facial geometric features of the target object.
4. The method according to claim 1, characterized in that, The step of acquiring a facial image of the target object and inputting the facial image into a 3D face reconstruction network to obtain the facial geometric features of the target object further includes: The image acquisition device acquires an image of the target object and sends the image to the terminal device. After receiving the object image, the face region of the target object is determined from the object image, and the face region is cropped to obtain the face image of the target object; The face image is input into a 3D face reconstruction network to obtain the facial geometric features of the target object.
5. The method according to claim 1, characterized in that, The step of determining the target coordinate information corresponding to the target facial key points from the coordinate information of the facial key points in the geometric features of the target region, and determining the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, includes: The coordinate information of the first target corresponding to the first target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region; Based on the first target coordinate information, the first neighborhood coordinate information corresponding to the first neighborhood facial key points is determined from the facial geometric features of the target object. The coordinate information of the second target corresponding to the second target facial key point is determined from the coordinate information of the facial key points of the geometric features of the target region; Based on the second target coordinate information, the second neighborhood coordinate information corresponding to the second neighborhood facial key points is determined from the facial geometric features of the target object; The first target coordinate information corresponding to the first target facial key point and the second target coordinate information corresponding to the second target facial key point constitute the target coordinate information corresponding to the target facial key point.
6. The method according to claim 1, characterized in that, The terminal device stores the correlation between facial deformity categories and the degree of dispersion. Determining the facial deformity category of the target object based on the degree of dispersion includes: Based on the degree of dispersion and the correlation, the facial deformity category of the target object is determined and the facial deformity category is output.
7. A facial anomaly recognition device, characterized in that, include: The data acquisition module is used to acquire the face image of the target object and input the face image into the three-dimensional face reconstruction network to obtain the facial geometric features of the target object. The facial geometric features include at least the coordinate information of multiple preset facial key points and multiple face patches. Each face patch is determined by the coordinate information of three adjacent facial key points, and the multiple face patches constitute the reconstructed three-dimensional face model of the target object. The data processing module is used to determine the geometric features of the target region corresponding to the region to be analyzed in the face image of the target object from the facial geometric features; The data collection module is used to determine the target coordinate information corresponding to the target facial key point from the coordinate information of the facial key points of the geometric features of the target region, and to determine the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, wherein the neighboring facial key points are adjacent to the target facial key point; The data calculation module is used to calculate the distance information between the normal of the target coordinate information and the normal of each neighboring coordinate information, and to determine the dispersion of facial key points in the geometric features of the target region based on the distance information; The data analysis module is used to determine the facial deformity category of the target object based on the degree of dispersion. Specifically, determining the target coordinate information corresponding to the target facial key points from the coordinate information of the facial key points in the geometric features of the target region, and determining the neighborhood coordinate information corresponding to the neighboring facial key points from the facial geometric features of the target object based on the target coordinate information, is as follows: The key points of the target face and their corresponding neighboring face key points were determined multiple times. The data calculation module is further used for: Calculate the cosine distance between the normal of the target coordinate information and the normal of each neighboring coordinate information, and calculate the cosine mean and cosine variance corresponding to the cosine distance, and use the cosine mean and cosine variance as the distance information; The dispersion of facial key points in the geometric features of the target region is calculated based on the distance information and preset values.
8. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and, in executing the computer program, implement the facial anomaly recognition method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the computer-readable storage medium is executed by one or more processors, the one or more processors perform the steps of facial anomaly recognition as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Traditional Chinese medicine remote facial consultation processing system and method based on multi-view image stereo reconstruction
CN110047597A
Three-dimensional model establishment method based on two-dimensional key points, computer and storage medium
CN115375835A