A face motion capture method, system, electronic device and medium
By combining the Facemoji universal interface and facial landmark matching algorithm with a target detection model, dynamic secondary mask-like images are generated, which solves the problem of accurate exposure of lesion areas in online diagnosis of skin disease patients and improves the diagnostic efficiency of doctors.
Patent Information
- Application Number
- CN202310330575.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-03-31
AI Technical Summary
When patients with skin diseases consult doctors online, it is difficult to accurately expose the affected areas on their face, which affects the efficiency of doctors' diagnosis. Moreover, existing methods of masking are not precise or convenient enough.
By using the Facemoji universal interface and facial landmark matching algorithm, combined with the target detection model, dynamic secondary mask-like images are generated to accurately cover non-lesion areas and expose lesion areas.
It enables precise detection and exposure of skin disease areas, improves doctors' consultation efficiency, and allows patients to accurately cover non-lesion areas even when their faces are moving dynamically.
Smart Images

Figure CN116403256B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial motion capture technology, and in particular to a facial motion capture method, system, electronic device, and medium for marking skin disease areas. Background Technology
[0002] Due to geographical location and other factors, many dermatology patients have a need to consult doctors online. However, when facial skin diseases occur, patients often feel anxious, ashamed, or unwilling to expose their facial information, making them reluctant to have face-to-face contact with doctors and unable to show their affected areas. This affects communication between doctors and patients, hindering doctors' understanding of the patient's condition.
[0003] In existing technologies, patients with skin diseases can cover and cover areas of their face other than the affected area, but this is mostly done manually and cannot accurately reveal the affected area. Furthermore, patients often turn their faces, requiring the covering to be moved simultaneously, which affects the doctor's efficiency in diagnosing the condition. Summary of the Invention
[0004] The purpose of this invention is to provide a method, system, electronic device, and medium for facial motion capture, which can dynamically detect faces and accurately reveal skin disease areas on the face.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] In a first aspect, the present invention provides a method for capturing facial motion, comprising:
[0007] Acquire dynamic facial information to be processed; the dynamic facial information to be processed includes multiple frames of facial images to be processed; each frame of the facial image to be processed includes lesion areas and non-lesion areas;
[0008] Based on the Facemoji universal interface, the face image to be processed in each frame is captured and masked to generate a primary mask image;
[0009] Based on the facial key point matching algorithm, key point matching is performed on the primary mask image to determine the facial key points to be processed.
[0010] The face image to be processed is input into the target detection model to obtain the lesion region to be processed; the target detection model is trained based on the training sample set and the neural network; each sample in the training sample set includes a face sample image and the lesion region corresponding to the face sample image;
[0011] Based on the lesion area to be processed, the key points of the face to be processed, and the primary mask-like image, a secondary mask-like image is determined; the secondary mask-like image is used to cover the non-lesion areas in the face image to be processed with a mask, while displaying the lesion areas in the face image to be processed.
[0012] Optionally, the step of performing keypoint matching on the primary mask-like image based on the facial keypoint matching algorithm to determine the facial keypoints to be processed specifically includes:
[0013] The initial mask-like image is input into a convolutional neural network for feature extraction to obtain a first face feature map;
[0014] Perform Fourier pooling on the first face feature map to obtain the second face feature map;
[0015] The second facial feature map is input into a pre-trained logistic regression classifier to obtain the facial key points to be processed.
[0016] Optionally, the facial motion capture method further includes:
[0017] The Delaunay algorithm is used to perform triangulation on the key points of the face to be processed, so as to obtain the triangular network of the face to be processed.
[0018] Optionally, the training process of the target detection model includes:
[0019] Acquire multiple face sample images;
[0020] Each of the aforementioned face sample images is labeled to determine the lesion region corresponding to the face sample image; the multiple face sample images and the lesion region corresponding to each face sample image constitute a training sample set;
[0021] The training sample set is input into the neural network model for training to obtain the optimal neural network model; the optimal neural network model is the object detection model.
[0022] Optionally, a secondary mask-like image is determined based on the lesion area to be treated, the facial key points to be treated, and the primary mask-like image, specifically including:
[0023] The lesion area to be processed is matched with the facial key points to be processed in order to determine the key points of the lesion.
[0024] Based on the key points of the lesion, a lesion image segmentation operation is performed on the primary mask-like image to obtain a secondary mask-like image.
[0025] Secondly, the present invention provides a facial motion capture system, comprising:
[0026] The dynamic information acquisition module is used to acquire dynamic information of the face to be processed; the dynamic information of the face to be processed includes multiple frames of face images to be processed; each frame of the face image to be processed includes lesion areas and non-lesion areas;
[0027] The mask image generation module is used to capture and mask each frame of the face image to be processed based on the Facemoji universal interface to generate a primary mask image.
[0028] The facial key point determination module is used to perform key point matching on the primary mask-like image based on the facial key point matching algorithm to determine the facial key points to be processed.
[0029] A facial lesion determination module is used to input the facial image to be processed into a target detection model to obtain the lesion region to be processed; the target detection model is trained based on a training sample set and a neural network; each sample in the training sample set includes a facial sample image and the lesion region corresponding to the facial sample image;
[0030] A facial lesion mask display module is used to determine a secondary mask image based on the lesion area to be processed, the key points of the face to be processed, and the primary mask image; the secondary mask image is used to cover the non-lesion areas in the face image to be processed with a mask, while displaying the lesion area in the face image to be processed.
[0031] Thirdly, the present invention provides an electronic device, the electronic device comprising a memory and a processor;
[0032] The memory is used to store computer programs, and the processor is used to run the computer programs to perform a face motion capture method.
[0033] A computer-readable storage medium storing a computer program;
[0034] The computer program implements the steps of the face motion capture method when executed by the processor.
[0035] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0036] This invention provides a method, system, electronic device, and medium for facial motion capture. It processes detected facial motion information frame by frame, capturing and masking each frame of the face image based on the FaceMoji universal interface to obtain a primary mask-like image. Because it uses the FaceMoji universal interface, image processing is more convenient. Then, a facial keypoint matching algorithm is used to determine the keypoints of the face to be processed, and a target detection model is used to obtain the corresponding lesion region in the face image. Combined with the primary mask-like image, a secondary mask-like image is obtained that can mask non-lesion regions of the face while simultaneously displaying lesion regions. This invention can accurately detect and expose skin disease areas while masking other non-skin disease areas, achieving accurate facial motion capture with skin disease area marking. Furthermore, this invention acquires facial motion information, performing the above detection and processing on each frame of the face image, thus enabling accurate facial masking even when a skin disease patient turns their face, further improving the efficiency of doctors' consultations. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a flowchart illustrating the face motion capture method of the present invention;
[0039] Figure 2 This is a schematic diagram of the facial motion capture system of the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] This invention proposes a facial motion capture method, system, electronic device, and medium. The face of a skin disease patient is enveloped by a virtual avatar mask, allowing them to communicate with a doctor about their condition without exposing their own face. The system dynamically displays the affected area without exposing other facial areas.
[0042] To make the objectives, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0043] Example 1
[0044] like Figure 1 As shown, this embodiment provides a face motion capture method, including:
[0045] Step 100: Obtain dynamic facial information to be processed; the dynamic facial information to be processed includes multiple frames of facial images to be processed; each frame of the facial image to be processed includes lesion areas and non-lesion areas. Among them, lesion areas include rosacea, folliculitis, etc. on the face.
[0046] Specifically, video data of the patient and doctor is acquired during video consultations for dermatology patients.
[0047] Step 200: Capture and mask each frame of the face image to be processed based on the FaceMoji universal interface to generate a primary mask-like image.
[0048] The Facemoji API is used to capture each frame of a dermatologist's face in real time during video calls between the patient and doctor, generating a dynamic mask-like camouflage effect, rather than personalizing static images. Furthermore, the specific mask-like dynamic camouflage animation effect can be chosen by the dermatologist themselves.
[0049] Step 300: Based on the facial landmark matching algorithm, landmark matching is performed on the initial mask-like image to determine the facial landmarks to be processed; after obtaining the anime character-style face, the facial landmark matching algorithm is used to determine the specific location of facial features, such as the nose and mouth. The facial landmark matching algorithm is based on a convolutional neural network.
[0050] Step 300 specifically includes:
[0051] 1) Input the primary mask-like image into a convolutional neural network for feature extraction to obtain the first face feature map.
[0052] 2) Perform Fourier pooling on the first face feature map to compress the first face feature map to a one-dimensional frequency domain to obtain the second face feature map, thereby better capturing the nonlinear features of the patient's facial information.
[0053] 3) The second facial feature map is input into a pre-trained logistic regression classifier to obtain the facial key points to be processed. The pre-trained logistic regression classifier is linear and can map the captured non-linear features (i.e., the second facial feature map) onto the facial key points, thereby obtaining the key point information of the target image. Among them, the facial key points include the position of the nose, the position of the mouth, and the position of the eyes, etc.
[0054] The facial motion capture method also includes:
[0055] After obtaining the locations of facial key points, the Delaunay algorithm is used to perform triangulation on the facial key points to be processed, resulting in a triangular mesh of the face to be processed. In this triangular mesh, the face is divided into multiple triangular regions.
[0056] Step 400: Input the face image to be processed into the target detection model to obtain the lesion region to be processed; the target detection model is trained based on the training sample set and the neural network; each sample in the training sample set includes a face sample image and the lesion region corresponding to the face sample image.
[0057] The training process of the target detection model includes:
[0058] 1) Obtain multiple face sample images.
[0059] 2) Each of the face sample images is labeled to determine the lesion region corresponding to the face sample image; the multiple face sample images and the lesion region corresponding to each face sample image constitute a training sample set. Specifically, the face sample images are manually labeled.
[0060] 3) Input the training sample set into the neural network model for training to obtain the optimal neural network model; the optimal neural network model is the target detection model.
[0061] Step 500: Based on the lesion area to be processed, the key points of the face to be processed, and the primary mask-like image, a secondary mask-like image is determined. The secondary mask-like image is used to mask the non-lesion areas in the face image to be processed while simultaneously displaying the lesion areas in the face image to be processed.
[0062] Step 500 specifically includes:
[0063] 1) Match the lesion area to be treated with the facial key points to be treated to determine the key points of the lesion. For example, if the lesion is on the cheek, the edge key points corresponding to the lesion area on the cheek are drawn out by the facial key points to be treated, and then the facial key points inside the lesion area are found.
[0064] 2) Based on the key points of the lesion, perform lesion image segmentation on the primary mask-like image to obtain a secondary mask-like image.
[0065] In a specific application, after the lesion area to be treated is determined, when a dermatologist clicks on the facial capture area (primary mask-like image), only the actual facial features related to the lesion are revealed for the doctor to diagnose. For example, if a patient has rosacea on their cheek, when the patient clicks on their cheek, the actual skin containing the lesion (including rosacea) will be displayed in the video, while the rest of the face is covered by the facial capture mask. Compared to static, animated image generation, this invention allows the lesion area to be presented to the doctor from multiple angles as the patient moves their head during the video chat, without exposing all of the patient's facial information.
[0066] Example 2
[0067] like Figure 2 As shown, in order to execute the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, this embodiment provides a face motion capture system, including:
[0068] The dynamic information acquisition module 101 is used to acquire dynamic information of the face to be processed; the dynamic information of the face to be processed includes multiple frames of face images to be processed; each frame of the face image to be processed includes lesion areas and non-lesion areas.
[0069] The mask image generation module 201 is used to capture and mask each frame of the face image to be processed based on the FaceMoji universal interface to generate a primary mask image.
[0070] The facial key point determination module 301 is used to perform key point matching on the primary mask-like image based on the facial key point matching algorithm to determine the facial key points to be processed.
[0071] The facial lesion determination module 401 is used to input the facial image to be processed into the target detection model to obtain the lesion region to be processed; the target detection model is trained based on the training sample set and the neural network; each sample in the training sample set includes a facial sample image and the lesion region corresponding to the facial sample image.
[0072] The facial lesion mask display module 501 is used to determine a secondary mask image based on the lesion area to be processed, the key points of the face to be processed, and the primary mask image; the secondary mask image is used to cover the non-lesion area in the face image to be processed with a mask, while displaying the lesion area in the face image to be processed.
[0073] Example 3
[0074] This embodiment provides an electronic device, including a memory and a processor.
[0075] The memory is used to store computer programs, and the processor is used to run the computer programs to execute the face motion capture method in Embodiment 1.
[0076] Optionally, the electronic device is a server.
[0077] A computer-readable storage medium storing a computer program; when executed by a processor, the computer program implements the steps of the face motion capture method in Embodiment 1.
[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0079] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for capturing facial motion, characterized in that, The facial motion capture method includes: Acquire dynamic facial information to be processed; the dynamic facial information to be processed includes multiple frames of facial images to be processed; each frame of the facial image to be processed includes lesion areas and non-lesion areas; Based on the Facemoji universal interface, the face image to be processed in each frame is captured and masked to generate a primary mask image. Based on a facial landmark matching algorithm, landmark matching is performed on the initial mask-like image to determine the facial landmarks to be processed; specifically including: The initial mask-like image is input into a convolutional neural network for feature extraction to obtain a first face feature map; Fourier pooling is performed on the first face feature map to obtain a second face feature map; the second face feature map is input into a pre-trained logistic regression classifier to obtain the key points of the face to be processed. The face image to be processed is input into the target detection model to obtain the lesion region to be processed; the target detection model is trained based on the training sample set and the neural network; each sample in the training sample set includes a face sample image and the lesion region corresponding to the face sample image; Based on the lesion area to be processed, the key points of the face to be processed, and the primary mask-like image, a secondary mask-like image is determined; the secondary mask-like image is used to cover the non-lesion areas in the face image to be processed with a mask, while displaying the lesion areas in the face image to be processed. Based on the lesion region to be processed, the facial key points to be processed, and the primary mask-like image, a secondary mask-like image is determined, specifically including: matching the lesion region to be processed with the facial key points to be processed to determine the lesion key points; and performing lesion image segmentation on the primary mask-like image based on the lesion key points to obtain the secondary mask-like image.
2. The facial motion capture method according to claim 1, characterized in that, The facial motion capture method also includes: The Delaunay algorithm is used to perform triangulation on the key points of the face to be processed, so as to obtain the triangular network of the face to be processed.
3. The facial motion capture method according to claim 1, characterized in that, The training process of the target detection model includes: Acquire multiple face sample images; Each of the aforementioned face sample images is labeled to determine the lesion region corresponding to the face sample image; the multiple face sample images and the lesion region corresponding to each face sample image constitute a training sample set; The training sample set is input into the neural network model for training to obtain the optimal neural network model; the optimal neural network model is the object detection model.
4. A facial motion capture system, characterized in that, The facial motion capture system includes: The dynamic information acquisition module is used to acquire dynamic information of the face to be processed; the dynamic information of the face to be processed includes multiple frames of face images to be processed; each frame of the face image to be processed includes lesion areas and non-lesion areas; The mask image generation module is used to capture and mask each frame of the face image to be processed based on the Facemoji universal interface to generate a primary mask image. A facial landmark determination module is used to perform landmark matching on the primary mask-like image based on a facial landmark matching algorithm to determine the facial landmarks to be processed; specifically, it includes: The initial mask-like image is input into a convolutional neural network for feature extraction to obtain a first face feature map; Fourier pooling is performed on the first face feature map to obtain a second face feature map; the second face feature map is input into a pre-trained logistic regression classifier to obtain the key points of the face to be processed. A facial lesion determination module is used to input the facial image to be processed into a target detection model to obtain the lesion region to be processed; the target detection model is trained based on a training sample set and a neural network; each sample in the training sample set includes a facial sample image and the lesion region corresponding to the facial sample image; A facial lesion mask display module is used to determine a secondary mask image based on the lesion area to be processed, the key points of the face to be processed, and the primary mask image; the secondary mask image is used to cover the non-lesion area in the face image to be processed with a mask, while displaying the lesion area in the face image to be processed. Based on the lesion region to be processed, the facial key points to be processed, and the primary mask-like image, a secondary mask-like image is determined, specifically including: matching the lesion region to be processed with the facial key points to be processed to determine the lesion key points; and performing lesion image segmentation on the primary mask-like image based on the lesion key points to obtain the secondary mask-like image.
5. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store a computer program, and the processor is used to run the computer program to perform the face motion capture method according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program; When the computer program is executed by the processor, it implements the steps of the face motion capture method according to any one of claims 1-3.
Citation Information
Patent Citations
Face cartoonalization method and device and computer storage medium
CN112907708A
Face image protection method and device, electronic equipment and readable storage medium
CN115272534A
API engine for discrimination of facial skin disease based on artificial intelligence that discriminates skin disease by using image captured through facial skin photographing device
KR102041906B1
Method and apparatus for protecting privacy of ophthalmic patient and storage medium
US20230076853A1