Facial nerve paralysis detection method and device based on dynamic structured light technology
Through the facial nerve paralysis detection method based on dynamic structured light technology, facial point clouds are collected in real time and topological symmetry models are generated, which solves the shortcomings of facial muscle deep movement and whole-face muscle movement evaluation in the existing technology, and achieves more accurate facial paralysis diagnosis and evaluation.
Patent Information
- Application Number
- CN202510131655.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-13
AI Technical Summary
There are errors in the existing diagnosis methods for facial paralysis, especially the inability to effectively evaluate the depth of facial muscles and the lack of movement of the entire facial muscles, resulting in inaccurate diagnosis results.
The facial nerve paralysis detection method based on dynamic structured light technology is adopted to collect high-precision facial point clouds in real time, generate topological symmetry models, and calculate the motion vector of each point to quantify facial symmetry, providing objective diagnostic tools.
It improves the accuracy of facial paralysis assessment and provides an objective and visual diagnostic tool, which improves the efficiency and accuracy of patient visits and follow-up visits.
Smart Images

Figure CN120147229A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and particularly relates to a method and device for detecting facial paralysis based on dynamic structured light technology. Background Art
[0002] Facial paralysis is a facial nerve disease caused by functional disorders of muscle movement. The main symptoms include facial stiffness, deviation of the mouth and eyes, movement of the frontal wrinkles, and incomplete closure of the eyelids, etc., which have a great impact on the lives of patients. The treatment of facial paralysis requires long-term monitoring and evaluation. The traditional evaluation method is that doctors guide patients to control different facial muscle movements according to the location and severity of facial paralysis, and judge the severity of facial paralysis through the symmetry and amplitude of the movements. In this facial paralysis diagnosis method, the evaluation of the severity of facial paralysis in patients is based on the subjective judgment of doctors, which has uncertainty. Therefore, in order to improve the accuracy of the evaluation of the severity of facial paralysis in patients, an objective and quantifiable method for evaluating facial paralysis is very necessary.
[0003] Currently, non-traditional methods for diagnosing facial paralysis can generally be divided into two categories: one is the method based on 2D image sequences. This method uses a camera to continuously capture the patient's face at a fixed distance to form a 2D image sequence, and then uses deep learning methods or manual annotation of key points for each 2D face image, and then processes these key points for diagnosis. However, this type of method can only evaluate horizontal and vertical movements, and cannot evaluate the movement of these points in depth. In addition, these key points cannot cover the muscle groups of the entire face and cannot evaluate the movement of the entire facial muscles. Therefore, the diagnostic results obtained by this type of method may have a large error due to the movement of facial muscles in depth and too few key points.
[0004] The other type of method is the method based on 3D models. This type of method uses a depth sensor to scan the human face to obtain a sequence of human face images of a specified action, and then uses the relative distance between the patient's facial 3D model and its mirror model as an indication of facial symmetry, and then evaluates facial asymmetry. However, the human face is asymmetric, and using this method to diagnose patients with inherently asymmetric faces will produce a large error. For the case of dynamic evaluation (collecting 3D models of patients making continuous expressions), this method cannot reflect the facial movement situation either. In addition, the depth sensors used in these methods generally have low precision and sparse point clouds. Analyzing with low-quality point clouds will undoubtedly also produce a large error. Summary of the Invention
[0005] The embodiments of the present application provide a method and device for detecting facial paralysis based on dynamic structured light technology, which can reconstruct the human face with high precision and good imaging integrity; use a non-rigid registration method to complete the transformation from the standard face model to the actual face model; and calculate the motion vectors of each point in the transformed topological symmetric model, and quantify facial symmetry through this information. Subsequently, an objective visualization tool is provided for doctors to diagnose facial paralysis, improving the efficiency and accuracy of patient visits and follow-up consultations.
[0006] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method for detecting facial paralysis based on dynamic structured light technology, including the following steps: First, perform real-time data collection on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model; then, perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face; the 3D model of the patient's face is a topological symmetric model; finally, calculate the motion vectors of each point in the 3D model of the patient's face, and quantify facial symmetry based on the symmetry operator.
[0007] In some exemplary embodiments, performing real-time data collection on the patient's face to obtain a high-precision and complete facial point cloud includes: using three sets of structured light systems to simultaneously reconstruct the human face from three directions to obtain a high-precision and complete facial point cloud.
[0008] In some exemplary embodiments, each set of structured light systems includes a projector and a camera, and the projector and the camera are connected by a trigger wire; the projectors of the three sets of structured light systems are connected by an external trigger; after the external trigger sends a signal, the three projectors are simultaneously triggered to project a set of the above-mentioned codes; after each projector projects a code, the camera is triggered to take a coded image to ensure synchronous control of the three systems.
[0009] In some exemplary embodiments, performing real-time data collection on the patient's face includes: the patient starts from a non-expression state and makes a specified symmetric expression according to the instruction, and during this process, the point cloud data of all frames is collected; the specified symmetric expressions include frowning, closing the eyes to the maximum extent, and grinning to the maximum extent.
[0010] In some exemplary embodiments, performing non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face includes: using a non-rigid registration algorithm to deform the template surface of the mesh model into the same shape as the target surface; by optimizing the loss function, find the best deformation position, and achieve the entire non-rigid registration through iteration to unify the topological structures of the standard human face and the patient's facial model.
[0011] In some exemplary embodiments, a non-rigid registration algorithm is adopted to deform the template surface of the mesh model into the same shape as the target surface, including: assuming that S is the template surface, T is the target surface, n is the number of points on the template surface, and S(X) is the affine transformation relationship from the entire surface S to the surface T; for each vertex v i After undergoing a local affine transformation X i After continuous iteration, it will move to a place u closer to the surface T i , thereby obtaining an optimal deformation from the template surface S to the target surface T; the optimization loss function is as follows:
[0012] E = E d (X) + αE s (X) + βE l (X) (1)
[0013] where E d (X) is the data loss function, E S (X) is the smoothing loss function, and E l (X) is the landmark loss function;
[0014] The data loss function is as follows:
[0015]
[0016] where v i is the i-th vertex in the template surface, with a total of n; X i is the displacement transformation matrix of each vertex, and ω i is the weight; the data loss function is used to find the point on the target mesh closest to the point v i , calculate the distance between the two, and then make the points in the template surface as close as possible to the target surface;
[0017] The smoothing loss function is used to measure whether the template surface after transformation is smooth; the smoothing loss function is as follows:
[0018]
[0019] The landmark loss function is used to adjust according to the preset feature points to make the feature points in the template surface fit those in the target surface; the landmark loss function is as follows:
[0020]
[0021] where G is a parameter used to measure rotation and translation, L is the marked feature point, and l is the normal direction vector of the facial symmetry plane.
[0022] In some exemplary embodiments, by optimizing the loss function, the best deformation position is found, and the entire non-rigid registration is achieved through iteration to unify the topological structures of the standard face and the patient's facial model, including: using non-rigid registration to transform the standard face I 0 as the source model, and the expressionless face model T of the patient 0 as the target model, deforming the standard face to complete the fitting of the expressionless face model S 0 ; for subsequent frames of the 3D model sequence, taking the model collected in the current frame as the target model, and combining with 3D key points, using the deformed model of the previous frame as the source model for deformation; among them, the expressionless face model T of the patient 0 is the first frame in the 3D model sequence.
[0023] In some exemplary embodiments, calculating the motion vector of each point in the patient's facial 3D model includes: obtaining the set of motion directions and the set of distances of each point relative to the first topological symmetric model; mirroring the expressionless topological symmetric model, and calculating the distances between all symmetric points between the two models, and screening out the parts with high facial symmetry by setting a threshold; fitting the facial symmetry plane through the symmetric points of the parts with high facial symmetry; using the HouseHolder operator to mirror the moving direction of the left face part in the topological symmetric model using the symmetric plane method phase direction; after mirroring the left face points in the topological symmetric model, obtaining the moving vector of each point in the model relative to the expressionless topological symmetric model, and transforming the facial symmetry problem into the vector similarity problem of corresponding points; the moving vector of each point in the model relative to the expressionless topological symmetric model is as follows:
[0024]
[0025] where, H i is the moving vector of each point relative to the expressionless topological symmetric model S 0 ; is the moving direction of mirroring the moving direction of the left face part in the topological symmetric model using the symmetric plane method phase direction; is the moving direction of the left face part in the topological symmetric model; N is the number of points on the left and right faces in the point cloud set of the standard face model, and M is the number of points on the midline.
[0026] In some exemplary embodiments, after transforming the facial symmetry problem into the vector similarity problem of corresponding points, constraint conditions are used to evaluate the similarity of vectors to obtain a symmetry operator based on the topological structure symmetric model;
[0027] The constraint conditions are:
[0028] 0 = f (α,α)<f (α,β) = f (β,α) <1 (6)
[0029] The symmetry operator based on the topological structure symmetry model is expressed as:
[0030]
[0031] where ε is a parameter for adjusting the weights of angles and distances; λ and γ are parameters for adjusting the rising rate and the symmetry center respectively.
[0032] In a second aspect, the embodiments of the present application further provide a facial paralysis detection device based on dynamic structured light technology, which uses the facial paralysis detection method based on dynamic structured light technology in the above embodiments for detection, including: a face 3D imaging module, a topology unification module, and a facial symmetry analysis module connected in sequence; wherein, the face 3D imaging module is used to perform real-time data acquisition on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, a mesh model is obtained; the topology unification module is used to perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard face, and the 3D model of the patient's face is a topological symmetry model; the facial symmetry analysis module is used to calculate the motion vector of each point in the 3D model of the patient's face and quantify the facial symmetry based on the symmetry operator.
[0033] The technical solutions provided by the embodiments of the present application have at least the following advantages:
[0034] The embodiments of the present application provide a facial paralysis detection method and device based on dynamic structured light technology. The detection method includes the following steps: First, perform real-time data acquisition on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model; then, perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard face; the 3D model of the patient's face is a topological symmetry model; finally, calculate the motion vector of each point in the 3D model of the patient's face and quantify the facial symmetry based on the symmetry operator.
[0035] The facial paralysis detection method and device based on dynamic structured light technology provided by this application. The hardware device of this application is developed based on the principle of coded structured light three-dimensional scanning and is mainly used to obtain a complete sequence of 3D face models in real time. To ensure the integrity of imaging, this application uses three structured light devices to reconstruct the face from different angles. Since the real-time dynamic structured light system projects codes onto the object to be measured quickly, considering that the fast-strobing visible light will cause discomfort to the human eye, the light source of this system uses infrared light that is invisible to the human eye. In addition, to avoid interference between the structured light projected by different devices on the face, three infrared projection light sources with different wavelength bands are used; the real-time three-dimensional point cloud reconstruction is realized based on the built-in GPU module of the imaging device, and the face point clouds from three perspectives are stitched into a complete face point cloud model through the pre-calibrated position parameters. This application asks the patient to make a specified expression from a non-expression state, and the 3D models are collected during this process.
[0036] Compared with the method of deforming a standard face by using deep learning to judge facial shapes and parameters, this application uses the truly collected facial 3D model sequence and can better perceive the details of facial movements. This application uses time-coded infrared structured light to reconstruct the burned area. By reconstructing the face from three angles, a high-precision, high-frame-rate, and complete 3D model of the patient's face can be obtained. Compared with 2D methods, low-precision 3D model methods, and incomplete 3D model methods, this application can obtain high-precision data that more conforms to the actual movement law. This application takes into account the natural asymmetry of the human face. The facial symmetry plane is obtained by fitting facial points with strong symmetry, rather than simply directly analyzing the symmetry between the facial model and its mirror model. This method of this application can perform point-by-point motion symmetry analysis on all regions of the face and is also applicable to patients and doctors to evaluate the prognosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] One or more embodiments are illustrated by way of example in the accompanying drawings, and these illustrative descriptions do not limit the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.
[0038] Figure 1 It is a schematic flow chart of a facial paralysis detection method based on dynamic structured light technology provided by an embodiment of this application.
[0039] Figure 2 It is a schematic structural diagram of a facial paralysis detection device based on dynamic structured light technology provided by an embodiment of this application.
[0040] Figure 3 It is a schematic diagram of a structured light acquisition device provided by an embodiment of this application.
[0041] Figure 4The flowchart of automatic registration provided by an embodiment of the present application.
[0042] Figure 5 The schematic diagram of the registration effect provided by an embodiment of the present application.
[0043] Figure 6 The schematic diagram of the standard face topological structure provided by an embodiment of the present application.
[0044] Figure 7 The heat map for quantifying facial symmetry provided by an embodiment of the present application. Detailed implementation manners
[0045] As can be seen from the background art, there are problems in the diagnosis results obtained by the existing facial paralysis diagnosis methods, such as large errors in the diagnosis results due to the movement of facial muscles in depth and too few key points. Dynamic assessment cannot reflect the facial movement situation either. In addition, the depth sensors used in these methods generally have low accuracy and sparse point clouds. Analyzing with low-quality point clouds will undoubtedly also generate large errors.
[0046] Facial paralysis is the sixth most common neurological disease, with a prevalence of approximately 0.03% in the population, and the incidence is still increasing year by year. Among these patients, only 15.1% of the patients are treated in time and cured. The vast majority of facial paralysis patients have their conditions aggravated and have difficulty in prognosis due to the inability to be diagnosed and treated in time. In addition, the treatment of facial paralysis is a long-term process. When doctors diagnose and evaluate the rehabilitation situation, they mostly use the facial paralysis assessment scale, requiring patients to make different facial movements, and evaluating the function of the peripheral facial nerve by observing the facial movement situation. This diagnostic method is mainly the subjective diagnosis of doctors and is extremely dependent on their own experience, which is not objective and intuitive. Currently, the diagnosis of facial paralysis based on the 3D face model also has certain applications in the industry. Its basic strategy is to use a depth sensor to obtain multiple frames of 3D face models with different expressions, and directly perform symmetry analysis on the original model and the mirror model of the face to diagnose and identify facial paralysis.
[0047] A related technology proposes an artificial intelligence-based facial paralysis degree assessment system, which includes an image acquisition module, an image cropping module, a 3D model construction module of the face to be measured, etc. The system finds the face model with the highest similarity in the face database as the normal face according to the constructed 3D face model of the measured person, and obtains the shape parameters of the 3D face model of the measured person. The parametric face is deformed according to this parameter. Then, a clustering algorithm is used to divide the ROI regions of the standard 3D face model and the 3D face model to be measured according to weights, and the assigned weights and the center point distances of the same ROI regions of the two faces are calculated. Finally, the degree of facial paralysis is evaluated according to the said weights and distances.
[0048] Another related technology proposes a modeling, grading method and system for self-supervised pre-training and facial paralysis grading. The method includes the following steps: collecting and processing facial paralysis data; a facial paralysis grade evaluation algorithm based on the combination of multi-convolution features and video frame context information; this algorithm mainly includes the following steps: using the potplay software to frame the collected facial paralysis video data, and then performing related processing on this series of frames using a unified standard; when using deep learning for facial paralysis recognition and facial paralysis grade evaluation, first use the series of frames preprocessed in the first step as the input of the deep learning network, and use the powerful sample essential feature extraction ability of deep learning to learn and extract features from this series of frames.
[0049] Another related technology proposes a facial paralysis detection method based on visual perception and audio information, which improves the problem that the detection and diagnosis of facial nerve paralysis diseases cannot be made quickly and simply. The method includes the following steps: collecting RGB and Depth images; obtaining the RGB image I1 and the depth map I2 of the face region; obtaining the 2D key points and 3D key points of the face; guiding the person to be tested to read text; specifying the feature values reflecting the mouth movement to obtain the feature sequence of the mouth movement; collecting audio information and performing frame segmentation and windowing operations, extracting the Mel cepstral coefficients of each frame, and correcting the mouth movement sequence; performing feature extraction on the corrected mouth movement sequence Cnew to obtain the motion similarity index S; analyzing the frequency distribution of the audio to obtain the speech clarity D; comprehensively using S and D to obtain the facial paralysis detection result.
[0050] Another related technology proposes a multi-feature fusion-based automatic facial paralysis evaluation method, including the following steps: obtaining facial image samples of facial paralysis patients. The patient faces forward in front of the photographing detection module, and the facial photographing module performs three-dimensional correction of the facial image according to the patient's facial features, correcting the patient's facial image from the three dimensions of raw, roll, and pitch, and performing focal length correction when the patient's head moves offset. After obtaining the facial image samples of the patient in this application, the data after the multi-region feature fusion of the patient's face effectively increases the reliability and relevance of the data during comparison. The comparison of multiple facial region features and the multi-feature fusion comparison of the side seat surface partial region features and regional perception features greatly enhance the change accuracy obtained by the comparison under the comparison of the computer deep learning system. While improving the comparison efficiency, compared with the prior art, it can effectively obtain the accurate evaluation of the patient's facial paralysis grade, which is convenient for subsequent treatment based on this evaluation.
[0051] Another related technology proposes a facial paralysis diagnosis and rating method and system based on deep learning. The facial paralysis diagnosis and rating method includes the steps of establishing a motion model of normal people under specific facial movements; obtaining the motion data of the user under actual facial movements; comparing the motion data of the user with the motion model of normal people to obtain a comparison result; and then analyzing the comparison result to generate a conclusion on the degree of facial paralysis of the user. The facial paralysis diagnosis and rating system based on deep learning is used to implement the above method, including a construction module, an image acquisition module, a diagnosis module, an evaluation module, and a display module. This solution solves the problems in the prior art of lacking an intelligent facial paralysis recognition solution that supports multiple platforms and multiple models, and lacking a unified high-precision and high-accuracy facial paralysis diagnosis and rating system, and reduces the manpower and material costs of patients going out to seek medical treatment frequently.
[0052] The main problems existing in the existing facial paralysis diagnosis methods based on computer algorithms include: 1) Using the 2D image sequence of the human face can only reflect the horizontal and vertical movements of the face, and cannot accurately reflect the movement of the facial skin surface in three-dimensional space, so there will be a large error. 2) It is inaccurate to represent the facial paralysis of the whole face muscles by analyzing the movement directions of key points. Because these key points mainly represent the contour, eyebrows, eyes, nose, mouth and other parts of the face, while there are no key points in the cheek area where facial paralysis patients often get sick. Therefore, diagnosing facial paralysis by analyzing these key points will also produce a large error. 3) Using the 3D model sequence of the human face and detecting facial paralysis by using the asymmetry between the original face and the mirror image face of facial paralysis patients will also produce a large error. Because the faces of normal people and their expressions are not symmetric, using this characteristic to detect facial paralysis will produce a large error in the detection results of some people. 4) At present, in the part of obtaining the 3D model in this kind of method for detecting facial paralysis based on the 3D model, there are situations of low precision and incomplete facial reconstruction. Therefore, using such a 3D model for subsequent operations will also produce a large error.
[0053] To solve the above technical problems, an embodiment of the present application provides a facial paralysis detection method and device based on dynamic structured light technology. The detection method includes the following steps: First, perform real-time data acquisition on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model. Then, perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face; the 3D model of the patient's face is a topologically symmetric model. Finally, calculate the motion vector of each point in the 3D model of the patient's face, and quantify facial symmetry based on the symmetry operator. The present application provides a facial paralysis detection method based on dynamic structured light technology, aiming to complete the reconstruction of the human face with a method of high precision and good imaging integrity; use the method of non-rigid registration to complete the transformation from the standard human face model to the actual human face model; and count the motion vector of each point in the transformed topologically symmetric model, and quantify facial symmetry through this information. Subsequently, provide an objective visualization tool for doctors to diagnose facial paralysis, improving the efficiency and accuracy of patient visits and follow-up consultations.
[0054] The following will elaborate on each embodiment of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0055] See Figure 1 , an embodiment of the present application provides a facial paralysis detection method based on dynamic structured light technology, including the following steps:
[0056] Step S101: Perform real-time data acquisition on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model.
[0057] Step S102: Perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face; the 3D model of the patient's face is a topologically symmetric model.
[0058] Step S103: Calculate the motion vector of each point in the 3D model of the patient's face, and quantify facial symmetry based on the symmetry operator.
[0059] The embodiment of the present application provides a method for detecting facial paralysis based on dynamic structured light technology, and proposes to use a method with high precision and good imaging integrity to complete the reconstruction of the human face; use the method of non-rigid registration to complete the transformation from the standard face model to the actual face model; and count the motion vectors of each point in the transformed topological symmetry model, and quantify the facial symmetry through this information. Subsequently, an objective visualization tool is provided for doctors to diagnose facial paralysis, improving the efficiency and accuracy of patient visits and follow-up consultations.
[0060] In some embodiments, real-time data collection of the patient's face is performed to obtain a high-precision and complete facial point cloud, including: using three sets of structured light systems to simultaneously reconstruct the human face from three directions to obtain a high-precision and complete facial point cloud.
[0061] In some embodiments, each set of structured light systems includes a projector and a camera, and the projector and the camera are connected by a trigger wire; the projectors of the three sets of structured light systems are connected by an external trigger; after the external trigger sends a signal, the three projectors are simultaneously triggered to project a set of the above-mentioned codes; after each projector projects a code, the camera is triggered to take a coded image to ensure the synchronous control of the three systems.
[0062] In some embodiments, real-time data collection of the patient's face includes: the patient starts to make a specified symmetric expression from the expressionless state according to the instruction, and during this process, the point cloud data of all frames is collected; the specified symmetric expressions include frowning, closing eyes to the maximum extent, and grinning to the maximum extent.
[0063] The hardware device of the present application is developed based on the principle of coded structured light three-dimensional scanning, and is mainly used to obtain a complete sequence of human face 3D models in real time. To ensure the integrity of imaging, the present application uses three sets of structured light devices to reconstruct the human face at different angles. Since the real-time dynamic structured light system projects codes onto the object to be measured quickly, considering that the rapidly flickering visible light will cause discomfort to the human eye, the light source of this system uses infrared light that is invisible to the human eye. In addition, to avoid interference between the structured light projected by different devices on the human face, three infrared projection light sources with different wavelength bands are used; real-time three-dimensional point cloud reconstruction is realized based on the GPU module built into the imaging device, and the facial point clouds from three perspectives are stitched into a complete facial point cloud model through the pre-calibrated position parameters. The present application intends to let the patient make a specified expression from the expressionless state, and collect the 3D models during this process.
[0064] First, these face models are converted into Mesh models using surface reconstruction algorithms. Then, 2D face keypoint detection algorithms are used to complete facial keypoint detection, and the correspondence between 2D images and 3D models is used to automatically complete 3D keypoint annotation. Based on the 3D keypoints, parts that do not belong to the facial skin (such as eyes and mouths) are removed. Secondly, a standard face model is used as the initial model, and the captured expressionless face (the first frame of the 3D model sequence) is used as the target model. Non-rigid registration is used to obtain an initial face model with the same topological structure as the standard model and the same shape as the target face. Then, in this sequence, the model obtained in the previous frame is used as the original model, and the model captured in this frame is used as the target point cloud, and non-rigid registration of this frame is completed in combination with the 3D keypoints. In this way, the present application can complete the topological unification of all captured face model sequences. Finally, the vectors of the movement of each point from the first-frame topologically symmetric model to subsequent frames are calculated, and the facial symmetry is analyzed according to the symmetry operator proposed in the present application.
[0065] The facial paralysis detection method based on dynamic structured light technology provided by the present application includes three core parts:
[0066] 1) Real-time high-precision complete 3D face imaging.
[0067] 2) Using non-rigid registration to complete the topological unification of the captured 3D model sequence of the patient's face.
[0068] 3) Analyzing facial symmetry using a facial symmetry operator based on topological structure symmetry.
[0069] Specifically, as Figure 2 shown. In the process of the first part (real-time high-precision complete 3D face imaging), the present application performs real-time data acquisition on the patient's face through the established three-infrared structured light reconstruction system to obtain high-precision and complete facial point clouds, and further obtains its mesh model. Then, non-rigid registration is used to generate a 3D model of the patient's face (unified topological structure) with the same topological structure as the standard face. Finally, the movement vectors of the patient's face are calculated point by point, and the facial symmetry is quantified through the symmetry operator proposed in the present application.
[0070] In current 3D imaging methods, structured light technology has been widely used in many fields such as industrial inspection, virtual reality, and medicine due to its simple hardware structure, large measurement range, large field of view, and dense point cloud. Compared with traditional binocular vision, structured light technology replaces one camera with a projector. The projector projects one or a set of pre-set light stripes onto the surface of the object to be measured. By capturing the distorted stripe pattern on the object surface with a camera from another angle and combining the calibration parameters of the projector and the camera, the surface of the object to be measured can be reconstructed. Since a single structured light system may produce blind spots (the part of the projector encoding that the camera cannot capture), in order to simultaneously acquire a complete face model, this application uses three structured light systems to reconstruct the face from three directions simultaneously. Considering that high-speed stroboscopic visible light can cause discomfort to the human eye, the projector light source of this system uses infrared light. In addition, to prevent the light stripes projected by different systems from interfering with each other, this application installs narrow-band filters on the cameras that are adapted to the projection light sources (infrared light with wavelengths of 730nm, 850nm, and 940nm).
[0071] To achieve high-frame-rate projection of light stripes, this application uses a DLP projector of TI-4500, whose projection frame rate can reach over 2800Hz, ensuring the real-time performance of 3D scanning. The projection pattern adopts a spatial encoding scheme of Gray code plus line shift method. An external trigger is used to connect the projectors of the three structured light systems, and trigger wires are used to connect the projectors and the cameras. After the trigger issues a signal, the three projectors are simultaneously triggered to project a set of the above-mentioned codes; after each projector projects a code, the camera is triggered to capture a coded image, ensuring the synchronous control of the three systems. The structured light acquisition device of this patent application is as Figure 3 shown. This system can perform real-time (30Hz) 3D scanning of the patient's face, obtaining a high-resolution (three cameras with a resolution of 720*540) and high-precision (precision less than 0.1mm) facial point cloud sequence.
[0072] In addition, to ensure the robustness of the test, the patient needs to perform specified symmetric expressions (such as frowning, closing eyes to the maximum extent, grinning to the maximum extent, etc.) starting from a non-expression state according to the instructions. During this process, the system acquires the point cloud data of all frames.
[0073] In the process of the second part (unifying the topological structures of the standard face and the patient's facial model), since the structured light system can generate a dense point cloud (the camera resolution used in this system is 720*540, and a single camera can generate up to 350,000 three-dimensional points in one imaging), and the topological structures of each frame of the point cloud are not unified, the patient's facial movement cannot be directly calculated. Therefore, it is necessary to unify the topological structures of these models.
[0074] In some embodiments, non-rigid registration is performed on the mesh model to generate a 3D patient facial model with the same topological structure as the standard face, including: using a non-rigid registration algorithm to deform the template surface of the mesh model into the same shape as the target surface; finding the optimal deformation position by optimizing the loss function, and realizing the entire non-rigid registration through iteration to unify the topological structures of the standard face and the patient facial model.
[0075] The non-rigid registration algorithm can deform the registration template into the same shape as the target surface. Assume that S is the template surface, T is the target surface, n is the number of points on the template surface, and S(X) is the affine transformation relationship from the entire surface S to the surface T; for each vertex v i After undergoing a local affine transformation X i After continuous iteration, it will move to a place u closer to the surface T i , thus obtaining an optimal deformation from the template surface S to the target surface T.
[0076] To achieve the above purpose and find the optimal deformation position, for each vertex v in the template surface i Find a most suitable local affine transformation X i , the non-rigid registration defines an optimized loss function, which includes three loss functions, namely: data loss function, smoothing loss function, and landmark loss function.
[0077] Among them, the optimized loss function is as follows:
[0078] E = E d (X) + αE s (X) + βE l (X) (1)
[0079] Among them, E d (X) is the data loss function, E S (X) is the smoothing loss function, and E l (X) is the landmark loss function;
[0080] The data loss function is as follows:
[0081]
[0082] Among them, v i is the i-th vertex in the template surface, with a total of n; X i is the displacement transformation matrix of each vertex, and ω i is the weight; the data loss function is used to find the point on the target mesh closest to the point v i , calculate the distance between the two, and then make the points on the template surface as close as possible to the target surface;
[0083] The smooth loss function is used to measure whether the template surface after transformation is smooth; the smooth loss function is as follows:
[0084]
[0085] The landmark loss function is used to adjust according to the preset feature points, so that the feature points on the template surface fit the feature points on the target surface; the landmark loss function is as follows:
[0086]
[0087] Wherein, G is a parameter used to measure rotation and translation, L is the marked feature points, and l is the normal direction vector of the facial symmetry plane.
[0088] After obtaining the optimized loss function of the algorithm, finally through iteration, the entire non-rigid registration can be achieved.
[0089] In some embodiments, by optimizing the loss function, the best deformation position is found, and the entire non-rigid registration is achieved through iteration, unifying the topological structures of the standard face and the patient's facial model, including: using non-rigid registration to deform the standard face I 0 as the source model, and the patient's expressionless face model T 0 as the target model, deforming the standard face to complete the fitting of the expressionless face model S 0 ; for subsequent frames of the 3D model sequence, taking the model collected in the current frame as the target model, and combining with 3D key points, deforming the model deformed in the previous frame as the source model, and the process is as Figure 4 shown; wherein, the patient's expressionless face model T 0 is the first frame in the 3D model sequence.
[0090] In this way, the present application can unify the topological structures of all 3D models. According to the above registration method, the registration effect obtained is as Figure 5 shown. Figure 5 Among them, (a), (b), and (c) are the expressionless face, the expression face, and the result after registration respectively.
[0091] It should be noted that the standard face model used in the present application is a left-right symmetric model, that is, it is divided into the left face and the right face; assuming that the point cloud set of this model is I 0 , the number of points on the left face and the right face is N each, and the number of points on the midline is M, and its topological structure is as Figure 6 shown, and the points with the same subscript on the left and right faces are mirror image points in the 3D model. Therefore, for the face model after unifying the topological structure, the points on the left face and the right face have a corresponding relationship (the information of its corresponding point can be quickly found for any point).
[0092] In some embodiments, calculating the motion vectors of each point in the 3D model of the patient's face includes: obtaining the set of motion directions and the set of distances of each point relative to the first topological symmetric model; mirroring the expressionless topological symmetric model and calculating the distances between all symmetric points between the two models, and screening out the parts with high facial symmetry by setting a threshold; fitting a facial symmetry plane through the symmetric points of the parts with high facial symmetry; using the HouseHolder operator to mirror the moving direction of the left face part in the topological symmetric model using the direction of the symmetric plane normal; after mirroring the left face points in the topological symmetric model, obtaining the moving vectors of each point in the model relative to the expressionless topological symmetric model, and transforming the facial symmetry problem into a vector similarity problem of corresponding points.
[0093] In the process of the third part (analyzing the facial paralysis area), it is necessary to analyze the specific facial paralysis area. It is not enough to simply use the symmetry of a single motion distance to judge the facial paralysis area, and the motion direction also needs to be combined for judgment. After the above processing, for each 3D model T in the human face expression sequence i , this application can use a topological symmetric model S 0 with the same topology as the topological structure and the standard human face model I i . For each point of the i-th topological symmetric model in the 3D model sequence, the set of motion directions D 0 and the set of distances A i of each point relative to the first topological symmetric S i model can be expressed as:
[0094]
[0095] In addition, considering that the human face is inherently asymmetric, this application mirrors the expressionless topological symmetric model S 0 and calculates the distances between all symmetric points between the two models, and screens out the parts with high facial symmetry by setting a threshold. A facial symmetry plane is fitted through these symmetric points, and this plane can be expressed as:
[0096]
[0097] The normal direction vector l of this plane can be expressed as:
[0098]
[0099] Through the HouseHolder operator, this application can mirror the moving direction of the left face part in the topological symmetric model using the direction of the symmetric plane normal l:
[0100]
[0101] After mirroring the left - face points in the topological symmetry model S i the movement vectors of each point in the model relative to the expressionless topological symmetry model are as follows:
[0102]
[0103] wherein, H i is the movement vector of each point relative to the expressionless topological symmetry model S 0 ; is the movement direction after mirroring the movement direction of the left - face part in the topological symmetry model using the symmetry - plane method; is the movement direction of the left - face part in the topological symmetry model; N is the number of points of the left - face and right - face in the point - cloud set of the standard face model, and M is the number of points on the mid - line.
[0104] In this way, the facial symmetry problem is transformed into a vector similarity problem of corresponding points. To evaluate the similarity of these vectors, constraint conditions are used to evaluate the vector similarity.
[0105] In some embodiments, after transforming the facial symmetry problem into a vector similarity problem of corresponding points, constraint conditions are used to evaluate the vector similarity, and a symmetry operator based on the topological - structure symmetry model is obtained;
[0106] The constraint conditions are:
[0107] 0 = f (α,α) < f (α,β) = f (β,α) < 1 (12)
[0108] The symmetry operator based on the topological - structure symmetry model is expressed as:
[0109]
[0110] wherein, ε is a parameter for adjusting the weights of angles and distances; λ and γ are parameters for adjusting the rising rate and the symmetry center respectively.
[0111] Specifically, this operator mainly uses the parameter ε to adjust the weights of angles and distances, and adopts a sigmoid function as an adjustment module to reduce the situation where the symmetry value becomes too large due to significant differences in distances and angles when the distance is small. In addition, the parameters λ and γ can be used to adjust the rising rate and the symmetry center of this adjustment component.
[0112] For each point in the topological symmetry model S i this application can use the above - mentioned operator to calculate its movement relative to the topological symmetry model S in the expressionless state0 Calculate the facial movement symmetry, and the results of its quantitative analysis can be expressed as:
[0113]
[0114] This application collected 3D model sequences of normal people with maximum eyebrow raising (symmetrical expression, expression A), pursing the lips (symmetrical expression, expression B), and left grinning (asymmetrical expression, expression C) to test the operators of this application. Among them, this application adopted ε = 0.3, λ = 1.0, and γ = 5, and the calculation results are as Figure 7 shown. From the results, obvious differences appear between symmetrical and asymmetrical expressions on the heat map. It should be noted that the human face and wrinkles themselves have asymmetry. By analyzing the results of expressions A and B, it can be seen that our operator can also well identify the asymmetrical parts with subtle changes (such as the forehead and cheek parts).
[0115] In addition, the embodiment of this application also provides a facial paralysis detection device based on dynamic structured light technology, which uses the facial paralysis detection method based on dynamic structured light technology as described in the above embodiment for detection, including: a human face 3D imaging module, a topology unification module, and a facial symmetry analysis module connected in sequence; wherein, the human face 3D imaging module is used to perform real-time data collection on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model; the topology unification module is used to perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face; the 3D model of the patient's face is a topologically symmetric model; the facial symmetry analysis module is used to calculate the motion vector of each point in the 3D model of the patient's face and quantify the facial symmetry based on the symmetry operator.
[0116] This application has been certified by the experimental system, and the effect is very ideal, consistent with the design expectations. The facial paralysis detection method and device based on dynamic structured light technology provided by this application optimize the process of non-rigid registration based on continuous 3D facial data frames. First, manually mark the 3D key points corresponding to the first frame in the standard human face, and combine the corresponding 3D key points of the acquisition frame to deform the standard human face into the acquisition human face using non-rigid registration. Then, for the non-rigid registration of subsequent frames, the model obtained from the previous frame can be used as the source model, the acquisition model as the target model, and combined with the corresponding 3D key points of the acquisition frame for automatic non-rigid registration.
[0117] This application proposes to use a standard human face that is topologically symmetric as the initial human face, and use the above method to represent all the collected 3D human face model sequences with this symmetric topological result.
[0118] The present application proposes a facial symmetry operator based on a symmetric topological structure model, which can quantify the symmetry of the topological structure symmetric model. Moreover, the operator can also adjust the calculated value by modifying parameters according to actual application requirements.
[0119] The present application uses time-coded infrared structured light to reconstruct the burned area. By reconstructing the face from three angles, a high-precision, high-frame-rate, and complete 3D model of the patient's face can be obtained. The present application uses a sequence of real 3D models of the test subject's face and completes the topological unification of the collected data using the method of non-rigid registration, which can better perceive the details of the full-face movement. The invention uses the moving distance on a topologically symmetric human face to provide a basis for facial paralysis symmetry and can intuitively calculate the movement of the full face.
[0120] With the above technical solutions, the embodiments of the present application provide a method and device for detecting facial paralysis based on dynamic structured light technology. The detection method includes the following steps: First, perform real-time data acquisition on the patient's face to obtain a high-precision and complete facial point cloud; and based on the facial point cloud, obtain a mesh model; then, perform non-rigid registration on the mesh model to generate a 3D model of the patient's face with the same topological structure as the standard human face, and the 3D model of the patient's face is a topological symmetric model; finally, calculate the motion vector of each point in the 3D model of the patient's face and quantify the facial symmetry based on the symmetry operator.
[0121] The method and device for detecting facial paralysis based on dynamic structured light technology provided by the present application. The hardware device of the present application is developed based on the principle of coded structured light three-dimensional scanning and is mainly used to obtain a complete sequence of 3D models of the human face in real time. To ensure the integrity of imaging, the present application uses three structured light devices to reconstruct the human face at different angles. Since the real-time dynamic structured light system projects codes onto the object to be measured quickly, considering that the fast-flashing visible light will cause discomfort to the human eye, the light source of this system uses infrared light that is invisible to the human eye. In addition, to avoid the interference of the structured light projected between devices on the human face, three infrared projection light sources with different wavelength bands are used; real-time three-dimensional point cloud reconstruction is realized based on the GPU module built in the imaging device, and the facial point clouds from three perspectives are stitched into a complete facial point cloud model through pre-calibrated position parameters. The present application intends to let the patient make a specified expression from a non-expression state, and collect 3D models during this process.
[0122] Compared with the method of deforming a standard face to judge facial shape and parameters using deep learning, the present application uses the actually captured sequence of facial 3D models, which can better perceive the details of facial movements. The present application uses dynamic structured light technology to reconstruct the face from three angles, and can obtain a high-precision, high-frame-rate, and complete 3D model of the patient's face. Compared with 2D methods, low-precision 3D model methods, and incomplete 3D model methods, the present application can obtain high-precision data that more conforms to the actual movement law. The present application takes into account the asymmetry of the human face. The facial symmetry plane is obtained by fitting facial points with stronger symmetry, rather than simply directly analyzing the symmetry between the facial model and its mirror model. This method of the present application can perform point-by-point motion symmetry analysis on all regions of the face, and is also applicable to patients and doctors to evaluate the prognosis effect.
[0123] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present application, and in actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make respective changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A facial nerve paralysis detection method based on dynamic structured light technology, characterized in that: The following steps are involved: Performing real-time data collection on the patient's face to obtain a high-precision, complete facial point cloud; and obtaining a mesh model based on the facial point cloud; Performing non-rigid registration on the mesh model to generate a 3D facial model of the patient having the same topological structure as a standard human face; the 3D facial model of the patient is a topologically symmetric model; The motion vector of each point in the 3D facial model of the patient is calculated, and the facial symmetry is quantified based on a symmetry operator.
2. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 1 is characterized in that: Real-time data collection of the patient's face is performed to obtain a high-precision, complete facial point cloud, including: Using three sets of structured light systems, the face is reconstructed from three directions simultaneously to obtain a high-precision and complete facial point cloud.
3. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 2 is characterized in that: Each structured light system includes a projector and a camera, which are connected by a trigger line; the projectors of the three structured light systems are connected by an external trigger; After the external trigger sends a signal, it triggers the three projectors to project a set of the above codes at the same time; after each projector projects a code, it triggers the camera to take a coded image to ensure the synchronous control of the three systems.
4. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 1, characterized in that: Real-time data collection of the patient's face, including: The patient starts from a neutral state and makes a designated symmetrical expression according to the instruction. During this process, the point cloud data of all frames are collected; The specified symmetrical expressions included frown, maximal eye closure, and maximal grin.
5. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 1, characterized in that: The mesh model is non-rigidly registered to generate a patient's facial 3D model having the same topological structure as a standard human face, including: Using a non-rigid registration algorithm, the template surface of the mesh model is deformed into the same shape as the target surface; By optimizing the loss function, the best deformation position is found, and the entire non-rigid registration is achieved through iteration to unify the topological structures of the standard face and patient face models.
6. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 5, characterized in that: Using a non-rigid registration algorithm, the template surface of the mesh model is deformed into the same shape as the target surface, including: Assume that S is the template surface, T is the target surface, n is the number of points on the template surface, and S(X) is the affine transformation relationship from the entire surface S to the surface T; for each vertex v i After a local radiation change X i After continuous iteration, it will move to a place u that is closer to the surface T. i , thus obtaining an optimal deformation from the template surface S to the target surface T; The optimization loss function is as follows: E=E d (X)+αE s (X)+βE l (X) (1) Among them, E d (X) is the data loss function, E S (X) is the smoothing loss function, E l (X) is the landmark loss function; The data loss function is as follows: Among them, v i is the i-th vertex in the template surface, there are n of them; X i is the displacement transformation matrix of each vertex, ω i is the weight; the data loss function is used to find the target grid to point v i The closest point is calculated, and the distance between the two is calculated, and then the points in the template surface are made to fit as close to the target surface as possible; The smooth loss function is used to measure whether the template surface is smooth after transformation; the smooth loss function is as follows: The landmark point loss function is used to adjust according to the preset feature points so that the feature points in the template surface and the target surface fit together; the landmark point loss function is as follows: Among them, G is a parameter used to measure rotation and translation, L is the marked feature point, and l is the normal vector of the facial symmetry plane.
7. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 5, characterized in that: By optimizing the loss function, the optimal deformation position is found, and the entire non-rigid registration is implemented through iteration to unify the topological structure of the standard face and patient face model, including: Non-rigid registration is used to take the standard face I0 as the source model and the patient's expressionless face model T0 as the target model, and the standard face is deformed to complete the fitting S0 of the expressionless face model; for subsequent frames of the 3D model sequence, the model collected by the current frame is taken as the target model, and combined with the 3D key points, the deformed model of the previous frame is used as the source model for deformation; among them, the patient's expressionless face model T0 is the first frame in the 3D model sequence.
8. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 1, characterized in that: Calculating the motion vector of each point in the 3D model of the patient's face, including: Obtain the movement direction set and distance set of each point relative to the first topological symmetric model; The expressionless topological symmetric model is mirrored, and the distance between all symmetric points between the two models is calculated. The parts with high facial symmetry are screened out by setting a threshold. The facial symmetry plane is fitted through the symmetric points of the parts with high facial symmetry. Through the HouseHolder operator, the movement direction of the left face part in the topological symmetric model is mirrored using the normal direction of the symmetry plane; After mirroring the left face points in the topological symmetry model, the movement vector of each point in the model relative to the expressionless topological symmetry model is obtained, and the facial symmetry problem is transformed into a vector similarity problem of corresponding points; The movement vector of each point in the model relative to the expressionless topological symmetric model is as follows: Among them, H i For each point, the topological symmetric model S is 0 The movement vector of To move the left face part of the topological symmetry model The moving direction of the mirror operation using the normal direction of the symmetry plane; is the moving direction of the left face part in the topological symmetric model; N is the number of points on the left and right faces in the point cloud set of the standard face model, and M is the number of points on the midline.
9. The facial nerve paralysis detection method based on dynamic structured light technology according to claim 8, characterized in that: After converting the facial symmetry problem into the vector similarity problem of corresponding points, the constraint conditions are used to evaluate the vector similarity, and the symmetry operator based on the topological structure symmetry model is obtained; The constraints are: 0=F (α,α) <f (α,β) =f (β,α) <1 (6) The symmetry operator based on the topological structure symmetry model is expressed as: Among them, ε is the parameter for adjusting the weights of angle and distance; λ and γ are the parameters for adjusting the rise rate and the symmetry center respectively.
10. A facial nerve paralysis detection device based on dynamic structured light technology, using the facial nerve paralysis detection method based on dynamic structured light technology as described in any one of claims 1 to 9 for detection, characterized in that: include: The face 3D imaging module, the topology unification module and the facial symmetry analysis module are connected in sequence; wherein, The face 3D imaging module is used to collect data on the patient's face in real time to obtain a high-precision, complete facial point cloud; and to obtain a mesh model based on the facial point cloud; The topology unification module is used to perform non-rigid registration on the mesh model to generate a patient's facial 3D model having the same topological structure as a standard human face, wherein the patient's facial 3D model is a topologically symmetric model; The facial symmetry analysis module is used to calculate the motion vector of each point in the patient's facial 3D model and quantify the facial symmetry based on a symmetry operator.
Citation Information
Cited By
An objective analysis method for facial observation based on multimodal data and a multispectral stereoscopic imaging observation instrument.
CN122827616A