Computer vision-based facial nerve disease rehabilitation condition detection method
By using binocular synchronous acquisition and epipolar correction, combined with posture correction and affected-side semantic unification, a regional rehabilitation index table and a movement quality table are generated. This solves the problem of inconsistent three-dimensional deformation indicators in the rehabilitation assessment of facial nerve diseases and achieves stability in longitudinal rehabilitation trend assessment and efficacy assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XUZHOU MEDICAL UNIVERSITY
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-09
AI Technical Summary
Existing methods for assessing facial nerve disorders are difficult to obtain consistent and comparable three-dimensional deformation indicators for different regions, and cannot assess longitudinal rehabilitation trends. They are also susceptible to head movements, occlusion, and scale drift, and lack a stable mechanism for semantic unification and functional region aggregation on the affected side.
By synchronously acquiring facial motion sequences with binoculars and performing epipolar correction and temporal labeling, segmented binocular motion sequences are generated. Face localization, pose correction, and scale normalization are performed to generate key point trajectories and functional partition semantic masks. Anisotropic Gaussian sets with partition attributes are constructed, and visibility-perceived Gaussian sputtering rendering and binocular reprojection consistency updates are performed to generate a four-dimensional Gaussian representation of the face. Combined with healthy side mirror comparison and healthy sample library template comparison, partitioned rehabilitation index tables and motion quality tables are generated, and follow-up alignment and trend calculation are performed.
It achieves a highly consistent imaging foundation, reduces interference from head movement and visual angle differences, precisely quantifies the differences between the affected and healthy sides, improves the sensitivity and interpretability of mild functional recovery, supports efficacy assessment and treatment plan adjustment, and enhances the robustness and reproducibility of quantitative assessment.
Smart Images

Figure CN122177452A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method for detecting the rehabilitation status of facial nerve diseases based on computer vision. Background Technology
[0002] Facial nerve disease rehabilitation assessment technology has gradually evolved from clinical subjective grading to objective quantification based on imaging and sensing. Existing methods typically collect standard action sequences such as resting, closing eyes, smiling, and raising eyebrows under terminal prompts, combine face detection and key point tracking to normalize posture and scale, and can use binocular / depth information to achieve three-dimensional reconstruction, thereby extracting indicators such as symmetry, displacement, and texture for rehabilitation interpretation.
[0003] However, assessments based on two-dimensional trajectories or single depths are susceptible to head movements, occlusion, and scale drift, making it difficult to obtain consistent three-dimensional deformation representations across follow-ups. At the same time, the lack of stable semantic unification and functional zoning aggregation mechanisms for the affected side results in incomparability and weak interpretability of indicators at different action segments and time points, and makes it impossible to form a zoning quantitative table and longitudinal trend results that can be compared with the healthy template. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a computer vision-based method for detecting the rehabilitation status of facial nerve diseases, which solves the problem of difficulty in obtaining consistent and comparable three-dimensional facial deformation indicators and assessing longitudinal rehabilitation trends.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] This invention provides a computer vision-based method for detecting the rehabilitation status of facial nerve diseases, comprising:
[0008] Under the guidance of the terminal, the binocular synchronous acquisition of facial motion sequences was completed, and epipolar correction and temporal marking were performed to obtain motion segmentation binocular sequences;
[0009] Perform face localization, pose correction and scale normalization on the action segmented binocular sequence, and complete semantic unification of the affected side to generate key point trajectories and functional partition semantic masks.
[0010] Based on the key point trajectory and functional partition semantic mask, the disparity initial value of the action segment binocular sequence is initialized in 3D and an anisotropic Gaussian set with partition attributes is constructed. The anisotropic Gaussian set is then subjected to visibility-perceptual Gaussian sputtering rendering and binocular reprojection consistency update to obtain a four-dimensional face Gaussian representation, a three-dimensional deformation vector field and a midline plane.
[0011] The healthy side mirror comparison is performed using the midline plane, and the three-dimensional deformation vector field is summarized according to the partition attributes to generate dynamic deformation features. At the same time, texture and gradient features are extracted in the functional partition semantic mask and compared with the template of the healthy sample library to obtain the partition rehabilitation index table and the movement quality table.
[0012] Follow-up alignment and trend calculation were performed on the regional rehabilitation index table and the movement quality table, and reference surface uniformity was performed by combining the four-dimensional Gaussian representation of the face to obtain the longitudinal rehabilitation trend results.
[0013] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, wherein obtaining the action segmentation binocular sequence specifically includes:
[0014] The terminal displays action commands such as natural rest, slightly closed eyes, showing teeth and smiling, and raising eyebrows, forming a sequence of action segments.
[0015] The binocular camera synchronously acquires left and right eye image frames according to the action segment sequence table and generates frame pairing identifiers.
[0016] Epipolar correction is performed on the left and right eye image frames. The epipolar constraints of the pixels are calculated using the camera's intrinsic and extrinsic parameters, and the action start and end markers are written according to the action segment sequence table.
[0017] The action start marker, action end marker, and frame pairing identifier are associated to form an action segmentation binocular sequence.
[0018] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the generation of key point trajectories and functional partition semantic masks specifically includes:
[0019] Face localization and key point localization are performed on the left and right eye image frames of the action segmented binocular sequence to obtain the key point set and generate the key point trajectory.
[0020] Extract the center points of both eyes based on the key point trajectory, determine the direction of the line connecting the centers of both eyes, and drive attitude correction to obtain attitude-corrected image frames.
[0021] The cropping box is determined based on the set of key points and the scale of the pose-corrected image frame is normalized to obtain the scale-normalized image frame.
[0022] During the natural resting phase, determine and mark the direction of the affected side based on the direction of mouth drooping.
[0023] Based on the affected side markers, semantic unification of the affected side is performed on the scale-normalized image frames to obtain semantically unified image frames;
[0024] Based on the set of key points, the semantically unified image frame is divided into eye region, mouth region and eyebrow region, and corresponding functional partition semantic mask is generated.
[0025] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, wherein: the construction of an anisotropic Gaussian set with partitioning attributes specifically includes:
[0026] Feature extraction is performed on the left and right eye image frames of the action segmented binocular sequence to obtain binocular features;
[0027] Under the constraint of functional partition semantic mask, the initial disparity value is estimated on the binocular features to obtain the initial disparity map;
[0028] Under the constraint of key point trajectory, the effective disparity region of the initial disparity map is filtered to obtain the facial disparity region;
[0029] Depth inversion is performed on the facial parallax region to obtain a sparse 3D point set;
[0030] Associating each 3D point in the sparse 3D point set with a functional partition semantic mask yields a 3D point set with partition attributes.
[0031] Anisotropic parameterization is performed on the 3D point set with partitioning properties to obtain anisotropic 3D Gaussian elements, while retaining the partitioning properties to form an anisotropic Gaussian set with partitioning properties.
[0032] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the step of performing visibility-perceived Gaussian sputtering rendering and binocular reprojection consistency update on the anisotropic Gaussian set specifically includes:
[0033] The visibility of each Gaussian element in the anisotropic Gaussian set is calculated according to the binocular camera view, forming a visibility set;
[0034] Gaussian sputtering rendering is driven by visibility sets, and rendered image frames are generated in both left and right viewpoints.
[0035] Perform a binocular reprojection consistency check on the rendered image frames and the left and right eye image frames of the action segment binocular sequence to form a consistency residual.
[0036] Based on the consistent residual-driven anisotropic Gaussian set, update the position and shape parameters of the Gaussian elements while keeping the partitioning properties unchanged;
[0037] The updated anisotropic Gaussian set is used to establish inter-frame correlations along the action segment time sequence to form a four-dimensional Gaussian representation of the face.
[0038] Based on inter-frame correlation, the displacement of Gaussian elements corresponding to adjacent frames in the four-dimensional face Gaussian representation is calculated to generate a three-dimensional deformation vector field.
[0039] Based on the data from the Gaussian representation of a four-dimensional face during natural resting motion, the facial midline structure is fitted to generate a midline plane.
[0040] In a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the generation of dynamic deformation features specifically includes:
[0041] The mirror mapping is defined by the midline plane, and the healthy side is mirrored on the four-dimensional Gaussian representation of the face to obtain the mirror Gaussian representation;
[0042] The mirror Gaussian representation and the four-dimensional face Gaussian representation are matched for partition attribute consistency to form a mirror comparison relationship;
[0043] Based on the mirror comparison relationship, the three-dimensional deformation vector field is partitioned and its attributes are summarized. The difference, magnitude and time difference of the deformation vector of the corresponding Gaussian element at the corresponding time are calculated to generate dynamic deformation features.
[0044] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the healthy sample library template is formed by extracting texture and gradient features from the action segment binocular sequence that is consistent with the facial action sequence collected by healthy people under the guidance of the terminal, and summarizing them according to the action segment to form a partitioned texture template, a partitioned gradient template and a dynamic deformation template.
[0045] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the step of extracting texture and gradient features within a functional partition semantic mask and comparing them with a template from a healthy sample database specifically includes:
[0046] Under the constraint of functional partitioning semantic mask, semantically unified image frames are cropped to obtain partitioned image blocks;
[0047] Texture encoding and gradient calculation are performed on the partitioned image blocks to form texture histogram features and gradient histogram features;
[0048] Similarity calculations are performed on the texture histogram features and the partitioned texture templates of the healthy sample library to generate texture alignment results;
[0049] Similarity calculations are performed between the gradient histogram features and the partition gradient templates of the healthy sample library to generate gradient alignment results;
[0050] Similarity calculation is performed between dynamic deformation features and dynamic deformation templates in the healthy sample database, and deformation comparison results are generated.
[0051] The texture comparison results, gradient comparison results, and deformation comparison results are summarized according to the partition attributes to form a partition rehabilitation index table.
[0052] The rehabilitation index tables for each region are summarized in order of movement segment to form a movement quality table.
[0053] As a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the step of performing follow-up alignment and trend calculation on the partitioned rehabilitation index table and the movement quality table specifically includes:
[0054] A follow-up index is generated using patient identifiers and collection timestamps as indexes;
[0055] Link the regional rehabilitation index table and the movement quality table to the follow-up index to form a follow-up record;
[0056] Based on the semantic uniformity rules for the affected side, follow-up records are matched against the same partition to generate a partition trend sequence;
[0057] The trend series of the partitions is sorted by time and the changes are statistically analyzed to generate trend calculation results.
[0058] In a preferred embodiment of the computer vision-based facial nerve disease rehabilitation detection method of the present invention, the step of reference surface homogenization specifically includes:
[0059] The reference Gaussian surface is extracted from the natural resting action segment of the four-dimensional face Gaussian representation and registration is performed to form a reference-consistent follow-up record;
[0060] By combining consistent follow-up records with trend calculation results, longitudinal recovery trend results are generated.
[0061] The beneficial effects of this invention are as follows: By achieving a highly consistent imaging foundation through binocular synchronous acquisition and epipolar correction, and combining posture correction, scale normalization, and semantic unification of the affected side, the interference of head movement, perspective differences, and inconsistent left-right labeling on rehabilitation assessment is significantly reduced, making data from different movement segments and different follow-up time points comparable. The use of a three-dimensional representation with functional zoning attributes and a mirror comparison mechanism allows for precise quantification of differences between the affected and healthy sides at local levels such as the eye, mouth, and eyebrow regions, improving the sensitivity and interpretability of minor functional recovery and compensatory movements. Simultaneously, the introduction of a healthy sample database template for comparison integrates multimodal features such as deformation and texture gradients into a zonal rehabilitation index table and a movement quality table, thereby achieving a structured output of "local indicators—movement quality—overall trend." Through follow-up alignment and reference surface consistency, longitudinal trend results are formed, stably reflecting rehabilitation progress and supporting efficacy assessment and program adjustments. Overall, this improves the robustness, repeatability, and clinical readability of quantitative assessment, reducing reliance on subjective manual grading. Attached Figure Description
[0062] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 This is a flowchart of a computer vision-based method for detecting the rehabilitation status of facial nerve diseases.
[0064] Figure 2 Flowchart for determining the direction of the affected side.
[0065] Figure 3 Flowchart for visibility-aware Gaussian sputtering rendering and updating.
[0066] Figure 4 Flowchart for follow-up alignment and trend calculation. Detailed Implementation
[0067] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0068] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0069] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0070] Reference Figures 1-4 As one embodiment of the present invention, this embodiment provides a method for detecting the rehabilitation status of facial nerve diseases based on computer vision, including the following steps:
[0071] S1. Under the guidance of the terminal, complete the binocular synchronous acquisition of facial action sequences and perform epipolar correction and temporal marking to obtain action segmentation binocular sequences.
[0072] S1.1. On the guided interface of a mobile phone or tablet, four action commands are displayed in sequence: "Natural Rest", "Slightly Close Eyes", "Smile with Teeth", and "Raise Eyebrows". The display order is recorded as an action segment sequence table. According to the received action segment sequence table, the binocular camera synchronously triggers the left and right cameras to acquire images during the duration of each action command, thereby obtaining a series of time-aligned left and right image frames.
[0073] Each pair of left and right eye image frames acquired at the same time is assigned a unique frame pairing identifier, and the acquisition timestamp of each frame is recorded simultaneously. The intrinsic parameter matrix, distortion coefficients, rotation matrix, and translation vector describing the relative positional relationship between the left and right cameras are also recorded as camera calibration parameters. Throughout the acquisition process, the exposure time, aperture gain, and white balance parameters of the stereo cameras are locked to maintain consistent illumination within the sequence.
[0074] S1.2. Stereo epipolar correction is performed on each pair of left and right eye image frames using camera calibration parameters. Specifically, a correction mapping matrix for image plane reprojection is calculated based on the camera's intrinsic parameter matrix, distortion coefficients, rotation matrix, and translation vector. Then, the corresponding correction mapping is applied to the left and right eye image frames respectively. New corrected left and right eye image frames are generated through pixel resampling. The correction makes the epipolar lines of the two images horizontally aligned, meaning that any pair of corresponding pixels in the corrected image frame is located in the same row, thus constraining the two-dimensional search space of corresponding points to a one-dimensional horizontal line.
[0075] To further explain, the calculation of the correction mapping matrix is based on the principle of stereo correction in stereo vision. The core is to calculate a rotation matrix that rotates the left camera coordinate system to be coplanar with the right camera coordinate system, specifically:
[0076] ;
[0077] in, The correction rotation matrix for the left camera is the rotation matrix of the left camera coordinate system relative to the right camera coordinate system after the correction process is applied. This represents the correction rotation matrix, used to rotate the left camera coordinate system to a position coplanar with the right camera coordinate system, so that the optical axes of the two cameras are parallel; This represents the rotation matrix of the left camera, which is the original pose of the left camera relative to the world coordinate system.
[0078] S1.3. Based on the start and end times of each action instruction displayed on the guide interface in the action segment sequence table, and the recorded image frame acquisition timestamps, locate the start and end frames of each action in the image frame sequence after epipolar correction is completed, and write the action start mark indicating the start of the action and the action end mark indicating the end of the action on the frame pairing identifier of the corresponding frame, respectively.
[0079] All corrected left and right eye image frames, corresponding frame pairing identifiers, action start markers, action end markers, acquisition timestamps, and camera calibration parameters are structurally associated and encapsulated, and the output is a segmented stereo sequence of actions.
[0080] S2. Perform face localization, pose correction and scale normalization on the action segmented binocular sequence, and complete semantic unification on the affected side to generate key point trajectories and functional partition semantic masks.
[0081] S2.1. For each corrected left-eye image frame in the action segmentation binocular sequence, perform face detection, locate the face region in the image, and further locate facial key points within the face region to obtain a set of coordinates containing sixty-eight preset feature points (such as the corners of the eyes, the tip of the nose, the corners of the mouth, etc.), which is denoted as the key point set; arrange the key point sets of all image frames in the action segmentation binocular sequence in chronological order to generate a key point trajectory describing the change of position of each key point over time.
[0082] Based on the keypoint set, the coordinates of the left and right eye center points are extracted. The direction vector of the line connecting the center points of both eyes is calculated, and the angle by which the image needs to be rotated to horizontal is calculated based on the direction vector. The rotation angle can be calculated from the pixel coordinates of the center points of both eyes, using the following formula:
[0083] ;
[0084] in, and These represent the coordinates of the center points of the left and right eyes in the image coordinate system, respectively. The rotation angle; This represents the arctangent function.
[0085] An affine transformation is performed on the corrected left eye image frame based on the rotation angle to rotate the image so that the line connecting the centers of both eyes in the rotated image is horizontal, thus obtaining the pose-corrected image frame. Simultaneously, the same rotation transformation is applied to the coordinates of each point in the keypoint set, updating the keypoint set.
[0086] S2.2. Based on the updated set of key points, determine a rectangular region that can contain all facial key points as the clipping box. Specifically, calculate the coordinates of all key points in... shaft and The minimum and maximum values along the axis are used to define an outer rectangle. This rectangle is then expanded according to a preset aspect ratio to serve as the uniform clipping frame.
[0087] The pose correction image frame is cropped using a uniform cropping box, and the cropped image area is scaled to a preset fixed size to achieve scale normalization and obtain a scale-normalized image frame. At the same time, the coordinates of the key point set are cropped and scaled accordingly, and the key point set is updated again.
[0088] To further explain, the preset aspect ratio is a fixed ratio determined based on the input requirements of the standard face detection model, the compatibility of subsequent processing steps, and to reserve appropriate boundaries to ensure the integrity of the facial area. It is usually set to 1:1 (square) to ensure that the proportion of the cropped facial area is coordinated and to avoid geometric distortion caused by image stretching during subsequent feature extraction.
[0089] The fixed size is set to 256×256 pixels because this size is a commonly used resolution in computer vision tasks, which can strike a balance between preserving sufficient facial detail information and reducing subsequent computational complexity.
[0090] Within the action segmentation binocular sequence, a scale-normalized image frame and its corresponding keypoint set are selected from the action segment marked as natural resting. Based on the vertical coordinates of the left and right corner points (e.g., points 48 and 54) in the keypoint set... (Coordinates) to determine the drooping direction of the corner of the mouth. The side with the larger vertical coordinate value (i.e., the lower position in the image) is identified as the drooping side and marked as the affected side.
[0091] S2.3. Based on the determined affected side marker, perform a semantic unification operation on the affected side for all scale-normalized image frames corresponding to all action segments in the action segmentation binocular sequence. If the affected side is determined to be the right side in the original image, perform a horizontal flip operation on the image so that the affected side is presented on the same side of the image (e.g., the left side) in all image sequences. Perform a corresponding mirror transformation on the coordinates in the updated keypoint set to obtain semantically unified image frames and the final keypoint set.
[0092] Based on the final set of keypoints, functional zones are divided on the semantically unified image frame. Specifically, the regions enclosed by keypoints belonging to the left and right eye contours are merged and defined as the eye region; the region enclosed by keypoints of the outer lip contour is defined as the mouth region; and the regions enclosed by keypoints of the left and right eyebrow contours are merged and defined as the eyebrow region. A binary mask image with the same size as the semantically unified image frame is generated for each zone. In the mask image, the pixel positions belonging to the corresponding zone have a value of 1, and the other positions have a value of 0, thus generating the semantic masks for the functional zones of the eye region, mouth region, and eyebrow region.
[0093] S3. Based on the key point trajectory and functional partition semantic mask, perform 3D initialization of disparity for the action segmented binocular sequence and construct an anisotropic Gaussian set with partition attributes. Perform visibility-perceived Gaussian sputtering rendering and binocular reprojection consistency update on the anisotropic Gaussian set to obtain a 4D face Gaussian representation, a 3D deformation vector field and a midline plane.
[0094] S3.1. Perform feature extraction on each pair of corrected left and right eye image frames in the action segmentation binocular sequence to obtain a binocular feature map containing multi-scale texture and structural information.
[0095] Under the constraint of the functional partition semantic mask, stereo matching is performed on the binocular feature map to calculate the initial disparity value. Specifically, the matching cost of the left and right eye features on the horizontal epipolar line is calculated only at the pixel positions of the eye, mouth, and eyebrow regions indicated by the functional partition semantic mask. A disparity value is assigned to each pixel to generate the initial disparity map. The disparity values of pixels outside the partition are set to zero or invalid values.
[0096] To further explain, the expression for calculating the matching cost is:
[0097] ;
[0098] in, Represents the pixel coordinates in the left eye feature map Above, the disparity value of a pixel is At that time, the matching cost of the left and right eyes at corresponding positions, Indicates the channel index. Represents the pixel coordinates in the left eye feature map Up, passage eigencomponent values, This indicates the pixel coordinates of the feature map in the right eye. Above, the channel index is The characteristic component values.
[0099] Using the keypoint coordinates of the corresponding frames in the keypoint trajectory, an effective region is defined on the initial disparity map. With each keypoint coordinate as the center, a neighborhood range (e.g., 5×5 pixels) is preset, and the pixels within these neighborhoods are regarded as the effective region. The disparity points located within these effective regions in the initial disparity map are retained, and the disparity points in the remaining regions are set to invalid, thus obtaining the facial disparity region.
[0100] S3.2. Perform depth inversion on the facial disparity region. Using the intrinsic parameter matrix and baseline length in the camera calibration parameters, invert the 2D image coordinates and disparity value of each effective disparity point to 3D space through triangulation. The calculation formula is as follows:
[0101] ;
[0102] in, Represents the two-dimensional coordinates of a pixel on the image plane (usually taken as the coordinates of the image after left eye correction). Represents the disparity value of a pixel. Indicates the camera's focal length. Indicates the baseline length of the stereo camera. Indicates the coordinates of the principal point. This represents the coordinates of the three-dimensional point obtained through inversion.
[0103] The sparse 3D point set is formed by the 3D points obtained by inverting all effective disparity points.
[0104] Each 3D point in the sparse 3D point set is associated with a functional partition semantic mask to obtain a 3D point set with partition attributes. Specifically, each 3D point is back-projected back to the 2D image coordinates. The functional partition semantic mask is queried based on the 2D image coordinates to obtain the partition label (eye area, mouth area, or eyebrow area) corresponding to the pixel position, and this partition label is assigned as an attribute to the 3D point.
[0105] S3.3. Perform anisotropic parameterization on the 3D point set with partitioning attributes. For each 3D point, initialize a 3D Gaussian element. Each Gaussian element is defined by position parameters (mean, i.e., 3D point coordinates), a covariance matrix (representing the shape and direction of anisotropy), opacity, and color features. The covariance matrix is constructed from a scaling factor and an anisotropic rotation matrix. Each Gaussian element retains its associated partitioning attributes, and all Gaussian elements constitute an anisotropic Gaussian set with partitioning attributes.
[0106] Based on key point trajectories and image gradient information, within an anisotropic Gaussian set with partitioning attributes, Gaussian elements are further distinguished into surface layer Gaussian and boundary layer Gaussian according to partitioning attributes and 3D position. Specifically, Gaussian elements located in the main regions of the eye, mouth, and eyebrow areas with gentle gradients are marked as surface layer Gaussian. Based on the image gradient magnitude at the edge of the functional partitioning semantic mask, Gaussian elements located next to high-gradient anatomical structures such as the eye fissure contour, lip margin contour, and nasal alar groove contour are selected and marked as boundary layer Gaussian.
[0107] A visibility-aware Gaussian sputtering rendering and binocular reprojection consistency update are performed on the anisotropic Gaussian set. Specifically, firstly, based on the viewing parameters of the binocular camera, the visibility of each Gaussian element in the coordinate systems of the left and right cameras is calculated. Visibility is determined by whether the Gaussian element is within the camera's view frustum and its depth order. Each Gaussian element is projected onto the image plane. For each pixel covered by the projection, the depth value of the current Gaussian element is compared with that of other Gaussian elements projected onto the same pixel. The Gaussian element with the smallest depth value (closest to the camera) is considered visible at the pixel. The visibility results of all Gaussian elements on all pixels of the image are summarized to form a visibility set.
[0108] The Gaussian sputtering rendering process is driven by the visibility set. Each Gaussian element is superimposed and blended according to its opacity, color and projection on the image plane, generating corresponding rendered image frames in the left and right virtual viewpoints respectively.
[0109] S3.4. Compare the generated left and right eye rendered image frames with the corrected left and right eye image frames of the corresponding frames in the action segmented binocular sequence pixel by pixel, calculate the photometric error, and form the binocular reprojection consistency residual.
[0110] Furthermore, the expression for calculating photometric error is:
[0111] ;
[0112] in, This represents the consistency residual (photometric error). This represents the total number of valid pixels. Indicates pixel index, and These represent the rendered image frame and the image frame after true correction, respectively, in pixels. The color vector at that location.
[0113] Based on the calculated consistency residuals, the position parameters (mean), shape parameters (covariance matrix), and color feature parameters of each Gaussian element in the anisotropic Gaussian set are updated using the backpropagation optimization algorithm. During this optimization process, the partition attributes associated with the Gaussian elements remain unchanged.
[0114] The optimized and updated anisotropic Gaussian set is organized according to the time frame order of the action segmented binocular sequence. A cross-time frame correspondence is established for each Gaussian element to form a four-dimensional Gaussian representation of the face that can characterize the dynamic appearance and structure of the face.
[0115] Based on the inter-frame correlation of the established four-dimensional Gaussian representation of the face, the displacement vectors of Gaussian elements with corresponding relationships between adjacent time frames in three-dimensional space are calculated; the displacement vectors of all corresponding Gaussian elements are organized in three-dimensional space to generate a three-dimensional deformation vector field describing the overall motion of the face.
[0116] In the four-dimensional Gaussian representation of a face, all Gaussian units identified as natural resting motion segments are extracted. Based on the three-dimensional positions of these Gaussian units, a plane that reflects the approximate left-right symmetry of the face is fitted using principal component analysis to generate a midline plane. Specifically, the coordinates of the three-dimensional center points of all Gaussian units are collected, the covariance matrix is calculated, the covariance matrix is decomposed into eigenvalues, and the eigenvector corresponding to the smallest eigenvalue is taken as the normal vector of the fitted plane. The plane position is determined by the centroid of all points, generating a midline plane that reflects the main symmetric direction of the facial point cloud distribution.
[0117] S4. Perform healthy side mirror comparison using the midline plane and summarize the three-dimensional deformation vector field according to the partition attributes to generate dynamic deformation features. At the same time, extract texture and gradient features within the functional partition semantic mask and compare them with the template of the healthy sample library to obtain the partition rehabilitation index table and the movement quality table.
[0118] S4.1. By defining a mirror mapping through the midline plane, perform a healthy-side mirroring on the four-dimensional face Gaussian representation to obtain a mirrored Gaussian representation. Specifically, for each Gaussian cell in the four-dimensional face Gaussian representation located on the healthy side of the midline plane (i.e., the side opposite to the marked affected side), calculate the coordinates of the mirror point of the three-dimensional center point with respect to the midline plane, and use the mirror point as the position of the new Gaussian cell, while keeping the shape parameters, opacity, color features, and partitioning attributes of the Gaussian cell unchanged, thereby generating a mirrored Gaussian representation that is spatially symmetrical with the original four-dimensional face Gaussian representation.
[0119] The mirror Gaussian representation is matched with the four-dimensional face Gaussian representation to form a mirror pair. Specifically, in three-dimensional space, for each Gaussian cell in the original four-dimensional face Gaussian representation, the mirror Gaussian representation is searched for Gaussian cells with the same partition attributes (eye area, mouth area, or eyebrow area) and the closest spatial distance, and a one-to-one mirror pair is established.
[0120] S4.2. Based on the mirror pair relationship, the 3D deformation vector field is partitioned and its attributes are summarized to generate dynamic deformation features. For each time frame within each action segment, all mirror pairs are traversed, and for each Gaussian element pair, the following calculations are performed:
[0121] First, read the deformation vector of each pair of Gaussian elements at the corresponding time in the three-dimensional deformation vector field, and calculate the vector difference between the two deformation vectors. The magnitude of the vector difference is used to characterize the difference in left and right symmetry of the partition at this time.
[0122] Next, the magnitudes of the deformation vectors on the original side and the mirror side are calculated to characterize the deformation intensity on each side.
[0123] Finally, by analyzing the temporal evolution of the three-dimensional deformation vector field within the action segment, the spatial location (anchor point) where the deformation first appears in each partition and the time difference between the deformation wave propagating to the mirror counterpart position are determined, which are used to characterize the initiation response of the partition.
[0124] The calculation results (symmetry difference, deformation intensity, and start-up response time difference) of all mirror pairs are categorized and summarized according to the partition attributes (eye zone, mouth zone, eyebrow zone) to form a structured dynamic deformation feature.
[0125] S4.3. Cropping semantically unified image frames under the constraints of functional partition semantic masks to obtain partitioned image blocks. Specifically, logical AND operations are performed on each semantically unified image frame using functional partition semantic masks for the eye region, mouth region, and eyebrow region respectively to extract image regions containing only the corresponding partition pixels. The bounding rectangle of each region is then cropped to obtain eye partitioned image blocks, mouth partitioned image blocks, and eyebrow partitioned image blocks.
[0126] Texture encoding and gradient calculation are performed on the partitioned image blocks to form texture histogram features and gradient histogram features. Specifically, for all eye region partitioned image blocks extracted within the action segment labeled "slightly closing eyes", texture features are calculated using the local binary mode method and statistically analyzed into texture histograms to form eye region texture histogram features. For mouth region partitioned image blocks extracted within the action segment labeled "showing teeth and smiling", and eyebrow region partitioned image blocks extracted within the action segment labeled "raising eyebrows", their directional gradient histograms are calculated respectively to form mouth region gradient histogram features and eyebrow region gradient histogram features.
[0127] S4.4. Similarity calculations are performed on the texture histogram features, gradient histogram features, and dynamic deformation features with the healthy sample database template. The healthy sample database template is obtained by extracting texture and gradient features from the action segment binocular sequence that is consistent with the facial action sequence collected by healthy people under terminal guidance, after face localization, posture correction, scale normalization and semantic unification of the affected side. It includes the partition texture template, partition gradient template and dynamic deformation template corresponding to each action segment and each partition.
[0128] The chi-square distance between the eye region texture histogram features and the eye region texture template of the "slightly closed eyes" action segment in the healthy sample database is calculated to form the texture comparison result.
[0129] The cosine similarity between the histogram features of the mouth region gradient and the gradient template of the mouth region segment of the "smiling with teeth showing" action segment in the healthy sample database is calculated. At the same time, the cosine similarity between the histogram features of the eyebrow region gradient and the gradient template of the eyebrow region segment of the "raising eyebrows" action segment in the healthy sample database is calculated. Together, they form the gradient comparison result.
[0130] Calculate the Euclidean distance between the dynamic deformation features and the corresponding dynamic deformation templates of the action segments in the healthy sample library in terms of symmetry difference, deformation intensity, and start-up response time difference, and form deformation comparison results.
[0131] The texture matching results, gradient matching results, and deformation matching results are summarized according to the regional attributes (eye region, mouth region, eyebrow region). A set of indicators including appearance texture similarity, motion gradient similarity, and three-dimensional deformation similarity is generated for each region to form a regional rehabilitation index table.
[0132] The zonal rehabilitation index tables are summarized according to the sequence of movement segments (natural rest, light eye closure, showing teeth and smiling, raising eyebrows) to form a movement quality table that comprehensively describes the quality of completion of each movement segment.
[0133] S5. Perform follow-up alignment and trend calculation on the regional rehabilitation index table and the movement quality table, and combine the four-dimensional Gaussian representation of the face to perform reference surface uniformity to obtain the longitudinal rehabilitation trend results.
[0134] S5.1. Generate a follow-up index using the patient's unique identifier and the current collection timestamp. Associate the zonal rehabilitation index table and the movement quality table with the generated follow-up index to form a follow-up record containing all assessment results of this study.
[0135] For current follow-up records and historically stored follow-up records, perform same-partition matching. Specifically, ensure that image data in all follow-up records have been semantically unified according to the same side (affected side is uniformly on the left side of the image), and that all images and key points have been scaled to the same size. Based on this, align rehabilitation indicators and movement quality data under the same partition attribute (eye area, mouth area, eyebrow area) in different follow-up records to generate a time series of each partition across multiple follow-ups, i.e., a partition trend series.
[0136] For each partition trend series, sort the time series according to the follow-up timestamp; perform statistical analysis on the changes of the sorted time series, calculate the slope, mean difference or cumulative change of each evaluation index (such as texture similarity, symmetry difference, etc.) over time, and generate quantitative trend calculation results.
[0137] S5.2. For the four-dimensional Gaussian representation of the face, extract all Gaussian elements that are identified as natural resting action segments to form the reference Gaussian surface for this follow-up; using the iterative nearest point algorithm, register the reference Gaussian surface for this follow-up with the reference Gaussian surface of the first follow-up (or the specified baseline follow-up) in three-dimensional point cloud, and calculate the optimal rigid transformation matrix (including rotation and translation).
[0138] The optimal rigid transformation matrix obtained by application is used to perform spatial transformation on all three-dimensional data (such as three-dimensional vectors in dynamic deformation features) in this follow-up record, so that the three-dimensional data of this follow-up are aligned with the baseline follow-up in the same coordinate system, forming a follow-up record with consistent reference.
[0139] By combining consistent follow-up records and trend calculation results, a structured longitudinal rehabilitation trend result is generated. The longitudinal rehabilitation trend result includes a chart showing the longitudinal changes of indicators by region, a quality evolution analysis showing the movement segment, and a spatial change summary based on the alignment of the three-dimensional deformation vector field.
[0140] In summary, this invention achieves a highly consistent imaging foundation through binocular synchronous acquisition and epipolar correction. Combined with posture correction, scale normalization, and semantic unification of the affected side, it significantly reduces the interference of head movement, visual angle differences, and inconsistencies in left-right labeling on rehabilitation assessment, making data from different movement segments and follow-up time points comparable. Employing a three-dimensional representation with functional zoning attributes and a mirror comparison mechanism, it can precisely quantify the differences between the affected and healthy sides at local levels such as the eye, mouth, and eyebrow regions, improving the sensitivity and interpretability of minor functional recovery and compensatory movements. Simultaneously, it introduces template comparison from a healthy sample database, fusing multimodal features such as deformation and texture gradients into a zonal rehabilitation index table and a movement quality table, thereby achieving a structured output of "local indicators—movement quality—overall trend." Through follow-up alignment and reference surface consistency, it forms longitudinal trend results, which can stably reflect rehabilitation progress and support efficacy assessment and program adjustment. Overall, it improves the robustness, repeatability, and clinical readability of quantitative assessment, reducing reliance on subjective manual grading.
[0141] Example 2, referring to Table 1, is the second embodiment of the present invention. To further verify the technical solution of the present invention, experimental simulation data of a computer vision-based method for detecting the rehabilitation status of facial nerve diseases are provided.
[0142] One binocular camera and one mobile phone were selected as the terminal guidance devices. The binocular camera resolution was set to 1280×720 pixels, the frame rate to 30 frames per second, the baseline length to 60 mm, and the acquisition distance to 60 cm. The ambient illumination was controlled within the range of 450 to 550 lux and kept constant. The terminal guidance interface displayed "natural rest," "slightly closed eyes," "toothy smile," and "raised eyebrows" sequentially. Each action command lasted for 3 seconds, and the start and end times of the display were recorded. During the duration of the action command, the binocular camera simultaneously triggered left and right eye acquisition and assigned a frame pairing identifier to each pair of synchronized frames. At the same time, the acquisition timestamp and camera calibration parameters were recorded. The exposure time, aperture gain, and white balance parameters remained locked during the acquisition. The healthy sample library template was collected from 20 healthy individuals in the same action segment sequence. After face localization, pose correction, scale normalization, and semantic unification of the affected side, the samples were summarized by action segment to form a partitioned texture template, a partitioned gradient template, and a dynamic deformation template for subsequent similarity comparison.
[0143] After binocular synchronous acquisition is completed, stereo epipolar correction is performed on each pair of left and right eye image frames according to the camera calibration parameters, generating corrected left and right eye image frames. Based on the action segment sequence table, the start and end times and acquisition timestamps are displayed, and action start and end markers are written to the frame pairing identifier, encapsulating the data to obtain an action-segmented binocular sequence. Face detection and 68-key point localization are performed on each corrected left eye image frame of the action-segmented binocular sequence, generating keypoint trajectories. The rotation angle is calculated based on the center points of both eyes, and affine rotation is performed on the corrected left eye image frame to obtain a pose-corrected image frame. Simultaneously, the keypoint set is updated with the same rotation. The bounding rectangle is calculated based on the updated keypoint set, and the cropping box is expanded at a 1:1 aspect ratio. After cropping, it is scaled to 256×256 pixels to obtain a scale-normalized image frame. In the natural resting motion segment, the ordinates of point 48 and point 54 are compared. The side with the larger ordinate value is determined as the drooping side, and the affected side direction is marked. When the affected side direction is to the right, the scale-normalized image frame is horizontally flipped to achieve semantic unification of the affected side, and the keypoint set is updated by mirroring the coordinates. Based on the final keypoint set, semantic masks for the eye region, mouth region, and eyebrow region are generated.
[0144] Stereo matching is performed on the binocular feature map under the constraint of the functional partition semantic mask to generate an initial disparity map. The effective region is then defined using a 5×5 pixel neighborhood of key points to obtain the facial disparity region. Triangulation depth inversion is performed on the facial disparity region to obtain a sparse 3D point set. This sparse 3D point set is then queried using the functional partition semantic mask according to the back-projected 2D coordinates to obtain a 3D point set with partition attributes. An anisotropic Gaussian set with partition attributes is initialized using the 3D point set with partition attributes. Visibility-aware Gaussian sputtering rendering is performed, and the photometric error is calculated pixel-by-pixel with the corrected real image. The anisotropic Gaussian set is then optimized and updated through backpropagation to obtain a four-dimensional face Gaussian representation and a 3D deformation vector field. Finally, principal component analysis is performed on the coordinates of the Gaussian center point during a natural resting motion segment to fit the midline plane. Based on the midline plane definition of mirror mapping, the healthy side is mirrored to obtain the mirrored Gaussian representation of the four-dimensional face Gaussian representation. Partition attribute consistency matching is performed to form a mirror comparison relationship. The three-dimensional deformation vector field is summarized according to partition attributes to obtain dynamic deformation features. Simultaneously, under the functional partition semantic mask constraint, partition image blocks are cropped, and local binary pattern texture histogram features and directional gradient histogram features are extracted. Similarity calculations are performed with healthy sample library templates, and the results of dynamic deformation feature comparison are combined to form a partition rehabilitation index table, which is then summarized according to the action segment sequence table to obtain an action quality table. The comparison scheme uses the same acquired data and the same key point localization process, but cancels pose correction, affected side semantic unification, functional partition semantic mask constraint stereo matching, visibility-perceived Gaussian sputtering rendering, and binocular reprojection consistency update. Deformation is directly estimated and indexes are calculated using a one-time disparity inversion sparse three-dimensional point set.
[0145] The details are shown in Table 1 below:
[0146] Table 1 Comparative Data of Facial Nerve Disease Rehabilitation Testing
[0147] Parameter name Subject A01 Subject B02 Subject C03 Subject D04 Subject E05 Subject F06 Invention Scheme: Positioning Error of Start and End Frames of Action Segment (ms) 45 52 38 57 41 49 Comparison Scheme - Positioning Error of Action Segment Start and End Frames (ms) 210 185 240 195 225 205 Invention Solution: Keypoint Trajectory Jitter RMS (px) 1.8 2.1 1.6 2.3 1.9 2 Comparison of solutions_Keypoint trajectory jitter RMS (px) 4.6 4.1 5 4.4 5.2 4.7 Invention Scheme: Effective Parallax Ratio (%) within Functional Zones 78 74 82 70 76 73 Comparison Scheme_Effective Parallax Percentage within Functional Zones (%) 52 48 55 46 58 50 Invention Scheme: Standard Deviation of Natural Resting Depth Noise (mm) 2.4 2.7 2.1 2.9 2.5 2.6 Comparison Scheme_Standard Deviation of Natural Resting Depth Noise (mm) 6.1 5.4 6.8 5.9 7.2 6 Invention Scheme: Binocular Reprojection Consistency Residual E(0-1) 0.033 0.037 0.028 0.041 0.035 0.039 Comparison Scheme_Binocular Reprojection Consistency Residual E(0-1) 0.082 0.075 0.089 0.078 0.095 0.084 Invention Scheme: Four-dimensional face Gaussian representation cross-frame correspondence preservation rate (%) 96 94 97 92 95 93 Comparison Scheme: 4D Face Gaussian Representation Cross-Frame Retention Rate (%) 79 74 82 71 77 75 Invention Solution - Motion Quality Table Overall Score (0-100) 86 83 88 79 84 81 Comparison Scheme - Overall Score of Motion Quality Table (0-100) 68 62 71 58 65 60
[0148] As can be seen from the table data, the positioning error of the start and end frames of the action segment decreased from an average of 210 milliseconds in the comparison scheme to an average of 47 milliseconds in the invention scheme, and the time consistency between the action segment boundary positioning and the action segment sequence table was significantly improved; the jitter of the key point trajectory decreased from an average of 4.67 pixels in the comparison scheme to an average of 1.95 pixels in the invention scheme. The posture correction and scale normalization directly improved the stability of the key point trajectory, and the subsequent facial parallax region definition based on the key point neighborhood obtained a more stable effective region input.
[0149] The functional partition semantic mask constraint stereo matching increases the effective disparity ratio within the functional partition from an average of 51.5% in the comparative scheme to an average of 75.5% in the invention scheme, and simultaneously improves the density of sparse 3D point sets and the coverage of partition attribute annotations; the standard deviation of natural resting depth noise decreases from an average of 6.23 mm in the comparative scheme to an average of 2.53 mm in the invention scheme, and the 3D initialization of disparity initial values combined with anisotropic Gaussian set parameterization reduces the 3D inversion jitter in the resting segment, providing a more stable geometric basis for midline plane fitting and healthy side mirror comparison.
[0150] The binocular reprojection consistency residual decreased from the average of 0.0838 in the comparative scheme to the average of 0.0355 in the invention scheme. At the same time, the cross-frame correspondence retention rate of the four-dimensional face Gaussian representation increased from the average of 76.3% in the comparative scheme to the average of 94.5% in the invention scheme. The contribution of visibility-perceived Gaussian sputtering rendering and binocular reprojection consistency update to cross-frame consistency forms a quantifiable advantage at the data level. The comprehensive score of the action quality table increased from the average of 64.0 in the comparative scheme to the average of 83.5 in the invention scheme. The joint comparison of dynamic deformation features and partition texture and partition gradient showed higher discrimination and consistency at the action segment level, reflecting the technical effect of healthy side mirror comparison and partition attribute summary on intra-individual comparison evaluation.
[0151] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A computer vision-based method for detecting the rehabilitation status of facial nerve diseases, characterized in that: include, Under the guidance of the terminal, the binocular synchronous acquisition of facial motion sequences was completed, and epipolar correction and temporal marking were performed to obtain motion segmentation binocular sequences; Perform face localization, pose correction and scale normalization on the action segmented binocular sequence, and complete semantic unification of the affected side to generate key point trajectories and functional partition semantic masks. Based on the key point trajectory and functional partition semantic mask, the disparity initial value of the action segment binocular sequence is initialized in 3D and an anisotropic Gaussian set with partition attributes is constructed. The anisotropic Gaussian set is then subjected to visibility-perceptual Gaussian sputtering rendering and binocular reprojection consistency update to obtain a four-dimensional face Gaussian representation, a three-dimensional deformation vector field and a midline plane. The healthy side mirror comparison is performed using the midline plane, and the three-dimensional deformation vector field is summarized according to the partition attributes to generate dynamic deformation features. At the same time, texture and gradient features are extracted in the functional partition semantic mask and compared with the template of the healthy sample library to obtain the partition rehabilitation index table and the movement quality table. Follow-up alignment and trend calculation were performed on the regional rehabilitation index table and the movement quality table, and reference surface uniformity was performed by combining the four-dimensional Gaussian representation of the face to obtain the longitudinal rehabilitation trend results.
2. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The obtained action segmentation binocular sequence is specifically as follows: The terminal displays action commands such as natural rest, slightly closed eyes, showing teeth and smiling, and raising eyebrows, forming a sequence of action segments. The binocular camera synchronously acquires left and right eye image frames according to the action segment sequence table and generates frame pairing identifiers. Epipolar correction is performed on the left and right eye image frames. The epipolar constraints of the pixels are calculated using the camera's intrinsic and extrinsic parameters, and the action start and end markers are written according to the action segment sequence table. The action start marker, action end marker, and frame pairing identifier are associated to form an action segmentation binocular sequence.
3. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The generation of key point trajectories and functional partition semantic masks specifically involves: Face localization and key point localization are performed on the left and right eye image frames of the action segmented binocular sequence to obtain the key point set and generate the key point trajectory. Extract the center points of both eyes based on the key point trajectory, determine the direction of the line connecting the centers of both eyes, and drive attitude correction to obtain attitude-corrected image frames. The cropping box is determined based on the set of key points and the scale of the pose-corrected image frame is normalized to obtain the scale-normalized image frame. During the natural resting phase, determine and mark the direction of the affected side based on the direction of mouth drooping. Based on the affected side markers, semantic unification of the affected side is performed on the scale-normalized image frames to obtain semantically unified image frames; Based on the set of key points, the semantically unified image frame is divided into eye region, mouth region and eyebrow region, and corresponding functional partition semantic mask is generated.
4. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The construction of the anisotropic Gaussian set with partitioning attributes specifically involves: Feature extraction is performed on the left and right eye image frames of the action segmented binocular sequence to obtain binocular features; Under the constraint of functional partition semantic mask, the initial disparity value is estimated on the binocular features to obtain the initial disparity map; Under the constraint of key point trajectory, the effective disparity region of the initial disparity map is filtered to obtain the facial disparity region; Depth inversion is performed on the facial parallax region to obtain a sparse 3D point set; Associating each 3D point in the sparse 3D point set with a functional partition semantic mask yields a 3D point set with partition attributes. Anisotropic parameterization is performed on the 3D point set with partitioning properties to obtain anisotropic 3D Gaussian elements, while retaining the partitioning properties to form an anisotropic Gaussian set with partitioning properties.
5. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The process of performing visibility-aware Gaussian sputtering rendering and binocular reprojection consistency updates on anisotropic Gaussian sets specifically involves: The visibility of each Gaussian element in the anisotropic Gaussian set is calculated according to the binocular camera view, forming a visibility set; Gaussian sputtering rendering is driven by visibility sets, and rendered image frames are generated in both left and right viewpoints. Perform a binocular reprojection consistency check on the rendered image frames and the left and right eye image frames of the action segment binocular sequence to form a consistency residual. Based on the consistent residual-driven anisotropic Gaussian set, update the position and shape parameters of the Gaussian elements while keeping the partitioning properties unchanged; The updated anisotropic Gaussian set is used to establish inter-frame correlations along the action segment time sequence to form a four-dimensional Gaussian representation of the face. Based on inter-frame correlation, the displacement of Gaussian elements corresponding to adjacent frames in the four-dimensional face Gaussian representation is calculated to generate a three-dimensional deformation vector field. Based on the data from the Gaussian representation of a four-dimensional face during natural resting motion, the facial midline structure is fitted to generate a midline plane.
6. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The generation of dynamic deformation features specifically includes: The mirror mapping is defined by the midline plane, and the healthy side is mirrored on the four-dimensional Gaussian representation of the face to obtain the mirror Gaussian representation; The mirror Gaussian representation and the four-dimensional face Gaussian representation are matched for partition attribute consistency to form a mirror comparison relationship; Based on the mirror comparison relationship, the three-dimensional deformation vector field is partitioned and its attributes are summarized. The difference, magnitude and time difference of the deformation vector of the corresponding Gaussian element at the corresponding time are calculated to generate dynamic deformation features.
7. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The health sample database template is formed by extracting texture and gradient features from the action segment binocular sequence that is consistent with the facial action sequence collected by healthy people under the guidance of the terminal. After face localization, posture correction, scale normalization and semantic unification of the affected side, the template is summarized according to the action segment to form a partitioned texture template, a partitioned gradient template and a dynamic deformation template.
8. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The step of extracting texture and gradient features within the functional partition semantic mask and comparing them with templates from the healthy sample library specifically involves: Under the constraint of functional partitioning semantic mask, semantically unified image frames are cropped to obtain partitioned image blocks; Texture encoding and gradient calculation are performed on the partitioned image blocks to form texture histogram features and gradient histogram features; Similarity calculations are performed on the texture histogram features and the partitioned texture templates of the healthy sample library to generate texture alignment results; Similarity calculations are performed between the gradient histogram features and the partition gradient templates of the healthy sample library to generate gradient alignment results; Similarity calculation is performed between dynamic deformation features and dynamic deformation templates in the healthy sample database, and deformation comparison results are generated. The texture comparison results, gradient comparison results, and deformation comparison results are summarized according to the partition attributes to form a partition rehabilitation index table. The rehabilitation index tables for each region are summarized in order of movement segment to form a movement quality table.
9. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The following steps involve performing follow-up alignment and trend calculation on the zonal rehabilitation index table and the movement quality table: A follow-up index is generated using patient identifiers and collection timestamps as indexes; Link the regional rehabilitation index table and the movement quality table to the follow-up index to form a follow-up record; Based on the semantic uniformity rules for the affected side, follow-up records are matched against the same partition to generate a partition trend sequence; The trend series of the partitions is sorted by time and the changes are statistically analyzed to generate trend calculation results.
10. The method for detecting facial nerve disease rehabilitation based on computer vision as described in claim 1, characterized in that: The process of achieving reference surface uniformity specifically involves: The reference Gaussian surface is extracted from the natural resting action segment of the four-dimensional face Gaussian representation and registration is performed to form a reference-consistent follow-up record; By combining consistent follow-up records with trend calculation results, longitudinal recovery trend results are generated.