A Facial Paralysis Auxiliary Diagnosis Method Based on a Generative Model

By constructing a generative model to generate normal face images of patients with facial paralysis and assessing the degree of disease, the problems of lack of data and unstable feature extraction in facial paralysis diagnosis are solved, and efficient and accurate auxiliary diagnosis of facial paralysis is achieved, reducing data acquisition costs and improving the reliability of the diagnosis.

CN119920448BActive Publication Date: 2025-07-25ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510401997.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-25
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing facial paralysis diagnosis methods rely on clinical symptoms and imaging examinations, and have problems such as strong subjectivity, long diagnosis time and difficulty in early diagnosis. The deep learning model lacks sufficient data and lacks robustness in feature extraction, resulting in insufficient diagnostic accuracy and reliability.

Method used

A facial paralysis assisted diagnosis method based on the generative model is constructed, and normal face images of patients with facial paralysis are generated through the image repair model. The difference is calculated using structural similarity index and absolute difference, and the degree of disease of patients with facial paralysis is evaluated based on multi-scale statistical similarity indicators. A two-stage training method with gradual mask optimization is used to improve the generalization ability of the model.

Benefits of technology

It realizes efficient generation and accurate diagnosis of facial images after facial paralysis patients recovered, reduces the difficulty and cost of data collection, protects patient privacy, improves diagnosis accuracy and reliability, and enhances the generalization ability of the model and the feasibility of clinical application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920448B_ABST
    Figure CN119920448B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for auxiliary diagnosis of facial paralysis based on a generative model, comprising the following steps: S1, constructing a model framework for repairing missing images of normal human faces; S2, loading a portrait dataset, and adopting a two-stage training method with step-by-step mask optimization to train the model until convergence; S3, collecting digital images of facial paralysis patients and masking the diseased parts, and inputting the processed images into the model; S4, comparing the generated normal human face images with the original images, and jointly using the structural similarity index and absolute difference to calculate the differences, obtaining a diseased mask, and pasting it back to the original image to obtain a facial paralysis marked image; S5, using the SWD evaluation result as a quantitative index to evaluate the degree of illness of facial paralysis patients. The present invention can generate the facial images after rehabilitation of facial paralysis patients by using the facial images on the healthy side, and provide corresponding quantitative indexes, enabling doctors to more intuitively evaluate the severity of facial paralysis and formulate a more accurate treatment plan accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and deep learning, and particularly relates to a method for auxiliary diagnosis of facial paralysis based on a generative model. Background Art

[0002] Facial paralysis, also known as Bell's palsy, is a common neurological disorder characterized by unilateral or bilateral paralysis of the facial muscles. Currently, the diagnosis of facial paralysis mainly relies on clinical symptoms, physical examinations, and auxiliary examinations such as electrophysiological examinations and imaging examinations. However, these diagnostic methods have many limitations, such as strong subjectivity, long diagnostic time, and difficulty in early diagnosis.

[0003] In recent years, with the development of deep learning and generative model technologies, automatic diagnostic methods based on images and videos have gradually attracted attention. Generative models based on image inpainting, such as generative adversarial networks (GANs), variational autoencoders (VAEs), and Transformer models, have shown excellent performance in image generation and feature extraction, providing new ideas and methods for the auxiliary diagnosis of facial paralysis.

[0004] Feature extraction for facial paralysis needs to consider the subtle changes in facial muscles, including asymmetry of eyebrows, degree of eyelid closure, drooping or raising of the corners of the mouth, etc. These subtle changes are crucial for accurate diagnosis of facial paralysis because they can reflect the functional state of the facial nerve. However, existing methods have some deficiencies.

[0005] On the one hand, the image and video data of facial paralysis patients are relatively scarce, making it difficult to meet the requirements of deep learning models for a large amount of data. This may lead to overfitting when the model learns the features of facial paralysis and is unable to effectively generalize to new, unseen data, thus affecting the diagnostic accuracy of the model and limiting its application in clinical practice.

[0006] On the other hand, in practical applications, facial images may be affected by various factors such as lighting conditions, shooting angles, and expression changes. When existing methods deal with these interference factors, the robustness of features is often insufficient, resulting in unstable feature extraction results and further affecting the reliability of diagnosis. At the same time, facial images usually have high dimensionality, which makes feature extraction and processing complex. When existing methods deal with high-dimensional data, they often face the problem of dimensionality curse, resulting in high computational costs and low feature extraction efficiency.

[0007] The generation model based on image inpainting can automatically learn the latent distribution of data. Through the multi-scale feature extraction method, it can generate high-quality normal face images when half of the face of a facial paralysis patient is masked. This method can effectively utilize limited data to train a high-performance model and improve the accuracy and robustness of feature extraction. Although generation models have been widely used in many fields, they have not been used for the auxiliary diagnosis of facial paralysis yet. Therefore, there is an urgent need to develop an auxiliary diagnosis method for facial paralysis based on the generation model to solve the above problems. Summary of the Invention

[0008] The present invention proposes an auxiliary diagnosis method for facial paralysis based on a generation model. By generating a normal face image of a facial paralysis patient through an image inpainting model and then extracting facial paralysis features by comparing the generated normal face image with the actual facial paralysis image, it aims to solve the deficiencies of existing methods in terms of data and feature extraction.

[0009] To this end, the present invention provides an auxiliary diagnosis method for facial paralysis based on a generation model, including: S1. Construct the framework of a normal face missing image inpainting model to generate a predicted normal face image after masking the diseased part in the facial paralysis patient's image; S2. Load the portrait dataset and use a two-stage training method with step-by-step mask optimization to train the model until convergence and save the model weights to obtain a generation model; S3. Collect digital images of facial paralysis patients, normalize them, mask the diseased parts in the images while retaining the normal areas of the face, and input the processed images into the normal face missing image inpainting model; S4. Compare the generated normal face image with the original image, jointly use the structural similarity index SSIM and absolute difference absdiff to calculate the difference, obtain the lesion mask, and paste it back to the original image to obtain a facial paralysis marked map; S5. Use the multi-scale statistical similarity index SWD to evaluate the similarity between the generated image and the input image, and use the obtained SWD evaluation result as a quantitative index to evaluate the degree of illness of facial paralysis patients; S6. Calculate the perceptual evaluation index FID and the local similarity evaluation index SSIM to evaluate the quality of the generated image.

[0010] The present invention also provides a computer-readable storage medium, including computer programs / instructions, which are used to implement the steps of the above-mentioned auxiliary diagnosis method for facial paralysis based on a generation model when executed.

[0011] The present invention also provides an auxiliary diagnosis system for facial paralysis based on a generation model, including computer programs / instructions, which are used to implement the steps of the above-mentioned auxiliary diagnosis method for facial paralysis based on a generation model when executed.

[0012] The present invention also provides a computer program product, including computer programs / instructions, which are used to implement the steps of the above-mentioned auxiliary diagnosis method for facial paralysis based on a generation model when executed.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0014] (1) The present invention generates a normal face image of a facial paralysis patient through a generative model, compares it with the real image of the patient, and calculates the total difference value, providing a new and effective solution for the auxiliary diagnosis of facial paralysis. This method can not only quantify the severity of facial paralysis, but also provide a more intuitive diagnostic basis for clinicians, thereby improving the accuracy and reliability of diagnosis.

[0015] (2) Although previous studies have made progress in generating facial images after the rehabilitation of facial paralysis patients, existing models usually rely on datasets of facial information of facial paralysis patients, which have significant limitations. First, the scarcity of facial paralysis patient samples makes data collection difficult and costly, and it is difficult to obtain a sufficient number of high-quality data. Second, the collection of facial data of facial paralysis patients may involve privacy issues and trigger scientific research ethics issues, further restricting the collection and application of data. In contrast, the present invention proposes an innovative method that does not require directly collecting the images of facial paralysis patients as the training set. The method of the present invention uses a normal portrait dataset and analyzes the normal face information through a deep learning model to predict the facial image after the rehabilitation of facial paralysis patients. This method has significant advantages in generating facial images after the rehabilitation of facial paralysis patients because it avoids the need to directly collect the images of facial paralysis patients and can achieve excellent image restoration effects only using a normal face image dataset, thereby reducing the difficulty and cost of data collection and protecting the privacy of patients.

[0016] (3) Although existing generative models based on image inpainting perform well in repairing small-area missing images, their performance significantly degrades when dealing with large-area occluded regions. These models usually have difficulty precisely reconstructing the features of large-area missing regions on the face. The method of the present invention involves occluding a large area on the affected side of a facial paralysis patient and generating a normalized facial image, which requires the image inpainting model to be able to precisely process large-area facial occluded regions. Therefore, developing an image inpainting model that can effectively process such large-area facial occluded regions is of great significance for improving the accuracy and practicality of the auxiliary diagnosis of facial paralysis.

[0017] (4) The present invention adopts a two-stage training method with progressive mask optimization. In the first stage, through random masking processing, the model can learn the facial restoration ability under different missing conditions, laying a foundation for subsequent more complex tasks. In the second stage, through specific semi-occlusive masks, the adaptability and restoration accuracy of the model to specific missing patterns are further improved, enhancing the pertinence and effectiveness of the model in the auxiliary diagnosis task of facial paralysis. This phased training method not only improves the generalization ability of the model, but also can more accurately generate normal face images when facing new and unseen facial paralysis images, thereby improving the accuracy of diagnosis.

[0018] (5) The method of the present invention is not only innovative in theory, but also highly feasible in actual clinical applications. By generating images of facial paralysis patients after recovery, doctors can more intuitively evaluate the severity of facial paralysis and formulate more precise treatment plans. In addition, the efficiency and accuracy of this method enable it to be quickly applied to large-scale clinical diagnoses, improving medical efficiency and patient prognosis.

[0019] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings to further elaborate on the present invention in detail. Brief Description of the Drawings

[0020] The specification drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0021] Figure 1 It is a schematic flowchart of a method for auxiliary diagnosis of facial paralysis based on a generative model of the present invention;

[0022] Figure 2 It is a generative model diagram of the present invention;

[0023] Figure 3 It is a structural block diagram of the Transformer main part in the generative model of the present invention;

[0024] Figure 4 It is an example of the mask update strategy in the generative model of the present invention;

[0025] Figure 5 It is an execution effect diagram of the method for auxiliary diagnosis of facial paralysis based on a generative model in an embodiment of the present invention. Detailed Embodiments

[0026] The following will refer to the drawings and combine with embodiments to elaborate on the present invention in detail.

[0027] Please refer to Figure 1 、 Figure 2 andFigure 5 , the facial paralysis assisted diagnosis method based on the generative model in this embodiment includes the following steps S1 - S6.

[0028] S1. Construct a model framework for repairing missing images of normal faces to generate predicted normal face images after masking the diseased parts in the images of facial paralysis patients.

[0029] Select a generative model framework based on image inpainting. This model is specifically designed for the large - hole image inpainting task, such as CoModGAN, LaMa, DeepFill v2, etc. Different from the good performance of most image inpainting models on small - area missing images, the existing models have unsatisfactory effects when dealing with large - area masked regions. After comparative analysis of multiple large - hole image inpainting models, the model framework constructed in the present invention is carefully designed and optimized. It can not only extract the half - face features of facial paralysis patients more efficiently, but also generate highly realistic normal face images when half of the face is masked. The architectural features of this framework are that it includes a convolutional head, a Transformer body, and a convolutional tail, as Figure 2 shown.

[0030] Constructing this model specifically includes the following steps S101 - S103.

[0031] S101. Construct the convolutional head part: This part contains five convolutional layers. The first convolutional layer is used to adjust the input dimension, expanding the number of input channels from 4 (3 channels for RGB images and 1 channel for the mask) to 180. The next three convolutional layers are used for downsampling the resolution. The stride of each convolutional layer is 2, gradually reducing the size of the feature map from downsampling to . The number of channels of the fifth convolutional layer is 90, which is specifically used to extract half - face features to enhance the model's ability to capture the half - face features of facial paralysis patients, thereby improving the diagnostic accuracy. The advantage of the fifth convolutional layer is to enhance the model's sensitivity to local features, making the generated normalized images closer to the real situation in terms of half - face details, which is beneficial for subsequent diagnostic analysis;

[0032] The convolutional head part receives the incomplete face image and the given mask M:

[0033]

[0034] where I is the original image, M is the mask, denotes element - wise multiplication.

[0035] S102. Construct the main part of the Transformer: This part contains five stages of adjusted Transformer blocks. Each stage of the adjusted Transformer block has the following adjustments: removing layer normalization, adopting feature interaction and fusion learning, applying an optimized multi-head attention mechanism, performing window shifting, and adding global residual connections. Specifically, stage 1 contains two adjusted Transformer blocks with a feature size of ; stage 2 contains three adjusted Transformer blocks with a feature size of ; stage 3 contains four adjusted Transformer blocks with a feature size of ; stage 4 contains three adjusted Transformer blocks with a feature size of ; stage 5 contains two adjusted Transformer blocks with a feature size of ;

[0036] Figure 3 And Figure 4 respectively show the structure of an optimized single Transformer module and an example of the mask update strategy in the multi-head context attention mechanism. These optimizations aim to improve the accuracy of distance-dependent modeling. In particular, the multi-head context attention mechanism effectively identifies and utilizes valid tokens for distance-dependent modeling using dynamic masks. In addition, the improved Transformer module enhances the stability of the training process when dealing with complex masks.

[0037] In traditional Transformer blocks, layer normalization is applied before each sub-module (such as the multi-head self-attention module and the MLP module), and residual connections are applied afterwards. The specific structure is as follows:

[0038]

[0039] Among them, x is the input feature tensor, represents applying layer normalization operation to the input feature tensor x, represents processing the input feature tensor x with a sub-module, represents processing the input feature tensor that has undergone layer normalization with a sub-module.

[0040] This structure helps to stabilize the training, but in the facial paralysis auxiliary diagnosis task, especially when dealing with large-hole masks, layer normalization may amplify useless tokens. In the facial images of facial paralysis patients, there are significant differences between the normal area and the damaged area (i.e., the area affected by facial paralysis), and large-hole masks usually cover the damaged area.

[0041] During the training process, layer normalization performs unified normalization on all tokens (including tokens in normal and damaged regions), which may cause useless tokens in the damaged regions (tokens carrying less useful information) to be over-amplified, thus interfering with the model's learning of key features and making the training process unstable. Therefore, to better adapt to the facial paralysis auxiliary diagnosis task, the adjusted Transformer block removes layer normalization. This adjustment enables the model to focus more on learning and extracting information that is truly useful for facial paralysis diagnosis, avoiding interference from useless information, and thus improving the training stability and diagnostic accuracy of the model in the facial paralysis auxiliary diagnosis task.

[0042] Instead of traditional residual learning, the adjusted Transformer block concatenates the output of the multi-head context attention module with the input features and performs fusion learning through the feature interaction module; the specific operations are as follows:

[0043]

[0044] Among them, is the output of the feature interaction module of the l th block in the k-th stage, represents the output of the feature interaction module of the l -1-th block in the k-th stage, represents the fully connected layer, represents the multi-head context attention mechanism, is the feature interaction module, and the specific implementation can be a multi-layer perceptron (MLP):

[0045]

[0046] Among them, is the activation function, , are learnable weights, , is the learnable bias.

[0047] Meanwhile, a dropout mechanism is introduced in the feature interaction module to randomly discard some features and enhance the generalization ability of the model:

[0048]

[0049] Among them, is the output feature tensor, is a regularization technique used to prevent the model from overfitting to the training data and improve the generalization ability of the model.

[0050] The multi-head context attention mechanism (MCA) uses a dynamic mask to select valid tokens and calculates the weighted sum of the valid tokens; the attention in the multi-head context attention module is calculated using the following formula :

[0051]

[0052] where Q, K, and V are the query matrix, key matrix, and value matrix respectively, is the scaling factor, which is used to prevent the dot product result from being too large and causing the gradient to vanish. is a dynamically generated mask, which is used to represent the valid tokens in the input image that are not covered by the mask. This mask is initialized by the input mask and is automatically updated during the propagation process. The valid tokens will be used for attention calculation, while the invalid tokens will be ignored.

[0053] Dynamic mask is represented as:

[0054]

[0055] where the aggregation weights of the invalid tokens are basically 0.

[0056] The generative model proposed in the present invention based on image inpainting adopts a dynamic mask update strategy to optimize the facial feature reconstruction, especially for dealing with large-area missing facial features in the auxiliary diagnosis of facial paralysis. Figure 4 The implementation process of this strategy is described in detail, and the specific steps are as follows: The feature map is initially divided into 2×2 windows (shown in red in the figure), and each window contains several tokens. Figure (a) shows the initial state, where some tokens are marked as invalid (the gray slashed squares represent the invalid areas); Figure (b) shows that under the action of the attention mechanism, all the tokens in the upper left window first become valid (the squares not covered by the slashes represent the valid areas), which marks the start of the mask update, and the valid tokens will participate in the subsequent image reconstruction; Figure (c) shows that the mask window moves pixels to the lower right, preparing to process the next window. At this time, the original upper left window becomes the middle window, and its token state remains unchanged; Figure (d) shows the mask update of the tokens in the new window to make them valid, repeating the process from Figure (a) to Figure (b), but applied to the new position; Figure (e) shows that the mask window continues to move pixels to the lower right to ensure that all the 2×2 windows in the initial state are processed, realizing the global feature reconstruction; Figure (f) shows the last mask update of the original lower right window to ensure that all the tokens finally become valid, completing the mask update process.

[0057] After each attention operation, the mask is adjusted according to the following strategy: The mask will be automatically updated based on the results of the attention operation to reflect which regions have been successfully reconstructed during the calculation process. As long as there is at least one valid token in the window, all tokens in the window will be updated to valid after the attention. If all tokens in the window are invalid, they will remain invalid after the attention. This update rule helps to gradually expand the region of valid pixels, thereby incorporating more context information in subsequent processing. By this method, the model can more accurately reconstruct large areas of missing facial features, which is particularly important for the auxiliary diagnosis of facial paralysis.

[0058] In addition, after each attention operation, the position of the window of size is moved by

[0059] pixels to achieve cross-window connection. This cross-window connection mechanism further enhances the model's understanding of the global structure of the image, thereby improving the accuracy and efficiency of inpainting images with large occlusion areas.

[0059] After multiple Transformer blocks, the model described in S1 achieves global residual connection through a convolutional layer to maintain the integrity of features. Specifically, the global residual connection uses the convolutional layer to transform the output features of the last Transformer block, and then adds the transformed features to the input features element-wise to generate the final output features. This mechanism not only ensures the effective transfer of features between different levels, but also improves the optimization stability of the model and the quality of the generated images.

[0060] S103. Construct the convolutional tail part.

[0061] This part uses a convolutional-based reconstruction module that can restore the spatial resolution of the feature map to the size of the input image through upsampling operations. The output of the Transformer main body is upsampled to match the size of the input image, thereby generating the inpainted face image.

[0062] S2. Load the portrait dataset, adopt a two-stage training method with gradual mask optimization, train the model until convergence, and save the model weights to obtain the generative model.

[0063] The CelebA-HQ portrait dataset used in this invention is derived from the original CelebA dataset and contains more than 200,000 celebrity facial images. This dataset consists of high-definition normal face images, which are preprocessed in multiple steps to improve quality. The specific preprocessing steps include: eliminating JPEG artifacts to reduce the impact of artifacts during JPEG compression; performing super-resolution processing to enhance image resolution and clarity; and high-quality resampling to ensure high quality during image sampling. The high-quality images in the preprocessed CelebA-HQ dataset reduce noise and artifact interference, enabling the trained generative model to effectively learn the details and structures of the face, thereby improving the quality and authenticity of face-generated images. The generative model trained using this CelebA-HQ dataset can effectively restore facial features and can thus be well used to infer the normal faces of patients with facial paralysis.

[0064] In this example, the dataset images are randomly divided into a training image set and a validation image set in a ratio of 7:3.

[0065] The two-stage training method using step-by-step mask optimization includes the following steps S201-S202.

[0066] S201. In the first stage, according to a preset rule, a mask is randomly generated on the training dataset to block part of the image area and simulate the missing content of the image. The training process is repeated continuously until the generative model reaches the preset convergence standard on the validation set, and then the model weights are saved. The specific training process is as follows.

[0067] The CelebA-HQ dataset is divided into multiple batches, and the images in each batch are randomly masked and input into the model constructed in S1 for face image restoration;

[0068] Traverse all batches of the dataset. In each training step, the model described in S1 generates a face image and calculates the loss function;

[0069] The adversarial loss is calculated using the following formula:

[0070] ,

[0071]

[0072] where, is the loss of the generator, is the loss of the discriminator, x and are the real image and the generated image respectively, and respectively represent the average of all possible values of the real image x and the generated image , and They respectively represent the probabilities that the real image and the generated image are judged as real images.

[0073] The generator aims to generate realistic images that make it difficult for the discriminator to distinguish the generated images from the real images; the discriminator is committed to accurately distinguishing between the two.

[0074] The feature matching loss is calculated using the following formula :

[0075]

[0076] where represents the activation value of the i-th layer of the pre-trained VGG-19 network, is the number of features in this layer, n is the number of selected layers, is the L1 norm.

[0077] The feature matching loss ensures that the generated images match the real images on multiple feature layers.

[0078] At the same time, an R1 regularization term is introduced into the loss function to improve the final performance of the generative model and thus generate higher-quality images:

[0079]

[0080] where is the output of the discriminator for the image x, is the gradient of the discriminator for the real image x, represents the L2 norm of the gradient.

[0081] The R1 regularization term prevents the discriminator from being overconfident by punishing the discriminator for having too large a gradient for the real image, thereby improving the stability of the generative model and the quality of the generated images.

[0082] The total loss is calculated using the following formula:

[0083]

[0084] where , and are hyperparameters, which are set to 0.5, 20, and 0.2 respectively in this example, and are used to balance the influence of different loss terms.

[0085] The loss of the model is calculated according to the loss function, and the parameters of the model are updated through backpropagation.

[0086] S202. In the second stage, using the weights of the generative model obtained from the first stage training, generate a semi - occluded mask for facial paralysis. Specifically, control the mask to cover 50% of the image, that is, occlude the left or right half of the face image to further improve the generalization ability of the model. Subsequently, retrain the generative model using the dataset with the specific mask. The training continues until the generative model meets the optimization criteria on the target task, and save the model weights. The specific training steps are as follows.

[0087] Extract images from the dataset of facial paralysis patients and apply the generated semi - occluded mask;

[0088] Input the masked image into the model constructed in step S1 for face image restoration;

[0089] Traverse all batches of the dataset. In each training step, the model generates a restored face image and calculates the loss function;

[0090] Calculate the loss of the model according to the total loss function described in S201, and update the model parameters through backpropagation;

[0091] Continue training until the generative model meets the optimization criteria on the target task, and save the model weights.

[0092] The model training uses the Adam optimizer, sets the first - order moment estimate ( ), and the second - order moment estimate ( ), the batch size is 64, and the initial learning rate is set to .

[0093] S3. Collect digital images of facial paralysis patients, normalize them, and occlude the diseased parts in the images while retaining the normal areas of the face. Input the processed images into the normal face missing image restoration model.

[0094] Step S3 specifically includes the following steps S301 - S302.

[0095] S301. Use a high - resolution camera device to collect facial images of facial paralysis patients, then normalize them, and occlude the abnormal side of the patients' faces;

[0096] The occlusion process aims to simulate the absence of facial content so that the model can more accurately reconstruct these areas.

[0097] S302. Input the occluded facial paralysis patient images into the generative model constructed in step S1, and set a random seed to generate multiple predicted normal face images.

[0098] S4. Compare the generated normal face image with the original image, jointly use the structural similarity index SSIM and absolute difference absdiff to calculate the difference, obtain the lesion mask, and paste it back onto the original image to obtain the facial paralysis marking map.

[0099] Step S4 specifically includes the following steps S401 - S407.

[0100] S401: Calculate the structural similarity index (SSIM) to quantify the structural similarity between the generated normal face image and the original image;

[0101] SSIM is an index used to measure the structural similarity between two images, especially suitable for measuring the structural similarity of images. The calculation formula is as follows:

[0102]

[0103] where x and y are two images to be compared, and are the means of images x and y respectively, and are the variances of image x and image y respectively, is the covariance of images x and y, and are small constants used to avoid the denominator being zero, usually taking and where L is the dynamic range of the image, and are small constants, which are taken as and .

[0104] S402: Calculate the absolute difference (absdiff) to quantify the pixel - level difference between the generated normal face image and the original image;

[0105] The absolute difference (absdiff) is a simple method for calculating pixel - level differences, used to measure the pixel - level difference between two images. The calculation formula is as follows:

[0106]

[0107] where x and y are the generated normal face image and the original image respectively, represents the absolute difference between the two images at each pixel position.

[0108] S403: Generate a difference map to highlight the abnormal areas between the normal face image and the original image;

[0109] Specifically, the SSIM value is converted into a difference map ssim_diff, and 1 - SSIM is used to represent the difference. The calculation result of the absolute difference absdiff is normalized to the range of [0, 1] for visualization, and the normalized absolute difference map is used as the difference map diff_map:

[0110] ,

[0111] 。

[0112] S404. Use the method of weighted sum to combine the SSIM difference map and the absdiff difference map to generate a comprehensive difference map. The calculation formula is as follows:

[0113]

[0114] where combined_diff represents the comprehensive difference map, is a weight parameter, usually between [0, 1], used to balance the influence of SSIM and absdiff. In this example, takes the value of 0.5.

[0115] S405: Combine the results of SSIM and absdiff, and perform thresholding on the comprehensive difference map to generate a lesion mask, which highlights the regions in the original image that are significantly different from the generated normal face images;

[0116] Select a threshold T. The selection of the threshold T can adopt a data-driven method or an adaptive method. Mark the part of the comprehensive difference map greater than T as the lesion region:

[0117] 。

[0118] S406. Repeat steps S401 to S405, compare multiple normal face images generated in step S3 with the original image, generate multiple different masks and fuse them to generate a more accurate and comprehensive lesion mask.

[0119] Vote on each mask, and select the regions that most masks consider to be lesions to reduce the misjudgment of a single mask:

[0120]

[0121] where N is the number of generated masks, represents the i-th generated mask.

[0122] S407. Paste the generated lesion mask back onto the original image to visualize the lesion area and assist doctors in diagnosis. Specifically, first read the original image y, and then apply the lesion mask final_mask to the original image y using the following formula to highlight the lesion area:

[0123]

[0124] where highlighted_image represents a visualized image, and color_mask is a color mask used to highlight the lesion area. In this example, red is selected.

[0125] S5. Use the multi-scale statistical similarity metric SWD to evaluate the similarity between the generated image and the input image, and use the obtained SWD evaluation result as a quantitative metric to evaluate the degree of illness of facial paralysis patients.

[0126] In this step, use the model that has been trained to convergence, input the images of facial paralysis patients into the model, and generate predicted normal face images. These images will be used for subsequent similarity evaluation.

[0127] Use the multi-scale statistical similarity metric SWD to evaluate the similarity between the generated image and the input image.

[0128] Evaluate the similarity between the generated normal face image and the original facial paralysis patient image through the multi-scale statistical similarity metric SWD. SWD can quantify the similarity between images and provide a numerical evaluation result.

[0129] Use the obtained similarity evaluation result as a quantitative metric to evaluate the degree of illness of facial paralysis patients.

[0130] Use the SWD evaluation result as a quantitative metric to evaluate the degree of illness of facial paralysis patients. The lower the SWD value, the higher the similarity between the generated image and the input image, and thus the lower the degree of illness; on the contrary, the higher the SWD value, the lower the similarity and the higher the degree of illness.

[0131] The multi-scale statistical similarity metric SWD is a metric that measures the distance between two probability distributions and is particularly suitable for high-dimensional data. SWD approximates the distance of high-dimensional distributions by projecting high-dimensional data onto multiple random directions and then calculating the distances of one-dimensional distributions on these projections. The specific calculation process is as follows.

[0132] First, convert the generated image and the real image into Laplacian pyramid representations, and start from the low resolution and gradually increase the resolution until the resolution of the original image is reached.

[0133] At each scale, local image patches are randomly extracted from the generated images and the patient images. These image patches can be fixed-size windows, such as 7x7 pixels. A large number of image patches are extracted from each image to form two sets: the set of generated image patches and the set of patient image patches.

[0134] Then, the extracted image patches are normalized to have a mean of 0 and a standard deviation of 1, eliminating the scale differences between different image patches.

[0135] For each scale, the Sliced Wasserstein Distance (SWD) between the set of generated image patches and the set of real image patches is calculated. Specifically, multiple random directions are selected, the image patches are projected onto these directions, and then the distance between the projected one-dimensional distributions is calculated. The SWD is calculated using the following formula:

[0136]

[0137] where, and are the sets of the i-th generated image patch and the j-th patient image patch at the s-th scale respectively, and are the inner products of the projections in the k-th random direction with the i-th generated image patch and the j-th patient image patch at the s-th scale respectively, is the K random directions, K is the number of random directions, and N is the number of image patches.

[0138] Next, the SWD values at different scales are aggregated by weighted averaging to obtain a comprehensive multi-scale statistical similarity metric. When aggregating the SWD values at different scales, a weight factor is introduced, and this weight factor is adjusted according to the symmetry difference of the image patches at each scale. Specifically, if the symmetry difference of the image patches at a certain scale is large, it indicates that the features at this scale are more important for the diagnosis of facial paralysis, so a higher weight is given. The features at different scales contribute differently to the diagnosis of facial paralysis. By introducing the weight factor, the importance of the features at each scale can be more accurately reflected, thereby improving the diagnostic value of the comprehensive metric. The specific formula is as follows:

[0139]

[0140] where, is the weight factor at the s-th scale, is the SWD value at the s-th scale. The weight factor is adjusted according to the symmetry difference at each scale, and the specific calculation formula is as follows:

[0141]

[0142] Among them, is the symmetry difference at the s-th scale, and the calculation formula is as follows:

[0143]

[0144] Among them, and are the left and right image patches of the i-th image patch at the s-th scale respectively, is the Euclidean norm.

[0145] This index can be used to quantitatively evaluate the degree of illness of patients with facial paralysis. Specifically, the lower the SWD value, the higher the similarity between the generated image and the patient's image at multiple scales, and thus the lower the degree of illness; on the contrary, the higher the SWD value, the lower the similarity between the generated image and the patient's image at multiple scales, and thus the higher the degree of illness.

[0146] S6. Calculate the perceptual evaluation index FID and the local similarity evaluation index SSIM to evaluate the quality of the generated image.

[0147] Step S6 specifically includes the following steps S601 - S603.

[0148] S601: Calculate the FID perceptual evaluation index.

[0149] Calculate FID using the following formula:

[0150]

[0151] Among them, represents the squared Euclidean distance between the mean vectors, represents the trace of the matrix (i.e., the sum of the diagonal elements of the matrix), represents the square root of the product of two covariance matrices.

[0152] FID evaluates the overall distribution of the image set and can more comprehensively measure the quality of the generated image.

[0153] S602. Calculate the SSIM evaluation index.

[0154] Calculate SSIM using the following formula:

[0155]

[0156] Among them, x and y are two images to be compared, and are the means of image x and image y respectively, and are the variances of image x and image y respectively, is the covariance of images x and y, and is a small constant used to avoid a zero denominator.

[0157] SSIM evaluates the local similarity of images and has good robustness to image noise and small perturbations, enabling more accurate evaluation of image quality.

[0158] S603. Calculate the comprehensive quantization error.

[0159] Calculate the comprehensive quantization error TotalError using the following formula:

[0160]

[0161] where and are weight parameters, usually between [0, 1], used to balance the influence of FID and SSIM.

[0162] In this example and both take the value of 0.5.

[0163] The above is only an embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A facial paralysis assisted diagnosis method based on a generative model, characterized in that, Including: S1. Construct the framework of a normal face missing image restoration model to generate a predicted normal face image after masking the lesion part in the image of a facial paralysis patient; S2. Load the portrait dataset, adopt a two-stage training method with step-by-step mask optimization, train the model until convergence, and save the model weights to obtain a generation model; S3. Collect digital images of facial paralysis patients, normalize them, mask the lesion parts in the images, while retaining the normal areas of the face, and input the processed images into the normal face missing image restoration model; S4. Compare the generated normal face image with the original image, jointly use the structural similarity index SSIM and the absolute difference absdiff to calculate the difference, obtain the lesion mask, and paste it back to the original image to get a facial paralysis marking map; S5. Use the multi-scale statistical similarity index SWD to evaluate the similarity between the generated image and the input image, and use the obtained SWD evaluation result as a quantitative index to evaluate the degree of illness of facial paralysis patients; S6. Calculate the perceptual evaluation index FID and the local similarity evaluation index SSIM to evaluate the quality of the generated image. In step S2, the two-stage training method includes the following steps: S201. The first stage: Randomly generate masks on the training dataset to occlude part of the face image, simulate the missing content of the face image, and repeat the training process until the generation model reaches the preset convergence standard on the validation set, and then save the model weights; S202. The second stage: Use the weights of the generation model obtained in the first stage of training to generate a masking mask for half of the face, that is, mask the left or right half of the face image, and repeat the training process until the generation model reaches the preset convergence standard on the validation set, and then save the model weights.

2. The facial paralysis assisted diagnosis method based on a generative model according to claim 1, wherein The framework of the normal face missing image restoration model in step S1 introduces an improved Transformer structure and a dynamic mask update strategy after each attention calculation. Step S1 includes: S101. The model architecture consists of a convolutional head, an improved Transformer body, and a convolutional tail. The steps to construct the framework of the normal face missing image restoration model are as follows: Construct a convolutional head part, which includes five convolutional layers. The first convolutional layer is used to adjust the input dimension, expanding the number of input channels from 4 channels to 180 channels. The next three convolutional layers are used for downsampling the resolution, with a stride of 2 for each convolutional layer, gradually reducing the size of the feature map from downsampled to ; The fifth convolutional layer has 90 channels and is used to extract half-face features; Build the main part of the Transformer, which includes five processing stages. In the first stage, the feature size of the Transformer block is ; in the second stage, the feature size of the Transformer block is ; in the third stage, the feature size of the Transformer block is ; in the fourth stage, the feature size of the Transformer block is ; in the fifth stage, the feature size of the Transformer block is ; Construct the convolutional tail part, which adopts a convolution-based reconstruction module to restore the spatial resolution of the feature map to the size of the input image through upsampling operations, so as to generate a restored normal face image; S102. The improved Transformer blocks all use a non-normalized structure, which includes a multi-head context attention module MCA and a feature interaction module. Among them, the output of the multi-head context attention module is concatenated with the input features and then fused and learned through the feature interaction module. The MCA module uses a dynamic mask to select effective tokens, calculates the weighted sum of the effective tokens, and a dropout mechanism is introduced in the feature interaction module; S103. The dynamic mask update strategy follows the following steps: After each attention operation, the mask is automatically updated to reflect the successfully reconstructed regions. If there is at least one valid token in the window, all tokens within that window are updated to valid; if all tokens within the window were originally invalid, they remain invalid after the attention operation. Through continuous window movement and repeated application of the attention mechanism, the mask will be gradually updated until it becomes completely valid. This process helps to gradually expand the region of valid pixels and integrate more context information, thereby improving the accuracy of facial feature reconstruction in facial paralysis assisted diagnosis.

3. The facial paralysis assisted diagnosis method based on a generative model according to claim 2, wherein The method for the MCA module to calculate the weighted sum of valid tokens is as follows: , Among them, is a dynamically generated mask used to represent valid tokens in the input image that are not covered by the mask. The dynamic mask is expressed as: , Among them, is the attention in the multi-head context attention module, where Q, K, and V are the query matrix, key matrix, and value matrix respectively, is the scaling factor.

4. The facial paralysis assisted diagnosis method based on a generative model according to claim 1, wherein Step S3 includes: S301. Use a high-resolution camera device to capture facial images of facial paralysis patients, and select the abnormal side of the face for masking. S302. Input the masked facial paralysis patient image into the trained model and set a random seed to generate a predicted normal face image.

5. The method for auxiliary diagnosis of facial paralysis based on a generative model according to claim 1, wherein, Step S4 includes: S401. Calculate the structural similarity index SSIM to quantify the structural similarity between the generated normal face image and the original image. Here, SSIM measures the structural similarity of the image, and the calculation formula is as follows: , Where x and y are two images to be compared, and are the means of images x and y respectively, and are the variances of images x and y respectively, is the covariance of images x and y, and are small constants used to avoid a zero denominator, usually taking and , where L is the dynamic range of the image, and are small constants; S402. Calculate the absolute difference absdiff to quantify the pixel-level difference between the generated normal face image and the original image. Here, absdiff is used to measure the pixel-level difference between two images, and the calculation formula is as follows: , where x and y are the generated normal face image and the original image respectively, indicating the absolute difference between the two images at each pixel position, S403. Generate a difference map to highlight the abnormal regions between the normal face image and the original image. Here, convert the SSIM value into a difference map ssim_diff, use 1 - SSIM to represent the difference, normalize the calculation result of the absolute difference absdiff to the range [0, 1] for visualization, and use the normalized absolute difference map as the difference map diff_map: , , S404. Use the method of weighted sum to combine the SSIM difference map and the absdiff difference map to generate a comprehensive difference map, and the calculation formula is as follows: , Among them, combined_diff represents the comprehensive difference map, is a weight parameter, usually between [0, 1], used to balance the influence of SSIM and absdiff; S405. Combine the results of SSIM and absdiff, perform thresholding on the comprehensive difference map to generate a lesion mask. This mask highlights the regions in the original image that are significantly different from the generated normal face image. Select a threshold T. The selection of the threshold T can adopt a data-driven method or an adaptive method, and mark the part of the comprehensive difference map greater than T as the lesion region: , S406. Repeat steps S401 to S405 to generate multiple different masks and fuse them to generate a more accurate and comprehensive lesion mask. Vote on each mask and select the regions that most masks consider to be lesions to reduce misjudgment of a single mask: , where N is the number of generated masks, denotes the i-th generated mask; S407. Paste the generated lesion mask back onto the original image to visualize the lesion region and assist the doctor in diagnosis. Here, first read the original image y, and then apply the lesion mask final_mask to the original image y using the following formula to highlight the lesion region: , Among them, highlighted_image represents a visualization image, and color_mask is a color mask used to highlight the lesion area.

6. The facial paralysis assisted diagnosis method based on a generative model according to claim 1, characterized in that Step S5 includes: Converting the generated image and the real image into Laplacian pyramid representations, and starting from the low resolution, gradually increasing the resolution until the resolution of the original image is reached; At each scale, randomly extract local image patches from the generated image and the patient image. These image patches are windows of a fixed size, and a large number of image patches are extracted from each image to form two sets: the generated image patch set and the patient image patch set; Normalize the extracted image patches so that their mean is 0 and the standard deviation is 1, eliminating the scale differences between different image patches; For each scale, calculate the SWD between the generated image patch set and the real image patch set, where the SWD is calculated using the following formula: , Among them, is the set of the i-th generated image patch at the s-th scale, is the set of the j-th patient image patch at the s-th scale, is the inner product of the projection in the k-th random direction and the i-th generated image patch at the s-th scale, is the inner product of the projection in the k-th random direction and the j-th patient image patch at the s-th scale, are K random directions, where K is the number of random directions and N is the number of image patches, Aggregate the SWD values at different scales through weighted averaging to obtain a comprehensive multi-scale statistical similarity index. When aggregating the SWD values at different scales, a weight factor is introduced , and this weight factor is adjusted according to the symmetry difference of the image patches at each scale. The specific formula is as follows: , Among them, is the weight factor of the s-th scale, is the SWD value at the s-th scale, and the weight factor is adjusted according to the symmetry difference at each scale. The specific calculation formula is as follows: , Among them, is the symmetry difference at the s-th scale, and the calculation formula is as follows: , Among them, and are the left and right image patches of the $i$-th image patch at the $s$-th scale respectively, is the Euclidean norm.

7. A computer-readable storage medium, comprising computer programs / instructions, characterized in that, When the program / instructions are executed, they are used to implement the steps of the facial paralysis auxiliary diagnosis method based on the generative model according to any one of claims 1-6.

8. A facial paralysis assisted diagnosis system based on a generative model, comprising computer programs / instructions, characterized in that, When the program / instructions are executed, they are used to implement the steps of the facial paralysis auxiliary diagnosis method based on the generative model according to any one of claims 1-6.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the program / instructions are executed, they are used to implement the steps of the facial paralysis auxiliary diagnosis method based on the generative model according to any one of claims 1-6.

Citation Information

Patent Citations

  • Semantic-guided face image restoration method

    CN113112416A

  • Object surface damage detection method, storage medium and electronic equipment

    CN117975061A