Three-dimensional face data prediction method and device based on multi-scale generative adversarial network

By extracting model data from CBCT data and annotating it using a multi-scale generative adversarial network, a three-dimensional facial data prediction model is constructed. This solves the problem of inaccurate prediction of facial soft tissue changes in existing technologies and achieves highly accurate facial data prediction and treatment decision support.

CN121502849APending Publication Date: 2026-02-10PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511982315.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing deep learning technologies based on 3D point cloud or 3D mesh data cannot accurately predict changes in facial soft tissues after orthognathic surgery, resulting in insufficient accuracy in predicting facial appearance after treatment.

Method used

A multi-scale generative adversarial network was used to extract model data from CBCT data through a soft tissue isosurface extraction algorithm. Key soft tissue landmarks were labeled and data preprocessed to construct a three-dimensional facial data prediction model. The model was then trained using the multi-scale generative adversarial network to predict changes in facial soft tissue after correction.

Benefits of technology

It improves the accuracy of predicting changes in facial soft tissue after correction, provides intuitive and reliable facial prediction results, enhances the effectiveness and efficiency of doctor-patient communication, and supports accurate facial correction decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502849A_ABST
    Figure CN121502849A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of maxillofacial malocclusion correction, and discloses a three-dimensional face data prediction method and device based on a multi-scale generative adversarial network, and the method comprises the steps: extracting a model data set of a patient through a soft tissue equal face value extraction algorithm, and carrying out the processing of face key soft tissue mark points, the processing accuracy and reliability of the three-dimensional point cloud used as a model training basis can be improved, then according to the three-dimensional point cloud and key point labeling information, the multi-scale generative adversarial network is trained to obtain the prediction model, the training accuracy of the prediction model can be improved, and the prediction accuracy of the prediction model can be improved. Then, through the prediction model, according to the pre-correction three-dimensional point cloud of the malocclusion patient, the face point cloud form of the patient after correction is predicted, the prediction accuracy of the face soft tissue change possibly occurring after correction can be improved, and a visual and credible face prediction result can be provided; therefore, an accurate and reliable malocclusion correction decision can be made.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of maxillofacial malocclusion correction technology, and in particular to a three-dimensional facial data prediction method and device based on multi-scale generative adversarial networks. Background Technology

[0002] In addition to achieving a normal occlusal relationship between teeth and jawbone, the results of malocclusion correction based on orthodontic / orthognathic treatment techniques also include improving the facial appearance of the patient as much as possible. Predicting the facial data of the patient after treatment is a common pre-treatment data processing method.

[0003] In recent years, deep learning has made significant progress in the field of computer vision. Many studies have attempted to use deep learning technology to model and predict facial features. For example, deep learning technology based on 3D point cloud or 3D mesh data can model and predict facial features, which can reflect spatial structure information, making the facial features learned by the model more complete, thus facilitating facial data prediction.

[0004] However, practice has shown that traditional deep learning techniques based on 3D point clouds or 3D mesh data are only applicable to scenarios with significant facial morphological changes, such as orthognathic surgery, and cannot accurately predict potential post-treatment changes in facial soft tissues. Therefore, proposing a novel facial data prediction technique to improve the accuracy of predicting potential post-treatment changes in facial soft tissues is of paramount importance. Summary of the Invention

[0005] This invention provides a method and apparatus for predicting three-dimensional facial data based on multi-scale generative adversarial networks, which can improve the accuracy of predicting possible changes in facial soft tissue after correction.

[0006] The first aspect of this invention discloses a method for predicting three-dimensional facial data based on multi-scale generative adversarial networks, the method comprising: Obtain CBCT data sets for each of the multiple corrected patients, wherein each CBCT data set for the corrected patient includes the patient’s pre-correction CBCT data and the patient’s post-correction CBCT data. Based on the soft tissue isosurface extraction algorithm, the model data set of each corrected patient is extracted from the CBCT data set of each corrected patient. The model data set includes pre-correction model data and post-correction model data. For each of the corrected patients' model data sets, annotation processing is performed on key soft tissue landmarks on the face to obtain key point annotation information for the corrected patient; and based on the key point annotation information for the corrected patient, data preprocessing operation is performed on the model data set of the corrected patient to obtain the three-dimensional facial point cloud data set of the corrected patient. Based on the three-dimensional point cloud data set of all the patients who have undergone correction and the key point annotation information of the patients, the constructed multi-scale generative adversarial network is trained to obtain a three-dimensional facial data prediction model. Using the three-dimensional facial data prediction model, based on the obtained three-dimensional point cloud data of the patient before treatment, a facial data prediction operation is performed on the patient to obtain the prediction result of the three-dimensional facial point cloud data of the patient after treatment.

[0007] A second aspect of this invention discloses a three-dimensional facial data prediction device based on a multi-scale generative adversarial network, the device comprising: The acquisition module is used to acquire CBCT data sets for each of the multiple corrected patients. Each CBCT data set for the corrected patient includes the patient's pre-correction CBCT data and post-correction CBCT data. The extraction module is used to extract the model data set of each of the corrected patients from the CBCT data set of each corrected patient based on the soft tissue isosurface extraction algorithm. The model data set includes pre-correction model data and post-correction model data. The processing module is used to perform annotation processing on the model data group of each of the corrected patients for key soft tissue landmarks on the face to obtain key point annotation information about the corrected patient; and to perform data preprocessing operation on the model data group of the corrected patient based on the key point annotation information about the corrected patient to obtain the three-dimensional facial point cloud data group of the corrected patient. The training module is used to train the constructed multi-scale generative adversarial network based on the three-dimensional point cloud data set of all the corrected patients and the key point annotation information of the corrected patients, so as to obtain a three-dimensional facial data prediction model. The prediction module is used to perform facial data prediction operations on the patient to be treated based on the obtained three-dimensional point cloud data of the patient before treatment using the three-dimensional facial data prediction model, and to obtain the prediction result of the three-dimensional facial point cloud data of the patient after treatment.

[0008] As an optional implementation, in a second aspect of the present invention, the three-dimensional point cloud data set includes pre-treatment three-dimensional point cloud data and post-treatment three-dimensional point cloud data. Furthermore, the training module trains the constructed multi-scale generative adversarial network based on the three-dimensional point cloud data set of all the treated patients and the key point annotation information of the treated patients to obtain the three-dimensional facial data prediction model. Specifically, this includes: Multiple downsampling operations are performed on the pre-treatment three-dimensional point cloud data of each of the corrected patients to obtain the pre-treatment three-dimensional point cloud data after each downsampling. The pre-treatment three-dimensional point cloud data of each of the corrected patients and the pre-treatment three-dimensional point cloud data after each downsampling are used as point cloud input data at the corresponding resolution. Based on a preset division ratio, all point cloud input data of all the corrected patients and the key point annotation information corresponding to the corrected patients are divided to obtain a division result, which includes a model training set, a model validation set and a model test set. The point cloud input data of all the corrected patients in the model training set and the key point annotation information of the corrected patients are input into the constructed multi-scale generative adversarial network for training, and the model training result of this training cycle is obtained. The model training result includes the predicted three-dimensional point cloud data of each corrected patient in the model training set at different resolutions. Based on the model training results of this training cycle and the post-treatment 3D point cloud data of all corrected patients in the model training set, the target loss function of the multi-scale generative adversarial network with respect to this training cycle is calculated. Then, through gradient optimization of the target loss function, the network parameters of the multi-scale generative adversarial network trained with respect to this training cycle are updated to obtain the target model corresponding to this training cycle. Based on the model validation set, the target model corresponding to this training cycle is evaluated to obtain the evaluation result of the target model corresponding to this training cycle. Based on the evaluation results of multiple target models trained based on a preset number of training cycles, the target model to be tested is determined from all the target models. Based on the model test set, the target model to be tested is tested to obtain test results. When the test results meet the preset target test results after training is completed, the target model to be tested is determined as a three-dimensional facial data prediction model.

[0009] As an optional implementation, in a second aspect of the present invention, the training module inputs all point cloud input data of all treated patients in the model training set, along with key point annotation information about the treated patients, into a pre-constructed multi-scale generative adversarial network for training, and obtains the model training results for this training cycle in the following specific ways: For each corrected patient in the model training set, all point cloud input data of the corrected patient and key point annotation information about the corrected patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, so as to obtain the key feature vector of the corrected patient in this training cycle. The key feature vector of the corrected patient in this training cycle is input into the multi-scale decoder of the multi-scale generative adversarial network to generate the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle.

[0010] As an optional implementation, in a second aspect of the present invention, the training module, for each corrected patient in the model training set, inputs all point cloud input data of the corrected patient and key point annotation information about the corrected patient into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, and obtains the key feature vector of the corrected patient in this training cycle in the following specific ways: For each corrected patient in the model training set, the key point annotation information of the corrected patient is input into the multi-scale encoder of the constructed multi-scale generative adversarial network, so as to extract the key point feature information corresponding to the key point annotation information through the feature extractor in the multi-scale encoder. The point cloud input data at each resolution of all point cloud input data of the corrected patient is input into the multi-scale encoder so that the feature extractor in the multi-scale encoder can extract the point cloud feature information corresponding to each point cloud input data. The attention module in the multi-scale encoder determines the potential feature vector corresponding to each type of point cloud input data by establishing the information association between key points and facial point clouds based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data. The multi-scale encoder is used to concatenate and fuse the potential feature vectors corresponding to the point cloud input data at all the resolutions of the corrected patient to obtain the key feature vector of the corrected patient in this training cycle.

[0011] As an optional implementation, in a second aspect of the present invention, the training module calculates the target loss function of the multi-scale generative adversarial network with respect to the current training cycle based on the model training results of the current training cycle and the post-treatment 3D point cloud data of all treated patients in the model training set. Specifically, this includes: The predicted 3D point cloud data at each resolution included in the model training results of this training cycle are taken as the predicted point cloud data set corresponding to that resolution, and the post-correction 3D point cloud data of all corrected patients in the model training set are taken as the real point cloud data set. The distance loss parameter between the predicted point cloud data set corresponding to that resolution and the real point cloud data set is calculated. Based on the distance loss parameters corresponding to all the resolutions in this training cycle and the preset weight parameters corresponding to each resolution, calculate the distance loss function corresponding to this training cycle; Based on the predicted point cloud data set corresponding to all the resolutions and the real point cloud data set, calculate the adversarial loss function of the multi-scale generative adversarial network for this training cycle; Based on the distance loss function corresponding to this training cycle and the adversarial loss function of the multi-scale generative adversarial network for this training cycle, calculate the target loss function of the multi-scale generative adversarial network for this training cycle.

[0012] As an optional implementation, in a second aspect of the invention, the processing module performs annotation processing on key facial soft tissue landmarks for each of the model data sets of the corrected patients to obtain key point annotation information about the corrected patients. Specifically, this includes: For each type of model data contained in the model data set of each of the corrected patients, the annotation results of the key facial soft tissue landmarks of each of the multiple annotation objects for the model data are obtained, and the annotation results include the initial annotation position of each of the multiple key annotation points. For each key annotation point, the annotation position difference between each annotation object for the key annotation point is calculated based on the initial annotation positions of all annotation objects for that key annotation point. The target annotation position of the key annotation point is determined based on the difference in annotation position between all the annotated objects for the key annotation point; Based on the target annotation location of all key annotation points in all model data contained in the model data group of each corrected patient, the key point annotation information of the corrected patient is determined.

[0013] As an optional implementation, in a second aspect of the invention, the apparatus further includes: The analysis module is used to analyze each model data in the model data group of each of the treated patients, based on the model data and preset facial structure information, to obtain the degree of facial structural integrity that the model data can express; and to analyze the scan quality corresponding to the model data. The filtering module is used to filter the model data sets of all the corrected patients based on the degree of facial structural integrity that can be expressed by all model data in the model data sets of all the corrected patients, as well as the scanning quality corresponding to the model data, to obtain target model data sets that meet preset data quality conditions; and to trigger the processing module to perform the operation of annotating key soft tissue landmarks of the face in the model data sets of each corrected patient to obtain key point annotation information about the corrected patient.

[0014] A third aspect of this invention discloses another three-dimensional facial data prediction device based on a multi-scale generative adversarial network, the device comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps in the three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to any of the first aspects of the present invention.

[0015] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the three-dimensional facial data prediction method based on multi-scale generative adversarial networks as described in any of the first aspects of the present invention.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention employs a soft tissue isosurface extraction algorithm to extract the model data set for each corrected patient from their acquired CBCT data set. It then performs annotation processing on key soft tissue landmarks on the face to obtain key point annotation information for that patient. Furthermore, it preprocesses the model data set to obtain 3D facial point cloud data. This improves the accuracy and reliability of the 3D facial point cloud data used for model training. Subsequently, based on the accurately processed 3D point cloud data set and the corresponding key point annotation information, a pre-constructed multi-scale generative adversarial network is trained to obtain 3D facial data predictions. The model can improve the training accuracy of the 3D facial data prediction model. Then, through the 3D facial data prediction model obtained through accurate training, the facial data prediction operation is performed on the patient before the treatment based on the obtained 3D point cloud data of the patient before the treatment. The prediction results of the 3D facial point cloud data of the patient after the treatment are obtained. This can improve the prediction accuracy of possible changes in facial soft tissue after the treatment and help to provide intuitive and reliable facial prediction results. This is conducive to making accurate and reliable facial correction decisions based on accurate and reliable facial prediction results, and also helps to improve the effectiveness and efficiency of communication between doctors and patients. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating a three-dimensional facial data prediction method based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Figure 2 This is a flowchart illustrating another method for predicting 3D facial data based on multi-scale generative adversarial networks disclosed in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a multi-scale generative adversarial network disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a three-dimensional facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of another three-dimensional facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Figure 6 This is a schematic diagram of the structure of another three-dimensional facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.

[0021] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0022] This invention discloses a method and apparatus for predicting 3D facial data based on multi-scale generative adversarial networks (GANs). It can extract the model data set of each corrected patient from the acquired CBCT data set using a soft tissue face value extraction algorithm. It then performs annotation processing on key soft tissue landmarks to obtain key point annotation information for that patient. Furthermore, it performs data preprocessing on the model data set to obtain 3D facial point cloud data. This improves the accuracy and reliability of the 3D facial point cloud data used for model training. Subsequently, based on the accurately processed 3D point cloud data set and the corresponding key point annotation information, it predicts the constructed multi-scale generative adversarial network. Training the network yields a 3D facial data prediction model, improving its training accuracy. This accurately trained model, based on pre-treatment 3D point cloud data of the patient, performs facial data prediction, resulting in a post-treatment prediction of the patient's 3D facial point cloud data. This improves the accuracy of predicting potential post-treatment changes in facial soft tissue and provides intuitive and reliable facial predictions. This allows for accurate and reliable facial correction decisions and enhances communication effectiveness and efficiency between doctors and patients. These details are explained in detail below.

[0023] Example 1 Please see Figure 1 , Figure 1 This is a flowchart illustrating a three-dimensional facial data prediction method based on multi-scale generative adversarial networks disclosed in an embodiment of the present invention. Figure 1 The described 3D facial data prediction method based on multi-scale generative adversarial networks can be applied to a 3D facial data prediction device based on multi-scale generative adversarial networks. This device may include one of a prediction device, a prediction system (including a local system or a cloud system), and a prediction server (including a local server or a cloud server). This embodiment of the invention does not limit the scope of the method. Figure 1 As shown, the 3D facial data prediction method based on multi-scale generative adversarial networks can include the following operations: 101. Obtain CBCT data sets for each of the multiple corrected patients.

[0024] In this embodiment of the invention, each CBCT data set of a treated patient includes the patient's pre-treatment CBCT data and post-treatment CBCT data. The CBCT data can be obtained using cone-beam computed tomography (CBCT), an oral and maxillofacial imaging technique. Furthermore, the CBCT data can be stored in DICOM format.

[0025] 102. Based on the soft tissue isosurface extraction algorithm, extract the model data set of each corrected patient from the CBCT data set of each corrected patient.

[0026] In this embodiment of the invention, the model data set includes pre-treatment model data and post-treatment model data. Specifically, for each type of CBCT data (i.e., pre-treatment or post-treatment CBCT data) in the CBCT data set of each treated patient, a soft tissue isosurface extraction algorithm is used to extract the soft tissue isosurfaces from the CBCT data according to a preset isosurface threshold. The extracted soft tissue isosurfaces are then meshed to obtain a meshed soft tissue model. The corresponding model data for the CBCT data is then extracted from the meshed soft tissue model (e.g., the meshed soft tissue model is saved as an STL file as STL model data). Optionally, the soft tissue isosurface extraction algorithm can be the FlyingEdges3D algorithm, the Marching Cubes algorithm, or any other algorithm capable of isosurface extraction. Optionally, the isosurface threshold can be a grayscale value (e.g., -750), a gradient magnitude (e.g., edge grayscale change rate), or any other threshold that can serve as the isosurface boundary; this embodiment of the invention does not limit the threshold.

[0027] 103. For each corrected patient's model data set, perform annotation processing on key soft tissue landmarks of the face to obtain key point annotation information about the corrected patient.

[0028] In this embodiment of the invention, specifically, for each treated patient, the annotation position of each key annotation point in all key annotation points of each model data contained in the model data set of the treated patient by a first annotation object at a first preset qualification level is obtained. Based on the annotation positions of all key annotation points of the treated patient by the first annotation object, the key point annotation information of the treated patient is determined. Alternatively, the annotation position of each key annotation point in all key annotation points of each model data contained in the model data set of the treated patient by multiple second annotation objects at a second preset qualification level is obtained. Based on the annotation positions of all key annotation points of each type of key annotation point by all second annotation objects for the treated patient, the target annotation position of that key annotation point is obtained by fusing them. Then, based on the target annotation positions of all key annotation points corresponding to the treated patient, the key point annotation information of the treated patient is determined. The first preset qualification level is higher than the second preset qualification level. This allows for flexible selection of annotation objects at different qualification levels for key point annotation, which is beneficial for improving the flexibility and accuracy of annotation of key soft tissue landmarks on the patient's face.

[0029] 104. Based on the key point annotation information of the treated patient, perform data preprocessing operations on the model data set of the treated patient to obtain the three-dimensional facial point cloud data set of the treated patient.

[0030] In this embodiment of the invention, the three-dimensional point cloud data set includes pre-treatment and post-treatment three-dimensional point cloud data. Specifically, the data preprocessing operation can be divided into four steps in sequence: registration, cropping, sampling, and normalization. The registration step uses labeled key points to rigidly register the pre- and post-treatment STL data, unifying them to the same reference coordinate system to ensure consistency and comparability of morphological changes during modeling. The cropping step effectively crops the facial region based on key point information, focusing on the analysis area and removing irrelevant parts to reduce interference. The sampling step uniformly samples the cropped STL model to obtain pre- and post-treatment three-dimensional point clouds. The normalization step normalizes the obtained three-dimensional point clouds to avoid calculation errors caused by scale differences.

[0031] 105. Based on the three-dimensional point cloud data set of all treated patients and the key point annotation information of the treated patients, the constructed multi-scale generative adversarial network is trained to obtain a three-dimensional facial data prediction model.

[0032] 106. Using a three-dimensional facial data prediction model, based on the three-dimensional point cloud data of the patient to be treated before treatment, perform facial data prediction operations on the patient to be treated to obtain prediction results of the three-dimensional facial point cloud data of the patient to be treated after treatment.

[0033] In this embodiment of the invention, the patient to be treated can be a malocclusion patient, and the aforementioned treated patient can be a malocclusion patient who has already undergone orthodontic treatment.

[0034] It is evident that implementation Figure 1 The described 3D facial data prediction method based on multi-scale generative adversarial networks can extract the model data set of each corrected patient from the acquired CBCT data set using a soft tissue isosurface extraction algorithm. It then performs annotation processing on key soft tissue landmarks to obtain key point annotation information for that patient. Furthermore, it performs data preprocessing on the model data set to obtain 3D facial point cloud data, improving the accuracy and reliability of the 3D facial point cloud data used for model training. Subsequently, based on the accurately processed 3D point cloud data set and the corresponding key point annotation information, the constructed scale-based generative adversarial network is trained. Training to obtain a 3D facial data prediction model can improve the training accuracy of the 3D facial data prediction model. Then, using the accurately trained 3D facial data prediction model, based on the obtained 3D point cloud data of the patient before treatment, facial data prediction operations are performed on the patient to obtain the prediction results of the 3D facial point cloud data of the patient after treatment. This can improve the prediction accuracy of possible changes in facial soft tissue after treatment and help provide intuitive and reliable facial prediction results. This is conducive to making accurate and reliable facial correction decisions based on accurate and reliable facial prediction results, and also helps to improve the effectiveness and efficiency of communication between doctors and patients.

[0035] In an optional embodiment, step 103 above involves annotating key facial soft tissue landmarks in the model data set of each treated patient to obtain key point annotation information for that treated patient, including: For each type of model data contained in the model data set of each corrected patient, obtain the annotation results of the key facial soft tissue landmarks of each of the multiple annotation objects for the model data. The annotation results include the initial annotation position of each key annotation point among the multiple key annotation points. For each key annotation point, calculate the annotation position difference between each annotation object for that key annotation point based on the initial annotation positions of all annotation objects for that key annotation point; The target annotation position of the key annotation point is determined based on the difference in annotation positions between all annotated objects for that key annotation point; Based on the target annotation location of all key annotation points in all model data contained in the model data set of each corrected patient, determine the key point annotation information of that corrected patient.

[0036] In this embodiment of the invention, specifically, determining the target annotation position of the key annotation point based on the difference in annotation positions among all annotated objects for that key annotation point includes: Determine whether the difference in the annotation position of each annotation object for the key annotation point is less than or equal to the preset annotation position difference; When it is determined that the difference in the annotation position of each annotation object relative to the key annotation point is less than or equal to the preset annotation position difference, the average value of the initial annotation positions of all annotation objects relative to the key annotation point is calculated and used as the target annotation position of the key annotation point. When it is determined that there is a case where the difference in the annotation position of the key annotation point among all annotated objects is greater than the preset annotation position difference, the position correction result of the target annotation object for the key annotation point is obtained, and the target annotation position of the key annotation point is obtained.

[0037] Optionally, each labeled object (e.g., an orthodontic specialist) has a corresponding qualification level (e.g., seniority), with the target labeled object corresponding to the highest data level. For example: Two orthodontic specialists (with 5+ years of orthodontic experience) manually locate soft tissue isosurfaces using STL. When the difference in labeled positions for the same key label point is less than or equal to a preset difference (e.g., 2mm), the average of all initial labeled positions for that key label point is taken as the target labeled position. When the difference in labeled positions for that key label point is greater than the preset difference, a senior physician (10+ years of orthodontic experience) checks and corrects the error.

[0038] For example, as shown in Table 1, Table 1 is a table of correspondence between the names and abbreviations of key annotation points disclosed in an embodiment of the present invention:

[0039] As can be seen, this optional embodiment can obtain the annotation results of key facial soft tissue landmarks for each model data of each corrected patient by different annotation objects, such as the initial annotation position, and determine the target annotation position of the key annotation point based on the error of the initial annotation position of all annotation objects for the same key annotation point. This can improve the accuracy and reliability of key annotation point annotation. Then, based on the target annotation position of all key annotation points, the key point annotation information of the corresponding patient can be determined, which can improve the accuracy and reliability of key point annotation information determination. This is conducive to improving the accuracy of subsequent data preprocessing of model data based on the accurately determined key point annotation information.

[0040] In another alternative embodiment, the method may further include: For each model data in the model data set of each treated patient, the model data is analyzed based on the model data and the preset facial structure information to obtain the degree of facial structural integrity that the model data can express; and to analyze the scan quality corresponding to the model data. Based on the degree of facial structural integrity that can be expressed by all model data in the model data set of all treated patients, and the corresponding scan quality of the model data, the model data sets of all treated patients are screened to obtain target model data sets that meet the preset data quality conditions; and the operation of annotating key soft tissue landmarks of the face in step 103 is triggered to obtain key point annotation information about the treated patient.

[0041] In this embodiment of the invention, optionally, the data quality condition is used to represent the condition that the degree of facial structural integrity that the corresponding model data can express is greater than or equal to a preset integrity level and the corresponding scan quality is greater than or equal to a preset scan quality. For example: Suppose data from 525 patients before and after treatment (i.e., before and after correction) are collected. Through quality screening, samples with incomplete facial structures or poor scan quality are removed, and finally 511 high-quality data are retained as the basis for model training.

[0042] As can be seen, this optional embodiment can analyze the model data for each patient based on each model data and preset facial structure information to obtain the degree of facial structural integrity that the model data can express, and analyze the scan quality corresponding to the model data. Then, based on the degree of facial structural integrity that all model data in the model data group of all treated patients can express, and the scan quality corresponding to the model data, the model data group of all treated patients is screened to obtain the target model data group that meets the preset data quality conditions. This can improve the accuracy and reliability of screening patient model data, and is conducive to further improving the accuracy and credibility of subsequent model training.

[0043] Example 2 Please see Figure 2 , Figure 2 This is a flowchart illustrating a three-dimensional facial data prediction method based on multi-scale generative adversarial networks disclosed in an embodiment of the present invention. Figure 2 The described 3D facial data prediction method based on multi-scale generative adversarial networks can be applied to a 3D facial data prediction device based on multi-scale generative adversarial networks. This device may include one of a prediction device, a prediction system (including a local system or a cloud system), and a prediction server (including a local server or a cloud server). This embodiment of the invention does not limit the scope of the method. Figure 2As shown, the 3D facial data prediction method based on multi-scale generative adversarial networks can include the following operations: 201. Obtain CBCT data sets for each of the multiple corrected patients.

[0044] 202. Based on the soft tissue isosurface extraction algorithm, extract the model data set of each corrected patient from the CBCT data set of each corrected patient.

[0045] 203. For each corrected patient's model data set, perform annotation processing on key soft tissue landmarks of the face to obtain key point annotation information about the corrected patient.

[0046] 204. Based on the key point annotation information of the treated patient, perform data preprocessing operations on the model data set of the treated patient to obtain the three-dimensional facial point cloud data set of the treated patient.

[0047] 205. Perform multiple downsampling operations on the pre-treatment 3D point cloud data of each treated patient to obtain the pre-treatment 3D point cloud data after each downsampling.

[0048] In this embodiment of the invention, the pre-treatment 3D point cloud data of each treated patient and the pre-treatment 3D point cloud data after each downsampling are used as point cloud input data at the corresponding resolution. Specifically, for each treated patient, the pre-treatment 3D point cloud data of that patient is downsampled to obtain the pre-treatment 3D point cloud data after this downsampling, and the pre-treatment 3D point cloud data after this downsampling is downsampled again until a preset number of downsampling times is reached. For example, the pre-treatment 3D point cloud data is denoted as pre-treatment point cloud 1 (including N points), and pre-treatment point cloud 1 needs to be downsampled twice. Specifically, pre-treatment point cloud 1 is downsampled once to obtain pre-treatment point cloud 2 (including N / 2 points), and pre-treatment point cloud 2 is downsampled again to obtain pre-treatment point cloud 3 (including N / 4 points). The number of points contained in different point clouds can be used to represent the corresponding resolution.

[0049] 206. Based on the preset division ratio, divide all point cloud input data of all treated patients and the key point annotation information corresponding to the treated patients to obtain the division result.

[0050] In this embodiment of the invention, the partitioning result includes a model training set, a model validation set, and a model test set, for example, partitioning them in a ratio of 8:1:1 to obtain the model training set, model validation set, and model test set.

[0051] 207. Input all point cloud input data of all treated patients in the model training set and the key point annotation information of the treated patients into the constructed multi-scale generative adversarial network for training, and obtain the model training results of this training cycle.

[0052] In this embodiment of the invention, the multi-scale generative adversarial network may include a generator and a discriminator, wherein the generator may include a multi-scale encoder and a multi-scale decoder. The model training results include predicted 3D point cloud data at different resolutions for each corrected patient in the model training set. The entire network is trained using supervised learning, with the optimization objective being to minimize the multi-scale difference between the predicted point cloud and the true post-correction point cloud. During training, the Adam optimizer is used to alternately and jointly optimize the generator and discriminator.

[0053] 208. Based on the model training results of this training cycle and the post-treatment 3D point cloud data of all corrected patients in the model training set, calculate the target loss function of the multi-scale generative adversarial network with respect to this training cycle.

[0054] In this embodiment of the invention, the target loss function can be composed of the distance loss function between the predicted 3D point cloud data at all resolutions and the corrected 3D point cloud data in the model training set, as well as the adversarial loss function of the multi-scale adversarial generative network.

[0055] 209. By optimizing the gradient of the target loss function, update the network parameters of the multi-scale generative adversarial network with respect to the training of this training cycle, and obtain the target model corresponding to this training cycle.

[0056] 210. Based on the model validation set, evaluate the target model corresponding to this training cycle, obtain the evaluation result of the target model corresponding to this training cycle, and determine the target model to be tested from all target models based on the evaluation results of multiple target models trained for a preset number of training cycles.

[0057] In this embodiment of the invention, by setting a maximum training period, each training period is evaluated using a model validation set, and finally the model with the smallest distance error in the evaluation results of the model validation set is selected for testing.

[0058] 211. Based on the model test set, test the target model to be tested, obtain the test results, and when the test results meet the preset target test results after training is completed, determine the target model to be tested as the three-dimensional facial data prediction model.

[0059] In this embodiment of the invention, during the model testing phase, the three-dimensional point cloud and key point information of the model test set are input, and the corresponding post-treatment facial soft tissue point cloud prediction results are output, thereby realizing the visual auxiliary prediction of the orthodontic treatment effect.

[0060] 212. Using a three-dimensional facial data prediction model, based on the obtained three-dimensional point cloud data of the patient to be treated before treatment, perform facial data prediction operation on the patient to be treated to obtain the prediction results of the three-dimensional facial point cloud data of the patient to be treated after treatment.

[0061] In this embodiment of the invention, for other descriptions of steps 201-204 and step 212, please refer to the detailed description of steps 101-104 and step 106 in Embodiment 1. These descriptions will not be repeated in this embodiment of the invention.

[0062] It is evident that implementation Figure 2The described 3D facial data prediction method based on multi-scale generative adversarial networks (GANs) can extract the model data set of each corrected patient from the acquired CBCT data set using a soft tissue isosurface extraction algorithm. It then performs annotation processing on key soft tissue landmarks to obtain key point annotation information for that patient. Furthermore, it performs data preprocessing on the model data set to obtain 3D facial point cloud data, improving the accuracy and reliability of the 3D facial point cloud data used for model training. Subsequently, based on the accurately processed 3D point cloud data set and the corresponding key point annotation information, the constructed multi-scale GAN is... Training to obtain a 3D facial data prediction model can improve the training accuracy of the 3D facial data prediction model. Then, using the accurately trained 3D facial data prediction model, based on the obtained 3D point cloud data of the patient before treatment, facial data prediction operations are performed on the patient to obtain the prediction results of the 3D facial point cloud data of the patient after treatment. This can improve the prediction accuracy of possible changes in facial soft tissues after treatment and help provide intuitive and reliable facial prediction results. This is conducive to making accurate and reliable facial correction decisions based on accurate and reliable facial prediction results, and also helps to improve the effectiveness and efficiency of communication between doctors and patients. Furthermore, it can divide all patients' point cloud input data and key point annotation information into training, validation, and test sets based on a preset division ratio. The point cloud input data and key point annotation information of the training set are then input into a pre-constructed multi-scale generative adversarial network for training, obtaining the training results for this training cycle, such as predicted point clouds at different resolutions. The target loss function is calculated based on the predicted and real point clouds of this training cycle, and the network parameters are updated through gradient optimization of the target loss function to obtain the target model corresponding to this training. This can improve the training accuracy and reliability of the target model in each training cycle. Subsequently, the target model of each training cycle is evaluated based on the validation set to obtain the model evaluation results for each training cycle, thereby determining the target model to be tested. The evaluation of the validation set can improve the accuracy and reliability of the target model to be tested. Finally, the target model to be tested is tested based on the test set. When the test results meet the preset target test results, the training is considered complete. The testing of the test set can improve the accuracy and reliability of the 3D facial data prediction model obtained through training.

[0063] In an optional embodiment, step 207 above, which involves inputting all point cloud input data of all treated patients in the model training set, along with key point annotation information about those patients, into the constructed multi-scale generative adversarial network for training, yields the model training results for this training cycle, including: For each corrected patient in the model training set, all point cloud input data of the corrected patient and key point annotation information about the corrected patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, so as to obtain the key feature vector of the corrected patient in this training cycle. The key feature vectors of the corrected patient in this training cycle are input into the multi-scale decoder of the multi-scale generative adversarial network to generate the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle.

[0064] In this embodiment of the invention, optionally, the multi-scale encoder may include one or more feature extractors. Further, when the number of feature extractors is one, all point cloud input data and keypoint annotation information for each patient are extracted using this feature extractor. When the number of feature extractors is greater than one, each type of data can be assigned to a separate feature extractor, thus ensuring that each type of point cloud input data and keypoint annotation information is input to its corresponding feature extractor. Further optionally, the multi-scale encoder may also include an attention module (e.g., one attention module for each feature extractor). By introducing an attention module, the feature extraction process for point clouds at different resolutions can be guided using keypoint features, thereby capturing richer structural detail information.

[0065] In this embodiment of the invention, optionally, the multi-scale decoder may consist of convolutional layers and fully connected layers, used to generate predicted 3D facial point cloud data of corresponding resolution.

[0066] As can be seen, this optional embodiment can input all point cloud input data of each corrected patient in the model training set, as well as the key point annotation information of the corrected patient, into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, to obtain the key feature vector of the corrected patient in this training cycle. The multi-scale encoder can improve the extraction accuracy of key feature vectors. Subsequently, the key feature vector of the corrected patient in this training cycle is input into the multi-scale decoder of the multi-scale generative adversarial network for generation, to obtain the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle. The multi-scale decoder can improve the prediction accuracy and reliability of the predicted point cloud data of the patient at different resolutions.

[0067] In this optional embodiment, as an optional implementation, for each corrected patient in the model training set, all point cloud input data of the corrected patient and key point annotation information about the corrected patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, to obtain the key feature vector of the corrected patient in this training cycle, including: For each corrected patient in the model training set, the key point annotation information of the corrected patient is input into the multi-scale encoder of the constructed multi-scale generative adversarial network, so that the key point feature information corresponding to the key point annotation information can be extracted by the feature extractor in the multi-scale encoder. The point cloud input data at each resolution of all point cloud input data of the corrected patient is input into the multi-scale encoder so that the feature extractor in the multi-scale encoder can extract the point cloud feature information corresponding to each point cloud input data. By using the attention module in the multi-scale encoder, based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data, the potential feature vector corresponding to each type of point cloud input data is determined by establishing the information association between key points and facial point clouds. By using a multi-scale encoder, the latent feature vectors corresponding to the point cloud input data of the corrected patient at all resolutions are spliced ​​and fused to obtain the key feature vector of the corrected patient in this training cycle.

[0068] In this embodiment of the invention, the input of the multi-scale generative adversarial network consists of two parts. The first part is the 3D point cloud data before correction and the point cloud after multiple (e.g., 2) downsampling, i.e., 3D point cloud data before correction at multiple resolutions. The second part is the key annotation points of the 3D point cloud data before correction. Specifically, the multi-scale generative adversarial network first extracts and fuses features of point clouds at multiple resolutions through a multi-scale encoder. Specifically, the feature extractor extracts key point features and geometric features (point cloud feature information) of point clouds at multiple resolutions. The feature extractor can adopt an improved PointNet structure to encode each point into a 3D feature vector. By performing max pooling on the outputs of different depth layers and concatenating them along the channel dimension, a feature encoding that fuses rich low-level and high-level semantic information is finally obtained. Subsequently, an attention module is introduced to guide the feature extraction process of point clouds at each resolution using key point features, thereby capturing richer structural detail information. Specifically, the attention module uses keypoint features as queries and point cloud features as keys and values, establishing informational connections between keypoints and the overall facial point cloud through point-to-point attention scores. The core idea of ​​this mechanism is to allow keypoints to actively focus and extract the most critical information regions for structural representation from the point cloud. Next, the network concatenates and fuses the latent feature vectors of the pre-correction point cloud at multiple resolutions to generate the final feature representation (i.e., key feature vectors), which is then input into the multi-scale decoder. For example, as shown... Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Figure 3Taking the training of three-dimensional point cloud data with three resolutions by performing two downsampling operations as an example, point cloud 1 before treatment, point cloud 2 before treatment, and point cloud 3 before treatment refer to the three-dimensional point cloud data before treatment in the training set, and point cloud 1 after treatment, point cloud 2 after treatment, and point cloud 3 after treatment refer to the corresponding three-dimensional point cloud data after treatment in the training set. The discriminator consists of three fully connected layers and a feature extractor. The result of the predicted point cloud is determined by the fully connected layers (the result includes true or false). Specifically, the pre-treatment point cloud 1, pre-treatment point cloud 2, and pre-treatment point cloud 3 are processed by a multi-scale encoder to obtain the final feature vector (i.e., key feature vector), which is then input into the fully connected layer of the multi-scale decoder. Subsequently, the predicted point cloud 1 corresponding to the pre-treatment point cloud 1 is output through the feature layer 1, the fully connected layer, and the convolutional layer. The predicted point cloud 2 corresponding to the pre-treatment point cloud 2 is output through the feature layer 1, the fully connected layer, the feature layer 2, the fully connected layer, and the convolutional layer. The predicted point cloud 3 corresponding to the pre-treatment point cloud 3 is output through the feature layer 1, the fully connected layer, the feature layer 2, the fully connected layer, the feature layer 3, the fully connected layer, and the convolutional layer. The predicted point cloud 3 is then distinguished from the pre-treatment point cloud 3, the predicted point cloud 2 from the post-treatment point cloud 2, and the predicted point cloud 1 from the post-treatment point cloud 1 by a discriminator.

[0069] As can be seen, this optional implementation can input key point annotation information for each patient and point cloud input data at each resolution into the constructed multi-scale encoder of the multi-scale generative adversarial network. The feature extractor in the multi-scale encoder can extract key point feature information corresponding to the key point annotation information and point cloud feature information corresponding to each type of point cloud input data, thereby improving the accuracy of key feature information and point cloud feature information extraction. Subsequently, through the attention module in the multi-scale encoder, based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data, the potential feature vector corresponding to each type of point cloud input data is determined by establishing the information association between key points and facial point clouds, thereby improving the accuracy of determining the potential feature vector corresponding to each resolution of point cloud input data. Finally, through the multi-scale encoder, the potential feature vectors corresponding to the point cloud input data at all resolutions of the corrected patient are concatenated and fused to obtain the key feature vector of the corrected patient in this training cycle, thereby improving the accuracy and reliability of key feature vector extraction.

[0070] In another optional embodiment, step 208 above, based on the model training results of this training cycle and the post-treatment 3D point cloud data of all treated patients in the model training set, calculates the target loss function of the multi-scale generative adversarial network with respect to this training cycle, including: The predicted 3D point cloud data at each resolution included in the model training results of this training cycle are used as the predicted point cloud data set corresponding to that resolution, and the post-correction 3D point cloud data of all corrected patients in the model training set are used as the real point cloud data set. The distance loss parameter between the predicted point cloud data set and the real point cloud data set corresponding to that resolution is calculated. Based on the distance loss parameters corresponding to all resolutions in this training cycle and the preset weight parameters corresponding to each resolution, calculate the distance loss function corresponding to this training cycle. Based on the predicted point cloud data set and the real point cloud data set corresponding to all resolutions, calculate the adversarial loss function of the multi-scale generative adversarial network for this training cycle; Based on the distance loss function corresponding to this training cycle and the adversarial loss function of the multi-scale generative adversarial network for this training cycle, calculate the target loss function of the multi-scale generative adversarial network for this training cycle.

[0071] In this embodiment of the invention, the formula for calculating the distance loss parameter between the predicted point cloud dataset and the actual point cloud dataset for each resolution is as follows: ; in, Represents the predicted point cloud dataset. Represents a set of real point cloud data. This represents the key feature vector corresponding to the predicted 3D point cloud data contained in the predicted point cloud dataset. This represents the key feature vector corresponding to the corrected 3D point cloud data contained in the real point cloud dataset. This represents the distance loss parameter.

[0072] When a multi-scale decoder predicts three types of 3D point cloud data at different resolutions, the distance loss function (i.e., CD loss) between the predicted point cloud data set and the real point cloud data set corresponding to all resolutions consists of three terms, calculated as follows: ; in, Represents the distance loss function. , and These represent the distance loss parameters at different resolutions. , and These are sets of predicted point cloud data at different resolutions. , and These are sets of real point cloud data at different resolutions. This represents hyperparameters (used for weighting). Specifically, when the resolution is divided into high, medium, and low resolutions, This represents the distance loss parameter at high resolution. This represents the distance loss parameter at medium resolution. This represents the distance loss parameter at low resolution.

[0073] The formula for calculating the adversarial loss function of a multi-scale generative adversarial network is as follows: ; in, This represents the adversarial loss function, where F stands for generator (multi-scale encoder + multi-scale decoder) and D stands for discriminator. Represents the first element in the predicted point cloud dataset. A key feature vector corresponding to the predicted 3D point cloud data. Represents the first in the real point cloud dataset The key feature vectors corresponding to the three-dimensional point cloud data after correction.

[0074] The formula for calculating the target loss function is as follows: ; in, Represents the target loss function. and This is a hyperparameter used for weighting.

[0075] As can be seen, this optional embodiment can use the predicted 3D point cloud data at each resolution included in the model training results of this training cycle as the predicted point cloud data set corresponding to that resolution, and use the post-correction 3D point cloud data of all corrected patients in the model training set as the real point cloud data set. It calculates the distance loss parameter between the predicted point cloud data set and the real point cloud data set corresponding to that resolution, and calculates the distance loss function corresponding to this training cycle based on the distance loss parameters corresponding to all resolutions in this training cycle and the preset weight parameters corresponding to each resolution. This improves the accuracy and reliability of the distance loss function calculation for this training cycle. Furthermore, based on the predicted point cloud data sets and real point cloud data sets corresponding to all resolutions, it calculates the adversarial loss function of the multi-scale generative adversarial network (GAN) for this training cycle, improving the accuracy and reliability of the adversarial loss function calculation. Finally, based on the distance loss function and the adversarial loss function of the multi-scale GAN for this training cycle, it calculates the target loss function of the multi-scale GAN for this training cycle, improving the accuracy and reliability of the target loss function calculation based on the distance loss function and the adversarial loss function.

[0076] Example 3 Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a three-dimensional facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Figure 4 The described 3D facial data prediction device based on multi-scale generative adversarial networks may include one of a prediction device, a prediction system (including a local system or a cloud system), and a prediction server (including a local server or a cloud server), and the embodiments of the present invention are not limited thereto. Figure 4 As shown, the 3D facial data prediction device based on multi-scale generative adversarial networks may include: The acquisition module 301 is used to acquire CBCT data sets for each of the multiple corrected patients. Each CBCT data set for a corrected patient includes the patient's pre-correction CBCT data and post-correction CBCT data.

[0077] Extraction module 302 is used to extract the model data set of each corrected patient from the CBCT data set of each corrected patient based on the soft tissue isosurface extraction algorithm. The model data set includes pre-correction model data and post-correction model data.

[0078] The processing module 303 is used to perform annotation processing on the model data group of each corrected patient for key soft tissue landmarks on the face to obtain key point annotation information about the corrected patient; and based on the key point annotation information about the corrected patient, to perform data preprocessing operation on the model data group of the corrected patient to obtain the three-dimensional facial point cloud data group of the corrected patient.

[0079] Training module 304 is used to train the constructed multi-scale generative adversarial network based on the three-dimensional point cloud data set of all the treated patients and the key point annotation information of the treated patients, so as to obtain a three-dimensional facial data prediction model.

[0080] The prediction module 305 is used to perform facial data prediction operations on the patient to be treated based on the obtained three-dimensional point cloud data of the patient before treatment using a three-dimensional facial data prediction model, and obtain the prediction results of the three-dimensional facial point cloud data of the patient after treatment.

[0081] It is evident that implementation Figure 4The described 3D facial data prediction device based on a multi-scale generative adversarial network can extract the model data set of each corrected patient from the acquired CBCT data set using a soft tissue isosurface extraction algorithm. It then performs annotation processing on key soft tissue landmarks to obtain key point annotation information for that patient. Furthermore, it performs data preprocessing on the model data set to obtain 3D facial point cloud data, improving the accuracy and reliability of the 3D facial point cloud data used for model training. Subsequently, based on the accurately processed 3D point cloud data set and the corresponding key point annotation information, it applies the constructed multi-scale generative adversarial network... Training to obtain a 3D facial data prediction model can improve the training accuracy of the 3D facial data prediction model. Then, using the accurately trained 3D facial data prediction model, based on the obtained 3D point cloud data of the patient before treatment, facial data prediction operations are performed on the patient to obtain the prediction results of the 3D facial point cloud data of the patient after treatment. This can improve the prediction accuracy of possible changes in facial soft tissues after treatment and help provide intuitive and reliable facial prediction results. This is conducive to making accurate and reliable facial correction decisions based on accurate and reliable facial prediction results, and also helps to improve the effectiveness and efficiency of communication between doctors and patients.

[0082] In an optional embodiment, the 3D point cloud data set includes pre-treatment and post-treatment 3D point cloud data. Furthermore, the training module 304 trains the constructed multi-scale generative adversarial network based on the 3D point cloud data set of all treated patients and keypoint annotation information about those patients to obtain a 3D facial data prediction model. Specifically, this includes: Multiple downsampling operations are performed on the pre-treatment 3D point cloud data of each corrected patient to obtain the pre-treatment 3D point cloud data after each downsampling. The pre-treatment 3D point cloud data of each corrected patient and the pre-treatment 3D point cloud data after each downsampling are used as point cloud input data at the corresponding resolution. Based on the preset division ratio, all point cloud input data of all corrected patients and the key point annotation information corresponding to the corrected patients are divided to obtain the division results, which include the model training set, the model validation set and the model test set. The point cloud input data of all the corrected patients in the model training set and the key point annotation information of the corrected patients are input into the constructed multi-scale generative adversarial network for training, and the model training results of this training cycle are obtained. The model training results include the predicted three-dimensional point cloud data of each corrected patient in the model training set at different resolutions. Based on the model training results of this training cycle and the post-treatment 3D point cloud data of all corrected patients in the model training set, the target loss function of the multi-scale generative adversarial network with respect to this training cycle is calculated. The network parameters of the multi-scale generative adversarial network trained with respect to this training cycle are updated through gradient optimization of the target loss function, and the target model corresponding to this training cycle is obtained. Based on the model validation set, the target model corresponding to this training cycle is evaluated to obtain the evaluation result of the target model corresponding to this training cycle. Based on the evaluation results of multiple target models trained for a preset number of training cycles, the target model to be tested is determined from all target models. Based on the model test set, the target model to be tested is tested to obtain the test results. When the test results meet the preset target test results after training is completed, the target model to be tested is determined as the 3D facial data prediction model.

[0083] As can be seen, this optional embodiment can divide the point cloud input data and key point annotation information of all patients according to a preset division ratio to obtain a training set, a validation set, and a test set. The point cloud input data and key point annotation information of the training set are then input into the constructed multi-scale generative adversarial network for training, obtaining the training results of this training cycle, such as predicted point clouds at different resolutions. The target loss function is calculated based on the predicted point cloud and the real point cloud of this training cycle, and the network parameters are updated through gradient optimization of the target loss function to obtain the target model corresponding to this training. This can improve the training accuracy and reliability of the target model in each training cycle. Subsequently, the target model of each training cycle is evaluated based on the validation set to obtain the model evaluation results of each training cycle, from which the target model to be tested is determined. The evaluation of the validation set can improve the accuracy and reliability of the determination of the target model to be tested. Then, the target model to be tested is tested based on the test set. When the test results meet the preset target test results, the training is considered complete. The testing of the test set can improve the accuracy and reliability of the 3D facial data prediction model obtained by training.

[0084] In this optional embodiment, as an optional implementation, the training module 304 inputs all point cloud input data of all treated patients in the model training set, along with key point annotation information about the treated patients, into the constructed multi-scale generative adversarial network for training. The specific methods for obtaining the model training results for this training cycle include: For each corrected patient in the model training set, all point cloud input data of the corrected patient and key point annotation information about the corrected patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, so as to obtain the key feature vector of the corrected patient in this training cycle. The key feature vectors of the corrected patient in this training cycle are input into the multi-scale decoder of the multi-scale generative adversarial network to generate the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle.

[0085] As can be seen, this optional implementation can input all point cloud input data of each corrected patient in the model training set, as well as the key point annotation information of the corrected patient, into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, to obtain the key feature vector of the corrected patient in this training cycle. The multi-scale encoder can improve the extraction accuracy of key feature vectors. Then, the key feature vector of the corrected patient in this training cycle is input into the multi-scale decoder of the multi-scale generative adversarial network for generation, to obtain the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle. The multi-scale decoder can improve the prediction accuracy and reliability of the predicted point cloud data of the patient at different resolutions.

[0086] In this optional implementation, optionally, the training module 304, for each corrected patient in the model training set, inputs all point cloud input data of the corrected patient and key point annotation information about the corrected patient into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, specifically including the following methods to obtain the key feature vector of the corrected patient in this training cycle: For each corrected patient in the model training set, the key point annotation information of the corrected patient is input into the multi-scale encoder of the constructed multi-scale generative adversarial network, so that the key point feature information corresponding to the key point annotation information can be extracted by the feature extractor in the multi-scale encoder. The point cloud input data at each resolution of all point cloud input data of the corrected patient is input into the multi-scale encoder so that the feature extractor in the multi-scale encoder can extract the point cloud feature information corresponding to each point cloud input data. By using the attention module in the multi-scale encoder, based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data, the potential feature vector corresponding to each type of point cloud input data is determined by establishing the information association between key points and facial point clouds. By using a multi-scale encoder, the latent feature vectors corresponding to the point cloud input data of the corrected patient at all resolutions are spliced ​​and fused to obtain the key feature vector of the corrected patient in this training cycle.

[0087] As can be seen, this optional implementation can also input the key point annotation information for each patient and the point cloud input data at each resolution into the constructed multi-scale encoder of the multi-scale generative adversarial network. The feature extractor in the multi-scale encoder can then extract the key point feature information corresponding to the key point annotation information and the point cloud feature information corresponding to each type of point cloud input data, thereby improving the accuracy of key feature information and point cloud feature information extraction. Subsequently, the attention module in the multi-scale encoder, based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data, establishes the information association between key points and facial point clouds to determine the potential feature vector corresponding to each type of point cloud input data, thereby improving the accuracy of determining the potential feature vector corresponding to each resolution of point cloud input data. Finally, the multi-scale encoder concatenates and fuses the potential feature vectors corresponding to the point cloud input data at all resolutions for the corrected patient to obtain the key feature vector for the corrected patient in this training cycle, thereby improving the accuracy and reliability of key feature vector extraction.

[0088] In this optional embodiment, as another optional implementation, the training module 304 calculates the target loss function of the multi-scale generative adversarial network with respect to the current training cycle based on the model training results of this training cycle and the post-treatment 3D point cloud data of all treated patients in the model training set. Specifically, this includes: The predicted 3D point cloud data at each resolution included in the model training results of this training cycle are used as the predicted point cloud data set corresponding to that resolution, and the post-correction 3D point cloud data of all corrected patients in the model training set are used as the real point cloud data set. The distance loss parameter between the predicted point cloud data set and the real point cloud data set corresponding to that resolution is calculated. Based on the distance loss parameters corresponding to all resolutions in this training cycle and the preset weight parameters corresponding to each resolution, calculate the distance loss function corresponding to this training cycle. Based on the predicted point cloud data set and the real point cloud data set corresponding to all resolutions, calculate the adversarial loss function of the multi-scale generative adversarial network for this training cycle; Based on the distance loss function corresponding to this training cycle and the adversarial loss function of the multi-scale generative adversarial network for this training cycle, calculate the target loss function of the multi-scale generative adversarial network for this training cycle.

[0089] As can be seen, this optional implementation can use the predicted 3D point cloud data at each resolution included in the model training results of this training cycle as the predicted point cloud data set corresponding to that resolution, and use the post-correction 3D point cloud data of all corrected patients in the model training set as the real point cloud data set. It calculates the distance loss parameter between the predicted point cloud data set and the real point cloud data set corresponding to that resolution, and calculates the distance loss function corresponding to this training cycle based on the distance loss parameters corresponding to all resolutions in this training cycle and the preset weight parameters corresponding to each resolution. This improves the accuracy and reliability of the distance loss function calculation for this training cycle. Furthermore, based on the predicted point cloud data sets and real point cloud data sets corresponding to all resolutions, it calculates the adversarial loss function of the multi-scale generative adversarial network (GAN) for this training cycle, improving the accuracy and reliability of the adversarial loss function calculation. Finally, based on the distance loss function and the adversarial loss function of the multi-scale GAN for this training cycle, it calculates the target loss function of the multi-scale GAN for this training cycle, improving the accuracy and reliability of the target loss function calculation based on the distance loss function and the adversarial loss function.

[0090] In another optional embodiment, the processing module 303 performs annotation processing on key facial soft tissue landmarks for each group of model data of the treated patient, specifically obtaining key point annotation information about the treated patient in the following ways: For each type of model data contained in the model data set of each corrected patient, obtain the annotation results of the key facial soft tissue landmarks of each of the multiple annotation objects for the model data. The annotation results include the initial annotation position of each key annotation point among the multiple key annotation points. For each key annotation point, calculate the annotation position difference between each annotation object for that key annotation point based on the initial annotation positions of all annotation objects for that key annotation point; The target annotation position of the key annotation point is determined based on the difference in annotation positions between all annotated objects for that key annotation point; Based on the target annotation location of all key annotation points in all model data contained in the model data set of each corrected patient, determine the key point annotation information of that corrected patient.

[0091] As can be seen, this optional embodiment can obtain the annotation results of key facial soft tissue landmarks for each model data of each corrected patient by different annotation objects, such as the initial annotation position, and determine the target annotation position of the key annotation point based on the error of the initial annotation position of all annotation objects for the same key annotation point. This can improve the accuracy and reliability of key annotation point annotation. Then, based on the target annotation position of all key annotation points, the key point annotation information of the corresponding patient can be determined, which can improve the accuracy and reliability of key point annotation information determination. This is conducive to improving the accuracy of subsequent data preprocessing of model data based on the accurately determined key point annotation information.

[0092] In yet another alternative embodiment, such as Figure 5 As shown, Figure 5 This is a schematic diagram of another 3D facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention, wherein the device further includes: The analysis module 306 is used to analyze each model data in the model data group of each treated patient, based on the model data and preset facial structure information, to obtain the degree of facial structural integrity that the model data can express; and to analyze the scan quality corresponding to the model data.

[0093] The filtering module 307 is used to filter the model data groups of all treated patients based on the degree of facial structural integrity that can be expressed by all model data in the model data groups of all treated patients, as well as the scanning quality corresponding to the model data, to obtain the target model data group that meets the preset data quality conditions; and to trigger the processing module 303 to perform the annotation processing of key soft tissue landmarks on the face for each model data group of treated patients, so as to obtain the key point annotation information about the treated patient.

[0094] It is evident that implementation Figure 5 The described device can also analyze the model data for each patient based on each model data and preset facial structure information to obtain the degree of facial structural integrity it can express, and analyze the corresponding scan quality of the model data. Then, based on the degree of facial structural integrity expressed by all model data in the model data group of all treated patients, and the corresponding scan quality of the model data, the device can screen the model data groups of all treated patients to obtain target model data groups that meet preset data quality conditions. This can improve the accuracy and reliability of screening patient model data, and is conducive to further improving the accuracy and credibility of subsequent model training.

[0095] Example 4 Please see Figure 6 , Figure 6This is a schematic diagram of another 3D facial data prediction device based on a multi-scale generative adversarial network disclosed in an embodiment of the present invention. Figure 6 As shown, the 3D facial data prediction device based on multi-scale generative adversarial networks may include: Memory 401 storing executable program code; Processor 402 coupled to memory 401; The processor 402 calls the executable program code stored in the memory 401 to execute some or all of the steps in the three-dimensional face data prediction based on multi-scale generative adversarial networks described in Embodiment 1 or Embodiment 2 of the present invention.

[0096] Example 5 This invention discloses a computer storage medium storing computer instructions. When these computer instructions are invoked, they are used to execute some or all of the steps in the three-dimensional facial data prediction method based on multi-scale generative adversarial networks described in Embodiment 1 or Embodiment 2 of this invention.

[0097] Example 6 This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform some or all of the steps in the three-dimensional face data prediction method based on multi-scale generative adversarial networks described in Embodiment 1 or Embodiment 2.

[0098] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0099] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0100] Finally, it should be noted that the above embodiments are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting 3D facial data based on multi-scale generative adversarial networks, characterized in that, The method includes: Obtain CBCT data sets for each of the multiple corrected patients, wherein each CBCT data set for the corrected patient includes the patient’s pre-correction CBCT data and the patient’s post-correction CBCT data. Based on the soft tissue isosurface extraction algorithm, the model data set of each corrected patient is extracted from the CBCT data set of each corrected patient. The model data set includes pre-correction model data and post-correction model data. For each of the corrected patients' model data sets, annotation processing is performed on key soft tissue landmarks on the face to obtain key point annotation information for the corrected patient; and based on the key point annotation information for the corrected patient, data preprocessing operation is performed on the model data set of the corrected patient to obtain the three-dimensional facial point cloud data set of the corrected patient. Based on the three-dimensional point cloud data set of all the patients who have undergone correction and the key point annotation information of the patients, the constructed multi-scale generative adversarial network is trained to obtain a three-dimensional facial data prediction model. Using the three-dimensional facial data prediction model, based on the obtained three-dimensional point cloud data of the patient before treatment, a facial data prediction operation is performed on the patient to obtain the prediction result of the three-dimensional facial point cloud data of the patient after treatment.

2. The three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to claim 1, characterized in that, The three-dimensional point cloud data set includes pre-treatment three-dimensional point cloud data and post-treatment three-dimensional point cloud data; And, the step of training the constructed multi-scale generative adversarial network based on the three-dimensional point cloud data set of all the treated patients and the key point annotation information of the treated patients to obtain a three-dimensional facial data prediction model includes: Multiple downsampling operations are performed on the pre-treatment three-dimensional point cloud data of each of the corrected patients to obtain the pre-treatment three-dimensional point cloud data after each downsampling. The pre-treatment three-dimensional point cloud data of each of the corrected patients and the pre-treatment three-dimensional point cloud data after each downsampling are used as point cloud input data at the corresponding resolution. Based on a preset division ratio, all point cloud input data of all the corrected patients and the key point annotation information corresponding to the corrected patients are divided to obtain a division result, which includes a model training set, a model validation set and a model test set. The point cloud input data of all the corrected patients in the model training set and the key point annotation information of the corrected patients are input into the constructed multi-scale generative adversarial network for training, and the model training result of this training cycle is obtained. The model training result includes the predicted three-dimensional point cloud data of each corrected patient in the model training set at different resolutions. Based on the model training results of this training cycle and the post-treatment 3D point cloud data of all corrected patients in the model training set, the target loss function of the multi-scale generative adversarial network with respect to this training cycle is calculated. Then, through gradient optimization of the target loss function, the network parameters of the multi-scale generative adversarial network trained with respect to this training cycle are updated to obtain the target model corresponding to this training cycle. Based on the model validation set, the target model corresponding to this training cycle is evaluated to obtain the evaluation result of the target model corresponding to this training cycle. Based on the evaluation results of multiple target models trained based on a preset number of training cycles, the target model to be tested is determined from all the target models. Based on the model test set, the target model to be tested is tested to obtain test results. When the test results meet the preset target test results after training is completed, the target model to be tested is determined as a three-dimensional facial data prediction model.

3. The three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to claim 2, characterized in that, The process involves inputting all point cloud data of all treated patients in the model training set, along with key point annotation information about those patients, into a pre-constructed multi-scale generative adversarial network for training. The resulting model training output for this training cycle includes: For each corrected patient in the model training set, all point cloud input data of the corrected patient and key point annotation information about the corrected patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, so as to obtain the key feature vector of the corrected patient in this training cycle. The key feature vector of the corrected patient in this training cycle is input into the multi-scale decoder of the multi-scale generative adversarial network to generate the predicted 3D point cloud data of the corrected patient at different resolutions in this training cycle.

4. The three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to claim 3, characterized in that, For each treated patient in the model training set, all point cloud input data of the treated patient and key point annotation information about the treated patient are input into the multi-scale encoder of the constructed multi-scale generative adversarial network for extraction, to obtain the key feature vector of the treated patient in this training cycle, including: For each corrected patient in the model training set, the key point annotation information of the corrected patient is input into the multi-scale encoder of the constructed multi-scale generative adversarial network, so as to extract the key point feature information corresponding to the key point annotation information through the feature extractor in the multi-scale encoder. The point cloud input data at each resolution of all point cloud input data of the corrected patient is input into the multi-scale encoder so that the feature extractor in the multi-scale encoder can extract the point cloud feature information corresponding to each point cloud input data. The attention module in the multi-scale encoder determines the potential feature vector corresponding to each type of point cloud input data by establishing the information association between key points and facial point clouds based on the key point feature information and the point cloud feature information corresponding to each type of point cloud input data. The multi-scale encoder is used to concatenate and fuse the potential feature vectors corresponding to the point cloud input data at all the resolutions of the corrected patient to obtain the key feature vector of the corrected patient in this training cycle.

5. The three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to any one of claims 2-4, characterized in that, The step of calculating the target loss function of the multi-scale generative adversarial network with respect to the current training cycle, based on the model training results of this training cycle and the post-treatment 3D point cloud data of all treated patients in the model training set, includes: The predicted 3D point cloud data at each resolution included in the model training results of this training cycle are taken as the predicted point cloud data set corresponding to that resolution, and the post-correction 3D point cloud data of all corrected patients in the model training set are taken as the real point cloud data set. The distance loss parameter between the predicted point cloud data set corresponding to that resolution and the real point cloud data set is calculated. Based on the distance loss parameters corresponding to all the resolutions in this training cycle and the preset weight parameters corresponding to each resolution, calculate the distance loss function corresponding to this training cycle; Based on the predicted point cloud data set corresponding to all the resolutions and the real point cloud data set, calculate the adversarial loss function of the multi-scale generative adversarial network for this training cycle; Based on the distance loss function corresponding to this training cycle and the adversarial loss function of the multi-scale generative adversarial network for this training cycle, calculate the target loss function of the multi-scale generative adversarial network for this training cycle.

6. The method for predicting three-dimensional facial data based on multi-scale generative adversarial networks according to any one of claims 1-4, characterized in that, The process of annotating key facial soft tissue landmarks in the model data set of each of the treated patients yields key landmark annotation information for that patient, including: For each type of model data contained in the model data set of each of the corrected patients, the annotation results of the key facial soft tissue landmarks of each of the multiple annotation objects for the model data are obtained, and the annotation results include the initial annotation position of each of the multiple key annotation points. For each key annotation point, the annotation position difference between each annotation object for the key annotation point is calculated based on the initial annotation positions of all annotation objects for that key annotation point. The target annotation position of the key annotation point is determined based on the difference in annotation position between all the annotated objects for the key annotation point; Based on the target annotation location of all key annotation points in all model data contained in the model data group of each corrected patient, the key point annotation information of the corrected patient is determined.

7. The three-dimensional facial data prediction method based on multi-scale generative adversarial networks according to claim 6, characterized in that, The method further includes: For each type of model data in the model data group of each treated patient, the model data is analyzed based on the model data and the preset facial structure information to obtain the degree of facial structure integrity that the model data can express; and the scan quality corresponding to the model data is analyzed. Based on the degree of facial structural integrity that can be expressed by all model data in the model data group of all the treated patients, and the corresponding scan quality of the model data, the model data groups of all the treated patients are screened to obtain target model data groups that meet the preset data quality conditions; and the operation of annotating key soft tissue landmarks of the face for each model data group of the treated patients is triggered to obtain key point annotation information about the treated patients.

8. A three-dimensional facial data prediction device based on multi-scale generative adversarial networks, characterized in that, The device includes: The acquisition module is used to acquire CBCT data sets for each of the multiple corrected patients. Each CBCT data set for the corrected patient includes the patient's pre-correction CBCT data and post-correction CBCT data. The extraction module is used to extract the model data set of each of the corrected patients from the CBCT data set of each corrected patient based on the soft tissue isosurface extraction algorithm. The model data set includes pre-correction model data and post-correction model data. The processing module is used to perform annotation processing on the model data group of each of the corrected patients for key soft tissue landmarks on the face to obtain key point annotation information about the corrected patient; and to perform data preprocessing operation on the model data group of the corrected patient based on the key point annotation information about the corrected patient to obtain the three-dimensional facial point cloud data group of the corrected patient. The training module is used to train the constructed multi-scale generative adversarial network based on the three-dimensional point cloud data set of all the corrected patients and the key point annotation information of the corrected patients, so as to obtain a three-dimensional facial data prediction model. The prediction module is used to perform facial data prediction operations on the patient to be treated based on the obtained three-dimensional point cloud data of the patient before treatment using the three-dimensional facial data prediction model, and to obtain the prediction result of the three-dimensional facial point cloud data of the patient after treatment.

9. A three-dimensional facial data prediction device based on multi-scale generative adversarial networks, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the three-dimensional facial data prediction method based on multi-scale generative adversarial networks as described in any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the three-dimensional facial data prediction method based on multi-scale generative adversarial networks as described in any one of claims 1-7.