Orthodontic effect prediction method, device, equipment, storage medium and program product
By obtaining the original tooth-exposed video before orthodontics, predicting and generating the tooth-exposed video after orthodontics, the problem of inaccurate orthodontic effect simulation in the prior art is solved, and a higher degree of orthodontic effect display is achieved.
Patent Information
- Application Number
- CN202510358157.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing orthodontic effect simulation technology is difficult to accurately reflect the real changes after orthodontic treatment, and it is difficult for users to intuitively understand the effects after orthodontics.
By obtaining the original tooth-exposed video before orthodontics, predict the tooth-exposed video after orthodontics, and using this as a constraint condition to generate the tooth-exposed video after orthodontics, the control network and diffusion model are used for feature extraction and video generation, and the authenticity of the video is improved by combining the tooth-exposed and appearance constraint features.
The orthodontic treatment effect display with a higher degree of authenticity is achieved, and users can intuitively feel the changes in teeth after orthodontics, improving the accuracy and intuitiveness of orthodontic simulation.
Smart Images

Figure CN120267424A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of medical technology, and in particular, to an orthodontic effect prediction method, apparatus, device, storage medium, and program product. Background Art
[0002] Orthodontics is an important branch of stomatology, mainly correcting teeth and jaws through scientific methods to improve tooth alignment, occlusion relationship, and facial aesthetics.
[0003] In recent years, orthodontics has received increasing attention and importance, and more and more people choose to undergo orthodontic treatment. Before orthodontic treatment, people often want to know the effect after orthodontics first, but the current orthodontic effect simulation technology has many deficiencies and is difficult to accurately reflect the real changes after orthodontic treatment. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:
[0005] According to a first aspect of one or more embodiments of this specification, an orthodontic effect prediction method is proposed, including:
[0006] Obtain the original tooth-exposing video before orthodontics;
[0007] Predict the tooth contour constraint features after orthodontics according to the original tooth-exposing video;
[0008] Using the tooth contour constraint features as the first constraint condition, generate an effect tooth-exposing video after orthodontics for the original tooth-exposing video.
[0009] Optionally, the process of predicting the tooth contour constraint features after orthodontics according to the original tooth-exposing video includes:
[0010] For each original tooth-exposing image frame in the original tooth-exposing video, predict the orthodontic tooth contour image after orthodontics to obtain the orthodontic tooth contour image sequence corresponding to the original tooth-exposing video;
[0011] Input the orthodontic tooth contour image sequence into a control network for feature extraction to obtain the initial tooth contour features after orthodontics output by the control network;
[0012] Determine the tooth contour constraint features based on the initial tooth contour features.
[0013] Optionally, the process of determining the tooth contour constraint features based on the initial tooth contour features includes:
[0014] Determine the initial tooth contour features as the tooth contour constraint features.
[0015] Optionally, the process of determining the tooth profile constraint features based on the initial tooth profile features includes:
[0016] Input the initial tooth profile features into a temporal network for temporal enhancement to obtain the temporally enhanced initial tooth profile features output by the temporal network;
[0017] Determine the temporally enhanced initial tooth profile features as the tooth profile constraint features.
[0018] Optionally, the initial tooth profile features include a sequence of tooth profile feature units corresponding to the orthodontic tooth profile image sequence. The temporal network includes a global temporal sub-network and a local temporal sub-network. The process of inputting the initial tooth profile features into the temporal network for temporal enhancement to obtain the temporally enhanced initial tooth profile features output by the temporal network includes:
[0019] Input the sequence of tooth profile feature units into the temporal network. The global temporal sub-network captures the global temporal correlation of the sequence of tooth profile feature units, and the local temporal sub-network captures the local motion features between the tooth profile feature units;
[0020] Integrate the global temporal correlation and the local motion features to obtain the temporally enhanced initial tooth profile features.
[0021] Optionally, the process of predicting the orthodontic tooth profile image for each original tooth-exposed image frame in the original tooth-exposed video includes:
[0022] For each original tooth-exposed image frame, extract the oral region image from the original tooth-exposed image frame;
[0023] Based on the oral region image, determine the original tooth profile image before orthodontics;
[0024] Input the original tooth profile image into the orthodontic contour prediction model to obtain the orthodontic tooth profile image output by the orthodontic contour prediction model.
[0025] Optionally, the process of extracting the oral region image from the original tooth-exposed image frame includes:
[0026] For each original tooth-exposed image frame in the original tooth-exposed video, determine the original oral center point of the original tooth-exposed image frame;
[0027] Smoothing the original oral center points in the original tooth-exposing video to update the positions of the original oral center points, and obtaining the updated oral center points for each original tooth-exposing image frame;
[0028] Extracting the oral region image from the corresponding original tooth-exposing image frame according to the updated oral center points.
[0029] Optionally, the process of generating the orthodontic effect tooth-exposing video for the original tooth-exposing video with the tooth contour constraint feature as the first constraint condition includes:
[0030] Using the inversion technique to generate a corresponding noise video for the original tooth-exposing video;
[0031] Inputting the tooth contour constraint feature and the noise video into a diffusion model, and the diffusion model denoises the noise video with the tooth contour constraint feature as the first constraint condition to generate the effect tooth-exposing video.
[0032] Optionally, it further includes:
[0033] Determining the tooth appearance constraint feature according to the original tooth-exposing video;
[0034] The process of generating the orthodontic effect tooth-exposing video for the original tooth-exposing video with the tooth contour constraint feature as the first constraint condition includes:
[0035] Generating the orthodontic effect tooth-exposing video for the original tooth-exposing video with the tooth contour constraint feature as the first constraint condition and the tooth appearance constraint feature as the second constraint condition.
[0036] Optionally, the process of determining the tooth appearance constraint feature according to the original tooth-exposing video includes:
[0037] Extracting the i-th frame of the original tooth-exposing image from the original tooth-exposing video;
[0038] Extracting the tooth appearance constraint feature from the i-th frame of the original tooth-exposing image.
[0039] According to the second aspect of one or more embodiments of this specification, an orthodontic effect prediction device is proposed, including:
[0040] An original acquisition unit, which acquires the original tooth-exposing video before orthodontics;
[0041] A constraint prediction unit, which predicts the tooth contour constraint feature after orthodontics according to the original tooth-exposing video;
[0042] An effect generation unit, which generates the orthodontic effect tooth-exposing video for the original tooth-exposing video with the tooth contour constraint feature as the first constraint condition.
[0043] According to a third aspect of one or more embodiments of this specification, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the steps of the foregoing method by running the executable instructions.
[0044] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the foregoing method are realized.
[0045] According to a fifth aspect of one or more embodiments of this specification, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the foregoing method are realized.
[0046] As can be seen from the above description, this specification can obtain the original tooth-exposing video before orthodontics, predict the tooth contour constraint features after orthodontics according to the original tooth-exposing video, and then use the tooth contour constraint features as the first constraint condition to generate the tooth-exposing video with orthodontic effect for the original tooth-exposing video. By adopting the above technical solution, the orthodontic treatment effect can be dynamically displayed through the tooth-exposing video with orthodontic effect, and the authenticity is higher, enabling users to more intuitively feel the real changes brought by orthodontic treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a schematic structural diagram of an orthodontic effect prediction service system provided by an exemplary embodiment.
[0048] Figure 2 is a flowchart of an orthodontic effect prediction method provided by an exemplary embodiment.
[0049] Figure 3 is a flowchart of a prediction method for tooth contour constraint features provided by an exemplary embodiment.
[0050] Figure 4 is a flowchart of a prediction method for an orthodontic tooth contour image provided by an exemplary embodiment.
[0051] Figure 5 is a schematic structural diagram of a timing network provided by an exemplary embodiment.
[0052] Figure 6 is a flowchart of a method for generating a tooth-exposing video with effect provided by an exemplary embodiment.
[0053] Figure 7 is a flowchart of another orthodontic effect prediction method provided by an exemplary embodiment.
[0054] Figure 8 It is a schematic structural diagram of a device provided by an exemplary embodiment.
[0055] Figure 9 It is a block diagram of an orthodontic effect prediction device provided by an exemplary embodiment. Detailed implementation manners
[0056] Orthodontics is an important branch of stomatology. It mainly corrects teeth and jaws through scientific methods to improve tooth arrangement, occlusion relationship and facial aesthetics.
[0057] In recent years, orthodontics has received increasing attention and importance, and more and more people choose to undergo orthodontic treatment. Before orthodontic treatment, people often want to know the effect after orthodontics first. However, there are many deficiencies in the current orthodontic effect simulation technology, and it is difficult to accurately reflect the real changes after orthodontic treatment.
[0058] This specification provides an orthodontic effect prediction solution, which can effectively improve the orthodontic simulation effect and then accurately reflect the real changes after orthodontic treatment.
[0059] Figure 1 It is a schematic architecture diagram of an orthodontic effect prediction service system provided by an exemplary embodiment. As Figure 1 shown, the system may include a server 11, a network 12, and several electronic devices, such as a PC (Personal Computer) 13, a mobile phone 14, etc.
[0060] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server hosted by a host cluster. During operation, the server 11 may run the server-side program of an application to implement the related functions of the application. For example, when the server 11 runs the program of the orthodontic effect prediction service, it can be implemented as a corresponding orthodontic effect prediction service platform.
[0061] PC13 and mobile phone 14 are only some types of electronic devices that users can use. In fact, users can obviously also use electronic devices of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the electronic device can run the program on the client side of a certain application to implement the related functions of the application. For example, when the electronic device runs the program of the orthodontic effect prediction service, it can be implemented as the client of the orthodontic effect prediction service. Specifically, the client can send the original tooth-exposing video before orthodontics of the user to the server 11, and the server 11 generates the tooth-exposing video after orthodontics effect, and can display the tooth-exposing video returned by the server 11 to the user, etc. Among them, the application program of the client of the above orthodontic effect prediction service can be started and run on the electronic device. The program on the client side can be a native application installed on the electronic device, or the program on the client side can be a small program, a fast application or other similar forms. Of course, when using web technologies such as HTML5 or similar, the related functions can be implemented through the page displayed by the browser. Here, the browser can be an independent browser application or a browser module embedded in some applications.
[0062] For the network 12 for interaction between electronic devices such as PC13 and mobile phone 14 and the server 11, it can be specifically selected to use wired or wireless networks to achieve communication based on the communication methods supported by the corresponding electronic devices. This specification does not limit this. For example, PC13 can support both wired and wireless communications, so wired or wireless networks can be used to achieve communication according to needs, while mobile phone 14 usually only supports wireless communication, so wireless networks can be used to achieve communication.
[0063] Of course, in other examples, the orthodontic effect prediction method provided in this specification can also be only applied in the electronic device, and the electronic device generates the tooth-exposing video after orthodontics effect. This specification does not make special restrictions on this.
[0064] Figure 2 is a flowchart of an orthodontic effect prediction method provided by an exemplary embodiment.
[0065] Please refer to Figure 2 , the orthodontic effect prediction method provided in this embodiment may include the following steps:
[0066] Step 202, obtain the original tooth-exposing video before orthodontics.
[0067] In this embodiment, the tooth-exposing video of the user before orthodontics can be taken as the original tooth-exposing video.
[0068] For example, a video of the user showing their teeth and speaking before orthodontics can be taken as the original toothed video.
[0069] For another example, a video of the user showing their teeth and laughing before orthodontics can also be taken as the original toothed video, etc.
[0070] Step 204: Predict the tooth contour constraint features after orthodontics based on the original toothed video.
[0071] Based on the aforementioned step 202, after obtaining the original toothed video, the tooth contour features after orthodontics can be predicted according to the original toothed video to obtain the tooth contour constraint features. The tooth contour constraint features can be used to constrain the tooth contour in the effective toothed video so that it is as close as possible to the predicted orthodontic effect.
[0072] In one example, an original toothed image frame can be extracted from the original toothed video, and then the orthodontic tooth contour image after orthodontics can be predicted based on this original toothed image frame, and features can be extracted from the orthodontic tooth contour image to obtain the tooth contour constraint features after orthodontics.
[0073] In another example, several original toothed image frames can be extracted from the original toothed video, and the orthodontic tooth contour images after orthodontics can be predicted respectively based on these original toothed image frames to obtain multiple orthodontic tooth contour images, and then features can be extracted from these orthodontic tooth contour images to obtain the tooth contour constraint features, etc.
[0074] Among them, the extracted original toothed image frames can be partial image frames in the original toothed video. For example, the original toothed image frames can be extracted every other frame or every two frames. All the original toothed image frames in the original toothed video can also be extracted for the prediction of the orthodontic tooth contour. Generally speaking, when predicting the orthodontic tooth contour, the more original toothed image frames are used, the more accurate the subsequent tooth contour constraint features will be.
[0075] Step 206: Generate an effective toothed video after orthodontics for the original toothed video with the tooth contour constraint features as the first constraint condition.
[0076] In this embodiment, the tooth contour constraint features that can be predicted are used as the first constraint condition to generate the effective tooth-exposing video. The effective tooth-exposing video corresponds to the original tooth-exposing video and can show the tooth shape after orthodontics. Taking the original tooth-exposing video as a video of a user showing teeth while speaking before orthodontics as an example, the generated effective tooth-exposing video while speaking is the corresponding video of the user showing teeth while speaking after orthodontics. Since the tooth shape after orthodontics of the user is shown in the effective tooth-exposing video while speaking, the user can intuitively feel the real changes of the teeth during speaking after orthodontics.
[0077] As can be seen from the above description, this specification can obtain the original tooth-exposing video before orthodontics, predict the tooth contour constraint features after orthodontics based on the original tooth-exposing video, and then use the tooth contour constraint features as the first constraint condition to generate the effective tooth-exposing video after orthodontics for the original tooth-exposing video. By adopting the above technical solution, the orthodontic treatment effect can be dynamically displayed through the effective tooth-exposing video after orthodontics, with higher authenticity, enabling the user to more intuitively feel the real changes brought by the orthodontic treatment.
[0078] The implementation process of this specification will be described in detail below from two aspects: the prediction of the tooth contour constraint features and the generation of the effective tooth-exposing video.
[0079] I. Prediction of Tooth Contour Constraint Features
[0080] In this embodiment, the tooth contour constraint features can be used to constrain the tooth contour in the generated effective tooth-exposing video, so that the tooth contour in the generated effective tooth-exposing video is closer to the orthodontic effect.
[0081] Please refer to Figure 3 , the prediction process of the tooth contour constraint features may include the following steps:
[0082] Step 302, for each original tooth-exposing image frame in the original tooth-exposing video, predict the orthodontic tooth contour image after orthodontics based on the original tooth-exposing image frame to obtain the orthodontic tooth contour image sequence corresponding to the original tooth-exposing video.
[0083] In this embodiment, taking the prediction of the tooth contour constraint features using all the original tooth-exposing image frames in the original tooth-exposing video as an example, since the original tooth-exposing video usually includes the entire face of the user, before predicting the tooth contour constraint features, the original tooth-exposing video can be preprocessed to extract the oral cavity region image, and then the oral cavity region image is used to predict the tooth contour constraint features, thereby reducing the calculation amount.
[0084] Please refer to Figure 4 , the prediction process of the orthodontic tooth contour image in this step may include the following steps:
[0085] Step 3022: For each original toothed image frame, extract the oral cavity region image from the original toothed image frame.
[0086] In this embodiment, for each original toothed image frame in the original toothed video, a face feature point detection algorithm can be used to detect the mouth region of the face in the original toothed image frame, and then the center point of the oral cavity region in the original toothed image frame can be determined based on the mouth region of the face, which is called the original oral cavity center point.
[0087] Among them, the detected mouth region of the face is usually the position coordinates of multiple feature points in the mouth region of the face, and then the average value of the position coordinates of these feature points can be calculated as the original oral cavity center point.
[0088] In this embodiment, since the user in the original toothed video may be speaking or laughing, these behaviors may cause a large change in the position of the oral cavity region of the user in the video. If the original oral cavity center point is used for cropping the oral cavity region image, the situation of not capturing the complete oral cavity region may occur, which will increase the difficulty of subsequent processing. Therefore, in this embodiment, after determining the original oral cavity center point of the original toothed image frame, the original oral cavity center point can be smoothed. Through the smoothing process, the position of the original oral cavity center point can be adjusted and updated to obtain the updated oral cavity center point, which is called the updated oral cavity center point.
[0089] Specifically, when updating the original oral cavity center point, for each original oral cavity center point, its position can be dynamically adjusted by using the original oral cavity center points of adjacent frames. For example, for each original oral cavity center point, the original oral cavity center point of the previous frame of the original toothed image and the original oral cavity center point of the next frame of the original toothed image can be obtained, and then the updated oral cavity center point can be obtained by weighted summation according to the original oral cavity center points of the previous and next frames and this original oral cavity center point. Among them, the weights can be (W i-1 , W i , W i+1 ). W i-1 represents the weight of the original oral cavity center point of the previous frame, W i represents the weight of the original oral cavity center point of the current frame (i.e., the original oral cavity center point to be updated), and W i+1 represents the weight of the original oral cavity center point of the next frame. Of course, in other examples, when smoothing the original oral cavity center point, a larger sliding window can also be selected. For example, for each original oral cavity center point, the original oral cavity center points of the previous two frames of the original toothed image and the original oral cavity center points of the next two frames of the original toothed image can be obtained to update this original oral cavity center point, etc. This specification does not make special restrictions on this.
[0090] In this embodiment, after obtaining the updated oral center point of each original toothed image frame after the update, the oral region image can be cropped based on the updated oral center point. For example, for each original toothed image frame, a region with a preset size can be cropped centered on its updated oral center point as the corresponding oral region image. The preset size can be 512×512 pixels or the like.
[0091] In this embodiment, by smoothing the original oral center point to obtain the updated oral center point, and then extracting the oral region image based on the updated oral center point, the extracted oral region image can adaptively follow the position change of the oral region in the original toothed video, improving the integrity and accuracy of the extraction of the oral region image, and providing high-quality basic data for subsequent orthodontic tooth contour image prediction and effect toothed video generation.
[0092] Step 3024, determine the original tooth contour image before orthodontics based on the oral region image.
[0093] Based on the foregoing step 3022, for each original toothed image frame, after extracting the corresponding oral region image, the contour image of the user's teeth before orthodontics can be determined based on the oral region image, which is called the original tooth contour image.
[0094] In this embodiment, for each oral region image, an instance segmentation model can be used to perform instance segmentation on it to obtain the contour line of each tooth, and then the tooth contour image corresponding to the oral region image can be obtained, which is called the original tooth contour image. The original tooth contour image can be a two-dimensional array, where the pixel value of the tooth contour line is 255 and the pixel value of the background is 0.
[0095] Step 3026, input the original tooth contour image into the orthodontic contour line prediction model to obtain the orthodontic tooth contour image output by the orthodontic contour line prediction model.
[0096] Based on the foregoing step 3024, for each original toothed image frame, after segmenting to obtain the corresponding original tooth contour image, the original tooth contour image can be input into the trained orthodontic contour line prediction model, and the orthodontic contour line prediction model is used to predict the contour line after orthodontics to obtain the orthodontic tooth contour image output by it. The orthodontic tooth contour image is similar to the original tooth contour image and can also be a two-dimensional array, where the pixel value of the tooth contour line is 255 and the pixel value of the background is 0.
[0097] Among them, when generating sample data for training the orthodontic contour prediction model, the 3D coordinates of the tooth model can be projected onto a 2D plane, and then the tooth contours before and after orthodontics can be paired by matching tooth key points such as cusp tips to ensure that the positions of the tooth key points before and after orthodontics remain consistent, thereby generating the 2D tooth contour line C before orthodontics. before and the 2D tooth contour line C after orthodontics. after as sample data.
[0098] In this embodiment, the orthodontic contour prediction model can adopt the network architecture of an image-to-image diffusion model. In order to better focus on the generation effect of the oral cavity area, Gaussian noise C r ∈R 512×512 is introduced into the orthodontic contour prediction model and the noise is limited to the oral cavity area.
[0099] When training the orthodontic contour prediction model, the 2D tooth contour line C before orthodontics before and the Gaussian noise C r are spliced as the input of the orthodontic contour prediction model, thereby providing clear generation guidance for the model.
[0100] In this embodiment, for each original tooth-exposed image frame in the original tooth-exposed video, a corresponding orthodontic tooth contour image can be predicted, and then multiple orthodontic tooth contour images can be obtained, that is, an orthodontic tooth contour image sequence corresponding to the original tooth-exposed video is obtained.
[0101] Step 304: Input the orthodontic tooth contour image sequence into the control network for feature extraction to obtain the initial tooth contour feature after orthodontics output by the control network.
[0102] Based on the foregoing step 302, after predicting the orthodontic tooth contour image sequence, it can be input into the ControlNet for feature extraction to obtain the tooth contour feature after orthodontics output by the control network, which is called the initial tooth contour feature. Among them, the control network is usually a neural network architecture that can be used to enhance the control ability of a generation model (such as a diffusion model).
[0103] In an example, the control network can extract the tooth contour features for each orthodontic tooth contour image in the orthodontic tooth contour image sequence respectively, and can output the tooth contour features corresponding to each tooth contour image. For easy distinction, it is called the tooth contour feature unit. That is, through the feature extraction of the control network, a sequence of tooth contour feature units corresponding to the orthodontic tooth contour image sequence can be output as the initial tooth contour feature, and each tooth contour feature unit in the tooth contour feature unit sequence corresponds to an orthodontic tooth contour image.
[0104] In another example, the control network can also globally model the orthodontic tooth contour image sequence and extract a unified tooth contour feature therefrom as the initial tooth contour feature, and this specification does not impose special restrictions on this.
[0105] Step 306: Determine the tooth contour constraint feature based on the initial tooth contour feature.
[0106] Based on the foregoing step 304, after obtaining the orthodontic tooth contour conditional feature output by the control network, the tooth contour constraint feature for constraining the effect showing teeth video can be determined based on this tooth contour conditional feature.
[0107] In one example, relatively simply, the orthodontic initial tooth contour feature output by the foregoing control network can be directly determined as the tooth contour constraint feature.
[0108] In another example, a temporal network can be used to perform temporal enhancement on the initial tooth contour feature, and the temporally enhanced initial tooth contour feature can be determined as the tooth contour constraint feature.
[0109] Specifically, when the initial tooth contour feature output by the control network is a sequence of tooth contour feature units corresponding to the orthodontic tooth contour image sequence, since the control network extracts the tooth contour features for each orthodontic tooth contour image respectively, it cannot guarantee the consistency of the extracted sequence of tooth contour feature units in the time series. Therefore, a temporal network can be used to perform temporal enhancement on the initial tooth contour feature.
[0110] Please refer to Figure 5 the schematic diagram of the temporal network structure shown. The temporal network can include a global time sub-network and a local time sub-network, and the global time sub-network and the local time sub-network can perform parallel processing on the initial tooth contour feature.
[0111] Specifically, after inputting the initial tooth contour features (i.e., the sequence of tooth contour feature units corresponding to the orthodontic tooth contour image sequence) into the temporal network, on the one hand, the global temporal sub-network can capture the global temporal correlation of the sequence of tooth contour feature units. For example, the global temporal sub-network may include a Content-aware Cross-attention Block and a Temporal Attention Block. The Content-aware Cross-attention Block can capture the content correlation between different tooth contour feature units through the cross-attention mechanism, and the Temporal Attention Block can capture the temporal dependence relationship between tooth contour feature units through the attention mechanism, thereby capturing the global temporal correlation of the sequence of tooth contour feature units.
[0112] After inputting the initial tooth contour features into the temporal network, on the other hand, the local temporal sub-network can capture the local motion features between the sequences of tooth contour feature units. For example, the local temporal sub-network may include one or more Temporal Convolution Blocks, and the Temporal Convolution Block can capture the local motion features between adjacent tooth contour feature units through convolution operations.
[0113] Then, the initial tooth contour features enhanced by time series can be output by integrating the global temporal correlation and the local motion features to serve as the tooth contour constraint features.
[0114] It can be seen that in this embodiment, the initial tooth contour features output by the control network can be input into the temporal network, and the temporal network performs temporal enhancement, thereby obtaining temporally continuous initial tooth contour features. Subsequently, the generation of the effective tooth-exposing video is carried out with the tooth contour constraint features as the constraint conditions, which helps to improve the coherence of the effective tooth-exposing video.
[0115] II. Generation of Effective Tooth-Exposing Video
[0116] Please refer to Figure 6 , the generation process of the effective tooth-exposing video may include the following steps:
[0117] Step 602, use the inversion technology to generate a corresponding noise video for the original tooth-exposing video.
[0118] In this embodiment, the DDIM (Denoising Diffusion Implicit Models) inversion technology can be used to add noise to the original tooth-exposing video first to generate a corresponding noise video.
[0119] Step 604: Input the tooth contour constraint feature and the noise video into the diffusion model. The diffusion model denoises the noise video with the tooth contour constraint feature as the first constraint condition to generate the effect tooth-exposing video.
[0120] In this embodiment, the effect tooth-exposing video can be generated based on the video-to-video diffusion model (Diffusion Models). The diffusion model can be a model based on the 3D UNet structure. By adding sinusoidal positional encoding in its temporal attention layer, the diffusion model can understand the temporal position information of each frame in the video, thereby making the generated effect tooth-exposing video more continuous.
[0121] In this embodiment, the noise video and the tooth contour constraint feature can be input into the diffusion model. The diffusion model can denoise the noise video with the tooth contour constraint feature as the constraint condition, and then generate the corresponding effect tooth-exposing video. Due to the constraint of the tooth contour constraint feature, the tooth shape in the effect tooth-exposing video generated by the diffusion model can be close to the orthodontic prediction effect.
[0122] Optionally, in another embodiment of this specification, to further improve the authenticity of the effect tooth-exposing video, the tooth appearance constraint feature of the user's teeth can also be determined according to the original tooth-exposing video. The tooth appearance constraint feature can represent the appearance features of the user's teeth in the original tooth-exposing video, such as color features, texture features, etc. Then, when generating the effect tooth-exposing video, on the basis of using the tooth contour constraint feature as the first constraint condition, the tooth appearance constraint feature is used as the second constraint condition, so that the appearance of the teeth in the generated effect tooth-exposing video is closer to the user's real teeth, avoiding distortion caused by the difference between the tooth appearance and the user's teeth.
[0123] In this embodiment, the i-th frame of the original tooth-exposing image can be extracted from the original tooth-exposing video, for example, the first frame of the original tooth-exposing image. Then, the appearance encoder can be used to extract the tooth appearance feature of the teeth from the first frame of the original tooth-exposing image as the appearance constraint feature.
[0124] Exemplarily, the appearance encoder can be a copy of the 2D UNet in the diffusion model. For example, this copy can be fine-tuned and trained as the appearance encoder. Of course, the appearance encoder can also be a convolutional neural network model (Convolutional Neural Networks, CNN), etc. This specification does not make special restrictions on this.
[0125] Exemplarily, when extracting the tooth appearance constraint feature, the tooth appearance constraint feature can also be extracted from the oral region image corresponding to the first frame of the original tooth-exposing image.
[0126] In this embodiment, when generating the effect showing teeth video, the tooth contour constraint feature, the tooth appearance constraint feature, and the noise video can be input into the diffusion model together. The diffusion model can use the tooth contour constraint feature as the first constraint condition and the tooth appearance constraint condition as the second constraint condition to denoise the noise video, and then generate the corresponding effect showing teeth video. Due to the constraint of the tooth contour constraint feature, the tooth shape in the effect showing teeth video generated by the diffusion model can be close to the orthodontic prediction effect. Due to the constraint of the tooth appearance constraint feature, the appearance of the teeth in the effect showing teeth video generated by the diffusion model, such as color and texture, can be closer to the user's real teeth, avoiding distortion of the tooth appearance in the effect showing teeth video, so that the user can more intuitively feel the real changes brought by orthodontic treatment.
[0127] Figure 7 It is a flowchart of another orthodontic effect prediction method provided by an exemplary embodiment.
[0128] Please refer to Figure 7 , after obtaining the original showing teeth video of the user before orthodontics, on the one hand, the DDIM inversion technology can be used to generate the corresponding noise video. On the other hand, the tooth contour constraint feature and the tooth appearance constraint feature can be generated based on the original showing teeth video respectively. Then, the noise video, the tooth contour constraint feature, and the tooth appearance constraint feature can be input into the diffusion model. The diffusion model can use the tooth contour constraint feature as the first constraint condition and the tooth appearance constraint feature as the second constraint condition to denoise the noise video, and then generate the effect showing teeth video.
[0129] In the above process, the determination methods of the tooth contour constraint feature and the tooth appearance constraint feature can refer to the detailed description of the foregoing embodiments. It should be noted that the original showing teeth video described in this specification can be a video including only the oral cavity area. For example, the oral cavity area video can be extracted from the received video including the user's entire facial image as the original showing teeth video to be input into the model for generating the effect showing teeth video, so as to reduce the processing amount of the model. The generated effect showing teeth video is also a video including only the oral cavity area (refer to the example of Figure 7 ), and then it can be pasted back into the video including the user's entire facial image to obtain the final video and return it to the user. Optionally, the original showing teeth video described in this specification can also be a video including the user's entire facial image, and the original showing teeth video can also be directly input into the model to generate the corresponding effect showing teeth video. This specification does not make special restrictions on this.
[0130] Figure 8 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 8, at the hardware level, the device includes a processor 802, an internal bus 804, a network interface 806, a memory 808, and a non-volatile memory 810. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in software. For example, the processor 802 reads the corresponding computer program from the non-volatile memory 810 into the memory 808 and then runs it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.
[0131] Please refer to Figure 9 , the orthodontic effect prediction device 900 can be applied to a device as shown in Figure 8 to implement the technical solutions of this specification. Among them, the orthodontic effect prediction device 900 may include:
[0132] An original acquisition unit 902, which acquires the original tooth-exposing video before orthodontics;
[0133] A constraint prediction unit 904, which predicts the tooth contour constraint features after orthodontics according to the original tooth-exposing video;
[0134] An effect generation unit 906, which uses the tooth contour constraint features as the first constraint condition to generate an effect tooth-exposing video after orthodontics for the original tooth-exposing video.
[0135] Optionally, the process in which the constraint prediction unit 904 predicts the tooth contour constraint features after orthodontics according to the original tooth-exposing video includes:
[0136] For each original tooth-exposing image frame in the original tooth-exposing video, predict the orthodontic tooth contour image after orthodontics to obtain an orthodontic tooth contour image sequence corresponding to the original tooth-exposing video;
[0137] Input the orthodontic tooth contour image sequence into a control network for feature extraction to obtain the initial tooth contour features after orthodontics output by the control network;
[0138] Determine the tooth contour constraint features based on the initial tooth contour features.
[0139] Optionally, the process of determining the tooth contour constraint features based on the initial tooth contour features includes:
[0140] Determine the initial tooth contour features as the tooth contour constraint features.
[0141] Optionally, the process of determining the tooth profile constraint features based on the initial tooth profile features includes:
[0142] Input the initial tooth profile features into a temporal network for temporal enhancement to obtain the temporally enhanced initial tooth profile features output by the temporal network;
[0143] Determine the temporally enhanced initial tooth profile features as the tooth profile constraint features.
[0144] Optionally, the initial tooth profile features include a sequence of tooth profile feature units corresponding to the orthodontic tooth profile image sequence. The temporal network includes a global temporal sub-network and a local temporal sub-network. The process of inputting the initial tooth profile features into the temporal network for temporal enhancement to obtain the temporally enhanced initial tooth profile features output by the temporal network includes:
[0145] Input the sequence of tooth profile feature units into the temporal network. The global temporal sub-network captures the global temporal correlation of the sequence of tooth profile feature units, and the local temporal sub-network captures the local motion features between the tooth profile feature units;
[0146] Combine the global temporal correlation and the local motion features to obtain the temporally enhanced initial tooth profile features.
[0147] Optionally, the process of predicting the orthodontic tooth profile image for each original tooth-exposed image frame in the original tooth-exposed video based on the original tooth-exposed image frame includes:
[0148] For each original tooth-exposed image frame, extract the oral cavity region image from the original tooth-exposed image frame;
[0149] Determine the original tooth profile image before orthodontics based on the oral cavity region image;
[0150] Input the original tooth profile image into an orthodontic contour prediction model to obtain the orthodontic tooth profile image output by the orthodontic contour prediction model.
[0151] Optionally, the process of extracting the oral cavity region image from the original tooth-exposed image frame includes:
[0152] For each original tooth-exposed image frame in the original tooth-exposed video, determine the original oral cavity center point of the original tooth-exposed image frame;
[0153] Smooth the original oral cavity center points in the original tooth-exposed video to update the positions of the original oral cavity center points and obtain the updated oral cavity center points for each original tooth-exposed image frame;
[0154] Extract the oral cavity region image from the corresponding original tooth-exposed image frame according to the updated oral cavity center point.
[0155] Optionally, the effect generation unit 906 uses the tooth contour constraint feature as the first constraint condition to generate the orthodontic effect tooth-exposed video for the original tooth-exposed video. The process includes:
[0156] Use the inversion technique to generate the corresponding noise video for the original tooth-exposed video;
[0157] Input the tooth contour constraint feature and the noise video into the diffusion model, and the diffusion model denoises the noise video with the tooth contour constraint feature as the first constraint condition to generate the effect tooth-exposed video.
[0158] Optionally, the constraint prediction unit 904 is further configured to determine the tooth appearance constraint feature according to the original tooth-exposed video;
[0159] The effect generation unit 906 is further configured to use the tooth contour constraint feature as the first constraint condition and the tooth appearance constraint feature as the second constraint condition to generate the orthodontic effect tooth-exposed video for the original tooth-exposed video.
[0160] Optionally, the process of determining the tooth appearance constraint feature according to the original tooth-exposed video includes:
[0161] Extract the i-th frame of the original tooth-exposed image from the original tooth-exposed video;
[0162] Extract the tooth appearance constraint feature from the i-th frame of the original tooth-exposed image.
[0163] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above embodiments.
[0164] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.
[0165] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.
Claims
1. An orthodontic effect prediction method, comprising: Obtaining an original tooth-exposing video before orthodontics; Predicting the tooth contour constraint features after orthodontics based on the original tooth-exposing video; Using the tooth contour constraint features as the first constraint condition to generate an effect tooth-exposing video after orthodontics for the original tooth-exposing video.
2. The method according to claim 1, wherein the process of predicting the tooth contour constraint features after orthodontics based on the original tooth-exposing video comprises: For each original tooth-exposing image frame in the original tooth-exposing video, predicting an orthodontic tooth contour image after orthodontics according to the original tooth-exposing image frame to obtain an orthodontic tooth contour image sequence corresponding to the original tooth-exposing video; Inputting the orthodontic tooth contour image sequence into a control network for feature extraction to obtain the initial tooth contour features after orthodontics output by the control network; Determining the tooth contour constraint features based on the initial tooth contour features.
3. The method according to claim 2, wherein the process of determining the tooth contour constraint features based on the initial tooth contour features comprises: Determining the initial tooth contour features as the tooth contour constraint features.
4. The method according to claim 2, wherein the process of determining the tooth contour constraint features based on the initial tooth contour features comprises: Inputting the initial tooth contour features into a temporal network for temporal enhancement to obtain the temporally enhanced initial tooth contour features output by the temporal network; Determining the temporally enhanced initial tooth contour features as the tooth contour constraint features.
5. The method according to claim 4, wherein the initial tooth contour features include a sequence of tooth contour feature units corresponding to the orthodontic tooth contour image sequence, and the temporal network includes a global time sub-network and a local time sub-network. The process of inputting the initial tooth contour features into the temporal network for temporal enhancement to obtain the temporally enhanced initial tooth contour features output by the temporal network comprises: Inputting the sequence of tooth contour feature units into the temporal network, and having the global time sub-network capture the global time correlation of the sequence of tooth contour feature units, and having the local time sub-network capture the local motion features between the sequence of tooth contour feature units; Integrating the global time correlation and the local motion features to obtain the temporally enhanced initial tooth contour features.
6. The method according to claim 2, wherein the process of predicting an orthodontic tooth contour image after orthodontics for each original tooth-exposing image frame in the original tooth-exposing video according to the original tooth-exposing image frame comprises: For each original tooth-exposing image frame, extracting an oral cavity region image from the original tooth-exposing image frame; Determining an original tooth contour image before orthodontics based on the oral cavity region image; Inputting the original tooth contour image into an orthodontic contour line prediction model to obtain the orthodontic tooth contour image output by the orthodontic contour line prediction model.
7. The method according to claim 6, wherein the process of extracting an oral cavity region image from the original tooth-exposing image frame comprises: For each original toothed image frame in the original toothed video, determine the original oral center point of the original toothed image frame; Smooth the original oral center points in the original toothed video to update the positions of the original oral center points, and obtain the updated oral center points of each original toothed image frame; Extract the oral region image from the corresponding original toothed image frame according to the updated oral center point.
8. The method according to claim 1, wherein the process of generating the orthodontic toothed video with the orthodontic effect for the original toothed video with the tooth contour constraint feature as the first constraint condition includes: Use inversion technology to generate a corresponding noise video for the original toothed video; Input the tooth contour constraint feature and the noise video into a diffusion model, and the diffusion model denoises the noise video with the tooth contour constraint feature as the first constraint condition to generate the toothed video with the orthodontic effect.
9. The method according to claim 1 further includes: Determine the tooth appearance constraint feature according to the original toothed video; The process of generating the orthodontic toothed video with the orthodontic effect for the original toothed video with the tooth contour constraint feature as the first constraint condition includes: Generate the orthodontic toothed video with the orthodontic effect for the original toothed video with the tooth contour constraint feature as the first constraint condition and the tooth appearance constraint feature as the second constraint condition.
10. The method according to claim 9, wherein the process of determining the tooth appearance constraint feature according to the original toothed video includes: Extract the i-th frame of the original toothed image from the original toothed video; Extract the tooth appearance constraint feature from the i-th frame of the original toothed image.
11. An orthodontic effect prediction device includes: An original acquisition unit that acquires the original toothed video before orthodontics; A constraint prediction unit that predicts the tooth contour constraint feature after orthodontics according to the original toothed video; An effect generation unit that generates the orthodontic toothed video with the orthodontic effect for the original toothed video with the tooth contour constraint feature as the first constraint condition.
12. An electronic device, comprising: A processor; A memory for storing instructions executable by the processor; wherein, the processor realizes the steps of the method according to any one of claims 1-10 by running the executable instructions.
13. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1-10 are realized.
14. A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1-10 are realized.