Style migration network and style tooth image generation method and device
By constructing a style migration network and combining a three-dimensional dentition grid model and tooth real-shoot images for training, the problem of difficult to generate tooth images close to the real-shoot effects in the prior art is solved, and high-quality and consistent style tooth image generation is achieved.
Patent Information
- Application Number
- CN202411783632.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to generate teeth images close to the real-life shooting effect, especially when dealing with unstable teeth image quality and diverse shooting angles and lighting environments.
Build a style migration network, generate scene rendering images by moving and/or rotating the three-dimensional dental grid model, and train the network with the real-shot image to generate style teeth images with a style close to the real-shot image of the tooth.
It has achieved the generation of a large number of style teeth images with styles close to the real-time teeth shots, overcoming the problem that traditional methods are difficult to deal with complex textures and lighting environments, and improving the quality and consistency of teeth image generation.
Smart Images

Figure CN119941965A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital medical technology, and in particular to a style transfer network and a method and device for generating styled tooth images. Background Art
[0002] Dental images are a kind of low-cost and easy-to-collect medical data, which contains important oral medical information such as tooth surface status and position. However, their unstable quality and diverse camera shooting angles and lighting environments make them difficult to be used by traditional image recognition methods.
[0003] Although accurate modeling of the oral cavity can be obtained through professional oral scanning equipment, due to the complex texture, lighting environment and soft tissue morphology of the oral scene, it is difficult to produce dental images close to the real effect even if traditional graphics coloring is used based on accurate modeling. Therefore, a method that can batch obtain dental images close to the real effect is urgently needed. Summary of the invention
[0004] In view of this, one of the technical problems solved by the embodiments of the present application is to provide a style transfer network and a method and device for generating style teeth images to overcome the problems in the prior art.
[0005] A first aspect of an embodiment of the present application discloses a method for generating a style transfer network, the method comprising: constructing an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image; performing movement and / or rotation operations on a three-dimensional dentition mesh model of a target sampling object to obtain a first preset number of scene rendering images; obtaining a second preset number of real-shot tooth images of the target sampling object; wherein different real-shot tooth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles; and training the style transfer network using the first preset number of scene rendering images and the second preset number of real-shot tooth images to obtain a trained style transfer network.
[0006] Optionally, performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images includes:
[0007] Performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images, and shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image;
[0008] Correspondingly, the method further includes:
[0009] According to the shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image, the image tags corresponding to all style tooth images are obtained; wherein the image tags are used to characterize the shooting parameter information and / or tooth posture transformation information.
[0010] Optionally, the three-dimensional dentition mesh model of the target sampling object is moved and / or rotated to obtain a first preset number of scene rendering images, and shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image, including:
[0011] Performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a third preset number of three-dimensional textureless models; wherein the third preset number is less than or equal to the first preset number;
[0012] Shooting parameter information corresponding to all three-dimensional texture-free models is randomly set, and rendering processing is performed on all three-dimensional texture-free models to obtain a first preset number of scene rendering images.
[0013] Optionally, the three-dimensional dentition mesh model is composed of a plurality of tooth sub-models, each tooth sub-model being used to simulate one tooth;
[0014] Correspondingly, the three-dimensional dentition mesh model of the target sampling object is moved and / or rotated to obtain a first preset number of scene rendering images, including:
[0015] Random movement and / or rotation operations are performed on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0016] Optionally, performing random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images includes:
[0017] A preset uniformly distributed random function is used to perform random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0018] Optionally, obtaining a second preset number of real-shot images of teeth of the target sampling object includes:
[0019] Using a photographing device to photograph the oral cavity of the target sampling object, to obtain a fourth preset number of real-shot images of teeth; wherein the fourth preset number is greater than or equal to the second preset number;
[0020] Performing image grayscale processing on a fourth preset number of real-shot teeth images to obtain a fifth preset number of grayscale processed images;
[0021] Processing all grayscale processed images using a Laplace operator to obtain a fifth preset number of sharpness quantized images;
[0022] According to the sharpness quantization values corresponding to all the sharpness quantized images, a second preset number of real-shot teeth images are obtained by screening from a fourth preset number of real-shot teeth images.
[0023] Optionally, the method further comprises:
[0024] According to the variance values of all the sharpness quantized images, sharpness quantized values corresponding to all the sharpness quantized images are determined.
[0025] Optionally, training a style transfer network using a first preset number of scene rendering images and a second preset number of real-shot teeth images to obtain a trained style transfer network includes:
[0026] Randomly scaling and / or rotating all or part of the real-shot teeth images to obtain a sixth preset number of real-shot teeth images; wherein the sixth preset number is greater than the second preset number;
[0027] The style transfer network is trained using a first preset number of scene rendering images and a sixth preset number of real-shot teeth images to obtain a trained style transfer network.
[0028] Optionally, training a style transfer network using a first preset number of scene rendering images and a sixth preset number of real-shot teeth images to obtain a trained style transfer network includes:
[0029] A sixth preset number of real-shot tooth images are randomly sampled to obtain a seventh preset number of image training sets; wherein each image training set includes a plurality of real-shot tooth images; wherein the seventh preset number is greater than or equal to 2;
[0030] A seventh preset number of style transfer networks are obtained using the first preset number of scene rendering images and the seventh preset number of image training sets.
[0031] A second aspect of an embodiment of the present application discloses a method for generating a styled dental image, the method comprising: obtaining a three-dimensional dentition mesh model of a target patient; performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images; and using the style transfer network generated as described in the first aspect to generate a corresponding styled dental image for each of the eighth preset number of scene rendering images.
[0032] According to a third aspect of an embodiment of the present application, a device for generating a style transfer network is disclosed, the device comprising: a construction module, configured to construct an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image; a processing module, configured to perform movement and / or rotation operations on a three-dimensional dentition mesh model of a target sampling object to obtain a first preset number of scene rendering images; an acquisition module, configured to obtain a second preset number of real-shot tooth images of the target sampling object; wherein different real-shot tooth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles; and a training module, configured to train the style transfer network using the first preset number of scene rendering images and the second preset number of real-shot tooth images to obtain a trained style transfer network.
[0033] A fourth aspect of an embodiment of the present application discloses a style tooth image generating device, the device comprising: an acquisition module, configured to obtain a three-dimensional dentition mesh model of a target patient; a processing module, configured to perform move and / or rotation operations on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images; a generation module, configured to generate a style tooth image corresponding to each of the eighth preset number of scene rendering images using the style transfer network generated as described in the first aspect.
[0034] Compared with the prior art, the present application constructs an initial style transfer network, the input and output of which are scene rendering images and style tooth images, respectively; next, the style transfer network is trained using a first preset number of scene rendering images obtained by moving and / or rotating the three-dimensional dentition mesh model of the target sampling object, and a second preset number of real-shot tooth images of the target sampling object, to obtain a trained style transfer network, and different real-shot tooth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles. The style transfer network trained in the above manner can generate a large number of style tooth images whose styles are close to the real-shot tooth images. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 It is a flowchart of a method for generating a style transfer network disclosed in Example 1 of the present application;
[0037] Figure 2is a synthesis result of the image labels corresponding to the style tooth image disclosed in the embodiment of the present application;
[0038] Figure 3 It is a flowchart of a method for generating a style tooth image disclosed in the second embodiment of the present application;
[0039] Figure 4 It is a structural schematic diagram of a device for generating a style transfer network disclosed in Embodiment 3 of the present application;
[0040] Figure 5 This is a schematic block diagram of the structure of a style tooth image generation device disclosed in Example 4 of the present application. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0042] It should be noted that the terms "first", "second", "third" and "fourth" in the specification and claims of the present application are used to distinguish different objects rather than to describe a specific order. The terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0043] Embodiment 1
[0044] like Figure 1 As shown, Figure 1 This is a schematic flow chart of a method for generating a style transfer network disclosed in Embodiment 1 of the present application. The method for generating a style transfer network includes:
[0045] Step S101, constructing an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image.
[0046] Step S102: performing a move and / or rotation operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images.
[0047] Step S103, obtaining a second preset number of real-shot teeth images of the target sampling object; wherein different real-shot teeth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles.
[0048] Step S104: training a style transfer network using a first preset number of scene rendering images and a second preset number of real-shot teeth images to obtain a trained style transfer network.
[0049] The present application uses a first preset number of scene rendering images and a second preset number of real-shot teeth images to train a style transfer network, and obtains a trained style transfer network that can generate a large number of styled teeth images that are close to the real-shot effects (i.e., real-shot teeth images).
[0050] The following describes in detail step S101, i.e., "constructing an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image", in conjunction with an embodiment.
[0051] The embodiment of the present application pre-constructs the parameters (such as bias and weight) of the initial style transfer network, and uses the scene rendering image and the style tooth image as the input and output of the style transfer network, respectively.
[0052] Among them, the scene rendering and the style tooth image are used as the input and output of the style transfer network respectively, which are used to make the style tooth image close to the real tooth image.
[0053] The following describes in detail step S102, i.e., "moving and / or rotating the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images", in conjunction with the embodiments.
[0054] The embodiments of the present application move the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images; and / or rotate the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images.
[0055] The three-dimensional dentition mesh model is acquired by an oral scanning device. Optionally, the oral scanning device may be a CBCT (cone beam computed tomography), an intraoral scanner, or the like.
[0056] It should be noted that the number of target sampling objects is one or more. The first preset number may also be multiple.
[0057] Here, by moving and / or rotating the three-dimensional dentition mesh model of one or more target sampling objects to obtain a first preset number of scene rendering images, the number of inputs into the style transfer network can be increased, which can significantly improve the performance of the style transfer network, capture more information and enhance the flexibility of the style transfer network.
[0058] The following describes in detail the method of obtaining the first preset number of scene rendering images.
[0059] In one example, performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images includes:
[0060] The three-dimensional dentition mesh model of the target sampling object is moved and / or rotated to obtain a first preset number of scene rendering images, as well as shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image.
[0061] Here, after the three-dimensional dentition mesh model of the target sampling object is moved and / or rotated, while obtaining a first preset number of scene rendering images, shooting parameter information and / or tooth posture transformation information corresponding to each rendering image is also obtained.
[0062] The moving operation processing may refer to using a matrix to move the three-dimensional dentition mesh model.
[0063] Rotation operations can refer to rotating a model using a rotation matrix. Rotation can be performed around the X, Y, Z axis or a combination of these.
[0064] The shooting parameter information may be the shooting parameter information of the corresponding virtual camera recorded in each scene rendering image after each translation and / or rotation operation, and the shooting parameter information is used to characterize at least one of the camera position, camera shooting angle, and shooting brightness (corresponding to the scene light source that does not coincide with the camera position).
[0065] The tooth posture transformation information may refer to recording the translation parameters and / or rotation parameters of the teeth in the three-dimensional dentition mesh model after each translation and / or rotation operation.
[0066] In the embodiment of the present application, after obtaining a first preset number of scene rendering images, and shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image, the method further includes:
[0067] According to the shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image, the image tags corresponding to all style tooth images are obtained; wherein the image tags are used to characterize the shooting parameter information and / or tooth posture transformation information.
[0068] It should be noted that after the three-dimensional dentition mesh of the target sampling object is moved and / or rotated, the shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image is recorded in order to obtain the image labels corresponding to all style tooth images; wherein the image label is used to represent the shooting parameter information (such as Figure 2 The camera pose corresponding to the lower left in the figure) and / or tooth pose transformation information (such as Figure 2 The tooth position corresponding to the lower left in the figure).
[0069] In one example, the image tag also includes: tooth posture transformation information (such as Figure 2 in the lower left); shooting parameter information (such as Figure 2 ); depth information (such as Figure 2 ).
[0070] The depth information may refer to the distance of the three-dimensional surface corresponding to each pixel from the camera (normalized to 0 to 1, with 0 being near). The underlying graphics API (Application Programming Interface) provides a z-buffer (hidden surface algorithm) interface for processing surface occlusion relationships (pixels with small depth generated by rasterization can cover pixels with large depth, and vice versa are discarded after rasterization and are not processed by the subsequent graphics pipeline). After rendering is completed, the value of the Z-buffer is the depth value of the scene. For each rendered pixel, the depth value of the corresponding position can indicate the distance from the camera.
[0071] It should also be noted that the depth information is for an infrared RGBD camera that uses structured light measurement. The image it produces contains "D", i.e., the depth value, which can more accurately express the spatial structure. The difference is that here the depth information is used as input rather than input to train and test the ability of the network model using this dataset to restore depth information through RGB (red, green, and blue) image reasoning.
[0072] In one example, a three-dimensional dentition mesh model of a target sampling object is moved and / or rotated to obtain a first preset number of scene rendering images, and shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image includes:
[0073] Performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a third preset number of three-dimensional textureless models; wherein the third preset number is less than or equal to the first preset number;
[0074] Shooting parameter information corresponding to all three-dimensional texture-free models is randomly set, and rendering processing is performed on all three-dimensional texture-free models to obtain a first preset number of scene rendering images.
[0075] Here, by randomly setting the shooting parameter information corresponding to all three-dimensional textureless models and rendering all three-dimensional textureless models, a first preset number of scene rendering images are obtained, and a first preset number (such as 2000) of scene rendering images can be obtained.
[0076] It should be noted that the relevant record data of the three-dimensional dental mesh model may include: vertex list data (used to represent the position of each vertex and the normal direction of the surface where the vertex is located), triangle list data (used to represent the three vertex subscripts contained in each triangle).
[0077] The three-dimensional dentition mesh model in the embodiment of the present application is composed of a plurality of tooth sub-models, each of which is used to simulate a tooth;
[0078] Correspondingly, performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images includes:
[0079] Random movement and / or rotation operations are performed on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0080] Specifically, the three-dimensional tooth mesh model of the target sampling object is segmented at the tooth level to obtain multiple tooth sub-models; then, all or part of the tooth sub-models of the target sampling object are randomly moved and / or rotated to obtain a first preset number of scene rendering images. The above-mentioned multiple tooth sub-models can correspond to the image area corresponding to each numbered tooth on the two-dimensional image (i.e., corresponding to Figure 2 Lower right, instance segmentation mask for each.
[0081] The tooth number is the digital number of teeth in the orthodontic field. Normally, it includes 11-17, 21-27, 31-37, and 41-47, which represent the teeth in the four quadrants of the upper and lower dentitions. The image area corresponding to each numbered tooth on the two-dimensional image is the pixel area occupied by the tooth on the two-dimensional image (a black and white mask image with the same size as the whole image, white represents the area occupied by the tooth, and black represents the area not occupied by the tooth).
[0082] Among them, tooth segmentation is a common technical process in the fields of stomatology, dental treatment and three-dimensional reconstruction. Its main purpose is to accurately separate the crown from the overall structure of the tooth.
[0083] Correspondingly, in this example, obtaining a first preset number of scene rendering images includes:
[0084] A preset uniformly distributed random function is used to perform random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0085] For example, (1) import the 3D tooth mesh model into Blender, use the bpy interface script to perform random translation and / or rotation operations, and randomly set the camera shooting parameter information within a preset range (the shooting parameter information includes the camera shooting position, camera shooting angle, and shooting brightness (corresponding to the scene light source that does not coincide with the camera shooting position). (2) use Blender to render the 3D textureless model in the scene graph, and obtain the scene rendering graph (such as Figure 2 The top left one in Figure 2 The above is the data corresponding to the right eye, and the data corresponding to the right eye and the left eye are the same, which will not be repeated here.) Here, the textureless object model can represent the color and texture information of the three-dimensional tooth mesh model.
[0086] Among them, the random setting is set by using a random function, and the formula of the random function rand is:
[0087] Q←30·0.001
[0088] r←rand(0,1)·Q
[0089] α←rand(0,1)·2π
[0090]
[0091] Among them, pos i represents tooth displacement (in meters); rot i represents the Euler angle rotation of the tooth (in radians); rand represents a uniformly distributed random function in the value range; i represents the tooth subscript; Q is a constant used to specify the size of the space for generating random displacement; r is a random number less than Q.
[0092] The embodiment of the present application generates a first preset number of scene rendering images by randomly setting the shooting parameter information corresponding to all three-dimensional textureless models, thereby increasing the richness of the input to the style transfer network, significantly improving the generalization ability and accuracy of the style transfer network and reducing the risk of overfitting, and making the use of the style transfer network obtained by the final training more effective.
[0093] The following is a detailed description of step S103, i.e., "obtaining a second preset number of real-shot teeth images of the target sampling object; wherein different real-shot teeth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles" in conjunction with an embodiment.
[0094] The real-shot tooth image in the embodiment of the present application does not have any label and does not correspond to the three-dimensional dentition mesh model; the real-shot tooth image is used to describe the style characteristics of the real-shot image, and the style characteristics may refer to the overall style of the real-shot image in the oral cavity, such as including the white or yellow color of the tooth surface. Optionally, the first preset number and the second preset number may be the same or different.
[0095] It should be noted that different real-shot tooth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles, which are used to ensure that the real-shot tooth images for training the style transfer network are different.
[0096] In one example, obtaining a second preset number of real-shot images of teeth of the target sampling object may include:
[0097] Using a photographing device to photograph the oral cavity of the target sampling object, to obtain a fourth preset number of real-shot images of teeth; wherein the fourth preset number is greater than or equal to the second preset number;
[0098] Performing image grayscale processing on a fourth preset number of real-shot teeth images to obtain a fifth preset number of grayscale processed images;
[0099] Processing all grayscale processed images using a Laplace operator to obtain a fifth preset number of sharpness quantized images;
[0100] According to the sharpness quantization values corresponding to all the sharpness quantized images, a second preset number of real-shot teeth images are obtained by screening from a fourth preset number of real-shot teeth images.
[0101] It should be noted that the fifth preset number may be greater than or equal to the second preset number and less than or equal to the fourth preset number. The purpose is to filter out images with poor processing effects during the image processing process.
[0102] Here, the sharpness quantization values corresponding to all the sharpness quantization images may be determined according to the variance values of all the sharpness quantization images.
[0103] For example, after obtaining a real patient oral scan video through an oral scanning device, each frame image in the oral scan video is sampled, and then the third-order Laplace operator of the grayscale image is convolved with the grayscale image using the opencv interface to calculate the variance; then, the variance is compared with a preset clarity threshold to determine the clarity of each frame image; the target frame image whose clarity exceeds the preset clarity threshold is used as the real tooth image: where R(t,x,y) represents the image color vector (arranged in red, green, and blue components) at time x,y of the video frame t.
[0104] I(x,y)=(0.299,0.587,0.114)·R(t,x,y)
[0105]
[0106] M=Var(L(x,y))
[0107] The variance M of the image L(x,y) processed by the Laplace operator is used to measure the image clarity, and I(x,y) is the grayscale conversion of a color vector of length 3 to 1 brightness value. The above method was used to screen the clarity of the patient's oral video frames, and 300 real-shot images with high clarity (satisfying M>50) were obtained.
[0108] The following describes in detail step S104, i.e., "training the style transfer network using a first preset number of scene rendering images and a second preset number of real-shot teeth images to obtain a trained style transfer network" in conjunction with the embodiments.
[0109] In the embodiment of the present application, a first preset number of scene rendering images and a second preset number of real-shot teeth images are used as input and output of a style transfer network, respectively, to adjust the parameters of the style transfer network and obtain a trained style transfer network.
[0110] In the process of training the style transfer network, the embodiment of the present application performs data enhancement on the training samples to increase the diversity of the training samples.
[0111] In one example, training a style transfer network using a first preset number of scene rendering images and a sixth preset number of real-shot teeth images to obtain a trained style transfer network includes:
[0112] A sixth preset number of real-shot tooth images are randomly sampled to obtain a seventh preset number of image training sets; wherein each image training set includes a plurality of real-shot tooth images; wherein the seventh preset number is greater than or equal to 2;
[0113] A seventh preset number of style transfer networks are obtained using the first preset number of scene rendering images and the seventh preset number of image training sets.
[0114] In the process of training the style transfer network lattice, data augmentation is used to improve the diversity of generated training samples.
[0115] The real-shot images (i.e., real-shot images of teeth) are randomly scaled and rotated for data augmentation to obtain a sixth preset number of image training sets, specifically:
[0116] Using a first preset number of scene rendering images and a second preset number of real-shot teeth images to train a style transfer network, and obtaining a trained style transfer network includes:
[0117] Randomly scaling and / or rotating all or part of the real-shot teeth images to obtain a sixth preset number of real-shot teeth images; wherein the sixth preset number is greater than the second preset number;
[0118] The style transfer network is trained using a first preset number of scene rendering images and a sixth preset number of real-shot teeth images to obtain a trained style transfer network.
[0119] After obtaining a sixth preset number of real-life tooth images, a bootstrap method may be used to randomly select training subsets, thereby training multiple different style transfer networks, such as three.
[0120] In one example, CycleGAN is used as a style transfer network, and training three CycleGANs is used as an example to illustrate how to use the bootstrap method to train CycleGAN. Specifically:
[0121] Three different CycleGANs generated by bootstrapping. Due to the characteristics of the bootstrapping method (when the number of samples with replacement sampling approaches positive infinity, about one-third of the samples are not sampled once), the three CycleGANs will randomly generate their own preferences, which are manifested in different background colors, textures, gum shapes, etc., thereby improving the richness of the training samples.
[0122] Among them, CycleGAN (Cycle Generative Adversarial Neural Network) is used as the basic architecture of the style transfer network. The framework structure is as follows: Computer Graphics (CG);
[0123] A real =sample(D render )
[0124] B real =sample(D camera )
[0125] B fake =f(A real ,θ)
[0126] A fake =g(B real ,θ)
[0127] A cycle =g(B fake ,θ)=g(f(A real ,θ),θ)
[0128] B cycle =f(A fake ,θ)=f(g(B real ,θ),θ)
[0129] Where D render The dataset corresponding to the scene rendering image, D camera is the dataset corresponding to the real-life dental images, sample is an independent random sample in the dataset, where sampling is to randomly select one of them. For example, Areal randomly selects one of the 2000 scene renderings, and Breal randomly selects one of the 300 real-life images (i.e., real-life dental images). f, g are generators. d render ,d camera is the discriminator, and the output is 0, 1. Among them, 1 represents the same style as the real tooth image, and 0 represents the style is different from the real tooth image. θ is the network parameter. The formula of the loss function is as follows:
[0130] Loss f (θ) = L MSE (d camera (B fake ,θ),1)
[0131] Loss g (θ) = L MSE (d render (A fake ,θ),1)
[0132] Loss cycle (θ)=L1(A cycle ,A real )+L1(B cycle ,B real )
[0133] Loss identity (θ)=L1(f(B real ,θ),B real )+L1(g(A real ,θ),A real )
[0134] L1 and L MSE Represents L1 distance and mean square error loss function, Loss identity is the ontology mapping loss. The style transfer network training is achieved by minimizing the weighted sum of the following loss functions:
[0135] θ * =argmin θ (λ1Loss f (θ)+λ2Loss g (θ)+λ3Losscycle (θ)+λ4Loss identity (θ)
[0136] Among them, Loss cycle (θ) is the cyclic loss function, and λ1, λ2, λ3 and λ4 are weights respectively.
[0137] The beneficial effects of this application are:
[0138] This application takes advantage of the fact that unlabeled and unmatched real-life dental images are easy to obtain, and uses a style transfer network based on overall domain transfer, which has sufficient training data and overcomes the pain point of difficulty in obtaining labeled data.
[0139] This application generates a large number of virtual images (style transfer images) with precise standards by generating real labels based on real tooth models (i.e., three-dimensional dentition mesh data) and using a stylized synthesis method to generate virtual images, which has high practical value.
[0140] Embodiment 2
[0141] Embodiment 2 of the present application provides a method for generating a style tooth image. Figure 3 This is a schematic diagram of a method for generating a style tooth image disclosed in Example 2 of the present application, the method comprising:
[0142] Step 301: Obtain a three-dimensional dentition mesh model of a target patient;
[0143] Step 302: performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images;
[0144] Step 303: Using the pre-trained style transfer network, generate a style tooth image (such as Figure 2 , corresponding real-shot style input in the middle and upper part).
[0145] In the embodiment of the present application, the style of the scene rendering is transferred through the style transfer network to obtain a style tooth image, making it close to the real tooth image. The real tooth image can be a real tooth image obtained by scanning the oral cavity of the target object through an oral scanning device, such as a real tooth image obtained by scanning the oral cavity of 300 real patients.
[0146] Embodiment 3
[0147] Embodiment 3 of the present application provides a device for generating a style transfer network. Figure 4 This is a schematic diagram of the structure of a device for generating a style transfer network disclosed in Embodiment 3 of the present application, the device comprising:
[0148] A construction module 401 is configured to construct an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image;
[0149] The processing module 402 is configured to perform a movement and / or rotation operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images;
[0150] The acquisition module 403 is configured to obtain a second preset number of real-shot images of teeth of the target sampling object; wherein different real-shot images of teeth of the same target sampling object correspond to different shooting device positions, camera shooting positions and / or shooting angles;
[0151] The training module 404 is configured to train the style transfer network using a first preset number of scene rendering images and a second preset number of real-shot teeth images to obtain a trained style transfer network.
[0152] In some examples, the processing module 402 is further configured to:
[0153] Performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images, and shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image;
[0154] Correspondingly, the device also includes:
[0155] The label generation module is configured to obtain image labels corresponding to all styles of tooth images based on the shooting parameter information and / or tooth posture transformation information corresponding to each scene rendering image; wherein the image labels are used to characterize the shooting parameter information and / or tooth posture transformation information.
[0156] In some examples, the processing module 402 is further configured to:
[0157] Performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a third preset number of three-dimensional textureless models; wherein the third preset number is less than or equal to the first preset number;
[0158] Shooting parameter information corresponding to all three-dimensional texture-free models is randomly set, and rendering processing is performed on all three-dimensional texture-free models to obtain a first preset number of scene rendering images.
[0159] In some examples, the three-dimensional dentition mesh model is composed of a plurality of tooth sub-model arrangements, each tooth sub-model being used to simulate a single tooth;
[0160] Correspondingly, the processing module 402 is further configured to:
[0161] Random movement and / or rotation operations are performed on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0162] In some examples, the processing module 402 is further configured to:
[0163] A preset uniformly distributed random function is used to perform random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain a first preset number of scene rendering images.
[0164] In some examples, the obtaining module 403 is further configured to:
[0165] Using a photographing device to photograph the oral cavity of the target sampling object, to obtain a fourth preset number of real-shot images of teeth; wherein the fourth preset number is greater than or equal to the second preset number;
[0166] Performing image grayscale processing on a fourth preset number of real-shot teeth images to obtain a fifth preset number of grayscale processed images;
[0167] Processing all grayscale processed images using a Laplace operator to obtain a fifth preset number of sharpness quantized images;
[0168] According to the sharpness quantization values corresponding to all the sharpness quantized images, a second preset number of real-shot teeth images are obtained by screening from a fourth preset number of real-shot teeth images.
[0169] In some examples, the apparatus further includes:
[0170] The determination module is configured to determine the sharpness quantization values corresponding to all the sharpness quantization images according to the variance values of all the sharpness quantization images.
[0171] In some examples, the training module 404 is further configured to:
[0172] Randomly scaling and / or rotating all or part of the real-shot teeth images to obtain a sixth preset number of real-shot teeth images; wherein the sixth preset number is greater than the second preset number;
[0173] The style transfer network is trained using a first preset number of scene rendering images and a sixth preset number of real-shot teeth images to obtain a trained style transfer network.
[0174] In some examples, the training module 404 is further configured to:
[0175] A sixth preset number of real-shot tooth images are randomly sampled to obtain a seventh preset number of image training sets; wherein each image training set includes a plurality of real-shot tooth images; wherein the seventh preset number is greater than or equal to 2;
[0176] A seventh preset number of style transfer networks are obtained using the first preset number of scene rendering images and the seventh preset number of image training sets.
[0177] Through the style transfer network generation device of this embodiment, the corresponding style transfer network generation methods in the aforementioned multiple method embodiments can be implemented and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0178] Embodiment 4
[0179] Embodiment 4 of the present application provides a device for generating a style tooth image. Figure 5 This is a schematic diagram of the structure of a style tooth image generation device disclosed in Example 4 of the present application, the device comprising:
[0180] An acquisition module 501 is configured to obtain a three-dimensional dentition mesh model of a target patient;
[0181] The processing module 502 is configured to perform a movement and / or rotation operation on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images;
[0182] The prediction module 503 is configured to generate a style tooth image corresponding to each of the eighth preset number of scene rendering images using a pre-trained style transfer network.
[0183] The style tooth image generation device of this embodiment can realize the corresponding style tooth image generation methods in the aforementioned multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0184] So far, specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recorded in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing can be advantageous.
[0185] The present application is described with reference to the flowcharts and / or block diagrams of the methods according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0186] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0187] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods and devices. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0188] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0189] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for generating a style transfer network, characterized in that: The method comprises: Constructing an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image; Performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images; Obtaining a second preset number of real-shot teeth images of the target sampling object; wherein different real-shot teeth images of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles; The style transfer network is trained using the first preset number of the scene rendering images and the second preset number of the real-shot teeth images to obtain the trained style transfer network.
2. The method according to claim 1, characterized in that The step of moving and / or rotating the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images includes: Performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target sampling object to obtain the first preset number of scene rendering images, as well as shooting parameter information and / or tooth posture transformation information corresponding to each of the scene rendering images; Correspondingly, the method further includes: According to the shooting parameter information and / or tooth posture transformation information corresponding to each of the scene rendering images, the image tags corresponding to all the style tooth images are obtained; wherein the image tags are used to characterize the shooting parameter information and / or tooth posture transformation information.
3. The method according to claim 2, characterized in that The moving and / or rotating operation processing is performed on the three-dimensional dentition mesh model of the target sampling object to obtain the first preset number of scene rendering images, and the shooting parameter information and / or tooth posture transformation information corresponding to each of the scene rendering images, including: Performing a moving and / or rotating operation on the three-dimensional dentition mesh model of the target sampling object to obtain a third preset number of three-dimensional textureless models; wherein the third preset number is less than or equal to the first preset number; Shooting parameter information corresponding to all the three-dimensional texture-free models is randomly set, and rendering processing is performed on all the three-dimensional texture-free models to obtain the first preset number of scene rendering images.
4. The method according to claim 1, characterized in that: The three-dimensional dentition mesh model is composed of a plurality of tooth sub-models, each of which is used to simulate a tooth; Correspondingly, the three-dimensional dentition mesh model of the target sampling object is moved and / or rotated to obtain a first preset number of scene rendering images, including: Random movement and / or rotation operations are performed on all or part of the tooth sub-models of the target sampling object to obtain the first preset number of scene rendering images.
5. The method according to claim 4, characterized in that The step of performing random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain the first preset number of scene rendering images includes: A preset uniformly distributed random function is used to perform random movement and / or rotation operations on all or part of the tooth sub-models of the target sampling object to obtain the first preset number of scene rendering images.
6. The method according to claim 1, characterized in that The step of obtaining a second preset number of real-shot images of teeth of the target sampling object comprises: Using a photographing device to photograph the oral cavity of the target sampling object, to obtain a fourth preset number of the real-shot images of the teeth; wherein the fourth preset number is greater than or equal to the second preset number; Performing image grayscale processing on the fourth preset number of the real-shot teeth images to obtain a fifth preset number of grayscale processed images; Processing all the grayscale processed images using a Laplacian operator to obtain the fifth preset number of sharpness quantized images; According to the sharpness quantization values corresponding to all the sharpness quantized images, the second preset number of the real-shot teeth images are obtained by screening from the fourth preset number of the real-shot teeth images.
7. The method according to claim 1, characterized in that The method of training the style transfer network by using the first preset number of scene rendering images and the second preset number of real-shot teeth images to obtain the trained style transfer network includes: Randomly scaling and / or rotating all or part of the real-shot teeth images to obtain a sixth preset number of the real-shot teeth images; wherein the sixth preset number is greater than the second preset number; The style transfer network is trained using the first preset number of the scene rendering images and the sixth preset number of the real-shot teeth images to obtain the trained style transfer network.
8. A method for generating a style tooth image, characterized in that: The method comprises: Obtain a three-dimensional dentition mesh model of the target patient; Performing movement and / or rotation operations on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images; Using the style transfer network generated as described in any one of claims 1 to 7, a style tooth image corresponding to each of the eighth preset number of scene rendering images is generated.
9. A device for generating a style transfer network, characterized in that: The device comprises: A construction module is configured to construct an initial style transfer network; wherein the input of the style transfer network is a scene rendering image, and the output is a style tooth image; A processing module configured to perform a movement and / or rotation operation on the three-dimensional dentition mesh model of the target sampling object to obtain a first preset number of scene rendering images; An acquisition module is configured to acquire a second preset number of real-shot images of teeth of the target sampling object; wherein different real-shot images of teeth of the same target sampling object correspond to different camera shooting positions and / or camera shooting angles; The training module is configured to train the style transfer network using the first preset number of the scene rendering images and the second preset number of the real-shot teeth images to obtain the trained style transfer network.
10. A device for generating a style tooth image, characterized in that: The device comprises: An acquisition module configured to obtain a three-dimensional dentition mesh model of a target patient; a processing module configured to perform a movement and / or rotation operation on the three-dimensional dentition mesh model of the target patient to obtain an eighth preset number of scene rendering images; The prediction module is configured to generate a style tooth image corresponding to each of the eighth preset number of scene rendering images using the style transfer network generated as described in any one of claims 1 to 7.