A method for data enhancement of oral CT images
By using feature control vectors for directional generation in the generative adversarial network, the problem of oral CT image data enhancement in the prior art is solved, especially when the data volume is small and the category is unbalanced, the generation of high-resolution and high-quality images is achieved.
Patent Information
- Application Number
- CN202111644473.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The prior art is difficult to effectively enhance oral CT image data with a small amount of original data, especially in the case of category imbalance, and it is impossible to generate high-resolution high-quality images.
Generative adversarial network (GAN) is used for data augmentation, and principal component analysis is performed on the latent space, and directional generation is recognized and used to use feature control vectors to achieve the generation of high-resolution oral CT images.
In the case of small amount of original data and unbalanced categories, high-resolution high-quality oral CT images can be generated, reducing the workload of manual annotation, and suitable for multiple different classification tasks.
Smart Images

Figure CN114359090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and more specifically, to a data enhancement method for oral CT images. Background Art
[0002] To complete the big data intelligent analysis of oral CT images, it is necessary to collect a large number of oral CT images, annotate the data by professionals, construct an oral CT image dataset, and perform subsequent intelligent analysis based on this dataset. The collection and annotation of the dataset are time-consuming and laborious. More importantly, due to the different incidence rates of different diseases, the sample quantity of common diseases is very different from that of rare diseases, resulting in the problem of class imbalance in the dataset.
[0003] In the prior art, some processing methods rely on the Euclidean distance between the image to be classified and the prototype images of each category to judge the category of the image to be classified. However, this method is not applicable to oral CT image data because the differences between images of different categories in the oral CT image dataset are only reflected in a relatively small area in many cases, and the differences between the prototype images of each category obtained by this method are not significant. Some processing methods apply the generative adversarial network WGAN to the ROI region dataset of thyroid ultrasound data, realizing the expansion of the ROI region datasets of benign and malignant thyroid ultrasound data, and comparing the performance of the VGG-16 network on the enhanced dataset and the original dataset, proving the feasibility of this data enhancement method. However, it cannot generate image data of a specified category in the original dataset directionally because it lacks a module for controlling the generated image category. The images randomly generated by the generative adversarial network need additional manual classification and annotation to achieve data enhancement. The annotation of medical images needs to be completed by professionals such as doctors, which takes a lot of time when the number of categories is large. Moreover, this method requires a large amount of original data, and requires the sample distribution of each category in the original training set to be relatively uniform. When the amount of original data is small and the original dataset has a serious class imbalance, this method may not be competent. At the same time, this method cannot generate high-quality images with a high resolution such as 512*512, and performs poorly in medical image classification tasks with high requirements for medical image data quality and rich dataset details. Summary of the Invention
[0004] In order to overcome the above defects in the prior art, the present invention provides a data enhancement method for oral CT images, which can generate high-resolution oral CT images under the condition of less original data to assist the oral CT images to achieve big data intelligent analysis.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: A data enhancement method for oral CT images, comprising the following steps:
[0006] S1. Enhance data based on the generative adversarial network: Input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model.
[0007] S2. Use the interpretable control discovery algorithm in the generative adversarial network space to obtain the feature control vector: Use the generative adversarial network space algorithm to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector.
[0008] S3. Directionally generate oral CT images using the feature control vector.
[0009] Furthermore, the generative adversarial network includes a first generator, a second generator, a discriminator, and a mapping network; the first generator divides the original oral CT image into N×N units, and designs an objective function to train the rectangular annotation of the ROI area on the oral CT image, obtaining an objective recognition model for the maxillary sinus ROI area of the oral CT image, which is used to output the coordinates of the maxillary sinus ROI area of the oral CT image; the second generator and the mapping network cooperate with each other, the mapping network decouples the latent space, generates an intermediate latent variable w from the hidden variable z, adds the affine transformation and random noise converted from the variable w to each layer of the second generator, to achieve the purpose of controlling the features of the generated image at each scale and realize controllable image generation; the discriminator accepts the input of the coordinate information from the first generator and the fake image information generated by the second generator, and judges the authenticity of the image generated by the second generator.
[0010] Furthermore, step S1 specifically includes:
[0011] S11. Input the original oral CT image into the generative adversarial network. A part of the original oral CT image has a rectangular annotation for the maxillary sinus ROI area, and the other part does not; after being processed by the first generator, the unannotated oral CT image obtains a rectangular annotation centered on the maxillary sinus ROI area with the same size.
[0012] S12. After the mapping network normalizes the hidden variable z, it passes through multiple fully connected layers to obtain the intermediate hidden variable w, and cooperates with random noise as the input of the initial and each convolutional layer of the second generator. The input process uses the AdaIN algorithm to generate controllable images.
[0013] S13. The outputs of the first generator and the second generator are jointly used as the input of the discriminator to judge whether the input picture is real or fake.
[0014] Among them, AdaIN is an off-the-shelf algorithm, and adaptive instance normalization is a literal translation of it. It is an algorithm for style transfer. By inputting a style image and a content image, some features of the style image can be transferred to the content image. IN is a term in machine learning, that is, instance normalization, which is used to calculate the mean and standard deviation of all pixels of a single image. Ada is an abbreviation of adaptive. As the name implies, it adaptively adjusts IN.
[0015] Furthermore, the objective function of the first generator is as follows:
[0016]
[0017] The output of the first generator is a 1×5 vector, that is, the predicted In the formula, when there is an object in the grid, p ij is 1, q ij is 0. When there is no object in the grid, p ij is 0, q ij is 1; x i and y i are the coordinates of the true bounding box, and are the coordinates of the predicted bounding box; w i and h i are the width and height of the true bounding box, and are the width and height of the predicted bounding box; c i is the confidence value, is the intersection value of the predicted bounding box and the true bounding box, that is, the area of the intersection of the predicted bounding box and the true bounding box divided by the area of the union.
[0018] Furthermore, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as follows:
[0019]
[0020] In the formula, λ is a constant, denotes the mathematical expectation, denotes the gradient, D is a probability function, which judges the probability that the input parameter is true, and the result is in the interval [0,1]; denotes the L2 norm of the gradient; x is the real data, is the data generated by the generator, denotes the distribution of the real data, denotes the model distribution implicitly defined by ; Let x r and x gFor a pair of true and false samples randomly sampled from and respectively, introducing a random number ∈ with a value in [0, 1], then refers to linearly interpolating and sampling randomly between x r and x g , and is the distribution satisfied by obtained by this sampling process.
[0021] Furthermore, the specific steps of step S3 include: projecting the real or generated oral CT image to the latent space of the second generator. The latent space is a representation of compressed data, where similar data points are closer in space. A weight parameter set on the feature control vector is used for movement to generate high-resolution oral CT image data of the required type.
[0022] The present invention also provides a data enhancement system for oral CT images, including:
[0023] An acquisition module: used to input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model;
[0024] A processing module: used to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space using the generative adversarial network space algorithm, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector;
[0025] A generation module: used to generate oral CT images in a feature control vector-directed manner.
[0026] Furthermore, the acquisition module includes:
[0027] A first generator unit: used to input the original oral CT image into the generative adversarial network. A part of the original oral CT image has a rectangular annotation for the maxillary sinus ROI area, and the other part does not; after being processed by the first generator, the unannotated oral CT image obtains a rectangular annotation with the maxillary sinus ROI area as the center and the same size; among them, the objective function of the first generator is:
[0028]
[0029] The output of the first generator is a 1×5 vector, that is, the predicted In the formula, when there is a target in the grid, p ij is 1, q ij is 0, when there is no target in the grid, p ij is 0, qij is 1; x i and y i are the real bounding box coordinates, and are the predicted bounding box coordinates; w i and h i are the width and height of the real bounding box, and are the width and height of the predicted bounding box; c i is the confidence value, is the intersection value between the predicted bounding box and the real bounding box, that is, the area of the intersection of the predicted bounding box and the real bounding box divided by the area of the union;
[0030] Mapping network and the second generator unit: After the mapping network normalizes the latent variable z, it passes through multiple fully connected layers to obtain the intermediate latent variable w, and together with random noise, it serves as the input for the initial and each convolutional layer of the second generator. The AdaIN algorithm is used in the input process to generate controllable images;
[0031] Discriminator unit: It is used to take the outputs of the first generator and the second generator as the inputs of the discriminator to determine whether the input image is real or fake;
[0032] Among them, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as:
[0033]
[0034] In the formula, λ is a constant, denotes the mathematical expectation, denotes the gradient, D is a probability function that judges the probability that the input parameter is true, and the result is in the interval [0,1]; denotes the L2 norm of the gradient; x is the real data, is the data generated by the generator, denotes the distribution of the real data, denotes the model distribution implicitly defined by ; Let x r and x g be a pair of real and fake samples randomly sampled from and respectively. Introduce a random number ∈ with a value of [0,1], then denotes linear random interpolation sampling between x r and x g ; then is the distribution satisfied by the obtained from this sampling process.
[0035] The present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the data enhancement method for oral CT images as described above.
[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data enhancement method for oral CT images as described above is implemented.
[0037] Compared with the prior art, the beneficial effects are as follows:
[0038] 1. The present invention adopts an unsupervised learning method. Different from ordinary image data, since oral CT image data has a standardized acquisition process, is taken from a fixed angle, and has high data quality, the unsupervised learning method can be used in the scenario of oral CT images. After obtaining sufficient original oral CT image data, only simple cropping needs to be performed on it in advance, and then the training of the generative adversarial network can be carried out to complete the data enhancement task, which can save the cost of manual classification and annotation, and only one training is required to meet the data enhancement needs of multiple different classification tasks, without separately training the network for each classification task. For different medical CT images, the network of the algorithm can be changed, and parameters can be set specifically for training to meet the requirements of different tasks. Compared with the prior art, the workload of manual annotation can be significantly reduced;
[0039] 2. The present invention can generate high-resolution and high-quality images of 512*512. Compared with the prior art that can only generate images with lower resolution, it can retain more effective information. In some classifications of oral CT images, the differences between different types of images are small, and richer detailed information is required for classification. Therefore, the high-resolution image generation method can be used in the data enhancement scenario of oral CT images, and this feature makes the present invention have stronger versatility and also has application value in other medical image classification tasks with higher requirements for image data quality;
[0040] 3. The algorithm of the present invention adopts the idea of style transfer. Compared with the prior art that can only randomly generate specific types of data, the intensity of generating specific features in the present invention is controllable. For oral CT images that do not exist in the newly acquired original data set, the trained model and the saved feature control vector can also be directly used to edit the image attributes. Since in some classifications of oral CT images, the boundaries between different types of images are not clear, and it is required to be directional and controllable when generating specific types of data, the present invention can be used in the scenario of oral CT images;
[0041] 4. The method provided by the present invention performs excellently in the case of fewer training samples and unbalanced data categories. In the field of medical image processing, the challenge of insufficient sample size is often encountered. For the oral CT image dataset with a small original sample size, the present invention needs to utilize all the information of all original samples as much as possible. Compared with the prior art which requires a relatively balanced original dataset, both positive and negative samples contribute during the training process of the present invention. When generating data of the category with a smaller original sample size, the information of data in other categories can be utilized, enabling the training of the generative adversarial network in a challenging environment with a small amount of original data and serious category imbalance, achieving data augmentation and improving the accuracy of the classification task. Description of the Drawings
[0042] Figure 1 is a schematic flowchart of the method of the present invention.
[0043] Figure 2 is a schematic overview of the generative adversarial network of the present invention.
[0044] Figure 3 is a detailed parameter diagram of the generative adversarial network of the present invention, where the arrows represent the data flow.
[0045] Figure 4 is an example diagram of the original oral CT image and its ROI area processed by the present invention.
[0046] Figure 5 is a schematic flowchart of step S2 of the present invention.
[0047] Figure 6 is a schematic flowchart of step S3 of the present invention.
[0048] Figure 7 is an effect diagram of the oral CT image of a certain category generated directionally by the method of the present invention. Detailed Embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention will be described in one of the embodiments in conjunction with the specific embodiments. Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, rather than physical diagrams, and should not be construed as a limitation to this patent; for better illustrating the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0050] In the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation of this patent. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances. Additionally, if the embodiments of the present invention involve descriptions such as "first", "second", etc., the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the meaning of "and / or" appearing throughout the text is that it includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, or scenario B, or the scenario where both A and B are satisfied simultaneously.
[0051] Embodiment 1:
[0052] As Figure 1 shown, a method for data enhancement of oral CT images includes the following steps:
[0053] S1. Enhance the data based on a generative adversarial network: Input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model;
[0054] S2. Obtain a feature control vector by using an interpretable control discovery algorithm in the generative adversarial network space: Use the generative adversarial network space algorithm to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector;
[0055] S3. Generate oral CT images directionally using the feature control vector.
[0056] A method for data enhancement of oral CT images provided by the present invention can learn each attribute of the CT image data in the oral CT image dataset (such as cysts, effusions, the bottom line of the maxillary sinus, etc.) without supervision and progressively. The trained model can be used to generate high-resolution oral CT images with a resolution of 512×512.
[0057] As shown in Table 1, Table 2 and Table 3, the generative adversarial network includes a first generator, a second generator, a discriminator and a mapping network; the implementation principle of the first generator is that the maxillary sinus ROI region, which plays an important role in the classification task, has a relatively fixed position and size on the original oral CT image; the first generator divides the original oral CT image into 8×8 units and designs an objective function to train the rectangular annotation of the ROI region on the oral CT image to obtain an objective recognition model for the maxillary sinus ROI region of the oral CT image, which is used to output the coordinates of the maxillary sinus ROI region of the oral CT image; the second generator cooperates with the mapping network, the mapping network decouples the latent space, generates an intermediate latent variable w from the hidden variable z, and adds the affine transformation and random noise converted from the variable w to each layer of the second generator to achieve the purpose of controlling the features of each scale of the generated image and realize controllable image generation; the discriminator receives the input of the coordinate information from the first generator and the fake image information generated by the second generator, and judges the authenticity of the image generated by the second generator.
[0058] The present invention uses an interpretable control discovery algorithm in the generative adversarial network space, performs principal component analysis on the latent space of the generative adversarial network model after training with the oral CT image dataset, and unsupervised identifies important latent feature control vectors to discover the interpretable control therein. After finding each feature control vector, it is manually screened and saved, and the specific attributes controlled by each feature control vector are determined. The real or generated oral CT image is projected onto the model latent space, and multiplied by each feature control vector and the weight parameter, so as to generate the data of the desired type, such as converting the oral CT image without fluid into the oral CT image with fluid.
[0059] Table 1 Parameter Table of Each Layer of the First Generator
[0060]
[0061]
[0062] Table 2 Parameter Table of Each Layer of the Second Generator
[0063]
[0064]
[0065] Table 3 Parameter Table of Each Layer of the Discriminator
[0066]
[0067] Further, the step S1 specifically includes:
[0068] S11. Input the original oral CT images into the generative adversarial network. Part of the original oral CT images are marked with rectangles for the ROI area of the maxillary sinus, and the other part is not. After being processed by the first generator, the unmarked oral CT images obtain rectangles with the same size centered on the ROI area of the maxillary sinus. The objective function of the first generator is as follows:
[0069]
[0070] The output of the first generator is a 1×5 vector, that is, the predicted In the formula, when there is a target in the grid, p ij is 1, q ij is 0. When there is no target in the grid, p ij is 0, q ij is 1; x i and y i are the coordinates of the true bounding box. and are the coordinates of the predicted bounding box; w i and h i are the width and height of the true bounding box. and are the width and height of the predicted bounding box; c i is the confidence value. is the intersection value of the predicted bounding box and the true bounding box, that is, the area of the intersection of the predicted bounding box and the true bounding box divided by the area of the union.
[0071] S12. After the mapping network normalizes the latent variable z, it passes through multiple fully connected layers to obtain the intermediate latent variable w, and together with the random noise, it is used as the input of the second generator and each convolutional layer. The AdaIN algorithm is used in the input process to generate controllable images.
[0072] S13. The outputs of the first generator and the second generator are jointly used as the input of the discriminator to determine whether the input pictures are real or fake.
[0073] Among them, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as follows:
[0074]
[0075] In the formula, λ is a constant. denotes the mathematical expectation. denotes the gradient. D is a probability function that judges the probability that the input parameter is true, and the result is in the interval [0,1]. denotes the L2 norm of the gradient; x is the real data. is the data generated by the generator. Refers to the distribution of real data, Refers to the model distribution implicitly defined by Let x r and x g be a pair of true and false samples randomly sampled from and respectively. Introduce a random number ∈ with a value in [0,1], then Refers to linear random interpolation sampling between x r and x g , Then is the distribution satisfied by the obtained by this sampling process.
[0076] In addition, as Figure 3 shown, the specific steps of step S3 include: projecting the real or generated oral CT image to the latent space of the second generator. The latent space is a representation of compressed data, where similar data points are closer in space. Move with the weight parameters set on the feature control vector to generate the high-resolution oral CT image data of the required type. As Figure 7 shown, adjusting the shape control vector of the bottom line of the maxillary sinus can directionally control the change of the shape of the bottom line of the maxillary sinus in the oral CT image, so as to directionally generate the oral CT image data of the required type.
[0077] Embodiment 2
[0078] This embodiment provides a data enhancement system for oral CT images, including:
[0079] Acquisition module: used to input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model;
[0080] Processing module: used to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space using the generative adversarial network space algorithm, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector;
[0081] Generation module: used to directionally generate oral CT images using the feature control vector.
[0082] Specifically, the acquisition module includes:
[0083] The first generator unit: It is used to input the original oral CT image into the generative adversarial network. A part of the original oral CT image has a rectangular annotation for the maxillary sinus ROI region, and the other part does not. After being processed by the first generator, the unannotated oral CT image gets a rectangular annotation centered on the maxillary sinus ROI region with the same size. Among them, the objective function of the first generator is:
[0084]
[0085] The output of the first generator is a 1×5 vector, that is, the predicted In the formula, when there is a target in the grid, p ij is 1, q ij is 0. When there is no target in the grid, p ij is 0, q ij is 1; x i and y i are the real bounding box coordinates, and are the predicted bounding box coordinates; w i and h i are the width and height of the real bounding box, and are the width and height of the predicted bounding box; c i is the confidence value, is the intersection value of the predicted bounding box and the real bounding box, that is, the area of the intersection of the predicted bounding box and the real bounding box divided by the area of the union;
[0086] The mapping network and the second generator unit: The mapping network is used to normalize the hidden variable z and then obtain the intermediate hidden variable w through multiple fully connected layers, and cooperate with random noise as the input of the second generator and each convolutional layer. The AdaIN algorithm is adopted in the input process to generate controllable images;
[0087] The discriminant unit: It is used to take the outputs of the first generator and the second generator as the inputs of the discriminator to discriminate whether the input pictures are real or fake;
[0088] Among them, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as:
[0089]
[0090] In the formula, λ is a constant, denotes the mathematical expectation, denotes the gradient, D is a probability function, which judges the probability that the input parameter is true, and the result is in the interval [0,1]; denotes the L2 norm of the gradient; x is the real data, The data generated by the generator, refers to the distribution of real data, refers to the model distribution implicitly defined by ; Let x r and x g be a pair of real and fake samples randomly sampled from and respectively. Introduce a random number ∈ with a value in [0, 1], then refers to linearly randomly interpolated sampling between x r and x g , then is the distribution satisfied by the sampling process obtained.
[0091] Embodiment 3
[0092] This embodiment provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the data enhancement method for oral CT images described in Embodiment 1.
[0093] Embodiment 4
[0094] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data enhancement method for oral CT images described in Embodiment 1.
[0095] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for data enhancement of oral CT images, characterized in that, it includes the following steps: S1. Enhance the data based on a generative adversarial network: Input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model; S2. Use an interpretable control discovery algorithm in the generative adversarial network space to obtain a feature control vector: Use the generative adversarial network space algorithm to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector; The generative adversarial network includes a first generator, a second generator, a discriminator, and a mapping network; The first generator divides the original oral CT image into N×N units, and designs an objective function to train the rectangular annotation of the ROI area on the oral CT image, and obtains an objective recognition model for the ROI area of the maxillary sinus in the oral CT image, which is used to output the coordinates of the ROI area of the maxillary sinus in the oral CT image; The second generator cooperates with the mapping network, the mapping network decouples the latent space, generates an intermediate latent variable w from the hidden variable z, adds the affine transformation and random noise converted from the variable w to each layer of the second generator, to achieve the purpose of controlling the features of each scale of the generated image and realize controllable image generation; The discriminator receives the input of the coordinate information from the first generator and the fake image information generated by the second generator, and judges the authenticity of the image generated by the second generator; S3. Use the feature control vector to directionally generate oral CT images; Project the real or generated oral CT image into the latent space of the second generator. The latent space is a representation of compressed data, where similar data points are closer in space. Move with the set weight parameters on the feature control vector to generate the required type of high-resolution oral CT image data.
2. The method for data enhancement of oral CT images according to claim 1, characterized in that, the specific steps of step S1 include: S11. Input the original oral CT image into the generative adversarial network. A part of the original oral CT image has a rectangular annotation for the ROI area of the maxillary sinus, and the other part does not; After being processed by the first generator, the unannotated oral CT image gets a rectangular annotation centered on the ROI area of the maxillary sinus with the same size; S12. After the mapping network normalizes the hidden variable z, it passes through multiple fully connected layers to obtain an intermediate hidden variable w, and cooperates with random noise as the input of the initial and each convolutional layer of the second generator. The AdaIN algorithm is used in the input process to generate controllable images; S13. The outputs of the first generator and the second generator are jointly used as the input of the discriminator to judge whether the input picture is real or fake.
3. The method for data enhancement of oral CT images according to claim 2, characterized in that, the objective function of the first generator is: The output of the first generator is a 1×5 vector, i.e., the predicted where p ij is 1 when there is a target in the grid, q ij is 0, p ij is 0 when there is no target in the grid, q ij is 1; x i and y i are the true bounding box coordinates, and are the predicted bounding box coordinates; w i and h i are the width and height of the true bounding box, and are the width and height of the predicted bounding box; c i is the confidence value, is the intersection value between the predicted bounding box and the true bounding box, i.e., the area of the intersection of the predicted bounding box and the true bounding box divided by the area of the union.
4. The method for data enhancement of oral CT images according to claim 3, characterized in that, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as follows: where λ is a constant, denotes the mathematical expectation, denotes the gradient, D is the probability function, which is the probability of the input parameter being true, and the result is in the interval [0, 1]; denotes the L2 norm of the gradient; x is the real data, is the data generated by the generator, denotes the distribution of the real data, denotes the model distribution implicitly defined by r Let x g and x and be a pair of true and false samples randomly sampled from denotes r between x g and x, linearly random interpolation sampling, then is the distribution satisfied by the 5. A data enhancement system for oral CT images, characterized in that, it includes: An acquisition module: used to input the original oral CT image data into the generative adversarial network, and train the original oral CT image data through the generative adversarial network to obtain an oral CT image data generation model; A processing module: used to perform principal component analysis on the latent space of the oral CT image data generation model obtained in step S1 in the Euclidean space using the generative adversarial network space algorithm, unsupervised identify the latent feature control vector, and manually screen and save the latent feature control vector; the generative adversarial network includes a first generator, a second generator, a discriminator and a mapping network; the first generator divides the original oral CT image into N×N units, and designs an objective function to train the rectangular annotation of the ROI area on the oral CT image to obtain an oral CT image maxillary sinus ROI area target recognition model for outputting the coordinates of the oral CT image maxillary sinus ROI area; the second generator cooperates with the mapping network, the mapping network decouples the latent space, generates an intermediate latent variable w from the hidden variable z, and adds the affine transformation and random noise converted from the variable w to each layer of the second generator to achieve the purpose of controlling the features of each scale of the generated image and realizing controllable image generation; the discriminator accepts the input of the coordinate information from the first generator and the fake image information generated by the second generator, and judges the authenticity of the image generated by the second generator; A generation module: used to generate oral CT images directionally using the feature control vector; project the real or generated oral CT image onto the latent space of the second generator, the latent space is a representation of compressed data, where similar data points are closer in space, and move with the set weight parameters on the feature control vector to generate the required type of high-resolution oral CT image data.
6. The data enhancement system for oral CT images according to claim 5, characterized in that, the acquisition module includes: A first generator unit: used to input the original oral CT image into the generative adversarial network, a part of the original oral CT image has a rectangular annotation on the maxillary sinus ROI area, and the other part does not; after being processed by the first generator, the unannotated oral CT image obtains a rectangular annotation with the maxillary sinus ROI area as the center and the same size; among them, the objective function of the first generator is: The output of the first generator is a 1×5 vector, i.e., the predicted where p ij is 1 when there is a target in the grid, q ij is 0, and p ij is 0 when there is no target in the grid, q ij is 1; x i and y i are the true bounding box coordinates, and are the predicted bounding box coordinates; w i and h i are the width and height of the true bounding box, and are the width and height of the predicted bounding box; c i is the confidence value, is the intersection value between the predicted bounding box and the true bounding box, i.e., the area of the intersection of the predicted bounding box and the true bounding box divided by the area of the union; A mapping network and a second generator unit: used for the mapping network to normalize the hidden variable z, and obtain the intermediate hidden variable w through multiple fully connected layers, and cooperate with random noise as the input of the initial and each convolutional layer of the second generator, and the AdaIN algorithm is used in the input process to generate controllable images; Discrimination unit: It is used to jointly use the outputs of the first generator and the second generator as the input of the discriminator to discriminate whether the input picture is real or fake; Among them, both the second generator and the discriminator adopt the WGAN-GP objective function, which is expressed as: where λ is a constant, denotes the mathematical expectation, denotes the gradient, D is the probability function, which is the probability of judging that the input parameter is true, and the result is in the interval [0, 1]; denotes the L2 norm of the gradient; x is the real data, is the data generated by the generator, denotes the distribution of the real data, denotes the model distribution implicitly defined by r Let x g and x and be a pair of true and false samples randomly sampled from denotes r between x g and x is the distribution satisfied by the sampling process.
7. An electronic device, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the data enhancement method for oral CT images according to any one of claims 1 to 4.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by the processor, it implements the data enhancement method for oral CT images according to any one of claims 1 to 4.
Citation Information
Patent Citations
Face continuous feature generation method based on information maximization generation antagonism network
CN108985464A
Method for generating medical ultrasonic image data based on adversarial network
CN111724344A