Human skull CT image preprocessing method and device based on deep learning
Through deep learning WSRGAN network and micro CT image registration technology, the problem of insufficient resolution and signal-to-noise ratio of human skull CT images is solved, and high-precision skull CT image reconstruction and denoising are achieved, meeting the research needs of paleontologists.
Patent Information
- Application Number
- CN202510792676.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to obtain high-precision, high signal-to-noise ratio human skull CT data, resulting in insufficient accuracy in the identification of ancient human skulls.
A deep learning-based method is used to super-resolution reconstruction and denoising of CT images through generative adversarial network (WSRGAN), and a high signal-to-noise ratio skull CT images are generated.
The resolution and signal-to-noise ratio of human skull CT images are improved, meeting paleontologists' observation needs for the three-dimensional structure of the skull, and the denoising effect is achieved without the need for a noiseless data set.
Smart Images

Figure CN120339350A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of human skull CT image preprocessing, and specifically to a method and device for preprocessing human skull CT images based on deep learning. Background Art
[0002] The human skull is a protective structure for important organs such as the brain, eyes, ears, and respiratory tract. It is not only an important object of morphological research but also of great significance in forensic medicine, paleoanthropology, anatomy, and clinical medicine. The human skull consists of 23 bones, divided into cranial vault bones (skullcap) and facial bones. The cranial vault bones include: frontal bone, parietal bone, occipital bone, temporal bone, sphenoid bone, and ethmoid bone. The facial bones include: maxilla, mandible, nasal bone, zygomatic bone, and lacrimal bone. The study of the human skull can not only help us better understand the evolutionary history of humans but also provide valuable scientific data for related fields and promote the development of multiple disciplines.
[0003] In the identification work of human skulls, although identification based on evidence such as DNA or other molecules has made great progress, in fields such as biology, forensic medicine, archaeology, and paleoanthropology, morphological identification is still the basic means of research and identification. Especially in cases where DNA or other molecular evidence cannot be directly obtained, morphological features can often provide important clues. Traditional skull identification usually relies on directly observing and manually measuring the shape, size, and features of the skull. For example, forensic doctors usually analyze the size, shape, and bone features of the skull through visual inspection, ruler measurement, and measuring tools (such as calipers and protractors). These measurement data are used to infer characteristics such as the gender and age of an individual. In some cases, the surface information of the skull may be missing due to time, environmental factors, or damage. In such cases, traditional skull identification methods may not provide enough information for accurate identification of identity, gender, age, etc. CT-based skull identification can effectively address this challenge and provide more comprehensive and accurate analysis.
[0004] In summary, how to obtain high-precision and high-signal-to-noise-ratio human skull CT data for more accurate identification of human, especially ancient human skulls, is a technical problem to be solved by scientific research personnel. Summary of the Invention
[0005] In order to overcome the deficiencies of traditional methods for preprocessing human skull CT images, the present invention provides a method and device for preprocessing human skull CT images based on deep learning, which can efficiently and objectively perform functions such as registration, super-resolution reconstruction, and noise removal on human skull CT images to solve problems such as low resolution and low signal-to-noise ratio of human skull CT images.
[0006] To achieve the above object, the technical solution of the present invention includes the following content.
[0007] A preprocessing method for human skull CT images based on deep learning, the method comprising: For training samples, obtaining the overall CT image of the human skull and the CT image of the region of interest in the human skull; wherein, the overall CT image is a low-definition CT image, and the CT image of the region of interest is a high-definition CT image; Respectively based on the overall CT image and the CT image of the region of interest, generating an overall three-dimensional model of the human skull and a local three-dimensional model of the region of interest; Performing registration of the local three-dimensional model of the region of interest on the overall three-dimensional model of the human skull to obtain a general CT data - micro-CT data image pair of the region of interest; Based on the general CT data - micro-CT data image pair of the region of interest, training a CT image super-resolution reconstruction model, the network structure of the CT image super-resolution reconstruction model being a weight-normalizable generative adversarial network; Combining the trained CT image super-resolution reconstruction model to generate a high signal-to-noise ratio human skull CT image of the test sample.
[0008] Further, for training samples, obtaining the overall CT image of the human skull and the CT image of the region of interest in the human skull includes: Placing the training sample on a positioning sample stage; Using a general industrial CT device to perform CT scanning on the training sample located on the positioning sample stage to obtain the overall CT image of the human skull; Calculating the relative position relationship between the center position of the field of view and the center point position of the region of interest in the overall CT image; Adjusting the spatial position of the training sample according to the relative position relationship so that the center position of the field of view and the center point position of the region of interest coincide; Using a micro industrial CT device to perform CT scanning on the training sample with adjusted space to obtain the CT image of the region of interest in the human skull.
[0009] Further, the positioning sample stage includes: a chassis (1), a tray (2), an X-axis direction platform (3) and a Y-axis direction platform (4); wherein, The top surface of the chassis (1) is fixedly connected to the bottom surface of the tray (2); The top surface of the tray (2) and the bottom surface of the X-axis direction platform (3) are slidably connected in the X-axis direction through a first adjustment structure (5), wherein the first adjustment structure (5) includes a first micrometer for displaying the sliding distance between the tray (2) and the X-axis direction platform (3); The top surface of the X-axis direction platform (3) and the bottom surface of the Y-axis direction platform (4) are slidably connected in the Y-axis direction through a second adjustment structure (6), wherein the second adjustment structure (6) includes a second micrometer for displaying the sliding distance between the X-axis direction platform (3) and the Y-axis direction platform (4).
[0010] Further, based on the overall CT image, a three-dimensional model of the entire human skull is generated, including: Performing semantic segmentation on the overall CT image to obtain the semantic region of the human skull; Performing isosurface extraction and surface volume rendering on the semantic segmentation result to generate a three-dimensional model of the entire human skull.
[0011] Further, the registration of the three-dimensional model of the local region of interest on the three-dimensional model of the entire human skull to obtain a pair of general CT data - micro-CT data images of the region of interest includes: After converting the three-dimensional model of the entire human skull and the three-dimensional model of the local region of interest into point cloud data respectively, uniformly sampling the point cloud data to obtain the overall point cloud and the local point cloud; Saving the overall point cloud and the local point cloud in ply format, visualizing them using Open3D or CloudCompare, and cropping the overall point cloud according to the manually set ROI scanning frame to obtain a local subset; Aligning the center point of the local subset point cloud with the center point of the local point cloud; For the local subset point cloud and the local point cloud, performing rough registration through the fast point feature histogram algorithm and the random sample consensus key point matching algorithm; Performing fine adjustment based on ICP rigid registration on the rough registration result to obtain the point cloud registration result; According to the point cloud registration result, obtaining a pair of general CT data - micro-CT data images of the region of interest.
[0012] Further, the backbone network of the CT image super-resolution reconstruction model is a WSRGAN network, and the WSRGAN network includes: a generation network and an adversarial network. The generation network consists of an encoding structure and a decoding structure. The encoding structure includes several encoding modules, and each encoding module includes: a residual block, a convolutional layer, a model weight normalization layer, a ReLU activation layer, a channel attention block, a spatial attention block, and an upsampling function; the decoding structure includes several decoding modules. The adversarial network is established based on the VGG architecture and includes a VGG19 network model, a LeakyReLU activation function, a convolutional layer, and a model weight normalization layer structure.
[0013] Further, based on the general CT data - micro - CT data image pair of the region of interest, train a CT image super - resolution reconstruction model, including: Construct an experimental data set based on the general CT data - micro - CT data image pair of the region of interest, and divide all the images in the experimental data set into several batches; Load the original weight file of the pre - trained CT image super - resolution reconstruction model; Input the general CT data of the region of interest into the generation network, so that the encoding module extracts feature vectors from the general CT data of the region of interest and generates an output vector to be input into the next encoding module, and the encoding module extracts feature vectors from the output vector; wherein, the feature vectors output by the encoding module in the next stage all obtain higher - dimensional image feature information than the feature vectors output in the previous stage; Input the final feature vectors extracted by the encoding structure into the decoding structure, and the last decoding module generates the super - resolution image corresponding to the general CT data of the region of interest. Based on the Euclidean distance of the feature vectors between the general CT data of the region of interest and the super - resolution image corresponding to the general CT data of the region of interest, obtain the perceptual loss; Input the general CT data of the region of interest into the generation network, obtain the reconstruction result of the super - resolution network model, input it together with the micro - CT data image into the adversarial network, and based on the discriminator, discriminate between the micro - CT data image and the reconstruction result of the super - resolution network model to obtain the adversarial loss; Adopt a variable step size And the stochastic gradient descent algorithm to update the weights of each batch , and use the LRS strategy to restart the optimization of the weights to obtain the optimal weight file of the WSRGAN network; wherein, represents the batch, represents the number of training epochs; By minimizing the total change amount of the reconstruction result of the super - resolution network model and the micro - CT data image in the horizontal and vertical direction gradient operators at the pixel level, obtain the total variation loss; Calculate the pixel - level distance between the reconstruction result of the super - resolution network model and the micro - CT data image based on the mean - square error loss to obtain the content loss; Train and optimize the WSRGAN network according to the perceptual loss, adversarial loss, total variation loss and content loss, and then obtain the CT image super - resolution reconstruction model.
[0014] Further, combined with the trained CT image super - resolution reconstruction model, generate high - signal - to - noise human skull CT images of test samples, including: Obtain the overall CT image of the test sample; Input the overall CT data of the human skull of the test sample into the trained CT image super-resolution reconstruction model to obtain a super-resolution image of the CT image; Input the super-resolution image of the CT image into the trained denoising network model to obtain a human skull CT image with high signal-to-noise ratio.
[0015] Further, inputting the super-resolution image of the CT image into the trained denoising network model to obtain a human skull CT image with high signal-to-noise ratio includes: Construct a denoising network model. The backbone network of the denoising network model is an improved Noise to Noise network. The improved Noise to Noise network includes: a downsampling part, a skip connection part, and an upsampling part; wherein, the feature map sizes and channel numbers of the same stages of the downsampling part and the upsampling part are the same, and the downsampling part introduces a residual module to solve gradient disappearance and network degradation, and the skip connection part introduces an attention mechanism to fuse different-level multi-scale feature maps extracted by the convolutional neural network; Train the improved Noise to Noise network by using human skull CT images with different levels of Gaussian random noise with a mean of zero as input data and label data respectively, and monitor the training process of the improved Noise to Noise network through the signal-to-noise ratio of the denoised image to obtain the denoising loss ; According to the denoising loss Perform the training and optimization of the improved Noise to Noise network, and then obtain the denoising network model; Input the super-resolution image of the CT image into the denoising network model to obtain a human skull CT image with high signal-to-noise ratio.
[0016] A device for preprocessing human skull CT images based on deep learning. The device includes: Obtain the overall CT image of the human skull and the CT image of the region of interest in the human skull; wherein, the overall CT image is a low-definition CT image, and the CT image of the region of interest is a high-definition CT image; respectively generate an overall three-dimensional model of the human skull and a local three-dimensional model of the region of interest based on the overall CT image and the CT image of the region of interest; perform registration of the local three-dimensional model of the region of interest on the overall three-dimensional model of the human skull to obtain a pair of general CT data - micro-CT data images of the region of interest; train a CT image super-resolution reconstruction model based on the pair of general CT data - micro-CT data images of the region of interest, and the network structure of the CT image super-resolution reconstruction model is a weight-normalizable generative adversarial network; Used to combine the trained CT image super-resolution reconstruction model to generate high signal-to-noise ratio human skull CT images for test samples.
[0017] A computer device, characterized in that the computer device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the preprocessing method for human skull CT images based on deep learning described in any one of the above is implemented.
[0018] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the preprocessing method for human skull CT images based on deep learning described in any one of the above is implemented.
[0019] Compared with the prior art, the present invention has the following advantages: 1) Develop a special positioning sample stage for registering general CT images and micro-CT images, accurately position the region of interest, such as human teeth, to meet the research needs of paleontologists to quickly and accurately obtain high-resolution local images of the human skull; 2) Apply deep learning technology to carry out super-resolution reconstruction of human skull CT images, improve the spatial resolution and density resolution of human skull CT images, and meet the research needs of paleontologists for fine observation of the three-dimensional microscopic structure of the human skull; 3) The traditional method of removing noise from human skull CT images requires experienced staff to repeatedly adjust parameters to maximize the performance of the algorithm. The CT image denoising method based on deep learning adopted by the present invention can maximize the performance of the deep learning method without constructing a noise-free data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of the preprocessing method for human skull CT images provided by an embodiment of the present invention.
[0021] Figure 2 is a general CT image of the whole human skull in an embodiment of the present invention.
[0022] Figure 3 is a schematic diagram of positioning the region of interest of the human skull in an embodiment of the present invention.
[0023] Figure 4 is a design drawing of the positioning sample stage for the human skull in an embodiment of the present invention.
[0024] Figure 5 is a micro-CT image of a local part of the human skull in an embodiment of the present invention.
[0025] Figure 6 is an image after registration of the human skull in an embodiment of the present invention.
[0026] Figure 7 It is a flow chart showing the creation process of a super-resolution reconstruction network model for human skull CT images based on deep learning provided by an embodiment of the present invention.
[0027] Figure 8 It is a flow chart showing the creation process of a denoising network model for human skull CT images based on deep learning provided by an embodiment of the present invention. Detailed implementation manners
[0028] To make the objectives, solutions and advantages of the present invention clearer, taking the experiments conducted on a real dataset as an example, the present invention will be further described in detail. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0029] The preprocessing method for human skull CT images of the present invention includes: obtaining CT images of human skulls, generating three-dimensional models of human skulls, registering three-dimensional models of human skulls, super-resolution reconstruction of human skull CT images based on deep learning, and denoising of human skull CT images based on deep learning.
[0030] Specifically, as Figure 1 shown, the present invention has the following steps: Step 1: Obtain CT images of human skulls.
[0031] According to the human skull CT image acquisition standard of the present invention, a general-purpose CT is used to obtain the overall CT image of the human skull. The format of the overall CT image of the human skull can be a DICOM-format overall CT image of the human skull, or a non-DICOM ordinary image in BMP / TIF format of the overall CT image of the human skull.
[0032] In one example, a 450kV general-purpose industrial CT is used to obtain the overall CT image of the human skull. Among them, the parameter settings of the 450kV general-purpose industrial CT are: voltage: 380kV; current: 1.5mA; exposure time: 1000ms; number of superimposed frames: 2 frames; number of projection images: 1440; voxel size: 160um. As Figure 2 shown is the overall CT image of the human skull obtained by the general-purpose CT.
[0033] As Figure 3As shown, the cross - section is the exact center position of the field of view of the whole CT image of the human skull. The red frame is the region of interest of the human skull, such as human teeth. By calculating the distance from the center point of the region of interest to the center position of the image field of view, that is, the horizontal and vertical distances, the scales of the micrometers for the horizontal and vertical rotations of the positioning sample stage are obtained. In this embodiment, it is necessary to translate 5 cm in the negative horizontal direction and move 7 cm in the positive vertical direction. The positioning sample stage includes a chassis 1 combined with the turntable of the CT device. Above the chassis 1, there are three - layer structures: a tray 2, an X - axis direction platform 3, and a Y - axis direction platform 4. Among them, the top surface of the chassis 1 is fixedly connected to the bottom surface of the tray 2. The top surface of the tray 2 and the bottom surface of the X - axis direction platform 3 are slidably connected in the X - axis direction through a first adjustment structure 5. Among them, the first adjustment structure 5 includes a first micrometer that shows the sliding distance between the tray 2 and the X - axis direction platform 3. The top surface of the X - axis direction platform 3 and the bottom surface of the Y - axis direction platform 4 are slidably connected in the Y - axis direction through a second adjustment structure 6. Among them, the second adjustment structure 6 includes a second micrometer that shows the sliding distance between the X - axis direction platform 3 and the Y - axis direction platform 4. Finally, the center point of the region of interest of the human skull coincides with the center point of the field of view through the positioning sample stage. The design drawing of the positioning sample stage of the present invention is as shown in Figure 4 shown.
[0034] According to the human skull CT image acquisition standard, the present invention uses a micro - CT to obtain local CT images of the human skull. The format of the local CT image of the human skull can be a DICOM - format local CT image of the human skull or a non - DICOM ordinary image in BMP / TIF format of the human skull.
[0035] In one example, a 225kV micro - industrial CT is used to obtain local CT images of the human skull, such as human teeth. Among them, the parameter settings of the 225kV micro - industrial CT are: voltage: 140kV; current: 120μA; exposure time: 1000ms; number of superimposed frames: 2 frames; number of projection images: 1440; voxel size: 40um. As shown in Figure 5 is the CT image of the local region of interest of the human skull obtained by the micro - CT, such as human teeth.
[0036] Read the original data of the general - type CT and the original data of the micro - CT (DICOM / TIFF / RAW), normalize the gray - scale values to (0 - 65535). For the general - type CT (the HU value from - 1000 - 3000 is converted to 0 - 65535) and the micro - CT (the original X - ray absorption value is normalized and converted to 0 - 65535), and uniformly save them as 16 - bit TIFF files. The unification of the file format and the gray - scale value is convenient for subsequent 3D reconstruction, registration, and analysis.
[0037] Step 2: Generate a three-dimensional model of the human skull based on the CT image of the human skull.
[0038] Step 2.1: Perform semantic segmentation on the CT image of the human skull to obtain the semantic region of the human skull.
[0039] In this embodiment, the Otsu method is used to confirm the within-class variance between the rigid skull structure and the background in the CT image of the human skull, and the optimal threshold is determined according to the maximum value of the traversed within-class variance, and the binary image threshold segmentation of the overall CT image is performed. The optimal threshold can divide the gray histogram of the overall CT image into two categories (rigid skull structure and background), making the variance between the two categories the largest. The gray value of the overall CT image in this embodiment is in the range of 0 - 65535 levels, which is a 16-bit gray-scale image. Generally, when the gray distribution histogram of the overall CT image is bimodal, the optimal threshold T should fall at the position of the trough between the two peaks. The calculation formula for the optimal threshold of the overall CT image is as follows: Where, is the threshold; is the between-class variance when the threshold is ; , are the probabilities of the human skull and the background respectively; and are the average gray values of the human skull and the background respectively; is the gray value of the overall CT image.
[0040] Apply the optimal threshold to perform binary image segmentation on the overall CT image and select the human skull as the region of interest. Using the morphological operation method, external small particles and internal small gaps in the binary CT image of the human skull are removed through operations such as erosion and dilation, and the main structures in the CT image of the human skull are retained, including the cranial vault bone and facial bones, etc. The region growing algorithm is used to combine adjacent pixels with similar attributes by selecting the seed points of the human skull and combining with the manual selection method to gather pixels with similar properties to form a semantic region, realizing the semi-automatic semantic segmentation of the CT image of the human skull.
[0041] Step 2.2: Perform isosurface extraction and surface volume rendering on the semantic segmentation result to generate a three-dimensional model of the human skull.
[0042] The present invention applies a surface volume rendering method for human skull CT images based on marching cubes, extracts isosurfaces and renders surface volume for the semantic segmentation results of human skull CT images, and generates an isothreshold triangular facet model of human skull CT images. By finding the voxels intersecting with the isosurface in the three-dimensional volume data, the intersection points of the isosurface and the edges of the intersecting voxels are solved by means of linear interpolation, and finally all the intersection points are connected to extract the isosurface, generating the isothreshold triangular facet model of the CT image. The marching cubes algorithm can extract the voxels with the same gray threshold in the human skull CT image stack and connect them in a topological form to form human skull triangular facets.
[0043] In one example, the visualization toolkit VTK is used to provide a support environment for generating the three-dimensional model of the human skull. The present invention uses the vtkMarchingCubes class in VTK to implement the surface volume rendering of human skull CT images. The output three-dimensional model file of the human skull is usually in formats such as STL, OBJ, PLY, etc. The cross-platform application tool Blender is used to optimize the model, improve the rendering quality and visualization effect of the three-dimensional model of the human skull. The three-dimensional model of the human skull is composed of voxels containing color and measurement value information. The appearance of the three-dimensional model can be changed arbitrarily (such as color, transparency, etc.), and it can be used for research work on three-dimensional data visualization. At present, the STL file, as one of the most popular three-dimensional model format files for sharing, can be read by three-dimensional data visualization software such as Meshlab or ImageJ, and can also be applied to digital-physical tasks such as 3D printing. As Figure 2 and Figure 5 shown in the lower right corner, they are the three-dimensional models of the whole human skull and the local region of interest respectively.
[0044] Step 3: Register the general CT image and the micro-CT image.
[0045] Register the three-dimensional model of the local subset of the complete human skull of the general CT and the three-dimensional model of the small-range region of interest of the micro-CT to obtain the registered pair of general CT data and micro-CT data images of the human skull.
[0046] Read the 3D model mesh file of the human skull, such as ply / stl / obj file, and convert it into point cloud data. Calculate the normal vectors of the human skull point cloud data, which is helpful for subsequent point cloud registration. Uniformly sample the human skull point cloud data. By default, 100,000 points are sampled from the point cloud, and save the point cloud data in ply format for visualization using Open3D or CloudCompare. Crop out a local subset of the overall mesh file of the human skull according to the manually set ROI scanning frame. The overall point cloud (global_pcd) represents the point cloud data of the local subset of the complete human skull in general-purpose CT. The local point cloud (local_pcd) represents the data of the small range of the region of interest in micro-CT that needs to be registered. Calculate the center of the point cloud of the local subset of the human skull. The center (centroid) of the point cloud is defined as the average value of all point coordinates: where: is each point in the point cloud; N is the total number of points; when calculating, the average values are taken for the x, y, and z coordinates respectively: , , , that is: , , where represents the point cloud data of the local subset of the complete human skull in general-purpose CT, is the center point of the point cloud of the local subset of the complete human skull in general-purpose CT, is the average value function, represents the point cloud data of the small range of the region of interest in micro-CT, is the center point of the point cloud of the small range of the region of interest in micro-CT.
[0047] Calculate the translation vector to align the center of the point cloud of the small range of the region of interest in micro-CT to the point cloud of the local subset of the complete human skull in general-purpose CT: T = - , where: T = (T x , T y , T z ) is the translation vector.
[0048] Translate the point cloud to offset all points of the point cloud of the small range of the region of interest in micro-CT by the translation vector T: This means that all points of the point cloud of the small range of the region of interest in micro-CT will be offset by (T x , T y , T zTranslate the general CT's complete human skull local subset point cloud's center and the micro-CT's small range region of interest point cloud's center to coincide by translational movement in a certain direction. If the initial pose deviation between the two is large, or the noise is large, or the point cloud densities are different, direct ICP may fall into a local optimal solution, and it is necessary to perform rough registration through FPFH (Fast Point Feature Histogram) + RANSAC (Random Sample Consensus) key point matching. The main steps are as follows: Downsample the point cloud to reduce the point cloud density and accelerate the calculation; Calculate the FPFH key point features and calculate the local geometric features of each point; Perform rough registration with RANSAC to perform global registration through RANSAC to avoid local optimality. Use ICP rigid registration for fine adjustment, and calculate the optimal rotation matrix and translation matrix by continuously iterating the closest point matching to best align the two point clouds, obtain the registration transformation matrix T (4×4), and use the registration transformation matrix T to transform the CT data. As Figure 6 Shown are the CT images and 3D models after registering the overall 3D model of the human skull and the 3D model of the local region of interest. Crop the coverage range of the local CT data of the human skull in the micro-CT from the overall CT data of the human skull in the general CT, and match it into a pair of CT images of the local human skull in the general CT and the region of interest of the human skull in the micro-CT.
[0049] Step 4: Based on the CT image super-resolution reconstruction model, obtain the super-resolution image of the CT image.
[0050] The present invention uses local high-resolution CT data of the human skull to assist in reconstructing the high-resolution CT data of the overall human skull.
[0051] Directly input the overall CT image of the human skull in the general CT into the trained deep learning-based human skull CT image super-resolution reconstruction WSRGAN network model, that is, the trained generation network, to obtain the overall CT image of the human skull with high resolution, realize the super-resolution reconstruction of the overall CT image of the human skull, and output the super-resolution image of the overall CT image of the human skull.
[0052] The super-resolution reconstruction WSRGAN network model uses a generative adversarial network with weight standardization. The generation network part is used to generate the super-resolution overall CT image of the human skull, and the adversarial network part judges the authenticity of the image generated by the generation network and feedbacks it to the generation network. Weight normalization layers are added to both the generation network and the adversarial network parts. The definition of the weight normalization layer WN is as follows: Among them, the vector represents the Euclidean norm of , both are K-dimensional vectors. is a scalar, W is the next state of V. The introduction of WN can achieve better training and testing accuracy.
[0053] The super-resolution reconstruction WSRGAN network model adopts a dual attention mechanism, that is, it combines a channel attention mechanism and a spatial attention mechanism. The channel attention mechanism can adaptively adjust the weights of channels, effectively improving the reconstruction metrics of super-resolution. The spatial attention mechanism is an effective complement to the channel attention mechanism. Introducing the spatial attention mechanism can make the feature extraction focus on the details in the high-frequency region, while suppressing the weights in the low-frequency region, which helps to reconstruct the high-frequency details in the image, especially structures such as cranial sutures, osteons, and cancellous bone in the human skull.
[0054] In one example, when training the super-resolution reconstruction network model, an experimental dataset is constructed by inputting multiple pairs of high-definition (2048×2048)-low-definition (512×512) CT images of human skulls, and then the WSRGAN network model for super-resolution reconstruction of human skull CT images based on deep learning is trained, verified, and tested. Load the original weight file of the pre-trained WSRGAN network model, and through the adversarial learning of high- and low-resolution human skull CT images, continuously regress to train the WSRGAN network model for super-resolution reconstruction of human skull CT images based on deep learning, adjust the reconstruction accuracy of the WSRGAN network model, and save the optimized parameter weights of the WSRGAN network model. In the loss function part, in order to reduce the adverse effects caused by feature matching errors, the present invention adds a total variation loss on the basis of the context loss function, which can make the gradient in the image change in a smaller area, making the image have a certain sharpness, and to a certain extent, it can improve the detail texture of the image, which helps to reconstruct structures such as cranial sutures, osteons, and cancellous bone in the human skull. Use a relative discriminator to judge the relative authenticity of human skull CT images and optimize the super-resolution effect of human skull CT images.
[0055] As Figure 7 shown, the creation process of the WSRGAN network model for super-resolution reconstruction of human skull CT images based on deep learning provided by the embodiment of the present invention. Input the low-resolution human skull CT image with a resolution of 512×512 in the experimental dataset into the generator based on the encoder-decoder structure. The generator has 6 modules, and each module includes a residual block, a convolutional layer, a model weight normalization layer, a ReLU activation layer, a channel attention block, a spatial attention block, and an upsampling function, etc. Each module can extract feature vectors from the low-resolution image and then generate an output vector to be input into the next module. The feature vector represents the main features of the input low-resolution image, and the decoder is responsible for upsampling the feature vector. The last decoding module can reconstruct the human skull CT image with a resolution of 2048×2048.
[0056] Load the human skull CT images with a high resolution of 2048×2048 in the experimental dataset and the super-resolution reconstruction results of the human skull CT images into the discriminator of the VGG19 network model pre-trained in the ImageNet dataset to judge the relative authenticity of the human skull CT images. The discriminator includes structures such as the VGG19 network model, the LeakyReLU activation function, the convolutional layer, and the model weight normalization layer.
[0057] During the training process, in order to save the best weights of the WSRGAN network model for super-resolution reconstruction of human skull CT images. The present invention adopts an optimization weight algorithm learning rate decay strategy LRS. LRS is a strategy that can restart the optimization weight at each epoch stage, further improving the convergence speed of the network model. In order to make the network model easier to converge to the local minimum, the present invention adopts a variable step size and the Stochastic Gradient Descent (SGD) algorithm to update the weights at each epoch stage, and at the same time divide all the images in the training set into batches. The SGD updates the weight value W of the WSRGAN network model as follows: where, is the gradient value of the i-th generation of training, is the set of original high-resolution images, is the set of high-resolution images generated by the generative network.
[0058] Use the LRS strategy to restart the optimization weight and minimize the loss function of the WSRGAN network model. The variable step size is defined as follows: where, is the i-th batch of the j-th generation, The value range of is is the number of human skull CT images in each batch, is the cycle length of the cosine function, and S is the number of iterations of the restart step.
[0059] Train the super-resolution reconstruction WSRGAN network model of human skull CT images based on deep learning, set the number of iterations to 500, the optimizer to the Adam SGD function, the batch size to 16, and the initial learning rate to . The loss function of the adversarial network model is a non-convex loss function, and the optimal super-resolution reconstruction WSRGAN network model of human skull CT images based on deep learning is obtained. The loss function includes four parts: perceptual loss, adversarial loss, content loss, and total variation loss. The perceptual loss It is established based on the VGG architecture. The VGG19 network is used to extract the feature vectors of the original high-resolution image and the reconstructed super-resolution image respectively, and the Euclidean distance between the two is calculated to focus on the perceptual information of the image. Adversarial loss The discriminator is used to identify the original high-resolution image and the reconstructed super-resolution image, and the generator is constrained to improve the visual effect of the reconstructed image. Total variation loss It can make the gradient in the image change in a smaller area, generate a certain sharpness, and improve the detail texture of the image to a certain extent. Content loss During the image super-resolution reconstruction process, the mean square error loss is used to calculate the pixel-level distance between the reconstructed super-resolution image and the original high-resolution image, so as to constrain the training of the generator. The definition formula of the loss function L for the final model training is as follows: Among them, is the perceptual loss part, is the adversarial loss part, is the total variation loss, is the content loss part.
[0060] Step 5: Based on the denoising network model, obtain the high signal-to-noise ratio CT image corresponding to the super-resolution image.
[0061] In the present invention, the whole CT image of the human skull to be denoised is input into the trained Noise to Noise network model for denoising of the human skull CT image based on deep learning to obtain a high signal-to-noise ratio human skull CT image, realizing the denoising process of the human skull CT image and outputting a high signal-to-noise ratio human skull CT image.
[0062] An improved Noise to Noise network model for processing the denoising problem of human skull CT images mainly includes the following three parts: downsampling with a residual module, skip connection with an attention mechanism module, and upsampling with a transposed convolution module. In the improved Noise to Noise network model, the feature map sizes and channel numbers in the same stages of the downsampling and upsampling parts are the same. When the feature map is input to the next stage of the downsampling part, the size is half of the original size, and the channel number is twice the original scale. At the same time, a deep learning-based image denoising method for human skull CT images introduces a residual module in the downsampling part of the original Noise to Noise network model. The residual module solves problems such as gradient disappearance and network degradation caused by the deepening of the traditional convolutional neural network layers, making it easier to obtain a network model with excellent performance. An attention mechanism module is introduced in the skip connection part of the original Noise to Noise network model. The attention mechanism module can fuse different-level multi-scale feature maps extracted by the convolutional neural network, which helps to remove different types of noise in human skull CT images, such as streak artifacts, salt-and-pepper noise, and Gaussian noise. The network model is trained, validated, and tested by inputting human skull CT images with different levels of Gaussian random noise with a mean of zero as input data and label data respectively. Since the loss function loss does not gradually decrease with iteration, but the Noise to Noise network model still converges, the present invention monitors the network training process through the signal-to-noise ratio (SNR) of the denoised image. Load the original weight file of the pre-trained Noise to Noise network model, compare the human skull CT images with different levels of random noise added, continuously regress, train the Noise to Noise network model for denoising human skull CT images based on deep learning, adjust and improve the signal-to-noise ratio of the denoised human skull CT images, and save the parameter weights of the optimized improved Noise to Noise network model.
[0063] In one example, the improved Noise to Noise network model uses the ResNet34 network model as the backbone network for downsampling. The downsampling part consists of four stages, each stage is composed of a number of residual modules based on the ResNet34 network model. The first stage contains 3 residual modules, the second stage contains 4 residual modules, the third stage contains 6 residual modules, and the fourth stage contains 3 residual modules. The skip connection part adopts a structure that combines a feature pyramid and an attention mechanism, connecting the feature maps of the same scale extracted in the same stage of downsampling encoding and upsampling decoding, which can double the number of channels of the feature maps in the upsampling stage to 128, 256, 512, and 1024 respectively. Each stage of upsampling includes a deconvolution operation with a deconvolution kernel size of 2×2, feature map connection, and two convolution operations with a convolution kernel size of 3×3, and outputs a high signal-to-noise ratio human skull CT image through a 1×1 convolutional layer.
[0064] As Figure 8 shown, the present invention provides an implementation example of a Noise to Noise network model creation process for human skull CT image denoising based on deep learning. Add human skull CT images with different levels of random noise in the input experimental dataset to the downsampling stage of the improved Noise to Noise network model, which is composed of a number of residual modules based on the pre-trained ResNet34 network model. Load the original weight file of the ResNet34 network model pre-trained in the ImageNet dataset, set the number of iterations to 100, the optimizer to the Adam function, the batch size to 16, and the initial learning rate to 1e-3, and train the Noise to Noise network model for human skull CT image denoising based on deep learning to obtain the optimal deep learning-based human skull CT image denoising network model. The loss function of the improved Noise to Noise network model is function, and the definition formula of the loss function is as follows: Among them, represents the noisy human skull CT data, represents the network parameters, represents another noisy distribution of the noisy human skull CT data.
[0065] In summary, paleontologists can only obtain low-precision whole CT data of human skulls through general industrial CT or high-precision local CT data of human skulls through micro industrial CT. In the present invention, an experimental dataset is established by collecting low-precision whole CT data of human skulls through a registered general industrial CT and high-precision local CT data of human skulls through micro CT, which is used to train, validate, and test the super-resolution reconstruction method of human skull CT images based on deep learning proposed in the present invention. For paleontologists to further observe the ancient tissue details of human skulls, high signal-to-noise ratio CT images are required. In the present invention, an experimental dataset of human skull CT images with different levels of random noise is established, which is used to train, validate, and test the denoising method of human skull CT images based on deep learning proposed in the present invention. The preprocessing method and device for human skull CT images based on deep learning developed in the present invention are used to more accurately identify human, especially ancient human skulls.
[0066] Compared with the existing preprocessing methods for human skull CT images, experiments show that: During the super-resolution reconstruction of human skull CT images, compared with the traditional bilinear interpolation and bicubic interpolation methods, as well as the EDSR, WDSR, and SRGAN network models based on deep learning, as shown in Table 1, the WSRGAN network model developed using an end-to-end architecture in the present invention improves the peak signal-to-noise ratio (PSNR) of similarity measurement (the larger the value, the higher the image quality) and the Learned Perceptual Image Patch Similarity (LPIPS) of image similarity measurement (the smaller the value, the higher the image quality). Table 1 During the denoising of human skull CT images, compared with the traditional non-local means (NLM) image denoising method, as shown in Table 2, the improved Noise to Noise network model adopted in the present invention significantly improves the signal-to-noise ratio (SNR) (the larger the value, the higher the image quality). Table 2 In an exemplary embodiment, a computer device is further provided. The computer device includes a memory and a processor. A computer program is stored in the memory and is loaded and executed by the processor to implement the above-mentioned preprocessing method for human skull CT images based on deep learning.
[0067] In an exemplary embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the preprocessing method for human skull CT images based on deep learning as described above.
[0068] In an exemplary embodiment, a computer program product is further provided. When the computer program product runs on a computer device, the computer device is caused to execute the above-mentioned preprocessing method for human skull CT images based on deep learning.
[0069] The above is only one embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A preprocessing method for human skull CT images based on deep learning, characterized in that, The method includes: For training samples, obtaining the overall CT image of the human skull and the CT image of the region of interest in the human skull; wherein, the overall CT image is a low-definition CT image, and the CT image of the region of interest is a high-definition CT image; Based on the overall CT image and the CT image of the region of interest respectively, generating an overall three-dimensional model of the human skull and a local three-dimensional model of the region of interest; Performing registration of the local three-dimensional model of the region of interest on the overall three-dimensional model of the human skull to obtain a general-purpose CT data - micro-CT data image pair of the region of interest; Based on the general-purpose CT data - micro-CT data image pair of the region of interest, training a CT image super-resolution reconstruction model, and the network structure of the CT image super-resolution reconstruction model is a weight-normalizable generative adversarial network; Combining the trained CT image super-resolution reconstruction model to generate a high signal-to-noise ratio CT image of the human skull for the test sample.
2. The method according to claim 1, wherein For training samples, obtaining the overall CT image of the human skull and the CT image of the region of interest in the human skull includes: Placing the training sample on a positioning sample stage; Using a general-purpose industrial CT device to perform CT scanning on the training sample located on the positioning sample stage to obtain the overall CT image of the human skull; Calculating the relative positional relationship between the center position of the field of view and the center point position of the region of interest in the overall CT image; Adjusting the spatial position of the training sample according to the relative positional relationship so that the center position of the field of view coincides with the center point position of the region of interest; Using a micro-industrial CT device to perform CT scanning on the training sample with adjusted space to obtain the CT image of the region of interest in the human skull.
3. The method according to claim 2, characterized in that, The positioning sample stage includes: a chassis (1), a tray (2), an X-axis direction platform (3), and a Y-axis direction platform (4); wherein, The top surface of the chassis (1) is fixedly connected to the bottom surface of the tray (2); The top surface of the tray (2) and the bottom surface of the X-axis direction platform (3) are slidably connected in the X-axis direction through a first adjustment structure (5), wherein the first adjustment structure (5) includes a first micrometer showing the sliding distance between the tray (2) and the X-axis direction platform (3); The top surface of the X-axis direction platform (3) and the bottom surface of the Y-axis direction platform (4) are slidably connected in the Y-axis direction through a second adjustment structure (6), wherein the second adjustment structure (6) includes a second micrometer showing the sliding distance between the X-axis direction platform (3) and the Y-axis direction platform (4).
4. The method according to claim 1, wherein Based on the overall CT image, generating an overall three-dimensional model of the human skull includes: Performing semantic segmentation on the overall CT image to obtain the semantic region of the human skull; Performing isosurface extraction and surface volume rendering on the semantic segmentation result to generate an overall three-dimensional model of the human skull.
5. The method according to claim 1, characterized in that, The performing registration of the local three-dimensional model of the region of interest on the overall three-dimensional model of the human skull to obtain a general-purpose CT data - micro-CT data image pair of the region of interest includes: After converting the overall three-dimensional model of the human skull and the local three-dimensional model of the region of interest into point cloud data respectively, uniformly sampling the point cloud data to obtain an overall point cloud and a local point cloud; Save the overall point cloud and local point cloud in ply format, visualize them using Open3D or CloudCompare, and crop the overall point cloud according to the manually set ROI scanning frame to obtain a local subset; Align the center point of the local subset point cloud with the center point of the local point cloud; For the local subset point cloud and the local point cloud, perform rough registration through the Fast Point Feature Histogram algorithm and the Random Sample Consensus key point matching algorithm; Based on the rough registration result, perform fine adjustment based on ICP rigid registration to obtain the point cloud registration result; According to the point cloud registration result, obtain the general CT data - micro-CT data image pair of the region of interest.
6. The method according to claim 1, wherein The backbone network of the CT image super-resolution reconstruction model is a WSRGAN network, and the WSRGAN network includes: a generation network and an adversarial network. The generation network consists of an encoding structure and a decoding structure. The encoding structure includes several encoding modules, and each encoding module includes: a residual block, a convolutional layer, a model weight normalization layer, a ReLU activation layer, a channel attention block, a spatial attention block, and an upsampling function; the decoding structure includes several decoding modules. The adversarial network is based on the VGG architecture and includes a VGG19 network model, a LeakyReLU activation function, a convolutional layer, and a model weight normalization layer structure.
7. The method according to claim 6, characterized in that, Based on the general CT data - micro-CT data image pair of the region of interest, train a CT image super-resolution reconstruction model, including: Construct an experimental dataset based on the general CT data - micro-CT data image pair of the region of interest, and divide all the images in the experimental dataset into several batches; Load the original weight file of the pre-trained CT image super-resolution reconstruction model; Input the general CT data of the region of interest into the generation network, so that the encoding module extracts feature vectors from the general CT data of the region of interest and generates an output vector to be input into the next encoding module, and the encoding module extracts feature vectors from the output vector; among them, the feature vectors output by the encoding module in the next stage have higher-dimensional image feature information than the feature vectors output in the previous stage; Input the final feature vector extracted by the encoding structure into the decoding structure, and the last decoding module generates the super-resolution image corresponding to the general CT data of the region of interest. Based on the Euclidean distance of the feature vectors between the general CT data of the region of interest and the super-resolution image corresponding to the general CT data of the region of interest, obtain the perceptual loss; Input the general CT data of the region of interest into the generation network, obtain the reconstruction result of the super-resolution network model, input it into the adversarial network together with the micro-CT data image, and based on the discriminator, discriminate between the micro-CT data image and the reconstruction result of the super-resolution network model to obtain the adversarial loss; Adopt variable step size and update the weights of each batch using the stochastic gradient descent algorithm , and use the LRS strategy to restart the optimization of the weights to obtain the optimal weight file of the WSRGAN network; where represents the batch represents the number of training epochs Minimize the total change amount of the reconstruction result of the super-resolution network model and the micro-CT data image in the pixel-level horizontal and vertical direction gradient operators to obtain the total variation loss; Calculate the pixel-level distance between the reconstruction result of the super-resolution network model and the micro-CT data image based on the mean square error loss to obtain the content loss; Train and optimize the WSRGAN network according to the perceptual loss, adversarial loss, total variation loss, and content loss, and then obtain a CT image super-resolution reconstruction model.
8. The method according to claim 1, wherein Combine the trained CT image super-resolution reconstruction model to generate high signal-to-noise ratio human skull CT images of test samples, including: Obtain the overall CT image of the test sample; Input the overall CT image of the human skull of the test sample into the trained CT image super-resolution reconstruction model to obtain a super-resolution image of the CT image; Input the super-resolution image of the CT image into the trained denoising network model to obtain a high signal-to-noise ratio human skull CT image.
9. The method according to claim 8, wherein Input the super-resolution image of the CT image into the trained denoising network model to obtain a high signal-to-noise ratio human skull CT image, including: Construct a denoising network model. The backbone network of the denoising network model is an improved Noise to Noise network. The improved Noise to Noise network includes: a downsampling part, a skip connection part, and an upsampling part; wherein, the feature map sizes and channel numbers of the same stage of the downsampling part and the upsampling part are the same, and the downsampling part introduces a residual module to solve the problem of gradient disappearance and network degradation, and the skip connection part introduces an attention mechanism to fuse different-level multi-scale feature maps extracted by the convolutional neural network; By using human skull CT images with different levels of Gaussian random noise with a mean of zero as input data and label data respectively, the improved Noise to Noise network is trained, and the training process of the improved Noise to Noise network is monitored by the signal-to-noise ratio of the denoised images to obtain the denoising loss ; According to the denoising loss Perform the training and optimization of the Noise to Noise network for the above improvement, and then obtain the denoising network model; Input the super-resolution image of the CT image into the denoising network model to obtain a high signal-to-noise ratio CT image.
10. A preprocessing device for human skull CT images based on deep learning, characterized in that, The device includes: A training module. For training samples, obtain the overall CT image of the human skull and the CT image of the region of interest in the human skull; wherein, the overall CT image is a low-resolution CT image, and the CT image of the region of interest is a high-resolution CT image; respectively generate an overall three-dimensional model of the human skull and a local three-dimensional model of the region of interest based on the overall CT image and the CT image of the region of interest; register the local three-dimensional model of the region of interest on the overall three-dimensional model of the human skull to obtain a general CT data - micro-CT data image pair of the region of interest; based on the general CT data - micro-CT data image pair of the region of interest, train a CT image super-resolution reconstruction model, and the network structure of the CT image super-resolution reconstruction model is a weight-normalizable generative adversarial network; A testing module, which is used to combine the trained CT image super-resolution reconstruction model to generate high signal-to-noise ratio human skull CT images of test samples.
Citation Information
Patent Citations
Super-resolution imaging method based on oral cavity CBCT reconstruction point cloud
CN112184556A
CBCT image reconstruction method based on deep learning and electronics noise simulation
CN114241074A
Dark field super-resolution imaging method, model evaluation method and system
CN115908126A
Micro fossil CT image preprocessing method and device based on deep learning
CN117078780A
Unmanned aerial vehicle image super-resolution reconstruction method, system and equipment
CN117893410A