An Improved U-Net Skull X-ray Denoising Method Based on Lightweight Simulation Pre-training

By constructing a lightweight simulation pre-trained improved U-Net network, utilizing depthwise separable convolution and SE channel attention mechanisms, combined with a two-stage training mechanism, the problem of scattering noise removal in skull X-ray images was solved, achieving efficient denoising and preservation of anatomical details.

CN122089595APending Publication Date: 2026-05-26ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2026-02-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for removing scattering noise from skull X-ray images often result in the loss of anatomical details due to traditional methods. General-purpose deep learning models are not good at distinguishing skull features, and the scarcity of real data makes model training difficult and generalization performance poor.

Method used

An improved U-Net network with lightweight simulation pre-training is constructed. A simulation dataset is generated through parameterized random modeling. Combining depthwise separable convolution and SE channel attention mechanism, a two-stage training mechanism is adopted for pre-training and fine-tuning to achieve denoising of skull X-ray images.

Benefits of technology

It effectively removes scattering noise while preserving key anatomical details such as cranial sutures and edges, improving the model's generalization ability and training efficiency, and ensuring the system's robustness and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089595A_ABST
    Figure CN122089595A_ABST
Patent Text Reader

Abstract

This invention relates to the field of medical image processing technology and discloses an improved U-Net skull X-ray denoising method based on lightweight simulation pre-training. The method first generates a dedicated simulation dataset for pre-training based on parametric random modeling and simplified physical projection. Then, a network integrating depthwise separable convolution and channel attention mechanisms is constructed. This network adopts a two-stage architecture with configurable channel counts and embeds channel-level attention fusion modules in skip connections. The network with half the channels is pre-trained using simulation data, and then weights are loaded and fine-tuned with real data. The trained model is used to denoise real skull X-ray images. This invention solves the problems of difficult model training under scarce real data and the difficulty in balancing noise removal and detail preservation by constructing a complete process from lightweight simulation to two-stage transfer learning, thus improving denoising performance and model practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to an improved U-Net skull X-ray denoising method based on lightweight simulation pre-training. Background Technology

[0002] In the field of medical imaging, X-ray imaging is a crucial diagnostic tool for conditions such as skull fractures and intracranial lesions due to its convenience and speed. However, phenomena such as Compton scattering generated when X-rays penetrate human tissue create scattering noise superimposed on the original signal, leading to decreased image contrast and blurred details. To address this issue, current technologies have primarily focused on both hardware and software development. Hardware correction methods, including the use of anti-scattering grids and increased air gaps, aim to physically reduce the number of scattered photons reaching the detector, but these methods have drawbacks such as increasing patient radiation dose or limiting imaging geometry. Software correction methods, on the other hand, focus on post-processing image processing. Early methods employed traditional algorithms based on statistical models or image filtering (such as Gaussian filtering and median filtering) to attempt to separate noise from mixed signals. In recent years, with the powerful capabilities of deep learning, particularly encoder-decoder architectures like U-Net, in medical image segmentation and denoising tasks, convolutional neural network-based X-ray and CT image denoising methods have become a mainstream research direction. In addition, to address the challenges of labeling medical data and the high privacy requirements, a transfer learning framework that uses simulation technology to generate data for pre-training models and then fine-tunes them with a small amount of real data has been widely used to improve the generalization ability of models in real-world scenarios.

[0003] However, the aforementioned existing technologies still have significant limitations when applied to the removal of scattering noise from skull X-ray images. Traditional filtering algorithms struggle to distinguish noise from the fine structure of the skeleton, easily leading to blurred edges and loss of detail. Existing deep learning models are mostly designed for common areas and do not fully consider the high density, complex morphology, and uneven distribution of scattering noise in skull tissue, resulting in poor denoising accuracy and detail preservation. Furthermore, models trained directly on simulated data struggle to accurately match the complex physical imaging process of reality, exhibiting limited generalization ability; while training models from scratch using only scarce real clinical data faces challenges such as slow convergence and unstable performance, failing to achieve a good balance between effectiveness and efficiency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that the existing technology has the disadvantages of traditional denoising algorithms that easily lose key anatomical details of the skull, general deep learning models that are not good at distinguishing skull scattering noise features, and model training difficulties and poor generalization performance due to the scarcity of real medical data and the stringent requirements for the fidelity of simulation data generation. To this end, we propose an improved U-Net skull X-ray denoising method with lightweight simulation pre-training.

[0005] To achieve the above objectives, this application adopts the following technical solution: a lightweight simulation pre-trained improved U-Net skull X-ray denoising method, comprising the following steps: S1, Construct and prepare a dataset of simulated and real skull X-ray images, specifically including: A three-dimensional skull model was built based on the GATE simulation platform. X-ray attenuation parameters and projection geometry were set to generate simulation image pairs containing primary clear signals and scatter noise. At the same time, real clinical skull X-ray images were acquired, and after normalization, noise annotation and data augmentation processing, a real training sample set was formed to complete the pairing and division of simulation and real data.

[0006] S2, constructing an improved U-Net denoising network integrating attention and a two-stage training mechanism, specifically including: The design introduces a feature extraction and restoration module with depthwise separable convolution and SE channel attention mechanism, constructs a two-stage training architecture that supports configurable channel number, and embeds a channel-level attention fusion mechanism between encoder and decoder to achieve adaptive weighting and multi-scale fusion of simulated features and real features.

[0007] S3 performs simulation pre-training, real-world fine-tuning, and denoised inference and evaluation, specifically including: The network is pre-trained in a lightweight manner using a simulation dataset, then the pre-trained weights are loaded and the network parameters are fine-tuned based on real data. The model is optimized through dynamic learning rate and early stopping strategy. The trained model is applied to real skull X-ray images for noise removal, outputting clear images. The denoising effect is verified based on PSNR, SSIM index and clinical visual assessment. If the input quality is abnormal or the output confidence is lower than the threshold, the process is interrupted and backup filtering is enabled, and abnormal events are recorded.

[0008] Preferably, step S1, constructing and preparing the simulation dataset and the real dataset for network pre-training, specifically includes: S11, generating a simulated skull phantom based on parametric random modeling. A geometrically simplified model is employed: an ellipsoid is used to simulate brain tissue, with its three semi-axis lengths a, b, and c randomly and independently selected within preset ranges [13,16] cm, [11,14] cm, and [9,12] cm, respectively; an elliptical shell surrounding this ellipsoid is used to simulate the skull's outer shell, with its thickness randomly selected within a preset range [6,8] mm. This modeling method does not strive for perfect replication of real skull anatomical details, but rather covers the range of structural variations common in adult skulls through randomization of key dimensions, generating diverse parametric phantoms.

[0009] S12, based on a simplified physical model and random imaging parameters, performs projection simulation to generate paired noise data. A parallel beam or conical beam model is used, with the light source photon energy randomly selected within the range of [50, 150] keV and the tube current randomly selected within the range of [100, 500] mA to simulate fluctuations in imaging conditions. The projection calculation uses a simplified physical model: by performing the above calculations on the phantom generated in S11 under multiple viewpoints, such as rotation around an axis, paired image pairs of primary clear signal and scatter noise signal are generated in batches. This process emphasizes the simulation of scattering noise characteristics rather than pixel-level accurate reconstruction of the real image.

[0010] S13. Acquire and preprocess real skull X-ray images to construct a supervised training dataset. Obtain real clinical skull X-ray DICOM images and perform preprocessing such as size normalization and grayscale standardization. Paired noisy-clear image data can be obtained through expert manual denoising or other processing methods, or relevant open-source datasets can be used directly.

[0011] Furthermore, the batch-generated paired simulation image data in S1 is a dataset specifically constructed for the network pre-training stage. When generating the simulation data, the modeling of the three-dimensional skull phantom adopts a simplified method that is parameterized and allows key geometric dimensions to be randomly set within a certain range. The physical conditions for X-ray projection imaging are set with relaxed settings, randomly taking values ​​within a preset range. The generated paired simulation image data is used as a pre-training dataset specifically for the network's pre-training stage. Simultaneously, the acquired and preprocessed real skull X-ray images are divided into a real training set and a real test set, used for the network's fine-tuning stage and performance evaluation, respectively. The constraints followed in the generation process of the simulation dataset regarding the model's geometric fidelity, material property accuracy, and completeness of physical effect simulation are more relaxed than data generation methods that require pixel-level consistency with real clinical images and are directly used for the network's final training.

[0012] Preferably, step S2, constructing an improved U-Net denoising network integrating attention and a two-stage training mechanism, specifically includes: S21. A feature extraction module and a feature restoration module are designed to address the distribution characteristics of scattering noise in skull X-ray images. The feature extraction module uses depthwise separable convolutions for spatial feature extraction and embeds a Squeeze-and-Excitation channel attention mechanism. It learns channel dependencies through global average pooling and fully connected layers, generating channel weights to recalibrate features, thereby enhancing the network's ability to distinguish between noise and skeletal details. The feature restoration module consists of 1×1 convolutions, 3×3 convolutions, and a Squeeze-and-Excitation channel attention mechanism connected sequentially, followed by batch normalization and a ReLU activation function, used to restore image details during the decoding stage.

[0013] S22, construct a two-stage training architecture that supports configurable channel count; in the pre-training stage, set the output channel count of all convolutional layers in the network to half of the target network; in the denoising stage, restore the channel count of each layer of the network to the full configuration, load pre-trained weights for the first half of the channels of each convolutional layer, and randomly initialize the second half of the channels.

[0014] S23 introduces a channel-level attention fusion mechanism in the skip connection between the encoder and decoder. This mechanism divides the feature map output by the encoder according to the channel source, performs global average pooling on the feature channels from the pre-training stage and the newly added feature channels from the fine-tuning stage, concatenates the resulting description vectors and inputs them into a shared fully connected network to generate channel attention weight vectors. After Sigmoid activation and normalization, the two parts of the features are weighted and fused. Then, the fused features are concatenated with the upsampling results of the corresponding layer of the decoder in the channel dimension. Finally, multi-scale feature integration is completed through 3×3 standard convolution.

[0015] Furthermore, in step S21, a feature extraction module and a feature restoration module are designed based on the distribution characteristics of scattering noise in the skull X-ray image: The feature extraction module uses a 3×3 depth separable convolutional layer to extract spatial features from the input noisy image to reduce the computational complexity of the model, and then connects to the Squeeze-and-Excitation channel attention submodule. The Squeeze-and-Excitation channel attention submodule compresses the spatial dimension of the feature map through a global average pooling layer to obtain a channel description vector. This vector is then passed through two fully connected layers to learn the non-linear dependencies between channels and outputs weight coefficients equal to the number of channels. These weight coefficients are then used to reweight the initial features extracted by depthwise separable convolutions in terms of channel dimension, thereby enhancing the network's ability to distinguish and select features from strong scattering noise in high-density skeletal regions and weak scattering noise in soft tissue regions. The feature restoration module is used to gradually restore image details in the decoder stage. Its structure consists of a 1×1 convolutional layer, a 3×3 convolutional layer, a Squeeze-and-Excitation channel attention submodule, a batch normalization layer, and a ReLU activation function. The 1×1 convolutional layer is used to adjust the number of channels, the 3×3 convolutional layer is used for spatial detail reconstruction, and the Squeeze-and-Excitation submodule performs importance recalibration on the reconstructed features.

[0016] Furthermore, step S22, which constructs a two-stage training architecture supporting configurable channel numbers, specifically includes: During the simulation data pre-training stage, the number of output channels of all convolutional layers in the improved U-Net network is set to half the number of corresponding channels of the target complete network, forming a lightweight network structure for training. In the real data denoising and fine-tuning stage, the number of channels in each layer of the network is restored to the complete standard configuration, and the parameters of each convolutional layer in the network are initialized: the convolutional kernel weights of the first C / 2 channels of the convolutional layer are loaded with the parameters learned in the pre-training stage, and the convolutional kernel weights of the last C / 2 channels adopt a random initialization strategy, where C is the complete number of output channels of the layer.

[0017] Furthermore, step S23 in constructing the improved U-Net denoising network integrating attention and a two-stage training mechanism specifically includes: In the skip connection path between the encoder and decoder of the U-Net network, a channel-level attention fusion mechanism is introduced; For the feature map output by a certain layer of the encoder, its channels are divided into two parts according to their source: the first half of the channels are features inherited from the pre-trained weights, and the second half of the channels are features added and initialized during the fine-tuning stage. The channel-level attention fusion mechanism first performs global average pooling on the two feature parts respectively to obtain two channel-level global description vectors. After concatenating the two description vectors, they are input into a shared two-layer fully connected network, which outputs an attention weight vector with the same number of channels as the original feature map. The attention weight vector is activated by the Sigmoid function along the channel dimension and then normalized to ensure that the sum of the weight coefficients corresponding to the first half of the channel and the second half of the channel is 1 at their respective positions. The two feature maps output by the encoder are multiplied by their corresponding normalized attention coefficients to complete adaptive weighting, and then the two weighted feature maps are concatenated along the channel dimension. Finally, the fused features obtained by splicing are spliced ​​together with the feature maps upsampled by the corresponding layer of the decoder in the channel dimension, and the information is integrated across channels and space through a 3×3 standard convolutional layer and passed to the next decoding layer. This ensures that high-frequency details of key anatomical structures such as cranial sutures and bone edges are effectively preserved while suppressing scattering noise.

[0018] Preferably, step S3 involves performing a two-stage training and evaluation based on targeted simulation pre-training and fine-tuning, specifically including: S31, network pre-training using a simulation dataset generated based on the physical features of the target data: The improved U-Net network with half the number of channels is trained by loading the simulation dataset generated in S1, which specifically simulates the physical features of skull X-ray scattering noise. The purpose of this pre-training is to enable the network to directly learn the low-level physical feature representation of skull X-ray scattering noise removal, rather than general image features. Training uses a loss function combining L1 and mean squared error to allow the network to initially establish a mapping relationship between noise and clear signals. S32, fine-tuning the network with targeted pre-trained weights using a small amount of real data: restoring the number of channels in the pre-trained network to its full configuration, and adopting a strategy of inheriting pre-trained weights for the first half of each convolutional layer and randomly initializing the second half of the channels. The network is fine-tuned using the relatively small amount of real skull X-ray training set prepared in S1. Since the network backbone has already grasped the noise-related physical characteristics through targeted simulation data, this fine-tuning process converges faster and requires less real data. It mainly focuses on adapting the network to the specific distribution details of real data, avoiding the long convergence and domain adaptation problems commonly encountered when training from scratch or transferring from an unrelated source domain. S33 applies the trained model to real skull X-ray images for denoising inference. The real image to be denoised is input into the improved U-Net model trained in S31 and S32, forward propagation is performed, and the output is a clear image with suppressed scattering noise; S34. Quantitative and qualitative evaluation of the denoising results. The peak signal-to-noise ratio and structural similarity index of the denoised image on the test set are calculated for quantitative evaluation. Simultaneously, experts conduct a qualitative evaluation of the preservation of details of key anatomical structures such as cranial sutures and edges in the denoised image, comprehensively validating the model's performance. S35 sets up an exception handling and process assurance mechanism. During inference, if abnormal input image quality or low model output confidence is detected, the standard process is automatically interrupted, switching to a traditional filtering algorithm as a backup, and the exception event is recorded to ensure system robustness.

[0019] Furthermore, in step S31, a simulation dataset generated based on the physical characteristics of the target data is used for network pre-training, specifically including: Load the simulation dataset generated by S1, which is generated by simulating the physical process of skull X-ray scattering noise and contains paired images of primary clear signal and scatter noise signal; The number of output channels of each convolutional layer in the improved U-Net network is set to half the number of channels of the corresponding layer in the full network, forming a lightweight network structure with half the number of channels. The lightweight network with halved channels was trained using the simulation dataset. During training, a weighted combination of L1 loss function and mean squared error loss was used as the optimization objective, enabling the network to learn the mapping relationship from noisy images to clear images and to initially grasp the physical characteristic distribution of skull X-ray scattering noise. The pre-training process does not require the simulated data to be consistent with the real data at the pixel level. Instead, it focuses on enabling the network to learn the common physical characteristics of scattering noise, providing targeted initial feature representations for the subsequent fine-tuning stage.

[0020] Furthermore, in step S32, a small amount of real data is used to fine-tune the network loaded with targeted pre-trained weights, specifically including: The number of channels in each layer of the lightweight network, which has been pre-trained by simulation, is restored to its full configuration to form a complete improved U-Net network architecture. For each convolutional layer in the network that has been restored to its full configuration, the parameters are initialized: for a convolutional layer with C output channels, the kernel weights of the first C / 2 channels are loaded with the parameters learned during the pre-training stage, and the kernel weights of the last C / 2 channels are randomly initialized. The complete network was fine-tuned using the relatively small amount of real skull X-ray training set prepared in S1. During the fine-tuning process, a dynamic learning rate scheduling strategy was adopted, and an early stopping mechanism was set to prevent overfitting. The fine-tuning process, based on the physical feature representation of scattering noise learned in pre-training, uses a small amount of real data to enable the network to quickly adapt to the specific distribution details of real skull X-ray data, significantly reducing the amount of real data and training time required for model convergence, and improving the model's ability to remove scattering noise from real skull X-ray images.

[0021] Furthermore, the fine-tuning in S32 also includes: The contribution weights of feature channels inherited from pre-training and newly added randomly initialized feature channels in the feature extraction and reconstruction process are dynamically adjusted through a channel-level attention fusion mechanism. During the fine-tuning training process, the channel-level attention fusion mechanism learns adaptive attention coefficients for both inherited feature channels and newly initialized feature channels, and achieves effective integration of the two types of features through weighted fusion. This mechanism ensures that the network can retain the basic physical characteristics of scattering noise learned from simulation data during the fine-tuning phase, while also flexibly learning the noise patterns and detailed features unique to real data.

[0022] The technical effects and advantages of this invention are as follows: This invention constructs a two-stage training architecture that supports configurable channel count. In the pre-training stage, a lightweight network with half the channels is used to learn the physical characteristics of scattering noise. In the fine-tuning stage, the network is restored to its full configuration and an initialization strategy that inherits the pre-training weights from the first half of the channels is adopted. This achieves targeted feature transfer and efficient convergence from simulation data to real data, solving the problems of model training difficulties, insufficient generalization ability, and unstable training process caused by the scarcity of real data.

[0023] This invention introduces a channel-level attention fusion mechanism in the skip connection between the encoder and decoder of the improved U-Net network, dynamically adjusting the contribution weights of feature channels from the pre-training and fine-tuning stages. This achieves adaptive weighting and multi-scale fusion of the two types of features, solving the problem of accurately preserving high-frequency details of key anatomical structures such as cranial sutures and edges while suppressing scattering noise.

[0024] This invention, by setting up anomaly handling and process assurance mechanisms, monitors the input image quality and output confidence during the denoising inference process, and switches to a backup filtering algorithm in case of anomalies. This achieves high robustness and continuity of system operation and solves the problems of process interruption and output reliability risks caused by input anomalies or model uncertainties. Attached Figure Description

[0025] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts: Figure 1 This is a system flowchart of the present invention; Figure 2 This is a system data flow diagram of the present invention; Figure 3 This is a diagram of the improved U-Net network structure of the present invention. Detailed Implementation

[0026] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0027] Reference Figure 1 , Figure 2 As shown, this invention provides a technical solution: a lightweight simulation pre-training improved U-Net skull X-ray denoising method, specifically including: Step S1: Build and prepare the simulation dataset and real dataset for network pre-training.

[0028] This step aims to generate a simulation dataset for network pre-training that does not pursue pixel-level fidelity but emphasizes the coverage of scattering noise features through parametric random modeling and simplified physical projection. Simultaneously, real clinical data is collected and preprocessed for fine-tuning and evaluation to address the problems of scarce real data and high simulation costs.

[0029] Sub-step S11: Generate a skull simulation model based on parametric random modeling.

[0030] This sub-step uses a geometrically simplified model to generate diverse skull phantoms. Specifically, to simulate brain tissue, a material with uniform aqueous properties, such as water, is used for modeling, and its X-ray attenuation coefficient is approximately expressed as: ; in The attenuation coefficient at standard temperature is 0.21 cm. -1 , This is a temperature-dependent constant, approximately 0.0015. The current temperature. The standard temperature is 300K.

[0031] The shape of this brain tissue is defined as an ellipsoid, and its volume is: ; The main size variations in the adult brain were covered by randomly and independently selecting half-axis a, b, and c within a reasonable range. The range is 13-16cm. It is 11-14cm. It measures 9-12 cm. Subsequently, an ellipsoidal shell surrounding this ellipsoid was used to simulate the outer shell of the skull, its volume calculated using the formula: ; Among them, the shell thickness , , Random selections were made within a 6-8 mm range, while the gap between the brain tissue and the outer shell of the skull was filled with a material whose decay properties are consistent with water to simulate cerebrospinal fluid. This parametric random modeling approach aims to cover common structural variations by randomizing key geometric dimensions, rather than striving for consistency with the anatomy of a specific patient, thereby generating diverse phantoms.

[0032] Sub-step S12: Perform projection simulation based on the simplified physical model and random imaging parameters to generate paired noise data.

[0033] This sub-step aims to generate image pairs with simulated coherent noise characteristics in batches using a simplified physical process. First, the projection geometry is set: using a parallel beam or conical beam model, the X-ray source is placed at coordinates (0, -50, 0) cm, and the receiver is placed at (0, 50, 0) cm, simulating a standard symmetrical scanning layout. To enhance data diversity, imaging parameters are randomly set within typical clinical ranges: the source photon energy is randomly selected within 50-150 keV, and the tube current is randomly selected within 100-500 mA. The projection calculation uses a simplified physical model: for each projection angle, the primary clear signal intensity is calculated using an exponential decay model. ; in For incident intensity, The attenuation coefficient is... The path length is the distance to the point of penetration; the scattering noise signal is characterized by the scattering attenuation coefficient and calculated by linear superposition. .

[0034] By performing multi-view batch projection calculations on multiple phantoms generated in sub-step S11, a large number of paired image pairs of Primary sharp signal and Scatter noise signal are finally generated. This process focuses on simulating the physical characteristics of scattering noise rather than accurately reconstructing the pixel values ​​of real images, thus forming a dataset specifically for the network pre-training stage.

[0035] Sub-step S13: Acquire and preprocess real skull X-ray images to construct a supervised training dataset.

[0036] This sub-step is responsible for building a real-world data benchmark for network fine-tuning and final evaluation. First, clinically real DICOM skull X-ray images are acquired. Then, the raw images undergo preprocessing such as size normalization and grayscale normalization to eliminate differences in equipment and protocols. A crucial step is to obtain paired noisy-to-clear images by manually denoising the preprocessed images by experts or through other processing methods; alternatively, relevant open-source datasets can be used directly. Finally, all real-world paired data are divided into training and testing sets in approximately an 8:2 ratio. The training set is used for subsequent network fine-tuning, while the testing set is used for objective performance evaluation.

[0037] Step S2: Construct an improved U-Net denoising network that integrates attention and a two-stage training mechanism.

[0038] like Figure 3As shown, this step aims to design a dedicated network architecture with strong feature discrimination capabilities and support for knowledge transfer from simulation to real data, specifically targeting the characteristics of skull X-ray scattering noise.

[0039] Sub-step S21: Design the feature extraction module and the feature restoration module.

[0040] This sub-step improves upon the basic convolutional units of traditional U-Net to enhance the ability to distinguish between noise and skeletal details. The feature extraction module is specifically designed as follows: input features are first processed through a 3×3 depthwise separable convolutional layer for spatial feature extraction; subsequently, a Squeeze-and-Excitation channel attention submodule is applied. This module compresses the feature map into channel description vectors through global average pooling, then passes them through two fully connected layers to learn the non-linear dependencies between channels and outputs the weight coefficients of each channel. This reweights the features extracted by the aforementioned depthwise separable convolution in terms of channel dimensions, thereby adaptively enhancing important features such as skeletal edges and suppressing redundant noise information. In the decoding path, the feature restoration module consists of a 1×1 convolutional layer for adjusting the number of channels, a 3×3 convolutional layer for reconstructing spatial details, an SE channel attention submodule for recalibrating feature importance, a batch normalization layer, and a ReLU activation function, progressively and accurately restoring image details.

[0041] Sub-step S22: Construct a two-stage training architecture that supports configurable channel numbers.

[0042] This sub-step designs a flexible channel configuration strategy to achieve two-stage training. In the simulated data pre-training stage, to reduce computational costs, the number of output channels for all convolutional layers in the network is set to half the number of channels for the corresponding layer in the target full network. For example, a 64-channel layer is set to 32 channels, forming a lightweight network for training. In the real data denoising and fine-tuning stage, the number of channels in each layer of the network is restored to the full configuration, and a specific parameter initialization strategy is applied to each convolutional layer: for a layer with C output channels, the kernel weights of its first C / 2 channels are directly loaded with the parameters learned in the pre-training stage, while the kernel weights of the last C / 2 channels are randomly initialized. This allows the network to simultaneously possess the basic feature representation capabilities inherited from the simulated data and the plasticity to learn new features from real data during fine-tuning.

[0043] Sub-step S23: Introduce a channel-level attention fusion mechanism in the skip connection between the encoder and decoder.

[0044] This sub-step designs an attention fusion mechanism to intelligently integrate feature channels from pre-training and newly added feature channels during fine-tuning. Specifically, in skip connections, the feature map output by the encoder is divided into two halves: the first half, representing channels inherited from pre-training weights, and the second half, representing channels added during fine-tuning. First, both halves of the feature map are subjected to global average pooling to obtain two channel-level global description vectors. These two description vectors are concatenated and fed into a shared two-layer fully connected network to generate an attention weight vector with the same number of channels as the original feature map. This weight vector is then subjected to Sigmoid activation and normalization to ensure that the sum of the attention coefficients for the two halves of the channel is 1. Next, the two halves of the encoder's feature map are multiplied by their corresponding normalized attention coefficients to achieve adaptive weighting. Finally, the weighted fused features are concatenated with the upsampled feature map of the corresponding layer in the decoder along the channel dimension, and a 3×3 standard convolutional layer is used for cross-channel and spatial information integration. This mechanism ensures that the network can dynamically adjust and effectively utilize both types of features during decoding, accurately preserving key anatomical details such as sutures and edges while suppressing noise.

[0045] Step S3: Perform a two-stage training and evaluation based on targeted simulation pre-training and fine-tuning.

[0046] This step aims to train a high-performance denoising model with a small amount of real data by first simulating pre-training and then fine-tuning in real data, and to conduct comprehensive and reliable performance verification.

[0047] Sub-step S31: Pre-train the network using the simulation dataset.

[0048] This sub-step pre-trains the lightweight network using the generated simulation data, aiming to enable it to learn the underlying physical characteristics of scattering noise. The specific training configuration is as follows: The improved U-Net network with half the number of channels is trained using the simulation dataset generated in sub-step S12. The optimizer is Adam, with an initial learning rate of 1e-4, a batch size of 16, and a total of 100 training epochs. A cosine annealing strategy is used to dynamically adjust the learning rate, for example, decreasing it to 0.5 times the current rate every 20 epochs. The loss function is a weighted combination of L1 loss and mean squared error loss. This pre-training does not aim for pixel-level consistency between the simulation data and real data, but rather focuses on enabling the network to establish a fundamental mapping relationship from noise to clear signals, mastering the common characteristics of scattering noise.

[0049] Sub-step S32: Fine-tune the network loaded with targeted pre-trained weights using a small amount of real data.

[0050] This sub-step, based on pre-training, uses a small amount of real data to quickly adapt the network. First, the number of network channels is restored to the full configuration, and the weights of each layer are initialized according to the strategy in sub-step S22, i.e., the first half of the channels inherits the pre-trained weights, and the second half of the channels are randomly initialized. Then, the complete network is fine-tuned using the real training set partitioned in sub-step S13. The fine-tuning process also employs a dynamic learning rate and an early stopping mechanism to prevent overfitting. During this process, the channel-level attention fusion mechanism described in sub-step S23 dynamically adjusts the contribution of inherited features and newly added features, enabling the network to both stably retain the noise physical priors learned from the simulation data and flexibly learn specific noise patterns and detail distributions in the real data, thereby achieving rapid convergence and efficient domain adaptation.

[0051] Sub-step S33: Apply the trained model to real skull X-ray images for denoising inference.

[0052] After the model training is completed, this sub-step performs the actual denoising task: the real skull X-ray image to be processed is input into the improved U-Net model trained by S31 and S32, and through forward propagation calculation, a clear image after the scattering noise is suppressed is directly output.

[0053] Sub-step S34: Perform quantitative and qualitative evaluation of the denoising results.

[0054] This sub-step comprehensively evaluates the model's denoising performance. Quantitatively, the peak signal-to-noise ratio (PSNR) and structural similarity index of the denoised images are calculated on a reserved real test set. Qualitatively, experts are invited to conduct a visual evaluation of the denoised images, focusing on the preservation of details and clarity of key anatomical structures such as cranial sutures and edges. The combined quantitative and qualitative results comprehensively validate the model's denoising effectiveness and practicality.

[0055] Sub-step S35: Set up exception handling and process assurance mechanisms.

[0056] To ensure the system's robustness in actual deployment, this sub-step incorporates an anomaly handling process. During inference, the quality of the input image, such as contrast, and the confidence level of the model output are monitored in real time. If an anomaly is detected, such as excessively low image quality or a confidence level below a preset threshold, the system will automatically interrupt the standard deep learning inference process and seamlessly switch to a traditional filtering algorithm, such as Gaussian filtering, as a backup solution. Simultaneously, the anomaly event is recorded, thus ensuring service continuity and system reliability.

[0057] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. A lightweight simulation pre-training-based improved U-Net skull X-ray denoising method, characterized in that, Includes the following steps: S1, construct and prepare a dataset of simulated and real skull X-ray images, the dataset including simulated image pairs generated based on a simplified model and preprocessed real clinical images; S2, Construct an improved U-Net denoising network that integrates attention and a two-stage training mechanism. The network includes feature modules that introduce depthwise separable convolution and SE channel attention mechanism, as well as a two-stage architecture that supports configurable channel number. S3, perform simulation pre-training, real-world fine-tuning, and denoising inference and evaluation. The simulation pre-training uses a simulation dataset to perform lightweight pre-training on the network. The real-world fine-tuning loads the pre-trained weights and fine-tunes the network based on real data. The denoising inference applies the trained model to real images to output clear images.

2. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 1, characterized in that, S1 includes: S11, Generate a skull simulation model based on parametric random modeling, wherein the parametric random modeling uses an ellipsoid to simulate brain tissue and an ellipsoidal shell to simulate the skull shell. S12, projection simulation is performed based on a simplified physical model and random imaging parameters to generate paired noise data. The simplified physical model uses an exponential decay formula to calculate the primary signal and a linear superposition model of scattering coefficients to calculate the scatter noise. S13. Collect and preprocess real skull X-ray images to construct a supervised training dataset. The preprocessing includes size normalization, grayscale standardization, and noise region annotation.

3. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 2, characterized in that, In the parametric random modeling, the lengths of the three semi-axis of the ellipsoid are randomly and independently selected within the preset ranges of [13,16]cm, [11,14]cm, and [9,12]cm, respectively; the shell thickness of the ellipsoid is randomly selected within the preset range of [6,8]mm; and the random imaging parameters include the photon energy of the light source randomly selected within the range of [50,150]keV, and the tube current randomly selected within the range of [100,500]mA.

4. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 1, characterized in that, S2 includes: S21, Design a feature extraction module and a feature restoration module. The feature extraction module uses depthwise separable convolution to extract spatial features and embeds an SE channel attention mechanism. The feature restoration module is composed of 1×1 convolution, 3×3 convolution and SE channel attention mechanism connected in sequence. S22, Construct a two-stage training architecture that supports configurable channel count. In the pre-training stage, the number of channels in each layer of the network is set to half of the full configuration, and in the denoising fine-tuning stage, it is restored to the full configuration. S23, In the skip connection between the encoder and the decoder, a channel-level attention fusion mechanism is introduced, which adaptively weights and multi-scale fuses the feature channels from the pre-training stage and the newly added feature channels in the fine-tuning stage.

5. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 4, characterized in that, Step S21, based on the distribution characteristics of scattering noise in skull X-ray images, designs a feature extraction module and a feature restoration module: The feature extraction module uses a 3×3 depth separable convolutional layer to extract spatial features from the input noisy image to reduce the computational complexity of the model, and then connects to the Squeeze-and-Excitation channel attention submodule. The Squeeze-and-Excitation channel attention submodule compresses the spatial dimension of the feature map through a global average pooling layer to obtain a channel description vector. This vector is then passed through two fully connected layers to learn the non-linear dependencies between channels and outputs weight coefficients equal to the number of channels. These weight coefficients are then used to reweight the initial features extracted by depthwise separable convolutions in terms of channel dimension, thereby enhancing the network's ability to distinguish and select features from strong scattering noise in high-density skeletal regions and weak scattering noise in soft tissue regions. The feature restoration module is used to gradually restore image details in the decoder stage. Its structure consists of a 1×1 convolutional layer, a 3×3 convolutional layer, a Squeeze-and-Excitation channel attention submodule, a batch normalization layer, and a ReLU activation function. The 1×1 convolutional layer is used to adjust the number of channels, the 3×3 convolutional layer is used for spatial detail reconstruction, and the Squeeze-and-Excitation channel attention submodule performs importance recalibration on the reconstructed features.

6. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 4, characterized in that, Step S22, which constructs a two-stage training architecture supporting configurable channel numbers, specifically includes: During the simulation data pre-training stage, the number of output channels of all convolutional layers in the improved U-Net network is set to half the number of corresponding channels of the target complete network, forming a lightweight network structure for training. In the real data denoising and fine-tuning stage, the number of channels in each layer of the network is restored to the complete standard configuration, and the parameters of each convolutional layer in the network are initialized: the convolutional kernel weights of the first C / 2 channels of the convolutional layer are loaded with the parameters learned in the pre-training stage, and the convolutional kernel weights of the last C / 2 channels adopt a random initialization strategy, where C is the complete number of output channels of the layer.

7. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 4, characterized in that, The channel-level attention fusion mechanism in S23 performs the following operations: The feature map output by the encoder is divided into two parts according to its channel source. The two parts include feature channels inherited from the pre-training weights and feature channels newly added and initialized during the fine-tuning stage. Global average pooling is performed on the two feature sets respectively to obtain two channel-level global description vectors; The two description vectors are concatenated and then input into a shared fully connected network to generate an attention weight vector with the same number of channels as the original feature map. The attention weight vector is subjected to Sigmoid activation and normalization to obtain normalized attention coefficients corresponding to the two parts of features; The two feature maps are multiplied by their corresponding normalized attention coefficients to complete the adaptive weighted fusion.

8. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 1, characterized in that, S3 includes: S31, The network is pre-trained using a simulation dataset. The pre-training uses a loss function that combines L1 and mean square error to enable the network to learn the physical characteristics of scattering noise. S32, Fine-tuning the network loaded with pre-trained weights using real data, wherein the fine-tuning adopts a dynamic learning rate and an early stopping strategy to adapt the network to the distribution of real data. S33 applies the trained model to real skull X-ray images for denoising inference and outputs clear images. S34, Quantitative and qualitative evaluations are performed on the denoising results. The quantitative evaluation is based on PSNR and SSIM indices, and the qualitative evaluation is performed by experts through visual evaluation. S35, set up an exception handling and process protection mechanism. When the input quality is detected to be abnormal and the output confidence is too low, the mechanism interrupts the process and enables backup filtering processing.

9. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 8, characterized in that, In step S32, a small amount of real data is used to fine-tune the network loaded with targeted pre-trained weights, specifically including: The number of channels in each layer of the lightweight network, which has been pre-trained by simulation, is restored to its full configuration to form a complete improved U-Net network architecture. For each convolutional layer in the network that has been restored to its full configuration, the parameters are initialized: for a convolutional layer with C output channels, the kernel weights of the first C / 2 channels are loaded with the parameters learned during the pre-training stage, and the kernel weights of the last C / 2 channels are randomly initialized. The complete network was fine-tuned using the relatively small amount of real skull X-ray training set prepared in S1. During the fine-tuning process, a dynamic learning rate scheduling strategy was adopted, and an early stopping mechanism was set to prevent overfitting. The fine-tuning process, based on the physical feature representation of scattering noise learned in pre-training, uses a small amount of real data to enable the network to quickly adapt to the specific distribution details of real skull X-ray data, significantly reducing the amount of real data and training time required for model convergence, and improving the model's ability to remove scattering noise from real skull X-ray images.

10. The improved U-Net skull X-ray denoising method based on lightweight simulation pre-training according to claim 8, characterized in that, The fine-tuning in S32 also includes: The channel-level attention fusion mechanism dynamically adjusts the contribution weights of inherited features and newly initialized features in the feature extraction and reconstruction process. This mechanism ensures that the network can retain the basic physical features learned from simulation data and flexibly learn features unique to real data during the fine-tuning stage.