Cryoelectron microscope image noise removal system and method based on deep learning
Through the diffusion model optimized by deep learning algorithm, the problem of high-frequency information loss in cryo-electron microscope image denoising is solved, efficient denoising is achieved and key details are retained, image quality and classification accuracy are improved, and it is suitable for structural biology research.
Patent Information
- Application Number
- CN202510665472.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-11
AI Technical Summary
When removing cryo-electron microscope image noise, it is difficult to retain key high-frequency information while denoising, resulting in poor image classification effect, and traditional methods have high calculation costs, making it difficult to meet practical application needs.
Deep learning algorithms, especially the improved version of the diffusion model, are adopted to optimize the training process and modular design, and three-step denoising is achieved, improving computing efficiency and retaining high-frequency details of the image, and training is combined with UNet network to adapt to different experimental conditions.
It significantly improves the signal-to-noise ratio of cryo-electron microscopy images, retains key high-frequency details, improves the accuracy of two-dimensional and three-dimensional classification, shortens processing time, and is suitable for structural biology research.
Smart Images

Figure CN120298249A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cryo-electron microscopy single-particle analysis, and particularly relates to a cryo-electron microscopy image noise removal system and method based on deep learning. Background Art
[0002] Cryo-electron microscopy (Cryo-EM) is currently widely used to resolve the high-resolution three-dimensional structures of proteins and macromolecular complexes, making significant contributions to structural biology. The core component of this technology is the multi-conformation macromolecule classification and modeling, which can reveal fundamental problems in structural biology and enhance the ability to design therapeutic molecules that can induce specific functional changes in target proteins. One of the main limitations in the multi-conformation modeling of biological macromolecules is the challenge of obtaining a sufficient number of high-precision target conformation particle image stacks. To prevent protein damage, the electron beam irradiation dose is often strictly limited. Therefore, cryo-electron microscopy images usually have a very low signal-to-noise ratio (SNR), ranging from 0.01 to 0.1. The high-intensity noise introduces significant errors in Bayesian probability calculations, greatly limiting the two-dimensional and three-dimensional classification effects of the original images. This usually requires continuous parameter adjustment and multiple rounds of classification to obtain high-quality target conformation particle image stacks.
[0003] To improve the contrast of cryo-electron microscopy images and reduce the noise level, various methods have been developed, such as BM3D, band-pass filters, and Wiener filters. However, the noise models in cryo-electron microscopy images are often unknown and very complex. In addition, differences in experimental parameters may lead to changes in the noise model. Therefore, the pre-defined image priors in these traditional methods cannot accurately match the noise models in cryo-electron microscopy images, limiting the performance of many traditional denoising techniques in cryo-electron microscopy data.
[0004] In the field of cryo-EM image denoising, Topaz-denoise is based on the self-supervised Noise2Noise (N2N) framework. This framework does not rely on paired clean image data, which is a highly innovative application in cryo-EM data processing. Since it is almost impossible to obtain noise-free real images in cryo-EM experiments, Topaz-denoise greatly reduces the need for manually labeled data by dividing multi-frame images into odd frames and even frames and using them as denoising labels for each other, simplifying the entire denoising process. Through this method, Topaz-denoise can effectively remove random noise in cryo-EM images while maintaining signal consistency among multi-frame images. Its excellent denoising effect significantly improves the contrast of the images, making subsequent analysis more accurate. Different from Topaz-denoise, NT2C adopts a denoising algorithm based on simulated data. It trains the model by simulating noise models under different experimental conditions. This denoising method based on simulated data is not only more powerful in performance but also shows higher flexibility in dealing with complex noise. NT2C can adjust the denoising model according to different experimental conditions by simulating the noise characteristics under various experimental parameters, thus achieving a more accurate denoising effect.
[0005] Disadvantages of Topaz-denoise:
[0006] High-frequency information in cryo-EM images usually contains key molecular structure details, especially in particle images used for two-dimensional and three-dimensional classification. These details are crucial for correctly identifying molecular conformations. While Topaz-denoise removes noise indiscriminately, it often inadvertently deletes this high-frequency structural information, resulting in the denoised images being clearer to the naked eye but losing important molecular structure details, which affects the final classification effect.
[0007] Disadvantages of NT2C:
[0008] The denoising method of NT2C based on simulated data also faces the dilemma of losing high-frequency information. Although NT2C has a significant denoising effect, its "powerful" denoising strategy often inevitably eliminates high-frequency information in the images while removing noise. In addition, there are certain differences between the simulated data and the actual experimental conditions. Although NT2C trains the denoising model by simulating multiple noise models, the complexity of the simulated noise models may not be sufficient to fully capture all kinds of noise characteristics that appear in real experiments, resulting in its denoising effect under certain experimental conditions may not meet expectations. Summary of the Invention
[0009] In view of this, the objective of the present invention is to provide a cryo-electron microscopy image noise removal system and method based on deep learning, aiming to achieve efficient denoising while maximizing the retention of key information in the image. The core innovation of this system lies in the introduction of deep learning algorithms, particularly an efficient improved version of the diffusion model, which breaks through the limitations of traditional denoising methods. Different from traditional diffusion models, the denoising system in the present invention is specially designed to greatly improve the computational efficiency and the training speed of the model, solving the bottleneck that diffusion models are difficult to popularize in practical applications due to a large number of computational steps. Traditional diffusion models often require hundreds or even thousands of diffusion steps, and each step requires complex calculations and repeated training, resulting in excessively high computational costs when dealing with large-scale cryo-electron microscopy data and being unable to meet the actual application requirements. However, through the analysis and optimization of the structure of the diffusion model, the present invention proposes an efficient diffusion model that only requires three steps, which not only greatly shortens the training time but also is no less effective in denoising than traditional diffusion models. By significantly reducing the training steps and optimizing the computational process, the present invention makes the processing of large-scale cryo-electron microscopy data more practical and feasible, and high-quality denoising results can be obtained in a short time. In addition, due to the powerful generalization ability of deep learning models, this system can not only remove random noise but also maintain good adaptability and robustness under different experimental conditions, further improving the actual effect of denoised particle images in two-dimensional and three-dimensional classifications.
[0010] A cryo-electron microscopy image noise removal system based on deep learning, comprising:
[0011] A data preprocessing module that preprocesses the input original electron microscopy image to obtain a pure noise image and a particle image;
[0012] An autoencoder data processing module that adds random perturbations to the particle coordinates to obtain random coordinates, and crops the cryo-electron microscopy image according to the random coordinates; the cropped image is copied and forms a data pair with itself, with one of them being the training input of the UNet network and the other being the training label, which is input to the training module;
[0013] An N2N data processing module that is used to receive the pure noise image, matches any two pure noise images to form a noise data pair, with one of them being the training input of the UNet network and the other being the training label of the UNet network, which is input to the training module;
[0014] A random sample data processing module that obtains a noisy image and an image before noise addition by adding a pure noise image to the particle image; constructs a noisy image data pair from images with different noise intensities, with the high-noise image as the training input of the UNet network and the low-noise image as the training label of the UNet network, which is input to the training module;
[0015] The training module trains the UNet network based on the training inputs and training labels output by the autoencoder data processing module, the random sample data processing module, and the N2N data processing module to obtain a pre-trained model;
[0016] The denoising module inputs the original cryo-electron microscopy image into the pre-trained model and outputs the denoised image.
[0017] Preferably, the preprocessing includes selecting pure noise images, contrast enhancement, and normalization processing.
[0018] Preferably, the random perturbation added by the autoencoder data processing module is not greater than the particle diameter.
[0019] Preferably, when the UNet network is trained, the mean square error is set as the loss function, and the model gradually optimizes the weight parameters of the network to minimize the difference between the prediction result and the target value; the loss function continuously monitors the error between the network output and the real data, and uses the backpropagation algorithm for gradient update, thereby gradually reducing the training error until the loss function converges to obtain the pre-trained model.
[0020] Preferably, the autoencoder data processing module, the N2N data processing module, and the random sample data processing module are all based on the simplified diffusion model formula:
[0021] img t+1 =αimg t +(1-α)τ n
[0022] img t represents the image after the t-th noise addition based on the original image, τ n represents a random one selected from N pure noise images, and α represents the diffusion model parameter;
[0023] For the autoencoder data processing module, α is 1, and at this time lim α→1 img t =img t+1 , and the label of the image is itself;
[0024] For the N2N data processing module, α is 0, and at this time lim α→0 img t =τ, and at this time img t will become a noise image, and its corresponding label is also a pure noise image; for the random sample data processing module, α≥0.5.
[0025] A method for removing noise from cryo-electron microscopy images based on deep learning, including:
[0026] First step: Identify the images containing particles and pure noise images from the original electron microscope images; perform normalization processing on the input original electron microscope images and output the normalized original electron microscope images to the autoencoder data processing module, the N2N data processing module, and the random sample data processing module;
[0027] Second step: The autoencoder data processing module receives the particle images, randomly crops the images, activates specific regions in the latent space, and outputs the processed particle images to the training module;
[0028] Third step: The N2N data processing module receives the pure noise images, introduces different noise data pairs, and outputs the processed pure noise images to the training module;
[0029] Fourth step: The random sample data processing module is used to receive the pure noise and particle images, introduce random noise into the particle images to construct a multi-conformation particle image space, and output the processed data pairs to the training module;
[0030] Fifth step: The training module receives the image data output from the autoencoder data processing module, the random sample data processing module, and the N2N data processing module, passes it as input to the UNet network for training until the loss function converges, obtains the pre-trained model and outputs it to the denoising module;
[0031] Sixth step: The denoising module receives the pre-trained model from the training module and processes the original cryo-electron microscope images through this model to obtain the denoised electron microscope images.
[0032] The present invention has the following beneficial effects:
[0033] The present invention significantly improves the signal-to-noise ratio (SNR) of cryo-electron microscope (cryo-EM) images through an advanced denoising algorithm. Through experimental verification, the present invention can effectively separate useful signals from noise, thereby making the particle images of target proteins or macromolecular complexes clearer. This improved signal-to-noise ratio provides more reliable image data for subsequent structural analysis and classification, and thus helps to obtain more accurate three-dimensional structure reconstruction.
[0034] While removing noise, the present invention is specially designed to retain the high-frequency details in the images, which is crucial for analyzing the fine structure of proteins. The present invention optimizes the algorithm to ensure that the morphology and structural features of the particles are maximally retained during the denoising process, avoiding the common problem of detail loss in traditional denoising methods. The retention of such high-frequency details enables the present invention to more accurately distinguish different protein conformations during 2D and 3D classification.
[0035] The present invention realizes the efficient classification of target proteins or macromolecular complexes by combining with the multi-conformation classification process (DMCCP). By using the images processed by the denoising model of the present invention, particle picking, 2D classification, and 3D classification can be performed quickly and accurately, significantly improving the quality and quantity of classification.
[0036] The present invention fully considers computational efficiency during design and can run efficiently on a computer equipped with a GPU, greatly shortening the denoising processing time. In addition, the algorithm framework of the present invention has good scalability and can be adjusted and optimized according to different requirements to adapt to the ever-developing structural biology research.
[0037] The present invention adopts a modular design and can be easily integrated into existing structural biology research processes, such as the relion software. This design makes the present invention not only applicable to specific research environments but also has wide applicability. In addition, the present invention provides detailed tutorials and pre-trained models, enabling users without a deep learning background to easily use the model for image denoising and structural analysis. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of the system composition of the present invention;
[0039] Figure 2 It is a schematic diagram of the result after denoising. Detailed Embodiments
[0040] The following takes embodiments in conjunction with the drawings and describes the present invention in detail.
[0041] A cryo-EM image noise removal system based on deep learning includes a training data generation module and a training result generation module;
[0042] The training data generation module includes a preprocessing module, an autoencoder data processing module, an N2N data processing module, and a random sample data processing module;
[0043] The training result generation module includes a training module and a denoising module.
[0044] The training data generation module is the core of the system and is responsible for accurately generating the data required by the model;
[0045] The data preprocessing module preprocesses the input original electron microscope images and respectively outputs the preprocessed images to three modules: the autoencoder data processing module, the random sample data processing module, and the N2N data processing module. The preprocessing includes steps such as selecting pure noise, contrast enhancement, and normalization to improve the accuracy of subsequent data generation and denoising;
[0046] The autoencoder data processing module is used to generate input data and label data in the autoencoder mode. Specifically, this module adds random perturbations no greater than the particle diameter to the particle coordinates to obtain random coordinates, and crops the cryo-EM images according to the random coordinates. This can make the particle images appear at any position after cropping, thus preventing the neural network from forming an obvious central attention bias during training and improving the overall robustness. Subsequently, the cropped images will be copied and paired with themselves. The same particle image is both the training input and the training label of the neural network, which enables the neural network to learn the common features of the particle images;
[0047] The N2N data processing module is used to process the pure noise images selected in the data preprocessing module, match any two pure noise images to form noise data pairs, where one is used as the training input of the neural network and the other is used as the training label. This training method can enhance the denoising effect of the images, and when used in combination with the data pairs generated by the autoencoder data processing module, enables the neural network to distinguish particle images from noise images, thus improving the denoising ability while maintaining the image details;
[0048] The random sample data processing module is used to receive pure noise and particle images. By continuously and randomly adding a certain proportion of pure noise images to the particle images according to the diffusion formula, the noisy images and the images before adding noise are obtained, and a particle image sequence with the noise intensity as a variable is constructed. Subsequently, this module sequentially extracts images with different noise intensities to construct noisy image data pairs, using the high-noise images as the neural network input and the low-noise images as the training labels. Such data pairs can enable the neural network to retain particle details during image denoising and simplify the diffusion process to improve the calculation efficiency. The use of randomly sampled noise further enhances the robustness of the model, enabling it to achieve excellent denoising effects with fewer training steps;
[0049] The training module receives the image data output from the autoencoder data processing module, the random sample data processing module, and the N2N data processing module, and uses it as input to the UNet network for training. By setting the mean squared error (MSE) as the loss function, the model gradually optimizes the weight parameters of the network to minimize the difference between the prediction result and the target value. During the training process, the loss function continuously monitors the error between the network output and the real data, and uses the backpropagation algorithm for gradient update, thereby gradually reducing the training error until the loss function converges to obtain the pre-trained model. The trained UNet model can effectively retain the high-frequency detail information of the image while denoising, and is used for subsequent denoising tasks.
[0050] The denoising module receives the pre-trained model from the training module and takes the original cryo-EM image as input. By loading the pre-trained UNet model, this module processes the input original image and suppresses and removes the noise in the image using the model's denoising ability. Finally, the denoising module outputs the processed image, which has a higher signal-to-noise ratio than the original image and retains the necessary structural information for subsequent 2D and 3D classification. These denoised images can be directly used for subsequent analysis, such as structure resolution or classification tasks.
[0051] The autoencoder data processing module, N2N data processing module, and random sample data processing module are all based on the simplified diffusion model formula:
[0052] img t = αimg t+1 +(1 - α)τ n
[0053] img represents the original image, τ n represents the pure noise image, and α represents the diffusion model parameter.
[0054] The autoencoder data processing module is the extreme case where α is 1. At this time, lim α→1 img t = img t+1 , and the label of the image is itself;
[0055] The N2N data processing module is the extreme case where α is 0. At this time, lim α→0 img t = τ. At this time, img t will become a noise image, and its corresponding label is also the pure noise image, which is consistent with the noise autoencoder in the autoencoder module.
[0056] The random sample data processing module is the general case. A part of random noise is added to the label to form a mapping from the high-noise image to the latent space to the low-noise image, generating stronger output heterogeneity. The random noise scheduling can increase the sampling range of the multi-conformation particle image space, making the image belong to the multi-conformation particle image space as much as possible. In addition, we set an upper limit for the diffusion parameter, that is, α ≥ 0.5. This is because when α < 0.5, the noise signal intensity is greater than the particle signal intensity, and the whole image is closer to the noise domain.
[0057] A method for removing noise from cryo-EM images based on deep learning, the method includes:
[0058] In the first step, images containing particles and pure noise images are identified from the original electron microscope images. The data preprocessing module performs normalization processing on the input original electron microscope images and outputs the normalized original electron microscope images to the autoencoder data processing module, the N2N data processing module, and the random sample data processing module;
[0059] In the second step, the autoencoder data processing module receives the particle images, randomly crops the images, and then copies the particle images to construct input-label data pairs, that is, the same particle image is both the input of the neural network and its training label. This form of data pair is input to the training module for training, which can help the neural network learn the common features of the particle images and protect the particle information;
[0060] In the third step, the N2N data processing module receives the selected pure noise images, randomly shuffles and matches them, constructs random noise-random noise data pairs, and outputs this data pair to the training module, which can help the neural network learn the feature distribution of the noise and improve the denoising ability;
[0061] In the fourth step, the random sample data processing module is used to receive pure noise and particle images, gradually introduce random noise into the particle images to construct a sequence of noisy particle image data. These sequences are constructed into data pairs according to the noise intensity order, and these data help the neural network analyze the multi-conformation particle image space so that it can capture different conformations of the particles;
[0062] In the fifth step, the training module receives the image data pairs output from the autoencoder data processing module, the random sample data processing module, and the N2N data processing module, randomly shuffles them, and passes them to the UNet network for training until the loss function converges, obtaining a pre-trained model and outputting it to the denoising module;
[0063] In the sixth step, the denoising module receives the pre-trained model from the training module and processes the original cryo-electron microscope image through this model to obtain the denoised electron microscope image.
[0064] Embodiment
[0065] Figure 1 is the overall flowchart of the system. As Figure 1 shown, first, the original electron microscope image enters the training data generation module, and training data is obtained through the content preprocessing module, the autoencoder data processing module, the N2N data processing module, and the random sample data processing module in the training data generation module. Figure 1 The left side is the original image.
[0066] As Figure 1 shown, the training data is input into the training result generation module, and the original image is denoised through the training module and the denoising module to obtain the denoised image, as Figure 2as shown
[0067] In summary, the above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cryo-electron microscopy image noise removal system based on deep learning, characterized in that, Including: A data preprocessing module that preprocesses the input original electron microscope image to obtain a pure noise image and a particle image; An autoencoder data processing module that adds random perturbations to the particle coordinates to obtain random coordinates, and crops the cryo-electron microscope image according to the random coordinates; the cropped image is copied and forms a data pair with itself, and one of them is used as the training input of the UNet network, and the other is used as the training label and input to the training module; An N2N data processing module for receiving the pure noise image, matching any two pure noise images to form a noise data pair, where one is used as the training input of the UNet network, and the other is used as the training label of the UNet network and input to the training module; A random sample data processing module that adds a pure noise image to the particle image to obtain a noisy image and the image before adding noise; constructs a noisy image data pair from images with different noise intensities, and uses the high-noise image as the training input of the UNet network and the low-noise image as the training label of the UNet network and inputs them to the training module; The training module trains the UNet network according to the training input and training label output by the autoencoder data processing module, the random sample data processing module, and the N2N data processing module to obtain a pre-trained model; A denoising module that inputs the original cryo-electron microscope image into the pre-trained model and outputs the denoised image.
2. The noise removal system for cryo-electron microscopy images based on deep learning according to claim 1, wherein The preprocessing includes selecting a pure noise image, contrast enhancement, and normalization processing.
3. A noise removal system for cryo-electron microscopy images based on deep learning according to claim 1, characterized in that, The random perturbation added by the autoencoder data processing module is not greater than the particle diameter.
4. A cryo-electron microscopy image noise removal system based on deep learning according to claim 1, characterized in that, When the UNet network is trained, the mean square error is set as the loss function, and the model gradually optimizes the weight parameters of the network to minimize the difference between the prediction result and the target value; The loss function continuously monitors the error between the network output and the real data, and uses the backpropagation algorithm for gradient update, thereby gradually reducing the training error until the loss function converges to obtain a pre-trained model.
5. A cryo-electron microscope image noise removal system based on deep learning according to claim 1, wherein The autoencoder data processing module, the N2N data processing module, and the random sample data processing module are all based on the simplified diffusion model formula: img t+1 = αimg t + (1 - α)τ n img t represents the image after the t-th noise addition based on the original image, τ n represents a random one selected from N pure noise images, and α represents the diffusion model parameter; For the autoencoder data processing module, when α is 1, then lim α→1 img t = img t+1 , and the label of the image is itself; For the N2N data processing module, α is 0, and at this time lim α→0 img t = τ, and at this time img t will become a noisy image, and its corresponding label is also a pure noisy image; For the random sample data processing module, α≥0.
5.
6. A method for removing noise from cryo-electron microscopy images based on deep learning, characterized in that, Including: The first step is to identify the image containing particles and the pure noise image from the original electron microscope image; Perform standardization processing on the input original electron microscope image and output the standardized original electron microscope to the autoencoder data processing module, the N2N data processing module, and the random sample data processing module; The second step is that the autoencoder data processing module receives the particle image, randomly crops the image, activates specific regions in the latent space, and outputs the processed particle image to the training module; The third step is that the N2N data processing module receives the pure noise image, introduces different noise data pairs, and outputs the processed pure noise image to the training module; Fourthly, the random sample data processing module is used to receive pure noise and particle images, introduce random noise into the particle images to construct a multi-conformation particle image space, and output the processed data pair to the training module; Fifthly, the training module receives the image data output from the autoencoder data processing module, the random sample data processing module, and the N2N data processing module, and uses it as input to be passed to the UNet network for training until the loss function converges, obtaining a pre-trained model and outputting it to the denoising module; Sixthly, the denoising module receives the pre-trained model from the training module, and processes the original cryo-electron microscopy image through this model to obtain a denoised electron microscopy image.