Method and apparatus for reconstructing structure of substance, and device and medium

By acquiring noise density maps and reference density maps, and utilizing machine learning models and diffusion posterior sampling techniques, the problems of noise interference and wedge-shaped defects in cryo-electron microscopy were solved, achieving more accurate reconstruction of material structures and improving cryo-electron microscopy image quality and reconstruction accuracy.

WO2026065549A1PCT designated stage Publication Date: 2026-04-02BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing cryo-electron microscopy techniques face severe noise interference, wedge-shaped loss, and prior distribution limitations when reconstructing material structures, resulting in insufficient reconstruction accuracy. In particular, at low resolution, they are prone to introducing illusory details and artifacts.

Method used

By acquiring a noise density map and a reference density map, and utilizing machine learning models and diffusion posterior sampling techniques, the motion of the noise density map is corrected based on the reference density map to generate a more accurate target density map. The reconstruction process is then optimized by combining flow matching and diffusion models.

Benefits of technology

It improves the quality and reconstruction accuracy of cryo-electron microscopy images, solves the wedge-shaped missing image problem, reduces artifacts, and enhances the accuracy of material structure reconstruction and the ability to restore details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123103_02042026_PF_FP_ABST
    Figure CN2024123103_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for reconstructing the structure of a substance, and a device and a medium. In the method, a noise density map related to the structure of a substance and a reference density map representing the structure of a target substance are acquired; and on the basis of the reference density map, the noise density map is converted into a target density map representing the structure of the target substance, wherein the noise density map, the reference density map and the target density map respectively indicate the distributions of substance particles, and the reference density map is used for correcting the motion of the substance particles during conversion. In this way, a noise density map may be used as a condition to obtain a more accurate target density map, thereby performing more accurate reconstruction of the structure of a target substance.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatuses, devices, and media for reconstructing a structure of a material TECHNICAL FIELD

[0001] Exemplary implementations of the present disclosure generally relate to computer technology, and particularly relate to methods, apparatuses, devices, and computer-readable storage media for reconstructing a structure of a material. BACKGROUND

[0002] Cryo-electron microscopy is an important technology in the field of structural biology and drug discovery that can determine high-resolution three-dimensional structures of biological molecules that are otherwise difficult to study by traditional methods. Cryo-electron microscopy provides insights into molecular mechanisms and aids in fields such as drug discovery. One major component of cryo-electron microscopy involves reconstructing a three-dimensional structure from noisy two-dimensional projections of particles, a task that can be formulated as an inverse problem aiming to recover a signal x e Rn(i.e., the noiseless protein density) from observations y e Rm. Following Bayesian statistics, the goal of this task is to sample the density p(x|y) from the posterior distribution, which can be decomposed as p(x|y) ∝ p(y|x)p(x). Therefore, the prior distribution is crucial in guiding the reconstruction process and improving the accuracy of the resulting structure.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for reconstructing a structure of a material is provided. In the method, a noisy density map related to the structure of the material and a reference density map representing a structure of a target material are obtained; and the noisy density map is converted into a target density map representing the structure of the target material based on the reference density map, the noisy density map, the reference density map, and the target density map respectively indicating a distribution of particles of the material, the reference density map being used to correct a movement of the particles of the material in the conversion.

[0005] In a second aspect of the present disclosure, an apparatus for reconstructing a structure of a material is provided. The apparatus comprises a density map obtaining module configured to obtain a noisy density map related to the structure of the material and a reference density map representing a structure of a target material; and a density map converting module configured to convert the noisy density map into a target density map representing the structure of the target material based on the reference density map, the noisy density map, the reference density map, and the target density map respectively indicating a distribution of particles of the material, the reference density map being used to correct a movement of the particles of the material in the conversion.

[0006] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the disclosure.

[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.

[0008] In a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.

[0009] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the various embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiment illustrated. Other BRIEF DESCRIPTION OF DRAWINGS

[0010] The above-mentioned and other features and advantages of various embodiments of the present disclosure will become more apparent by reference to the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals designate like elements, wherein:

[0011] FIG. 1 shows a schematic diagram of an example environment in which implementations of the present disclosure can be implemented;

[0012] FIG. 2 shows a schematic diagram of a process of reconstructing a material structure according to some embodiments of the present disclosure;

[0013] FIG. 3A shows a schematic diagram of anisotropic noise according to some embodiments of the present disclosure;

[0014] FIG. 3B shows a schematic diagram of a wedge-shaped missing repair of a noise signal according to some embodiments of the present disclosure;

[0015] FIG. 4 shows a schematic diagram of an example architecture of a machine learning model according to some embodiments of the present disclosure;

[0016] FIG. 5 shows a flowchart of a method for reconstructing a material structure according to some embodiments of the present disclosure;

[0017] FIG. 6 shows a block diagram of an apparatus for reconstructing a material structure according to some embodiments of the present disclosure; and

[0018] FIG. 7 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0023] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user, so that the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0024] As an optional but not limited implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be the manner of pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0026] As used herein, the term “model” can learn a relationship between a corresponding input and output from training data, such that after training is completed, a corresponding output can be generated for a given input. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes an input and provides a corresponding output by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a “model” can also be referred to as a “machine learning model,” a “learning model,” a “machine learning network,” or a “learning network,” which are used interchangeably herein.

[0027] A “neural network” is a machine learning network based on deep learning. A neural network is capable of processing an input and providing a corresponding output, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can include three phases, namely a training phase, a testing phase, and an application phase (also referred to as an inference phase). In the training phase, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to obtain consistent inferences from the training data that satisfy an expected goal. Through training, the model can be considered to have learned a relationship (also referred to as a mapping) from input to output from the training data. The parameter values of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether the model is able to provide correct outputs, determining the performance of the model. The testing phase can sometimes be merged into the training phase. In the application or inference phase, the trained model can be used to process actual model inputs based on the trained parameter values, determining corresponding model outputs.

[0029] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the environment 100, an electronic device 120 includes or is deployed with a material structure reconstruction system 110. The material structure reconstruction system 110 is configured to generate a target map 106 related to a structure of a target material based on a noise 102 and a reference map 104 related to the structure of the target material, thereby enabling reconstruction of the structure of the target material. In some embodiments, the reference map 104 and the target map 106 are three-dimensional images.

[0030] In the environment 100, the electronic device 120 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combinations of the aforementioned and the like, including accessories and peripherals related thereto, or any combinations thereof. In some embodiments, the terminal device 110 can also be capable of supporting any type of interface to a user (such as "wearable" circuitry, etc.). The recommendation system 110 can be implemented, for example, in various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like.

[0031] It should be understood that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0032] As mentioned previously, cryo-EM techniques can be used to reconstruct three-dimensional structures. A Gaussian distribution can be introduced as a prior p(x) for three-dimensional reconstruction for cryo-EM, effectively serving as a frequency-dependent low-pass filter. Building on this, recent work has explored more complex regularizers, rather than explicitly defining a prior distribution. While these methods were initially developed for the three-dimensional reconstruction task, their approach is closely related to methods of density map modification and post-processing. Both in the refinement process and in post-processing, there is a common goal of improving the quality of the maps generated by cryo-EM. Among all these methods, a series of methods have emerged that directly learn p(x|y) from data using data-driven pre-trained models. While these models are powerful, these models are often designed for a specific task and assume a consistent degradation operator, which limits their universal applicability and versatility. Some works related to cryo-EM will be introduced below.

[0033] Density correction and denoising can be used in cryo-EM to improve the quality of the maps generated by cryo-EM. Density correction involves correcting errors in observed cryo-EM density maps using known characteristics of the expected density of certain regions in the map. These methods are often heuristic-based, using pre-defined forward operators that contain local resolution estimates and filtering strategies. While these methods have a high degree of interpretability, their reliance on heuristic noise terms and limited incorporation of complex priors often limit their ability to handle more complex structural details. Deep learning is playing an increasingly important role in the denoising and correction of maps generated by cryo-EM. Some methods use pre-trained models trained on noise-noise data pairs to recover high-frequency details and refine density maps, thereby helping to build atomic models. These models can sometimes introduce hallucinated details, especially in low-resolution density maps, as they rely on learned mappings.

[0034] Cryo-electron tomography is used to reconstruct a three-dimensional volume of a biological specimen by capturing two-dimensional images at different tilt angles. The reconstructed three-dimensional volume is referred to as a tomogram, providing a detailed view of cellular structures and macromolecular complexes in their native environment. However, one major challenge of cryo-electron tomography is the missing wedge problem, which arises due to physical limitations of the microscope that limit the range of tilt angles during data acquisition. Typically, only angles between (-60°, +60°) can be captured, leaving a wedge-shaped region in Fourier space without information. This missing information can lead to anisotropic resolution in the reconstruction, resulting in artifacts such as elongation or distortion in certain directions, particularly along the missing tilt axis. Traditional methods to address the missing wedge problem rely mainly on various signal processing techniques, often using regularization strategies to compensate for the missing information. While these methods can improve the overall quality of the tomogram, they are often based on heuristic assumptions and have limited ability to fully recover the lost data. In recent years, deep learning has introduced new solutions to the missing wedge problem by leveraging data-driven models. These methods are able to capture more complex patterns, thereby improving the interpretability of the tomogram. However, these methods mostly target the problem at the level of the full tomogram, which makes it challenging to integrate specific prior knowledge of the protein structure. This limits their ability to fully exploit the available information that could be beneficial in improving the accuracy of sub-tomogram reconstructions.

[0035] Ab initio modeling in cryo-electron microscopy involves estimating the three-dimensional structure of a protein from a set of 2D particle images, where the orientation of the particles is unknown. Early methods relied on experimental techniques, such as using image tilt pairs, due to the low signal-to-noise ratio (SNR) and ill-posed nature of particle images. Some methods improved the SNR but sacrificed high-frequency information.

[0036] The denoising diffusion probability model achieves an iterative refinement process by learning to gradually denoise samples from a normal distribution, and this model has achieved results on some generative tasks. In some related work, this model has been introduced into cryo-electron microscopy. A flow-based model regresses the vector field of the probability path required for generation, and this model has achieved success in image generation and molecular generation.

[0037] The inverse problem aims to recover the original sample given some degraded observations, such as image super-resolution, inpainting, and deblurring. Diffusion models have demonstrated their ability to solve inverse problems in unsupervised methods. Since pre-trained denoising diffusion probability models have the ability to model the data distribution (prior distribution), they can help discover the posterior distribution given a degraded model (likelihood).

[0038] To address at least one of the problems in the aforementioned related work, an embodiment of this disclosure proposes a scheme for reconstructing the structure of matter. Specifically, a noise density map related to the structure of the matter and a reference density map representing the structure of the target matter are obtained. Based on the reference density map, the noise density map is converted into a target density map representing the structure of the target matter. The noise density map, the reference density map, and the target density map respectively indicate the distribution of matter particles, and the reference density map is used to correct the motion of matter particles during the conversion.

[0039] According to the scheme disclosed herein, a noise density map can be transformed towards a reference density map to obtain a target density map. In this way, the noise density map can be used as a condition to generate the posterior distribution of the target density map, resulting in a more accurate target density map and thus enabling a more precise reconstruction of the structure of the target material.

[0040] To better understand the embodiments of this disclosure, the basic concepts of flow matching and diffusion posterior sampling are first introduced.

[0041] Flow matching transforms samples in the noise distribution step by step. To define data points The generation process. This process is modeled based on ordinary differential equations (ODEs): dx t =v θ (t,x t )dt (1)

[0042] in is a time-dependent vector field parameterized by neural network Θ, which represents the sample velocity at some intermediate time step t.

[0043] The vector field can generate a probability density path p t (x t ), which reshapes the noise distribution p1(x1) to the data distribution p0(x0). Equation (1) can be solved via a differentiable neural ODE solver. In some examples, it is more efficient to learn a manually designed conditional vector field given a data point x0. The design principle of the vector field u|x0is to simplify the generated flow trajectory. For example, the conditional flow trajectory can be a straight line between x0and a noise sample x1, which is represented as follows: t |x0= (1 - t)x0+ tx1(2)

[0044] where the corresponding vector field is represented as: t u(t, x|x0) = x0- x1(3)

[0045] Therefore, the goal of (conditional) flow matching can be represented as follows:

[0046] Diffusion posterior sampling has become a method to solve inverse problems. Given a partial measurement y derived from some degenerate operator x can be sampled from the posterior distribution p(x|y). With a diffusion model that learns the score of the prior distribution, a likelihood term can be plugged in so that the score of the posterior distribution can be obtained, which can be represented as follows:

[0047] However, the computation of log p t (y|x t ) is usually difficult because the tractable likelihood is only defined in the data space (i.e., t = 0). The likelihood term requires marginalization over all possible x0∈ p0(x0). In some examples, a Laplace approximation of the likelihood term can be used, so that In this way, the conditional score can be approximated as:

[0048] If the observation distribution p0(y|x0) is assumed to be a Gaussian distribution, the derivative of the log probability will yield:

[0049] Putting all of the above together, the score of the posterior distribution can be approximated by the following equation:

[0050] where λ t denotes a hyperparameter controlling the step size of the control likelihood term.

[0051] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0052] FIG. 2 shows a schematic diagram 200 of a process of reconstructing a structure of a material according to some embodiments of the present disclosure. As shown in FIG. 2, to reconstruct the structure of a target material, a noisy density map 205 (denoted by xi) related to the structure of the material and a reference density map 210 (denoted by y) representing the structure of the target material can be first acquired. In some embodiments, the noisy density map 205 and the reference density map 210 can be three-dimensional volume maps.

[0053] In some embodiments, the reference density map 210 is a protein density map generated based on data collected by cryo-EM. In some examples, the density map 215 is a zoom-in of a portion of the reference density map 210. As shown in the density map 215, the protein density map generated based on data collected by cryo-EM is not accurate and does not accurately represent the structure of the protein (as an example of the target material). Therefore, the structure of the protein can be reconstructed based on the noisy density map 205 and the reference density map 210, thereby improving the accuracy of the density map generated by cryo-EM.

[0054] After the noisy density map 205 and the reference density map 210 are acquired, the noisy density map 205 can be converted into a target density map 220 (denoted by xo) representing the structure of the target material based on the reference density map 210. The noisy density map 205, the reference density map 210 and the target density map 220 respectively indicate the distribution of particles (e.g., electrons) of the material. The reference density map 210 is used to correct the movement of the particles of the material in the conversion. In some examples, the conversion of the noisy density map 205 into the target density map 220 is a denoising operation performed on the noisy density map 205, thereby obtaining the target density map 220. In some examples, the density map 222 is a zoom-in of a portion of the target density map 220. As shown in the density map 222, the reconstructed density map can accurately represent the structure of the protein, resulting in a more accurate target density map 220.

[0055] In some embodiments, converting the noise density map to the target density map includes a plurality of conversion steps, in a first conversion step of the plurality of conversion steps, an input density map of the first conversion step can be converted to an intermediate density map, the input density map being the noise density map or an output density map of a conversion step before the first conversion step. For example, in the case that the first conversion step is step 225, the input density map of the first conversion step is the noise density map. In the case that the first conversion step is step 230, the input density map of the first conversion step is an output density map of a conversion step before the first conversion step.

[0056] After obtaining the intermediate density map, correction information (e.g., correction information 240) for the motion of the material particles in the intermediate density map can be determined based on the difference between the intermediate density map and the reference density map. In some examples, the correction information indicates an adjustment to the velocity of the motion of the material particles. The motion of the material particles can be represented by a vector field.

[0057] The intermediate density map can then be converted to an output density map of the first conversion step based on the correction information, as an input density map or a target density map for a second conversion step after the first conversion step. For example, in the case that the first conversion step is step 225, the output density map of the first conversion step can be an input density map for a second conversion step (e.g., step 230). In the case that the first conversion step is step 235, the output density map of the first conversion step can be a target density map. In this way, based on the difference between the given reference density map and the intermediate density map, the intermediate density map can be corrected to guide the first conversion step to output the density map in the correct direction, such that the output density map of the first conversion step is more accurate.

[0058] Given a vector field v t (x t ) that generates a prior distribution p0(x0), this vector field can be converted to a vector field v t (x t |y) that generates a posterior distribution p0(x0|y), i.e., the motion of the material particles with added correction information. The vector field v t (x t |y) is also called a conditional vector field, where y represents the condition (i.e., the reference density map). First, the probability flow ODE converts the vector field v t (x t ) to the vector field v t (x t |y):

[0059] where f t represents the intercept term, g tdenotes the diffusion coefficient. Equation (9) is the inverse-time SDE with the same marginal as the standard form of the stochastic differential equation (SDE) dx t = f t (x t )dt + g t dw

[0060] Inserting the condition (i.e., y) into both sides of equation (10) gives

[0061] Equation (11) motivates adding a likelihood term weighted by the diffusion coefficient to generate a vector field that models the posterior distribution. The following settings can be made for the flow trajectory defined in equation (2): and Thus the conditional vector field can be expressed as

[0062] In computing the difference between the intermediate density map and the reference density map, attention can be focused on the difference between the features of the reference density map and the features of the intermediate density map. In some examples, the features can include texture features, shape features, keypoint features, and the like.

[0063] In some embodiments, a first feature position where a reference feature in the reference density map is located can be determined. For example, if the reference feature in the reference density map is at the top-left corner, the first feature position can be determined to be the top-left corner.

[0064] In some embodiments, the first feature position can be determined based on a generation mode of the reference density map. For example, if the generation mode of the reference density map includes a wedge-shaped missing generation mode, the reference feature of the reference density map will be generated within a certain angle, and thus the first feature position can be determined to be within the angle.

[0065] After the first feature position is determined, an intermediate feature can be extracted from the intermediate density map at a second feature position opposite the first feature position. The intermediate feature corresponds to the reference feature. Then, based on the difference between the reference feature and the intermediate feature, the correction information can be determined.

[0066] The likelihood term in equation (12) can be approximated by Laplace approximation to obtain

[0067] where denotes the extraction of the intermediate feature from the second feature position of the intermediate density map, and y denotes the reference feature of the reference density map.

[0068] After the target density map is generated, the quality of the target density map can also be evaluated. In some examples, the quality of the target density map can be evaluated using Fourier Shell Correlation (FSC), which compares the target density map and its ground truth in Fourier space.

[0069] In some embodiments, a ground truth density map for the target density map can be obtained. The ground truth density map is converted to Fourier space by Fourier transform and added with a noise signal to obtain a reference density map.

[0070] In some embodiments, the noise signal can include a first noise signal (also referred to as Gaussian noise), indicating the same level of noise added in each direction in the Fourier-transformed ground truth density map. In cryo-EM reconstruction, the spectrum noise can capture different noise characteristics at different spatial frequencies, which are caused by factors such as contrast transfer function (CTF) and detector defects. Consider introducing noise in the Fourier domain, where the variance of the added Gaussian noise is related to the frequency, the higher the frequency, the larger the noise variance. The Gaussian noise can be added uniformly to each direction in the Fourier-transformed ground truth density map.

[0071] Given a density map in Fourier domain Observation model which can be represented as follows:

[0072] where ∈ R D×D×D represents the noise amount, the value on the spherical shell with the same radius v is the same (v represents an index representing a component in the frequency space), and the distribution of ε can be represented as follows:

[0073] Alternatively or additionally, the noise signal can include a second noise signal indicating different levels of noise added in different directions in the true value density map after Fourier transform. In some embodiments, the second noise signal includes anisotropic noise. Anisotropic noise is generated when the distribution of noise has a direction dependence, affecting some directions more than others. FIG. 3A illustrates a schematic diagram 300A of anisotropic noise, according to some embodiments of the present disclosure. As shown in FIG. 3A, the noise level added in regions 302 and 304 is greater than the noise level added in regions 306 and 308. Anisotropic noise often causes a preferred direction problem, where particles in the sample tend to adopt a particular direction more frequently than others during the data collection process, resulting in uneven view distribution. This leads to over-represented orientations with better reconstruction quality and under-represented orientations with higher noise and lower resolution. To approximate this degradation, the noise can be amplified by a factor when the particle direction falls within a certain angular range, simulating the increased uncertainty in directions that are sampled less. Observation model may be represented as follows:

[0074] where ε ∈ R D×D×D As defined in equation (15), a > 1 is a scaler that increases the noise in certain directional angles, reflecting the increased uncertainty in these directions. min and θ max represent the angles that control the part of the noise that is amplified in FIG. 3A.

[0075] In some embodiments, the second noise signal includes a wedge missing repair noise signal. The wedge missing utility on a sub-tomogram is simulated by applying a wedge mask determined by a range of tilt angles to the density in the Fourier domain, effectively removing data from the unmeasured data. Observation model may be represented as follows:

[0076] where θ min and θ max are typically set to -60° and +60°, respectively. FIG. 3B illustrates a schematic diagram 300B of a wedge missing repair noise signal, according to some embodiments of the present disclosure. As shown in FIG. 3B, noise is added in regions 352 and 354.

[0077] In some embodiments, the noise density map 205 can be converted to the target density map 220 with a machine learning model. The machine learning model can be trained based on a predetermined training objective with a plurality of reference density maps and sample noise density maps of corresponding samples of the reference material. The predetermined training objective is configured to cause the machine learning model to recover the sample density map from the sample noise density map. In some examples, the machine learning model can be first trained with the sample density map, which can cause the machine learning model to learn structural representation of the sample density map. Then, the machine learning model can be trained to recover the sample density map from the sample noise density map.

[0078] With reference back to FIG. 2, during the training of the machine learning model, without the observation value (e.g., the reference density map 210) that can be referred to, the machine learning model generates the sample density map 245 from the noise density map 205.

[0079] FIG. 4 shows a schematic diagram 400 of an example architecture of a machine learning model according to some embodiments of the present disclosure. As shown in FIG. 4, the input 402 of the machine learning model has a size of D, and after a down-sampling operation 404, an intermediate result 406 with a size of D / 4 can be obtained. In some examples, the input 402 is a three-dimensional density map, and the input 402 can have a size of 64x64x64. A neighbor attention operation 408 is performed on the intermediate result 406, and a down-sampling operation 410 is continued, and an intermediate result 412 with a size of D / 8 can be obtained. Next, a global attention operation 414 is performed on the intermediate result 412, and an up-sampling operation 416 is continued, and an intermediate result 418 with a size of D / 4 can be obtained. Then, a neighbor attention operation 420 is performed on the intermediate result 418, and an up-sampling operation 422 is continued, and an output 424 of the machine learning model with a size of D can be obtained. In some examples, the output 424 can be a vector field. It is noted that the architecture shown in FIG. 4 is only an example architecture of the machine learning model in the present disclosure, and other machine learning model architectures can also be used according to actual needs.

[0080] FIG. 5 shows a schematic diagram of a process 500 for reconstructing a material structure according to some embodiments of the present disclosure. The process 500 can be implemented at the electronic device 120 of FIG. 1.

[0081] At block 510, the electronic device 120 obtains a noise density map related to the material structure and a reference density map representing a structure of a target material.

[0082] At block 520, the electronic device 120 converts the noise density map into a target density map representing a structure of the target matter based on the reference density map, the noise density map, the reference density map and the target density map respectively indicating a distribution of matter particles, the reference density map being used to correct a motion of the matter particles in the conversion.

[0083] In some embodiments, the converting the noise density map into the target density map includes a plurality of conversion steps, and a first conversion step of the plurality of conversion steps includes: converting an input density map of the first conversion step into an intermediate density map, the input density map being the noise density map or an output density map of a conversion step before the first conversion step; determining correction information for a motion of matter particles in the intermediate density map based on a difference between the intermediate density map and the reference density map; and converting the intermediate density map into an output density map of the first conversion step based on the correction information, as an input density map of a second conversion step after the first conversion step or the target density map.

[0084] In some embodiments, the determining the correction information for the motion of the matter particles in the intermediate density map includes: determining a first feature position at which a reference feature in the reference density map is located; extracting an intermediate feature from the intermediate density map at a second feature position opposite to the first feature position; and determining the correction information based on a difference between the reference feature and the intermediate feature.

[0085] In some embodiments, the first feature position is determined based on a generation mode of the reference density map.

[0086] In some embodiments, the process 500 further includes obtaining a ground truth density map for the target density map, the ground truth density map being converted to a Fourier space by a Fourier transform and being added with a noise signal to obtain the reference density map; and determining a generation quality of the target density map by comparing the target density map and the ground truth density map in the Fourier space.

[0087] In some embodiments, the noise signal includes at least one of: a first noise signal indicating a same level of noise being added in each direction in the Fourier transformed ground truth density map, or a second noise signal indicating different levels of noise being added in different directions in the Fourier transformed ground truth density map.

[0088] In some embodiments, the reference density map is a protein density map generated based on data collected by cryo-electron microscopy.

[0089] In some embodiments, the noise density map is converted into the target density map using a machine learning model, and the machine learning model is trained by: using a plurality of reference matters with respective sample density maps and sample noise density maps, training the machine learning model based on a predetermined training target configured to cause the machine learning model to recover the sample density map from the sample noise density map.

[0090] FIG. 6 illustrates a block diagram of an apparatus 600 for reconstructing a structure of a material, according to some embodiments of the present disclosure. The apparatus 600 can be implemented at or included in the electronic device 120 of FIG. 1. Various modules / components in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0091] As shown, the apparatus 600 includes a density map obtaining module 610 configured to obtain a noisy density map related to a structure of a material and a reference density map representing a structure of a target material. The apparatus 600 further includes a density map converting module 620 configured to convert, based on the reference density map, the noisy density map to a target density map representing the structure of the target material, the noisy density map, the reference density map, and the target density map respectively indicating a distribution of material particles, the reference density map being used to correct a motion of the material particles in the conversion.

[0092] In some embodiments, the density map converting module 620 is further configured to convert an input density map of a first conversion step to an intermediate density map, the input density map being the noisy density map or an output density map of a conversion step before the first conversion step; determine correction information for a motion of material particles in the intermediate density map based on a difference between the intermediate density map and the reference density map; and convert the intermediate density map to an output density map of the first conversion step based on the correction information as an input density map of a second conversion step after the first conversion step or the target density map.

[0093] In some embodiments, the density map converting module 620 is further configured to determine a first feature position at which a reference feature in the reference density map is located; extract an intermediate feature from the intermediate density map at a second feature position opposite to the first feature position; and determine the correction information based on a difference between the reference feature and the intermediate feature.

[0094] In some embodiments, the first feature position is determined based on a generation mode of the reference density map.

[0095] In some embodiments, the apparatus 600 further includes a quality evaluating module configured to obtain a ground truth density map for the target density map, the ground truth density map being converted to a Fourier space by a Fourier transform and being added with a noise signal to obtain the reference density map; and determine a generation quality of the target density map by comparing the target density map and the ground truth density map in the Fourier space.

[0096] In some embodiments, the noise signal comprises at least one of: a first noise signal indicating a same level of noise added in each direction in the Fourier transformed ground truth density map, or a second noise signal indicating different levels of noise added in different directions in the Fourier transformed ground truth density map.

[0097] In some embodiments, the reference density map is a protein density map generated based on data collected by cryo-electron microscopy.

[0098] In some embodiments, the noise density map is converted into the target density map by a machine learning model, and the machine learning model is trained by: training the machine learning model based on a predetermined training objective configured to cause the machine learning model to recover the sample density map from the sample noise density map, by utilizing respective sample density maps and sample noise density maps of a plurality of reference substances.

[0099] FIG. 7 illustrates a block diagram of a device 700 that is capable of implementing a number of implementations of the present disclosure. It should be understood that the computing device 700 illustrated in FIG. 7 is merely an example and should not be construed as any limitation of the functionality and scope of the implementations described herein. The computing device 700 illustrated in FIG. 7 can be used to implement the methods described above.

[0100] As illustrated in FIG. 7, the computing device 700 is in the form of a general- purpose computing device. Components of the computing device 900 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and is capable of executing a variety of processing functions according to the programs stored in the memory 720. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of the computing device 700.

[0101] The computing device 700 typically includes a plurality of computer storage media. Such media can be volatile and nonvolatile media and removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and other data. The memory 720 can be volatile memory (such as registers, cache, random access memory (RAM)), non-volatile memory (such as read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable media and can include machine readable media such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible by the computing device 700.

[0102] The computing device 700 can further include additional removable / non-removable, volatile / non-volatile storage devices. Although not shown, a floppy disk drive for reading from or writing to a removable, non-removable magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-removable optical disk (e.g., a CD-ROM) can be provided. In such instances, each drive can be connected to the bus by one or more data media interfaces. The storage device 720 can include a computer-program product that has one or more program modules configured to carry out the various methods or actions of the implementations of the present disclosure.

[0103] The communication unit 740 enables communications with other computing devices over a communication medium. Additionally, the functionality of the components of the computing device 700 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0104] The input device(s) 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device(s) 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 740, based on a requirement. The external devices can include one or more devices that enable a user to interact with the computing device 700, or any devices (e.g., a network card, a modem, etc.) that enable the computing device 700 to communicate with one or more other computing devices. Such communication can be carried out through an input / output (I / O) interface (not shown).

[0105] According to an example implementation of the present disclosure, a computer-readable storage medium is provided having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.

[0106] The computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all.

[0107] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all.

[0108] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all.

[0109] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0110] Having described various implementations of the disclosure above, the descriptions are not exhaustive and do not limit the disclosure to the disclosed implementations. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The scope of the disclosure includes all the implementations of which an equivalent would be apparent to those skilled in the art from the disclosure given and the associated drawings. The selection of the terms to be used in the written description is not intended to limit the scope of the present disclosure, but rather to best describe the principles of the various implementations in preference to a careful drawing of equivalent alternatives which are to be substituted for the preferred implementations illustrated and described.

Claims

1. A method of reconstructing a structure of a material, comprising: obtaining a noisy density map related to the structure of the material and a reference density map representing a structure of a target material; and converting the noisy density map into a target density map representing the structure of the target material based on the reference density map, the noisy density map, the reference density map and the target density map respectively indicating a distribution of material particles, the reference density map being used to correct a movement of material particles in the converting.

2. The method of claim 1, wherein converting the noisy density map into the target density map comprises a plurality of converting steps, and a first converting step of the plurality of converting steps comprises: converting an input density map of the first converting step into an intermediate density map, the input density map being the noisy density map or an output density map of a converting step before the first converting step; determining correction information for a movement of material particles in the intermediate density map based on a difference between the intermediate density map and the reference density map; and converting the intermediate density map into an output density map of the first converting step based on the correction information, as an input density map of a second converting step after the first converting step or the target density map.

3. The method of claim 2, wherein determining the correction information for the movement of material particles in the intermediate density map comprises: determining a first feature position at which a reference feature in the reference density map is located; extracting an intermediate feature from the intermediate density map at a second feature position opposite to the first feature position; and determining the correction information based on a difference between the reference feature and the intermediate feature.

4. The method of claim 3, wherein the first feature position is determined based on a generation mode of the reference density map.

5. The method of claim 1, further comprising: obtaining a ground truth density map for the target density map, the ground truth density map being converted to a Fourier space by a Fourier transform and being added with a noise signal to obtain the reference density map; and determining a generation quality of the target density map by comparing the target density map and the ground truth density map in the Fourier space.

6. The method of claim 4, wherein the noise signal comprises at least one of: a first noise signal indicating a same level of noise being added in each direction in the Fourier-transformed ground truth density map, or a second noise signal indicating different levels of noise being added in different directions in the Fourier-transformed ground truth density map.

7. The method of claim 1, wherein the reference density map is a protein density map generated based on data collected by cryo-electron microscopy.

8. The method of claim 1, wherein the noisy density map is converted into the target density map using a machine learning model, and the machine learning model is trained by: The machine learning model is trained based on a predetermined training objective configured to cause the machine learning model to recover the sample density map from the sample noise density map.

9. An apparatus for reconstructing a structure of a substance, comprising: a density map obtaining module configured to obtain a noise density map related to a structure of a substance and a reference density map representing a structure of a target substance; and a density map converting module configured to convert the noise density map into a target density map representing the structure of the target substance based on the reference density map, the noise density map, the reference density map and the target density map respectively indicating a distribution of substance particles, the reference density map being used to correct a motion of the substance particles in the conversion.

10. An electronic device, comprising: a set of processing units; and a set of memories coupled to the set of processing units and storing instructions for execution by the set of processing units, the instructions, when executed by the set of processing units, cause the electronic device to perform the method according to any one of claims 1 to 8.

11. A computer readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Local quality assessment method based on deep learning cryoelectron microscope three-dimensional density map

    CN116012537A

  • Particle screening method, device and equipment in cryoelectron microscope imaging and storage medium

    CN116434827A

  • Cryoelectron microscope image processing method and device, terminal and storage medium

    CN117197114A

  • Cryoelectron microscope image generation method, system and equipment guided by physical information, chip and medium

    CN118411439A

  • Cryo-electron microscope protein model building method based on neural network, and storage medium

    WO2024119597A1