Three-dimensional Gaussian modeling method for analyzing biomacromolecule components and conformational heterogeneity

By employing a 3D Gaussian modeling method and a dual encoder-single decoder architecture, the compositional and conformational heterogeneity of biomacromolecules is analyzed. This overcomes the limitations of cryo-electron microscopy in high-resolution reconstruction and complex motion processing, and achieves a high-fidelity, intuitive description of structural changes.

CN121564221APending Publication Date: 2026-02-24SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511748157.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing cryo-electron microscopy techniques struggle to achieve high-fidelity, intuitive joint analysis when dealing with the composition and conformational heterogeneity of biological macromolecules, and traditional methods have limitations in high-resolution reconstruction and handling complex motions.

Method used

A three-dimensional Gaussian modeling method is adopted, which combines an image encoder, a Gaussian encoder and a deformable decoder through a dual encoder-single decoder architecture. The multilayer perceptron is used to realize the variation of Gaussian parameters to represent the components and conformational heterogeneity. A geometric constraint loss function is designed to ensure the accuracy and physical rationality of the model.

Benefits of technology

It enables intuitive analysis of the compositional and conformational heterogeneity of biological macromolecules, supports high-resolution reconstruction and modeling of complex motions, and can naturally connect density models and atomic models, thus improving the processing capability of large-scale datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564221A_ABST
    Figure CN121564221A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional Gaussian modeling method for analyzing biomacromolecule components and conformation heterogeneity, and belongs to the technical field of cryoelectron microscope image processing. According to the method, the component and conformation heterogeneity can be directly represented through the change of Gaussian parameters, and an intuitive and explainable structure change description is provided; the consistency and integrity of the local structure along the transformation trajectory are ensured through geometric constraint; modeling of tens of thousands of Gaussian models on a pseudo-atom level is supported; gaussian variation is converted into an atomic conformation landscape through natural support, and a direct bridge is provided between density-based analysis and pseudo-atom level interpretation; the method is realized through an efficient Gaussian projection algorithm and parallel computing, and the processing capacity of a large-scale data set is remarkably improved. According to the method, the component and conformation heterogeneity can be effectively processed at the same time, the continuous conformation change is visually represented, the local structure consistency is kept, the density model and the atomic model are naturally connected, and the defects in the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cryo-electron microscopy image processing technology, and in particular to a three-dimensional Gaussian modeling method for resolving the composition and conformational heterogeneity of biological macromolecules from cryo-electron microscopy single-particle analysis data. Background Technology

[0002] The function of biomacromolecules is closely related to the heterogeneity of their dynamic structures, namely, the heterogeneity of their components (such as the assembly of subunits) and conformations (such as the relative motion of domains). Cryo-electron microscopy (cryo-EM) enables imaging of biomacromolecules at near-atomic resolution, providing the possibility of directly observing their dynamic processes. However, cryo-EM experimental images have extremely low signal-to-noise ratios and often contain mixed information of molecules in different states (discrete or continuously changing), making the resolution of structural heterogeneity from these images a significant challenge. Traditional single-particle analysis methods typically rely on classification and averaging techniques to group images of particles with similar structures into one class and then perform 3D reconstruction. These methods are effective for datasets with discrete-state heterogeneity but struggle to handle highly flexible molecules or continuous conformational changes. Many-body refinement techniques improve the estimation of local changes by dividing molecules into multiple rigid parts, but their limitation on the size of individual parts restricts their ability to handle highly flexible systems or small subunit motions. In recent years, neural network-based methods have been introduced to address the problem of continuous heterogeneity. For example, methods like CryoDRGN utilize the Neural Radiation Field (NeRF) framework to map projected images into a continuous latent space, thus revealing complex structural changes without predefined discrete categories. However, these NeRF-based methods inherently require explicit 3D volume reconstruction for image comparison, cannot directly visualize continuous structural transformation paths, and heavily rely on manual inspection for interpreting conformational landscapes, making the process time-consuming. On the other hand, methods based on Gaussian Mixture Models (GMMs), such as e2gmm, use Gaussian spheres to represent protein structures, intuitively describing structural heterogeneity through changes in Gaussian parameters. However, these methods are limited in handling the number of Gaussian spheres, making it difficult to support the large number of Gaussian models required for high-resolution reconstruction, and lack sensitivity to complex motions. Some subsequent methods (such as 3DFlex) focus on estimating continuous motion through deformation fields, but their ability to characterize component heterogeneity is limited, and the intuitive interpretation of discrete and continuous changes remains challenging. In summary, existing technologies, when dealing with the heterogeneity of cryo-EM data, either lack the intuitiveness of continuous changes and the preservation of local structure, or are deficient in the ability to handle the heterogeneity of mixed components and conformations, or are limited by the model representation capabilities and cannot simultaneously handle high resolution and complex motion.

[0003] Therefore, there is an urgent need in this field for a new method that can intuitively and faithfully combine the analysis of compositional and conformational heterogeneity, and can naturally connect density models and atomic models. Summary of the Invention

[0004] The purpose of this invention is to provide a three-dimensional Gaussian modeling method for analyzing the composition and conformational heterogeneity of biological macromolecules, so as to overcome the limitations of existing technologies in cryo-electron microscopy heterogeneity analysis.

[0005] To achieve the above objectives, the specific technical solution adopted by the present invention is as follows:

[0006] A three-dimensional Gaussian modeling method for resolving the compositional and conformational heterogeneity of biological macromolecules, comprising the following steps:

[0007] S1: Obtain protein density map data;

[0008] S2: Extract attitude parameter features;

[0009] S3: Calculate the contrast transfer function to obtain the CTF parameter file;

[0010] S4: Generate Gaussian point cloud coordinates and density values;

[0011] S5: Standardize the image data;

[0012] S6: Construct a heterogeneous reconstruction network architecture; this network architecture adopts a dual encoder-single decoder architecture, including an image encoder, a Gaussian encoder, and a deformable decoder, and all modules are implemented using a multilayer perceptron with ReLU activation function;

[0013] S7: Based on the heterogeneous reconstruction network, image data processing is performed to achieve 3D modeling.

[0014] Furthermore, in S1, the data preprocessing process begins with the parsing of command-line parameters and the initialization of the computing environment; let the set of input parameters be:

[0015] Θ={θ source ,θ data ,θ star ,θ size ,θ downsample ,θ apix ,θ map ,θ thres ,θ output};

[0016] The parameters represent the source directory, data directory, star file path, image size, downsampling size, pixel size, consensus density map, threshold, and output directory, respectively.

[0017] Furthermore, in S2, the particle attitude parameters are analyzed, including Euler angles and translation vectors; for the i-th particle, the attitude parameter extraction process is as follows:

[0018] Euler angle extraction:

[0019] φ i =(α i ,β i ,γ i )=(rlnAngleRot i ,rlnAngleTilt i ,rlnAnglePsi i );

[0020] Convert Euler angles to rotation matrices using rotation matrix transformation functions:

[0021]

[0022] Translation parameters are processed according to units; if the translation parameters are in angstroms, they are converted to pixel units.

[0023]

[0024] Finally, the complete attitude parameters are obtained:

[0025] Ψ i =[R i ,t i ];

[0026] Downsample and scale the translation parameters:

[0027]

[0028] Integrate the rotation matrix and translation vector into complete attitude parameters:

[0029] Ψ i ′=[vec(R i ),t i ′]

[0030] Where vec(·) means expanding the matrix into a vector.

[0031] Furthermore, in S3, based on the classical CTF theoretical model, the electron wavelength is first calculated:

[0032]

[0033] Where V is the accelerating voltage (volts);

[0034] For the spatial frequency vector f = (f x ,f y), Calculate the azimuth:

[0035] θ = arctan2(f y ,f x )

[0036] Effective defocus calculation is as follows:

[0037]

[0038] The phase offset term is:

[0039]

[0040] The final CTF function is:

[0041]

[0042] Optionally, the envelope function can be applied:

[0043]

[0044] Furthermore, in S4, the process of generating the initial Gaussian point cloud from the consensus density map is based on grid sampling and threshold filtering; let the density map be V(x,y,z), and the sampling interval be Δ=(Δ x ,Δ y ,Δ z );

[0045] Generate sampling grid:

[0046]

[0047] Apply density threshold screening:

[0048]

[0049] Optionally, downsampling can be performed when the number of points exceeds the target number N. target hour:

[0050]

[0051] Furthermore, in S5:

[0052]

[0053] Where μ i and σ i denoted as the mean and standard deviation of the i-th image, respectively.

[0054] The parameter files output by the S1-S5 data preprocessing include: attitude parameter file: CTF parameter file: Standardized image data; Gaussian point cloud coordinates and density values.

[0055] Furthermore, in step S6, the heterogeneous reconstruction network adopts a deep generative model architecture based on variational autoencoders, including the following modules:

[0056] (1) Basic building blocks

[0057] • ResidLinear: Residual linear layer, implementing residual connections of z = Linear(x) + x;

[0058] • MLP: Standard Multilayer Perceptron, which includes linear layers, activation functions, and normalization layers;

[0059] • ResidLinearMLP: A multilayer perceptron based on residual connections to enhance gradient flow;

[0060] (2) Core Network Architecture The network body HetGMMVAESimpleIndependent contains the following key components:

[0061] Image encoder: Input image Mapping to the latent space:

[0062] z μ ,z σ =E image (I)

[0063] z = z μ +∈·exp(0.5×z σ ),∈~N(0,1)

[0064] The encoder can be selected from either a standard MLP or a residual MLP architecture, with an output dimension of 2×zdim, corresponding to the logarithm of the mean and variance of the latent variables, respectively.

[0065] Gaussian encoder: processes parameter information for each Gaussian point.

[0066] g i =[d i ,s i ,p i (Density, scale, location)

[0067] e i =E gaussian (g i )

[0068] The Gaussian encoder maps the 5-dimensional parameters of each Gaussian (1-dimensional density, 1-dimensional scale, and 3-dimensional position) to a high-dimensional embedding space. The feature decoder integrates latent variables and Gaussian embeddings.

[0069] (Broadcast to all Gaussian points)

[0070] h = D feature ([z exp ,e])

[0071] Δd i =f d (h i ),Δp i =f p (h i ),Δs i =f s (h i )

[0072] Where f d ,f p ,f s These correspond to prediction networks for changes in density, location, and scale, respectively.

[0073] Furthermore, S7 specifically includes:

[0074] S7-1: Let the input image be... Where D is the image size, the network first passes through an image encoder Mapping the image to the latent space:

[0075]

[0076] in For d z θ-dimensional latent vector image These are encoder parameters.

[0077] Meanwhile, each Gaussian point i is encoded by a Gaussian encoder. Generate the corresponding embedding vector:

[0078]

[0079] Where g i =[d i ,s i ,p i Includes Gaussian density, scale, and location parameters;

[0080] Deformation Decoder By combining latent variables of the image with Gaussian embeddings, the parameter variations of each Gaussian are estimated:

[0081]

[0082] S7-2: The training process uses mini-batch gradient descent, with each batch containing B samples;

[0083] S7-3: Design a multi-objective loss function;

[0084] S7-4: The AdamW optimizer is used for parameter updates during the training process, and the learning rate adopts an exponential decay strategy.

[0085] Furthermore, S7-2 specifically refers to the following process for the j-th sample in the batch:

[0086] First, the image encoder processes the input image I. j Generate the corresponding latent representation z j This step enables the transformation from pixel space to semantic space, capturing the structural heterogeneity information contained in the image.

[0087] Next, for each Gaussian point i, the deformation decoder is based on z j and Gaussian embedding e i Calculate the parameter change Δg ij This process can be decomposed into a density change Δd. ij and position change Δp ij The estimate.

[0088] Then, the estimated parameter changes are applied to the initial Gaussian model. We obtain a structurally deformed version for the current image:

[0089]

[0090] Based on the deformed Gaussian model and the corresponding particle attitude parameter Ψ j =(R j ,t j The process involves generating a reconstructed image through 3D-to-2D projection rendering. The rendering process includes steps such as rigid body transformation, projection integration, and CTF modulation, ultimately resulting in a synthetic image of the same size as the original image.

[0091] Furthermore, S7-3 specifically includes:

[0092] To ensure the stability of the training process and the quality of reconstruction, a composite loss function with multiple constraint terms was designed. The reconstruction loss measures the difference between the synthetic image and the real image.

[0093]

[0094] The variational regularization term constrains the distribution of the latent space through KL divergence:

[0095]

[0096] Geometric constraints ensure that the learned structural changes are physically plausible. Neighbor embedding consistency constraints encourage spatially adjacent Gaussians to have similar embedding representations.

[0097]

[0098] Local rigid constraints maintain the continuity of the structure after deformation:

[0099]

[0100] Displacement direction consistency constraints ensure motion coordination between adjacent regions:

[0101]

[0102] Exclusion constraints prevent excessive clustering of Gaussian points:

[0103]

[0104] The latent spatial consistency constraint maintains the local smoothness of the latent representation through K-nearest neighbor loss:

[0105]

[0106] The total loss function is the weighted sum of the individual components:

[0107]

[0108] Furthermore, in S7-4, the training process uses the AdamW optimizer for parameter updates, and its update rule is as follows:

[0109]

[0110] The learning rate employs an exponential decay strategy:

[0111]

[0112] In actual training, a progressive training strategy is adopted. Initially, the focus is on optimizing the reconstruction loss and basic geometric constraints, while more complex regularization terms are gradually introduced as training progresses. This strategy helps prevent the model from getting trapped in local optima too early, improving the stability and convergence of the training.

[0113] After training using the above method, the model performance is evaluated by assessing the reconstruction quality on the validation set. Key evaluation metrics include the structural similarity between the reconstructed image and the real image, the clustering effect of the latent space, and the physical plausibility of the generated structure.

[0114] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0115] This invention can directly characterize compositional and conformational heterogeneity through changes in Gaussian parameters, providing an intuitive and interpretable description of structural changes; it ensures the consistency and integrity of local structures along the transformation trajectory through geometric constraints; it supports modeling tens of thousands of Gaussian models at the pseudo-atomic level, overcoming the limitation of traditional Gaussian methods on the number of Gaussian models; when atomic models are available, it naturally supports the transformation of Gaussian changes into atomic conformational landscapes, providing a direct bridge between density-based analysis and pseudo-atomic level interpretation; and through efficient Gaussian projection algorithms and parallel computing, it significantly improves the processing capacity of large-scale datasets. Attached Figure Description

[0116] Figure 1 This is the architecture of the present invention.

[0117] Figure 2 To capture large-scale motion in a simulated dataset for this invention; a) Consistent mapping and atomic model from the simulated dataset; b) Dimensionally reduced two-dimensional latent conformation space; c) Estimated Gaussian displacement directly mapped to atomic coordinates (red); d) Visualization of the reconstructed atomic model.

[0118] Figure 3 Analysis of heterojunctions in ribosome assembly; a) Visualization of the latent space encoded by this invention, where dark blue dots represent cluster centers corresponding to different conformational classes. b) Ten representative density maps generated from the cluster centers by this invention, each labeled with its class index. Red boxes highlight regions showing major conformational differences. c) Examples of present ribosomal proteins L16 and L25, partially present and completely absent. d) Visualization of the movement of the dynamic central bulge domain and the corresponding changes in the atomic model and CryoSPARC reconstruction plot. Fixed dotted envelopes are used to better illustrate structural movement. e) Distribution of the learned Gaussian embeddings.

[0119] Figure 4 Analysis of heterogeneity of the pre-catalytic spliceosome: a) Visualization of the latent space encoded by this invention; b) Six representative density maps generated from cluster centers by this invention, each labeled with its category index. Red boxes highlight regions showing major conformational differences. c) Map reconstructed using the clustering projection from this invention via a freezing method. d) Visualization of the segmented Gaussians. e) UMAP visualization of components of different colors, roughly responding to different principal components 1, 2, and 3. The reconstructed volumes and their corresponding atomic models are shown on the right. Detailed Implementation

[0120] To make the objectives, technical solutions, and advantages of this invention clearer, the following detailed description, in conjunction with specific embodiments and accompanying drawings, further illustrates the invention. Obviously, the described embodiments are only a portion, not all, of the embodiments disclosed in this invention. All other embodiments obtained by those skilled in the art based on the embodiments disclosed in this invention without inventive effort are within the scope of protection of this invention.

[0121] This invention employs a three-dimensional Gaussian function as a pseudo-atomic representation of protein structure and uses an innovative dual encoder-single decoder architecture to jointly estimate the changes in Gaussian parameters, thereby achieving accurate modeling of structural heterogeneity in cryo-electron microscopy images.

[0122] 1. Representation method of three-dimensional Gaussian density map

[0123] Representing protein density maps as learnable three-dimensional Gaussian sets Each Gaussian function The mathematical expression describing a local Gaussian shape density distribution is:

[0124]

[0125] Where d i p represents the density value. i Represents the position vector, Σ i Let represent the covariance matrix. Considering that protein molecules are composed of atomic spheres and are imaged at the same scale, the covariance matrix Σ is... i Simplified to scalar deviation s i Combined form with identity matrix I: This simplifies the anisotropic Gaussian to an isotropic Gaussian.

[0126] The target density value at any location x in three-dimensional space is the sum of the contributions of all nearby Gaussian functions, calculated using the following formula:

[0127]

[0128] Within this representational framework, component heterogeneity is expressed through the Gaussian density value d. i The changes are modeled, while conformational heterogeneity is modeled through Gaussian positions p. i Model the changes.

[0129] 2. Gaussian projection rendering mechanism

[0130] By leveraging the property that a Gaussian distribution can be obtained by integrating along the coordinate axes to obtain a two-dimensional Gaussian distribution, efficient computation from a three-dimensional Gaussian distribution to a two-dimensional projection can be achieved. Specifically, the two-dimensional Gaussian distribution is obtained by integrating the three-dimensional Gaussian distribution along the z-axis, and the parameters of the projected two-dimensional Gaussian distribution are... and Obtained by directly discarding specific rows and columns of the corresponding 3D version.

[0131] Rendered particle projection I k ∈R H×W Through the complete rendering process The mathematical expression for implementation is:

[0132]

[0133] Where C k This represents the contrast transfer function (CTF) modulator, where P represents the projection operator along the z-axis, and T represents the particle attitude-based modulator φ. k The transformation operation T is used to improve computational efficiency. The transformation operation T actually calculates the inverse matrix of the particle's attitude to align the observation direction with the coordinate axes.

[0134] 3. Heterogeneous Reconstruction Network Architecture

[0135] Employing a unique dual-encoder-single-decoder architecture, all modules are implemented using a multilayer perceptron with ReLU activation function, specifically including the following three main modules:

[0136] The image encoder consists of three hidden layers, responsible for converting the original particle image into i-dimensional latent variables. This encoder introduces a positional encoding mechanism to help the network determine particle pose and learn high-frequency features.

[0137] The Gaussian encoder assigns a unique identifier to each Gaussian and generates a j-dimensional latent embedding by taking the Gaussian's density, scale, position, and rotation parameters as input. This Gaussian-wise strategy has no limit on the number of Gaussians, allowing for the introduction of a given atomic model and obtaining more accurate parameter estimates.

[0138] The deformation decoder follows a Gaussian-by-Gaussian paradigm, receiving the concatenation result of the i-dimensional latent variables output by the image encoder and the j-dimensional Gaussian embeddings output by the Gaussian encoder, and outputs the predicted deformation for each Gaussian. This design allows for the flexible selection of Gaussians of interest and observation of their parameter variations without the need to re-estimate all Gaussians.

[0139] The estimated parameter variations are applied to an initial 3D Gaussian, followed by projection calculations based on the specified particle posture. The projection results of the deformed volume are then transformed to the Fourier domain and modulated using a contrast transfer function.

[0140] 4. Design of Geometric Constraint Loss Function

[0141] All loss functions are defined in real space, including reconstruction loss and geometric constraints, to ensure the accuracy and physical rationality of the model.

[0142] The reconstruction loss uses the L1 loss function to compare the deformed volume projection I′. k Compared with the original image y k Differences between them:

[0143]

[0144] in This represents the change in the Gaussian parameters.

[0145] The geometric constraint section comprises three sub-constraints designed to prevent overfitting and maintain the inherent continuity of protein motion:

[0146] Gaussian embedding constraints are based on the average nearest distance d. mean Considering the radius μ1 = 2.5 × d mean The nearest Gaussian within the range is achieved through the following formula:

[0147]

[0148] Among them, the weighting coefficient Used to control constraint strength, the closer the Gaussian pairs are, the greater the constraint strength.

[0149] Local rigid constraints enforce local rigidity in deformation by constraining the relative distance between adjacent Gaussians. in This represents the distance between Gaussians after deformation, with the threshold μ2 set to 1.5×d. mean .

[0150] Displacement direction constraints on displacement vector v n Additional constraints are applied to the direction of displacement to ensure consistency of displacement direction:

[0151]

[0152] In this invention, it is assumed that the heterogeneity present in the experimental particle images can be represented by a single consensus structure and its corresponding structural changes. Compositional heterogeneity manifests as partial structural loss, while conformational heterogeneity manifests as structural displacement. Figure 1This paper presents the overall workflow architecture of the invention, which uses 3D Gaussian representations of macromolecular structures. Its heterogeneity analysis network consists of two encoders and one decoder. Each particle image is encoded into a continuous latent space, while each Gaussian is independently encoded to obtain a unique identifier (embedding). The decoder then combines the latent variables and each Gaussian embedding to estimate the parameter variations of the Gaussian. This directly illustrates structural changes in proteins, capturing compositional and conformational heterogeneity. The method uses a set of 3D Gaussian functions to fit a consensus model and identifies parameter variations of the Gaussian functions by learning an adaptive variational autoencoder (VAE). Each Gaussian function is defined by seven parameters: a density value, a 3D scale vector for calculating its covariance, and 3D spatial coordinates. Local structural motion is described by the 3D positional variations of the Gaussian functions constituting the model, while the presence or absence of components is reflected by variations in their density values. Initial Gaussian functions are uniformly sampled from homogeneous reconstructions and filtered by provided contour levels. The number of Gaussian functions determines the achievable resolution of the reconstructed structure but also significantly impacts computational cost. In practice, this invention employs a sampling interval of two voxels, which performs well on a large number of experimental datasets.

[0153] The heterogeneous reconstruction network consists of two encoders (an image encoder and a Gaussian encoder) and a single decoder that integrates their representations. The image encoder processes a 2D particle image and transforms it into a continuous low-dimensional latent conformation space. Each point embedded in this low-dimensional manifold corresponds to a latent 3D conformation. Although the number of Gaussian functions is far less than the number of voxels, directly inputting all Gaussian parameters into the network remains computationally infeasible. To address this challenge, this invention introduces a learnable Gaussian encoding submodule, enabling the network to capture the contextual relationships between Gaussian functions. The Gaussian encoder takes attributes such as density, scale, and position for each Gaussian function and outputs a unique latent embedding as its identifier. Inspired by NeRF-based architectures, this method introduces positional encoding to help the network capture high-frequency spatial information, thereby enhancing its ability to reconstruct fine structural details. The latent variables from the image encoder are then concatenated with the Gaussian embeddings and passed to the decoder, which predicts the parameter variations for each Gaussian function. This Gaussian-wise paradigm enables efficient batch-level parallelization and facilitates the aggregation and differentiation of Gaussian functions based on the learned parameter variations.

[0154] The 3D Gaussian function is naturally defined in real space. Therefore, the difference between the experimental particle image and the rendered projection (given a known particle pose) is measured in the real space domain. Although two projection transformations are still required to apply the CTF modulator, comparisons in real space make the final structural changes more intuitive and interpretable. Once the network is trained, the learned conformational space can be visualized using standard dimensionality reduction techniques. By identifying cluster centers within this continuous latent space, representative 3D volumes can be generated to capture discrete conformational states. Furthermore, since the conformational manifold is continuous, smooth traversal between clusters can be achieved to produce dynamic structural transformations. Another advantage of the real-space Gaussian representation is its natural compatibility with atomic models. The shifts or omissions of the Gaussian function are easily mapped to atomic models via nearest-neighbor connections. This property enables the generation of deformable atomic models (PDB structures) corresponding to different conformational states inferred from cryo-electron microscopy data. Therefore, this invention provides an alternative perspective for linking density-based heterogeneity analysis with atomic-level structure.

[0155] Example 1:

[0156] I. Specific Implementation Steps for Data Preprocessing

[0157] This embodiment describes in detail the mathematical implementation process of the data preprocessing step in the method of the present invention, which is responsible for extracting two-dimensional particle images, particle orientation parameters, and contrast transfer function parameters from a .star file in RELION format.

[0158] Step 1: Parameter Initialization and Environment Configuration

[0159] The data preprocessing process begins with the parsing of command-line arguments and the initialization of the computation environment. Let the set of input arguments be:

[0160] Θ={θ source ,θ data ,θ star ,θ size ,θ downsample ,θ apix ,θ map ,θ thres ,θ output}

[0161] The parameters represent the source directory, data directory, star file path, image size, downsampling size, pixel size, consensus density map, threshold, and output directory, respectively.

[0162] Step 2: Attitude Parameter Analysis and Transformation

[0163] The particle attitude parameters, including Euler angles and translation vectors, are parsed from the RELION .star file. For the i-th particle, the attitude parameter extraction process is as follows:

[0164] Euler angle extraction:

[0165] φ i =(α i ,β i ,γ i )=(rlnAngleRot i ,rlnAngleTilt i ,rlnAnglePsi i )

[0166] Convert Euler angles to rotation matrices using rotation matrix transformation functions:

[0167]

[0168] Translation parameters are processed according to units. If the translation parameter is in angstroms, it is converted to pixel units:

[0169]

[0170] Finally, the complete attitude parameters are obtained:

[0171] Ψ i =[R i ,t i ]

[0172] Step 3: CTF Function Calculation

[0173] The contrast transfer function (CTF) is calculated based on the classic CTF theoretical model. First, the electron wavelength is calculated:

[0174]

[0175] Where V is the accelerating voltage (volts).

[0176] For the spatial frequency vector f = (f x ,f y ), Calculate the azimuth:

[0177] θ = arctan2(f y ,f x )

[0178] Effective defocus calculation is as follows:

[0179]

[0180] The phase offset term is:

[0181]

[0182] The final CTF function is:

[0183]

[0184] Optionally, the envelope function can be applied:

[0185]

[0186] Step 4: Gaussian Point Cloud Generation

[0187] The process of generating an initial Gaussian point cloud from the consensus density map is based on grid sampling and thresholding. Let the density map be V(x,y,z), and the sampling interval be Δ=(Δ x ,Δ y ,Δ z ).

[0188] Generate sampling grid:

[0189]

[0190] Apply density threshold screening:

[0191]

[0192] Optionally, downsampling can be performed when the number of points exceeds the target number N. target hour:

[0193]

[0194] Step 5: Coordinate Transformation and Parameter Integration

[0195] Downsample and scale the translation parameters:

[0196]

[0197] Integrate the rotation matrix and translation vector into complete attitude parameters:

[0198] Ψ i ′=[vec(R i ),t i ′]

[0199] Where vec(·) means expanding the matrix into a vector.

[0200] Step Six: Data Standardization and Output

[0201] Standardize the image data:

[0202]

[0203] Where μ i and σ i denoted as the mean and standard deviation of the i-th image, respectively.

[0204] The final output parameter file includes: Attitude parameter file: CTF parameter file: Standardized image data; Gaussian point cloud coordinates and density values. All output files use industry-standard formats to ensure compatibility with subsequent processing modules. The system also performs basic quality verification, checking the integrity and consistency of the data to ensure that the generated preprocessed results meet the requirements of subsequent heterogeneity analysis.

[0205] II. Specific Implementation Steps for Heterogeneous Network Training

[0206] This embodiment details the overall process of training the heterogeneous network in the method of the present invention. This training process is based on a variational autoencoder framework, using a dual-encoder-single-decoder architecture to learn the mapping relationship from cryo-electron microscopy images to Gaussian parameter variations, thereby achieving modeling of the compositional and conformational heterogeneity of biological macromolecules.

[0207] Step 1: Network Architecture and Training Framework

[0208] The core of the heterogeneity reconstruction network is a deep generative model whose goal is to learn low-dimensional representations of structural changes from input cryo-electron microscopy images and map these changes into a three-dimensional Gaussian parameter space. The network employs an encoder-decoder structure, where the encoder is responsible for compressing high-dimensional image data into low-dimensional latent variables, and the decoder generates corresponding Gaussian parameter changes based on these latent variables.

[0209] Network architecture design

[0210] The heterogeneity reconstruction network adopts a deep generative model architecture based on variational autoencoders. Its core components include an image encoder, a Gaussian encoder, and a deformable decoder, and it mainly includes the following modules:

[0211] 1. Basic building blocks

[0212] • ResidLinear: A residual linear layer that implements the residual connection z = Linear(x) + x.

[0213] MLP: Standard Multilayer Perceptron, consisting of linear layers, activation functions, and normalization layers.

[0214] • ResidLinearMLP: A multilayer perceptron based on residual connections to enhance gradient flow.

[0215] 2. Core Network Architecture The network entity HetGMMVAESimpleIndependent contains the following key components:

[0216] Image encoder: Input image Mapping to the latent space:

[0217] z μ ,z σ =E image (I)

[0218] z = z μ +·∈exp(0.5×z σ ),∈~N(0,1)

[0219] The encoder can be selected from either a standard MLP or a residual MLP architecture, with an output dimension of 2×zdim, corresponding to the logarithm of the mean and variance of the latent variables, respectively.

[0220] Gaussian encoder: processes parameter information for each Gaussian point.

[0221] g i =[d i ,s i ,p i (Density, scale, location)

[0222] e i =E gaussian (g i )

[0223] The Gaussian encoder maps the 5-dimensional parameters of each Gaussian (1-dimensional density, 1-dimensional scale, and 3-dimensional position) to a high-dimensional embedding space. The feature decoder integrates latent variables and Gaussian embeddings.

[0224] (Broadcast to all Gaussian points)

[0225] h = D feature ([z exp ,e])

[0226] Δd i =f d (h i ),Δp i =f p (h i ),Δs i =f s (h i )

[0227] Where f d ,f p ,f s These correspond to prediction networks for changes in density, location, and scale, respectively.

[0228] Data input:

[0229] Let the input image be Where D is the image size. The network first passes through an image encoder. Mapping the image to the latent space:

[0230]

[0231] in For d z θ-dimensional latent vector image These are encoder parameters.

[0232] Meanwhile, each Gaussian point i is encoded by a Gaussian encoder. Generate the corresponding embedding vector:

[0233]

[0234] Where g i =[d i ,s i ,p i It includes the density, scale, and location parameters of the Gaussian.

[0235] Deformation Decoder By combining latent variables of the image with Gaussian embeddings, the parameter variations of each Gaussian are estimated:

[0236]

[0237] Step 2: Data Flow During Training

[0238] The training process uses mini-batch gradient descent, with each batch containing B samples. For the j-th sample in a batch, the complete processing flow is as follows:

[0239] First, the image encoder processes the input image I. j Generate the corresponding latent representation z j This step enables the transformation from pixel space to semantic space, capturing the structural heterogeneity information contained in the image.

[0240] Next, for each Gaussian point i, the deformation decoder is based on z j and Gaussian embedding e i Calculate the parameter change Δg ij This process can be decomposed into a density change Δd. ij and position change Δp ij The estimate.

[0241] Then, the estimated parameter changes are applied to the initial Gaussian model. We obtain a structurally deformed version for the current image:

[0242]

[0243] Based on the deformed Gaussian model and the corresponding particle attitude parameter Ψj =(R j ,t j The process involves generating a reconstructed image through 3D-to-2D projection rendering. The rendering process includes steps such as rigid body transformation, projection integration, and CTF modulation, ultimately resulting in a synthetic image of the same size as the original image.

[0244] Step 3: Design of Multi-Objective Loss Function

[0245] To ensure the stability of the training process and the quality of reconstruction, we designed a composite loss function that includes multiple constraint terms. The reconstruction loss measures the difference between the synthetic image and the real image:

[0246]

[0247] The variational regularization term constrains the distribution of the latent space through KL divergence:

[0248]

[0249] Geometric constraints ensure that the learned structural changes are physically plausible. Neighbor embedding consistency constraints encourage spatially adjacent Gaussians to have similar embedding representations.

[0250]

[0251] Local rigid constraints maintain the continuity of the structure after deformation:

[0252]

[0253] Displacement direction consistency constraints ensure motion coordination between adjacent regions:

[0254]

[0255] Exclusion constraints prevent excessive clustering of Gaussian points:

[0256]

[0257] The latent spatial consistency constraint maintains the local smoothness of the latent representation through K-nearest neighbor loss:

[0258]

[0259] The total loss function is the weighted sum of the individual components:

[0260]

[0261] Step 4: Optimize strategies and update parameters

[0262] The training process uses the AdamW optimizer to update parameters, and its update rules are as follows:

[0263]

[0264] The learning rate employs an exponential decay strategy:

[0265]

[0266] In actual training, a progressive training strategy is adopted. Initially, the focus is on optimizing the reconstruction loss and basic geometric constraints, while more complex regularization terms are gradually introduced as training progresses. This strategy helps prevent the model from getting trapped in local optima too early, improving the stability and convergence of the training.

[0267] Step 5: Model Evaluation and Result Analysis

[0268] After training, the model performance is evaluated by assessing the reconstruction quality on the validation set. Key evaluation metrics include the structural similarity between the reconstructed image and the real image, the clustering effect of the latent space, and the physical plausibility of the generated structure.

[0269] Latent space analysis is an important approach to understanding the learning performance of a model. Principal component analysis (PCA) allows for dimensionality reduction and visualization of latent variables, enabling observation of the distribution of different conformational states within the latent space. Furthermore, by traversing the main directions of the latent space, continuous conformational change trajectories can be generated, providing a visual representation of the dynamic behavior of biomolecules.

[0270] Analysis of Gaussian parameter variations provides a fine-grained understanding of local structural changes. By examining the density and positional variation patterns of different Gaussians, flexible regions and key functional sites within the structure can be identified, laying the foundation for subsequent biological interpretation.

[0271] Monitoring of the training process includes the convergence of each component of the loss function, quantitative evaluation of reconstruction quality, and visual inspection of the generated structure. These monitoring methods collectively ensure the reliability of the training process and the interpretability of the results.

[0272] The following is an analysis of the results of example verification based on the method of the present invention provided in the foregoing embodiments:

[0273] 1. This invention captures large-scale motion in simulated datasets.

[0274] The ability of this invention to capture large-scale molecular motion was evaluated using a human immunoglobulin G (IgG) antibody complex dataset generated by CryoBench simulation. This dataset simulates one-dimensional continuous circular motion by rotating the dihedral angle between the antibody (Fab) domain of the linker fragment and the rest of the molecule. The dataset contains 100 conformations, each accompanied by a corresponding density map and atomic model, providing a total of 100,000 particle images with a bounding box size of 128 pixels and a pixel width of [missing information]. The Gaussian-based method e2gmm failed to describe this large-scale motion. Figure 2 In the diagram, a) is the consistent mapping and atomic model from the simulated dataset, along with the initial 3D Gaussian used in this invention. b) is the dimensionality-reduced 2D latent conformational space. c) is the estimated Gaussian displacement directly mapped to the atomic coordinates (red). In the real space, a deformed atomic model (purple) is generated. The deformed model is superimposed on the provided ground reality structure (blue), producing an RMSD of 2.26 μm, demonstrating the accuracy of this invention in capturing large-scale conformational motion. d) is a visualization of the reconstructed atomic model. The left panel shows representative reconstructed atomic models. The right panel shows two RMSD distribution plots: one for all estimated models, and the other classified by rotation angle. In comparison, this invention successfully recovers the complete 360° rotational motion. Figure 2 The reconstructed volume shown in b clearly demonstrates this. The potential conformational space exhibits a distinct toroidal distribution, consistent with the characteristics of continuous 3D rotational motion of the molecules. To better illustrate the motion trajectory, two different perspectives are provided, including a top view aligned along the rotation axis. The reconstructed density map maintains considerable resolution throughout the large-scale motion range. Four representative conformations are shown in the figure, each rotated 90° along the rotation axis.

[0275] Furthermore, the estimated Gaussian shifts are mapped onto the atomic model associated with the consensus structure, such as... Figure 2 The red dashed line in c illustrates this. Unlike multibody refinement methods, this invention does not explicitly impose strict rigid body constraints on the deformation field, yet the resulting atomic displacements still exhibit physically consistent motion characteristics. The Gaussian coding paradigm, combined with local rigidity regularization terms, causes adjacent Gaussians to exhibit similar parameter changes. While voxel-based methods can also capture rotational motion, they sometimes introduce inconsistencies in local regions. In contrast, this invention maintains the structural continuity of the deformation volume because all deformations are applied to the same consensus model. The two superimposed atomic models exhibit a high degree of structural overlap, corresponding to a root mean square deviation (RMSD) between atoms of [value missing]. This demonstrates the high accuracy of this invention in capturing large-scale conformational motion. Furthermore, we reconstructed 1000 atomic models to measure the RMSD distribution, with results as follows... Figure 2 As shown in d. The results of classification by rotation angle indicate that conformations closer to the original structure produce smaller RMSD values, indicating higher reconstruction accuracy. For composite conformations in the range of 144–180°, the average RMSD is approximately This level of precision is still sufficient to effectively capture conformational heterogeneity.

[0276] 2. Analysis of heterojunctions in ribosome assembly.

[0277] The bL17 restriction E. coli ribosome assembly intermediates dataset (EMPIAR-10841) contains a large number of discrete conformations validated by both classification and deep learning analyses. This dataset contains 123,804 particle images with a bounding box size of 160 pixels and a pixel width of [missing information]. In this experiment, the density and position parameters of the Gaussian function were allowed to vary freely without any masking operations. The latent variables were reduced to a two-dimensional space using uniform manifold approximation and projection (UMAP) to visualize the conformational distribution. The results showed clear separation of particle clusters in the latent space, and ten representative cluster centers were selected for detailed analysis. Figure 3 This is a heterojunction analysis of ribosome assembly. a) Visualization of the latent space encoded by this invention, where dark blue dots represent cluster centers corresponding to different conformational classes. b) Ten representative density maps generated from the cluster centers by this invention, each labeled with its class index. Red boxes highlight regions showing major conformational differences. c) Examples of present ribosomal proteins L16 and L25, partially present and completely absent. d) Visualization of the movement of the dynamic central bulge domain and the corresponding changes in the atomic model and CryoSPARC reconstruction plot. Fixed dotted envelopes are used to better illustrate structural movement. e) Distribution of the learned Gaussian embeddings. Four distinct clusters were identified and reconstructed. On the right, the corresponding volumes are shown, where each volume illustrates the effect of removing one component. Figure 3 As shown in b, significant structural differences exist between different clusters, while the areas marked in red highlight subtle compositional variations. Large-scale structural differences may involve the ribosome base and multiple stem regions, reflecting different stages or pathways of ribosome assembly; while localized, small-scale changes may correspond to the integration of ribosomal RNA (rRNA) fragments during maturation. Cluster 9 is far from other clusters in the two-dimensional latent space, and its reconstructed volume exhibits different orientations, which may stem from subtle errors in pose estimation.

[0278] The fine structural variability within the dataset was further investigated. Figure 3 c) The results highlight the multiple states of ribosomal proteins L16 and L25 in different conformations, exhibiting "complete presence, partial presence, or complete absence." To address the ambiguity observed in the dynamic central protrusion domain, the variation in this region was subsequently constrained to positional parameters. For example... Figure 3 As shown in d, this method clearly reveals the tilting motion of the structural domain. To optimize the visualization, the PDB 6PJ6 atomic model is aligned with the reconstructed volume, and Gaussian displacements are mapped to the corresponding atoms. The superimposed atomic model and reconstructed volume together confirm the fine conformational changes captured by this invention. In-depth analysis of the learned Gaussian embeddings is performed using k-means clustering. Figure 3e) The four identified clusters have been color-coded. Cluster-by-cluster elimination experiments show that the remaining structures can still highly reproduce the corresponding discrete conformations, proving that Gaussian embedding can effectively capture different structural features.

[0279] 3. Heterogeneity analysis of the pre-catalytic spliceosome.

[0280] This invention was further applied to analyze structural variations in the precatalytic spliceosome dataset (EMPIAR-10180). This dataset contains 327,490 particle images with a bounding box size of 320 pixels and a pixel width of [missing information]. The images were downsampled to 256×256 resolution, and the initial Gaussian function was sampled from the consensus density map stored in EMPIAR-10180. All Gaussian parameters were allowed to vary freely during training. Six distinct clusters were identified using the K-means algorithm. Figure 4 a). The density map reconstructed based on the cluster center shows clear compositional differences and continuous motion trajectories, with the area marked by the red box particularly highlighting the absence of the SF3b and foot structural domains. Figure 4 b). While the NeRF-based method OPUS-DSD can capture similar component heterogeneity, the Gaussian-based method e2gmm fails to achieve the same effect. To verify the effectiveness of this invention, the clustering projection was further reconstructed in CryoSPARC ( Figure 4 c), where the missing states of the SF3b and foot structural domains are also clearly identifiable.

[0281] The learned Gaussian embeddings were divided into five clusters. Figure 4 d) Each cluster roughly corresponds to a specific biological domain: the blue clusters mainly represent the SF3b region, while the purple clusters mainly correspond to the helicase domain. The results of principal component analysis (PCA) of the latent space ( Figure 4 (e) Further, it revealed different conformational variability along the first two principal component axes (PC1-PC2). The helicase and SF3b domain exhibited coordinated but distinct rotational movements along different component axes: along the first principal component axis, the spliceosome transitioned from an "open" to a "closed" state—the SF3b domain was initially located at the top of the core body, and eventually folded down to the core region; along the second principal component axis, both domains exhibited a lateral oscillating motion. Mapping the Gaussian shift to the atomic model PDB-5NRL fitted with the density map using CryoAlign resulted in smooth and coherent atomic motions, indicating that this method can suppress artificial structural distortions to some extent. It should be noted that although the voxel-based method can generate a globally continuous trajectory, it may exhibit local inconsistencies.

[0282] In summary, this invention provides an intuitive and interpretable framework for analyzing compositional and conformational heterogeneity in cryo-electron microscopy datasets. By modeling macromolecular structures using a three-dimensional Gaussian function, it bridges the gap between density-based analysis and atomic-level interpretation. When combined with corresponding atomic models, this invention generates a continuous landscape of atomic conformations. This method enables researchers to gain a deeper understanding of the dynamic structural behavior and functional mechanisms of macromolecules.

[0283] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A three-dimensional Gaussian modeling method for analyzing the composition and conformational heterogeneity of biological macromolecules, characterized in that, The method includes the following steps: S1: Obtain protein density map data; S2: Extract attitude parameter features; S3: Calculate the contrast transfer function to obtain the CTF parameter file; S4: Generate Gaussian point cloud coordinates and density values; S5: Standardize the image data; S6: Construct a heterogeneous reconstruction network architecture; this network architecture adopts a dual encoder-single decoder architecture, including an image encoder, a Gaussian encoder, and a deformable decoder, and all modules are implemented using a multilayer perceptron with ReLU activation function; S7: Based on the heterogeneous reconstruction network, image data processing is performed to achieve 3D modeling.

2. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, In step S2, the particle attitude parameters are analyzed, including Euler angles and translation vectors. For the i-th particle, the attitude parameter extraction process is as follows: Euler angle extraction: φ i =(α i ,β i ,γ i )=(rlnAngleRot i ,rlnAngleTilt i ,rlnAnglePsi i ); Convert Euler angles to rotation matrices using rotation matrix transformation functions: Translation parameters are processed according to units; if the translation parameters are in angstroms, they are converted to pixel units. Finally, the complete attitude parameters are obtained: P i =[R i ,t i ]; Downsample and scale the translation parameters: Integrate the rotation matrix and translation vector into complete attitude parameters: Ψ i ′=[vec(R i ),t i ′] Where vec(·) means expanding the matrix into a vector.

3. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, In S3, based on the classic CTF theoretical model, the electron wavelength is first calculated: Where V is the accelerating voltage (volts); For the spatial frequency vector f = (f x ,f y ), Calculate the azimuth: θ=arctan2(f y ,f x ) Effective defocus calculation is as follows: The phase offset term is: The final CTF function is: Select the envelope function to apply:

4. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, In step S4, the process of generating the initial Gaussian point cloud from the consensus density map is based on grid sampling and threshold filtering; let the density map be V(x,y,z), and the sampling interval be Δ=(Δ x ,Δ y ,Δ z Generate sampling grid: Apply density threshold screening: Perform downsampling when the number of points exceeds the target number N. target hour:

5. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, In S5: Where μ i and σ i denoted as the mean and standard deviation of the i-th image, respectively.

6. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, In step S6, the heterogeneity reconstruction network adopts a deep generative model architecture based on variational autoencoders, including the following modules: (1) Basic building blocks ResidLinear: Residual linear layer, implementing residual connections of z = Linear(x) + x; MLP: Standard Multilayer Perceptron, which includes linear layers, activation functions, and normalization layers; ResidLinearMLP: A multilayer perceptron based on residual connections to enhance gradient flow; (2) The core network architecture includes: Image encoder: Input image Mapping to the latent space: With μ ,With σ =E image (AND) from=from μ +∈·exp(0.5×z σ ),∈~N(0,1) The encoder can be selected from either a standard MLP or a residual MLP architecture, with an output dimension of 2×zdim, corresponding to the logarithm of the mean and variance of the latent variables, respectively. Gaussian encoder: processes parameter information for each Gaussian point. g i =[d i ,s i ,p i (Density, scale, location) yes i =E gaussian (g i ) The Gaussian encoder maps the 5-dimensional parameters of each Gaussian (1-dimensional density, 1-dimensional scale, and 3-dimensional position) to a high-dimensional embedding space. The feature decoder integrates latent variables and Gaussian embeddings. (Broadcast to all Gaussian points) h=D feature ([z exp ,e]) Δd o =f d (h i ),Δp i =f p (h i ),Δs i =f s (h i ) Where f d ,f p ,f s These correspond to prediction networks for changes in density, location, and scale, respectively.

7. The three-dimensional Gaussian modeling method as described in claim 1, characterized in that, Specifically, S7 is: S7-1: Let the input image be... Where D is the image size, the network first passes through an image encoder Mapping the image to the latent space: in For d z θ-dimensional latent vector image For encoder parameters; Meanwhile, each Gaussian point i is encoded by a Gaussian encoder. Generate the corresponding embedding vector: Where g i =[d i ,s i ,p i Includes Gaussian density, scale, and location parameters; Deformation Decoder By combining latent variables of the image with Gaussian embeddings, the parameter variations of each Gaussian are estimated: S7-2: The training process uses mini-batch gradient descent, with each batch containing B samples; S7-3: Design a multi-objective loss function; S7-4: The AdamW optimizer is used for parameter updates during the training process, and the learning rate adopts an exponential decay strategy.

8. The three-dimensional Gaussian modeling method as described in claim 7, characterized in that, Specifically, S7-2 refers to the following complete processing flow for the j-th sample in a batch: S7-21: First, the image encoder processes the input image I. j Generate the corresponding latent representation z j ; Next, for each Gaussian point i, the deformation decoder is based on z j and Gaussian embedding e i Calculate the parameter change Δg ij This process can be decomposed into a density change Δd. ij and position change Δp ij The estimate. S7-22: Then, the estimated parameter changes are applied to the initial Gaussian model. We obtain a structurally deformed version for the current image:

9. The three-dimensional Gaussian modeling method as described in claim 7, characterized in that, Specifically, S7-3 is: Reconstruction loss measures the difference between a synthetic image and a real image: The variational regularization term constrains the distribution of the latent space through KL divergence: Neighbor embedding consistency constraints encourage spatially adjacent Gaussians to have similar embedding representations: Local rigid constraints maintain the continuity of the structure after deformation: Displacement direction consistency constraints ensure motion coordination between adjacent regions: Exclusion constraints prevent excessive clustering of Gaussian points: The latent spatial consistency constraint maintains the local smoothness of the latent representation through K-nearest neighbor loss: The total loss function is the weighted sum of the individual components:

10. The three-dimensional Gaussian modeling method as described in claim 7, characterized in that, In S7-4, the training process uses the AdamW optimizer for parameter updates, and the update rule is as follows: The learning rate employs an exponential decay strategy: