Cryoelectron microscope density map similarity retrieval method

Through multi-process calculation and hybrid parallel strategies, combined with Meanshift and Dbscan algorithms for density map preprocessing and clustering, and point cloud registration is used to use ICP registration algorithm to solve the accuracy and efficiency of cryo-electron microscope density map search tools, and efficient and flexible density map similarity search is achieved.

CN120523988APending Publication Date: 2025-08-22SHANDONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510599504.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-11
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing cryo-electron microscope density map search tools are insufficient in terms of accuracy and efficiency, especially in the similarity comparison between high-resolution and medium-resolution density maps, and lack flexibility and personalized search capabilities.

Method used

Multi-process computing and hybrid parallel strategy are adopted, combined with Meanshift and Dbscan algorithms for preprocessing and clustering of density maps, point cloud registration is used using ICP registration algorithm, and the similarity of density maps is evaluated through a comprehensive scoring function, providing a multi-dimensional retrieval method.

Benefits of technology

It improves the search accuracy and efficiency of cryo-electron microscope density maps, provides flexible parameter adjustment space, adapts to different experimental conditions and sample characteristics, and is suitable for diverse experimental scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523988A_ABST
    Figure CN120523988A_ABST
Patent Text Reader

Abstract

The invention discloses a cryoelectron microscope density map similarity retrieval method, and belongs to the technical field of cryoelectron microscope image processing. The method comprises the steps of firstly constructing a point cloud library of a density map, performing feature extraction and similarity comparison performance optimization by using multiple parallel strategies, and performing efficient and accurate density map similarity retrieval. According to the method, a scoring function based on geometric errors, normal and cosine similarity, point cloud density consistency and direction histogram signature feature similarity is designed, a density map similar to a target density map is finally output, and a transfer matrix capable of being applied to a superposition density map is provided. According to the method, the multi-process calculation capability is fully utilized, and the local similarity registration efficiency of the cryoelectron microscope density map is improved; by innovatively designing a multi-dimensional scoring function, the retrieval precision is improved; and meanwhile, a more flexible parameter adjustment space is provided for a user, so that the density map similarity retrieval can be finely controlled according to different experiment conditions and sample characteristics, and the requirements of diversified experiment scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cryo-electron microscopy image processing, and in particular to a cryo-electron microscopy density map similarity retrieval method. Background Art

[0002] Cryo-electron microscopy (Cryo-EM) involves freezing a sample, imaging it with an electron microscope, and then reconstructing it using computer algorithms. This technique allows for the acquisition of high-resolution 3D structures without destroying the sample. In recent years, with advances in direct electron detectors, image processing algorithms, and computing power, Cryo-EM has become a leading method for elucidating the structures of biomacromolecules. With the widespread adoption of Cryo-EM, a large number of medium- to high-resolution biomacromolecule structures have been resolved and stored in public databases such as the Electron Microscopy Data Bank (EMDB). While this data contains a wealth of biological information, its retrieval and utilization remain challenging. EMDB has become a fundamental resource for the tertiary structure of biomacromolecules. However, to fully utilize this indispensable resource, the ability to search the EMDB by structural similarity is crucial. Effective retrieval methods can help researchers classify biomacromolecules, supporting structural analysis, function prediction, and drug development. In the field of structural analysis, searching and comparing Cryo-EM density maps allows researchers to better understand the functions and interactions of biomacromolecules. The diversity and complexity of biomacromolecules result in a high degree of heterogeneity in Cryo-EM density maps, complicating retrieval and comparison. The task of searching cryo-EM density maps primarily relies on similarity comparisons. This is particularly true for high- and medium-resolution density maps, which contain rich and clear structural information. This places higher demands on the accuracy and efficiency of similarity comparisons. Existing cryo-EM density map search tools include EM-SURFER, based on 3Dzernikke descriptors, and Omokage search, based on gmfit. Both tools have web-based interfaces capable of searching the entire EMDB. However, EM-SURFER does not provide overlaid density maps, while Omokage search is suitable for low-resolution density maps. Furthermore, the search process of EM-SURFER and Omokage search is transparent to the user, which hinders localization and personalized search library construction. CryoAlign utilizes local spatial feature descriptors to enable fast, accurate, and robust comparison of two cryo-EM density maps. Compared to other alignment methods, such as gmfit, fitmap, and VESPER, CryoAlign stands out by providing more accurate density map overlays while maintaining a low failure rate. However, CryoAlign only provides a method for aligning density maps and has not yet been applied to the retrieval of cryo-EM density maps. Furthermore, the computational process of its local registration method is a traversal process, which will become a performance bottleneck for CryoAlign's local registration. Summary of the Invention

[0003] The purpose of the present invention is to provide a cryo-electron microscopy density map similarity retrieval method to make up for the shortcomings of the existing technology.

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] A cryo-electron microscopy density map similarity retrieval method comprises the following steps:

[0006] S1: Collect cryo-electron microscopy density map data and perform preprocessing;

[0007] S2: Sampling and clustering each density map: First, the density map is sampled into a vector. Based on the voxel value and the vector, the three-dimensional coordinates of the entire density map are obtained. Then, the three-dimensional coordinates of the density map are clustered. The Meanshift and Dbscan algorithms are used to obtain the key point data of the density map. The key point cloud data of all density maps are persisted as a point cloud library.

[0008] S3: After extracting the query density map into point cloud data, the optimal rotation and translation matrix is ​​estimated using the ICP registration algorithm for each point cloud data in the point cloud library; and a one-to-many batch comparison is performed using multiple processes to improve retrieval efficiency;

[0009] S4: Evaluate the results of each pair of point cloud registrations. The evaluation dimensions are based on geometric error, normal cosine similarity, point cloud density consistency, and SHOT feature similarity, and perform a normalized comprehensive score on the score value.

[0010] S5: By judging the similarity of the point cloud data, the similarity of the cryo-electron microscopy density map is indirectly evaluated; all the comparison results are sorted, the density maps with higher scores (such as TOP10) are selected, and the translation and rotation matrix after each pair of point cloud data comparison is obtained, which is used as the overlay visualization of the target-source density map, and finally the search results are output.

[0011] Furthermore, the pre-processing in S1 is: applying low-pass filtering Remove high-frequency noise and retain the main structural features.

[0012] Furthermore, in S2, the density map is first sampled at a certain voxel interval, and then MeanShift clustering is performed on the sampled points. Finally, the key points are extracted by the Dbscan algorithm, which specifically includes:

[0013] S2-1: First, the input density map is converted into a point cloud by uniform sampling, and the density vector is assigned using the mean shift equation. Then, the key points in the point cloud are identified using clustering techniques, and the local spatial structure feature descriptor is calculated. The uniformly sampled grid points are regarded as point clouds, and the unit vector is calculated as the "density vector" of these points. The density vector of each grid point x is calculated.i = 1, ..., N density value is not less than the recommended contour level unit vector, direction unit vector Reflects the grid point x i Trend of surrounding density values, y i It is calculated by the following formula:

[0014]

[0015] where k(p) is the Gaussian kernel function, Φ(x i ) is the grid point x i The density value of k(p) adjusts the neighboring influence according to the input distance P and bandwidth σ;

[0016] S2-2: Parallel optimization using the mean shift algorithm

[0017] In the density map, the density value corresponds to the integration of the density function associated with the atoms, and the high-density areas can indicate the protein backbone; the MeanShift algorithm is used to efficiently identify these dense areas in the map. Unlike density vector generation, the frozen symbol determines the local density maximum point through the convergence result of the following iteration. The algorithm operates on a physical interval of (Δ x , Δ y , Δ z ) D×H×W The core of the three-dimensional voxel grid is the two basic tensor fields that drive the density-aware drift process: the first tensor Encodes the normalized electron density value, where the element Q[i,j,k] quantifies the relative density of voxel (i,j,k); the secondary tensor Explicitly store the physical coordinates of the voxel center, defined as X[i, j, k] = (iΔ x , jΔ y , kΔ z ) T .

[0018] The numerical basis of the algorithm comes from generating molecules and the denominator Two Gaussian weighted convolution operations. Capturing the weighted spatial distribution of the density-mass product via vector convolution Calculation, where ⊙ represents element-wise multiplication and * is the 3D convolution operator. Characterize the density-weighted nuclear response, calculated as where ∈ = 10 -6 Ensure numerical stability. Gaussian kernel The spatial weighting is implemented with bandwidth σ, which is formalized as W[i, j, k] = exp(-1.5||(i, j, k)-c||2 / σ 2 ), c is the nuclear center.

[0019] The above intermediate results are divided element by element to generate the drift vector field in Perform per-voxel vector normalization. Each The encoded voxel (i, j, k) is subjected to an optimized drift direction weighted by the local density gradient.

[0020] The master-slave architecture is used to decouple global convolution calculation from distributed point processing. The master process (Rank 0) first constructs a Gaussian kernel based on the physical interval parameters and then performs 3D convolution calculations accelerated by FFT. and These tensors reside in the master process memory in response to subsequent drift vector queries, while the derived drift field C is broadcast to all workers via MPI_Bcast to maintain consistency.

[0021] Point cloud processing adopts a cyclic decomposition strategy, and the initial point set Y (0) Divide into K blocks (K is the number of MPI processes). Each worker process k receives consecutive blocks Where N is the total number of points. At iteration t, the worker calculates the local drift vector by querying the cached C field through the nearest neighbor grid index:

[0022]

[0023] in Implement the nearest voxel coordinate rounding, through the amplitude threshold θ = lower_bound 2 Suppress drift below the noise level:

[0024]

[0025] The parallel update mechanism adopts a three-level reduction process to reduce synchronization overhead: first, cache-aligned buffers and Kahan compensated summation are used within the node to aggregate the partial sums of multiple CPU threads; second, MPI_Allreduce is used to implement inter-node reduction based on binary tree topology, and FP16 / FP32 precision is automatically selected according to the vector magnitude - for ||ΔY|| ∞ <10 3 The vectors are compressed to FP16 format using random rounding, while the others maintain FP32 precision. The global update rule integrates the reductions through adaptive normalization:

[0026]

[0027] Where η represents the learning rate, δ max is the maximum allowable displacement for a single iteration, the denominator ||ΔY (t)||+∈ maintains the directionality of small drift while preventing division by zero anomalies;

[0028] S2-3: Extract key point cloud using Dbscan algorithm

[0029] To enhance representation capabilities and reduce the size of the point cloud, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm was incorporated; this algorithm clusters points within a specified threshold distance, typically equal to the sampling distance. This effectively reduces redundancy and captures essential structural information in a more compact form. The remaining points serve as key points in the subsequent alignment stage.

[0030] First, the input voxel data (shape is (3, x, y, z)) is flattened into a (3, N) two-dimensional coordinate matrix (N is the total number of points), and after transposition, it forms an (N, 3) array structure, where each row represents a three-dimensional point; then the DBSCAN algorithm is called to perform density clustering on the point cloud by setting the neighborhood radius eps (distance threshold ∈) and the minimum cluster size min_samples (density threshold m); the algorithm traverses each point p and checks the number of points in its neighborhood | N ∈ (p) | whether it satisfies | N ∈ (p)|≥m, if satisfied, it is marked as a core point and the cluster is expanded, and the point is finally divided into several clusters or noise (labeled as -1); if the centroid extraction mode is enabled (centroid only=True), for each non-noise cluster C, its geometric centroid is calculated in parallel through multiple processes Finally, the coordinate matrix of all centroids is output; otherwise, the original data is retained. The centroid calculation is accelerated through parallelization.

[0031] Furthermore, in S3:

[0032] Based on the determined keypoints and the initial points assigned with "density vectors", a density-based signature of the histogram of directions (SHOT) feature description is calculated for each keypoint. The local neighborhood points around each keypoint are examined to calculate the modified SHOT descriptor. The directions of the assigned density vectors of these neighboring points are quantized into discrete bins, and a histogram is constructed to collect the distribution of these directions. The histogram effectively summarizes the local geometric characteristics of the density map. After the sampling and clustering stages, the two input density maps are effectively converted into point clouds and corresponding keypoints. In the first stage of alignment, the initial transformation parameters are estimated using a fast and provable point cloud registration method (TEASER).

[0033] When using point cloud registration for fine alignment, candidate descriptors from a larger density map can easily interfere with a smaller number of descriptor queries. To address the problem of accurate local alignment, we treat it as a global retrieval problem within a small "dataset" and employ a translation mask as a simple segmentation scheme for the larger point cloud. We then use a two-stage alignment process to compute a series of transfer matrices. Based on these transfer matrices, we compute similarity scores for all superpositions using a multi-dimensional scoring function, and select the top 1-n as output.

[0034] The computational process of the local point cloud registration method is a traversal process. The translation mask traverses the entire volume along the x, y, and z axes at a fixed step size, with each translation resulting in a registration. If serial computation is used, this becomes a performance bottleneck for local registration, resulting in a time complexity of O(x_step * y_step * z_step). This also reduces the efficiency of the registration when the density map is large. To improve efficiency, a hybrid parallelization strategy using MPI and OpenMP is employed. MPI provides coarse-grained parallelization support for the point cloud registration process, while OpenMP optimizes the translation mask triple loop for thread-level parallelism through shared memory parallelism, significantly reducing runtime overhead. The MPI environment is first initialized, and computational processes are configured and allocated to available resources. Subsequently, the input data (including B_pcd, B_key_pcd, A_key_pcd, A_key_feats, and B_key_feats) is partitioned and distributed via the MPI communication mechanism. Each process independently performs the following operations: extracting the point cloud structure and its keypoints, calculating the translation mask parameters, verifying the mask dimensionality, and performing local feature extraction to establish keypoint correspondences in the input dataset. The results of each process are aggregated and sent to the root process via the MPI collective communication protocol.

[0035] The root process constructs a correspondence matrix based on the aggregation results and establishes and solves the optimization problem for calculating the initial transformation matrix. To further improve registration accuracy, the root process applies the Iterative Closest Point (ICP) algorithm to iteratively minimize the geometric distance between the transformed source and target point clouds. After convergence, a registration score is calculated to quantitatively assess the quality of the transformation. Ultimately, the transformation matrix and registration score are broadcast to all processes via MPI communication primitives to maintain global consistency. The algorithm concludes with the termination of the MPI environment, ensuring resource release and process cleanup. This structure enables efficient parallel computing and robust point cloud registration, making it particularly suitable for large-scale datasets.

[0036] Furthermore, the hybrid MPI-OpenMP registration algorithm achieves a combination of distributed and shared memory parallelism through five phases of collaborative work:

[0037] Phase 1: MPI initialization and data preparation. The master process (rank 0) initializes the MPI environment and loads the source point cloud (A_pcd) and target point cloud (B_pcd). SHOT feature descriptors are calculated using the spherical histogram of the brain voxel size with the feature radius to capture local geometric patterns. Spatial validity verification is achieved by comparing the bounding box normalized dimension inequality:

[0038]

[0039] When the inequality holds, the execution is terminated and a dimension mismatch warning is returned;

[0040] Phase 2: Hierarchical parameter distribution, using a two-stage broadcast mechanism to transmit the BroadcastParams structure, scalar parameters (space boundary center, terminal search radius r, step size) are transmitted through standard MPI_Bcast; the feature matrix F A With F B Distribution through custom broadcast_matrix: first communicate the matrix dimensions (rows / columns), then broadcast the Eigen:MatrixXd data block in row-first format, supporting dynamic buffer adjustment of the receiving node;

[0041] Phase 3: Workload partitioning uses a block allocation strategy to divide the i-axis search space [0, terminal x ] is divided into MPI processes; the single process block size calculation formula is:

[0042]

[0043] Where δ is the indicator function that distributes the residual blocks to the low-rank processes; each process calculates the local boundary as start_i = rank × Δ + ξ and end_i = min (start_i + chunk, terminal x ), ξ is the initial offset compensation.

[0044] Phase 4: Multi-granularity parallel execution The three-dimensional (i, j, k) loop nest is parallelized using the collapse(3) instruction of OpenMP. Each thread generates a center c mask =center+(iΔ, jΔ, kΔ), a spherical mask of radius r, is used to perform constrained ICP registration via Registration_mask. This process includes super-max_correspondence_dist correspondence filtering, initial registration based on SHOT features, and voxel-adaptive threshold adjustment. The successful transformation matrix and the registration score s l By caching Eigen matrices into thread-local buffers, critical section protection is implemented when aggregating global results.

[0045] Phase 5: Result aggregation uses non-blocking MPI communication with four message tags; coordinate-score pairs (MPI_INT[3], MPI_DOUBLE) are transmitted via tags 1-2, and the transformation matrix is ​​serialized into an MPI_DOUBLE

[16] array and transmitted via tags 3-4, with column-major storage maintained via explicit transposition:

[0046]

[0047] The main process asynchronously receives and splices the data stream, and writes the final result to the file (coordinates / scores are separated by tabs), realizing Complexity scaling, where P is the number of MPI processes and T is the number of OpenMP threads per node.

[0048] Furthermore, in S4, a comprehensive scoring function is proposed to be constructed based on the rigid body transformation matrix (including rotation and translation) obtained by point cloud registration to evaluate the alignment quality of cryo-EM density maps. This scoring function integrates multiple metrics, including geometric error, normal vector cosine similarity, point cloud density consistency, and directional histogram signature (SHOT) feature similarity, to achieve multi-dimensional robustness evaluation. After point cloud registration using the ICP algorithm, the optimal transformation matrix is ​​superimposed on the target-source density map, and various metrics are calculated from different dimensions.

[0049] The geometric error quantifies the degree of spatial alignment between the source point cloud and the target point cloud. It is defined as the mean Euclidean distance between the transformed source point cloud point ai and its corresponding point bj in the target point cloud, where K is the total number of corresponding points. This metric reflects the positional deviation between point clouds:

[0050]

[0051] Normal vector similarity (NS(A,B)) evaluates the degree of alignment of surface orientations between point clouds by calculating the cosine similarity of the normal vectors nai and nbj of corresponding points and introducing the corresponding weight wi for weighting. This indicator reflects the directional consistency of local surface features:

[0052]

[0053] Density consistency (DC(A,B)) evaluates the degree of local density matching between point clouds and analyzes the consistency of neighborhood structure by comparing the number of neighborhood points within a predefined radius r. This metric can effectively identify registration deviations caused by sampling differences, noise, or inconsistent resolution. Where NbrA(pi,r) and NbrB(qi,r) represent the number of neighborhood points within radius r between the source point cloud point pi and the target point cloud point qi, respectively:

[0054]

[0055] The SHOT signature feature similarity (SHOT(A,B)) evaluates the similarity of local geometric features between point clouds by calculating the mean cosine similarity between the SHOT descriptors fai and fbj of all corresponding points, where K is the total number of corresponding points:

[0056]

[0057] Each indicator is normalized to ensure that its value range is in the interval [0,1]:

[0058] normalized_score=max(0,min(1,normalized_score))

[0059] The final similarity score is composed of the weighted sum of each normalized indicator. The weight ω reflects the relative importance of each indicator. The weight distribution strategy, that is, the multi-dimensional scoring function is as follows:

[0060] final_score=w1·nor NS +w2·nor SHOT -w3·nor d_geo -w4·nor DC Among them, ω1 is used to evaluate the consistency of surface normals. Since normal information reflects the local geometric structure of the point cloud, it is given a higher weight to enhance the sensitivity to local registration features; ω2 represents the matching degree of local geometric features. It is also given a higher weight for the local feature registration requirements of high-resolution cryo-electron microscopy point clouds; ω3 evaluates the global alignment of the point cloud in space. This indicator is crucial to the overall quality of registration, but may be affected by noise or uneven distribution of point clouds, so it is given a medium weight; ω4 reflects the difference in local density of the point cloud. Since the density distribution is easily affected by data sampling or noise, it is given a lower weight to reduce interference; nor NS is the normalized normal vector similarity score, nor SHOT is the normalized signature of direction histogram (SHOT) feature similarity score, nor d_geo is the normalized geometric error score, nor DC is the normalized density consistency score.

[0061] The scoring function integrates multiple metrics, including geometric alignment, surface orientation consistency, local density characteristics, and descriptor similarity, to comprehensively evaluate the global and local structural characteristics of the point cloud. This multi-level evaluation mechanism ensures accurate matching of cryo-EM density maps in various search scenarios, providing reliable technical support for structural biology and molecular modeling research.

[0062] Compared with the prior art, the present invention has at least the following beneficial effects:

[0063] This invention fully utilizes the multi-process computing power to accelerate the efficiency of local similarity registration of cryo-electron microscopy density maps; improves retrieval accuracy through the innovative design of multi-dimensional scoring functions; at the same time, provides users with more flexible parameter adjustment space, enabling them to fine-tune the control of density map similarity retrieval according to different experimental conditions and sample characteristics, meeting the needs of diverse experimental scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a time chart comparing the present invention and VESPER at different volumes.

[0065] Figure 2 This is a diagram of the speedup ratio before and after parallelization of the present invention.

[0066] Figure 3 It is a mixed (local + global) search result graph.

[0067] Figure 4 This is a basic data processing workflow diagram in the present invention. DETAILED DESCRIPTION

[0068] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described below in detail with reference to specific embodiments and the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments disclosed herein without inventive effort are intended to fall within the scope of protection of the present invention.

[0069] Example 1

[0070] In this embodiment, the density map is first subjected to feature extraction to construct a retrieval library, and then the constructed retrieval library is subjected to similarity retrieval of the target density map. The original cryo-electron microscopy density map is input, which can be from the EMDB library, or a multi-conformation density map data set can be customized as needed. Each input density map is sampled and clustered. First, the density map is sampled into a vector. According to the voxel value and the vector, the three-dimensional coordinates of the entire density map are obtained. Then, the three-dimensional coordinates of the density map are clustered, and the key point data of the density map are obtained using the Meanshift and Dbscan algorithms. The key point cloud data of all density maps are persisted as a retrieval library. The retrieval library of the density map is persisted in advance as a key point cloud of the density map and a sampling point cloud database. When comparing, only the target density map needs to be sampled and the key points extracted, which can greatly improve the retrieval performance. Since the point cloud data of the density map is finally stored, the storage size of the library will not be too large, which is conducive to the transplantation of the retrieval library and also convenient for the maintenance of the retrieval library.

[0071] Secondly, the retrieval process is to input the query density map, extract the query density map into point cloud data, and use the ICP registration algorithm to estimate the optimal rotation and translation matrix for each point cloud data in the constructed point cloud library. Then, the results of each pair of point cloud registrations are evaluated. The evaluation dimensions are based on geometric error, normal cosine similarity, point cloud density consistency, and SHOT feature similarity, and the score values ​​are normalized and comprehensively scored. By judging the similarity of the point cloud data, the similarity of the cryo-electron microscopy density map is indirectly evaluated. After that, all the comparison results are sorted, the density maps with the top 10 scores are picked out, and the translation and rotation matrix after the comparison of each pair of point cloud data is obtained, which is used as the overlay visualization of the target-source density map. Batch processing is used for comparison, and multiple processes are used for one-to-many batch comparison, thereby improving the retrieval efficiency. The retrieval method provided in this embodiment includes the following steps:

[0072] 1. Build a search library and use the meanshift and dbscan algorithms to extract key points from the point cloud features of the density map. The experimental data consists of 15 categories, with 5 similar density maps in each category, for a total of 75 cryo-electron microscopy density maps. The specific steps are as follows:

[0073] 1. Data preprocessing and feature initialization

[0074] Step 1.1 Input data configuration

[0075] (1) Experimental data:

[0076] Contains cryo-EM density maps of 15 biological complexes, each with 5 structural variants. The data format uses the EMD-3060 standard storage format with a resolution range of Preprocessing of each density map: Apply low-pass filtering Remove high-frequency noise and retain the main structural features

[0077] (2) Point cloud processing:

[0078] An adaptive threshold ρ>3σ is used to extract significant voxels, where σ is the standard deviation of the density map. The original spatial coordinates are retained when generating a 3D point cloud. Each point contains (x, y, z, ρ) four-tuple information. The octree structure is used for spatial indexing to improve the efficiency of subsequent processing.

[0079] Step 1.2 Parallel acceleration mechanism

[0080] (1) MPI data block:

[0081] The main process (Rank 0) reads all 75 density map metadata and sorts them topologically by category to ensure that the five maps of the same category are allocated to continuous storage space. Non-blocking communication is used: MPI_Isend / MPI_Irecv is used to implement asynchronous transmission. The block formula is:

[0082]

[0083] (2) Hierarchical extraction of key points.

[0084] Step 2.1 Density-aware drift field construction

[0085] (1) Definition of tensor field:

[0086] The density tensor Q stores the normalized relative density value, reflecting the compactness of the local structure. The coordinate tensor X records the position of the voxel in the physical space and is used to calculate the mass moment tensor dimension: a typical setting is D×H×W=512 3 , voxel spacing

[0087]

[0088] X[i,j,k]=(iΔ x , jΔ y ,kΔ z ) T

[0089] (2) Gaussian convolution kernel:

[0090] The core radius w = 5 corresponds to the physical size, covering the typical α-helix diameter, and the bandwidth σ = 2Δ ensures that the coverage area includes 3-5 adjacent voxels. The formula is as follows:

[0091]

[0092] (3) Drift field calculation:

[0093] molecular Calculate the weighted integral of the density-mass moment, reflecting the local density distribution, the denominator As a normalization factor to avoid divergence in low-density areas

[0094]

[0095] Step 2.2 Distributed keypoint extraction

[0096] (1) Data block strategy:

[0097] Divide data blocks based on the principle of spatial locality to reduce inter-process communication. The block formula ensures that adjacent points are assigned to the same process and maintains spatial continuity.

[0098]

[0099] (2) Drift iteration:

[0100] Nearest neighbor interpolation: Bilinear interpolation is used to improve coordinate mapping accuracy, and dynamic threshold mechanism: Correspondence Displacement noise threshold

[0101]

[0102] (3) Density clustering and retrieval database construction.

[0103] Step 3DBSCAN clustering optimization

[0104] (1) Neighborhood parameters:

[0105] The search radius is adaptive to the size of the map, adapting to different molecular scales, and the minimum number of points ensures the detection of true structural features and filters out noise artifacts.

[0106] ∈=0.3×diameter of the spectrum

[0107] min_samples=10

[0108] (2) Key point extraction:

[0109] The centroid of each density-connected cluster is calculated, and the most representative spatial position is retained. The final number of key points is about 10%-20% of the original point cloud, achieving efficient representation.

[0110]

[0111] 2. Query similar density maps in the constructed retrieval library for the query density map, sample the point cloud of the query density map and extract key points, and use the ICP algorithm to perform similarity comparison with the point cloud features in the retrieval library.

[0112] Step 1: Input the density map to be queried, sample the point cloud and extract key points from the density map to be queried. The specific steps are the same as those in Example 1.

[0113] Step 2: Parallel computing environment initializes the main process (rank 0) to load the pre-processed query point cloud key points K q and its SHOT feature F q , retrieve all target point clouds in the library Execute MPI_Init to create a communication domain containing P processes and verify the consistency of the characteristic matrix dimensions:

[0114]

[0115] If the dimensions do not match, MPI_Abort is triggered to terminate the calculation.

[0116] Step 3: Feature data broadcasting distributes query features through the improved tree broadcast algorithm:

[0117] The main process will F q Divide into blocks Data Packet Counting Relay Nodes Use MPI_Isend to asynchronously send data packets to the relay node. The relay node receives the data through MPI_Irecv and forwards it to the child node.

[0118] Step 4: Multi-level parallel alignment: Start the OpenMP thread group in each MPI process to execute: for each assigned translation vector t=(t x , t y , t z ), apply OpenMP's SIMD instructions to parallelize: Use the ICP registration algorithm.

[0119] 3. Evaluate the results of each pair of point cloud registrations. The evaluation dimensions are to perform a normalized comprehensive score based on geometric error, normal cosine similarity, point cloud density consistency, and SHOT feature similarity. By judging the similarity of the point cloud data, the similarity of the cryo-electron microscopy density map is indirectly evaluated. All comparison results are sorted, and the density maps with the top 10 scores are selected to obtain the translation and rotation matrix after the comparison of each pair of point cloud data, which are used as the overlay visualization of the target-source density map. Finally, the retrieval results are output.

[0120] Example 2: This example evaluates the effects of the present invention and existing retrieval methods through comparative experiments.

[0121] To test the ability of our method to handle hybrid global-local search tasks, we constructed a dataset that proportionally mixed global and local search tasks. This dataset contained 15 categories, each with five mutually similar density maps. Approximately 6,000 pairwise comparisons were performed, with each density map compared to all maps in its own category and all other categories.

[0122] As shown in Table 1, the first column indicates the global / local ratio and the type of registration method. The first column of each table indicates the type of registration method, and subsequent columns show the precision, recall, F1 score, and the time required to search 75 density maps after performing 6,000 alignments. This experiment used a cross-dataset testing strategy covering both global and local search tasks. The first row shows the density maps with 20% global search and 80% local search. Due to the predominance of local search data in the database, EM-SURFER and Omokage suffer from a large number of false positives, resulting in low F1 scores. Although EM-SURFER has the fastest matching speed, the inclusion of a large amount of local search data significantly reduces its accuracy. The second row shows a test scenario with a 50% global and 50% local search density maps. As the proportion of global search maps increases and the proportion of local search maps decreases, the retrieval accuracy of both EM-SURFER and Omokage improves, but their performance is still affected by the local search maps. The third row shows that when the dataset consists of 80% global search graphs and only 20% local search graphs, the retrieval accuracy of EM-SURFER and Omokage significantly improves compared to when the local search task dominates, but their performance is still affected by the proportion of local graphs. This indicates that EM-SURFER and Omokage are less capable of handling high-proportion local search tasks, and their accuracy decreases in mixed search scenarios.

[0123] Table 1 Comparison of retrieval accuracy and efficiency under different data ratios

[0124]

[0125] In contrast, the introduction of local retrieval task data in the present invention has little impact on the overall retrieval accuracy, and the F1 value reaches the highest, which proves that the present invention can efficiently handle mixed retrieval tasks.

[0126] The parallel efficiency of the present invention was verified in the experimental environment configured with an Intel(R) Core(TM) i9-10900X CPU @ 3.70GHz processor, 64G memory, and an NVIDIA GeForce RTX 3080 10G graphics card. Taking EMD-3661 and EMD-6647 as examples, the former samples 4,647 points and the latter samples 11,384 points. Single-threaded local registration will produce a large performance bottleneck, such as Figure 1As shown in the figure, parallel processing can achieve significant acceleration: when 20 threads are used, the acceleration ratio reaches 1:7, indicating that the efficiency of a single local registration after parallel optimization can reach 7 times that of the processing before optimization.

[0127] Performance comparison under the same environment and data shows that the accelerated local registration method of the present invention can reach twice the speed of VESPER under the same number of threads. Table 2 shows the performance of the parallel processing of the algorithm, and the performance tests before and after parallelization are carried out on the three tasks of automatic construction of point cloud library (75Classes), key point extraction algorithm and mask registration algorithm. The performance optimization of automatic construction of point cloud library is mainly aimed at the parallel processing of large-scale data, and the speed is increased by nearly 50% after parallelization; the key point extraction algorithm mainly optimizes the data transmission between CPU and GPU, and the speed is increased by nearly 45% after focusing on the optimization of Meanshift algorithm; the local registration algorithm can be regarded as a collection of multiple global registrations, which is naturally suitable for multi-core parallel acceleration.

[0128] Table 2. Runtime of operators before and after parallelization

[0129]

[0130] By comparing the runtime of the present invention and VESPER, we found that the VESPER registration process takes significantly longer when processing large-volume density maps. We categorized the density maps by file size and selected five pairs of maps from the ranges of 10M-50M, 50M-100M, 100M-200M, and 200M-500M, ensuring that both registration applications were configured with four threads. We then compared the runtime of Ours and VESPER. The results are shown in Figure 2. Figure 2 As shown in the figure, for large-volume density maps in the range of 200M-500M containing more complex feature information, the runtime consumption of VESPER increases significantly, while the present invention is much faster than VESPER. This also shows that although VESPER is suitable for registration tasks, its inefficiency in processing large-volume density maps makes it unsuitable for retrieval tasks.

[0131] exist Figure 4 middle: Figure 4 a. Build a retrieval library: sample the density map and cluster key points, associate the sampling point and key point data with the density map number, and save them persistently. Figure 4 b. Retrieval process: After inputting the density map to be queried, extract the point cloud features of the density map to be queried, align the point cloud features of the density map to be queried with all the point clouds in the retrieval database using ICP, and then use the scoring function to align and score, and output the top ten results. In order to verify the effectiveness of the present invention, a large number of unknown category density maps are mixed into the database to perform mixed retrieval tasks. Figure 3 and 4As shown in the figure, by comparing Ours / VESPER / Omokage, the density maps of the top 10 similarity scores are shown using EMD-6675 as an example. Although there are many low-resolution and small-volume density maps in the retrieval database (such as Figure 4 Although the algorithm can identify local similarities between these small-volume density maps and the upper half of EMD-6675 (EMD-5722, EMD-8399, EMD-2629, EMD-3089, EMD-3087, and EMD-3088 in the search results), the present invention still captures local information and identifies local similarities between these small-volume density maps and the upper half of EMD-6675. This is further demonstrated in the third "Our Alignment" example, where it is clearly observed that the upper half of EMD-6675 indeed shares local similarities with EMD-8399, which ranks highly in the search results. Density maps belonging to the same category as EMD-6675, such as EMD-6676, EMD-8779, EMD-1763, and EMD-6677, also receive high scores. Furthermore, the present invention avoids misjudgments when evaluating high-resolution, feature-rich density maps, and avoids cases where high-resolution density maps with completely different morphologies receive abnormally high scores, further demonstrating that the present invention is more suitable for retrieval of medium- and high-resolution density maps. Compared with the search results of VESPER and Omokage (such as Figure 4 b), VESPER's accuracy is lower than that of the present invention, there are misjudgments and it cannot effectively capture local features; Omokage relies on shape and size similarity for retrieval, and Omokage's retrieval results for EMD-1458 and EMD-5001 have no obvious local similarity with EMD-6675. Its high retrieval score is only due to the similarity in surface shape with EMD-6675, which further shows that Omokage is only suitable for global retrieval tasks, while the present invention is more suitable for mixed retrieval tasks.

[0132] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects disclosed in the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A cryo-electron microscopy density map similarity retrieval method, characterized in that: The retrieval method includes the following steps: S1: Collect cryo-electron microscopy density map data and perform preprocessing; S2: Sampling and clustering each density map: First, the density map is sampled into a vector. Based on the voxel value and the vector, the three-dimensional coordinates of the entire density map are obtained. Then, the three-dimensional coordinates of the density map are clustered. The Meanshift and Dbscan algorithms are used to obtain the key point data of the density map. The key point cloud data of all density maps are persisted as a point cloud library. S3: After extracting the query density map into point cloud data, the optimal rotation and translation matrix is ​​estimated using the ICP registration algorithm for each point cloud data in the point cloud library; and a one-to-many batch comparison is performed using multiple processes to improve retrieval efficiency; S4: Evaluate the results of each pair of point cloud registrations. The evaluation dimensions are based on geometric error, normal cosine similarity, point cloud density consistency, and SHOT feature similarity, and perform a normalized comprehensive score on the score value. S5: By judging the similarity of the point cloud data, the similarity of the cryo-electron microscopy density map is indirectly evaluated; all the comparison results are sorted, the density maps with higher scores are selected, and the translation and rotation matrix after each pair of point cloud data comparison is obtained, which is used as the overlay visualization of the target-source density map, and finally the search results are output.

2. The cryo-electron microscopy density map similarity retrieval method according to claim 1, characterized in that: The pre-processing in S1 is: applying low-pass filtering to remove high-frequency noise.

3. The cryo-electron microscopy density map similarity retrieval method according to claim 1, characterized in that: The S2 includes: S2-1: Treat the uniformly sampled grid points as a point cloud and calculate the x i = 1, ..., N density value is not less than the recommended contour level unit vector, direction unit vector Reflects the grid point x i Trend of surrounding density values, y i It is calculated by the following formula: where k(p) is the Gaussian kernel function, φ(x i ) is the grid point x i The density value of k(p) adjusts the neighboring influence according to the input distance P and bandwidth σ; S2-2: Parallel optimization using the mean shift algorithm The master-slave architecture is used to decouple global convolution calculation from distributed point processing. The master process (Rank 0) first constructs a Gaussian kernel based on the physical interval parameters and then performs 3D convolution calculations accelerated by FFT. and Point cloud processing adopts a cyclic decomposition strategy; the global update rule integrates the reduction through adaptive normalization: Where η represents the learning rate, δ max is the maximum allowable displacement for a single iteration, the denominator ||ΔY (t) ||+∈ maintains the directionality of small drift while preventing division by zero anomalies; S2-3: Use the DBSCAN algorithm to extract key point clouds: first flatten the input voxel data into a two-dimensional coordinate matrix and then transpose it; then call the DBSCAN algorithm to perform density clustering on the point cloud.

4. The cryo-electron microscopy density map similarity retrieval method according to claim 1, characterized in that: In S3: a hybrid parallel strategy of MPI and OpenMP is adopted; first, the MPI environment is initialized, the computing processes are configured and allocated to available resources, and then the input data is partitioned and distributed through the MPI communication mechanism; each process independently performs the following operations: extracting the point cloud structure and its key points, calculating the translation mask parameters, verifying the mask dimensional integrity, and performing local feature extraction to establish the correspondence relationship between the key points of the input dataset; the results of each process are aggregated and sent to the root process through the MPI collective communication protocol; the root process constructs the correspondence matrix based on the aggregated results, establishes the initial transformation matrix calculation optimization problem and solves it; The root process applies the Iterative Closest Point (ICP) algorithm to iteratively minimize the geometric distance between the transformed source point cloud and the target point cloud. After convergence, the registration score is calculated to quantitatively evaluate the transformation quality. Finally, the transformation matrix and registration score are broadcast to all processes through MPI communication primitives to maintain global consistency. The algorithm ends when the MPI environment terminates.

5. The cryo-electron microscopy density map similarity retrieval method according to claim 4, characterized in that: The hybrid MPI-OpenMP alignment algorithm achieves a combination of distributed and shared memory parallelism through five phases of collaborative work: Phase 1: MPI initialization and data preparation. The master process (rank 0) initializes the MPI environment and loads the source point cloud A_pcd and the target point cloud B_pcd. SHOT feature descriptors are calculated using the spherical histogram of the feature radius and voxel size to capture local geometric patterns. Spatial validity verification is achieved by comparing the bounding box normalized dimension inequality: When the inequality holds, the execution is terminated and a dimension mismatch warning is returned; Phase 2: Hierarchical parameter distribution uses a two-stage broadcast mechanism to transmit the BroadcastParams structure, and scalar parameters are transmitted through standard MPI_Bcast; the feature matrix F A With F B Distribution through custom broadcast_matrix: first communicate the matrix dimensions, then broadcast the Eigen:MatrixXd data block in row-major format, supporting the receiving node to dynamically adjust the buffer; Phase 3: Workload partitioning uses a block allocation strategy to divide the i-axis search space [0, terminal x ] is divided into MPI processes; the single process block size calculation formula is: Where δ is the indicator function that distributes the residual blocks to the low-rank processes; each process calculates the local boundary as start_i = rank × Δ + ξ and end_i = min (start_i + chunk, terminal x ), ξ is the initial offset compensation; Phase 4: Multi-granularity parallel execution The three-dimensional (i, j, k) loop nesting is parallelized through the collapse(3) instruction of OpenMP; each thread generates a center c mask =center+(iΔ, jΔ, kΔ), a spherical mask of radius r, performs constrained ICP registration through Registration_mask; this process includes super max_correspondence_dist correspondence filtering, initial registration based on SHOT features and voxel adaptive threshold adjustment; the successful transformation matrix and the registration score s l By caching Eigen matrices into thread-local buffers, critical section protection is implemented when aggregating global results; Phase 5: Result aggregation uses non-blocking MPI communication with four message tags; coordinate-score pairs (MPI_INT[3], MPI_DOUBLE) are transmitted via tags 1-2, and the transformation matrix is ​​serialized into an MPI_DOUBLE[16] array and transmitted via tags 3-4, with column-major storage maintained via explicit transposition: The main process asynchronously receives and splices the data stream, and writes the final result to the file to achieve Complexity scaling, where P is the number of MPI processes and T is the number of OpenMP threads per node.

6. The cryo-electron microscopy density map similarity retrieval method according to claim 1, characterized in that: In S4, a comprehensive scoring function is constructed based on the rigid body transformation matrix obtained by point cloud registration. The final similarity score is composed of the weighted sum of each normalized index. The weight ω reflects the relative importance of each index. The constructed multi-dimensional scoring function is as follows: final_score=w1·nor NS +w2·nor SHOT -w3·nor d_geo -w4·nor DC Among them, ω1 is used to evaluate the consistency of surface normal; ω2 represents the matching degree of local geometric features; ω3 evaluates the global alignment of point clouds in space; ω4 reflects the difference in local density of point clouds; nor NS is the normalized normal vector similarity score, nor SHOT is the normalized directional histogram signature SHOT feature similarity score, nor d_geo is the normalized geometric error score, nor DC is the normalized density consistency score.

7. The cryo-electron microscopy density map similarity retrieval method according to claim 1, characterized in that: In S4, the geometric error quantifies the degree of spatial alignment between the source point cloud and the target point cloud, and is defined as the mean Euclidean distance between the transformed source point cloud point ai and its corresponding point bj in the target point cloud, where K is the total number of corresponding points. This indicator reflects the positional deviation between the point clouds: Normal vector similarity (NS(A,B)) evaluates the degree of alignment of surface orientations between point clouds by calculating the cosine similarity of the normal vectors nai and nbj of corresponding points and introducing the corresponding weight wi for weighting. This indicator reflects the directional consistency of local surface features: Density consistency (DC(A,B)) evaluates the degree of matching of local density between point clouds and analyzes the consistency of neighborhood structure by comparing the number of neighborhood points within a predefined radius r; where NbrA(pi,r) and NbrB(qi,r) represent the number of neighborhood points within radius r of source point cloud point pi and target point cloud point qi, respectively: The SHOT feature similarity of the directional histogram signature (SHOT(A,B)) evaluates the similarity of local geometric features between point clouds by calculating the mean cosine similarity of the SHOT descriptors fai and fbj of all corresponding points, where K is the total number of corresponding points: Each indicator is normalized to ensure that its value range is in the interval [0,1]: normalized_score=max(0,min(1,normalized_score)).

Citation Information

Cited By

  • Unmanned aerial vehicle monitoring method and system for high slope displacement

    CN121767892A