High-throughput experimental data multi-dimensional index construction method based on material gene engineering
By constructing catalyst gene fingerprint (CGF) and multi-level index structure, data management problems in high-throughput catalyst screening experiments are solved, efficient multi-dimensional data retrieval and in-depth analysis are achieved, and query efficiency and similarity evaluation capabilities are improved.
Patent Information
- Application Number
- CN202510708122.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
The multi-dimensional data generated in high-throughput catalyst screening experiments are difficult to efficiently manage and retrieve, and existing indexing technologies cannot effectively support complex queries and potential associations of multi-dimensional features.
The catalyst gene fingerprint (CGF) is constructed by integrating composition features, experimental reaction condition features and thermal response features, using multi-level index structure and tensor decomposition technology, combined with query cost model, to achieve rapid retrieval and precise positioning of high-throughput experimental data.
The systematized integration and efficient management of high-throughput catalytic experimental data is realized, and the in-depth analysis and similarity evaluation of multi-dimensional features are supported, which improves the response efficiency in complex query scenarios.
Smart Images

Figure CN120578664A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of material index construction, and in particular to a method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering. Background Art
[0002] During high-throughput screening experiments, catalysts simultaneously generate and collect massive amounts of multidimensional data. This data includes not only the compositional characteristics of the catalyst itself (such as elemental composition, carrier information, additive type and content, and specific preparation process parameters), detailed experimental reaction conditions (such as temperature, pressure, space velocity, feed composition and flow rate), but also performance characterization data monitored in real time during the experiment, such as complex thermal response characteristics captured by infrared thermal imaging technology (such as ignition point, temperature distribution, and heating curve). This data is multidimensional, diverse, and massive, and often contains complex nonlinear correlations. Efficiently managing and retrieving this data and extracting valuable structure-activity relationship information from it are key challenges facing the field of materials genetic engineering.
[0003] In order to meet the above challenges, building an efficient indexing mechanism based on the characteristics of data in specific fields is a key technical means to improve data retrieval efficiency and support subsequent data analysis and knowledge discovery. Indexing technology can significantly reduce the amount of data that needs to be scanned during query by preprocessing and structuring the key features of the data, thereby speeding up retrieval. In the field of materials science, building an index that can effectively reflect the multi-dimensional characteristics of materials is crucial for achieving rapid positioning of material data, similarity comparison, and screening of potential excellent materials. A well-designed index should not only be able to process numerical and textual data, but also support queries based on complex feature combinations and adapt to the needs of the continuous growth and evolution of material data.
[0004] Patent application publication number CN113568906A discloses a distributed index structure and load balancing method for high-throughput data streams. This invention primarily targets multi-source, large-scale, and continuously generated data stream environments. Its core lies in the construction of a two-layer distributed index structure: a bottom-layer B+ tree index based on individual tuples in the data stream, used to index specific data records; a top-layer index based on data source and time window, typically using a hash table combined with a linked list, used to quickly locate data stream segments from a specific data source within a specific time range. This index construction fully incorporates the characteristics of its application area, high-throughput data streams, namely the significant temporal continuity and source diversity of data. It also emphasizes the use of multi-node collaboration (such as data stream receiving nodes, coordination nodes, query nodes, and storage nodes) and load balancing strategies in a distributed environment to ensure real-time data processing and system stability. However, the index construction technology proposed in this reference document focuses on the rapid reception and distribution of data streams, location by time / source, and overall system throughput and scalability. The deep indexing support for the complex semantic associations and multi-dimensional intrinsic features of the data content itself is relatively limited. For example, it does not perform special feature extraction and fusion for the diverse attribute combinations of the data entity itself to form a unified descriptor. Its B+ tree index mainly acts on certain preset fields of the stream tuple. There may be certain technical limitations in supporting complex retrieval based on the comprehensive similarity of multi-dimensional features or mining potential associations between different features. Summary of the Invention
[0005] In view of this, the present invention provides a method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering. By integrating the composition characteristics, experimental reaction condition characteristics and thermal response characteristics of the catalyst, a unified "catalyst gene fingerprint" is formed, and based on this fingerprint, an efficient multidimensional index structure is constructed to achieve rapid retrieval, precise positioning and intelligent analysis of massive catalytic experimental data.
[0006] The technical solution of the present invention is achieved as follows: The present invention provides a method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering, comprising: S1. Obtaining the composition characteristics of the predetermined catalytic material and the corresponding experimental reaction condition characteristics; S2. collecting infrared thermal imaging data of a predetermined catalytic material during a high-throughput experiment, and extracting thermal response characteristics of the predetermined catalytic material using a preset ignition criterion based on the infrared thermal imaging data; S3. Integrate composition characteristics, experimental reaction condition characteristics, and thermal response characteristics to generate a catalyst gene fingerprint that uniformly characterizes the multi-dimensional comprehensive properties of the predetermined catalytic material; S4. Based on the catalyst gene fingerprint, a multidimensional index of high-throughput experimental data is constructed. The multidimensional index includes a multi-level index structure, which is used to establish a mapping relationship between the catalyst gene fingerprint and the corresponding high-throughput experimental data.
[0007] Preferably, the catalyst gene fingerprint is defined as a combination of composition characteristics, experimental reaction condition characteristics and thermal response characteristics, and its definition formula is: , Where, is the genetic fingerprint of the predetermined catalytic material C, is the composition characteristic of the predetermined catalytic material C, is the experimental reaction condition characteristic of the predetermined catalytic material C, is the thermal response characteristic of the predetermined catalytic material C.
[0008] Preferably, the characterization of the composition characteristics, experimental reaction condition characteristics and thermal response characteristics is as follows: , , , Where, E is the element composition vector of the catalyst; L is the carrier characteristic vector; P is the catalyst preparation parameter vector; A is the additive component vector; is the reaction temperature, is the reaction pressure, is the gas space velocity, is the feed composition vector, is the flow parameter; is the ignition temperature; is the temperature peak eigenvector; is the temperature change rate eigenvector; is the spatial temperature distribution characteristic; is the feature descriptor vector of the temperature-time curve.
[0009] Preferably, the ignition temperature Determined by the normalized dynamic ignition criterion NDIC, which comprehensively considers the temperature rise rate and the degree of normalization of the current temperature relative to the reference baseline and the characteristic activation temperature; When the calculated normalized dynamic ignition criterion value continuously reaches the preset minimum confirmation point number and exceeds the preset judgment threshold, the ignition time is determined, and the temperature corresponding to the ignition time is used as the ignition temperature.
[0010] Preferably, the calculation formula of the normalized dynamic ignition criterion NDIC is as follows: , in, is the normalized dynamic ignition criterion value calculated at time i, is the average channel temperature at time i, is the average channel temperature at the previous moment, is the time interval between time i and time i-1, is the normalized temperature rise rate parameter, is the reference baseline temperature, is the characteristic activation temperature parameter, is the index adjustment factor.
[0011] Preferably, the multi-level index structure includes: The top-level index is constructed based on the catalyst gene fingerprint CGF and adopts indexing technology based on space partitioning or graph embedding to achieve rapid clustering and retrieval of similar catalyst gene fingerprints; The middle-level index associates catalyst gene fingerprints with key derived features generated by high-throughput experiments and is organized using a multidimensional B-tree or R-tree structure; The underlying index associates key derived features with the storage location and access path of the original experimental data.
[0012] Preferably, the method further comprises: Construct an associative index layer in the multi-dimensional index and use tensor decomposition technology to capture the nonlinear relationship between the features of each dimension in the catalyst gene fingerprint; Tensor decomposition represents the multidimensional features of catalyst gene fingerprints as a high-order tensor and decomposes it into a combination of a series of low-order feature vectors.
[0013] Preferably, the calculation formula for tensor decomposition is: , in, Represents an associated index, represents the tensor outer product operation, , , , They represent the characteristic vectors of the rth component in the four dimensions of composition, structure, performance and experimental conditions, and R represents the selected rank number.
[0014] Preferably, the method further comprises: For hybrid storage architectures that include relational databases and distributed object storage, an adaptive path selection mechanism based on a query cost model is introduced; When receiving a joint query request across storage systems, the query cost model evaluates the query costs of different index paths. The query cost model considers factors such as the average time to access the index and retrieve data, data transmission overhead, and the cost of cross-data source join operations. The optimal data access path is selected based on the evaluation results to execute the query.
[0015] Preferably, the calculation formula of the query cost model is: , in, is the total query cost, M is the number of data sources involved in the query, Index for accessing the jth data source and the average time to retrieve data, is the cost of transmitting data from the jth data source, and is the corresponding weight coefficient, K is the number of cross-data source connection operations involved in the query, is the cost of performing the k-th join operation, For its weight.
[0016] The present invention has the following beneficial effects compared to the prior art: (1) The present invention achieves systematic integration and efficient management of massive and heterogeneous data generated by high-throughput catalytic experiments by constructing a unified catalyst gene fingerprint and establishing a multi-dimensional and multi-level index structure based on it; (2) The "catalyst gene fingerprint (CGF)" proposed in this invention uniformly encodes multiple sources of information, such as catalyst composition, experimental reaction conditions, and key thermal response characteristics, into standardized digital descriptors. This fingerprint representation not only lays the foundation for quantitative comparison and similarity assessment between different catalysts, but also facilitates the use of data analysis methods such as machine learning to extract deep knowledge from multi-dimensional features, promoting the understanding of catalytic mechanisms and structure-activity relationships; (3) By using methods such as the Normalized Dynamic Ignition Criterion (NDIC) to accurately extract thermal response characteristics (such as ignition temperature) from infrared thermal imaging data, the present invention can more objectively and stably quantify the thermal behavior of catalysts under specific conditions; (4) The multi-level and multi-dimensional index structure constructed by the present invention is optimized for the different feature dimensions and query requirements of catalyst gene fingerprints. The top-level index can achieve rapid clustering and preliminary screening of similar CGFs, while the middle and bottom-level indexes can further associate specific derived features and original experimental data. This hierarchical design effectively balances the search scope and search depth, improving the response efficiency in complex query scenarios; (5) This invention introduces tensor decomposition technology to construct an associative index layer, which can effectively capture the complex nonlinear coupling relationships between the various dimensional features in the catalyst gene fingerprint. Compared with traditional indexes based on single features or simple combination of features, this method can reveal deeper feature patterns, thereby providing more accurate and insightful results when performing similarity retrieval or pattern recognition; (6) In view of the fact that data may be stored in a hybrid architecture such as a relational database and a distributed object storage in actual applications, the adaptive path selection mechanism based on the query cost model proposed in this invention can dynamically select the optimal data access path according to the query request and data distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a technical implementation diagram of the present invention; Figure 3 This is a schematic diagram of infrared thermal imaging data processing according to the present invention; Figure 4 Schematic diagram of the construction process of the catalyst gene fingerprint of the present invention Figure 5 Schematic diagram of the catalyst gene fingerprint structure of the present invention Figure 6 A simplified schematic diagram of the multi-dimensional index structure of the present invention Figure 7 This is a schematic diagram of an example of query path optimization of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] like Figure 1 As shown, the present invention provides a method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering, comprising: S1. Obtaining the composition characteristics of the predetermined catalytic material and the corresponding experimental reaction condition characteristics; S2. collecting infrared thermal imaging data of a predetermined catalytic material during a high-throughput experiment, and extracting thermal response characteristics of the predetermined catalytic material using a preset ignition criterion based on the infrared thermal imaging data; S3. Integrate composition characteristics, experimental reaction condition characteristics, and thermal response characteristics to generate a catalyst gene fingerprint that uniformly characterizes the multi-dimensional comprehensive properties of the predetermined catalytic material; S4. Based on the catalyst gene fingerprint, a multidimensional index of high-throughput experimental data is constructed. The multidimensional index includes a multi-level index structure, which is used to establish a mapping relationship between the catalyst gene fingerprint and the corresponding high-throughput experimental data.
[0021] like Figure 2 As shown, the technical approach of the present invention is based on the concept of materials genetic engineering. Aiming at the complex and multidimensional characteristics of high-throughput catalytic experimental data, the present invention first constructs a unified "catalyst gene fingerprint (CGF)" by integrating the composition characteristics of the catalyst (elemental composition, carrier characteristics, preparation parameters, auxiliary agent components, etc.), experimental reaction condition characteristics (reaction temperature, pressure, space velocity, feed composition, etc.), and thermal response characteristics extracted from infrared thermal imaging data (such as the ignition temperature determined by the normalized dynamic ignition criterion (NDIC)). Then, based on this standardized digital descriptor, a multi-dimensional index structure with multiple levels is constructed. The top-level index is used for rapid clustering and preliminary screening of similar CGFs, the middle-level index is associated with specific feature descriptions, and the bottom-level index points to the original experimental data storage location. Tensor decomposition technology is also introduced to construct an associative index layer to capture the complex nonlinear relationships between features. Furthermore, an adaptive path selection mechanism based on a query cost model and a query-driven index evolution mechanism are designed to achieve efficient data retrieval, similarity search, and pattern recognition in heterogeneous storage environments.
[0022] Specifically, in one embodiment of the present invention, step S1 includes: Obtain the composition characteristics of the intended catalytic material: Compositional characteristics refer to a series of parameters that describe the physical and chemical composition of the catalyst sample itself and its preparation history. These characteristics are determined or recorded during the catalyst design, synthesis, or preparation stage. Specifically, in this embodiment, the compositional characteristics obtained mainly include: a. Elemental composition of the catalytic material: This includes the identities and relative amounts or molar ratios of all elements that comprise the catalyst's active components and support. For example, a typical three-way catalyst might include the precise loading of precious metals (e.g., Pt, Pd, Rh), the ratios of primary support elements (e.g., Al, Si, Ce, Zr), and the identities and amounts of any other possible doping elements. This information can be derived from the catalyst's design formulation or from chemical analysis (e.g., X-ray fluorescence spectroscopy (XRF) or inductively coupled plasma optical emission spectroscopy (ICP-OES)).
[0023] b. Catalytic material support characteristics: The support is a crucial component of the catalyst, and its characteristics directly influence its performance. This study primarily focuses on the support type (e.g., γ-Al2O3, SiO2, TiO2, CeO2-ZrO2 composite oxides, etc.), physical structural parameters (e.g., specific surface area, pore volume, average pore diameter), and possible morphological features (e.g., nanoparticles, nanorods, mesoporous structures, etc.). These data are determined using materials characterization techniques (e.g., BET nitrogen adsorption, X-ray diffraction (XRD), and transmission electron microscopy (TEM).
[0024] c. Catalytic material preparation process parameters: The catalyst preparation method and process conditions significantly influence its final microstructure and catalytic performance. Therefore, key preparation process parameters must be obtained, such as calcination temperature, calcination time, impregnation method, pH value, precursor type, drying conditions, and reduction / activation conditions. These parameters should be recorded in the experimental operation flow or batch production records.
[0025] d. Catalytic additive components: Additives are small amounts of substances added to improve the activity, selectivity, or stability of the catalyst. The chemical composition, dosage, or ratio of the additive to the primary active component should be known. For example, specific information on additives such as alkali metals (e.g., K), alkaline earth metals (e.g., Ba), or certain transition metal oxides (e.g., La2O3) should be provided.
[0026] The above composition characteristic data will be structured after acquisition. For example, for elemental composition, it is expressed in the form of a chemical formula or a vector of mass percentage / molar percentage of each element; for carrier characteristics and preparation parameters, it can be a specific numerical value, text description or predefined classification code.
[0027] Get the corresponding experimental reaction condition characteristics: Experimental reaction condition characteristics refer to the specific external environmental parameters of a predetermined catalytic material during performance evaluation or screening experiments. These parameters directly determine the thermodynamic and kinetic processes of the catalytic reaction and, together with the compositional characteristics of the catalyst, influence its ultimate apparent performance. In high-throughput experiments, each catalyst sample (or each reaction well / channel) corresponds to a specific set of experimental conditions. Specifically, in this embodiment, the experimental reaction condition characteristics obtained mainly include: a. Experimental Reaction Temperature: This refers to the setpoint or measured temperature of the catalyst bed or reactor during the experiment. This may be a constant value or a sequence of temperatures during a temperature programming process. Units are Celsius (°C) or Kelvin (K).
[0028] b. Experimental reaction pressure: refers to the total pressure of the reaction system.
[0029] c. Experimental reaction gas space velocity: Gas Hourly Space Velocity (GHSV), defined as the volume flow rate of gas passing through a unit volume of catalyst per unit time.
[0030] d. Feed composition of the experimental reaction: refers to the concentration or mole fraction of each component in the gas mixture entering the reactor.
[0031] e. Feed flow rate of the experimental reaction: refers to the total volume flow rate or mass flow rate of the gas mixture entering the reactor per unit time, as well as the individual flow rate of each component.
[0032] These experimental reaction condition characteristic data are set and recorded by the control system of the high-throughput experimental platform, or are collected in real time through various sensors connected to the reaction device (such as thermocouples, pressure sensors, mass flow controllers, etc.).
[0033] After completing step S1, the system obtains a complete set of "original gene" information for each predetermined catalytic material, namely its intrinsic "composition characteristics" and the external "experimental reaction condition characteristics" it has experienced.
[0034] Specifically, if Figure 3 As shown, in one embodiment of the present invention, step S2 includes: Collect infrared thermal imaging data of the predetermined catalytic material during high-throughput experiments: This example uses a high-throughput catalyst screening device to test a predetermined catalytic material. This device comprises a multi-channel parallel reactor array, capable of simultaneously testing a variety of catalyst samples with varying compositions. During the experiment, a high-resolution infrared thermal imager was trained on the catalyst array to capture real-time temperature distribution information on the sample surface.
[0035] Specifically, an infrared thermal imager continuously monitors the catalyst samples in the array reactor at a sampling rate of 10 frames per second, recording changes in their surface temperature throughout the temperature-programmed experiment. For each predetermined catalytic material, the system defines a corresponding region of interest (ROI) to precisely track the temperature evolution of that catalyst channel during the experiment.
[0036] For each catalyst sample point, the acquired infrared thermal imaging data includes the following aspects: a time-varying 2D temperature distribution sequence; a time-varying curve of the average temperature within the catalyst ROI; temperature distribution statistics within the catalyst ROI (e.g., maximum temperature, minimum temperature, standard deviation, etc.); and the relative temperature difference with a reference area (e.g., an inactive area or a blank control channel). This raw infrared thermal imaging data undergoes preliminary preprocessing, including spatial registration, temperature calibration, and noise filtering, to eliminate systematic errors and random interference caused by the instrument itself.
[0037] Based on infrared thermal imaging data, the thermal response characteristics of the predetermined catalytic material are extracted using the preset ignition criteria: After obtaining the pre-processed infrared thermal imaging data, this embodiment uses the normalized dynamic ignition criterion (NDIC) to extract the thermal response characteristics of the catalyst, especially the ignition temperature Compared with the traditional fixed temperature threshold method, NDIC can more robustly identify the ignition moment and reduce the impact of fluctuations in experimental conditions.
[0038] The calculation formula of normalized dynamic ignition criterion (NDIC) is: , in, is the normalized dynamic ignition criterion value calculated at time i, is the average temperature in the sample channel ROI at time i, is the average channel temperature at the previous moment, is the time interval between time i and time i-1, that is, the sampling period of the infrared thermal imager, is the normalized temperature rise rate parameter, As the reference baseline temperature, the average value of the ROI temperature during a stable period after the reaction starts is taken. is the characteristic activation temperature parameter, which is the typical temperature threshold at which the catalyst is significantly activated or undergoes a violent reaction, set based on the prior knowledge of the catalyst system. is an exponential adjustment factor, set to a constant greater than 1, used to amplify the criterion response when the temperature approaches or exceeds the characteristic activation temperature.
[0039] The calculation process of the NDIC criterion is as follows: First, calculate the temperature rise rate at each moment i , and by normalizing the parameter Standardize; calculate the relative temperature at each moment i , characterizes the normalization degree of the current temperature relative to the reference baseline and characteristic activation temperature; using the exponential adjustment factor Enhance the response of the criterion when the temperature approaches or exceeds the characteristic activation temperature; combine the two factors of temperature rise rate and relative temperature to obtain the NDIC value at each moment.
[0040] Ignition judgment criteria are as follows: preset minimum confirmation points For 5 consecutive sampling points (corresponding to a 0.5 second time window in this embodiment), the judgment threshold is preset is 0.6. When the calculated Continuous values The sampling points all stably exceed the preset judgment threshold , the first time point that meets this condition The ignition time of the catalyst sample is determined. According to the experimental design, the temperature corresponding to the current experimental reaction conditions (usually the inlet gas temperature or furnace temperature in the programmed temperature experiment) is recorded as the ignition temperature of the catalyst ( ).
[0041] Except ignition temperature ( ), this embodiment also extracts other thermal response features from the infrared thermal imaging data to form a complete thermal response feature vector TRF(C): Temperature peak eigenvector ( ): Contains the highest temperature reached by the catalyst ROI during the entire experiment ( ) and its appearance time ( ), and the average temperature of the steady-state section after ignition ( ). This eigenvector can reflect the exothermic intensity and persistence of the catalytic reaction.
[0042] Temperature change rate eigenvector ( ): Describes the temperature rise rate at different stages, including the average temperature rise rate at the ignition stage ( ), the maximum temperature rise rate during the process of reaching the highest temperature ( ) and the temperature fluctuation rate during the stable reaction phase ( ). These parameters can characterize different aspects of the kinetics of catalytic reactions.
[0043] Spatial temperature distribution characteristics ( ): obtained through statistical analysis of infrared thermal imaging frames, including the standard deviation of temperature distribution within the ROI ( ), the difference between the highest temperature point and the average temperature ( ), the maximum temperature gradient ( ) and the size and shape characteristics of hot spots. These spatial characteristics can reflect the uniformity of catalyst surface activity distribution and the characteristics of hot spot formation.
[0044] The characteristic descriptor vector of the temperature-time curve ( ) : Frequency domain features extracted from the temperature curve via wavelet transform or Fourier analysis, as well as statistical features such as curve inflection points and integrated area. These descriptors can capture subtle features of the temperature change process and reflect the complexity of the catalytic reaction mechanism.
[0045] Through the above steps, this embodiment completes the comprehensive extraction of the thermal response characteristics of the predetermined catalytic material.
[0046] The advantages of this thermal response feature extraction method are: on the one hand, the ignition temperature determination method based on the normalized dynamic ignition criterion (NDIC) is more robust than traditional methods and can adapt to different experimental conditions and the characteristics of different types of catalysts; on the other hand, the extracted multidimensional thermal response features comprehensively characterize the behavioral characteristics of the catalyst during the reaction process.
[0047] Specifically, if Figure 4 As shown in Figure 1, step S3 is the core link in constructing the catalyst genetic fingerprint (CGF). Its main task is to organically integrate the composition characteristics, experimental reaction condition characteristics, and thermal response characteristics obtained in the first two steps S1 and S2 to form a unified digital descriptor that can comprehensively characterize the multidimensional properties of catalytic materials.
[0048] S3.1 Basic definition and composition of catalyst gene fingerprint According to the core concept of the present invention, the catalyst gene fingerprint (CGF) is a ternary combination of the composition characteristics, experimental reaction condition characteristics, and thermal response characteristics of the predetermined catalytic material C. Its mathematical definition is as follows: , Where, is the genetic fingerprint of the predetermined catalytic material C, is the composition characteristic of the predetermined catalytic material C, is the experimental reaction condition characteristic of the predetermined catalytic material C, is the thermal response characteristic of the predetermined catalytic material C. These three types of characteristics respectively characterize the information of the catalytic material in three dimensions: "intrinsic properties", "external environment" and "performance response".
[0049] Specifically, if Figure 5 As shown in Figure 2, the detailed composition of each dimension feature is as follows: Compositional characteristics CF(C) are characterized by the chemical composition and structural properties of the catalyst: , Here, E is the elemental composition vector of the catalyst, which records the quantitative information of each element contained in the catalyst and can be expressed as a vector of length 118 (corresponding to the number of elements in the periodic table). Each component in the vector represents the mass percentage or molar percentage of the corresponding element in the catalyst. L is the carrier characteristic vector, which describes the basic characteristics of the catalyst carrier, including: carrier type code, surface area, average pore size, pore volume, crystal phase composition, crystallinity and other parameters. P is the catalyst preparation parameter vector, which records the key process parameters of the catalyst preparation process, including: preparation method code, calcination temperature, calcination time, calcination atmosphere, reduction temperature, reduction time, reducing atmosphere, pH value and other parameters. A is the additive component vector, which records the information of the additives added to the catalyst. It uses a coding method similar to the elemental composition vector to indicate the type and content of the additives.
[0050] The experimental reaction conditions are characterized by PCF(C) by describing the process parameters of the catalytic reaction: , in, is the reaction temperature, which can be a single temperature value or a temperature program descriptor, is the reaction pressure, i.e. the system pressure during the experiment, is the gas space velocity, defined as the volume flow of gas passing through a unit volume of catalyst per unit time, is the feed composition vector, recording the content of each component in the reaction gas mixture, It is the flow parameter, including the total flow and the flow information of each component.
[0051] Thermal response feature (TRF) (C) is a key feature extracted from infrared thermal imaging and temperature time series data obtained from a catalyst rapid screening device: , in, is the light-off temperature, the catalyst light-off temperature determined by the normalized dynamic light-off criterion (NDIC) method in step S2; is the temperature peak feature vector, including the highest temperature When the maximum temperature is reached And the average temperature of the steady state after ignition ; is the temperature change rate feature vector, including the average temperature rise rate during the ignition stage , maximum temperature rise rate and the temperature fluctuation rate during the stable reaction phase ; is the spatial temperature distribution feature, including the standard deviation of the temperature distribution within the ROI , the difference between the highest temperature point and the average temperature , maximum temperature gradient and other statistical indicators; is the characteristic descriptor vector of the temperature-time curve, the frequency domain features extracted from the temperature curve by wavelet transform or Fourier analysis, as well as statistical features such as the curve inflection point and integral area.
[0052] S3.2 Preprocessing and Normalization of Feature Vectors Before integrating the three types of features, a series of preprocessing operations are required on the original feature vector to eliminate dimensional differences, handle missing values and outliers, and standardize the values.
[0053] For characteristic values that may be missing during actual measurement or recording, the following strategies are used to fill them: For missing values in composition characteristics, reasonable inference and filling are made based on stoichiometric relationships or the average values of similar catalysts; for missing values in experimental condition characteristics, the standard operating parameters of the experimental sequence are used to fill them; for missing values in thermal response characteristics, interpolation algorithms or corresponding characteristic values of similar thermal response curves are used to fill them.
[0054] In this embodiment, missing value detection applies an algorithm based on pattern recognition and is filled in with an expert knowledge system: , in, is the corresponding feature mean of the K most similar samples (in other feature dimensions) found by the K-nearest neighbor (KNN) algorithm, It is a value inferred based on the laws of catalytic chemistry. It is the global statistical value of the feature dimension (such as mean or median).
[0055] For possible abnormal measurement values, the improved Z-score method is used for detection: , in, is the absolute median difference: , when When , it is determined to be a potential outlier and further confirmed by combining domain knowledge. For confirmed outliers, you can choose to replace them with reasonable estimates or special marks.
[0056] To eliminate the dimensional differences between different features, Min-Max normalization is applied to all numerical features to map the values to the [0, 1] interval.
[0057] For categorical features such as carrier type and preparation method, one-hot encoding or embedding vector representation is used: One-Hot Encoding: Convert category features into binary vectors with a length equal to the number of categories, where only the positions corresponding to the category are 1 and the rest are 0. Embedding Vector: For category features with intrinsic relationships (such as the physicochemical associations between different carrier materials), pre-train low-dimensional embedding vectors to capture the semantic similarities between categories.
[0058] S3.3 Feature Dimensionality Reduction and Optimization Considering that the original feature dimensions are high and redundant, this embodiment adopts the following dimensionality reduction strategy: Domain knowledge-based feature screening: Based on the principles of catalytic chemistry, key features that have a significant impact on catalytic performance are retained, such as the type and content of active metals, and the properties of the support.
[0059] Principal Component Analysis (PCA): Apply PCA dimensionality reduction to the subset of features with high correlation, retaining the principal components with a cumulative contribution rate of 95%: X′=X·W PCA Among them, W PCA is the PCA transformation matrix, which consists of eigenvectors.
[0060] Through the above processing, the three types of eigenvectors are optimized into a more compact form: .
[0061] S3.4 Feature Fusion and Fingerprint Generation This example uses a hierarchical feature fusion strategy to generate the final catalyst gene fingerprint: According to the degree of influence of each feature on catalytic performance, weight coefficients are assigned to the three types of features. 、 and ,satisfy .
[0062] The initial fused feature vector is calculated as follows: , In order to capture the interaction between different types of features, a second-order interaction feature is constructed: , in represents the tensor product operation, retaining the most informative interaction terms through feature importance evaluation.
[0063] Combine the directly connected feature vectors with the interaction features to form enhanced features: , Then, the nonlinear dimensionality reduction technique t-SNE is applied to generate a fixed-length catalyst gene fingerprint: , in, Indicates that the dimension of the final CGF is set to 128.
[0064] S3.5 CGF Storage and Index Preprocessing The generated catalyst gene fingerprints are stored in a standardized format and subjected to the following preprocessing: Binary Representation: Quantize the 128-dimensional floating-point CGF vector to 8-bit precision, forming a compact binary representation. Similarity Hash Preprocessing: Apply locality-sensitive hashing technology to generate a hash code for each CGF, facilitating fast similarity retrieval. Metadata Annotation: Add metadata tags such as generation time, data source, and feature integrity score to each CGF to facilitate subsequent data management.
[0065] Through the above steps, this example successfully integrated the compositional characteristics, experimental reaction conditions, and thermal response characteristics of the catalytic material into a unified Catalyst Genetic Fingerprint (CGF). This multi-dimensional comprehensive characterization establishes a "digital identity" for the catalyst, enabling quantitative comparison and similarity assessment between different catalytic materials.
[0066] Specifically, if Figure 6 As shown, in one embodiment of the present invention, step S4 constructs a multidimensional index structure based on the catalyst gene fingerprint (CGF) generated in step S3, which can efficiently associate various types of heterogeneous data. This index structure must not only support rapid retrieval and similarity search of catalyst characteristics, but also achieve accurate mapping from catalyst digital fingerprints to original experimental data to meet the complex multi-level and multi-dimensional query requirements of high-throughput catalysis research.
[0067] S4.1 Overall Architecture of Multi-Level Index Structure Based on the logical hierarchy and storage characteristics of catalytic material data, this embodiment designs a three-layer vertical index structure, which is tightly coupled with the three-layer logical architecture of the material database (basic metadata layer, derived feature data layer, and raw data index layer): The top-level index is constructed based on the catalyst gene fingerprint CGF and adopts indexing technology based on space partitioning or graph embedding to achieve rapid clustering and retrieval of similar catalyst gene fingerprints; The middle-level index associates catalyst gene fingerprints with key derived features generated by high-throughput experiments and is organized using a multidimensional B-tree or R-tree structure; The underlying index associates key derived features with the storage location and access path of the original experimental data.
[0068] This vertical index structure allows users to start from the intrinsic properties of the catalyst and go deeper into specific experimental phenomena and original evidence layer by layer, achieving full-link traceability from "catalyst genes" to "performance" to "experimental data."
[0069] S4.2 Top-level index construction method The top-level index is a high-dimensional feature index built for the catalyst gene fingerprint (CGF). Its main function is to support the rapid clustering and retrieval of similar catalysts. Considering the high-dimensional nature of CGF (128 dimensions) and the complex similarity measurement requirements, this embodiment adopts a hybrid indexing strategy based on combining space partitioning and graph structure: The HNSW (Hierarchical Navigable Small World) algorithm is used to construct an approximate nearest neighbor index. HNSW achieves efficient indexing of high-dimensional spaces by constructing a multi-level nearest neighbor graph structure: First, the CGF vectors of all catalysts are normalized to apply the cosine similarity metric; a hierarchical graph structure with L layers is constructed, where L is dynamically determined according to the data size N: L ≈ log(N); A hierarchical graph By defining the maximum out-degree The index construction process follows the process of "entry point selection-greedy search-neighborhood update".
[0070] To improve the retrieval efficiency under large-scale data, this embodiment combines feature partitioning with locality sensitive hashing (LSH) technology: The 128-dimensional CGF is divided into k sub-vectors (in this embodiment, k=4, and each sub-vector has 32 dimensions). An independent LSH hash table is constructed for each sub-vector, and a random projection vector is generated using the p-stable distribution. The hash function is defined as: , where a is a random projection vector, b is a random offset, and w is the bucket width; construct L hash tables (in this embodiment, L=10), each hash table uses a cascade of k hash functions.
[0071] S4.3 Method for constructing mid-level index The middle-level index is responsible for associating the catalyst gene fingerprint with the key derived features generated by the experiment, allowing users to perform precise queries or range filtering based on performance indicators. Considering the multidimensionality and structured characteristics of derived features, this embodiment adopts an improved R* tree structure for organization: Feature vector construction: For each catalyst C, construct a multidimensional vector containing its CGF identifier and key derived features , where P1~Pn are key performance indicators (such as ignition temperature, conversion rate, selectivity, etc.); R-tree construction: based on Construct an R-tree and optimize the tree structure by minimizing the perimeter, overlapping area, and covered area of the bounding rectangles; dynamic balancing strategy: use forced reinsert technology to handle node overflow to maintain tree balance and query efficiency.
[0072] For frequent range queries and multi-condition combination queries, this embodiment also constructs a multi-dimensional B+ tree index group: For each important derived feature Build a B+ tree index with the key value ; Adopts composite key design to support single-dimensional range query and multi-dimensional joint query; introduces prefix compression and node caching mechanism to optimize index storage and access efficiency.
[0073] The design of the mid-level index takes into account common query patterns in materials research, such as "find all catalysts with an ignition temperature between 200-250°C and a CO conversion rate greater than 90%." The optimized multi-dimensional index structure enables efficient execution of such queries.
[0074] S4.4 Construction method of underlying index The underlying index is responsible for associating the derived features with the storage location and access path of the original experimental data, solving the problem of unified access to heterogeneous data sources in high-throughput catalysis experiments. This embodiment designs a mapping index structure that adapts to the hybrid storage environment: For each catalyst sample C, construct its raw data mapping table , including the following fields: data type identifier (such as infrared thermal imaging video, mass spectrometry data, gas chromatography, etc.); storage system identifier (such as object storage system OBS1, file system FS2, etc.); storage path / object ID (specific access address); data segmentation information (such as video time segment, sensor data sampling range); access permission tag (data security level and access control information).
[0075] In order to support refined access to raw data (such as extracting only the fragments of the ignition stage in the infrared video), this embodiment constructs a fine-grained data block index: logically divides large raw data files (such as infrared thermal imaging videos) into blocks; records metadata information for each data block, including timestamps, corresponding experimental stage tags, etc.; and constructs a block-level inverted index to support precise positioning based on experimental stages or phenomenon characteristics.
[0076] Taking into account the heterogeneity of hybrid storage environments, a unified data access adaptation layer is designed: it encapsulates the interface differences of different storage systems and provides standardized data access methods; implements intelligent caching strategies based on data characteristics to optimize the read performance of frequently accessed data; and supports asynchronous prefetching and batch reading to reduce the latency overhead of access across storage systems.
[0077] S4.5 Construction of the associated index layer In order to more accurately capture the complex nonlinear relationships between the multidimensional features of catalysts, this embodiment adds a dedicated correlation index layer in addition to the basic three-layer index structure, and uses tensor decomposition technology to establish multidimensional correlations: The multidimensional characteristics of the catalyst are represented as a high-order tensor T, whose dimensions include composition, structure, performance and experimental conditions: Composition dimension: captures characteristics such as elemental composition and carrier type; structural dimension: characterizes physical structural properties (such as specific surface area, pore size distribution, etc.); performance dimension: records catalytic performance indicators (such as ignition temperature, conversion rate, etc.); condition dimension: describes experimental condition parameters (such as reaction temperature, gas space velocity, etc.).
[0078] Apply the Tucker decomposition method to decompose the high-dimensional tensor T into the product of the core tensor G and the factor matrix: , in, Represents an associated index, represents the tensor outer product operation, , , , They represent the characteristic vectors of the rth component in the four dimensions of composition, structure, performance and experimental conditions, and R represents the selected rank number.
[0079] The specific implementation steps are as follows: solve Tucker decomposition by alternating least squares (ALS); use nuclear norm regularization to avoid overfitting; and determine the optimal rank number R (R=20 in this embodiment) through cross-validation.
[0080] Associative index constructed based on tensor decomposition results Supports the following functions: Material property prediction: predict potential catalytic performance based on composition and structural features; Condition optimization recommendation: recommend possible optimal experimental conditions for catalysts with specific compositions; Similarity metric enhancement: provide high-order similarity metrics that take feature interactions into account.
[0081] S4.6 Design and implementation of similarity metric functions To accurately assess the similarity between catalysts and support similarity-based retrieval and clustering, this example designs a configurable weighted comprehensive similarity metric function. For two catalysts A and B, whose gene fingerprints are CGF(A) and CGF(B), respectively, the similarity S(A, B) between them can be calculated using the following formula: , in, is the dimension of the gene fingerprint vector, and are the i-th fingerprint feature values of catalysts A and B respectively, is the weight factor of the corresponding feature, reflecting the importance of the feature in a specific application scenario. is the normalized scale parameter of the i-th feature, which is used to eliminate the influence of the dimension difference of different features. It can be set based on expert knowledge or learned from historical data through machine learning methods. The similarity metric function supports neighbor search in the top-level index, similar pattern discovery in the associated index, and user-oriented similar catalyst recommendation functions.
[0082] S4.7 Query Processing and Path Optimization To adapt to hybrid storage architectures and optimize multimodal query performance, this paper proposes an adaptive indexing strategy. When a user initiates a joint query across storage systems (for example, "Find all catalysts that possess a specific genetic signature X and an ignition temperature below Y under specific experimental conditions, and display infrared thermal imaging sequences of their ignition processes"), the query processor first decomposes the query into subquery components, including composition conditions, structural conditions, performance conditions, and data retrieval requirements. Leveraging the aforementioned multi-level indexing, candidate catalysts that meet the genetic signature and derived signature conditions, along with their associated metadata, are efficiently located within the relational database. Subsequently, based on the obtained raw data index information, the corresponding raw data segments are accurately batch-extracted from the distributed object storage system.
[0083] To optimize cross-storage access queries, this paper introduces an adaptive path selection mechanism based on a query cost model. For a query request Q, its optimal access path P* is determined by the following formula: , in, Represents the set of all possible index paths, is the query cost evaluation model: , Where M is the number of data sources (or data layers) involved in the query, is the index to access the jth data source and the average time to retrieve data, is the cost of transmitting data from the jth data source, and is the corresponding weight coefficient. K is the number of cross-data source join operations involved in the query. is the cost of performing the kth join operation, is its weight. The system dynamically adjusts the index structure by analyzing historical query logs and performance indicators. Or precompute some join results to minimize the overall query cost .
[0084] Execute the query according to the selected optimal execution plan, integrate the results from different data sources, and return the final query results. Figure 7 As shown, Figure 7 This is an example diagram of query path optimization. The adaptive path selection mechanism of this embodiment can intelligently select the most efficient data access path based on the specific query characteristics and current system load conditions, significantly improving the response performance of complex queries.
[0085] S4.8 Dynamic Evolution Mechanism of Indexes To achieve dynamic index optimization, this paper introduces a query-driven index evolution mechanism. This mechanism monitors the access frequency and query response time of different index paths. For frequently accessed raw data or specific query combinations, the system can dynamically create auxiliary indexes or materialized views, and even build lightweight metadata indexes or pre-aggregated results on the object storage side to accelerate responses to subsequent similar queries. The index evolution process can be expressed as: , in, Represents the index structure at time t, represents the index adjustment function, is the query set at time t, is a set of query paths, The index is the result of performance evaluation. Specific strategies for index tuning include: accelerating indexes for frequently accessed paths, reconstructing or eliminating inefficient indexes, and pre-building indexes for emerging query patterns.
[0086] When new catalyst sample data is added, the system performs incremental index updates: calculates the catalyst gene fingerprint CGF of the new sample; adds corresponding index items to each layer of the index structure; adjusts the index balance as needed (such as reorganization of the R* tree); and updates the tensor decomposition results of the associated index layer.
[0087] In practical applications, the implementation process of this index construction method is as follows: First, for each sample in the catalyst database, a complete catalyst gene fingerprint (CGF) is calculated based on its composition information, experimental conditions, and infrared thermal imaging data. Secondly, a multi-level vertical index structure is constructed, and a cross-storage index mapping relationship is established. Subsequently, the system continuously monitors query performance and access patterns during operation, selects the optimal access path based on the query cost model, and optimizes the index structure through an index evolution mechanism. When new catalyst sample data is added, the system automatically calculates its gene fingerprint and updates the relevant index nodes to maintain the integrity and validity of the index structure.
[0088] The present invention proposes a method for constructing a high-throughput catalyst data multi-dimensional index based on material genetic engineering, which realizes efficient indexing and retrieval of multimodal data of catalytic materials. First, the introduction of catalyst gene fingerprints provides a unified framework for systematically characterizing the multi-dimensional characteristics of catalysts, establishes a mapping relationship between composition characteristics, reaction conditions and thermal response performance, and makes the retrieval results more accurate and comprehensive. The concept of "catalyst gene fingerprint" expresses the physicochemical properties of catalysts and experimental observation data in the same framework, realizing the integration of full-chain data from catalyst design to performance evaluation. Secondly, the multi-level vertical index architecture supports different retrieval requirements from simple attribute queries to complex association analysis, allowing users to start from the intrinsic properties of the catalyst and go deeper into specific experimental phenomena and original evidence layer by layer. The association index layer constructed by tensor decomposition technology can effectively capture the nonlinear relationship between the multi-dimensional characteristics of the catalyst, greatly improving the retrieval accuracy under complex conditions. The adaptive indexing strategy is optimized for hybrid storage environments, effectively solving the performance bottleneck of joint queries in an environment where relational databases and distributed object storage coexist, and significantly reducing the response time of complex queries containing a large number of infrared thermal imaging videos. The query-driven index evolution mechanism and the adaptive path selection based on the query cost evaluation model ensure that the index structure can be continuously optimized as the data scale grows and the query pattern changes, maintaining the long-term efficient operation of the system. For catalytic material research and development, the indexing method of the present invention enables researchers to quickly find candidate materials under specific composition and performance conditions, and deeply analyze their reaction mechanisms, greatly accelerating the discovery and optimization process of new catalytic materials. Overall, the present invention provides a systematic solution for the efficient management and in-depth utilization of high-throughput catalyst experimental data by combining the concept of material genetic engineering with modern database indexing technology, thereby promoting the development of catalytic material informatics.
[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering, characterized in that: include: S1. Obtaining the composition characteristics of the predetermined catalytic material and the corresponding experimental reaction condition characteristics; S2. collecting infrared thermal imaging data of a predetermined catalytic material during a high-throughput experiment, and extracting thermal response characteristics of the predetermined catalytic material using a preset ignition criterion based on the infrared thermal imaging data; S3 Integration Composition characteristics, experimental reaction condition characteristics, and thermal response characteristics are used to generate catalyst gene fingerprints that uniformly characterize the multi-dimensional comprehensive properties of the predetermined catalytic material; S4. Based on the catalyst gene fingerprint, a multidimensional index of high-throughput experimental data is constructed. The multidimensional index includes a multi-level index structure, which is used to establish a mapping relationship between the catalyst gene fingerprint and the corresponding high-throughput experimental data.
2. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 1, characterized in that: Catalyst gene fingerprint is defined as the combination of composition characteristics, experimental reaction condition characteristics and thermal response characteristics, and its definition formula is: , Where, is the genetic fingerprint of the predetermined catalytic material C, is the composition characteristic of the predetermined catalytic material C, is the experimental reaction condition characteristic of the predetermined catalytic material C, is the thermal response characteristic of the predetermined catalytic material C.
3. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 2, characterized in that: The characterization of composition characteristics, experimental reaction condition characteristics and thermal response characteristics are as follows: , , , Where, E is the element composition vector of the catalyst; L is the carrier characteristic vector; P is the catalyst preparation parameter vector; A is the additive component vector; is the reaction temperature, is the reaction pressure, is the gas space velocity, is the feed composition vector, is the flow parameter; is the ignition temperature; is the temperature peak eigenvector; is the temperature change rate eigenvector; is the spatial temperature distribution characteristic; is the feature descriptor vector of the temperature-time curve.
4. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 3, characterized in that: Ignition temperature Determined by the normalized dynamic ignition criterion NDIC, which comprehensively considers the temperature rise rate and the degree of normalization of the current temperature relative to the reference baseline and the characteristic activation temperature; When the calculated normalized dynamic ignition criterion value continuously reaches the preset minimum confirmation point number and exceeds the preset judgment threshold, the ignition time is determined, and the temperature corresponding to the ignition time is used as the ignition temperature.
5. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 4, characterized in that: The calculation formula of normalized dynamic ignition criterion NDIC is as follows: , in, is the normalized dynamic ignition criterion value calculated at time i, is the average channel temperature at time i, is the average channel temperature at the previous moment, is the time interval between time i and time i-1, is the normalized temperature rise rate parameter, is the reference baseline temperature, is the characteristic activation temperature parameter, is the index adjustment factor.
6. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 1, characterized in that: The multi-level index structure includes: The top-level index is constructed based on the catalyst gene fingerprint CGF and adopts indexing technology based on space partitioning or graph embedding to achieve rapid clustering and retrieval of similar catalyst gene fingerprints; The middle-level index associates catalyst gene fingerprints with key derived features generated by high-throughput experiments and is organized using a multidimensional B-tree or R-tree structure; The underlying index associates key derived features with the storage location and access path of the original experimental data.
7. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 1, characterized in that: The method further comprises: Construct an associative index layer in the multi-dimensional index and use tensor decomposition technology to capture the nonlinear relationship between the features of each dimension in the catalyst gene fingerprint; Tensor decomposition represents the multidimensional features of catalyst gene fingerprints as a high-order tensor and decomposes it into a combination of a series of low-order feature vectors.
8. The method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering according to claim 7, characterized in that: The calculation formula for tensor decomposition is: , in, Represents an associated index, represents the tensor outer product operation, , , , They represent the characteristic vectors of the rth component in the four dimensions of composition, structure, performance and experimental conditions, and R represents the selected rank number.
9. The method for constructing a multidimensional index of high-throughput experimental data based on material genetic engineering according to claim 1, characterized in that: The method further comprises: For hybrid storage architectures that include relational databases and distributed object storage, an adaptive path selection mechanism based on a query cost model is introduced; When receiving a joint query request across storage systems, the query cost model evaluates the query costs of different index paths. The query cost model considers factors such as the average time to access the index and retrieve data, data transmission overhead, and the cost of cross-data source join operations. The optimal data access path is selected based on the evaluation results to execute the query.
10. The method for constructing a multi-dimensional index of high-throughput experimental data based on material genetic engineering according to claim 9, characterized in that: The calculation formula of the query cost model is: , in, is the total query cost, M is the number of data sources involved in the query, Index for accessing the jth data source and the average time to retrieve data, is the cost of transmitting data from the jth data source, and is the corresponding weight coefficient, K is the number of cross-data source connection operations involved in the query, is the cost of performing the k-th join operation, For its weight.
Citation Information
Patent Citations
Distributed index structure and load balancing method for high-throughput data flow
CN113568906A
Cited By
Catalyst prediction method and system based on molecular characterization contrast learning
CN121034442A