A biochemical information database extraction system and method

Through technologies such as multi-source interface units and adaptive convolutional neural networks, the problems of biochemical data are solved, efficient collection, integration, storage and query are achieved, and the efficiency and security of biochemical data processing are improved.

CN119917788BActive Publication Date: 2025-07-22CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510379371.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

Biochemical data sources are diverse and heterogeneous, and it is difficult for existing technology to efficiently collect, integrate, process and store, query optimization efficiency is low, privacy protection is insufficient, and version management is complex.

Method used

Multi-source interface units, adaptive convolutional neural networks, distributed storage management, intelligent query optimization and privacy protection modules are adopted, combined with graph structure dynamic priority scheduling, chaotic mapping encryption, reinforcement learning indexing and version management algorithms to achieve efficient data collection, integration, storage and secure query.

Benefits of technology

Improve data collection and integration efficiency, ensure data quality and security, improve query speed and accuracy, ensure data privacy, and support multi-branch version management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917788B_ABST
    Figure CN119917788B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of biochemical data processing, and discloses a biochemical information database extraction system and method. The system includes modules for biochemical data collection, preprocessing, feature extraction, distributed storage management, and intelligent query optimization, and may also include a privacy protection module and a version management module. The collection module integrates multi-source heterogeneous data; the preprocessing module detects anomalies and standardizes the data; the feature extraction module mines multi-modal features; the storage management module encrypts and stores the data and balances the load; the query optimization module improves the query efficiency. The privacy protection module ensures data security, and the version management module realizes data version traceability. The invention effectively solves the problems of biochemical data processing, improves the efficiency, security, and traceability of data processing, and provides strong support for fields such as biochemistry research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biochemical data processing, and in particular to a biochemical information database extraction system and method. Background Art

[0002] In the field of biochemistry, with the rapid development of experimental technology and information technology, biochemical data has shown explosive growth. These data come from a wide range of sources, covering multiple channels such as biochemical experimental equipment, public databases, and real-time sensors. The data from different data sources has significant heterogeneity, with huge differences in data format, structure, and semantics, which poses great challenges to the effective integration and utilization of data.

[0003] From the perspective of data acquisition, the data generated by biochemical experimental equipment is usually binary files or text files in a specific format, and the data structures and content specifications generated by different manufacturers and different models of equipment are different. Public databases, such as GenBank, PDB, etc., although provide a large amount of shared data for researchers, the data standards and access interfaces of these databases also have differences, making the data acquisition and integration process complex and cumbersome. Real-time sensors continuously generate a large amount of time-series data. How to efficiently collect and fuse these multi-source, heterogeneous, and time-series data has become the primary problem in biochemical information processing. The data preprocessing link also faces many problems. Due to factors such as the complexity of the experimental environment, equipment errors, and interference during data transmission, there are inevitably outliers in the original biochemical data. If these abnormal data are not processed in time, it will seriously affect the accuracy and reliability of subsequent data analysis. Traditional anomaly detection methods, such as detection algorithms based on fixed thresholds, cannot adapt to the dynamic change characteristics of biochemical data and are prone to missed detection or false detection. Moreover, the data from different data sources have differences in dimension, value range, etc. If no standardization processing is carried out, it will lead to a decline in the performance of subsequent data analysis models.

[0004] Feature extraction is a crucial step in biochemical data analysis. Biochemical data contains multiple modalities, such as text descriptions, spectral data, and molecular structure diagrams, etc. Each modality contains unique biochemical information. However, existing feature extraction methods often can only process single-modal data and are difficult to fully exploit the correlations and complementary information among multi-modal data. For complex biochemical data, traditional feature extraction algorithms are inefficient and cannot meet the requirements for data processing speed and accuracy in the big data era. In terms of data storage, with the continuous increase in the amount of biochemical data, traditional centralized storage methods face problems such as limited storage capacity, bottlenecks in data read and write performance, and single-point failures. Although distributed storage technology provides solutions, in a distributed storage environment, secure data storage and load balancing have become new challenges. How to perform efficient encrypted storage of biochemical data and ensure load balancing among distributed nodes to improve the overall performance and reliability of the storage system is an urgent problem to be solved.

[0005] Query optimization is an important part in the application of biochemical information databases. When conducting research, scientific researchers need to frequently query specific biochemical information from the database. Existing query methods are often inefficient and difficult to meet the needs of high-frequency queries. Moreover, users' query requirements are usually expressed in natural language. How to convert unstructured natural language queries into efficient structured query logic and optimize the query process is the key to improving the query efficiency of the database. In addition, biochemical data involves a large amount of sensitive information, such as personal health data, patent technology data, etc., and data privacy protection is of great importance. Traditional access control strategies and encryption technologies cannot meet the strict requirements for privacy protection of biochemical data. In the scenario of multi-user concurrent access, permission management and privacy protection face greater challenges. At the same time, with the in-depth research and data update, the management and traceability of data versions have become increasingly important, and existing version management methods have deficiencies in dealing with complex version dependencies and conflicts. Summary of the Invention

[0006] The purpose of the present invention is to provide a biochemical information database extraction system and method to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A biochemical information database extraction system, the system includes:

[0008] It includes a biochemical data collection module, a feature extraction module, a distributed storage management module, and an intelligent query optimization module, where:

[0009] The biochemical data acquisition module includes a multi-source interface unit and a heterogeneous data integration unit. The multi-source interface unit is used to access biochemical experimental equipment, public databases, and real-time sensor data sources. The heterogeneous data integration unit uses a dynamic priority scheduling algorithm based on a graph structure to dynamically weightedly fuse the timeliness and relevance of different data sources;

[0010] The feature extraction module includes an adaptive convolutional neural network unit. The adaptive convolutional neural network unit dynamically adjusts the convolutional kernel parameters according to the dimension of the input data and uses a sparse attention mechanism to extract key biochemical features;

[0011] The distributed storage management module includes a sharding encryption unit and a dynamic load balancing unit. The sharding encryption unit uses a lightweight sharding encryption algorithm based on chaotic mapping to encrypt the feature data in blocks and store it in distributed nodes;

[0012] The intelligent query optimization module includes a reinforcement learning index unit and a semantic parsing unit. The reinforcement learning index unit dynamically optimizes the multi-level index structure through the deep deterministic policy gradient algorithm.

[0013] Preferably, the system further includes a data preprocessing module, and the data preprocessing module includes an anomaly detection unit and a data normalization unit;

[0014] The anomaly detection unit is based on a dynamic threshold cleaning algorithm, and statistically analyzes the local data distribution through a sliding window to generate an adaptive cleaning threshold; further includes:

[0015] A local outlier factor calculation sub-unit, which is used to calculate the local outlier factor based on the KL divergence of the data distribution within the sliding window. The calculation formula is:

[0016]

[0017] where, is the data distribution within the sliding window, is the historical benchmark distribution, is the discretized interval index in the data distribution;

[0018] A dynamic threshold generation sub-unit, which is used to generate an adaptive cleaning threshold according to the outlier factor and the historical error distribution, and mark the abnormal data segments;

[0019] The data normalization unit is used to perform normalization processing on the data after anomaly detection.

[0020] Preferably, the feature extraction module further includes a multimodal fusion unit, which uses a joint embedding algorithm based on tensor decomposition to map text descriptions, spectral data, and molecular structure diagrams to a unified feature space and eliminates modal redundancy through low-rank constraints.

[0021] Preferably, the dynamic load balancing unit of the distributed storage management module includes:

[0022] A node status monitoring subunit for collecting the storage load and network latency of distributed nodes in real time;

[0023] A shard migration decision subunit that dynamically adjusts the distribution of data shards using a Nash equilibrium strategy based on game theory, and its utility function is:

[0024]

[0025] where is the utility value of node under the strategy combination, is the current shard storage strategy of node , is the set of shard storage strategies of other nodes, is the load coefficient of node , is the network latency coefficient, , are weight factors.

[0026] Preferably, the semantic parsing unit of the intelligent query optimization module includes:

[0027] A natural language processing subunit for converting unstructured query statements into structured query logic;

[0028] A syntax tree optimization subunit that generates a syntax tree with minimized query cost by pruning redundant nodes and merging similar paths.

[0029] Preferably, the system further includes a privacy protection module, which includes:

[0030] A differential noise injection unit for adding Laplace noise to sensitive fields during the data preprocessing stage;

[0031] An access control unit that adopts a dynamic permission allocation strategy based on attribute-based encryption to generate fine-grained access tokens according to user roles and query contexts.

[0032] Preferably, the access control unit of the privacy protection module further includes:

[0033] A policy conflict detection subunit for identifying and resolving permission policy conflicts during multi-user concurrent access;

[0034] The token dynamic update subunit automatically refreshes the validity period of the access token according to the time decay function and the access frequency. The time decay function is:

[0035]

[0036] Wherein, is the remaining value of the validity period of the access token at time , is the initial validity period, is the time decay factor, is the time the token has been used.

[0037] Preferably, the system further includes a version management module, and the version management module includes:

[0038] The data snapshot generation unit generates a version snapshot by using a differential compression algorithm based on incremental hashing;

[0039] The version backtracking unit records the version dependency relationship through a directed acyclic graph and supports multi-branch backtracking operations.

[0040] Preferably, the version backtracking unit of the version management module further includes:

[0041] The conflict merging subunit solves version conflicts by using a three-way merge algorithm based on operation transformation;

[0042] The metadata verification subunit verifies the integrity and consistency of the version snapshot through a Merkle tree, and its hash calculation is:

[0043]

[0044] Wherein, is the hash value of the parent node, , are the hash values of the left and right child nodes respectively, is the cryptographic hash function, represents the concatenation operation of the hash values.

[0045] Preferably, the present invention further includes a method for extracting a biochemical information database, which is applied to the above-mentioned biochemical information database extraction system. The method includes:

[0046] Collect heterogeneous biochemical data through a multi-source interface unit, and perform data fusion by using a graph structure dynamic priority scheduling algorithm;

[0047] Perform anomaly detection and standardization processing on the original data based on a dynamic threshold cleaning algorithm;

[0048] Extract multi-modal features using an adaptive convolutional neural network and complete feature fusion through a tensor decomposition algorithm;

[0049] Store the feature data in distributed nodes using a chaotic mapping piecewise encryption algorithm and dynamically adjust the piecewise distribution based on the Nash equilibrium strategy;

[0050] Optimize high-frequency queries through reinforcement learning indexing and generate a syntax tree with minimized cost by combining semantic parsing.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] The biochemical data acquisition module of the present invention accesses multiple data sources through a multi-source interface unit. The heterogeneous data integration unit uses a dynamic priority scheduling algorithm based on a graph structure, fully considering the timeliness and relevance of the data for dynamic weighted fusion. This enables the system to quickly and accurately collect and integrate multi-source heterogeneous biochemical data. Compared with traditional methods, it greatly improves the efficiency of data collection and the quality of fusion, providing a more comprehensive and accurate data basis for subsequent data analysis. For example, when integrating data from different experimental devices and public databases, it can better retain the internal connections between the data and avoid information loss.

[0053] The anomaly detection unit in the data preprocessing module is based on a dynamic threshold cleaning algorithm. It generates an adaptive cleaning threshold by statistically analyzing the local data distribution through a sliding window, and can accurately detect abnormal data, effectively overcoming the defects of traditional fixed threshold detection methods and reducing the cases of missed detection and false detection. The data standardization unit further standardizes the data to ensure the consistency and comparability of the data, providing high-quality data for subsequent feature extraction and analysis, and improving the accuracy and reliability of the analysis results. The adaptive convolutional neural network unit in the feature extraction module can dynamically adjust the convolution kernel parameters according to the input data dimension and extract key biochemical features using a sparse attention mechanism, significantly improving the efficiency and accuracy of feature extraction. The multi-modal fusion unit adopts a joint embedding algorithm based on tensor decomposition to map multiple modal data to a unified feature space and eliminate modal redundancy, fully mining the complementary information of multi-modal data, providing richer and more representative features for biochemical data analysis, and helping to improve the performance of subsequent data analysis models.

[0054] The sharding encryption unit of the distributed storage management module adopts a lightweight sharding encryption algorithm based on chaotic mapping to encrypt and store the feature data in blocks, effectively ensuring the security of the data. The dynamic load balancing unit dynamically adjusts the distribution of data shards by real-time monitoring of node status and using the Nash equilibrium strategy based on game theory, avoiding the problem of uneven node load, improving the overall performance and reliability of the storage system, and ensuring the efficient storage and fast access of data. The reinforcement learning indexing unit of the intelligent query optimization module dynamically optimizes the multi-level index structure through the deep deterministic policy gradient algorithm, which can well adapt to high-frequency query scenarios and significantly improve the query speed. The semantic parsing unit converts unstructured query statements into structured query logic and generates a scheme to minimize the query cost by optimizing the syntax tree, further improving the query efficiency and accuracy, and facilitating researchers to quickly obtain the required biochemical information.

[0055] The differential noise injection unit of the privacy protection module adds Laplace noise to sensitive fields during the data preprocessing stage, effectively protecting the privacy of the data. The access control unit adopts a dynamic permission allocation strategy based on attribute-based encryption, generates fine-grained access tokens by combining user roles and query contexts, can detect and resolve permission policy conflicts, and automatically refreshes the token validity period according to the time decay function and access frequency, comprehensively ensuring the secure access of data and preventing the leakage of sensitive information. The data snapshot generation unit of the version management module generates version snapshots using a differential compression algorithm based on incremental hashing, greatly saving storage space. The version backtracking unit records the version dependency relationship through a directed acyclic graph, supports multi-branch backtracking operations, and uses a three-way merge algorithm based on operation transformation to resolve version conflicts, and verifies the integrity and consistency of the version snapshot through a Merkle tree, facilitating researchers to trace and manage the historical versions of the data and ensuring the traceability and integrity of the data. Brief Description of the Drawings

[0056] Figure 1 It is the working principle diagram of the biochemical information database extraction system described in the present invention;

[0057] Figure 2 It is the working flow chart of the anomaly detection unit;

[0058] Figure 3 It is the working flow chart of the dynamic load balancing unit;

[0059] Figure 4 It is the working flow chart of the access control unit. Detailed Embodiments

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] Please refer to Figures 1-4 , the present invention provides a technical solution: a biochemical information database extraction system, and the system includes:

[0062] Biochemical data acquisition module: Access biochemical experimental equipment, public databases, and real-time sensor data sources through a multi-source interface unit to obtain various types of raw data in the biochemical field. The heterogeneous data integration unit adopts a dynamic priority scheduling algorithm based on a graph structure, and performs dynamic weighted fusion according to the timeliness and relevance of data from different data sources, so that the fused data can more accurately reflect the comprehensive characteristics of biochemical information.

[0063] Data preprocessing module: The anomaly detection unit is based on a dynamic threshold cleaning algorithm, uses a sliding window to statistically analyze the local data distribution and generates an adaptive cleaning threshold to identify and process abnormal data. The data normalization unit normalizes the data after anomaly detection to ensure the consistency and comparability of the data in subsequent analysis and processing.

[0064] Feature extraction module: The adaptive convolutional neural network unit dynamically adjusts the convolutional kernel parameters according to the dimension of the input data, and adopts a sparse attention mechanism to extract key biochemical features, improving the accuracy and efficiency of feature extraction. The multi-modal fusion unit adopts a joint embedding algorithm based on tensor decomposition to map multi-modal data such as text descriptions, spectral data, and molecular structure diagrams to a unified feature space, and eliminates modal redundancy through low-rank constraints to achieve effective fusion of multi-modal data.

[0065] Distributed storage management module: The sharding encryption unit adopts a lightweight sharding encryption algorithm based on chaotic mapping to encrypt the feature data in blocks and store it in distributed nodes to ensure the security of data storage. The dynamic load balancing unit real-time collects the storage load and network latency of distributed nodes, and adopts a Nash equilibrium strategy based on game theory to dynamically adjust the data sharding distribution to ensure the stability and efficiency of the overall system performance.

[0066] Intelligent query optimization module: The reinforcement learning index unit dynamically optimizes the multi-level index structure through a deep deterministic policy gradient algorithm to adapt to high-frequency query scenarios and speed up query speed. The semantic parsing unit converts unstructured query statements into structured query logic, and generates a syntax tree with minimized query cost by pruning redundant nodes and merging similar paths, further improving query efficiency and accuracy.

[0067] The present invention will be further described below in conjunction with Embodiments 1 to 5:

[0068] Embodiment 1:

[0069] This embodiment mainly elaborates on the working process of the anomaly detection unit in the data preprocessing module. In the process of biochemical data processing, abnormal data will seriously affect the accuracy of subsequent analysis results. Therefore, it is crucial to accurately detect and process abnormal data.

[0070] The local outlier factor calculation subunit of the anomaly detection unit calculates the local outlier factor based on the KL divergence of the data distribution within the sliding window. Specifically, in the real-time process of data collection and processing, a sliding window of appropriate size is set, and as data continuously flows in, the sliding window gradually moves along the data sequence. At each window position, the data distribution within the window is denoted as , and the historical reference distribution is denoted as . The discretized interval index in the data distribution is , and the KL divergence is calculated through the formula to obtain the local outlier factor. This factor can quantify the degree of difference between the data distribution within the current window and the historical reference distribution. The greater the difference, the more likely the data within the window is abnormal data.

[0071] The dynamic threshold generation subunit generates an adaptive cleaning threshold based on the calculated outlier factor and the historical error distribution. In actual operation, the historical error distribution can be obtained through statistical analysis of a large amount of historical data. Combining the outlier factor with the historical error distribution can more accurately determine a dynamically changing threshold. When the outlier factor of the data within the window exceeds this threshold, the dynamic threshold generation subunit marks this data segment as abnormal data, and the subsequent data preprocessing module will perform corresponding processing on these marked abnormal data, such as elimination, correction, etc., to ensure the data quality entering the subsequent processing link.

[0072] Embodiment 2:

[0073] In the feature extraction module, the multimodal fusion unit uses a joint embedding algorithm based on tensor decomposition to process multimodal data. Taking common text descriptions, spectral data, and molecular structure diagrams in biochemical experiments as examples, these data contain biochemical information from different angles. The multimodal fusion unit maps these data to a unified feature space, and in the mapping process, modal redundancy is eliminated through low-rank constraints. For example, when processing spectral data and molecular structure diagrams, there may be duplicate expressions of some information, and low-rank constraints can effectively remove these redundant parts, making the fused features more refined and representative, facilitating subsequent data analysis and model training.

[0074] In the dynamic load balancing unit of the distributed storage management module, the node status monitoring subunit collects the storage load and network latency of distributed nodes in real time. Through specialized monitoring programs and technical means, the storage usage and network transmission latency data of each node are obtained regularly. These data are fed back into the system in real time to provide a basis for subsequent decision-making.

[0075] The shard migration decision subunit dynamically adjusts the data shard distribution using the Nash equilibrium strategy based on game theory. Its utility function is , where is the utility value of node under the strategy combination, is the current shard storage strategy of node , is the set of shard storage strategies of other nodes, is the load coefficient of node , is the network latency coefficient, , are weight factors. In practical applications, according to the requirements and characteristics of the system, the weight factors and are set reasonably. For example, when the system is more sensitive to storage load, the value of can be appropriately increased; if more attention is paid to network latency, the weight of is increased. By continuously calculating the utility values of each node and according to the Nash equilibrium strategy, the distribution of data shards among different nodes is dynamically adjusted to balance the storage load and network latency of the system and improve the overall performance of the distributed storage system.

[0076] Example 3:

[0077] This example details the specific working process of the semantic parsing unit in the intelligent query optimization module. The semantic parsing unit plays a key role in improving the accuracy and efficiency of queries.

[0078] The natural language processing sub-unit is responsible for converting unstructured query statements into structured query logic. In actual application scenarios, users may input query requirements in various natural language forms, such as "Find the relevant experimental data of a certain protein with specific functions". The natural language processing sub-unit uses natural language processing techniques to perform lexical analysis, syntactic analysis, and semantic understanding on these query statements. Through lexical analysis, the sentence is split into individual words or phrases and their parts of speech are marked; syntactic analysis constructs the grammatical structure of the sentence; semantic understanding further clarifies the semantic relationships between various words and sentence fragments. After these processing steps, the unstructured natural language query is converted into structured query logic that the system can understand and process, such as being converted into a specific query statement format, clarifying information such as the target data type and conditions of the query.

[0079] The syntax tree optimization sub-unit generates a syntax tree with minimized query cost by pruning redundant nodes and merging similar paths. After generating the initial syntax tree, the syntax tree is analyzed. When querying "the activity data of a certain specific protein under different temperature conditions", there may be some redundant nodes in the syntax tree, which may be caused by repeated expressions or unnecessary logical branches in the query statement. The syntax tree optimization sub-unit identifies and prunes these redundant nodes. At the same time, for similar query paths, such as paths that point to the same target data under different condition combinations, they are merged. Through such optimization operations, the resulting syntax tree can execute the query task at the lowest cost, reduce the consumption of computing resources during the query process, and improve query efficiency.

[0080] Example 4:

[0081] This example focuses on introducing the working principles and specific implementation methods of the privacy protection module and its included differential noise injection unit and access control unit. This module is of great significance for ensuring the security and privacy of biochemical data.

[0082] In the data preprocessing stage, the differential noise injection unit adds Laplace noise to sensitive fields. Biochemical data may contain some sensitive information, such as experimental sample information involving personal privacy, specific experimental technical details, etc. To protect this sensitive information, the differential noise injection unit adds Laplace noise to sensitive fields when the data preprocessing module processes the data. The characteristics of Laplace noise enable it to hide the true value of the original data to a certain extent while maintaining the overall statistical characteristics of the data. For example, for a sensitive numerical value representing the concentration of an experimental sample, after adding Laplace noise to it, it is difficult for an attacker to directly obtain the accurate original concentration value from the processed data, thus protecting the privacy of the data.

[0083] The access control unit adopts a dynamic permission allocation policy based on attribute-based encryption and generates fine-grained access tokens according to user roles and query contexts. In an actual biochemical information database system, different users have different roles, such as researchers, administrators, students, etc., and their access requirements and permissions for data vary. The access control unit first determines the initial access permission range according to the user's role. For example, researchers may have higher data access permissions and can view and download detailed experimental data in specific fields; while students may only be able to access some publicly available data or processed teaching data. At the same time, combined with query contexts, such as factors like the query time, query frequency, and specific data content of the query, the permissions are further adjusted dynamically.

[0084] The access control unit also includes a policy conflict detection subunit and a token dynamic update subunit. The policy conflict detection subunit is used to identify and resolve permission policy conflicts during multi-user concurrent access. In the case of multiple users accessing the database simultaneously, there may be situations where the permission policies of different users conflict with each other. For example, when one user attempts to modify a certain piece of data while another user has read-only permission for the same data at the same time, a conflict occurs. The policy conflict detection subunit monitors these concurrent access requests in real-time, discovers and resolves conflicts in a timely manner through the analysis and comparison of permission policies, and takes corresponding solutions, such as coordinating according to factors like user priority and the urgency of access requests.

[0085] The token dynamic update subunit automatically refreshes the validity period of the access token according to the time decay function and access frequency. The time decay function is , where is the remaining value of the validity period of the access token at time , is the initial validity period, is the time decay factor, is the time the token has been used. As time goes by, the access requirements and permissions of users may change. The token dynamic update subunit dynamically adjusts the validity period of the access token according to this time decay function, combined with the user's access frequency. If a user has not performed access operations for a long time or has a low access frequency, the token dynamic update subunit will appropriately shorten the validity period of their access token; conversely, if the user accesses frequently and within their permission range, the validity period will be appropriately extended to ensure the convenience of user access and the security of data.

[0086] Example 5:

[0087] This example elaborates in detail the working methods of the version management module and its included data snapshot generation unit and version backtracking unit. The version management module is crucial for data traceability and historical record management.

[0088] The data snapshot generation unit generates version snapshots using a differential compression algorithm based on incremental hashing. During the continuous update and change of biochemical data, to record the different version states of the data, the data snapshot generation unit generates version snapshots regularly or when specific events occur (such as important experimental data updates, system configuration changes, etc.). The differential compression algorithm based on incremental hashing can efficiently capture the changed parts of the data. Specifically, it identifies the state of the data by calculating the hash value of the data. When the data changes, only the changed parts are hashed and stored, rather than re-storing the entire data. This can greatly reduce the occupancy of storage space and at the same time speed up the generation speed of version snapshots. For example, in a database containing a large amount of biochemical experimental data, after each update of the experimental results, only the data parts different from the previous version are stored through this algorithm, effectively reducing the storage cost.

[0089] The version backtracking unit records version dependencies through a directed acyclic graph and supports multi-branch backtracking operations. In practical applications, the version evolution of biochemical data may have multiple branch situations, such as data branches under different experimental conditions, data branches in different research directions, etc. The directed acyclic graph can clearly record the dependency relationships between each version. Each version serves as a node in the graph, and the update relationship between versions serves as a directed edge. In this way, the system can accurately trace back to any historical version.

[0090] The version backtracking unit also includes a conflict merging subunit and a metadata verification subunit. The conflict merging subunit uses a three-way merge algorithm based on operation transformation to solve version conflicts. During the process of multi-user collaboration or the merging of data from different branches, version conflicts may occur. For example, two users modify the same data simultaneously. The conflict merging subunit uses the three-way merge algorithm based on operation transformation to compare and merge the current version, the target version, and the common ancestor version. By analyzing the operation differences between different versions and performing merge operations according to certain rules, it tries to retain the valid modifications of all parties and solve the version conflict problem.

[0091] The metadata verification subunit verifies the integrity and consistency of the version snapshot through a Merkle tree, and its hash calculation is , where is the hash value of the parent node, , are the hash values of the left and right child nodes respectively, is a cryptographic hash function, Represents the concatenation operation of hash values. The Merkle tree hashes data blocks layer by layer to form a tree structure. During version backtracking and data verification, by calculating the hash values of each node and comparing them with the stored hash values, it is possible to quickly verify whether the data of the version snapshot is complete and whether it has been tampered with. If a hash value inconsistency is found, it indicates that there may be a problem with the data, and further investigation and repair are required to ensure the reliability and integrity of the data.

[0092] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0093] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A biochemical information database extraction system, characterized in that, It includes a biochemical data acquisition module, a feature extraction module, a distributed storage management module, and an intelligent query optimization module, where: The biochemical data acquisition module includes a multi-source interface unit and a heterogeneous data integration unit. The multi-source interface unit is used to access biochemical experimental equipment, public databases, and real-time sensor data sources. The heterogeneous data integration unit uses a dynamic priority scheduling algorithm based on a graph structure to perform dynamic weighted fusion on the timeliness and relevance of different data sources; The feature extraction module includes an adaptive convolutional neural network unit. The adaptive convolutional neural network unit dynamically adjusts the convolutional kernel parameters according to the dimension of the input data and uses a sparse attention mechanism to extract key biochemical features; The distributed storage management module includes a sharding encryption unit and a dynamic load balancing unit. The sharding encryption unit uses a lightweight sharding encryption algorithm based on chaotic mapping to encrypt the feature data in blocks and store it in distributed nodes; The intelligent query optimization module includes a reinforcement learning index unit and a semantic parsing unit. The reinforcement learning index unit dynamically optimizes the multi-level index structure through the deep deterministic policy gradient algorithm; The dynamic load balancing unit of the distributed storage management module includes: A node status monitoring subunit, which is used to collect the storage load and network latency of distributed nodes in real time; A sharding migration decision subunit, which uses a Nash equilibrium strategy based on game theory to dynamically adjust the data sharding distribution, and its utility function is: ; Among them, is the utility value of the node under the strategy combination, is the current shard storage strategy of the node , is the set of shard storage strategies of other nodes, is the load factor of the node , is the network delay coefficient, , are weight factors.

2. The biochemical information database extraction system according to claim 1, characterized in that, The system also includes a data preprocessing module, and the data preprocessing module includes an anomaly detection unit and a data normalization unit; The anomaly detection unit is based on a dynamic threshold cleaning algorithm, and statistically analyzes the local data distribution through a sliding window to generate an adaptive cleaning threshold; it further includes: A local outlier factor calculation subunit, which is used to calculate the local outlier factor based on the KL divergence of the data distribution within the sliding window, and the calculation formula is: ; Among them, is the data distribution within the sliding window, is the historical reference distribution, is the discretized interval index in the data distribution; A dynamic threshold generation subunit, which is used to generate an adaptive cleaning threshold according to the outlier factor and the historical error distribution, and mark abnormal data segments; The data normalization unit is used to perform normalization processing on the data after anomaly detection.

3. The biochemical information database extraction system according to any one of claims 1 to 2, characterized in that The feature extraction module also includes a multi-modal fusion unit. The multi-modal fusion unit uses a joint embedding algorithm based on tensor decomposition to map text descriptions, spectral data, and molecular structure diagrams to a unified feature space, and eliminates modal redundancy through low-rank constraints.

4. The biochemical information database extraction system according to claim 1, characterized in that, The semantic parsing unit of the intelligent query optimization module includes: A natural language processing subunit, which is used to convert unstructured query statements into structured query logic; A syntax tree optimization subunit, which generates a syntax tree with minimized query cost by pruning redundant nodes and merging similar paths.

5. The biochemical information database extraction system according to claim 1, characterized in that, It also includes a privacy protection module, and the privacy protection module includes: A differential noise injection unit, which is used to add Laplace noise to sensitive fields during the data preprocessing stage; An access control unit, which uses a dynamic permission allocation strategy based on attribute-based encryption to generate fine-grained access tokens according to user roles and query contexts.

6. The biochemical information database extraction system according to claim 5, characterized in that, The access control unit of the privacy protection module further includes: A policy conflict detection subunit, configured to identify and resolve permission policy conflicts during multi-user concurrent access; A token dynamic update subunit, which automatically refreshes the validity period of an access token according to a time decay function and an access frequency. The time decay function is: ; Among them, is the remaining value of the validity period of the access token at time , is the initial validity period, is the time decay factor, is the time the token has been used.

7. The biochemical information database extraction system according to claim 1, wherein, It further includes a version management module, and the version management module includes: A data snapshot generation unit, which generates a version snapshot by using a differential compression algorithm based on incremental hashing; A version backtracking unit, which records version dependency relationships through a directed acyclic graph and supports multi-branch backtracking operations.

8. The biochemical information database extraction system according to claim 7, characterized in that, The version backtracking unit of the version management module further includes: A conflict merging subunit, which uses a three-way merge algorithm based on operation transformation to resolve version conflicts; A metadata verification subunit, which verifies the integrity and consistency of the version snapshot through a Merkle tree. Its hash calculation is: ; Among them, is the hash value of the parent node, , are the hash values of the left and right child nodes respectively, is the password hash function, represents the concatenation operation of hash values.

9. A method for extracting biochemical information from a database, characterized in that, Applied to the system according to any one of claims 1 to 8, including: Collecting heterogeneous biochemical data through a multi-source interface unit, and performing data fusion by using a graph-structured dynamic priority scheduling algorithm; Performing anomaly detection and normalization processing on the original data based on a dynamic threshold cleaning algorithm; Extracting multi-modal features by using an adaptive convolutional neural network, and completing feature fusion through a tensor decomposition algorithm; Storing the feature data in distributed nodes by using a chaotic mapping sharding encryption algorithm, and dynamically adjusting the sharding distribution based on a Nash equilibrium strategy; Optimizing high-frequency queries through reinforcement learning indexing, and generating a syntax tree with minimized cost in combination with semantic parsing.

Citation Information

Patent Citations

  • Data interaction system for archive management

    CN117827743A