Pulmonary fibrosis drug target screening method fusing multiple omics data

By integrating multi-omics data and computational models based on biological networks, the limitations of single-omics data in drug target screening for pulmonary fibrosis have been overcome, enabling efficient and accurate target screening, shortening research and development time, and improving treatment efficacy.

CN121768482APending Publication Date: 2026-03-31THE SIXTH MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies for screening drug targets in pulmonary fibrosis have limitations due to the use of single-mathematical data analysis, resulting in biased and highly erroneous screening results. This makes it difficult to quickly and accurately identify effective drug targets, thus affecting the drug development process and treatment efficacy.

Method used

This study employs a multi-omics data fusion approach to acquire biological information from multiple levels, including genomics, transcriptomics, and proteomics. By integrating multi-omics data through data fusion technology, and using a computational model based on biological networks, the study analyzes the correlation between potential targets and the pathological process of pulmonary fibrosis, screens out high-priority drug targets, and stores them persistently.

Benefits of technology

It has improved the accuracy and efficiency of drug target screening, shortened the research and development time, provided a more comprehensive direction for drug development, laid a solid foundation for the development of drugs for the treatment of pulmonary fibrosis, and improved the treatment effect and the quality of life of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768482A_ABST
    Figure CN121768482A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological medicine, and discloses a pulmonary fibrosis drug target screening method fusing multi-omics data. The method comprises the following steps: acquiring multi-omics original data related to pulmonary fibrosis from a public database and an experimental data source; performing quality control and normalization processing on the data to generate a standardized multi-omics data set; integrating the data set into a unified multi-omics feature expression spectrum through a data fusion technology; analyzing the expression profile by using a calculation model based on a biological network, and deducing the correlation degree between each potential target spot and the pulmonary fibrosis pathological process; sorting the potential target spots according to the correlation degree, and screening out high-priority drug target spots; and importing the high-priority drug targets into a drug target management system for persistent storage. According to the system, multiple omics data are integrated through the system, the relevance between the targets and diseases is quantified, accurate and efficient screening of the pulmonary fibrosis drug targets is achieved, and reliable data support is provided for research and development of new drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a method for screening drug targets for pulmonary fibrosis that integrates multi-omics data. Background Technology

[0002] Pulmonary fibrosis is a lung disease characterized by fibroblast proliferation and massive accumulation of extracellular matrix. Abnormal repair of damaged alveolar tissue leads to structural abnormalities, resembling scarring in the lungs. The cause is unknown in most cases. Idiopathic pulmonary fibrosis is the most common type, a severe interstitial lung disease that leads to progressive loss of lung function.

[0003] The incidence and mortality rates of pulmonary fibrosis are increasing year by year. The average survival time after diagnosis of idiopathic pulmonary fibrosis is only 2.8 years, and the mortality rate is higher than that of most cancers, hence it is called a "tumor-like disease." It has a serious impact on patients' health and life. Patients will experience symptoms such as dry cough and progressive dyspnea. As the disease progresses and lung damage worsens, respiratory function deteriorates, leading to decreased exercise capacity, limited activity, and even simple daily activities such as dressing and washing can cause shortness of breath, severely reducing quality of life. It can also lead to mental health problems such as anxiety and depression, as well as a series of complications such as pulmonary heart disease, pulmonary hypertension, and lung infections.

[0004] Treatment for pulmonary fibrosis mainly includes medication and surgery. Commonly used medications include glucocorticoids, which suppress inflammation, but long-term use can lead to numerous side effects such as osteoporosis, weakened immunity, elevated blood sugar, and obesity. Antifibrotic drugs such as pirfenidone and nintedanib can slow the progression of pulmonary fibrosis to some extent, but they cannot completely cure the disease, and some patients have poor tolerance to these drugs, experiencing adverse reactions such as gastrointestinal discomfort. Surgical treatment primarily involves lung transplantation, an effective treatment for end-stage pulmonary fibrosis. However, it faces challenges such as donor shortages, high surgical risks, and postoperative immune rejection. Many patients experience a deterioration of their condition while waiting for a donor, and even after a successful lung transplant, long-term use of immunosuppressants is required, increasing the risk of infections and other complications.

[0005] Drug target screening is a crucial step in developing effective therapeutic drugs. Accurately identifying drug targets allows for more targeted drug development, enabling researchers to focus on specific molecular or biological processes, thereby increasing the success rate of drug development and saving time and costs. For complex and serious diseases like pulmonary fibrosis, finding key drug targets holds promise for developing drugs that can truly reverse or halt disease progression, bringing hope for overcoming the challenge of pulmonary fibrosis and fundamentally improving patients' prognosis and quality of life. Summary of the Invention

[0006] The purpose of this invention is to provide a method for screening drug targets for pulmonary fibrosis by integrating multi-omics data, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides a method for screening drug targets for pulmonary fibrosis by integrating multi-omics data, the method comprising:

[0008] Raw multi-omics data related to pulmonary fibrosis were collected from public databases and experimental data sources;

[0009] The collected multi-omics raw data are subjected to quality control and normalization processing to generate standardized multi-omics datasets;

[0010] Standardized multi-omics datasets are integrated into a unified multi-omics feature representation spectrum through data fusion technology;

[0011] We used a computational model based on biological networks to analyze the expression profiles of multi-omics features and deduce the degree of association between each potential target and the pathological process of pulmonary fibrosis.

[0012] Potential targets are sorted according to their relevance, and high-priority drug targets are selected. These high-priority drug targets are then imported into the drug target management system for persistent storage.

[0013] Preferably, the collection of multi-omics raw data related to pulmonary fibrosis includes the following steps: obtaining single-cell RNA sequencing data from pulmonary fibrosis patient tissues from a gene expression comprehensive database; collecting mass spectrometry proteomic data of pulmonary fibrosis-related models from a protein database; extracting whole-genome methylation data covering different disease stages of pulmonary fibrosis from a genomics data warehouse; and aligning the above three types of data with spatiotemporal labels to form a multi-omics raw data set with temporal and spatial labels.

[0014] Preferably, the quality control and normalization processing of the collected multi-omics raw data includes the following steps: using an automated quality scoring algorithm to evaluate the signal-to-noise ratio of each data sample; using adaptive threshold filtering technology to remove abnormal samples with scores below a preset threshold; applying quantile normalization method to correct the expression distribution of the retained samples; using a sliding window algorithm to eliminate batch effects on the corrected data; and finally generating a standardized multi-omics dataset that conforms to Gaussian distribution characteristics.

[0015] Preferably, the process of integrating standardized multi-omics datasets into a unified multi-omics feature expression spectrum through data fusion technology includes the following steps: constructing a multimodal data fusion framework based on tensor decomposition; mapping different omics data to a unified multidimensional feature space; aligning the feature space using alternating least squares; eliminating intermodal conflicts through a multi-source feature cross-validation algorithm; and generating a multi-omics feature expression spectrum with complete biological significance.

[0016] Preferably, the analysis of multi-omics feature expression profiles using a computational model based on biological networks includes the following steps: establishing a multi-level biological network model specific to pulmonary fibrosis; projecting the multi-omics feature expression profiles onto network nodes; using a random walk algorithm to calculate the connectivity strength between each potential target and the core pathological module; determining the functional contribution of the target in different biological processes through modular analysis; and finally outputting a quantitative correlation index.

[0017] Preferably, the derivation of the association degree between each potential target and the pathological process of pulmonary fibrosis includes the following steps: constructing an association degree synthesis algorithm based on weighted geometric mean; nonlinearly combining connectivity strength and functional contribution; evaluating the statistical significance of association degree using Monte Carlo simulation; iteratively optimizing the association degree estimate through a Bayesian update mechanism; and generating the final standardized association degree score.

[0018] Preferably, the process of ranking potential targets based on their correlation includes the following steps: establishing a dynamic priority queue management mechanism; analyzing the changing trend of correlation using a sliding time window; determining target priority groups using a percentile ranking algorithm; adjusting the final ranking results by combining clinical importance weights; and generating a time-sensitive target priority list.

[0019] Preferably, after screening high-priority drug targets, the method further includes a validation step: designing a target stability testing scheme based on cross-validation; conducting external validation using an independent dataset; evaluating the expression specificity of the target at different stages of the disease using subtype-specific analysis; and finally generating a target credibility report containing validation indicators.

[0020] Preferably, the process of importing high-priority drug targets into a drug target management system for persistent storage includes the following steps: constructing a target storage architecture based on a graph database; designing a multi-dimensional indexing mechanism to support fast retrieval; establishing a version control system to manage the update history of target data; realizing the mapping of relationships between targets and related biological entities; and completing the persistent storage of target data.

[0021] Preferably, the method further includes a data auditing step: deploying an audit trail system based on blockchain technology; setting up smart contracts to automatically perform data integrity checks; periodically generating target data change reports; verifying data consistency through encrypted hash values; and synchronizing the audit results to a distributed data management platform.

[0022] Compared with the prior art, the beneficial effects of the present invention are:

[0023] This patent employs a multi-omics data integration method that overcomes the limitations of traditional single-omics data, acquiring biological information from multiple levels, including the genome, transcriptome, and proteome. At the genomic level, it reveals differences and variations in gene sequences; this genetic information forms the intrinsic basis of disease development, and many gene mutations related to pulmonary fibrosis may directly or indirectly affect disease progression. At the transcriptome level, it reveals dynamic changes in gene expression, identifying which genes are upregulated or downregulated during pulmonary fibrosis, thus clarifying the regulatory mechanisms at the gene transcription level and revealing the characteristics of gene expression during disease development. At the proteome level, it focuses on protein expression, modification, and protein-protein interactions. As direct executors of life activities, protein changes more directly reflect cellular functional states and pathological changes in disease. By integrating these different omics data, it's like drawing a complete pathological atlas of pulmonary fibrosis from multiple perspectives, comprehensively reflecting the pathological process of pulmonary fibrosis, compensating for the limitations of traditional single-omics data, making the screened targets more representative and reliable, and providing a more comprehensive and accurate direction for subsequent drug development.

[0024] In this patented method, data fusion technology and a computational model based on biological networks play a crucial role. Data fusion technology integrates standardized data from different omics, eliminating inconsistencies and redundancies to form a unified and coherent dataset. The computational model based on biological networks enables in-depth analysis of this multi-omics feature expression profile. This model considers the complex interactions between biomolecules, such as gene-gene, protein-protein, and gene-protein interactions, as well as the synergistic effects of these molecules in various biological pathways. By simulating and analyzing these complex biological networks, the correlation between potential targets and the pathological process of pulmonary fibrosis can be accurately deduced, deeply uncovering the biological information hidden behind massive amounts of data. Compared to traditional methods, this in-depth analysis can more accurately determine the role of each potential target in the occurrence and development of pulmonary fibrosis, reducing screening errors caused by one-sided analysis or simple correlation judgments, greatly improving the accuracy of target screening, and laying a solid foundation for developing targeted and effective drugs for the treatment of pulmonary fibrosis.

[0025] With the rapid development of biotechnology, biological data is exploding, and traditional drug target screening methods are struggling to handle this massive amount of data. This patented method, however, integrates multi-omics data to acquire multi-level biological information in one go, avoiding repetitive research and analysis of single-omics data. Simultaneously, it utilizes a computational model based on biological networks for analysis, leveraging the powerful computing capabilities and efficient algorithms of computers to rapidly process large amounts of biological data. This efficient data processing and analysis method can evaluate and screen numerous potential targets in a short time, significantly shortening the drug target screening time. In drug development, time cost is a critical factor; rapidly screening high-priority drug targets can accelerate the development of drugs for pulmonary fibrosis, enabling new drugs to enter clinical trials more quickly and bringing more hope for treatment to patients. Attached Figure Description

[0026] Figure 1 This is a schematic diagram illustrating the working principle of the pulmonary fibrosis drug target screening method that integrates multi-omics data as described in this invention.

[0027] Figure 2 A flowchart for the collection of raw data from multi-omics studies related to pulmonary fibrosis;

[0028] Figure 3 A flowchart for quality control and normalization processing of raw data from multi-omics studies;

[0029] Figure 4 This is a comprehensive analysis diagram of potential drug targets for pulmonary fibrosis. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Please see Figure 1This invention provides a method for screening drug targets for pulmonary fibrosis by integrating multi-omics data. The method begins by collecting raw multi-omics data related to pulmonary fibrosis from public databases and experimental data sources. This raw multi-omics data includes various types of data sources such as genomics, transcriptomics, and proteomics, providing molecular-level information about the disease state of pulmonary fibrosis. The collected raw multi-omics data undergoes systematic quality control and normalization. The quality control system uses algorithms to assess data integrity, and the normalization process adjusts the data scale to eliminate technical variations, generating a standardized multi-omics dataset. This standardized multi-omics dataset is integrated using data fusion technology, which maps different omics data to a common feature space, forming a unified multi-omics feature expression profile. The multi-omics feature expression profile is input into a biological network-based computational model. This model simulates the biological network structure of the pathological process of pulmonary fibrosis, analyzing the degree of association between each potential target and the disease mechanism. The degree of association is used as a quantitative indicator for dynamically ranking potential targets. The ranking process considers the statistical significance of the association strength, selecting high-priority drug targets. High-priority drug targets are imported into the drug target management system, which uses database technology to achieve persistent storage and retrieval of target information.

[0032] Example 1: See Figure 2 The collection and preliminary integration of multi-omics raw data related to pulmonary fibrosis were conducted, with data acquisition based on multiple authoritative public databases. A comprehensive gene expression database served as the primary data source, providing single-cell RNA sequencing data from pulmonary fibrosis patient tissues. The acquisition of single-cell RNA sequencing data was achieved through an automated download of raw sequencing files via a programming interface. A breakpoint resumption mechanism was employed during the data download process to ensure the integrity of large-scale file transfers. Each data file included complete metadata descriptions, including sample number, basic patient information, and sequencing platform parameters. After parsing, the metadata was stored in a structured database table, which was indexed to support rapid retrieval and cross-validation. Mass spectrometry proteomics data collection from a protein database was conducted concurrently with gene expression data acquisition. This mass spectrometry proteomics data covered protein expression profiles from both animal and cell models of pulmonary fibrosis. Data acquisition scripts automatically filtered raw data files that did not meet quality standards, which were determined based on file size, format, and internal consistency checks. The selection of pulmonary fibrosis-related models was based on the degree to which the models represented disease characteristics, including pathological scores and fibrosis marker expression levels. The raw mass spectrometry data was converted to a standardized mzML format, preserving all original spectrogram and annotation information during the conversion process.

[0033] The whole-genome methylation data extraction in the genomics data warehouse considers the dynamic process of disease development. The whole-genome methylation data includes samples from different stages of pulmonary fibrosis. Disease staging is based on clinical diagnostic criteria and imaging features, establishing a precise correspondence between staging information and methylation data. The data download task scheduling system optimizes download order and network bandwidth usage, monitors download progress, and automatically retryes failed tasks. Methylation data preprocessing includes background correction and signal intensity normalization. Preprocessed data is stored in a columnar database to improve query efficiency. Spatiotemporal identifier alignment integrates the three types of data into a unified multi-omics raw dataset. Spatiotemporal identifiers include sample collection time points and tissue spatial location information. The alignment algorithm establishes a mapping relationship based on sample identifiers and clinical metadata, and the mapping relationship undergoes multiple verifications to ensure accuracy. The alignment of time-series data considers sampling time intervals and disease progression rates, while spatial location information distinguishes different lung lobe regions and cell type distributions. The dataset employs a hierarchical storage architecture to manage raw data and metadata, and the storage architecture implements data version management and access control.

[0034] The single-cell RNA sequencing data processing workflow includes cell quality control and gene expression quantification. Cell quality control filters out low-quality cells and multiple somatic cells. After gene expression matrix generation, preliminary correlation analysis is performed with proteomic data to identify correlation patterns at the RNA and protein levels. Mass spectrometry proteomic data library search results are cross-mapped with genome annotation information, using standard gene identifiers as a bridge. Probe annotations for whole-genome methylation data are updated to the latest genome version, including CpG island locations and chromatin state information. A data source tracing mechanism is implemented during the construction of the multi-omics raw dataset, recording the origin and transformation history of each data file. Each data point in the dataset is associated with a quality assessment score, which guides data weight allocation in subsequent analyses. The data storage system implements off-site backup and disaster recovery capabilities, with backup strategies ensuring data security and long-term availability. The data access interface provides multiple query methods supporting conditional filtering, and query results are exported in standard bioinformatics file formats.

[0035] The data acquisition module and quality control module are connected via a pipeline, enabling automatic data flow and monitoring of processing status. Data exchange between modules employs an adaptive buffering mechanism, balancing processing speed and memory efficiency. System logs record complete data acquisition and processing activities, used for performance optimization and fault diagnosis. The metadata schema definition for the multi-omics raw dataset strictly adheres to international standards, supporting semantic interoperability and data integration. Single-cell RNA sequencing data from pulmonary fibrosis patient tissues contains information on rare cell types, which may be involved in specific pathways of fibrosis progression. An incremental update strategy is implemented during data download, downloading only newly added or modified data files. Gene expression matrix generation uses the latest version of the cell barcode parsing algorithm, which distinguishes real cells from background noise. Expression level standardization considers sequencing depth and gene length factors, facilitating cross-sample comparisons. Mass spectrometry proteomic data from pulmonary fibrosis-related models contain quantitative information on post-translational modifications, requiring specific verification and validation steps. Reanalysis of raw mass spectrometry data uses a unified database and search parameters, ensuring comparability between different datasets. Protein inference results are processed using strict false discovery rate control standards to minimize the impact of incorrect identifications on subsequent analyses. The normalization method for quantitative data is selected based on the experimental design, eliminating technical variations while preserving biological differences.

[0036] Genome-wide methylation data covering different stages of pulmonary fibrosis reveals epigenetic dynamics. Methylation data preprocessing includes probe filtering and color bias correction. Batch effect correction algorithms identify and adjust for the influence of technical batches, preserving biological differences. Differential methylation analysis considers cellular heterogeneity, which may confound methylation signals. Integration of methylation and gene expression data needs to consider regulatory directions, and the integration analysis reveals the impact of epigenetic regulation on transcription. Spatiotemporal alignment algorithms were developed considering the specific needs of data types, handling time series and spatial coordinates at different resolutions. Temporal alignment considers the relationship between sample collection time and disease progression, while spatial alignment integrates macroscopic tissue location and microscopic cell adjacency information. Visualization tools for alignment results help verify the quality of data integration, demonstrating the consistency of multi-omics data across spatiotemporal dimensions. The completion of the construction of the multi-omics raw dataset marks the end of the data preparation phase, providing high-quality input for subsequent analysis.

[0037] Example 2: See Figure 3The process involves quality control and normalization of the raw multi-omics data, followed by data fusion to create a multi-omics feature expression spectrum. An automated quality scoring algorithm evaluates the signal-to-noise ratio (SNR) of each data sample, calculating the ratio of signal intensity to background noise and generating a quality score. Samples with quality scores below a preset threshold are removed using an adaptive threshold filtering technique, which dynamically adjusts the removal boundary based on the overall data distribution. The retained samples undergo quantile normalization, adjusting the expression distribution of different samples to a consistent quantile benchmark. The normalized data is then processed using a sliding window algorithm to eliminate batch effects, identifying and correcting technical variations introduced by experimental batches. The resulting standardized multi-omics dataset exhibits a Gaussian distribution and possesses statistical consistency to meet subsequent analytical needs. In the data fusion stage, a multimodal data fusion framework based on tensor decomposition is constructed, representing different omics data as high-dimensional tensor structures. Different omics data are mapped to a unified multidimensional feature space, the dimensions of which are optimized to preserve the maximum biological variance. Feature space alignment is optimized using alternating least squares, which iteratively updates the feature mapping matrix until convergence. A multi-source feature cross-validation algorithm detects and eliminates data conflicts between modalities, comparing feature consistency across different data sources. The resulting multi-omics feature expression spectrum contains complete biological significance information and serves as the input feature matrix for downstream analysis.

[0038] The quality control process implements a real-time monitoring mechanism to track each processing step, recording changes in quality indicators during data transformation. The quality scoring algorithm integrates multiple evaluation indicators, including detection limits and repeatability, with indicator weights dynamically adjusted based on data type. Adaptive threshold filtering uses percentile methods to determine cut-off points, reducing the impact of outliers on threshold settings. Quantile normalization establishes a reference distribution as a calibration benchmark, derived from merged data from a high-quality sample set. A sliding window algorithm configures adjustable window sizes to adapt to different data densities, with window size optimization based on the correlation strength between data points. Batch effect elimination considers confounding factors in the experimental design, including sequencing batches, experimental dates, and operator factors. The calibration model incorporates covariate adjustment techniques to balance interference from technical factors and biological signals. The standardized multi-omics dataset uses a hierarchical data format, supporting fast read / write and random access. Data validation checks the uniformity of the standardized data distribution, ensuring no new biases are introduced. The tensor decomposition model uses the Tucker decomposition structure to process multi-omics data, capturing interactions between different modalities. The feature space dimension is determined based on the eigenvalue decay curve, which identifies the principal components that contribute significantly. The alternating least squares computation process is accelerated using a parallelization strategy, distributing the computational tasks across multiple processor cores. Convergence criteria are set with a relative error threshold and a maximum number of iterations; the relative error threshold prevents overfitting. Feature cross-validation employs a leave-one-out method to verify modal consistency, hiding one modality at a time to test the reconstruction error.

[0039] The generation process of multi-omics feature expression profiles includes a feature selection step, based on analysis of variance and mutual information indices. The profile data structure design supports feature backtracking queries, allowing the tracing of the original data source for each feature. The data fusion system implements a redundant feature elimination algorithm, merging highly correlated feature terms. Intermediate results during the fusion process are cached to improve the efficiency of repeated computations. The final feature expression profile is standardized and converted to Z-score form, making different features comparable in scale. The anomaly detection module of the quality control system identifies deviation patterns, using the isolated forest algorithm to mark outliers. The data cleaning rule base defines correction strategies for various data problems, dynamically loading corresponding rules based on data type. The normalization parameter database stores historical normalization parameters, used for benchmark calibration of new data. The processing log records the complete quality trajectory of each sample, supporting traceability analysis of quality issues. The multimodal data fusion framework includes a modality weight allocation mechanism, based on the data quality and biological importance of each modality. The feature space alignment optimization objective function includes a smoothness constraint, which preserves the topological structure of the feature space. The conflict detection algorithm calculates the inter-modal feature correlation matrix, which identifies inconsistent feature pairs. The fusion result is evaluated using an internal consistency index, which measures the degree of harmony among the fused features. The output of the multi-omics feature representation spectrum includes a feature quality score, which guides feature weight settings in downstream analyses.

[0040] The data processing pipeline implements a fault-tolerance mechanism to handle abnormal situations, activating a backup plan when a step fails. A resource management system monitors memory and CPU usage, preventing processing interruptions caused by resource contention. Standardized multi-omics datasets and multi-omics feature expression profiles are linked to versions, ensuring the reproducibility of the data analysis process. The entire processing system adopts a modular design concept, allowing for independent upgrades and maintenance of individual components. Data flow facilitates inter-module communication via message queues, decoupling dependencies between processing modules. The automated quality control process integrates with a laboratory information management system interface, providing metadata for sample preparation. Normalization parameter selection is automatically optimized based on data distribution characteristics, using a grid search method to find the optimal parameter combination. The batch effect correction model includes variance partitioning analysis, quantifying the contributions of technical and biological factors. Tensor decomposition calculations employ a random initialization and multiple-run strategy to avoid local optima. A feature space visualization tool provides a low-dimensional projected view, aiding in the evaluation of the fusion results' quality. The generation of multi-omics feature expression profiles marks the completion of the data preprocessing stage, which provides qualified input for computational model analysis based on biological networks.

[0041] Example 3: Utilizing a biological network-based computational model to analyze multi-omics feature expression profiles and deduce the association between each potential target and the pathological process of pulmonary fibrosis. The construction of a pulmonary fibrosis-specific multi-level biological network model integrates protein-protein interactions, gene regulatory relationships, and signal transduction pathways. The network model data comes from several authoritative biological databases, including STRING, KEGG, and Reactome. Network nodes represent proteins, genes, and metabolites, and network edges characterize the strength of functional associations between molecules. The multi-omics feature expression profile is projected onto network nodes through node attribute mapping, which assigns expression levels and methylation levels to corresponding nodes. A random walk algorithm simulates the particle diffusion process on the biological network, traversing the network structure starting from a specific set of pathology-related seed nodes. Connectivity strength is calculated based on the stationary distribution probability of a random walker visiting each node; this stationary distribution probability reflects the topological affinity between nodes and core pathological modules.

[0042] Modular analysis employs a community detection algorithm to identify functional units within the network, dividing the network into functional modules based on edge density optimization. Functional contribution is quantified by the topological centrality of a target node within its module, including node degree, betweenness centrality, and eigenvector centrality indices. The synthesis of connectivity indices is based on a nonlinear integration of connectivity strength and functional contribution, using a weighted geometric mean algorithm to balance the contributions of both indices. The formula for the weighted geometric mean algorithm is as follows:

[0043]

[0044] in: Indicate target point Standardized correlation score, Indicate target point The connectivity strength, Indicate target point Functional contribution and This is a hyperparameter that adjusts the relative weights of connectivity strength and functional contribution. The Monte Carlo simulation method assesses the statistical significance of the association degree, generating a random network background distribution to calculate the p-value. The Bayesian update mechanism iteratively optimizes the association degree estimate, combining the prior distribution and the likelihood function to obtain the posterior distribution.

[0045] A multi-level biological network model specific to pulmonary fibrosis comprises three levels: molecular, pathway, and phenotypic. Connections between levels are established through hierarchical relationships. Network edge weights are assigned based on the strength of experimental evidence and evolutionary conservation, and are dynamically updated to reflect the latest research progress. Missing values ​​are imputed before projection of multi-omics feature expression profiles using the k-nearest neighbor algorithm based on network neighbor information. The projected node feature vectors contain multi-dimensional biological information and support complex network algorithm computations.

[0046] The transition probability matrix of the random walk algorithm is constructed based on edge weights, and a restart mechanism is introduced into the transition probability matrix to ensure walk convergence. The core pathological module is defined based on the known set of pulmonary fibrosis-related genes, and includes key components of the transforming growth factor β pathway. Connectivity strength calculation considers the length discount of the walk path, assigning higher weights to shorter paths. Stationary distribution calculation uses a power-law iteration method to solve for eigenvectors, and the power-law iteration method sets a convergence threshold to avoid infinite loops. Functional unit identification in modular analysis uses the Louvain algorithm to optimize modularity, and the Louvain algorithm recursively merges communities to maximize the modularity function. Functional contribution calculation integrates the local importance of nodes within a module and the bridging role between modules, with the bridging role quantified by the participation coefficient. The hyperparameters of the association degree synthesis algorithm are determined through grid search, which optimizes the hyperparameter combination on the validation set. The property of the weighted geometric mean ensures that the association degree score is insensitive to extreme values, and the score range is normalized to the interval between zero and one. The random network generation in Monte Carlo simulation maintains the original network degree distribution, which is maintained through an edge reconnection algorithm. Statistical significance assessment employs empirical p-value correction for multiple hypothesis testing, with false discovery rate control used for correction. The prior distribution of the Bayesian update mechanism is based on known target knowledge, and conjugate priors are used to simplify calculations. The expected value of the posterior distribution serves as a point estimate of the correlation degree, and the posterior variance measures the uncertainty of the estimate.

[0047] The visualization of the biological network model employs a force-directed layout algorithm, which clearly displays the network module structure. Parallel computing is implemented during network analysis to accelerate large-scale network computation, dividing the network into subgraphs for independent processing. The output includes a complete score profile for each target, supporting result traceability and sensitivity analysis. The correlation degree derivation process utilizes an automated batch processing mode, supporting large-scale target screening tasks. Parameter settings for the computational model are managed through configuration files, allowing parameter adjustments without modifying the code. The restart probability of the random walk algorithm is set to an empirical value, balancing the local and global exploration capabilities of the walk. Seed node selection for the core pathology module considers node degree centrality, as height nodes have a significant impact on walk results. Normalization is introduced in connectivity strength calculation to eliminate node degree bias, making nodes of different degrees comparable. Functional enrichment analysis is performed on the community structure generated by modular analysis, validating the biological rationale of the community.

[0048] Functional contribution calculation incorporates a connectivity density index within modules, which measures the tightness of connections between nodes and other nodes within the module. Association degree synthesis considers disease relevance weights for different modules, based on the strength of the association between the module and the pulmonary fibrosis phenotype. Hyperparameter optimization of the weighted geometric mean algorithm uses cross-validation to prevent overfitting to training data. The number of repetitions in the Monte Carlo simulation is determined according to accuracy requirements, as the number of repetitions affects the stability of the p-value estimation. The number of iterations in the Bayesian update mechanism sets a convergence criterion, avoiding premature stopping or overcomputation. The final standardized association degree score is quantile-normalized to ensure a consistent score distribution. A quality control step is implemented throughout the analysis process to verify intermediate results and identify computational errors and data anomalies. The association degree score serves as the primary basis for target prioritization, and is combined with subsequent validation indicators to assess target potential. Computational model analysis based on biological networks reveals the systematic location of potential targets within the pulmonary fibrosis pathological network, providing a quantitative basis for target prioritization.

[0049] See Figure 4 This study presents the comprehensive analysis results of potential drug targets for pulmonary fibrosis based on a biological network model. A scatter plot visually presents the systematic location and importance assessment of each target within the disease pathology network. Each bubble in the plot represents a potential drug target; its horizontal position reflects the network connectivity strength between the target and the core pathology module, while its vertical position reflects the topological contribution of the target within the functional module. The size of the bubble intuitively indicates the overall correlation score, and the color intensity represents the statistical significance level. The upper right corner of the plot clusters important targets with strong connectivity, high functional contribution, and significant correlation; these are priority candidate targets for pulmonary fibrosis drug development. The known core pathology module targets marked with red asterisks validate the reliability of the analysis method. The overall distribution pattern reveals the modular characteristics and hierarchical structure of disease-related targets within the biological network, providing important evidence for systematic drug target screening.

[0050] Example 4: Potential targets are ranked and high-priority drug targets are selected based on their correlation. A validation step is incorporated into the ranking process to assess target reliability. A dynamic priority queue management mechanism updates the target ranking in real time based on new data, and uses a max-heap data structure to maintain target priorities. A sliding time window analyzes the changing trends of correlation, and the width of the sliding time window is adaptively adjusted according to the data update frequency. A percentile ranking algorithm divides targets into priority groups based on their correlation distribution, and the ranking is determined based on the relative position of the target groups. Clinical importance weights integrate disease relevance and drugability information, and are quantified through expert review and literature evidence. A time-sensitive target priority list is marked with version timestamps, supporting historical version tracing and comparative analysis.

[0051] The target stability testing scheme for cross-validation divides the data into multiple subsets, and the cross-validation scheme uses k-fold cross-validation to assess the ranking stability. External validation using independent datasets employs unseen data from different research cohorts or experimental platforms. Subtype-specificity analysis assesses the expression specificity of the target at different disease stages, comparing differential expression among different clinical subgroups. The target confidence report includes sensitivity and specificity, providing a reliable measure of target priority. The dynamic priority queue management mechanism is implemented using a message queue architecture, which handles real-time data streams and update requests. Queue update triggers are activated based on new data arrival events, initiating a re-ranking calculation process. Sliding time window analysis configures multiple time scales, capturing short-term fluctuations and long-term trends. The calculation of correlation trends uses a linear regression model, fitting the slope of the time series data. Trend significance testing uses a t-test to assess the statistical significance of the changing trend. The percentile ranking algorithm's group boundaries are based on the quantiles of the data distribution, and the group boundaries are dynamically adjusted to adapt to changes in the target population distribution. Priority grouping is defined into three categories: high priority, medium priority, and low priority, with different subsequent processing strategies corresponding to each priority category. The clinical importance weights were determined using the Delphi method, which collected expert opinions and reached a consensus through multiple iterations. The weight allocation matrix integrates the weights of different clinical dimensions into a comprehensive weight, and the weight allocation matrix remains transparent and adjustable.

[0052] The timeliness management of the target priority list is based on a data freshness index, which measures the interval between the data collection time and the current time. A list version control system records a complete snapshot of each ranking update, supporting difference comparison and rollback operations. The stability test of k-fold cross-validation randomly divides the data into k mutually exclusive subsets, with the k value chosen based on data size and computational resources. The stability metric calculates the consistency of the ranking results across different subsets, using Kendall's coefficient of harmony. External validation datasets are obtained from a public data warehouse, providing datasets independent of the training data. The validation process calculates the predictive performance of target priorities on independent datasets, using the area under the receiver operating characteristic (ROC) curve. Subtype-specificity analysis grouping is based on clinicopathological features, including disease stage, pathological type, and treatment response. Differential expression analysis uses a linear model to account for confounding factors, including age and gender. The generation of target confidence reports automatically integrates validation results, which are summarized and visualized in a standardized format. The report includes confidence intervals for target priority ranking, reflecting the uncertainty of the ranking estimate. The entire sorting and validation process undergoes quality assurance checks, verifying the quality of input data and the correctness of the computation process. The combination of sorting results and validation metrics supports the target selection decision-making process, which considers the balance between scientific evidence and clinical needs. A max-heap data structure with a dynamic priority queue management mechanism maintains the target priority order, ensuring efficient query and update operations. Heap structure adjustment operations are performed after each data update, maintaining the heap property. The width of the sliding time window is optimized based on data volatility; a narrower window is used to capture rapid changes when data volatility is high. Analysis of correlation trends incorporates seasonal adjustments to eliminate the impact of periodic fluctuations.

[0053] The percentile ranking algorithm is implemented using an online computation algorithm, which supports streaming data scenarios. Grouping thresholds are dynamically adjusted based on the number of targets, ensuring a reasonable distribution of targets across groups. The Delphi method for clinical importance weights employs an anonymous feedback mechanism, reducing the herd effect of expert opinions. The weight matrix is ​​maintained through regular updates, reflecting advancements in medical knowledge. The timeliness markers for the target priority list use the ISO time format, ensuring unambiguous interpretation of time information. The version control system uses a differential storage strategy for snapshot storage, saving storage space. K-fold cross-validation subset partitioning maintains class proportions, achieved through stratified sampling. The Kendall's concordance coefficient, a stability measure, considers tie rankings, handled using a modified formula. Preprocessing of the external validation dataset follows the same process as the training data, ensuring data comparability. The area under the receiver operating characteristic (ROC) curve for predictive performance evaluation is calculated using the trapezoidal rule, which approximates the integral area under the curve. Differential expression tests for subtype specificity analysis were performed using the empirical Bayesian method, which reduces variance estimation and improves performance in small samples. Multiple test correction was used to control the false positive rate; the Benjamin-Hotchberg method was employed for multiple test correction.

[0054] The target confidence report is designed with machine readability in mind, facilitating subsequent automated processing. Confidence interval calculations utilize bootstrap resampling, which is independent of distribution assumptions. Quality assurance checks include data integrity verification, which checks for missing and outlier values. The sorting and verification processes are logged for detailed operational information, used for performance monitoring and troubleshooting. See Table 1 for the relationship between target priority ranking and verification metrics.

[0055] Table 1: Correspondence between Target Priority Grouping and Validation Metrics

[0056] Priority grouping Percentile range of correlation Stability coefficient range External validation AUC threshold Subtype specificity requirements High priority ≥85% ≥0.80 ≥0.75 Significant in at least two subgroups Medium priority 60%-85% 0.60-0.80 0.65-0.75 Significant in at least one subgroup low priority <60% <0.60 <0.65 No significance requirement

[0057] Priority grouping is based on percentile intervals of association, reflecting the relative ranking of targets. The stability coefficient range measures the consistency of the ranking results; high consistency indicates reliable ranking. The external validation AUC threshold sets the minimum standard for predictive performance, which affects the target's translational potential. Subtype specificity requires ensuring the target's relevance within specific disease subtypes; subtype relevance increases the clinical application value of the target.

[0058] The dynamic priority queue management mechanism optimizes throughput to support large-scale target screening, and this throughput optimization is achieved through parallel computing. The real-time requirements of sliding time window analysis determine the data update frequency; scenarios with high real-time requirements employ a stream processing architecture. The computational complexity of the percentile ranking algorithm is linearly related to the number of targets, ensuring algorithm scalability. The update cycle for clinical importance weights is synchronized with medical knowledge updates, typically every six months to one year. The storage of the target priority list utilizes database indexes to optimize query performance, accelerating time-based range queries. The computational load of k-fold cross-validation is distributed across multiple computing nodes, reducing the computational pressure on individual nodes. Data quality filtering is implemented for acquiring external validation datasets, excluding low-quality validation data. Multiple test corrections for subtype-specific analysis control the family error rate, maintaining the reliability of the overall conclusions. The target confidence report is organized using a modular structure, allowing for the flexible addition of new validation metrics. Automated report generation reduces manual intervention, improving efficiency and consistency. The entire ranking and validation system interface provides a standard data format, facilitating system integration and data exchange.

[0059] Example 5: High-priority drug targets are imported into a drug target management system for persistent storage. The drug target management system constructs a target storage architecture based on a graph database. The graph database uses an attribute graph model to represent targets and their relationships. The attribute graph model contains three basic elements: nodes, edges, and attributes. Nodes represent drug targets, biological entities, and compounds, while edges depict the biological relationships between entities, such as interactions, regulation, and participation. A multi-dimensional indexing mechanism supports rapid retrieval based on target name, function, and disease association. The multi-dimensional index uses a combination of B+ trees and inverted indexes. A version control system manages the update history of target data, recording the content, time, and operator of each data change. The mapping of the relationship between targets and related biological entities is based on standard biological ontologies, including GeneOntology and ChEBI ontologies. Persistent storage writes data to non-volatile storage media, ensuring long-term accessibility of the data. An audit trail system based on blockchain technology is deployed in a distributed network environment, recording all data access and modification operations. Smart contracts automatically execute data integrity check rules, encoding business logic and automatically triggering the check process. Target data change reports are generated periodically, including summary information, change statistics, and descriptions of anomalies. Data consistency is verified using cryptographic hash values, which are generated using the SHA-256 algorithm. Audit results are synchronized to a distributed data management platform, which maintains multiple data replicas to ensure fault tolerance.

[0060] The graph database storage architecture, using Neo4j graph database management system as an example, provides native graph storage and processing capabilities. Node attributes include target identifiers, official names, gene symbols, and descriptive information; edge attributes define relation types, evidence sources, and confidence scores. The indexing mechanism is created on node labels and attribute combinations, accelerating specific queries. The version control system is implemented based on Git version control principles, which manage the change history of data files. Relationship mapping uses standard identifiers to link records in different databases, including UniProtID and EnsemblID. Persistent storage employs a transaction mechanism to ensure data consistency, guaranteeing the atomicity of multiple operations. The blockchain network is maintained by multiple nodes as a distributed ledger, storing immutable audit records. Smart contracts are written using the Solidity programming language, which defines data integrity check conditions. Data change report generation can be configured daily or weekly, depending on the data update frequency. Cryptographic hash calculations are performed before and after data storage, and hash value comparisons verify whether the data has been tampered with.

[0061] Audit results are synchronously replicated asynchronously using message queues, which decouple the audit system from the storage system. The graph database uses the Cypher query language, which supports declarative graph pattern matching. Node and edge creation follows a standardized data model, ensuring data consistency and interoperability. Index maintenance is automatically triggered when data is updated, maintaining consistency between the index and the data. The version control system's branch management supports parallel development, allowing different versions to evolve independently. The accuracy of relation mappings is maintained through regular synchronization with authoritative databases, such as NCBIGene, which provides the latest annotations. Persistent storage backup strategies include full and incremental backups, with recovery point and time targets defined. The blockchain network's consensus algorithm employs a practical Byzantine fault-tolerant algorithm, which tolerates partial node failures. Smart contract deployments undergo rigorous testing to verify logical correctness, covering various boundary conditions and abnormal scenarios. Data change reports can be customized to meet different audit needs, including structured data and unstructured text.

[0062] Cryptographic hashes are stored in a secure area separate from the original data, preventing attackers from simultaneously tampering with both the data and the hashes. The visualization interface for audit results provides multi-dimensional analytical views, helping to identify data access patterns. The graph database's storage architecture is optimized for graph traversal query performance, achieved through adjacency node pointers. Node attributes support full-text search functionality, utilizing the Lucene search engine library. The indexing mechanism includes composite indexes covering multiple query conditions, reducing random I / O operations. The version control system's difference comparison algorithm identifies data changes, efficiently calculating textual or binary differences. The relational mapping update mechanism handles ontology version evolution, keeping mapping relationships synchronized with the ontology version. The persistent storage compression algorithm reduces storage space usage, balancing compression ratio and access speed. The blockchain network's smart contracts support an upgrade mechanism, allowing for vulnerability fixes and feature additions. The data change report distribution list can be configured to specify recipients, supporting dynamic addition and removal of members. The verification frequency of cryptographic hashes is set based on data sensitivity, with critical data requiring frequent verification being verified multiple times daily. The long-term archiving of audit results meets regulatory compliance requirements, and read-only storage media is used to prevent modification. The drug target management system's user interface provides a data import wizard that guides users through uploading target data. The access control system manages user access levels to data, assigning operation permissions based on roles. System logs record all user operations and system events, and are used for security auditing and fault diagnosis.

[0063] Cluster deployment of graph databases provides high availability, achieved through data sharding and replication. The query optimizer selects efficient execution plans and analyzes query patterns and data statistics. Transaction isolation levels ensure data consistency across concurrent accesses, preventing dirty reads and non-repeatable reads. Version control system conflict resolution strategies handle parallel modifications and prompt users to manually resolve conflicts. Relational mapping validation tools check identifier validity, querying external databases to verify identifier existence. Persistent storage integrity checks periodically scan for data errors, using checksums to verify data block integrity. Encrypted transmission in the blockchain network protects inter-node communication security, using TLS to prevent eavesdropping and tampering. Smart contract gas mechanisms prevent infinite loops and measure computational resource consumption. Data change report template engines support custom layouts, filling data into preset formats. Salting cryptographic hashes prevents rainbow table attacks, and adding random strings to the salt enhances hash security. Statistical analysis of audit results identifies abnormal access patterns, applying machine learning algorithms to detect anomalies. The drug target management system's application programming interface (API) follows RESTful design principles, which utilize standard HTTP methods to manipulate resources. The data export function supports multiple formats, including JSON and XML, while preserving data structure and semantics. The system integration interface interfaces with a laboratory information management system, which provides raw experimental data. A monitoring dashboard displays real-time system performance metrics, helping administrators understand the system's status.

[0064] In this example of a graph database storage architecture, a target node might contain attributes such as "TGFB1," whose official name is "TransformingGrowthFactorBeta1." This node is connected to another node representing "TGFBR2" via an edge named "INTERACTS_WITH," with the edge attributes recording the evidence source as "STRINGdatabase" and a confidence score of "0.98." Multi-dimensional indexing allows users to quickly query all targets related to the fibrosis pathway. A version control system records that the last modification to the "TGFB1" node was updating the confidence score from "0.95" to "0.98." Relational mapping links the "TGFB1" node to the terminology of "SMAD protein signal transduction" in GeneOntology. Persistent storage ensures this data remains available after a system restart. A blockchain audit trail records every query and modification operation on the "TGFB1" node data, and smart contracts check the data integrity rules of the "TGFB1" node, such as whether required fields are complete. Cryptographic hashes are recalculated when the "TGFB1" node data is updated, and audit results are synchronized to three geographically distributed storage nodes. This specific implementation demonstrates how a drug target management system can achieve secure and traceable target data management.

[0065] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0066] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for screening drug targets of pulmonary fibrosis by fusing multi-omics data, characterized in that, The method performs the following operations: Collecting multi-omics raw data related to pulmonary fibrosis from public databases and experimental data sources; Performing quality control and normalization processing on the collected multi-omics raw data to generate standardized multi-omics data sets; Integrating the standardized multi-omics data sets into a unified multi-omics feature expression profile through data fusion technology; Analyzing the multi-omics feature expression profile using a biological network-based computational model to deduce the correlation degree of each potential target point to the pathological process of pulmonary fibrosis; According to the correlation degree, the potential target points are sorted, and high-priority drug target points are screened out; the high-priority drug target points are imported into the drug target point management system for persistent storage. 2.The method of claim 1, wherein The collection of multi-omics raw data related to pulmonary fibrosis includes the following steps: obtaining single-cell RNA sequencing data of pulmonary fibrosis patient tissues from the Gene Expression Omnibus database; collecting mass spectrometry proteomic data of pulmonary fibrosis related models from the protein database; extracting whole genome methylation data covering different disease stages of pulmonary fibrosis from the genomics data warehouse; aligning the above three types of data with time and space labels to form a multi-omics raw data set with time and space labels. 3.The method of claim 1, wherein, The quality control and normalization processing of the collected multi-omics raw data includes the following steps: using an automatic quality scoring algorithm to evaluate the signal-to-noise ratio of each data sample; removing abnormal samples with a score below a preset threshold through adaptive threshold filtering technology; applying quantile normalization method to correct the expression distribution of the retained samples; using a sliding window algorithm to eliminate batch effects of the corrected data; finally generating a standardized multi-omics data set conforming to the characteristics of Gaussian distribution. 4.The method of claim 1, wherein, The integration of the standardized multi-omics data sets into a unified multi-omics feature expression profile through data fusion technology includes the following steps: constructing a multi-modal data fusion framework based on tensor decomposition; mapping different omics data to a unified multi-dimensional feature space; using an alternating least squares method to align the feature space; eliminating inter-modal conflicts through multi-source feature cross-validation algorithm; generating a multi-omics feature expression profile with complete biological significance.

5. The method of claim 1, wherein the method is a method of screening for drug targets for pulmonary fibrosis by fusing multi-omics data, characterized by, The analysis of the multi-omics feature expression profile using a biological network-based computational model includes the following steps: establishing a multi-level biological network model specific to pulmonary fibrosis; projecting the multi-omics feature expression profile onto the network nodes; using a random walk algorithm to calculate the connectivity strength of each potential target point to the core pathological module; determining the functional contribution of the target point in different biological processes through modular analysis; finally outputting the quantitative correlation degree index.

6. The method of claim 5, wherein the method is characterized by, The deducing of the correlation degree of each potential target point to the pathological process of pulmonary fibrosis includes the following steps: constructing an association degree synthesis algorithm based on weighted geometric mean; nonlinearly combining connectivity strength and functional contribution; using Monte Carlo simulation method to evaluate the statistical significance of the association degree; iteratively optimizing the association degree estimate value through Bayesian updating mechanism; generating the final standardized correlation degree score.

7. The method of claim 1, wherein the method is a method of screening for drug targets for pulmonary fibrosis by fusing multi-omics data, characterized by, The sorting of potential target points according to the correlation degree comprises the following steps: establishing a dynamic priority queue management mechanism; using a sliding time window to analyze the trend of the correlation degree; determining the priority grouping of the target points by a percentile sorting algorithm; adjusting the final sorting result in combination with the clinical importance weight; and generating a target priority list with time effectiveness. 8.The method of claim 7, wherein, The screening of high-priority drug targets further comprises a verification step: designing a target stability test scheme based on cross-validation; External verification is performed by using independent data sets; subtype-specific analysis is used to evaluate the expression specificity of the target points at different stages of the disease; and finally, a target reliability report containing verification indicators is generated. 9.The method of claim 1, wherein, The introduction of high-priority drug targets into the drug target management system for persistent storage comprises the following steps: constructing a target storage architecture based on a graph database; designing a multi-dimensional indexing mechanism to support fast retrieval; establishing a version control system to manage the update history of target data; realizing the relationship mapping of target points and related biological entities; and completing the persistent storage of target data.

10. The method of claim 9, wherein the method is a method of screening for drug targets for pulmonary fibrosis by fusing multi-omics data. The method further comprises a data auditing step: deploying an audit tracking system based on blockchain technology; setting up an intelligent contract to automatically perform data integrity checks; generating a target data change report at regular intervals; and verifying data consistency through cryptographic hash values; The audit results are synchronized to a distributed data management platform.

Citation Information

Patent Citations

  • Drug repositioning method based on multi-information fusion and random walk model

    CN107506591A

  • Repositioning drug discovery method based on integration of a plurality of transcriptome data sets and drug target information

    CN108694991A

  • Construction method and application of multi-scale and multi-level regulation and control network

    CN119541622A

  • ARDS treatment target analysis system combined with gene expression data

    CN119763651A

  • Unicellular organism network inference method based on multi-omics multilayer heterogeneous network

    CN120412759A