Dynamic biomarker analysis method and system based on big data

Through the attribute identification and debatch effect method of dynamic biomarker data, data stability and accuracy problems in dynamic biomarker analysis are solved, and efficient data analysis and risk assessment are achieved.

CN120260909AInactive Publication Date: 2025-07-04ZHEJIANG HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510325983.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, dynamic biomarker analysis cannot perform batch-effect processing based on the specific attributes of the data, resulting in low data stability and accuracy, large calculation volume, and cannot be used as an effective analysis standard.

Method used

By obtaining the attributes of dynamic biomarker data, identifying and applying corresponding debatch effect methods, the data is classified and processed, the biomarker processing data is obtained, and compared with historical data, the biomarker data related to the target disease is screened out, and a regulatory network for molecular interaction relationships is constructed.

Benefits of technology

It improves the comparability and stability of data, reduces the amount of calculation, provides a more comprehensive basis for risk assessment, and improves data analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260909A_ABST
    Figure CN120260909A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic biomarker analysis method and system based on big data, and relates to the technical field of data analys.The dynamic biomarker analysis method comprises the steps of obtaining dynamic biomarker data, obtaining a batch effect removing method corresponding to each dynamic biomarker data according to the dynamic biomarker data, and obtaining a batch effect removing method corresponding to each dynamic biomarker data based on the batch effect removing method; and processing the dynamic biomarker data to obtain biomarker processing data. According to the method, the data comparability is ensured and the reliability of subsequent analysis is improved through a data attribute intelligent matching batch removal effect method, the batch removal effect is verified through the differential gene number and the overlapping rate, the stability of the data is ensured, the biomarker data types related to the target disease are selected through the biomarker feature set, and the accuracy of the biomarker data types related to the target disease is improved. According to the method, the data calculation amount is reduced, the data analysis efficiency is improved, and a more comprehensive risk assessment basis is provided for clinic by constructing the regulation and control network containing the molecular interaction relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and specifically relates to a method and system for analyzing dynamic biomarkers based on big data. Background Art

[0002] Biomarkers are defined as measurable indicators of health status, disease presence, or intervention effects, and play a key role in clinical research. Studying biomarkers can assist in clinical trials, and the correlation between biomarkers and clinical outcomes is crucial in drug development and regulatory decision-making. Traditional methods rely on cross-sectional data (such as gene / protein expression at a single time point), which cannot capture the dynamic changes during disease progression and single-dimensional data is difficult to comprehensively reflect the complexity of diseases. Cloud computing and distributed algorithms (such as Spark) make it possible to process time-series data of millions of samples, and deep learning models (such as LSTM, DNB) have further improved the speed and accuracy of dynamic biomarker analysis.

[0003] Currently, for the analysis of dynamic biomarkers, there are still problems such as being unable to remove batch effects from dynamic biomarker data according to the specific attributes of the data. Often, the data is directly processed by a unified batch effect removal method or not processed for batch effect removal at all, which affects the stability of the data. It is impossible to select corresponding reference data based on historical data, and often the average value of the measurement indicators is directly used as the reference data, which not only has a huge data calculation amount but also has low data accuracy and cannot be used as an analysis standard. Summary of the Invention

[0004] To solve the above technical problems, a method and system for analyzing dynamic biomarkers based on big data are provided. The technical solution of the present invention solves the problems proposed in the above background art, such as being unable to remove batch effects from dynamic biomarker data according to the specific attributes of the data. Often, the data is directly processed by a unified batch effect removal method or not processed for batch effect removal at all, which affects the stability of the data. It is impossible to select corresponding reference data based on historical data, and often the average value of the measurement indicators is directly used as the reference data, which not only has a huge data calculation amount but also has low data accuracy and cannot be used as an analysis standard.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for analyzing dynamic biomarkers based on big data, comprising:

[0007] Obtaining dynamic biomarker data, where the dynamic biomarker data includes RNA sequencing data, serum proteome data, and genomic data;

[0008] Based on dynamic biomarker data, through data attribute recognition, obtain the batch effect removal method corresponding to each dynamic biomarker data;

[0009] Based on the batch effect removal method, process the dynamic biomarker data to obtain biomarker processed data;

[0010] Obtain biomarker historical data;

[0011] Based on the biomarker historical data and the batch effect removal method, obtain biomarker baseline data;

[0012] Compare the biomarker processed data with the biomarker baseline data to obtain biomarker abnormal data, where the biomarker abnormal data represents the data in the biomarker processed data that is different from the biomarker baseline data;

[0013] Obtain target disease information, where the target disease information includes the target disease type information and pathological information for dynamic biomarker analysis;

[0014] Based on the biomarker abnormal data and the target disease information, obtain target characteristic disease information;

[0015] Based on the target characteristic disease information, analyze the biomarker abnormal data to obtain disease risk information.

[0016] Preferably, the step of based on dynamic biomarker data, through data attribute recognition, obtain the batch effect removal method corresponding to each dynamic biomarker data specifically includes:

[0017] Based on the dynamic biomarker data, obtain the data attribute corresponding to each dynamic biomarker data, that is, the cell information during data collection;

[0018] Based on the data attribute corresponding to each dynamic biomarker data, using single-cell data and mixed-cell data as the discrimination criterion, classify the dynamic biomarker data to obtain biological single-cell data and biological mixed-cell data;

[0019] Based on the batch effect removal tool, with the single-cell data batch effect removal as the benchmark, obtain the batch effect removal method corresponding to the biological single-cell data;

[0020] Based on the analysis requirements of the batch effect removal tool, obtain the data dimension threshold;

[0021] Based on the data dimension threshold, divide the dimensions of the biological mixed-cell data to obtain biological mixed-cell high-dimensional data and biological mixed-cell low-dimensional data;

[0022] Based on the batch effect removal tool, select the batch effect removal methods corresponding to the high-dimensional data and low-dimensional data of the biological mixed cells respectively;

[0023] According to the batch effect removal method corresponding to the biological single-cell data, and the batch effect removal methods corresponding to the high-dimensional data and low-dimensional data of the biological mixed cells, obtain the batch effect removal method.

[0024] Preferably, based on the batch effect removal method, process the dynamic biomarker data to obtain biomarker processed data, specifically including:

[0025] Based on the batch effect removal method, perform batch effect removal on the dynamic biomarker data to obtain initial biomarker data, where the initial biomarker data includes biological single-cell processed data, biological mixed cell high-dimensional processed data, and biological mixed cell low-dimensional processed data;

[0026] Compare each initial biomarker data with the corresponding dynamic biomarker data to obtain the number of differential genes and the overlap rate;

[0027] Based on the batch effect removal method, obtain the differential gene number threshold and overlap rate threshold corresponding to each batch effect removal method;

[0028] Compare the number of differential genes and overlap rate corresponding to each initial biomarker data with the differential gene number threshold and overlap rate threshold corresponding to the batch effect removal method corresponding to each initial biomarker data to determine whether the batch effect removal effect of the initial biomarker data meets the analysis requirements;

[0029] Among them, if the number of differential genes corresponding to the initial biomarker data exceeds the differential gene number threshold or the overlap rate exceeds the overlap rate threshold, the batch effect removal effect of the initial biomarker data does not meet the actual requirements, and reselect the batch effect removal method to perform batch effect removal on the dynamic biomarker data;

[0030] If the number of differential genes corresponding to the initial biomarker data does not exceed the differential gene number threshold and the overlap rate does not exceed the overlap rate threshold, then use the initial biomarker data as the biomarker processed data.

[0031] Preferably, the method for obtaining the biomarker reference data according to the biomarker historical data and the batch effect removal method specifically includes:

[0032] According to the biomarker historical data, based on data attribute recognition, classify the biomarker historical data to obtain biological historical single-cell data, biological historical mixed cell high-dimensional data, and biological historical mixed cell low-dimensional data;

[0033] Based on the batch effect removal method corresponding to each type of dynamic biomarker data, using the data classification result as the matching criterion, perform batch effect removal on biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data to obtain biomarker historical processed data;

[0034] Obtain target disease information;

[0035] Based on the target disease information, obtain the biomarker characteristic data information corresponding to each target disease, and the biomarker characteristic data information includes biomarker data type information;

[0036] Aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, and obtain a biomarker characteristic set;

[0037] Use the biomarker historical processed data corresponding to the biomarker data types in the biomarker characteristic set as biomarker benchmark data.

[0038] Preferably, the obtaining of target characteristic disease information according to biomarker abnormal data and target disease information specifically includes:

[0039] According to the target disease information, based on the pathological analysis of the target disease, obtain the target disease abnormal data type information corresponding to each target disease, and the target disease abnormal data type information represents the biomarker data types affected by the pathological process of the target disease;

[0040] According to the biomarker abnormal data, obtain the biomarker abnormal data type information;

[0041] Compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease to obtain target characteristic disease information;

[0042] Among them, if the biomarker abnormal data type information contains any of the target disease abnormal data type information corresponding to a target disease, then this target disease is a target characteristic disease.

[0043] Preferably, the analysis of biomarker abnormal data based on the target characteristic disease information to obtain disease risk information specifically includes:

[0044] According to the target characteristic disease information, obtain the target disease abnormal data type information corresponding to the target characteristic disease;

[0045] Based on the pathological analysis of the target disease, taking the disease impact mode as the benchmark, classify the types of abnormal data of the target disease. Consider the types of abnormal data of the target disease directly affected by the disease as the types of directly affected abnormal data, and the types of abnormal data of the target disease indirectly affected by the disease as the types of indirectly affected abnormal data, to obtain the classification information of the types of abnormal data of the target disease;

[0046] According to the information of the target characteristic disease, based on the experimental database, obtain the pathological experimental data of the biomarker;

[0047] Based on the dynamic network biomarker (DNB) algorithm, use the pathological experimental data of the biomarker corresponding to the type of directly affected abnormal data as the core nodes, and the pathological experimental data of the biomarker corresponding to the type of indirectly affected abnormal data as the regulatory edges to construct the disease critical state network;

[0048] Based on the disease critical state network, analyze the abnormal data of the biomarker to obtain the disease stage information corresponding to each target characteristic disease;

[0049] Aggregate the disease stage information corresponding to each target characteristic disease to obtain the disease risk information.

[0050] Furthermore, a dynamic biomarker analysis system based on big data is proposed to implement the above analysis method, including:

[0051] The main control module is used to compare the number of differential genes and the overlap rate corresponding to the initial data of each biomarker with the threshold of the number of differential genes and the overlap rate corresponding to the batch effect removal method for the initial data of each biomarker to determine whether the batch effect removal effect of the initial data of this biomarker meets the analysis requirements, compare the information of the types of abnormal data of the biomarker with the information of the types of abnormal data of the target disease corresponding to each target disease to obtain the information of the target characteristic disease, consider the types of abnormal data of the target disease directly affected by the disease as the types of directly affected abnormal data, and the types of abnormal data of the target disease indirectly affected by the disease as the types of indirectly affected abnormal data, to obtain the classification information of the types of abnormal data of the target disease, based on the dynamic network biomarker (DNB) algorithm, use the pathological experimental data of the biomarker corresponding to the type of directly affected abnormal data as the core nodes, and the pathological experimental data of the biomarker corresponding to the type of indirectly affected abnormal data as the regulatory edges to construct the disease critical state network, and analyze the abnormal data of the biomarker to obtain the disease stage information corresponding to each target characteristic disease;

[0052] An information acquisition module, which is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, biomarker historical data, target disease information, target disease type information for dynamic biomarker analysis, and pathological information;

[0053] A data processing module, which is used to remove batch effects from the dynamic biomarker data to obtain initial biomarker data, compare each initial biomarker data with the corresponding dynamic biomarker data to obtain the number of differential genes and the overlap rate, classify the biomarker historical data based on data attribute recognition according to the biomarker historical data to obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain the biomarker characteristic data information corresponding to each target disease based on the target disease information, aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types to obtain a biomarker characteristic set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker characteristic set as biomarker reference data;

[0054] A display module, which interacts with the main control module to output and display biomarker processing data, biomarker reference data, biomarker abnormal data, target characteristic disease information, and disease risk information.

[0055] Optionally, the main control module specifically includes:

[0056] A control unit, which is used to compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease to obtain target characteristic disease information, use the target disease abnormal data types directly affected by the disease as directly affected abnormal data types, use the target disease abnormal data types indirectly affected by the disease as indirectly affected abnormal data types, obtain target disease abnormal data type classification information, based on the dynamic network biomarker (DNB) algorithm, use the biomarker pathological experimental data corresponding to the directly affected abnormal data types as core nodes, use the biomarker pathological experimental data corresponding to the indirectly affected abnormal data types as regulatory edges to construct a disease critical state network, analyze the biomarker abnormal data, and obtain the disease stage information corresponding to each target characteristic disease;

[0057] An information receiving unit, which interacts with the information acquisition module and the data processing module to receive data and transmit it to the judgment unit;

[0058] A judgment unit, which is used to compare the number of differential genes and the overlap rate corresponding to the initial data of each biomarker with the threshold values of the number of differential genes and the overlap rate corresponding to the batch effect removal method for the initial data of each biomarker, and judge whether the batch effect removal effect of the initial biomarker data meets the analysis requirements.

[0059] Optionally, the information acquisition module specifically includes:

[0060] A first acquisition unit, which is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, and biomarker historical data;

[0061] A second acquisition unit, which is used to acquire target disease information, target disease type information for dynamic biomarker analysis, and pathological information.

[0062] Optionally, the data processing module specifically includes:

[0063] A first data processing unit, which is used to perform batch effect removal on the dynamic biomarker data to obtain the initial biomarker data, compare each initial biomarker data with the corresponding dynamic biomarker data, and obtain the number of differential genes and the overlap rate;

[0064] A second data processing unit, which is used to classify the biomarker historical data based on data attribute recognition according to the biomarker historical data, obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain the biomarker characteristic data information corresponding to each target disease based on the target disease information, aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, obtain a biomarker characteristic set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker characteristic set as the biomarker reference data.

[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0066] The present invention provides a method and system for analyzing dynamic biomarkers based on big data. By intelligently matching the batch effect removal method through data attributes, the comparability of data is ensured, the reliability of subsequent analysis is improved, and the batch effect removal effect is verified through the number of differential genes and the overlap rate, ensuring the stability of the data. By selecting the biomarker data types related to the target disease through the biomarker characteristic set, the data calculation amount is reduced, the data analysis efficiency is improved, and by constructing a regulatory network including molecular interaction relationships, a more comprehensive risk assessment basis is provided for clinical practice. Description of the Drawings

[0067] Figure 1 Flow chart of a dynamic biomarker analysis method based on big data proposed by the present invention;

[0068] Figure 2 Flow chart for obtaining the method of removing batch effects in the present invention;

[0069] Figure 3 Flow chart for obtaining the processed biomarker data in the present invention;

[0070] Figure 4 Flow chart for obtaining the benchmark data of biomarkers in the present invention;

[0071] Figure 5 Block diagram of the structure of a dynamic biomarker analysis system based on big data proposed by the present invention. Detailed implementation manners

[0072] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.

[0073] Referring to Figure 1 - Figure 4 As shown, a dynamic biomarker analysis method based on big data in an embodiment of the present invention includes:

[0074] Obtaining dynamic biomarker data, where the dynamic biomarker data includes RNA sequencing data, serum proteome data, and genomic data;

[0075] According to the dynamic biomarker data, based on data attribute recognition, obtaining the method of removing batch effects corresponding to each dynamic biomarker data;

[0076] Specifically, according to the dynamic biomarker data, based on data attribute recognition, obtaining the method of removing batch effects corresponding to each dynamic biomarker data specifically includes:

[0077] According to the dynamic biomarker data, obtaining the data attribute corresponding to each dynamic biomarker data, that is, the cell information during data collection;

[0078] According to the data attribute corresponding to each dynamic biomarker data, using single-cell data and mixed-cell data as the discrimination criterion, classifying the dynamic biomarker data to obtain biological single-cell data and biological mixed-cell data;

[0079] Based on the batch effect removal tool, using the single-cell data batch effect removal as the benchmark, obtaining the method of removing batch effects corresponding to the biological single-cell data;

[0080] Based on the analysis requirements of the batch effect removal tool, obtain the data dimension threshold;

[0081] According to the data dimension threshold, divide the biological mixed cell data into dimensions to obtain high-dimensional biological mixed cell data and low-dimensional biological mixed cell data;

[0082] Based on the batch effect removal tool, select the batch effect removal methods corresponding to the high-dimensional biological mixed cell data and the low-dimensional biological mixed cell data respectively;

[0083] According to the batch effect removal methods corresponding to the biological single-cell data, the high-dimensional biological mixed cell data, and the low-dimensional biological mixed cell data, obtain the batch effect removal methods.

[0084] In this solution, by classifying the dynamic biomarker data according to the cell information (single cell or mixed cell) during data collection, the batch removal method is selectively used. For example, single-cell RNA sequencing data (such as on the 10XGenomics platform) requires a single-cell specific tool (such as Harmony or SCTransform of Seurat), while mixed cell data (such as bulk RNA-seq) can use traditional methods (such as ComBat). This avoids over- or under-correction caused by a "one-size-fits-all" method. Through data dimension threshold division, the mixed cell data is further subdivided. High-dimensional data (such as whole-genome methylation data) uses a dimension reduction-based method (such as BBKNN), and low-dimensional data (such as targeted proteome data) uses a statistical model (such as ComBat) to ensure that data of different complexities are effectively corrected. By dynamically selecting the batch effect removal method, it is ensured that biomarker data from different sources (such as the TCGA database and local experimental data) have a unified dimension. For example, in lung cancer research, after batch effect removal of single-cell sequencing data (GEO dataset) and serum proteome data (local detection), the association between SAA1 gene expression and protein level can be directly compared.

[0085] In this embodiment, in the data dimension threshold, the gene number threshold is 10,000, and the protein or metabolite threshold is 500. If the number of genes in the biological mixed cell data exceeds 10,000 or the number of proteins or metabolites is higher than 500, then this data is high-dimensional biological mixed cell data. If the number of genes in the biological mixed cell data is lower than 10,000 or the number of proteins or metabolites is lower than 500, then this data is low-dimensional biological mixed cell data;

[0086] It can be understood that for high-dimensional data, direct operation will cause a huge system load, that is

[0087] Directly processing a 10,000×10,000 covariance matrix requires approximately 10 8In the second operation, the memory usage exceeds the capacity of a general server (e.g., 10,000×10,000×4 bytes ≈ 400GB).

[0088] Therefore, the batch effect removal tools of Seurat and Harmony are selected, and through the dimensionality reduction optimization algorithm, directly processing high-dimensional matrices is avoided (e.g., Harmony is 2-3 times faster than Seurat in 10x Genomics data).

[0089] For low-dimensional data, it acts directly on the original dimensions, relying on the local similarity between variables (such as MNN) or linear models (such as ComBat). If the number of variables > 500, the computational complexity will increase exponentially (time complexity O(n 2 ))

[0090]

[0091] Based on the batch effect removal method, the dynamic biomarker data is processed to obtain biomarker processed data;

[0092] Specifically, based on the batch effect removal method, the dynamic biomarker data is processed to obtain biomarker processed data, which specifically includes:

[0093] Based on the batch effect removal method, the batch effect of the dynamic biomarker data is removed to obtain the initial biomarker data, and the initial biomarker data includes single-cell biological processed data, high-dimensional processed data of biological mixed cells, and low-dimensional processed data of biological mixed cells;

[0094] Each type of initial biomarker data is compared with the corresponding dynamic biomarker data to obtain the number of differential genes and the overlap rate;

[0095] Based on the batch effect removal method, the threshold of the number of differential genes and the overlap rate threshold corresponding to each batch effect removal method are obtained;

[0096] The number of differential genes and the overlap rate corresponding to each type of initial biomarker data are compared with the threshold of the number of differential genes and the overlap rate threshold corresponding to the batch effect removal method corresponding to each type of initial biomarker data to determine whether the batch effect removal effect of the initial biomarker data meets the analysis requirements;

[0097] Among them, if the number of differential genes corresponding to the initial biomarker data exceeds the threshold of the number of differential genes or the overlap rate exceeds the overlap rate threshold, the batch effect removal effect of the initial biomarker data does not meet the actual requirements, and the batch effect removal method is reselected to perform batch effect removal on the dynamic biomarker data;

[0098] If the number of differential genes corresponding to the initial biomarker data does not exceed the differential gene number threshold and the overlap rate does not exceed the overlap rate threshold, the initial biomarker data is used as the biomarker processing data.

[0099] In this solution, through the threshold comparison of the number of differential genes and the overlap rate, the quantification evaluation of the batch effect removal effect is realized, ensuring that the batch effect removal method does not affect the accuracy and stability of the data, and avoiding misjudgment in subsequent data analysis.

[0100]

[0101] Obtain biomarker historical data;

[0102] According to the biomarker historical data and the batch effect removal method, obtain the biomarker reference data;

[0103] Specifically, according to the biomarker historical data and the batch effect removal method, obtaining the biomarker reference data specifically includes:

[0104] According to the biomarker historical data, based on data attribute recognition, classify the biomarker historical data to obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data;

[0105] Based on the batch effect removal method corresponding to each dynamic biomarker data, using the data classification result as the matching standard, perform batch effect removal on the biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data to obtain the biomarker historical processing data;

[0106] Obtain target disease information;

[0107] Based on the target disease information, obtain the biomarker characteristic data information corresponding to each target disease, and the biomarker characteristic data information includes biomarker data type information; for example, in the study of lung cancer metastasis, the DNB module can distinguish the core genes (such as SAA1) of different metastasis paths;

[0108] Aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, and obtain the biomarker characteristic set;

[0109] Use the biomarker historical processing data corresponding to the biomarker data types in the biomarker characteristic set as the biomarker reference data.

[0110] In this solution, through data attribute recognition, the historical biomarker data is classified in exactly the same way as the dynamic biomarker data, ensuring the same data classification method, providing a data basis for subsequent removal of batch effects from the historical data. By using the batch effect removal method corresponding to each dynamic biomarker data and taking the data classification result as the matching criterion, batch effects are removed from the historical single-cell biomarker data, the historical high-dimensional mixed-cell biomarker data, and the historical low-dimensional mixed-cell biomarker data. (For example, if Harmony is selected for batch effect removal of the high-dimensional mixed-cell biomarker data in the dynamic biomarker data, then Harmony is also selected for batch effect removal of the historical high-dimensional mixed-cell biomarker data in the historical biomarker data), reducing data errors caused by different data processing steps, ensuring data accuracy, and using the historical biomarker processing data corresponding to the biomarker data types in the biomarker feature set as the biomarker reference data, and only using the biomarker reference data as the subsequent data comparison standard, avoiding directly comparing all the historical biomarker processing data with the biomarker processing data, reducing system load, and improving data analysis efficiency.

[0111] Compare the biomarker processing data with the biomarker reference data to obtain biomarker abnormal data, where the biomarker abnormal data represents the data in the biomarker processing data that is different from the biomarker reference data;

[0112] Obtain target disease information, where the target disease information includes the target disease type information and pathological information of the dynamic biomarker analysis;

[0113] Obtain target characteristic disease information based on the biomarker abnormal data and the target disease information;

[0114] Specifically, obtaining target characteristic disease information based on the biomarker abnormal data and the target disease information specifically includes:

[0115] Based on the target disease information and through target disease pathological analysis, obtain the target disease abnormal data type information corresponding to each target disease, where the target disease abnormal data type information represents the types of biomarker data affected by the target disease pathological process;

[0116] Obtain the biomarker abnormal data type information based on the biomarker abnormal data;

[0117] Compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease to obtain the target characteristic disease information;

[0118] Among them, if the biomarker abnormal data type information contains the target disease abnormal data type information corresponding to any one of the target diseases, then this type of target disease is the target characteristic disease.

[0119] In this solution, by comparing the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease, further screening is performed on the target diseases to be analyzed, eliminating impossible disease types, and obtaining the target characteristic disease information, thereby improving the efficiency of subsequent disease risk analysis.

[0120] Based on the target characteristic disease information, analyze the biomarker abnormal data to obtain the disease risk information.

[0121] Specifically, based on the target characteristic disease information, analyze the biomarker abnormal data to obtain the disease risk information, which specifically includes:

[0122] According to the target characteristic disease information, obtain the target disease abnormal data type information corresponding to the target characteristic disease;

[0123] Based on the pathological analysis of the target disease, taking the disease influence mode as the benchmark, classify the target disease abnormal data types, regarding the target disease abnormal data types directly affected by the disease as the directly affected abnormal data types, and regarding the target disease abnormal data types indirectly affected by the disease as the indirectly affected abnormal data types, to obtain the target disease abnormal data type classification information;

[0124] According to the target characteristic disease information, based on the experimental database, obtain the biomarker pathological experimental data;

[0125] Based on the dynamic network biomarker (DNB) algorithm, taking the biomarker pathological experimental data corresponding to the directly affected abnormal data types as the core nodes, and taking the biomarker pathological experimental data corresponding to the indirectly affected abnormal data types as the regulatory edges, construct a disease critical state network;

[0126] Based on the disease critical state network, analyze the biomarker abnormal data to obtain the disease stage information corresponding to each target characteristic disease;

[0127] Aggregate the disease stage information corresponding to each target characteristic disease to obtain the disease risk information.

[0128] In this solution, a disease critical state network is constructed by taking the pathological experimental data of biomarkers corresponding to abnormal data types that directly affect as the core nodes and the pathological experimental data of biomarkers corresponding to abnormal data types that indirectly affect as the regulatory edges. Based on the disease critical state network, the abnormal biomarker data is analyzed to obtain the disease stage information corresponding to each target characteristic disease, and the disease stage information corresponding to each target characteristic disease is aggregated to obtain the disease risk information.

[0129] Apply the dynamic network biomarker (DNB) algorithm, taking the directly affected data as the core nodes (such as the SAA1 protein in lung cancer metastasis) and the indirectly affected data as the regulatory edges to construct a disease critical state network. When the network correlation suddenly increases and the variance within the module exceeds the threshold (such as Z-score > 2), it indicates that the system enters the phase transition critical point.

[0130] Low-risk stage: The fluctuation of the directly affected data is within the baseline range (such as gene expression Z-score < 1.5), and the indirectly affected data (such as IL-6) increases slightly.

[0131] Medium-risk stage: The directly affected data shows periodic abnormalities (such as HER2 expression fluctuations), and the indirect data (such as circulating tumor DNA) continuously deviates from the normal value by 1.5 - 2 times.

[0132] High-risk stage: The directly affected data undergoes irreversible mutations (such as PIK3CA mutations), and the importance of the indirect data is significantly increased (such as AUC of GDF15 > 0.85).

[0133] Refer to Figure 5 As shown, further, in combination with the above-mentioned method for analyzing dynamic biomarkers based on big data, a system for analyzing dynamic biomarkers based on big data is proposed, including:

[0134] The main control module is used to compare the number of differential genes and the overlap rate corresponding to the initial data of each biomarker with the threshold of the number of differential genes and the overlap rate corresponding to the batch effect removal method for the initial data of each biomarker, determine whether the batch effect removal effect of the biomarker initial data meets the analysis requirements, compare the types of biomarker abnormal data with the types of target disease abnormal data corresponding to each target disease, obtain the target characteristic disease information, take the types of target disease abnormal data directly affected by the disease as the types of directly affected abnormal data, take the types of target disease abnormal data indirectly affected by the disease as the types of indirectly affected abnormal data, obtain the classification information of the types of target disease abnormal data, and based on the dynamic network biomarker (DNB) algorithm, use the biomarker pathological experimental data corresponding to the types of directly affected abnormal data as the core nodes and the biomarker pathological experimental data corresponding to the types of indirectly affected abnormal data as the regulatory edges to construct a disease critical state network, analyze the biomarker abnormal data, and obtain the disease stage information corresponding to each target characteristic disease;

[0135] The information acquisition module is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, biomarker historical data, target disease information, the types of target diseases for dynamic biomarker analysis, and pathological information;

[0136] The data processing module is used to remove the batch effect from the dynamic biomarker data to obtain the biomarker initial data, compare each biomarker initial data with the corresponding dynamic biomarker data to obtain the number of differential genes and the overlap rate, classify the biomarker historical data based on data attribute recognition according to the biomarker historical data, obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain the biomarker characteristic data information corresponding to each target disease based on the target disease information, aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, obtain the biomarker characteristic set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker characteristic set as the biomarker reference data;

[0137] The display module interacts with the main control module to output and display biomarker processing data, biomarker reference data, biomarker abnormal data, target characteristic disease information, and disease risk information.

[0138] The main control module specifically includes:

[0139] A control unit, which is used to compare the types of biomarker abnormal data information with the types of target disease abnormal data information corresponding to each target disease, obtain target characteristic disease information, take the types of target disease abnormal data directly affected by the disease as the directly affected abnormal data types, take the types of target disease abnormal data indirectly affected by the disease as the indirectly affected abnormal data types, obtain the classification information of the types of target disease abnormal data, and based on the dynamic network biomarker (DNB) algorithm, use the biomarker pathological experimental data corresponding to the directly affected abnormal data types as the core nodes and the biomarker pathological experimental data corresponding to the indirectly affected abnormal data types as the regulatory edges to construct a disease critical state network, analyze the biomarker abnormal data, and obtain the disease stage information corresponding to each target characteristic disease;

[0140] An information receiving unit, which interacts with the information acquisition module and the data processing module and is used to receive data and transmit it to the judgment unit;

[0141] A judgment unit, which is used to compare the number of differential genes and the overlap rate corresponding to each biomarker initial data with the threshold of the number of differential genes and the overlap rate corresponding to the batch effect removal method corresponding to each biomarker initial data, and judge whether the batch effect removal effect of the biomarker initial data meets the analysis requirements.

[0142] An information acquisition module, specifically including:

[0143] A first acquisition unit, which is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, and biomarker historical data;

[0144] A second acquisition unit, which is used to acquire target disease information, the types of target diseases for dynamic biomarker analysis, and pathological information.

[0145] A data processing module, specifically including:

[0146] A first data processing unit, which is used to perform batch effect removal on the dynamic biomarker data, obtain biomarker initial data, compare each biomarker initial data with the corresponding dynamic biomarker data, and obtain the number of differential genes and the overlap rate;

[0147] A second data processing unit, which is configured to classify biomarker historical data based on data attribute recognition according to the biomarker historical data, obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain biomarker feature data information corresponding to each target disease based on the target disease information, aggregate the biomarker feature data information corresponding to each target disease, remove duplicate data types, obtain a biomarker feature set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker feature set as biomarker reference data.

[0148] In summary, the advantages of the present invention are as follows: By using multi-source dynamic biomarker data such as RNA sequencing, serum proteome, and genome, and based on the intelligent matching of data attributes to remove batch effects, the data comparability is ensured, the reliability of subsequent analysis is improved, and the batch removal effect is verified by the number of differential genes and the overlap rate, ensuring the stability of the data. By aggregating the biomarker feature data information corresponding to each target disease, removing duplicate data types, obtaining a biomarker feature set, and selecting biomarker data types related to the target disease through the biomarker feature set, the data calculation amount is reduced, and the data analysis efficiency is improved. By constructing a regulatory network including molecular interaction relationships, a more comprehensive risk assessment basis is provided for clinical practice.

[0149] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for analyzing dynamic biomarkers based on big data, characterized in that, Including: Obtain dynamic biomarker data, where the dynamic biomarker data includes RNA sequencing data, serum proteome data, and genomic data; Based on the dynamic biomarker data and through data attribute identification, obtain the batch effect removal method corresponding to each type of dynamic biomarker data; Based on the batch effect removal method, process the dynamic biomarker data to obtain biomarker processed data; Obtain biomarker historical data; Based on the biomarker historical data and the batch effect removal method, obtain biomarker baseline data; Compare the biomarker processed data with the biomarker baseline data to obtain biomarker abnormal data, where the biomarker abnormal data represents the data in the biomarker processed data that is different from the biomarker baseline data; Obtain target disease information, where the target disease information includes the target disease type information and pathological information for dynamic biomarker analysis; Based on the biomarker abnormal data and the target disease information, obtain target characteristic disease information; Based on the target characteristic disease information, analyze the biomarker abnormal data to obtain disease risk information.

2. The dynamic biomarker analysis method based on big data according to claim 1, characterized in that The step of, based on the dynamic biomarker data and through data attribute identification, obtaining the batch effect removal method corresponding to each type of dynamic biomarker data specifically includes: Based on the dynamic biomarker data, obtain the data attribute corresponding to each type of dynamic biomarker data, that is, the cell information during data collection; Based on the data attribute corresponding to each type of dynamic biomarker data and using single-cell data and mixed-cell data as the discrimination criterion, classify the dynamic biomarker data to obtain biological single-cell data and biological mixed-cell data; Based on the batch effect removal tool and using the batch effect removal of single-cell data as the benchmark, obtain the batch effect removal method corresponding to the biological single-cell data; Based on the analysis requirements of the batch effect removal tool, obtain the data dimension threshold; According to the data dimension threshold, divide the dimension of the biological mixed-cell data to obtain biological mixed-cell high-dimensional data and biological mixed-cell low-dimensional data; Based on the batch effect removal tool, select the batch effect removal methods corresponding to the biological mixed-cell high-dimensional data and the biological mixed-cell low-dimensional data respectively; Based on the batch effect removal method corresponding to the biological single-cell data, and the batch effect removal methods corresponding to the biological mixed-cell high-dimensional data and the biological mixed-cell low-dimensional data, obtain the batch effect removal method.

3. A method for analyzing dynamic biomarkers based on big data according to claim 1, characterized in that The step of, based on the batch effect removal method, processing the dynamic biomarker data to obtain biomarker processed data specifically includes: Based on the batch effect removal method, perform batch effect removal on the dynamic biomarker data to obtain biomarker initial data, where the biomarker initial data includes biological single-cell processed data, biological mixed-cell high-dimensional processed data, and biological mixed-cell low-dimensional processed data; Compare each biomarker initial data with the corresponding dynamic biomarker data to obtain the number of differential genes and the overlap rate; Based on the batch effect removal method, obtain the differential gene number threshold and overlap rate threshold corresponding to each batch effect removal method; Compare the number of differential genes and the overlap rate corresponding to the initial data of each biomarker with the threshold of the number of differential genes and the overlap rate threshold corresponding to the batch effect removal method for the initial data of each biomarker to determine whether the batch effect removal effect of the initial data of this biomarker meets the analysis requirements; Among them, if the number of differential genes corresponding to the initial data of the biomarker exceeds the threshold of the number of differential genes or the overlap rate exceeds the overlap rate threshold, the batch effect removal effect of the initial data of this biomarker does not meet the actual requirements, and reselect the batch effect removal method to perform batch effect removal on the dynamic biomarker data; If the number of differential genes corresponding to the initial data of the biomarker does not exceed the threshold of the number of differential genes and the overlap rate does not exceed the overlap rate threshold, then use the initial data of the biomarker as the processed data of the biomarker.

4. A method for analyzing dynamic biomarkers based on big data according to claim 1, characterized in that The obtaining of the biomarker reference data according to the biomarker historical data and the batch effect removal method specifically includes: According to the biomarker historical data, based on data attribute recognition, classify the biomarker historical data to obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data; Based on the batch effect removal method corresponding to each dynamic biomarker data, using the data classification result as the matching criterion, perform batch effect removal on the biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data to obtain the biomarker historical processed data; Obtain the target disease information; Based on the target disease information, obtain the biomarker characteristic data information corresponding to each target disease, and the biomarker characteristic data information includes biomarker data type information; Aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, and obtain the biomarker characteristic set; Use the biomarker historical processed data corresponding to the biomarker data types in the biomarker characteristic set as the biomarker reference data.

5. A method for analyzing dynamic biomarkers based on big data according to claim 1, characterized in that The obtaining of the target characteristic disease information according to the biomarker abnormal data and the target disease information specifically includes: According to the target disease information, based on the pathological analysis of the target disease, obtain the target disease abnormal data type information corresponding to each target disease, and the target disease abnormal data type information represents the biomarker data types affected by the pathological process of the target disease; According to the biomarker abnormal data, obtain the biomarker abnormal data type information; Compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease to obtain the target characteristic disease information; Among them, if the biomarker abnormal data type information contains the target disease abnormal data type information corresponding to any one target disease, then this target disease is the target characteristic disease.

6. The dynamic biomarker analysis method based on big data according to claim 1, characterized in that The analysis of the biomarker abnormal data based on the target characteristic disease information to obtain the disease risk information specifically includes: According to the target characteristic disease information, obtain the target disease abnormal data type information corresponding to the target characteristic disease; Based on the pathological analysis of the target disease, taking the disease impact mode as the benchmark, classify the types of abnormal data of the target disease. The types of abnormal data of the target disease directly affected by the disease are used as the types of directly affected abnormal data, and the types of abnormal data of the target disease indirectly affected by the disease are used as the types of indirectly affected abnormal data, so as to obtain the classification information of the types of abnormal data of the target disease; According to the target characteristic disease information, based on the experimental database, obtain the biomarker pathological experimental data; Based on the dynamic network biomarker (DNB) algorithm, using the biomarker pathological experimental data corresponding to the types of directly affected abnormal data as the core nodes and the biomarker pathological experimental data corresponding to the types of indirectly affected abnormal data as the regulatory edges, construct the disease critical state network; Based on the disease critical state network, analyze the biomarker abnormal data to obtain the disease stage information corresponding to each target characteristic disease; Aggregate the disease stage information corresponding to each target characteristic disease to obtain the disease risk information.

7. A dynamic biomarker analysis system based on big data for implementing the analysis method according to any one of claims 1-6, characterized in that, Including: The main control module is used to compare the number of differential genes and the overlap rate corresponding to each biomarker initial data with the threshold of the number of differential genes and the overlap rate corresponding to the batch effect removal method corresponding to each biomarker initial data to determine whether the batch effect removal effect of the biomarker initial data meets the analysis requirements, compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease to obtain the target characteristic disease information, use the types of abnormal data of the target disease directly affected by the disease as the types of directly affected abnormal data, use the types of abnormal data of the target disease indirectly affected by the disease as the types of indirectly affected abnormal data, obtain the classification information of the types of abnormal data of the target disease, based on the dynamic network biomarker (DNB) algorithm, use the biomarker pathological experimental data corresponding to the types of directly affected abnormal data as the core nodes and the biomarker pathological experimental data corresponding to the types of indirectly affected abnormal data as the regulatory edges, construct the disease critical state network, and analyze the biomarker abnormal data to obtain the disease stage information corresponding to each target characteristic disease; The information acquisition module is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, biomarker historical data, target disease information, the types of target diseases for dynamic biomarker analysis, and pathological information; A data processing module, which is used to remove batch effects from dynamic biomarker data, obtain initial biomarker data, compare each initial biomarker data with the corresponding dynamic biomarker data, obtain the number of differential genes and the overlap rate, classify the biomarker historical data based on data attribute recognition according to the biomarker historical data, obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain the biomarker characteristic data information corresponding to each target disease based on the target disease information, aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, obtain a biomarker characteristic set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker characteristic set as biomarker reference data; A display module, which interacts with the main control module to output and display biomarker processing data, biomarker reference data, biomarker abnormal data, target characteristic disease information, and disease risk information.

8. A dynamic biomarker analysis system based on big data according to claim 7, characterized in that The main control module specifically includes: A control unit, which is used to compare the biomarker abnormal data type information with the target disease abnormal data type information corresponding to each target disease, obtain target characteristic disease information, use the target disease abnormal data types directly affected by the disease as directly affected abnormal data types, use the target disease abnormal data types indirectly affected by the disease as indirectly affected abnormal data types, obtain target disease abnormal data type classification information, and based on the dynamic network biomarker (DNB) algorithm, use the biomarker pathological experiment data corresponding to the directly affected abnormal data types as core nodes and the biomarker pathological experiment data corresponding to the indirectly affected abnormal data types as regulatory edges to construct a disease critical state network, analyze the biomarker abnormal data, and obtain the disease stage information corresponding to each target characteristic disease; An information receiving unit, which interacts with the information acquisition module and the data processing module to receive data and transmit it to the judgment unit; A judgment unit, which is used to compare the number of differential genes and the overlap rate corresponding to each initial biomarker data with the differential gene number threshold and overlap rate threshold corresponding to the batch effect removal method corresponding to each initial biomarker data to judge whether the batch effect removal effect of the initial biomarker data meets the analysis requirements.

9. The dynamic biomarker analysis system based on big data according to claim 7, characterized in that, The information acquisition module specifically includes: A first acquisition unit, which is used to acquire dynamic biomarker data, RNA sequencing data, serum proteome data, genomic data, and biomarker historical data; A second acquisition unit, which is used to acquire target disease information, target disease type information for dynamic biomarker analysis, and pathological information.

10. A dynamic biomarker analysis system based on big data according to claim 7, characterized in that, The data processing module specifically includes: The first data processing unit, which is used to remove batch effects from dynamic biomarker data, obtain initial biomarker data, compare each initial biomarker data with the corresponding dynamic biomarker data, and obtain the number of differential genes and the overlap rate. The second data processing unit, which is used to classify biomarker historical data based on data attribute recognition according to biomarker historical data, obtain biological historical single-cell data, biological historical mixed-cell high-dimensional data, and biological historical mixed-cell low-dimensional data, obtain biomarker characteristic data information corresponding to each target disease based on target disease information, aggregate the biomarker characteristic data information corresponding to each target disease, remove duplicate data types, obtain a biomarker characteristic set, and use the biomarker historical data corresponding to the biomarker data types in the biomarker characteristic set as biomarker reference data.

Citation Information

Cited By

  • Precise enclosure method for stem cell detection and analysis

    CN121837206A