Multi-biomarker medical examination data fusion and analysis method
By standardizing and semantically embedding biomarker data, quantifying biomarker coupling relationships, calculating change rates and risk scores, this approach addresses the shortcomings of traditional biomarker data fusion and analysis, enabling more accurate disease risk prediction and personalized treatment support.
Patent Information
- Application Number
- CN202511494835.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional medical laboratory data fusion and analysis methods lack the ability to combine the clinical background information of biomarkers with the data itself, fail to effectively enhance the clinical interpretability of the data through semantic embedding, and fail to consider the nonlinear coupling relationship and dynamic evolution between biomarkers. This results in an inability to accurately capture the complex interactions between biomarkers, especially during changes in disease state, and an inability to respond promptly to acute clinical changes. Furthermore, they give little consideration to the volatility and risk stability of biomarkers in risk assessment, and fail to provide dynamic and personalized disease prediction and decision support.
By collecting biomarker data, constructing structured datasets, and performing standardization and semantic embedding, the nonlinear coupling relationships between biomarkers are quantified, a coupling strength function is established, the rate of change and risk score of biomarkers are calculated, and combined with volatility analysis, disease stability is assessed, providing personalized disease prediction and decision support.
It improves the clinical interpretability of biomarker data, accurately captures the interactions and dynamic changes between biomarkers, provides timely disease progression analysis, and helps doctors develop personalized treatment plans.
Smart Images

Figure CN120998515A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical informatics, and in particular to a method for the fusion and analysis of medical test data using multiple biomarkers. Background Technology
[0002] With the continuous development of precision medicine, the demand for early diagnosis, accurate prediction, and personalized treatment of diseases in the medical field is increasing. Multi-biomarker detection technology has become an indispensable part of medical research and clinical diagnosis. The combined detection of multiple biomarkers (such as gene expression, protein levels, and metabolite concentrations) can provide more comprehensive disease information than a single biomarker, and is particularly valuable in the diagnosis and prognostic assessment of complex diseases such as cancer, cardiovascular disease, and neurodegenerative diseases.
[0003] However, with the continuous advancement of medical testing technology, biomarker data from different platforms exhibit high heterogeneity. These data differ significantly in units, measurement methods, and reference intervals, directly impacting the effective fusion and comprehensive analysis of multi-biomarker data. Therefore, standardizing and rationally integrating data from different testing platforms, and combining this with accurate analysis in the context of the biomarkers' clinical background, has become a major challenge in medical data processing.
[0004] Traditional medical laboratory data fusion and analysis methods suffer from the following technical problems: they lack the ability to combine the clinical background information of biomarkers with the data itself, fail to effectively enhance the clinical interpretability of the data through semantic embedding, and limit the practical application value of the data; they fail to consider the nonlinear coupling relationship and dynamic evolution between biomarkers, resulting in an inability to accurately capture the complex interactions between biomarkers, especially during changes in disease state, and an inability to respond promptly to acute clinical changes; and they give little consideration to the volatility and risk stability of biomarkers in risk assessment, failing to provide dynamic and personalized disease prediction and decision support. Summary of the Invention
[0005] This invention provides a method for the fusion and analysis of multi-biomarker medical laboratory data to address the technical problems of traditional methods that lack the ability to combine the clinical background information of biomarkers with the data itself, fail to effectively enhance the clinical interpretability of the data through semantic embedding, thus limiting the practical application value of the data; fail to consider the nonlinear coupling relationships and dynamic evolution between biomarkers, resulting in the inability to accurately capture the complex interactions between biomarkers, especially during changes in disease state, and fail to respond promptly to acute clinical changes; and rarely consider the volatility and risk stability of biomarkers in risk assessment, failing to provide dynamic and personalized disease prediction and decision support.
[0006] The present invention provides a method for fusing and analyzing medical laboratory data using multiple biomarkers, specifically comprising the following technical solutions: A method for fusing and analyzing medical laboratory data using multiple biomarkers includes the following steps: S1. Collect biomarker data and construct a structured biomarker dataset; standardize and semantically embed the data items in the structured biomarker dataset to obtain the semantic embedding vector of the biomarker; based on the semantic embedding vector of the biomarker, quantify the nonlinear coupling relationship between different biomarkers to obtain the coupling strength between biomarkers. S2. Based on the coupling strength between biomarkers, the changes of biomarkers over time are modeled to obtain the rate of change of biomarkers; based on the rate of change of biomarkers and the coupling strength between biomarkers, the risk score of biomarkers is calculated, and the stability of the disease is assessed through risk score volatility analysis, thereby obtaining the final risk score.
[0007] Preferably, S1 specifically includes: The data items in the structured biomarker dataset include the name, value, unit, and reference range of the biomarker to which the data item belongs; the values of the biomarkers are normalized according to their reference ranges to obtain standardized biomarker data, and a standardized biomarker dataset is constructed.
[0008] Preferably, S1 specifically includes: Medical label information of biomarkers is introduced, and semantic embedding vectors of medical label information are obtained; based on standardized biomarker data, combined with semantic embedding vectors of medical label information, semantic embedding vectors of biomarkers are obtained.
[0009] Preferably, S1 specifically includes: By calculating the differences between the semantic embedding vectors of different biomarkers and introducing the coupling coefficient between biomarkers, the relative volatility between biomarkers is quantified. Combined with the periodic changes between biomarkers, a coupling strength function is constructed to obtain the coupling strength between biomarkers.
[0010] Preferably, S2 specifically includes: Based on the coupling strength between biomarkers and the semantic embedding vector of biomarkers, and by introducing the decay coefficient of biomarkers, an evolution rate equation for biomarkers is constructed to obtain the rate of change of biomarkers.
[0011] Preferably, S2 specifically includes: Based on the rate of change of biomarkers and the coupling strength between biomarkers, and by introducing the weights and coupling strength coefficients of biomarkers, and combining integral calculations, a risk score for the biomarkers is obtained; the risk scores of the biomarkers are then weighted to obtain a comprehensive risk score.
[0012] Preferably, S2 specifically includes: Based on the risk score of the biomarker, and combined with the average risk score of the biomarker, the volatility measure of the risk score of the biomarker is obtained through integral calculation; when the volatility measure of the risk score of the biomarker exceeds the preset threshold, it indicates that the current risk status of the biomarker is unstable and needs to be closely monitored.
[0013] Preferably, S2 specifically includes: Based on the risk score volatility measurement of biomarkers, the overall risk score is adjusted for stability to obtain the final risk score, which helps doctors determine the patient's disease status and develop personalized treatment plans.
[0014] The beneficial effects of the technical solution of the present invention are: 1. By integrating the medical label information of biomarkers (such as biomarker category, related diseases, etc.) into the data processing process, each biomarker not only reflects its numerical information, but also includes its clinical background information. Semantic embedding can effectively improve the clinical interpretability of biomarker data, ensure that the data processing results have practical significance in medical decision-making, and provide doctors with more targeted and interpretable analysis results.
[0015] 2. By establishing a coupling strength function between biomarkers, this invention can more accurately capture the complex interactions and effects between biomarkers, especially the mutual influence between biomarkers during disease development. It can reflect the dynamic changes in biological systems, including periodic changes and relative fluctuations, thereby making disease risk prediction more accurate and dynamic.
[0016] 3. By introducing the evolution rate equation and decay coefficient of biomarkers, this invention can dynamically track the change process of biomarkers over time. It not only considers the interaction between biomarkers, but also accurately reflects the natural decay effect of biomarkers over time, providing more timely and prospective disease progression analysis for clinical monitoring, and helping doctors to assess patients' health status in real time.
[0017] 4. This invention comprehensively considers the rate of change of biomarkers and the coupling strength between biomarkers to calculate the risk score of each biomarker; by assessing the volatility of the risk scores of biomarkers, the stability of the disease is further analyzed, thereby making accurate predictions of the patient's risk status; finally, by combining comprehensive risk scores and stability analysis, it provides more comprehensive and accurate support for clinical decision-making and helps doctors develop personalized treatment plans. Attached Figure Description
[0018] Figure 1 This is a flowchart of a method for fusing and analyzing medical test data using multiple biomarkers, as described in this invention. Detailed Implementation
[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0021] The following description, in conjunction with the accompanying drawings, details the specific scheme of a method for fusing and analyzing medical laboratory data using multiple biomarkers provided by this invention.
[0022] See attached document Figure 1 The diagram illustrates a flowchart of a method for fusing and analyzing medical laboratory data using multiple biomarkers, provided by an embodiment of the present invention. The method includes the following steps: S1. Collect biomarker data and construct a structured biomarker dataset; standardize and semantically embed the data items in the structured biomarker dataset to obtain the semantic embedding vector of the biomarker; based on the semantic embedding vector of the biomarker, quantify the nonlinear coupling relationship between different biomarkers to obtain the coupling strength between biomarkers. Biomarker data were collected from multiple biomarker detection platforms, including gene expression data, proteomics data, metabolomics data, and routine medical test data. The gene expression data was obtained through RNA sequencing-based gene expression detection; the proteomics data was obtained through mass spectrometry; the metabolomics data was obtained through nuclear magnetic resonance or liquid chromatography; and the routine medical test data was obtained through routine medical testing. The methods used to collect the biomarker data are well-known to those skilled in the art and will not be elaborated upon here. The collected biomarker data underwent format standardization and unit conversion to generate a structured biomarker dataset. Each data item in the biomarker dataset includes the name, value, unit, and reference range of its corresponding biomarker. The methods used for format standardization and unit conversion are well-known to those skilled in the art and will not be elaborated upon here. Due to differences in measurement units and scales across platforms, and the varying clinical semantics of biomarkers in different medical fields, it is necessary to standardize and semantically embed the structured data (i.e., data items in the structured biomarker dataset) to ensure consistency and medical interpretability of biomarkers in data analysis. Specifically, the values of biomarkers are normalized according to their reference intervals using methods such as Z-scores to unify them to the range [0,1], resulting in standardized biomarker data, and a standardized biomarker dataset is constructed. ,in, Indicates the first A standardized biomarker data, It refers to the quantity of standardized biomarker data.
[0023] Since the clinical significance of medical data requires the integration of medical labels for biomarkers, such as proteins, genes, and metabolites, the medical label information of biomarkers is incorporated into the data processing. The first step is to determine the medical label set for each biomarker. This represents the medical background of the biomarker, such as its category and related diseases. The medical label information for each biomarker can be obtained from medical domain knowledge bases such as LOINC (Logical Naming and Coding System for Observational Indicators) and UMLS (Unified Medical Language System). To incorporate this medical label information into the biomarker's data representation, semantic embedding is used to obtain the biomarker's semantic embedding vector. Specifically, the semantic embedding vector for each biomarker is calculated using the following formula: ; in, It is the first The semantic embedding vector of the biomarker represents the _th biomarker. The semantic space representation of standardized biomarker data; For the first Biomarkers for medical labeling information The weight of represents the weight of the th Standardized biomarker data for medical labeling information The contribution level, with a value range of This is obtained using the existing TF-IDF method. Specifically, the medical label information corresponding to each biomarker is regarded as a term set, all biomarkers are used to form a document set, and the term frequency (TF) and inverse document frequency (IDF) are calculated. The formula is as follows: ,in, For medical label information In the Frequency of occurrence of biomarkers This indicates that it contains medical label information. The number of biomarkers; Medical label information The semantic embedding vectors, obtained from a medical domain knowledge base, represent medical label information. The semantic information of medical tags is extracted using medical knowledge bases such as LOINC or UMLS, and then embedded using word2vec or GloVe models. These are technical methods well known to those skilled in the art and will not be elaborated here. To establish nonlinear coupling relationships between biomarkers and capture the dynamic interaction effects among different biomarkers, a coupling strength function is needed to represent the mutual influence between biomarkers at a specific time point. This function is used to analyze the evolution of biomarkers and more accurately capture changes in disease states. The coupling strength between biomarkers obtained based on the coupling strength function is not only affected by the differences between biomarkers but also by factors such as the periodicity and relative volatility of the biomarkers. The mathematical formula for the coupling strength function is as follows: ; in, Indicates the first The first biomarker and the first A biomarker in time The coupling strength; The first is obtained by using a smooth S-curve. The first biomarker and the first Differences between semantic embedding vectors of individual biomarkers This is converted into a tractable coupling strength to represent the relative volatility between biomarkers; It is the first The first biomarker and the first The coupling coefficient between biomarkers, used to control the strength of their influence, is obtained using the least squares method and its value ranges from [value range missing]. Specifically, at different times Based on the semantic embedding vectors of biomarkers, the difference between two biomarkers is calculated. This difference is then substituted into the coupling strength function as the independent variable, while the actual coupling strength at the corresponding time point is used as the dependent variable. A parameter fitting model is established, and the coupling coefficient is obtained by minimizing the sum of squared errors between the predicted and actual coupling strengths. The estimated value is obtained by using the least squares method, which is a well-known technique to those skilled in the art and will not be described in detail here. For the first A biomarker in time semantic embedding vector; For the first A biomarker in time semantic embedding vector; It is a periodic correction coefficient used to reflect the periodic synergistic effect of biomarkers under different physiological states. It is obtained by the least squares method and its value ranges from [value missing]. Specifically, periodic components are extracted using Fourier transform, and then the periodic variation terms between biomarkers are substituted into the coupling strength function to construct the fitting target: Thus, the periodic correction coefficient is determined. The value of , where, The Fourier transform is a well-known technique for observing periodic changes in data, and will not be described in detail here. It is a periodic variation term among biomarkers, reflecting the synergistic effect between biomarkers. It simulates the periodic variation characteristics that biomarkers may have and is used to capture the influence of periodic variation on coupling strength when there are similar fluctuation patterns among biomarkers.
[0024] S2. Based on the coupling strength between biomarkers, the changes of biomarkers over time are modeled to obtain the rate of change of biomarkers; based on the rate of change of biomarkers and the coupling strength between biomarkers, the risk score of biomarkers is calculated, and the stability of the disease is assessed through risk score volatility analysis, thereby obtaining the final risk score. The coupling strength function between different biomarkers can effectively describe the complex interaction effects between them. Based on the coupling strength between biomarkers, an evolution rate equation for each biomarker is introduced to establish a dynamic model of its change over time, thereby ensuring effective simulation of the evolution process of biological systems in long-term observation. The specific formula for the evolution rate equation of biomarkers is as follows: ; in, It is the first A biomarker over time The rate of change of represents the th The evolution rate of individual biomarkers; It refers to the quantity of standardized biomarker data; It is the first The decay coefficient of a biomarker, representing the natural decay of the biomarker over time, is estimated using maximum likelihood estimation and has a value range of [value missing]. Specifically, for time series based on semantic embedding vectors of biomarkers, an exponential model is employed. The maximum likelihood function is then solved by maximum likelihood estimation, which is a technique well known to those skilled in the art and will not be elaborated here. Indicates time The semantic embedding vectors of all other biomarkers semantic embedding vectors of current biomarkers The sum of dynamic effects; The natural decay term represents the current biomarker's own decay trend over time. The above biomarker evolution rate equation describes the dynamic interaction effect between biomarkers and the decay effect of biomarkers, reflecting the evolution process of biomarkers over time.
[0025] To assess the disease risk of various biomarkers at different time points and further analyze the stability of disease risk, a risk score for each biomarker is calculated by combining the coupling strength between biomarkers and the rate of change of the biomarkers. This assesses the stability of patient status and provides a basis for disease prognosis. The risk score of a biomarker considers not only the rate of change of the biomarker at the current moment but also its interaction with other biomarkers. The formula for calculating the risk score of a biomarker is as follows: ; in, It is the first A biomarker in time The risk score indicates the risk level of the first The contribution of each biomarker to overall disease risk at the current moment; It is the first The weight of the i-th biomarker represents the weight of the i-th biomarker. The contribution of each biomarker to the overall risk score was calculated using the weighted least squares method. The weighted least squares method is a well-known technique in the art and will not be described in detail here. Indicates the first A biomarker in time The rate of change is used to reflect the acute progression or changes of a disease; This is the integral of the coupling effect between biomarkers, used to reflect the relationship between other biomarkers and the first biomarker. The impact of interactions between individual biomarkers on risk scores; These are the start and end times of the time interval, respectively. The specific values are obtained based on the time period required for observation or data collection, and are not limited here. It is the first The coupling strength coefficient of the first biomarker determines the first The extent to which the interaction between a biomarker and other biomarkers influences the risk score is determined through multiple regression analysis or fitting based on maximum likelihood estimation. The multivariate regression analysis is a technique well-known to those skilled in the art and will not be described in detail here. By calculating the risk score for each biomarker, the role of each biomarker in the disease development process can be quantified, and information support based on changes in biomarkers can be provided, which helps doctors make more accurate judgments based on changes in biomarkers during treatment. Furthermore, the variability of biomarker risk scores over a time interval is calculated to quantitatively describe disease stability; greater variability indicates a more unstable risk status for the current biomarker, and is more likely to indicate acute changes in the disease or fluctuations in the patient's condition. The formula for calculating the variability of biomarker risk scores is as follows: ; in, It is the first The risk score volatility measure of the first biomarker, representing the... A biomarker within a time interval The degree of fluctuation in the risk score within; For the first A biomarker within a time interval The average risk score, representing the first The average risk level of each biomarker over the entire time window; If the variability measure of a biomarker’s risk score exceeds the threshold preset by the rule of thumb based on statistical confidence intervals, the current risk status of the biomarker is considered unstable, suggesting that the disease may be entering an acute phase or that there may be changes in treatment efficacy, requiring close attention. The risk scores of biomarkers are weighted to obtain a comprehensive risk score; the formula for calculating the comprehensive risk score is as follows: ; in, This represents the overall risk score; Indicates the first The weighting coefficients for each biomarker, with values ranging from [value range missing], are [value range missing]. The results were obtained by fitting the data using the least squares method. Furthermore, measures of risk score volatility based on biomarkers Comprehensive risk score Stability adjustments are made to obtain the final risk score, which better captures unstable factors in disease progression, thereby improving the accuracy of the assessment. The final risk score is expressed as: ; in, This indicates the final risk score; It is an adjustment factor used to balance the weights between the comprehensive risk score and the risk score volatility measure. It is optimized through cross-validation, which is a technique well known to those skilled in the art and will not be described in detail here. The final risk score represents a patient's disease risk level and can be used as a support tool for clinical decision-making, helping doctors to assess a patient's disease status and develop personalized treatment plans.
[0026] In summary, a method for the fusion and analysis of medical laboratory data using multiple biomarkers has been developed.
[0027] The order of the embodiments is for illustrative purposes only and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0028] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0029] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for fusing and analyzing medical laboratory data using multiple biomarkers, characterized in that, Includes the following steps: S1. Collect biomarker data and construct a structured biomarker dataset; standardize and semantically embed the data items in the structured biomarker dataset to obtain the semantic embedding vector of the biomarker; based on the semantic embedding vector of the biomarker, quantify the nonlinear coupling relationship between different biomarkers to obtain the coupling strength between biomarkers. S2. Based on the coupling strength between biomarkers, the changes of biomarkers over time are modeled to obtain the rate of change of biomarkers; based on the rate of change of biomarkers and the coupling strength between biomarkers, the risk score of biomarkers is calculated, and the stability of the disease is assessed through risk score volatility analysis, thereby obtaining the final risk score.
2. The method for fusion and analysis of multi-biomarker medical test data according to claim 1, characterized in that, S1 specifically includes: The data items in the structured biomarker dataset include the name, value, unit, and reference range of the biomarker to which the data item belongs; the values of the biomarkers are normalized according to their reference ranges to obtain standardized biomarker data, and a standardized biomarker dataset is constructed.
3. The method for fusion and analysis of multi-biomarker medical test data according to claim 2, characterized in that, S1 specifically includes: Medical label information of biomarkers is introduced, and semantic embedding vectors of medical label information are obtained; based on standardized biomarker data, combined with semantic embedding vectors of medical label information, semantic embedding vectors of biomarkers are obtained.
4. The method for fusion and analysis of multi-biomarker medical test data according to claim 3, characterized in that, S1 specifically includes: By calculating the differences between the semantic embedding vectors of different biomarkers and introducing the coupling coefficient between biomarkers, the relative volatility between biomarkers is quantified. Combined with the periodic changes between biomarkers, a coupling strength function is constructed to obtain the coupling strength between biomarkers.
5. The method for fusion and analysis of multi-biomarker medical test data according to claim 1, characterized in that, S2 specifically includes: Based on the coupling strength between biomarkers and the semantic embedding vector of biomarkers, and by introducing the decay coefficient of biomarkers, an evolution rate equation for biomarkers is constructed to obtain the rate of change of biomarkers.
6. The method for fusion and analysis of multi-biomarker medical test data according to claim 5, characterized in that, S2 specifically includes: Based on the rate of change of biomarkers and the coupling strength between biomarkers, and by introducing the weights and coupling strength coefficients of biomarkers, and combining integral calculations, a risk score for the biomarkers is obtained; the risk scores of the biomarkers are then weighted to obtain a comprehensive risk score.
7. The method for fusion and analysis of multi-biomarker medical test data according to claim 6, characterized in that, S2 specifically includes: Based on the risk score of the biomarker, and combined with the average risk score of the biomarker, the volatility measure of the risk score of the biomarker is obtained through integral calculation; when the volatility measure of the risk score of the biomarker exceeds the preset threshold, it indicates that the current risk status of the biomarker is unstable and needs to be closely monitored.
8. The method for fusion and analysis of multi-biomarker medical test data according to claim 7, characterized in that, S2 specifically includes: Based on the risk score volatility measurement of biomarkers, the overall risk score is adjusted for stability to obtain the final risk score, which helps doctors determine the patient's disease status and develop personalized treatment plans.