Physiological state risk early warning in-vitro detection system and method based on multi-marker conjoint analysis
By using multi-dimensional biomarker synchronous detection and random forest joint inference model, the problems of unreasonable biomarker selection and data processing in existing technologies are solved, enabling early and accurate warning of physiological state risks and improving detection efficiency and signal sensitivity.
Patent Information
- Application Number
- CN202610181765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for physiological state detection suffer from problems such as a lack of systematic selection of biomarkers, simplistic data processing methods, a lack of integration of clinical diagnostic standards into risk warning models, and non-standardized testing procedures. These issues result in low testing efficiency, poor data consistency, and difficulty in achieving accurate risk warnings.
By employing multi-dimensional biomarker synchronous detection technology, combined with multi-biomarker collaborative association rules and clinical diagnostic criteria, a random forest joint inference early warning model is constructed. Through targeted capture, data standardization, outlier removal, and signal-to-noise ratio optimization, a standardized risk early warning fusion dataset is generated to achieve hierarchical identification and verification of physiological state risks.
It enables early and accurate warning of physiological risks such as metabolic abnormalities, inflammatory responses, and organ function decline, improves the sensitivity and specificity of detection signals, and covers physiological state changes caused by the synergistic effects of multiple systems.
Smart Images

Figure CN122067772A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an in vitro detection system and method for early warning of physiological state risks based on joint analysis of multiple biomarkers. Background Technology
[0002] With the aging population and the rising incidence of chronic diseases, early risk warning of physiological conditions has become a core need in the healthcare field. Current mainstream in vitro diagnostic technologies mostly focus on single biomarker detection, such as screening for diabetes risk solely through blood glucose testing or assessing inflammatory status solely through interleukin testing. However, abnormal physiological conditions are often the result of the synergistic effects of multiple systems. Single biomarker detection cannot comprehensively reflect the overall changes in physiological systems, leading to missed diagnoses and misdiagnoses, and thus failing to meet the need for accurate risk warning.
[0003] Some existing technologies attempt to use multi-biomarker joint detection, but significant shortcomings remain: First, biomarker selection lacks systematicity, failing to fully consider the specific expression patterns of each biomarker in different physiological systems, resulting in weak correlation between detection targets and risk types; second, data processing is simplistic, only analyzing or simply overlaying the detection results of each biomarker without considering synergistic associations and metabolic pathway correlations between biomarkers, making it difficult to establish a precise mapping between biomarker characteristics and risk levels; third, the construction of risk warning models lacks deep integration with clinical diagnostic standards, and the model algorithm parameters are not sufficiently optimized, resulting in poor accuracy and reliability of risk grading identification; fourth, the warning results lack a follow-up verification mechanism, failing to review the initial warning data with clinical pathological data, making it difficult to clearly identify risk types, core driving biomarkers, and associated physiological systems, resulting in insufficient practicality and guidance of the warning details.
[0004] In addition, the existing multi-marker detection system lacks standardized design in its detection process. The technical connections between sample preprocessing, marker capture, data fusion, and model building are not smooth, resulting in low detection efficiency and poor data consistency, which further limits its promotion and application in clinical practice.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0007] According to one aspect of this application, an in vitro detection system and method for physiological state risk early warning based on multi-marker joint analysis is provided, comprising: acquiring pretreated serum, urine, and tissue fluid samples; using multi-dimensional marker simultaneous detection technology to target and capture the pretreated serum, urine, and tissue fluid samples, generating raw data of each marker concentration, activity response curve, and molecular binding characteristic information; performing outlier removal, signal-to-noise ratio optimization, and data standardization calibration on the raw detection data; combining multi-marker collaborative association rules, physiological state risk judgment thresholds, and disease diagnosis reference standards to perform feature data fusion processing, establishing a precise mapping relationship between marker features and risk levels, and generating a standardized risk early warning fusion dataset; based on standardization... A risk warning fusion dataset is used to construct a multi-marker joint reasoning early warning model. Using the concentration, activity intensity, and molecular interaction characteristics of core markers as input dimensions and clinical diagnostic criteria as the basis for judgment, the model algorithm parameters are optimized to achieve graded identification of physiological state risks and generate initial risk warning data. Combining the specific expression patterns of each marker in different physiological systems, metabolic pathway correlations, and clinicopathological correlation data, the initial warning data undergoes risk level verification, concentration threshold review, and associated marker combination analysis to clarify the risk type, the time of warning threshold breach, the core driving marker, and the associated physiological system. Finally, the verified data are integrated to generate detailed physiological state risk warning information related to metabolic abnormalities, inflammatory responses, and organ function decline.
[0008] Another aspect of this application is an in vitro detection system for early warning of physiological state risk based on joint analysis of multiple biomarkers, the system being configured to execute the above-described in vitro detection system and method for early warning of physiological state risk based on joint analysis of multiple biomarkers by executing executable instructions.
[0009] According to another aspect of this application, an electronic device includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute, by executing the executable instructions, the above-described in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis.
[0010] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis.
[0011] This application provides an in vitro detection system and method for physiological state risk early warning based on multi-marker joint analysis. Using serum, urine, and tissue fluid as samples, this application employs multi-dimensional marker simultaneous detection technology to target and capture the concentration, activity response, and molecular binding characteristics of core markers of metabolites, cytokines, and enzymes. After outlier removal and standardized calibration, a random forest joint inference early warning model is constructed by combining multi-marker synergistic rules with clinical standard fusion data to achieve risk classification and identification. Furthermore, by verifying the specific patterns of markers with metabolic pathway correlation data, information such as risk type and core driving markers is clarified, ultimately generating a precise early warning detail including dimensions of metabolic abnormalities, inflammatory responses, and organ function decline, forming a complete technical closed loop of "sample capture - data processing - model early warning - verification output".
[0012] Multi-marker synergistic detection breaks through the limitations of single indicators, covering core dimensions of metabolism, inflammation, and organ function, and improving the comprehensiveness of risk warning; it integrates technologies such as fluorescence signal amplification and surface plasmon resonance to ensure the sensitivity and specificity of detection signals, accurately capture the characteristics of low-concentration markers, and achieve early and accurate warning of physiological risks such as metabolic abnormalities, inflammatory responses, and organ function decline.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0014] Figure 1 The flowchart illustrates an in vitro detection system and method for early warning of physiological state risk based on joint analysis of multiple biomarkers, provided in an embodiment of this application.
[0015] Figure 2 This illustration shows a schematic diagram of an in vitro detection system for early warning of physiological state risk based on joint analysis of multiple biomarkers, provided in an embodiment of this application. Detailed Implementation
[0016] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0017] The following is combined Figure 1 This application describes an in vitro detection system and method for early warning of physiological state risks based on joint analysis of multiple biomarkers, according to exemplary embodiments thereof. It should be noted that the application scenarios described below are merely illustrative for the purpose of understanding the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application are applicable to any suitable scenario.
[0018] In one implementation, Figure 1A schematic flowchart of an in vitro detection system and method for early warning of physiological state risk based on joint analysis of multiple biomarkers, according to an embodiment of this application, is shown.
[0019] S101, Obtain pre-processed serum, urine and tissue fluid samples.
[0020] In one embodiment, the timing, volume, collection container requirements, and contamination prevention procedures for collecting serum, urine, and tissue fluid samples are clearly defined. Serum must be collected in an empty stomach; urine must be processed promptly after collection; and tissue fluid collection must be performed under aseptic conditions. Sterile, interference-free dedicated containers must be used for collection, and contact between the sample and external impurities must be avoided during collection to ensure the initial purity of the sample.
[0021] After collection, serum samples are allowed to clot naturally before being centrifuged to separate the supernatant serum. Urine samples are centrifuged immediately after collection to remove sediment and impurities. Tissue fluid samples are filtered through a sterile membrane to remove tissue fragments and other particulate matter. Centrifugation operations must maintain stable speed and time, and filtration must use a membrane with an appropriate pore size to ensure that the sample is free of obvious impurities after preliminary processing.
[0022] Protein purification techniques were employed to remove interfering proteins from the samples. Serum and tissue fluid samples were separated from the target biomarker using gel filtration chromatography, utilizing the molecular sieving effect of the gel particles. Urine samples were treated with solid-phase extraction, where the target biomarker was adsorbed onto an extraction column, removing metabolic waste and impurities. During purification, parameters such as elution flow rate and extraction conditions needed to be carefully controlled to ensure the effective separation and retention of the target biomarker.
[0023] The initial concentration of the target biomarker in the processed sample was determined by ultraviolet spectrophotometry. If the concentration was too high, a gradient dilution with physiological saline was performed; if the concentration was too low, the sample was concentrated using freeze-drying. After concentration, the sample was reconstituted with physiological saline to ensure the sample concentration met the requirements for subsequent detection. Dilution and concentration operations must strictly adhere to the ratios and process requirements to avoid affecting the biomarker activity.
[0024] After purification and concentration adjustment, the samples were aliquoted into sterile cryovials, labeled with sample number, collection date, and processing time, and then stored in an ultra-low temperature freezer. Repeated freeze-thaw cycles should be avoided during storage. For long-term storage, glycerol should be added as a cryoprotectant to prevent loss of marker activity. Aliquoting must ensure consistent volume, labeling information must be clear and accurate, and the freezer temperature must be stably controlled.
[0025] S102 employs multi-dimensional biomarker synchronous detection technology to target and capture pre-processed serum, urine, and tissue fluid samples, generating raw data on the concentration of each biomarker, activity response curves, and molecular binding characteristic information.
[0026] In one embodiment, a targeted capture platform is built using multi-dimensional biomarker synchronous detection technology. The platform integrates a target-specific identification module, a signal acquisition module, and a data transmission module, with each module working collaboratively via a data bus. Core detection targets in serum, urine, and tissue fluid samples are clearly identified. Metabolites include glucose, triglycerides, and uric acid; cytokines include interleukin-6 and tumor necrosis factor-α; and enzymes include alanine aminotransferase (ALT), aspartate aminotransferase (AST), and lactate dehydrogenase (LDH), ensuring that the targets cover core physiological states related to metabolism, inflammation, and organ function. The target-specific identification module in the platform carries a dedicated probe array for each core target. The binding specificity of the probes to the targets has been verified through previous experiments. The signal acquisition module can simultaneously receive the reaction signals from each probe, and the data transmission module transmits the signals to the subsequent processing unit in real time.
[0027] Based on the principle of specific antigen-antibody binding and fluorescence signal amplification technology, a precise biomarker capture system is constructed to achieve specific recognition and signal conversion of target biomarkers in pretreated samples, ensuring the sensitivity and specificity of the detection signal. For core targets such as metabolites, cytokines, and enzymes, monoclonal antibodies with high specificity and high affinity are screened as capture probes. The screening process is verified through antigen-antibody binding experiments to ensure that the binding specificity of the probe to the target biomarker is more than 100 times higher than that of non-target substances, avoiding cross-reactions. The capture antibodies are covalently immobilized on the reaction area of the detection chip. After amination treatment of the chip surface, a glutaraldehyde cross-linking agent is used to achieve stable connection between the antibody and the chip surface. The immobilization density is controlled at 1000-2000 antibody molecules per square micrometer to ensure uniform probe coverage in the reaction area.
[0028] Fluorescein was selected as the labeling agent to label polyclonal antibodies that specifically bind to the target marker, thus preparing fluorescein-labeled secondary antibodies as signal amplification carriers. During the labeling process, the molar ratio of fluorescein to antibody was strictly controlled at 3:1 to avoid over-labeling and loss of antibody activity. Unbound free fluorescein was removed by dialysis to ensure that the purity of the secondary antibody was higher than 95% and the fluorescence quantum yield was consistently above 0.8, guaranteeing consistent signal amplification.
[0029] When the pretreated sample enters the reaction area, the target biomarker specifically binds to the capture probe to form an antigen-antibody complex. Subsequently, a fluorescein-labeled secondary antibody is added, which specifically binds to the target biomarker in the complex, forming a sandwich structure of "capture probe-target biomarker-fluorescein-labeled secondary antibody." This structure ensures detection specificity through dual specific binding, while the signal amplification effect of the fluorescein-labeled secondary antibody increases the detection signal intensity by 100-1000 times compared to single antibody binding, effectively lowering the detection limit and meeting the detection requirements for low-concentration biomarkers.
[0030] For example, targeting blood glucose, anti-glucose oxidase monoclonal antibodies are screened as capture probes. After immobilization by cross-linking with glutaraldehyde on an aminated chip, the probes are ensured to uniformly cover the reaction area. Fluorescein-labeled anti-glucose oxidase polyclonal antibodies are prepared as secondary antibodies, with the fluorescein-to-antibody molar ratio controlled at 3:1. After dialysis purification, they are ready for use. When blood glucose in the sample enters the reaction area, it specifically binds to the capture probe. The addition of the fluorescein-labeled secondary antibody forms a sandwich complex. The fluorescence signal intensity increases linearly with increasing blood glucose concentration, and a blood glucose concentration as low as 1 mmol / L can be detected. Furthermore, there is no cross-reactivity with other carbohydrates.
[0031] Standardized pretreated samples, after centrifugation, protein purification, and impurity removal, were sequentially injected into the reaction chamber of the targeted capture system according to a preset volume gradient. The volume gradient was set based on previous experiments to ensure that the biomarker concentration in the sample was within the detection linear range, while avoiding excessive sample volume that could cause overflow or insufficient sample volume that could affect the sufficiency of the reaction. The injection process was controlled by a high-precision peristaltic pump, with a stable flow rate of 10 μL / s, to avoid insufficient binding of the biomarker to the probe due to excessively high flow rate, or reduced detection efficiency due to excessively low flow rate.
[0032] The multi-channel parallel detection module includes independent detection channels corresponding to nine core targets. Each channel is equipped with a dedicated fluorescence detection probe, whose excitation wavelength matches that of the fluorescein, and whose emission wavelength detection range covers the emission wavelength range of the fluorescein. The module employs a synchronous triggering mechanism to ensure that each channel starts signal acquisition simultaneously, avoiding data deviations caused by time differences. The detection probe has a detection sensitivity of 0.1 fluorescence unit, enabling precise capture of even weak fluorescence signal changes.
[0033] Signal intensity-concentration calibration curves for each biomarker were established through preliminary experiments. These curves were plotted using a series of standards at various concentrations, with three parallel samples for each concentration gradient. The average value was taken as the signal intensity value corresponding to that concentration. The least squares method was used for curve fitting to ensure a good fit R² ≥ 0.99. After the multi-channel parallel detection module acquired fluorescence signal intensity data, it was substituted into the corresponding calibration curves to automatically generate the raw concentration data for each biomarker. The data retained four significant digits, providing accurate raw data for subsequent data processing.
[0034] For example, with preset volume gradients of 50 μL, 100 μL, and 150 μL, samples of different volume gradients are sequentially injected into the reaction chamber using a high-precision peristaltic pump, with the flow rate controlled at 10 μL / s. The nine channels of the multi-channel parallel detection module correspond to nine core targets, including blood glucose, triglycerides, and interleukin-6, with each channel simultaneously acquiring fluorescence signal intensity. For the blood glucose target, calibration curves were initially plotted using standards at concentrations of 1 mmol / L, 3 mmol / L, 5 mmol / L, 7 mmol / L, and 9 mmol / L, with a goodness of fit R² = 0.995. When a sample's fluorescence signal intensity was detected to be 500 fluorescence units, the raw blood glucose concentration of that sample was calculated to be 5.234 mmol / L by substituting it into the calibration curve.
[0035] The detection program is initiated using real-time dynamic detection technology. The program is set to start timing after the sample is injected into the reaction chamber and continuously record the binding kinetics of the marker and the detection probe within 0-120 minutes. The sampling frequency of the detection system is set according to experimental requirements to ensure complete capture of signal changes during the binding process. The time interval is set to 1 minute to ensure data continuity while avoiding processing pressure caused by excessive data volume.
[0036] The detection system automatically collects fluorescence signal intensity data for each target point at a set 1-minute interval. Signal data is collected three times at each time point, and the average value is taken as the final signal intensity value for that time point, reducing the impact of random errors on the data. Signal stability is monitored in real time during the acquisition process. If the coefficient of variation of the three collected data at a certain time point exceeds 5%, the system automatically re-acquires data to ensure data reliability.
[0037] Signal intensity data for each marker at 121 time points within 0-120 minutes were organized chronologically. Using time as the x-axis and signal intensity as the y-axis, activity response curves for each marker were automatically plotted using data analysis software. The rising phase of the curve reflects the binding rate between the marker and the probe; a steeper slope indicates a faster binding rate. The plateau phase reflects a balanced binding state with stable signal intensity, indicating good binding stability. The binding performance between the marker and the probe can be judged through the curve characteristics.
[0038] For example, the detection system starts timing after sample injection, automatically collecting fluorescence signal intensity data for nine biomarkers at each time point from minute 1 to minute 2 up to minute 120. Each data point is collected three times and the average value is taken. For the tumor necrosis factor-α target, the collected signal intensity data shows a continuous increase from 0 to 30 minutes, after which the signal intensity tends to stabilize. The slope of the rising phase is 10 fluorescence units / minute, and the coefficient of variation of the signal intensity in the stable phase is 2%. The activity response curve is plotted using data analysis software, clearly demonstrating the rapid binding process and good binding stability of tumor necrosis factor-α to the probe.
[0039] The binding process between biomarkers and the detection system is monitored in real time using surface plasmon resonance molecular interaction (SPR) analysis. This technique is based on the phenomenon of surface plasmon resonance. When a biomarker binds to a probe on the surface of the detection chip, it causes a change in the refractive index of the chip surface, resulting in a shift in the resonance angle. By monitoring the amount of resonance angle shift, the binding process can be reflected in real time. During the detection process, the capture probe is fixed to the surface of the sensor chip. After the sample is injected, the change curve of the resonance angle over time is recorded in real time. The binding affinity, binding rate, and dissociation rate parameters are obtained by analyzing the curve.
[0040] The binding kinetic curves were fitted using a 1:1 binding model. Binding rate constants and dissociation rate constants were calculated using data analysis software, and subsequently, the binding affinity constant was calculated. The binding rate constant reflects the speed of binding between the marker and the probe, the dissociation rate constant reflects the stability of the complex, and the binding affinity constant comprehensively reflects the binding ability of both. These parameters together constitute the molecular binding characteristic information, providing crucial evidence for subsequent analysis of the interaction between the marker and the detection system.
[0041] The raw concentration data, activity response curve data, and molecular binding characteristic parameters of each biomarker generated in the previous stage were summarized and organized according to a unified data format. The data format adopted was CSV, with each row corresponding to one biomarker and each column corresponding to one data indicator, including biomarker name, raw concentration data, signal intensity at each time point, binding rate constant, dissociation rate constant, binding affinity constant, etc. The integrated dataset was verified for completeness to ensure that there was no missing data or errors, forming a complete raw biomarker detection dataset. This provides comprehensive and systematic raw data support for subsequent data processing steps such as outlier removal and signal-to-noise ratio optimization.
[0042] For example, by monitoring the binding process of blood glucose with the detection system using surface plasmon resonance molecular interaction analysis, recording the change curve of the resonance angle over time, and fitting it with a 1:1 binding model, the binding rate constant was calculated. The dissociation rate constant is The affinity constant is The raw blood glucose concentration data (5.234 mmol / L), signal intensity data at 121 time points from 0 to 120 minutes, and the aforementioned molecular binding characteristic parameters were summarized and combined with the relevant data of eight other biomarkers, including triglycerides and interleukin-6, in a unified CSV format to form a complete raw biomarker detection dataset.
[0043] S103 performs outlier removal, signal-to-noise ratio optimization, and data standardization calibration on the original detection data. Combining multi-marker collaborative association rules, physiological state risk judgment thresholds, and disease diagnosis reference standards, feature data fusion processing is performed to establish a precise mapping relationship between marker features and risk levels, generating a standardized risk warning fusion dataset.
[0044] In one implementation, the Grubbs test is used to remove outliers. A significance level of 0.05 is set. By calculating the Grubbs statistic for each data point and comparing it to the critical value for the corresponding sample size, data points with a statistic greater than the critical value are removed, ensuring data validity. Wavelet denoising technology is used to optimize the signal-to-noise ratio. The db4 wavelet basis function is selected to perform a three-level decomposition of the original signal, retaining low-frequency approximation coefficients and reconstructing the signal while filtering high-frequency noise interference. Z-score normalization is used to calibrate the data. By calculating the mean and standard deviation of each indicator, the original data is converted into standardized data with a mean of 0 and a standard deviation of 1, constructing a standardized detection data base layer. For example, for raw blood glucose concentration data, a sample with a detection value of 15.6 mmol / L has a Grubbs statistic of 2.89 calculated using the Grubbs test. With a critical value of 2.75 for a sample size of 30, this data point is identified as an outlier and removed. The db4 wavelet basis function is then used to perform a three-level decomposition and reconstruction of the remaining data signal, improving the signal-to-noise ratio from 25 dB to 42 dB. Finally, the blood glucose concentration data were converted into standardized values through Z-score standardization, where 5.234 mmol / L corresponds to a standardized value of 0.86.
[0045] A multi-marker synergistic association rule base is introduced, covering interaction rules among metabolite, cytokine, and enzyme markers, formulated based on clinical research findings and pathological mechanisms. Physiological state risk assessment threshold ranges are integrated; these ranges are determined through statistical analysis of large-sample clinical data, clearly defining the concentration and activity characteristics of each marker corresponding to low, intermediate, and high risks. Clinical diagnostic reference standards are incorporated, including diagnostic thresholds and assessment procedures recommended by authoritative medical guidelines, establishing a feature fusion rule framework and clarifying the core basis for data fusion to ensure the scientific and standardized nature of the fusion process. For example, the rule base clearly defines the increased inflammatory risk when interleukin-6 and tumor necrosis factor-α are synergistically elevated; the physiological state risk assessment threshold range is set as low-risk (<5 pg / mL), intermediate-risk (5-10 pg / mL), and high-risk (>10 pg / mL), while also incorporating relevant diagnostic criteria from the "Guidelines for the Diagnosis of Inflammation-Related Diseases," forming a complete feature fusion rule framework and providing a clear basis for subsequent data fusion.
[0046] A weighted fusion algorithm is used to match calibrated data with feature fusion rules. Weights are assigned based on the impact of each biomarker on physiological risk, determined using the analytic hierarchy process (AHP). The total weight for metabolite biomarkers is 0.4, for cytokines 0.3, and for enzymes 0.3. Standardized biomarker concentration and activity data are then fed into the weighted fusion algorithm to calculate a comprehensive risk score. Based on the correlation between the comprehensive risk score and risk level, a precise mapping relationship is established between biomarker characteristics and low, medium, and high risk levels. For example, if a sample has standardized blood glucose of 0.86, triglycerides of 0.52, interleukin-6 of 1.23, and alanine aminotransferase (ALT) of 0.38, the weighted comprehensive risk score is calculated as: 0.86 × 0.15 + 0.52 × 0.15 + 1.23 × 0.3 + 0.38 × 0.1 + the sum of other biomarker scores, totaling 2.87. The overall risk score is defined as follows: less than 1.5 is low risk, 1.5-3.0 is medium risk, and greater than 3.0 is high risk. The overall risk score of this sample is 2.87, which corresponds to a medium risk level.
[0047] The system integrates standardized data, feature fusion rules, mapping results, and verification mechanisms. Standardized data includes standardized values of concentration and activity characteristics of each biomarker. Feature fusion rules cover risk assessment threshold ranges and clinical diagnostic criteria based on collaborative association rules. The mapping results clearly define the risk level corresponding to each sample. The verification mechanism includes data integrity verification, logical consistency verification, and clinical rationality verification, ensuring no missing data, correct rule application, and results consistent with clinical common sense. All content is organized according to a unified data format to generate a standardized risk warning fusion dataset with a unified structure and complete dimensions, providing high-quality data support for subsequent model construction. For example, it integrates standardized data of nine biomarkers for a sample, feature fusion rule application records, and intermediate-risk level mapping results. After verification confirms data integrity, correct rule application, and results consistent with clinical logic, the dataset is entered according to the prescribed format. The dataset includes fields such as sample number, biomarker standardized value, fusion rule application status, comprehensive risk score, and risk level, forming a standardized risk warning fusion dataset with a unified structure and complete dimensions.
[0048] S104 constructs a multi-marker joint reasoning early warning model based on a standardized risk warning fusion dataset. It takes the concentration, activity intensity and molecular interaction characteristics of core markers as input dimensions, and clinical diagnostic criteria as the judgment basis. The model algorithm parameters are optimized to achieve hierarchical identification of physiological state risks and generate initial risk warning data.
[0049] In one implementation, a deep feature association analysis is conducted on the standardized risk warning fusion dataset and clinical diagnostic criteria. The association strength between each standardized feature and the clinical diagnostic indicator is calculated using the Pearson correlation coefficient. Features with an absolute correlation coefficient greater than 0.6 are defined as strongly correlated features, while redundant features with an absolute correlation coefficient less than 0.3 are eliminated to ensure that the retained features are directly related to risk assessment. The key retained features include the standardized concentration values of nine core biomarkers, the slope of the activity response curve, the signal intensity during the plateau phase of the activity response curve, the binding affinity constant, the binding rate constant, and the dissociation rate constant, totaling 12 core features.
[0050] The 12 selected key features are preprocessed. For missing values, the median of the feature across all samples is used for imputation, avoiding the influence of outliers on mean imputation and ensuring the rationality of data distribution. Regarding feature dimensions, all sample feature data are uniformly converted into vectors of the same dimension. A feature alignment algorithm is used to eliminate dimensional differences between samples, ensuring that each sample's feature vector contains valid data across all 12 dimensions. The final result is a well-structured, feature-clear, and free-of-redundancy, complete model input dataset, providing high-quality data support for model training.
[0051] For example, from the standardized risk warning fusion dataset, the standardized concentration values of nine biomarkers—blood glucose, triglycerides, uric acid, interleukin-6, tumor necrosis factor-α, alanine aminotransferase (ALT), aspartate aminotransferase (AST), lactate dehydrogenase (LDH), and creatinine—were selected, along with the slope of the corresponding activity response curves and the binding affinity constant, totaling 12 key features. The correlation coefficients between each feature and the clinical risk diagnostic indicators were calculated to be above 0.65. For a sample lacking the interleukin-6 activity feature parameter, the median value of this parameter across all samples was calculated to be 0.72, and this median was used for imputation. A feature alignment algorithm was then used to convert all 12 features of the samples into 12-dimensional vectors, completing feature preprocessing and generating the model input dataset.
[0052] A multi-marker joint inference and early warning framework is constructed based on a feature-model-judgment three-dimensional architecture. The framework adopts a modular design, comprising a feature input module, an algorithm operation module, and a risk judgment module. These three modules achieve unidirectional data transmission through standardized data interfaces, ensuring the orderly and secure flow of data. The feature input module has a built-in data format conversion engine, supporting the conversion of the 12-dimensional feature vectors of the model input dataset into tensor formats recognizable by the algorithm operation module. It also features data validation, performing real-time detection of the dimension and numerical range of the input data. When data anomalies are detected, an alarm mechanism is automatically triggered to ensure the validity of the input data. The algorithm operation module is equipped with a preset random forest algorithm, with a built-in algorithm operation engine and data caching unit. The algorithm operation engine is responsible for inference calculations on the input tensor data, while the data caching unit temporarily stores intermediate data during the calculation process, improving computational efficiency. The risk judgment module has a built-in clinical diagnostic standard database, storing risk judgment thresholds and related diagnostic criteria for low, medium, and high risks. After receiving the risk probability values output by the algorithm operation module, it completes risk level matching through a threshold comparison algorithm and feeds the matching results back to the data output unit.
[0053] For example, the feature input module receives 12 key feature vectors from the model input dataset, converts them into 32-bit floating-point tensor format using a data format conversion engine, and after data verification confirms that the data dimensions are correct and the values are within a reasonable range, transmits them to the algorithm computation module through a standardized data interface. The algorithm computation module loads the basic random forest algorithm, and the computation engine performs inference calculations such as splitting and voting on the tensor data. The data caching unit temporarily stores the calculation results of each decision tree, and finally outputs the risk probability value. The risk determination module retrieves a risk threshold from the clinical diagnostic standard database, compares the received risk probability value with the threshold, and completes the risk level determination.
[0054] The model input dataset was divided into training and test sets in a 7:3 ratio using stratified sampling to ensure that the proportions of low-risk, medium-risk, and high-risk samples in both sets were consistent with the original dataset, thus avoiding training bias caused by imbalanced sample distribution. The training set was used for model parameter learning, while the test set was used for model performance validation. Model performance was comprehensively evaluated based on classification accuracy, precision, recall, and F1 score, with classification accuracy being the core objective function.
[0055] The training set was imported into the initial early warning model framework, and the random forest algorithm was used as the core inference engine. The initial values for the number of decision trees were set to 100, the maximum tree depth to 10, and the minimum number of samples for node splitting to 5. Combined with a gradient boosting optimization strategy, a grid search method was used to iteratively fine-tune the model's hyperparameters. The adjustment step size for the number of decision trees was 10, with an adjustment range of 50-200; the adjustment step size for the maximum tree depth was 1, with an adjustment range of 5-15; and the adjustment step size for the minimum number of samples for node splitting was 1, with an adjustment range of 2-10. One set of hyperparameters was adjusted in each iteration, for a total of 50 iterations. After each iteration, the model performance was verified using a test set, and evaluation metrics such as classification accuracy were recorded. When the iteration reached the 32nd iteration, the model achieved a classification accuracy of 92%, a precision of 91%, a recall of 90%, and an F1 score of 0.905 on the test set. All metrics reached their optimal levels, and the iteration was stopped, generating the optimized joint inference early warning model.
[0056] For example, the model input dataset was divided into 700 training samples and 300 test samples using stratified sampling. The training set contained 245 low-risk samples, 280 medium-risk samples, and 175 high-risk samples, while the test set contained 105 low-risk samples, 120 medium-risk samples, and 75 high-risk samples, maintaining the same sample ratio as the original dataset. The training set was imported into the initial early warning model framework, using the random forest algorithm as the core inference engine. The initial hyperparameters were set to 100 decision trees, a maximum tree depth of 10, and a minimum number of samples for node splits of 5. Iterative optimization was performed using a gradient boosting optimization strategy. When the number of decision trees was adjusted to 150, the maximum tree depth to 8, and the minimum number of samples for node splits to 3, the test set classification accuracy reached 92%, precision to 91%, recall to 90%, and the F1 score to 0.905. All performance indicators were optimal, generating the optimized joint inference early warning model.
[0057] Using clinical diagnostic criteria as the rigid basis for judgment, and combining the statistical analysis results of large-sample clinical data, risk probability thresholds are set: low-risk threshold below 0.3, medium-risk threshold between 0.3 and 0.7, and high-risk threshold above 0.7, forming a scientific and reasonable probability threshold division mechanism. The model's input dataset is imported into the optimized joint inference early warning model. The model uses its built-in inference engine to perform risk level mapping calculations on the input 12-dimensional feature vectors. Based on the ensemble decision-making mechanism using the random forest algorithm, it outputs the risk probability value for each sample. The risk probability values are judged according to the probability threshold division mechanism to determine the risk level of each sample. Simultaneously, a feature importance evaluation algorithm is used to calculate the weight ratio of each feature in the risk level determination process. The top three features with the highest weight ratios are identified as core features, clarifying the contribution of core features and the corresponding judgment criteria. Finally, initial risk warning data containing sample number, risk level, core feature name, core feature contribution, and judgment criteria is generated, providing a clear basis for subsequent early warning data verification.
[0058] After the 12-dimensional feature vector of a sample was imported into the optimized joint inference early warning model, the model calculated and output a risk probability value of 0.62 through the inference engine. Based on the probability threshold classification mechanism (0.3-0.7 is intermediate risk), the sample was determined to be at an intermediate risk level. The feature importance assessment algorithm calculated that the standardized value of tumor necrosis factor-α concentration had a weight of 0.28, the interleukin-6 activity parameter had a weight of 0.23, and the standardized value of triglyceride concentration had a weight of 0.18; these three were considered core features. The initial risk warning data clearly recorded the sample number, the intermediate risk level, and the core features as the standardized value of tumor necrosis factor-α concentration, the interleukin-6 activity parameter, and the standardized value of triglyceride concentration, with contribution values of 0.28, 0.23, and 0.18, respectively. The determination was based on the fact that the above core feature parameters exceeded the low-risk judgment range, reaching the feature threshold corresponding to intermediate risk.
[0059] S105 combines the specific expression patterns of each biomarker in different physiological systems, the correlation of metabolic pathways, and the correlation data of clinicopathology to perform risk level verification, concentration threshold verification, and biomarker combination analysis on the initial warning data, thereby clarifying the risk type, the time of warning threshold breach, the core driving biomarker, and the associated physiological system.
[0060] In one implementation, the system systematically analyzes the specific expression patterns of various biomarkers in the circulatory, digestive, immune, and urinary systems. Metabolic biomarkers are mainly responsible for substance transport in the circulatory system and participate in nutrient metabolism in the digestive system, exhibiting dual specific expression. Cytokine biomarkers, as core mediators of the immune response, are concentrated in the immune system for specific expression, mediating inflammation and immune regulation. Enzyme biomarkers have multiple functions, participating in excretion and metabolism in the urinary system, aiding digestion and decomposition in the digestive system, and regulating substance transformation in the circulatory system, all with corresponding specific expression.
[0061] Based on the aforementioned specific expression patterns, bioinformatics analysis was used to extract metabolic pathway association maps corresponding to each biomarker. This revealed that metabolites are deeply involved in glucose and lipid metabolism pathways, cytokines focus on inflammatory response and immune regulation pathways, and enzymes are involved in organ function-related pathways such as liver function regulation, kidney excretion, and myocardial metabolism. Simultaneously, a large amount of clinical case data was collected as clinicopathological control data, covering biomarker detection results and pathological diagnoses from patients with different risk levels and abnormalities in different physiological systems.
[0062] From specific expression patterns, metabolic pathway association maps, and clinicopathological control data, the system captures three core features: risk level verification criteria are determined based on statistical analysis of large-sample clinicopathological data, covering key verification indicators corresponding to metabolic abnormalities, inflammatory responses, and organ function decline, and clarifying the core judgment dimensions of each risk level; concentration threshold ranges are subdivided according to different physiological systems, distinguishing the normal physiological range and risk warning range of each biomarker in the corresponding system, ensuring the tissue specificity of the threshold-adapted biomarker; biomarker combination association rules identify biomarker combinations with synergistic or antagonistic effects in different physiological systems, such as cytokine combinations that promote each other in inflammatory response pathways, and metabolite combinations that influence each other in glucose and lipid metabolism pathways.
[0063] For example, metabolite markers such as blood glucose and triglycerides are specifically expressed in the circulatory and digestive systems, respectively, and participate in glucose and lipid metabolism pathways. Clinical and pathological control data show that when their concentrations consistently exceed the normal range, they are highly correlated with metabolic disorders such as obesity and diabetes. Cytokine markers such as interleukin-6 and tumor necrosis factor-α are specifically expressed in the immune system and are deeply involved in inflammatory response pathways. Clinical and pathological data show that when the concentrations of both are synergistically elevated, they are often accompanied by acute inflammation or acute exacerbations of chronic inflammation. The captured risk level verification criteria include glucose and lipid metabolism indicators corresponding to metabolic abnormalities, cytokine levels corresponding to inflammatory responses, and enzyme activity indicators corresponding to organ dysfunction. The concentration threshold range clarifies the normal range and intermediate- and high-risk ranges of interleukin-6 in the immune system. The characteristics of the marker combination association rules clearly show that when interleukin-6 and tumor necrosis factor-α act synergistically, the inflammatory risk is more significant than when a single marker is abnormally elevated.
[0064] A matching analysis was performed on the risk levels and specific expression patterns of the initial warning data to generate a level fit parameter. The risk level information of the samples recorded in the initial warning data was retrieved, including the classification results of low, intermediate, and high risk, and the corresponding list of core biomarkers. Based on the specific expression patterns of each core biomarker, the physiological system to which the biomarker belongs and the corresponding risk association logic were determined. For example, abnormalities in immune system biomarkers should be primarily associated with the risk of inflammatory responses, and abnormalities in metabolic system biomarkers should be primarily associated with the risk of metabolic abnormalities.
[0065] A cosine similarity matching algorithm was used to analyze the correlation between the risk level and specific expression patterns of initial early warning data. The algorithm uses the association between biomarkers, physiological systems, and risk types as the baseline vector, and the association between the risk level and core biomarkers of the initial early warning data as the matching vector. The degree of matching is quantified by calculating the cosine similarity between the two vectors. A matching degree greater than 0.8 is defined as a high degree of fit, indicating that the initial risk level and the specific expression patterns of the biomarkers are completely consistent; 0.6 to 0.8 is a moderate degree of fit, indicating that there are slight deviations but the core logic is consistent; and less than 0.6 is a low degree of fit, indicating that there is a significant contradiction between the initial risk level and the specific expression patterns. Based on the calculated matching degree, a level fit parameter is generated, with a value ranging from 0 to 1. A higher value indicates a better fit between the risk level and the specific expression patterns of the initial early warning data, providing a quantitative basis for subsequent risk level calibration.
[0066] For example, the initial risk level of a sample's early warning data is intermediate, with the associated core biomarkers being interleukin-6 and tumor necrosis factor-α. Both are specifically expressed in the immune system, and clinicopathological data show that patients at this concentration level often exhibit moderate inflammatory responses, consistent with the core logic of the initial intermediate-risk level. Using a cosine similarity matching algorithm, the cosine similarity between the baseline vector and the vector to be matched is 0.85, and the generated level fit parameter is also 0.85, indicating a high degree of consistency between the initial risk level and the specific expression patterns, requiring no significant adjustments.
[0067] The concentration threshold characteristics were reviewed and analyzed. The threshold judgment criteria were adjusted based on the correlation with metabolic pathways, and a threshold correction factor was generated. The concentration threshold characteristics of each biomarker involved in the initial warning data were extracted, including the upper and lower limits of the thresholds corresponding to different risk levels. Combining the correlation of the metabolic pathways to which each biomarker belongs, the analysis was conducted to determine whether the current threshold adequately considers the interactions of other biomarkers in that metabolic pathway. For example, in the glucose metabolism pathway, did the blood glucose threshold consider the influence of other related biomarkers such as insulin and glycated hemoglobin? In the inflammatory response pathway, did the interleukin-6 threshold consider the levels of co-markers such as tumor necrosis factor-α and C-reactive protein?
[0068] If the initial threshold does not adequately consider the synergistic or antagonistic effects of other biomarkers in the metabolic pathway, resulting in a threshold that is set too high or too low, the threshold determination criteria need to be adjusted according to the regulatory logic of the metabolic pathway. When there are biomarkers with synergistic effects in the metabolic pathway, if the concentrations of other synergistic biomarkers are within the normal range, the risk threshold of the current biomarker can be appropriately increased; if other synergistic biomarkers have become abnormal, the risk threshold of the current biomarker needs to be decreased. A threshold correction factor is generated based on the adjustment magnitude. A correction factor greater than 1 indicates that the threshold needs to be increased, and a correction factor less than 1 indicates that the threshold needs to be decreased. The specific value of the correction factor is determined based on the interaction strength of biomarkers in the metabolic pathway, ensuring that the adjusted threshold better conforms to physiological regulatory laws.
[0069] For example, the initial warning data for a certain sample did not consider the synergistic effect of alanine aminotransferase (ALT) and aspartate aminotransferase (AST) in the liver metabolic pathway. Both ALT and AST jointly reflect the degree of hepatocyte damage in the liver metabolic pathway. The initial threshold only considered ALT levels, leading to an excessively high threshold. By combining the statistical results of the synergistic effect of ALT and AST in clinical pathological data, when AST concentration is normal, the risk threshold for ALT can be adjusted to 0.9 times the original threshold, generating a threshold correction factor of 0.9. This correction factor optimizes the concentration threshold determination criteria, making the threshold more closely reflect the actual physiological mechanisms of liver metabolism.
[0070] The correlation of associated biomarker combinations is validated to detect anomalies where the biomarker combinations do not match the clinicopathological data, generating a combination-matching bias coefficient. Associated biomarker combination features are extracted from the initial warning data to clarify the biomarker combination forms corresponding to each risk level, such as the blood glucose + triglyceride combination for metabolic abnormality risk and the alanine aminotransferase + creatinine combination for organ dysfunction. The correlation of this biomarker combination with the corresponding clinicopathological data is validated, focusing on whether the synergistic effect of the biomarkers in the combination is consistent with the clinicopathological manifestations. For example, a combination labeled as metabolic abnormality risk may not show abnormalities accompanied by more liver function damage, which does not match the initially determined risk type.
[0071] For detected mismatch anomalies, a deviation quantification algorithm is used to calculate the degree of fit deviation. The algorithm uses the main risk type corresponding to the biomarker combination in clinical pathology data as a benchmark and the risk type labeled in the initial warning data as a comparison object to quantify the degree of deviation between the two. A fit deviation coefficient is generated, with the coefficient value ranging from 0 to 1. The larger the coefficient value, the more severe the fit deviation, where 0 represents perfect fit, below 0.3 is mild deviation, 0.3 to 0.7 is moderate deviation, and above 0.7 is severe deviation, providing a quantitative reference for subsequent biomarker combination rescreening.
[0072] For example, the initial warning data for a sample included a combination of biomarkers: blood glucose and alanine aminotransferase (ALT), indicating a risk of metabolic abnormality. However, clinical pathology data showed that when this combination was abnormal, patients often presented with combined symptoms of hepatocellular damage and glucose metabolism disorders, reflecting a combined risk of liver dysfunction and metabolic abnormalities, which did not match the initial assessment of a single metabolic abnormality. Using a bias quantification algorithm, the deviation was calculated to be 0.35, and the generated group fit bias coefficient was also 0.35, indicating a moderate fit bias, requiring adjustment of the biomarker combination.
[0073] Based on preset verification and analysis rules, the risk level fit parameters, threshold correction factors, and combination fit deviation coefficients are processed. A secondary risk level determination algorithm is used to calibrate the risk level results, dynamic threshold adjustment technology optimizes the determination criteria, and a biomarker combination re-screening algorithm corrects the fit deviation, generating verification and optimization results. The preset verification and analysis rules define the reasonable ranges and corresponding processing logic for the risk level fit parameters, threshold correction factors, and combination fit deviation coefficients: the reasonable range for the risk level fit parameter is above 0.6; a value below 0.6 requires initiating a secondary risk level determination. The reasonable range for the threshold correction factor is 0.8 to 1.2; values exceeding this range require focused verification of the adjustment's rationality. The reasonable range for the combination fit deviation coefficient is below 0.3; a value above 0.3 requires initiating a biomarker combination re-screening.
[0074] Based on pre-defined verification and analysis rules, the three parameters mentioned above are systematically processed: A secondary risk level determination algorithm is used, combined with the level fit parameter and clinicopathological control data, to calibrate the initial risk level. Highly matched risk levels remain unchanged, moderately matched levels are fine-tuned, and low-matched levels are re-determined. Through dynamic threshold adjustment technology, the concentration thresholds of each biomarker are optimized based on a threshold correction factor, while simultaneously verifying the suitability of the adjusted thresholds for metabolic pathway association. A biomarker combination re-screening algorithm is employed, based on the combination fit deviation coefficient, to eliminate biomarker combinations with severe fit deviations, and to re-screen biomarker combinations from the metabolic pathway association map whose synergistic effects match the clinicopathological data, ensuring a high degree of match between the combination and the risk type. The combined results of these processes generate a verification and optimization result, including the calibrated risk level, optimized concentration threshold, and re-screened biomarker combinations.
[0075] For example, the preset reasonable range for the grade fit parameter is above 0.6, the reasonable range for the threshold correction factor is 0.8 to 1.2, and the reasonable range for the combination fit deviation coefficient is below 0.3. A sample has a grade fit parameter of 0.85, which is within the reasonable range, so no adjustment to the risk grade is needed; the threshold correction factor is 0.9, which is within the reasonable range, so the concentration threshold of alanine aminotransferase (ALT) is adjusted according to this factor; the combination fit deviation coefficient is 0.35, which is outside the reasonable range. Therefore, the combination of blood glucose and ALT is removed using a biomarker combination re-screening algorithm, and specific combinations of abnormal blood glucose and triglyceride metabolism are re-screened from the glucose metabolism pathway, generating a validation and optimization result that includes the calibrated risk grade, optimized threshold, and new biomarker combinations.
[0076] The above validation and optimization results were integrated, including the calibrated risk level, optimized concentration threshold, and rescreened biomarker combination. Combining metabolic pathway correlation maps and clinicopathological control data, the risk type of the samples was identified, categorized into three types: metabolic abnormalities, inflammatory responses, and organ dysfunction. This was determined based on the metabolic pathways to which the core biomarkers belonged and their clinicopathological manifestations. For example, abnormalities in glucose and lipid metabolism pathway biomarkers correspond to metabolic abnormalities, abnormalities in inflammatory response pathway biomarkers correspond to inflammatory responses, and abnormalities in enzyme biomarkers of organ function regulation pathways correspond to organ dysfunction.
[0077] Based on the calibrated risk level and optimized concentration threshold, the change curve of biomarker concentration over time is traced back to determine the warning threshold breach time, i.e., the specific time point when the biomarker concentration first exceeds the adjusted threshold. This moment reflects the critical time point when the risk begins to manifest. Based on the characteristics of biomarker combination association rules and specific expression patterns, core driving biomarkers are identified. Core driving biomarkers are those that contribute the most to the risk level determination and conform to specific expression patterns. The contribution of each biomarker is calculated using a feature importance assessment algorithm, and the top two biomarkers in terms of contribution are selected as core driving biomarkers. The associated physiological systems corresponding to the core driving biomarkers are identified, and the physiological systems mainly affected by the risk are determined by combining specific expression patterns. Integrating the above information, accurate and verified risk association details are generated, including risk type, warning threshold breach time, core driving biomarkers, and associated physiological systems, providing accurate data support for the subsequent generation of risk warning details.
[0078] For example, integrating the validation and optimization results of a sample, the calibrated risk level is intermediate. The optimized concentration thresholds for interleukin-6 and tumor necrosis factor-α better reflect the synergistic effect of inflammatory response pathways. The rescreened biomarker combination is interleukin-6 + tumor necrosis factor-α. Combining metabolic pathway association maps and clinicopathological control data, the risk type is determined to be inflammatory response. By retrospectively analyzing the biomarker concentration change curves, it is determined that the time when interleukin-6 first exceeds the optimized threshold is 30 minutes after sample testing, i.e., the moment the warning threshold is exceeded. Calculated using a feature importance assessment algorithm, the contribution of interleukin-6 is 0.42, and the contribution of tumor necrosis factor-α is 0.38, making them the core driving biomarkers. Both are specifically expressed in the immune system, and the associated physiological system is the immune system. Integrating the above information, detailed risk association data after accurate validation is generated.
[0079] S106 integrates and verifies various data to generate detailed information on physiological state risk warnings related to metabolic abnormalities, inflammatory responses, and organ function decline.
[0080] In one implementation, meticulously validated risk-related detailed data is compiled, encompassing risk type, warning threshold breach time, core driving biomarkers, and associated physiological systems. Simultaneously, standardized detection data, feature fusion rules, model judgment criteria, and validation optimization records from previous processes are integrated. The integrated data undergoes a completeness check to ensure no critical information is missing. Data format standardization converts data from different sources and formats into a structured data format, laying the foundation for subsequent detailed information generation. For example, integrating validated data for a sample where the risk type is inflammatory response, the warning threshold breach time is 30 minutes after detection, the core driving biomarkers are interleukin-6 and tumor necrosis factor-α, and the associated physiological system is the immune system, along with the sample's standardized detection data, glucose and lipid metabolism and inflammatory response-related feature fusion rules, random forest model judgment criteria, and threshold correction records, confirms no missing information after a completeness check, thus completing the data format standardization conversion.
[0081] Based on the integrated data, risk types are further refined and labeled. If the risk type is metabolic abnormality, it is further clarified as abnormal glucose metabolism, abnormal lipid metabolism, or mixed metabolic abnormality; if it is an inflammatory response, it is differentiated into acute inflammation, chronic inflammation, or localized inflammation; if it is organ dysfunction, the specific associated organs are labeled, such as liver dysfunction, kidney dysfunction, myocardial dysfunction, etc., making the risk type more targeted. For example, if a sample's risk type is metabolic abnormality, it is further refined and labeled as mixed glucose and lipid metabolism abnormality based on the abnormal manifestations of core driver markers blood glucose and triglycerides and the correlation data of glucose and lipid metabolism pathways; if another sample's risk type is organ dysfunction, it is further refined and labeled as liver dysfunction based on the abnormalities of core driver markers alanine aminotransferase (ALT) and aspartate aminotransferase (AST) and liver metabolic pathway data.
[0082] The core information is presented in a structured manner: the integrated key information is presented according to a pre-defined structure, including three parts: basic sample information, core risk information, and supporting data information. Basic sample information includes sample number, testing date, and sample type; core risk information covers risk type and detailed annotations, warning threshold breach time, core driving biomarkers, and related physiological systems; supporting data information includes concentration data, activity characteristic parameters, molecular binding characteristic parameters, and clinical pathological references for the core driving biomarkers, ensuring clear and well-organized details and sufficient data support. For example, in the warning details for a sample, the basic sample information clearly states the sample number, testing date, and sample type as serum; the core risk information indicates the risk type as acute inflammation, the warning threshold breach time as 25 minutes after testing, the core driving biomarkers as interleukin-6 and tumor necrosis factor-α, and the related physiological system as the immune system; the supporting data information lists the standardized concentration values, activity response curve slopes, and affinity constants of the two core biomarkers, as well as the corresponding clinical pathological data and the reference range of biomarker characteristics in patients with acute inflammation.
[0083] Supplementary Clinical Reference Recommendations: Based on risk type, core driving biomarkers, and associated physiological systems, and in accordance with clinical diagnostic reference standards and multi-biomarker synergistic association rules, supplementary targeted clinical reference recommendations are provided. These recommendations include further examinations, lifestyle modifications, and key clinical interventions, ensuring that the warning details not only serve as risk alerts but also provide practical guidance for subsequent diagnosis and treatment. For example, for a sample with a risk type of renal dysfunction, associated physiological systems of the urinary system, and core driving biomarkers of creatinine and uric acid, supplementary clinical reference recommendations include further renal ultrasound examination, reduction of high-purine food intake, and regular monitoring of renal function indicators. It also emphasizes the need for clinical attention to electrolyte balance, providing clear direction for the patient's subsequent treatment.
[0084] The integrated information, detailed risk types, structured core information, and supplementary clinical recommendations are compiled into a final physiological state risk warning detail according to a standardized format. This standardized format ensures clear information hierarchy, concise and professional language, facilitating rapid retrieval of key information by clinical medical staff, while also supporting data export and sharing to meet clinical diagnostic data management needs. The final warning detail is presented in a hierarchical format, first listing basic sample information, then clearly displaying core risk information, followed by supporting data, and finally providing clinical reference recommendations. The overall format is consistent and logically coherent, allowing medical staff to quickly obtain key information such as the sample's risk type and core driving biomarkers, and also access detailed supporting data and clinical recommendations to meet their diagnostic and treatment reference needs.
[0085] In one implementation, such as Figure 2 As shown, this application also provides an in vitro detection system for early warning of physiological state risks based on joint analysis of multiple biomarkers, comprising:
[0086] The sample acquisition module 201 is used to acquire pre-processed serum, urine and tissue fluid samples to provide a standardized sample basis for subsequent testing.
[0087] The biomarker targeted capture module 202 is used to build a targeted capture platform using multi-dimensional biomarker synchronous detection technology. Based on the principle of antigen-antibody specific binding and fluorescence signal amplification technology, it constructs a precise capture system. Through multi-channel parallel detection, real-time dynamic recording and molecular interaction analysis, it generates raw data of each biomarker concentration, activity response curve and molecular binding characteristic information.
[0088] The detection data processing and fusion module 203 is used to process the raw data through Grubbs test, wavelet denoising technology and Z-score standardization algorithm, and combine the multi-marker collaborative association rule base, risk judgment threshold and clinical diagnostic criteria to establish the mapping relationship between marker features and risk level using a weighted fusion algorithm to generate a standardized risk warning fusion dataset.
[0089] The joint reasoning early warning model construction module 204 is used to build an early warning framework based on the feature-model-judgment three-dimensional architecture. It uses the random forest algorithm as the core reasoning engine, combines gradient boosting optimization strategy to fine-tune hyperparameters, and uses clinical diagnostic criteria as the judgment basis to realize the risk classification identification of physiological state and generate initial risk warning data.
[0090] The early warning data verification and analysis module 205 is used to combine the specific expression patterns of biomarkers in the circulatory, digestive, immune, and urinary systems, the correlation of metabolic pathways, and clinical pathological data to verify the risk level, review the concentration threshold, and analyze the combination of related biomarkers on the initial early warning data, so as to clarify the risk type, the time when the early warning threshold is exceeded, the core driving biomarkers, and the related physiological systems.
[0091] The risk warning detail generation module 206 is used to integrate the verified data and generate detailed information on physiological state risk warnings related to metabolic abnormalities, inflammatory responses, and organ function decline, providing accurate references for clinical diagnosis and treatment.
[0092] The computer-readable storage medium provided in the above embodiments of this application and the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application stored therein.
[0093] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis, electronic devices, electronic devices, and readable storage media are relatively simple in description because they are fundamentally similar to the embodiments of the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis described above. Relevant parts can be referred to in the descriptions of the embodiments of the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis described above.
Claims
1. An in vitro detection system and method for early warning of physiological state risks based on joint analysis of multiple biomarkers, characterized in that, include: Obtain pretreated serum, urine, and tissue fluid samples; Multidimensional biomarker synchronous detection technology was used to target and capture pre-treated serum, urine and tissue fluid samples to generate raw data of each biomarker concentration, activity response curves and molecular binding characteristics. Outlier removal, signal-to-noise ratio optimization, and data standardization calibration are performed on the raw detection data. Combined with multi-marker collaborative association rules, physiological state risk judgment thresholds, and disease diagnosis reference standards, feature data fusion processing is carried out to establish a precise mapping relationship between marker features and risk levels, and generate a standardized risk warning fusion dataset. A multi-marker joint reasoning early warning model is constructed based on a standardized risk early warning fusion dataset. The model takes the concentration, activity intensity and molecular interaction characteristics of core markers as input dimensions and clinical diagnostic criteria as the judgment basis. The model algorithm parameters are optimized to achieve hierarchical identification of physiological state risks and generate initial risk early warning data. By combining the specific expression patterns of each biomarker in different physiological systems, the correlation of metabolic pathways, and the correlation with clinicopathological data, the initial warning data are subjected to risk level verification, concentration threshold review, and combination analysis of associated biomarkers to clarify the risk type, the time of breach of the warning threshold, the core driving biomarker, and the associated physiological system. By integrating and verifying the various data, detailed information on physiological state risk warnings related to metabolic abnormalities, inflammatory responses, and organ function decline is generated.
2. The method as described in claim 1, characterized in that, Multidimensional biomarker simultaneous detection technology was used to target and capture pretreated serum, urine, and tissue fluid samples, generating raw concentration data, activity response curves, and molecular binding characteristics of each biomarker, including: A targeted capture platform was built using multidimensional biomarker synchronous detection technology to identify the core detection targets in serum, urine and tissue fluid samples, including metabolites such as blood glucose, triglycerides and uric acid, cytokines such as interleukin-6 and tumor necrosis factor-α, and enzymes such as alanine aminotransferase, aspartate aminotransferase and lactate dehydrogenase. Based on the principle of specific binding of antigen and antibody and fluorescence signal amplification technology, a biomarker precision capture system is constructed to achieve specific identification and signal conversion of target biomarkers in pretreated samples, ensuring the sensitivity and specificity of the detection signal; Standardized pretreated samples that have undergone centrifugation, protein purification, and impurity removal are injected into the targeted capture system according to a preset volume gradient. Signal intensity data of each target point are collected synchronously through a multi-channel parallel detection module to generate raw data of each biomarker concentration. By combining real-time dynamic detection technology to continuously record the binding dynamics of markers and detection probes within 0-120 minutes, and collecting signal change data at 1-minute intervals, the activity response curves of each marker are generated. By using surface plasmon resonance molecular interaction analysis technology, the binding affinity, binding rate, and dissociation rate parameters of the biomarker and the detection system are analyzed, molecular binding characteristic information is extracted, and raw concentration data, activity response curves, and binding characteristic parameters are integrated to form a complete raw biomarker detection dataset.
3. The method as described in claim 1, characterized in that, Outlier removal, signal-to-noise ratio optimization, and data standardization calibration are performed on the raw detection data. Combined with multi-marker collaborative association rules, physiological state risk assessment thresholds, and disease diagnostic reference standards, feature data fusion processing is conducted to establish a precise mapping relationship between marker features and risk levels, generating a standardized risk warning fusion dataset, including: Outliers were removed from the original detection data using the Grubbs test, the signal-to-noise ratio was optimized by combining wavelet denoising technology, and the data calibration was completed using the Z-score standardization algorithm to build a standardized detection data base layer. A multi-marker collaborative association rule base is introduced, integrating the threshold range for physiological state risk assessment with clinical diagnostic reference standards, building a feature fusion rule framework, and clarifying the core basis for data fusion; Based on the weighted fusion algorithm, the calibrated data and feature fusion rules are matched and calculated to establish a precise mapping relationship between biomarker concentration, activity characteristics and low-risk, medium-risk and high-risk levels; By integrating standardized data, fusion rules, mapping results, and verification mechanisms, a standardized risk warning fusion dataset with a unified structure and complete dimensions is generated.
4. The method as described in claim 1, characterized in that, A multi-marker joint inference early warning model is constructed based on a standardized risk early warning fusion dataset. Using the concentration, activity intensity, and molecular interaction characteristics of core markers as input dimensions and clinical diagnostic criteria as the judgment basis, the model algorithm parameters are optimized to achieve graded identification of physiological state risks and generate initial risk early warning data, including: Feature association analysis and preprocessing are performed on the standardized risk warning fusion dataset and clinical diagnostic criteria to generate the model input dataset; A multi-marker joint inference early warning framework is constructed based on the feature-model-judgment three-dimensional architecture, and an initial early warning model framework is generated. The initial early warning model framework includes a feature input module, an algorithm operation module, and a risk judgment module. The feature input module is used to receive standardized feature data, the algorithm operation module is used to perform inference calculations, and the risk judgment module is used to match clinical diagnostic criteria. The model input dataset is imported into the initial early warning model framework. The random forest algorithm is used as the core inference engine. The model hyperparameters are iteratively tuned by combining gradient boosting optimization strategy to generate an optimized joint inference early warning model. Using clinical diagnostic criteria as the rigid basis for judgment, the model performs risk level mapping calculation on the input feature data, and adopts a probability threshold division mechanism to achieve graded identification of low risk, medium risk and high risk, generating initial risk warning data that includes risk level, contribution of core features and judgment basis.
5. The method as described in claim 4, characterized in that, By combining the specific expression patterns of various biomarkers in different physiological systems, their correlation with metabolic pathways, and their correlation with clinicopathological data, the initial warning data were subjected to risk level verification, concentration threshold review, and combined analysis of associated biomarkers to clarify the risk type, the time of warning threshold breach, the core driving biomarkers, and the associated physiological systems, including: Based on the specific expression patterns of each biomarker in the circulatory, digestive, immune, and urinary systems, metabolic pathway association maps and clinicopathological control data were extracted to capture the risk level verification criteria, concentration threshold ranges, and biomarker combination association rules. A matching analysis was performed on the risk level and specific expression patterns of the initial early warning data to generate a level fit parameter. The concentration threshold characteristics were re-analyzed, and the threshold judgment criteria were adjusted according to the correlation of metabolic pathways to generate a threshold correction factor; The correlation of the characteristics of the associated biomarker combination is verified, and abnormalities in the mismatch between the biomarker combination and the clinicopathological data are detected, and the combination matching deviation coefficient is generated. Based on the preset verification and analysis rules, the level fit parameters, threshold correction factors and combination fit deviation coefficients are processed. The risk level secondary judgment algorithm is used to calibrate the level results, the threshold dynamic adjustment technology is used to optimize the judgment criteria, and the marker combination re-screening algorithm is used to correct the fit deviation and generate the verification optimization results. By integrating and optimizing the results, we can identify the risk types, warning threshold breach times, core driving biomarkers, and related physiological systems of metabolic abnormalities, inflammatory responses, and organ function decline, and generate detailed risk association data after accurate verification.
6. An in vitro detection system for early warning of physiological state risks based on joint analysis of multiple biomarkers, characterized in that, The system is configured to execute, via executing executable instructions, the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis as described in any one of claims 1 to 5.
7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis as described in any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the in vitro detection system and method for physiological state risk warning based on multi-marker joint analysis as described in any one of claims 1 to 5.