Clinical biochemical determination calibration model construction and bias correction method based on support vector machine

By introducing SVM algorithm and PLS to judge the linear characteristics of data, and building a calibration model, the problem that the calibration curve cannot be monitored and adjusted in real time in clinical biochemical detection is solved, real-time monitoring and automatic adjustment of detection results are achieved, detection accuracy and reliability are improved, and CLSI specifications are met.

CN120296370APending Publication Date: 2025-07-11李建红
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377780.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing clinical biochemical detection methods cannot monitor and automatically adjust the calibration curve in real time, resulting in high uncertainty in the detection results, unable to meet the calibration traceability requirements of CLSI EP05-A3 and other standards, and are susceptible to human errors.

Method used

The support vector machine (SVM) algorithm is used to combine partial least squares method (PLS) to judge the linear characteristics of the data, dynamically select the model type, select the linear core SVM or neural network algorithm through the principal component interpretation variance threshold, build a calibration model, and realize the visualization of calibration curves and real-time monitoring of bias through the Bohrium platform, and seamlessly connect it with the laboratory information management system (LIS) to form a full-process automation closed loop.

Benefits of technology

Real-time monitoring and automated adjustment of calibration curves are realized, the accuracy and reliability of detection results are improved, the CLSI EP09-A3 specifications are met, the incidence of misdiagnosis and missed diagnosis is reduced, and the medical examination level and social health level are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296370A_ABST
    Figure CN120296370A_ABST
Patent Text Reader

Abstract

The invention discloses a clinical biochemical determination calibration model construction and bias correction method based on a support vector machine, and belongs to the technical field of medical examination. The method is technically characterized in that calibration data, EQA data and IQC data are collected, Grubbs inspection and Z-score standardization processing are combined, a partial least square method is used for calculating a principal component interpretation variance, when the principal component interpretation variance is larger than or equal to 95%, an explicit calibration equation: Cadjuded = K * m + B is constructed based on a linear kernel SVM regression model, the slope K and the intercept B are directly output, and traceable calibration is achieved. And verifying the performance of the model through a mean square error, a decision coefficient and bias. Visualization and bias real-time monitoring of a calibration curve are achieved by means of a Bohrium platform, and integration with a laboratory information management system (LIS) is achieved through an API interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical testing, and particularly relates to a method for constructing a calibration model and correcting bias for clinical biochemical determination based on a support vector machine (SVM). This method significantly improves the accuracy and reliability of clinical biochemical testing through the artificial intelligence SVM algorithm, providing more accurate data support for test results. Background Art

[0002] Clinical biochemical testing occupies a key position in the modern medical system. Its test results are important bases for disease diagnosis, treatment plan formulation, and condition monitoring. Accurate clinical biochemical test results play a decisive role in improving medical quality and ensuring patient health. During the quality control process, since the specific parameters of the calibration equation cannot be obtained, it is difficult for laboratories to determine the accuracy and reliability of calibration, increasing the uncertainty of test results. Clinical biochemical testing is vulnerable to various factors, such as reagent batch change, environmental temperature and humidity change, instrument aging, etc. These factors will cause the calibration curve to drift, resulting in deviations in test results and failing to meet the strict requirements for calibration traceability in relevant standards such as Clinical and Laboratory Standards Institute (CLSI) EP05 - A3. Traditional calibration methods cannot monitor and automatically adjust the calibration curve in real time, require frequent manual intervention, and are prone to introducing human errors, making it difficult to ensure the timeliness and accuracy of calibration. The existing calibration methods cannot meet the increasing clinical needs, and there is an urgent need for an innovative calibration technology that can effectively solve the problems existing in traditional methods, improve the accuracy and reliability of clinical biochemical testing, and provide more accurate data support for clinical diagnosis and treatment. The present invention aims to solve the existing problems, introduce the artificial intelligence SVM algorithm into the field of clinical biochemical determination calibration, construct an accurate calibration model, and achieve bias correction, providing a more reliable technical guarantee for clinical biochemical testing.

[0003] In the prior art, the invention patent "Chemiluminescence Calibration Curve Optimization Method Based on Artificial Intelligence Algorithm" (Patent No.: CN2024118228054) proposes an algorithm based on a neural network for processing nonlinear data in chemiluminescence detection. The present invention innovatively introduces the SVM algorithm, uses the partial least squares method (PLS) to determine the linear characteristics of the data, and dynamically selects the model type through the principal component explained variance threshold (≥95%): when the data is approximately linear, the linear kernel SVM is used to directly output the calibration equation parameters (K and B) to ensure the interpretability and calibration traceability of the model; only switch to the neural network algorithm when the data presents significant nonlinearity (principal component explained variance <95%). This strategy takes into account the needs of both linear and nonlinear scenarios and solves the problem of non-traceability of parameters in the prior art. The two methods have different linear characteristics, adopt different artificial intelligence algorithms (chemiluminescence method uses neural network method), and have different scopes of application. This technology is mainly used for the detection of various biochemical items in clinical laboratories. It combines PLS with SVM algorithms for explicit parameter output of clinical biochemical calibration models, filling the gap in real-time calibration and automated adjustment of traditional methods. The full-process automated tool chain (Bohrium platform, Python script, laboratory information management system (LIS) integration) significantly improves calibration efficiency and standardization level.

[0004] The present invention innovatively proposes a principal component explained variance threshold judgment mechanism, dynamically selects a suitable model based on the linear characteristics of the data, and meets the strict requirements of clinical testing for traceability while ensuring the advancement of the algorithm, solving the problem that traditional models cannot be explained. Based on the linear characteristics of the biochemical test absorbance value (OD value) and the measured value, the SVM algorithm is used to directly output the calibration equation parameters (K and B), and the calibration curve visualization and real-time bias monitoring are achieved through the Bohrium platform. Combined with Python scripts and API interfaces, it is seamlessly connected to the LIS system to form a closed loop of the entire process of "detection-calibration-quality control". Summary of the invention

[0005] The present invention belongs to the field of medical laboratory technology, and provides a method for constructing a clinical biochemical measurement calibration model based on SVM and correcting bias, comprising the following steps: (1) multi-source data acquisition and preprocessing: integrating calibration data, EQA target values ​​and measurement values, and continuous IQC data, eliminating outliers through Grubbs test, and using Z-score standardization to eliminate instrument differences; (2) dynamic model construction: calculating the variance explained by the principal component based on PLS, selecting a linear kernel SVM regression model if it is ≥95%, directly outputting explicit calibration equation parameters (slope K and intercept B), and realizing traceable calibration; if the variance explained by the principal component is insufficient, switching to a neural network algorithm; (3) full process automation integration: realizing calibration curve visualization and real-time bias monitoring through the Bohrium platform, automatically generating a calibration report in combination with a Python script, and seamlessly connecting with the LIS through an API interface to form a "detection-calibration-quality control" closed-loop management; (4) performance verification: using mean square error (MSE<1.5), determination coefficient (R 2 >0.98) and bias % (Bias %) (±5%) were used as evaluation criteria to ensure that the model met the CLSI EP09-A3 specification.

[0006] The present invention focuses on the calibration link of clinical biochemical determination, and shows significant advantages in technical application, quality control and clinical practice, providing strong support for improving the level of medical inspection. It is widely and deeply applied, running through all aspects of medical inspection, and providing accurate data support for clinical diagnosis and treatment. In clinical laboratories such as hospitals at all levels and physical examination centers, it is used for the detection and calibration of various biochemical projects, such as liver function, kidney function, blood lipids, blood sugar and other projects, to ensure the accuracy of the test results, and help doctors to diagnose diseases, monitor the condition and formulate treatment plans. The present invention not only brings changes at the medical technology level, but also has a wide and positive impact in the social and economic fields, improves the overall health level of society, and creates significant economic benefits. By improving the accuracy and reliability of clinical biochemical detection, reducing the occurrence of misdiagnosis and missed diagnosis, providing patients with more accurate medical services, improving patients' medical experience and satisfaction, providing reliable data support for medical research, accelerating the progress of medical research, promoting the development of new diagnostic methods and treatment methods, promoting the development of medical science, and improving the overall medical level of society. Reduce repeated detection and erroneous treatment caused by detection system errors, reduce the waste of medical resources, and save medical costs. Improve the efficiency of medical services, improve the mutual recognition of results, shorten the patient's hospital stay, and reduce the economic burden on patients and society. The application of this invention will drive the development of related medical equipment, reagents and software industries, create new economic growth points, attract more companies and investments to enter the field of clinical biochemical testing, and promote industrial upgrading and innovation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1Patent implementation flowchart, which shows the entire process from data collection to model evaluation and automated integration, including key links such as data preprocessing, model construction, performance evaluation, real-time bias monitoring, etc., as well as the logical relationships and data flows between each link.

[0008] Figure 2 Original calibration curve and calibrated curve after SVM calibration taking total protein (TP) as an example. Among them, b is the original calibration curve, and a is the calibration curve after SVM calibration, visually presenting the differences in the calibration curves before and after calibration, and reflecting the optimization effect of the present invention on the calibration curve. Detailed implementation manners

[0009] The present invention will be further described in detail below through TP embodiments and accompanying drawings.

[0010] Embodiment 1

[0011] (1) Data collection and preprocessing

[0012] 1. Data collection

[0013] 1.1 Calibration data: Collect the OD values corresponding to different concentration gradients of TP and the concentration data of the calibration solution from the biochemical analyzer, as shown in Table 1.

[0014] Table 1 TP calibration solution concentration and OD value

[0015] Calibration solution Concentration (g / L) OD value OD1 0.00 -664.5 OD2 53.5 1107

[0016] 1.2 The third EQA TP data of the national routine chemistry A in 2024, as shown in Table 2.

[0017] Table 2 The measured values and target values of the third EQA of the national routine chemistry A in 2024 for TP data

[0018]

[0019] 1.3 IQC data: Collect the TP IQC data of Tianjin Baodi District People's Hospital in October 2024 continuously, as shown in Table 3.

[0020] Table 3 TP IQC data of Tianjin Baodi District People's Hospital in October 2024 continuously

[0021] Item Lot number Calculate target value Calculate standard deviation Calculate coefficient of variation Roche control product 1 89751 61.73 0.6117 0.9909 Roche control product 2 89752 41.37 0.4984 1.2047

[0022] 2. Data preprocessing

[0023] 2.1 Outlier removal: Use the Grubbs test (α = 0.05) to screen the collected data, and regard the calibration data deviating from the mean value by ±20% and the IQC data exceeding ±3SD as outliers and remove them.

[0024] 2.2 Data standardization: The Z-score standardization method is used to process the data generated by different instruments to make the data comparable.

[0025] The calculation formula is:

[0026]

[0027] Among them, X is the original data point; μ is the mean of the data set; σ is the standard deviation of the data set; Z is the standardized data point.

[0028] 2.3 Dataset division: The processed data is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. The training set is used for model training, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model.

[0029] (2) Construction of SVM calibration model

[0030] 1. Use PLS to judge the linear characteristics of the data. When the variance explained by the principal component is ≥95%, select the linear kernel SVM for SVM; if the variance explained by the principal component is insufficient and the data shows non-linear characteristics, then use the neural network algorithm, which can make the model better fit the data and improve the calibration accuracy.

[0031] 2. Hyperparameter tuning: The value ranges of the penalty coefficient C (1 - 100) and the kernel coefficient γ (0.001 - 1) are determined through preliminary experiments and cross-validation.

[0032] 2.1 Penalty coefficient C: The C value controls the model's tolerance for errors. A smaller C (such as C = 1) tends to select a simpler model (to prevent overfitting), and a larger C (such as C = 100) allows the model to fit the training data better (to improve the fitting ability). Through grid search verification, C within the range of 1 - 100 can balance the generalization and accuracy of the model.

[0033] 2.2 Kernel coefficient γ: The γ value affects the mapping ability of the SVM kernel function. A smaller γ (such as γ = 0.001) corresponds to a smoother decision boundary and is suitable for linear or approximately linear data; a larger γ (such as γ = 1) is suitable for complex non-linear data. Combining the results of principal component analysis (when the variance explained by the principal component is ≥95%, the data is approximately linear), choosing γ = 0.001 - 1 can cover the calibration requirements from linear to weakly non-linear.

[0034] 2.3 The above parameter ranges refer to the suggestions on SVM hyperparameter optimization in "Support Vector Machines for Classification: A Survey" (IEEE, 2019), and their effectiveness is verified on the training set through Monte Carlo cross-validation (MCCV).

[0035] 3. Calibration equation determination: The calibration equation C is obtained through SVM regression adjusted = K×m + B, where C adjusted is the calibrated concentration, m is the measured concentration before calibration, K is the slope of the model, and B is the intercept of the model.

[0036] (III) Performance evaluation and verification

[0037] 1. Evaluation index calculation - MSE: Calculate the mean square error between the predicted target value and the actual target value.

[0038] 2. R 2 : Evaluate the ability of the model to explain the variability of the data.

[0039] 3. Bias%: Bias% within ±5% is usually considered low bias, indicating that the predicted value is very close to the true value and the prediction accuracy of the model is relatively high.

[0040] 4. Calculate the slope, intercept, and calibration equation of the regression model through Python code, as shown in Table 4.

[0041] Table 4 Calibration equation and parameters of the TP regression model

[0042] Item K B Calibration equation TP 0.9894 2.0030 <![CDATA[C adjusted = 0.9894×m + 2.0030]]>

[0043] 5. Comparison before and after calibration: Retest the 2024 national routine chemistry A third EQA data with the adjusted curve and calculate Bias%, as shown in Table 5.

[0044] Table 5 MSE, R 2 and average bias before and after calibration

[0045] Item MSE <![CDATA[R 2 > Bias% before calibration Bias% after calibration TP 0.0002 0.9988 -1.42 -0.65

[0046] Calibration curve comparison diagram: As Figure 2 shown, where b is the original calibration curve and a is the calibration curve after SVM calibration, visually presenting the difference between the calibration curves before and after, and reflecting the optimization effect of the present invention on the calibration curve.

[0047] (IV) Application of the full-process automation tool chain

[0048] 1. Visualization display: Generate a calibration curve comparison diagram with the help of the Bohrium platform, visually display the difference between the original calibration curve and the optimized calibration curve, and monitor the bias trend in real time. Laboratory personnel can clearly understand the calibration process and bias changes through the visualization interface.

[0049] 2. System integration uses Python scripts to automatically generate calibration reports, and the report content includes key information such as calibration equation parameters, performance indicators, and visualization charts. The calibration parameters are synchronized to the LIS through the API interface to achieve full-process automation of data collection, model training, bias monitoring, and report generation.

[0050] Example 2

[0051] Example 2: For the calibration model construction and bias correction of other biochemical items such as albumin (ALB), urea (BUN), direct bilirubin (DBIL), magnesium (Mg), calcium (Ca), etc., replace TP with other biochemical items (such as ALB, BUN, DBIL, Mg, Ca, etc.), and perform data collection, preprocessing, model construction, performance evaluation and verification according to the steps of Example 1. The calibration equations and performance indicators of each item are shown in Table 6.

[0052] Table 6 Regression model calibration equations and performance indicators for items such as ALB, BUN, DBIL, Mg, Ca, etc.

[0053]

[0054] Through the implementation of the present invention, the calibration models of each biochemical item have reached the expected performance indicators, and the bias between the calibrated test results and the target values has been significantly reduced, all meeting the CLSI EP09-A3 standard, proving the reliability and effectiveness of the present invention.

[0055] (VI) Protection scope of the present invention

[0056] The present invention protects a technical method for calibration and bias correction of clinical biochemical assays based on SVM. This method calculates the principal component explained variance through PLS. When the principal component explained variance ≥ 95%, a linear SVM model is selected to construct a calibration equation (target value = K × OD value + B), and the slope K and intercept B are dynamically output to achieve traceable calibration. At the same time, multi-source data such as calibration, EQA, and IQC are processed by Grubbs test and Z-score standardization, and the calibration curve visualization and bias monitoring are realized with the help of the Bohrium platform. The Python script and API interface are used to seamlessly connect with the LIS to complete the full-process automation of "detection - calibration - quality control". Taking MSE < 1.5 and R 2 > 0.98 as the model evaluation criteria and Bias% (±5%) as the error interpretation criteria, the detection accuracy of biochemical items is improved, providing high-reliability support for clinical diagnosis and scientific research.

[0057] The present invention aims to provide a method for constructing a calibration model for clinical biochemical assays based on SVM and bias correction. By collecting multi-source data (including calibration data, EQA data, and IQC data), and combining steps such as data preprocessing, model construction, performance evaluation, and automated integration, it realizes the optimization of the calibration curve and the real-time monitoring of bias, significantly improving the accuracy and reliability of the detection results.

Claims

1. A method for constructing a calibration model for clinical biochemical assays based on support vector machines and correcting bias, characterized in that The following steps are involved: (1) Multi-source data acquisition: Collect calibration data, external quality assessment (EQA) data, and internal quality control (IQC) data of the biochemical analyzer, wherein the calibration data includes the correspondence between absorbance (OD value) and calibration solution concentration, the EQA data includes target value and laboratory measurement value, and the IQC data is the quality control measurement value of continuous monitoring; (2) Data preprocessing: outliers are removed and standardized on the collected data, and the data are divided into training set, validation set and test set; (3) Dynamic model construction: The linear characteristics of the data are determined through principal component analysis. When the variance explained by the principal component reaches the preset threshold, the support vector machine regression algorithm is used to construct an explicit calibration equation and output the slope K and intercept B. If the nonlinear characteristics of the data are significant, the nonlinear algorithm is switched to. (4) Model performance verification: Verify the model based on the mean squared error (MSE), coefficient of determination (R 2 ), and percentage bias (Bias%), and ensure compliance with clinical laboratory standardization specifications; (5) Full-process automated integration: Visual tools are used to achieve calibration curve comparison and real-time bias monitoring, automatically generate calibration reports, and synchronize calibration parameters to the laboratory information management system (LIS).

2. The method according to claim 1, wherein In the step (2), outlier elimination includes: using the Grubbs test (α=0.05) to eliminate data that deviates from the mean by ±20%; eliminating the measured values ​​exceeding ±3SD for the IQC data; and using the Z-score standardization formula to standardize the data of different instruments, the formula being: Among them, X is the original data point, μ is the mean of the data set, and σ is the standard deviation of the data set.

3. The method according to claim 1, wherein In the step (3), the preset threshold of the principal component analysis is explained variance ≥ 95%. At this time, the linear kernel support vector machine regression model is selected, and the penalty coefficient C and the kernel coefficient γ are optimized by the grid search algorithm, where the value range of C is 1-100 and the value range of γ is 0.001-1.

4. The method according to claim 1, characterized in that In the step (4), the criteria for model performance verification are as follows: mean square error MSE < 1.5; coefficient of determination R 2 > 0.98; bias percentage Bias% is within the range of ± 5%.

5. The method according to claim 1, wherein In step (5), the full process automation integration includes: generating a calibration curve comparison chart through the Bohrium platform and monitoring the bias trend in real time; automatically generating a calibration report using a Python script; and synchronizing the calibration parameters to the laboratory information management system (LIS) through an API interface.

6. The method according to claim 1, wherein The explicit calibration equation is: C adjusted = K × m + B, where C adjusted is the calibrated concentration value, m is the measured concentration value before calibration, K is the slope, and B is the intercept.

7. The method according to claim 1, wherein The method is applicable to all biochemical detection items in clinical laboratories in which the absorbance (OD value) and concentration value are linearly related, including but not limited to total protein (TP), albumin (ALB), urea (BUN), direct bilirubin (DBIL), calcium (Ca), and magnesium (Mg).