Laboratory ability verification statistical analysis method based on isolated forest and LOF
By integrating the isolation forest and LOF algorithms, the identification of global and local abnormal samples in laboratory proficiency testing is achieved, which improves the scientificity and fairness of the test results and is suitable for complex testing scenarios.
Patent Information
- Application Number
- CN202510718059.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing laboratory proficiency testing methods are unable to effectively identify outliers when faced with high-dimensional, nonlinear or complex structured data, which affects the scientificity and fairness of the test results.
The isolation forest and local outlier factor (LOF) algorithms are combined to generate a comprehensive anomaly score through global and local anomaly detection to identify abnormal samples in laboratory proficiency testing.
It improves the accuracy and stability of abnormal sample identification in laboratory proficiency testing, is applicable to multi-type and heterogeneous data, has efficient computing capabilities, and supports the visual expression of multi-dimensional detection data.
Smart Images

Figure CN120655148A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis, and in particular relates to a laboratory capability verification statistical analysis method based on isolation forest and LOF. Background Art
[0002] Proficiency Testing (PT), as an essential component of a quality management system, is a core method for measuring the technical capabilities of testing organizations and ensuring the accuracy and comparability of test results. Conducting PT not only effectively enhances laboratory capacity building but is also a key step in achieving quality assurance and continuous improvement. It plays a particularly crucial role in product quality inspection. The results of PT are directly related to the fairness of product quality assessments, impacting corporate reputation, product market access, and even international trade.
[0003] In recent years, laboratory proficiency testing has received increasing attention from the International Laboratory Accreditation Cooperation and laboratory accreditation organizations of various countries. The construction of proficiency testing programs and information platforms in various industries has also been accelerated, effectively improving laboratory quality management and technical levels, and is an important support path for building a national quality infrastructure. Internationally, some countries have explicitly required laboratories to actively participate in proficiency testing activities and use proficiency testing results as one of the core bases for accreditation applications, supervisory reviews, and maintenance of technical capabilities. Therefore, whether or not a laboratory passes proficiency testing has become a basic condition for obtaining accreditation qualifications.
[0004] With the development of detection technology and the growing demand for complex testing, laboratory proficiency testing faces multiple challenges, such as high data dimensionality, strong sample heterogeneity, and complex coupling of detection parameters. In particular, in scenarios such as trace analysis, multi-component quantification, and nonlinear detection, the volatility of test results between laboratories has increased, and outlier identification has become a key issue affecting the scientificity and fairness of proficiency testing. At present, mainstream statistical evaluation methods for proficiency testing mainly include the z-score method, the En value method, and the interquartile range (IQR) method. These methods are usually based on the normal distribution assumption or low-dimensional central tendency to evaluate the rationality of data, and have certain operability and comparability. However, when faced with high-dimensional, nonlinear, or complex data, such methods have obvious limitations in outlier identification and systematic error judgment.
[0005] In response to the above technical bottlenecks, the gradual introduction of artificial intelligence and machine learning into statistical analysis for proficiency testing provides new ideas for promoting its intelligence. Isolation Forest is an unsupervised anomaly detection method based on random partitioning. It has good computational efficiency and model interpretability and is suitable for global anomaly identification of medium and large-scale experimental data. The Local Outlier Factor (LOF) analyzes the degree of deviation of samples in the local density space and can effectively identify edge-type or locally perturbed anomaly data, making up for the shortcomings of the global model. The performance of a single model in different scenarios has limitations. Isolation Forest is more sensitive to global sparse anomalies, while LOF has advantages in identifying local perturbation anomalies. Therefore, combining the complementary characteristics of the two algorithms, an integrated outlier identification model is constructed to significantly improve the accuracy and stability of abnormal sample identification in laboratory proficiency testing.
[0006] In summary, laboratory proficiency testing, as a key link in quality management and accreditation review, is evolving from traditional statistical analysis models to intelligent and integrated ones. The development of an outlier identification method that integrates isolation forests and LOF can effectively take into account both global and local outlier detection capabilities, and provide strong technical support for the scientific and fair proficiency testing in complex testing scenarios. It has important research significance and broad application prospects. Summary of the Invention
[0007] The main purpose of the present invention is to overcome the shortcomings of the existing technology and provide an outlier identification method that integrates the isolation forest and LOF algorithms. First, laboratory data is collected and standardized, and an isolation forest model is built to achieve global outlier detection of the data. The isolation forest anomaly score of the sample is calculated, laying the foundation for subsequent comprehensive laboratory evaluation. Then, LOF is used to calculate the local outlier factor to avoid the inability to effectively detect abnormal data of marginal or local disturbance types, thereby improving the accuracy and adaptability of the detection method in practical application scenarios. Finally, the two statistical analysis methods are integrated to calculate the comprehensive score of laboratory samples by setting an appropriate ratio, which serves as an important basis for experts to judge laboratory capabilities.
[0008] The present invention is achieved through the following technical solution: a laboratory proficiency testing statistical analysis method based on isolation forest and LOF, comprising the following steps:
[0009] S1. Collect test result data of the same proficiency testing sample from several laboratories, perform standardization processing, generate standardized samples, and provide a unified data scale for subsequent outlier detection;
[0010] S2. Input the standardized sample data after the standardization process in step S1 into the isolation forest model, and use the isolation forest algorithm to perform global anomaly detection on each standardized sample to generate a global anomaly score;
[0011] S3, using the local outlier factor (LOF) algorithm to perform local density anomaly detection on the standardized samples after the standardization process in step S1, by comparing the density difference between the sample point and its neighboring samples, judging the local anomaly of each standardized sample and generating a local anomaly score;
[0012] S4. Normalize the global anomaly score determined in step S2 and the local anomaly score determined in step S3, and then perform a weighted calculation to determine a comprehensive anomaly score;
[0013] S5. Classify each sample into three categories: "normal value", "suspicious value" and "extreme outlier" according to the dual threshold mechanism;
[0014] S6. Output laboratory proficiency testing assessment report, including abnormal laboratory identification results, visual charts and abnormal sample feature analysis.
[0015] Furthermore, in step S2, the parameters of the isolation forest algorithm include: the number of trees is 50 to 500, and the number of sub-sampling samples is 128 to 512.
[0016] Furthermore, in step S3, the number of nearest neighbors of each standardized sample in the LOF algorithm is k, and its value range is 10 to 30.
[0017] Furthermore, in step S4, the value of the weight parameter α of the weighted calculation is 0.4 to 0.6, preferably 0.5.
[0018] Furthermore, in step S5, the first threshold in the dual-threshold mechanism is set to 0.85, and the second threshold is set to 0.95.
[0019] Furthermore, the present invention is used for capability verification data containing two or more detection indicators, and supports comprehensive analysis and visual expression of abnormalities in multi-dimensional detection data.
[0020] The beneficial effects of the present invention are:
[0021] (1) Strong model complementarity: The integration of isolation forest and LOF algorithms takes into account both global sparsity and local anomalies, and the identification is more comprehensive; (2) No distribution assumption: It does not rely on the normal distribution characteristics of the data and is applicable to multi-type and heterogeneous experimental results; (3) High efficiency and scalability: The present invention is adaptable to large-scale data, has high computational efficiency, and can be integrated into the information platform; (4) Visualization and interpretability: The output of comprehensive anomaly scores and classification results can intuitively present the source and distribution of anomalies. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION
[0023] This paper takes laboratory proficiency testing as the research background and designs a set of laboratory proficiency testing statistical analysis methods based on isolation forest and LOF algorithms. The designed method mainly consists of four parts: (1) data collection and standardization; (2) isolation forest model training and scoring; (3) LOF anomaly factor calculation; (4) anomaly score fusion and outlier classification. In order to assist in understanding the results of model analysis, the present invention has designed a visualization module to draw box plots, radar charts, etc. for each detection indicator to show the degree of deviation of abnormal laboratories in each indicator dimension and assist in identifying the source of anomalies. The final output analysis report includes: the comprehensive score of each laboratory, the abnormal level to which it belongs, the main deviation indicator description and processing suggestions. The report can be directly applied to proficiency testing summary and feedback, and improve the efficiency of laboratory proficiency testing experts in interpreting results and making decisions.
[0024] The present invention will be described in further detail below with reference to the accompanying drawings and examples.
[0025] like Figure 1 The statistical analysis method for laboratory proficiency testing based on isolation forest and LOF shown includes the following steps.
[0026] S1. Data collection and standardization: Collect the test result data of the same proficiency testing sample from several laboratories, perform standardization processing, and generate standardized samples to provide a unified data scale for subsequent outlier detection.
[0027] Laboratory proficiency testing (PCT) begins with collecting test results from multiple laboratories for the same batch of samples. For example, a typical sample contains four items: methanol, lead, ethyl acetate, and iron. Different laboratories measure their contents using established methods, forming an n×m raw data matrix, where n represents the number of participating laboratories and m represents the number of test items.
[0028] Since different items have different units and dimensions, direct statistical analysis may lead to model bias or failure, so the raw data needs to be normalized. In this example, the Z-score normalization method is used to convert the data of each item into a standard normal distribution with a mean of 0 and a standard deviation of 1. The data after Z-score normalization avoids scoring bias caused by absolute numerical differences and provides a unified data scale basis for subsequent outlier detection.
[0029] The Z-score standardization calculation formula is:
[0030]
[0031] In formula (1),
[0032] Z ij ——The standardized score of the i-th laboratory in the j-th project;
[0033] X ij ——The original measurement value of the jth item in the i-th laboratory;
[0034] μ j ——the average value of the jth item;
[0035] σ j ——The standard deviation of the j-th item.
[0036] S2, Isolation Forest Model Training and Scoring: Input the standardized sample data after the standardization process in step S1 into the Isolation Forest Model, use the Isolation Forest Algorithm to perform global anomaly detection on each standardized sample, and generate a global anomaly score;
[0037] The standardized data will be fed as input into the main model of the present invention - isolation forest for preliminary anomaly detection. Isolation forest is an unsupervised learning algorithm built on the principle of random segmentation. Its core idea is that "the easier it is to randomly isolate a sample, the more likely it is to be an anomaly." The isolation forest model divides the data step by step by constructing multiple isolated trees, each tree randomly selecting a feature and its segmentation point in each split. For each sample point, the average path length required for it to be individually isolated in all isolated trees is counted, and then its anomaly score is calculated. In this embodiment, the number of trees in the isolation forest algorithm is 200, and the number of sub-sampling samples is 256, which can make the model more robust within the limitations of computing resources.
[0038] The global anomaly scoring function is as follows:
[0039]
[0040] In formula (2),
[0041] s(x)——Isolation Forest Anomaly Score of sample x, ranging from [0,1];
[0042] E(h(x))——the average path length of sample x in the isolation forest;
[0043] c(n)——sample size normalization constant, which is related to the sample subset size n. The calculation formula is:
[0044]
[0045] H(i)≈ln(i)+0.5772; (4)
[0046] In formula (3) and formula (4),
[0047] n – the sample subset size for each tree in the isolation forest;
[0048] H(i) is the i-th harmonic number, approximately calculated as the sum of the natural logarithm of i and Euler's constant.
[0049] The Isolation Forest algorithm has excellent global modeling capabilities and can efficiently identify outlier laboratories that deviate significantly from the main population in high-dimensional, nonlinear data structures. When the score s(x) approaches 1, it indicates that the sample is easily isolated and is a potential outlier. When the score is less than 0.5, it indicates that the structure is relatively robust and is a normal sample. The score s(x) can be used to judge data quality.
[0050] S3, LOF anomaly factor calculation: Use the local outlier factor algorithm to perform local density anomaly detection on the standardized samples after the standardization process in step S1. By comparing the density difference between the sample point and its neighboring samples, the local anomaly of each standardized sample is determined and a local anomaly score is generated;
[0051] In order to improve the credibility of the outlier detection model results and the ability to identify local anomalies, the present invention introduces an auxiliary model, LOF, for verification based on the isolation forest judgment. LOF judges the local anomaly of the sample by comparing the density difference between the sample point and its neighboring samples. This method is suitable for identifying samples at the edge of the data or in low-density areas, and specifically includes the following steps.
[0052] First, the number of nearest neighbors of each standardized sample is k, which is 20, which can take into account the stability and sensitivity of the model. The k nearest neighbor set N of each sample is determined. k (x), calculate the reachable distance of each neighborhood sample:
[0053]
[0054] In formula (5),
[0055] reach_dist k (x,y) – the reachable distance from point x to its neighboring point y;
[0056] d(x,y) — Euclidean distance between point x and its neighboring point y;
[0057] k_dist(y) – the k-th nearest neighbor distance of point y.
[0058] Then, calculate the local reachability density (LRD) of sample x:
[0059]
[0060] In formula (6),
[0061] lrd k (x)——local reachability density of sample x;
[0062] N k (x)——the set of k nearest neighbors of sample x;
[0063] |N k (x)|——The number of samples in the nearest neighbor set.
[0064] Finally, the local outlier factor (LOF) of the sample is obtained:
[0065]
[0066] In formula (7),
[0067] LOF k (x)——local outlier factor score of sample x;
[0068] lrd k (y)——local reachability density of neighboring sample y;
[0069] lrd k (x)——the local reachable density of sample x itself.
[0070] When LOF k When (x) > 1 and is much larger than 1, it indicates that the density of sample x is much lower than that of its neighboring samples and is a local outlier. Otherwise, the sample is a normal point. The introduction of LOF can effectively capture marginal outliers that may be missed by the isolation forest, thus forming a "global-local" dual anomaly recognition structure and enhancing the robustness of the overall model.
[0071] S4, anomaly score fusion and outlier classification: normalize the global anomaly score determined in step S2 and the local anomaly score determined in step S3, and then perform weighted calculation to determine a comprehensive anomaly score;
[0072] In order to integrate the recognition results of the isolation forest and LOF models, the present invention designs an anomaly score fusion mechanism, which specifically includes the following steps.
[0073] First, the global anomaly score determined in step S2 and the local anomaly score determined in step S3 are subjected to minimum-maximum normalization processing respectively, and the calculation formula is as follows:
[0074]
[0075] In formula (8) and formula (9),
[0076] S IF ——the isolation forest raw anomaly score of the sample;
[0077] S LOF ——The LOF raw anomaly score of the sample;
[0078] ——Normalized score.
[0079] Then, a weighted fusion method is used to calculate the comprehensive anomaly score:
[0080]
[0081] In formula (10),
[0082] α——weight coefficient, α∈[0,1]; α preferably ranges from 0.4 to 0.6, and in this embodiment, a more preferred value is 0.5;
[0083] Score - the final comprehensive anomaly score of the sample.
[0084] S5. Classify each sample into three categories: "normal value", "suspicious value" and "extreme outlier" according to the dual threshold mechanism;
[0085] The present invention can be adjusted according to actual conditions when applied, for example, by tuning with reference to historical data. A dual-threshold classification mechanism is set according to the comprehensive score Score to determine whether a data point is outlier. This threshold can also be adjusted according to actual conditions: the first threshold in the dual-threshold mechanism is set to 0.85, and the second threshold is set to 0.95, that is, when the comprehensive score satisfies Score<0.85, it is judged as a normal sample; when the comprehensive score satisfies 0.85≤Score<0.95, it is judged as a suspicious sample; when the comprehensive score satisfies Score≥0.95, it is judged as an extremely abnormal sample. This classification mechanism can provide multi-level discrimination results for specific samples, which is convenient for experts to further analyze and judge the detection capabilities of each laboratory.
[0086] S6. Output laboratory proficiency testing assessment report, including abnormal laboratory identification results, visual charts and abnormal sample feature analysis, for proficiency testing data containing two or more test indicators, and support abnormal comprehensive analysis and visual expression of multi-dimensional test data.
[0087] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A statistical analysis method for laboratory proficiency testing based on isolation forests and LOF, characterized in that: The following steps are involved: S1. Collect test result data of the same proficiency testing sample from several laboratories, perform standardization processing, generate standardized samples, and provide a unified data scale for subsequent outlier detection; S2. Input the standardized sample data after the standardization process in step S1 into the isolation forest model, use the isolation forest algorithm to perform global anomaly detection on each standardized sample, and generate a global anomaly score; S3, using the local outlier factor algorithm to perform local density anomaly detection on the standardized samples after the standardization process in step S1, by comparing the density difference between the sample point and its neighboring samples, judging the local anomaly of each standardized sample and generating a local anomaly score; S4. Normalize the global anomaly score determined in step S2 and the local anomaly score determined in step S3, and then perform a weighted calculation to determine a comprehensive anomaly score; S5. Classify each sample into three categories: "normal value", "suspicious value" and "extreme outlier" based on the dual threshold mechanism; S6. Output laboratory proficiency testing assessment report, including abnormal laboratory identification results, visual charts and abnormal sample feature analysis.
2. The method according to claim 1, characterized in that In step S2, the parameters of the isolation forest algorithm include: the number of trees is 50 to 500, and the number of sub-sampling samples is 128 to 512.
3. The method according to claim 1, characterized in that In step S3, the number of nearest neighbors of each standardized sample in the LOF algorithm is k, and the value range is 10 to 30.
4. The method according to claim 1, wherein In step S4, the value of the weight parameter α of the weighted calculation is 0.4 to 0.6, preferably 0.
5.
5. The method according to claim 1, wherein In step S5, the first threshold in the dual-threshold mechanism is set to 0.85, and the second threshold is set to 0.
95.
6. The method according to claim 1, wherein The method is used for capability verification data containing two or more detection indicators, and supports comprehensive analysis and visual expression of abnormalities in multi-dimensional detection data.
Citation Information
Patent Citations
System and method for detecting data drift
CA3085092A1
Method for judging doubtful values of sample detection data
CN106557652A
Time series data anomaly monitoring system and method based on LOF and isolated forest
CN115577275A
Monitoring method and device for testing equipment stability, electronic equipment and storage medium
CN117827568A
Method for verifying and comprehensively evaluating multi-parameter capability of ecological environment based on covariance distance
CN117993774A