Method and device for calculating reference interval of biomarker

By extracting biomarker detection results from multiple databases, combining deduplication, removal and reference interval algorithms, the reference interval of biomarker is determined, which solves the problem of lack of systematic methods in the prior art, and achieves low-cost and reliable reference interval establishment.

CN115620819BActive Publication Date: 2025-06-24PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211394489.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-06-24
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

The lack of systematic methods and guidelines for establishing reference intervals for biomarkers using big data, resulting in difficulties and high costs for clinical workers when establishing reference intervals.

Method used

By extracting the detection results of the same biomarker from multiple databases, combining multiple deduplication and removal algorithm processing, the target data set that is most consistent with the normal distribution is determined, and a variety of reference interval algorithms are used to calculate the reference interval, and finally the final result is determined based on consistency and confidence intervals.

Benefits of technology

A low-cost and reliable establishment of reference intervals for biomarkers is achieved, reducing the cost of obtaining reference intervals and improving efficiency, avoiding subjective or one-sided methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620819B_ABST
    Figure CN115620819B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for calculating a reference interval of a biomarker, including: obtaining the detection results of the same biomarker from multiple databases respectively to obtain multiple detection result sets; respectively processing the multiple detection result sets by using multiple duplicate removal algorithms and multiple exclusion algorithms to obtain multiple target data sets of the biomarker; respectively determining the numerical distribution states of the numerical values of the biomarker in the multiple target data sets, and further determining one of the target data sets that most conforms to the normal distribution; respectively using multiple reference interval algorithms to calculate the reference interval of the biomarker based on the target data set; calculating the consistency data of all the reference interval calculation results, and determining whether the consistency data meets a predetermined condition; if the consistency data meets the predetermined condition, determining one of the reference intervals as the final result according to the confidence intervals of the respective reference intervals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biochemical test data processing, and particularly to a method and device for calculating a reference interval of a biomarker. Background Art

[0002] In the medical field, the clinical index reference interval refers to the distribution range of 95% of the measured values of a test index in the healthy population, which is an interval including two endpoints. Establishing a suitable reference interval is crucial for clinical decision-making.

[0003] For various biomarkers, it is necessary to determine the reference interval separately, and there are many choices for the data sources and algorithms used to determine the reference interval. Currently, most of the reference intervals used in domestic clinical laboratories are those from abroad. Although the industry standards currently publish the reference intervals of routine biomarkers, they only cover a small part of the biomarkers. Therefore, it is crucial for laboratories to establish a reference interval suitable for their own laboratories.

[0004] Directly recruiting healthy people to establish a reference interval is costly and infeasible for most laboratories. However, clinical laboratories generate a large amount of real-world data every day. Through a reasonable data cleaning and mining process, a stable and reliable reference interval can be established. This method is low-cost and feasible for most laboratories. However, there are multiple algorithms for establishing a reference interval using real data, which involve multiple links. Currently, there is no clear guideline or systematic method theory to guide researchers to use big data to establish a reference interval. In addition, before establishing a reference interval, a series of issues such as the data type to be used, the sample size, the proportion of pathological data, and the analysis characteristics of the test items need to be considered. This is undoubtedly very difficult for clinical workers. Summary of the Invention

[0005] In view of this, a method for calculating a reference interval of a biomarker according to the present invention includes:

[0006] Obtaining the detection results of the same biomarker from multiple databases respectively to obtain multiple detection result sets, where each of the detection result sets respectively includes the detection results of the biomarker of a large number of individuals;

[0007] Processing the multiple detection result sets respectively by using multiple duplicate removal algorithms and multiple exclusion algorithms to obtain multiple target data sets of the biomarker, the number of which is the product of the number of the detection result sets, the number of the duplicate removal algorithms, and the number of the exclusion algorithms;

[0008] Determining the numerical distribution state of the values of the biomarker in the multiple target data sets respectively, and then determining one of the target data sets that most conforms to the normal distribution;

[0009] Using a variety of reference interval algorithms respectively, calculate the reference interval of the biomarker based on the target data set;

[0010] Calculate the consistency data of all reference interval calculation results, and determine whether the consistency data meets the predetermined conditions;

[0011] If the consistency data meets the predetermined conditions, determine one of the reference intervals as the final result according to the confidence intervals of the respective reference intervals;

[0012] If the consistency data does not meet the predetermined conditions, select a reference interval algorithm according to the numerical distribution state of the target data set, and use the corresponding reference interval as the final result.

[0013] Optionally, respectively determine the normal distribution states of the numerical values of the biomarker in the multiple target data sets, and then determine one of the target data sets with the best normal distribution, specifically including:

[0014] Calculate the kurtosis coefficient and skewness coefficient of each target data set respectively;

[0015] Calculate the normal distribution numerical values of each target data set respectively, where the normal distribution numerical value = |skewness coefficient| + |kurtosis coefficient - 3|;

[0016] Determine the target data set with the smallest such normal distribution numerical value.

[0017] Optionally, calculate the consistency data of all reference interval calculation results, and determine whether the consistency data meets the predetermined conditions, specifically including:

[0018] Calculate the bias matrix of all reference intervals, and then calculate the bias median in the bias matrix;

[0019] The predetermined condition is that the bias median reaches the threshold.

[0020] Optionally, determine one of the reference intervals as the final result according to the confidence intervals of the respective reference intervals, specifically including:

[0021] Calculate the widths of the confidence intervals of the respective reference intervals respectively;

[0022] Use the reference interval with the smallest confidence interval width as the final result.

[0023] Optionally, the multiple databases include the following 7:

[0024] Physical examination individual database;

[0025] Outpatient patient database;

[0026] Inpatient database;

[0027] Combined database of physical examination individuals and outpatients;

[0028] Combined database of physical examination individuals and inpatients;

[0029] Combined database of outpatients and inpatients;

[0030] Combined database of physical examination individuals, outpatients and inpatients.

[0031] Optionally, the multiple deduplication algorithms include the following 4 types:

[0032] For multiple test results of the same individual in the same database, only the first test result is retained;

[0033] For multiple test results of the same individual in the same database, only the last test result is retained;

[0034] For multiple test results of the same individual in the same database, one test result is randomly retained;

[0035] All test results of the same individual are retained for the same database.

[0036] Optionally, the multiple elimination algorithms include the following 5 types:

[0037] Use the Tukey method to eliminate the value of a single test result;

[0038] Use the Tukey method to eliminate the test results of an individual;

[0039] Use the cyclic 4SD method to eliminate the value of a single test result;

[0040] Use the cyclic 4SD method to eliminate the test results of an individual;

[0041] Do not eliminate any test results.

[0042] Optionally, the multiple reference interval algorithms include the following 7 types:

[0043] Combination of Box-Cox algorithm and Hoffmann algorithm;

[0044] Combination of Box-Cox algorithm and Bhattacharya algorithm;

[0045] Combination of Box-Cox algorithm and expectation maximization algorithm;

[0046] kosmic algorithm;

[0047] refineR algorithm;

[0048] Combination of Box-Cox algorithm and parametric method;

[0049] Non-parametric method.

[0050] The present invention also provides a method for calculating a reference interval of a biomarker, including:

[0051] For each database, divide the database into multiple regions according to individual age and / or individual gender information;

[0052] Use the above calculation method to calculate the reference interval of the biomarker for each partition.

[0053] Correspondingly, the present invention also provides a device for calculating a reference interval of a biomarker, including: a processor and a memory connected to the processor; wherein, the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to execute the above method for calculating the reference interval of the biomarker.

[0054] According to the method and device for calculating the reference interval of the biomarker provided by the present invention, for the biomarker for which a reference interval needs to be established, by extracting data from multiple databases and combining the processing of various optional duplicate removal algorithms and various optional elimination algorithms, multiple target data sets can be obtained, covering as many combinations of databases and algorithms as possible; then, according to the numerical distribution state of the biomarker in the data set, a data set that best conforms to the normal distribution is determined to obtain a good data basis; then, all available reference interval algorithms are used to calculate the reference interval based on this data set respectively. If the consistency of all reference intervals is high enough, then an optimal reference interval is further selected according to the confidence interval. If the consistency is not high enough, then the reference interval obtained by using a recommended algorithm according to the pre-established knowledge base is used as the final result. This solution obtains a reliable reference interval through objective factors such as numerical distribution state and consistency, avoids subjectively or one-sidedly establishing a reference interval by a certain method, reduces the cost of obtaining the reference interval and has high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a flowchart of the method for calculating the reference interval of the biomarker provided by the present invention;

[0057] Figure 2Schematic diagram of the reference interval calculation process for a preferred embodiment of the present invention. Detailed implementation manners

[0058] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] The technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0060] The present invention provides a method for calculating a reference interval of a biomarker, which can be executed by an arithmetic device having a memory and a processor, such as a computer or a server, etc. As Figure 1 shown, the method includes the following steps:

[0061] S1. Obtain the detection results of the same biomarker from multiple databases respectively to obtain multiple detection result sets, where each detection result set includes the detection results of the biomarker of a large number of individuals. Specifically, the biomarker referred to in the present application refers to a biochemical index, also known as a biomarker. Regarding the database, it refers to a database that stores the biochemical test results of real individuals, such as the outpatient database of a hospital, the individual examination database of a physical examination center, etc. In actual applications, various databases can be selected according to actual factors, or multiple databases can be fused into one database.

[0062] Assume that the number of optional databases is N1, and the biomarker is denoted as bio, then the number of detection result sets of bio is N1.

[0063] S2. Process the multiple detection result sets respectively by using multiple duplicate removal algorithms and multiple exclusion algorithms to obtain multiple target data sets of the biomarker, and the number thereof is the product of the number of detection result sets, the number of duplicate removal algorithms, and the number of exclusion algorithms.

[0064] There may be the same patient or individual in each data set who has undergone n biochemical examinations at different times and has n detection results. The duplicate removal algorithm refers to adopting a certain rule to remove n - 1 detection results of the patient or individual and only retain 1 detection result of the patient or individual. There are various optional duplicate removal algorithms and can be designed according to actual needs.

[0065] Since the test results of individuals with physical abnormalities or diseases are usually inevitable in most datasets, the biomarker values of such individuals are abnormal in themselves. If the reference interval is calculated using the data including these abnormalities, it may have a greater or lesser impact on the calculation results. The elimination algorithm refers to adopting a certain rule to eliminate the data that may belong to abnormalities based on the biomarker values of all individuals in the same dataset. There are various optional elimination algorithms, and they can be designed according to actual needs.

[0066] Suppose there are N2 kinds of optional duplicate removal algorithms and N3 kinds of optional elimination algorithms, and the results are obtained by processing N1 test result sets respectively. Then the number of target datasets is N1 * N2 * N3.

[0067] S3. Respectively determine the numerical distribution status of the biomarker values of multiple target datasets, and then determine one target dataset that best conforms to the normal distribution. In this field, it is considered that the biomarker values of a large number of individuals should present a normal distribution curve, with smaller numerical differences among most individuals and a smaller number of individuals with large differences. The purpose of this step is to find one target dataset that best conforms to the state distribution from N1 * N2 * N3 target datasets.

[0068] There are various algorithms for judging whether a large amount of data conforms to the normal distribution. For example, generate the numerical distribution curve of the biomarker and judge according to the state of the curve. In a specific embodiment, it is judged based on the kurtosis coefficient and skewness coefficient. In a preferred embodiment, calculate the kurtosis coefficient and skewness coefficient of each target dataset respectively; calculate the normal distribution values of each target dataset, and the normal distribution value = |skewness coefficient| + |kurtosis coefficient - 3|; determine the target dataset with the smallest normal distribution value.

[0069] In the unique target dataset i obtained through the above steps, it is the numerical value of a certain biomarker of a large number of individuals obtained by using a certain elimination method and a certain duplicate removal method for a certain database.

[0070] S4. Respectively use a variety of reference interval algorithms to calculate the reference interval of the biomarker based on the target dataset. There are various known reference interval algorithms, and this method can select any available algorithm to calculate the reference interval for the target dataset i respectively.

[0071] Suppose there are N4 kinds of available reference interval algorithms, then N4 reference intervals can be calculated based on the target dataset i.

[0072] S5. Calculate the consistency data of all reference interval calculation results, and judge whether the consistency data reaches the predetermined condition. If the consistency data reaches the predetermined condition, then execute step S6, otherwise execute step S7.

[0073] As an example, the bias matrix method can be used to determine the consistency of N4 reference intervals. Calculate the median bias in the bias matrix. If the median bias reaches a threshold (the threshold is 0.4 in a specific embodiment), it indicates that these reference intervals are consistent.

[0074] S6. Determine a reference interval as the final result according to the confidence intervals of each reference interval. The confidence interval specifically includes the confidence interval of the upper limit and the confidence interval of the lower limit. For example, a reference interval is (bio L ~bio M ), where the confidence interval of bio L is bio Ll ~bio Lm , and the confidence interval of bio M is bio Ml ~bio Mm . Obviously, the sizes (widths) of the confidence intervals of each reference interval obtained according to the above process are different. In a preferred embodiment, calculate the widths of the confidence intervals of each reference interval respectively; select the reference interval with the smallest confidence interval width as the final result, that is, the one with the smallest confidence interval among the N4 reference intervals is considered the best reference interval.

[0075] S7. Select a reference interval algorithm according to the numerical distribution state of the target data set, and use the corresponding reference interval as the final result. Specifically, pre-specify the numerical distribution states that the various available reference interval algorithms in step S4 are most suitable for processing. More specifically, determine in advance the data sets with what kurtosis coefficients and skewness coefficients that the various reference interval algorithms are suitable for calculating, and thus establish a knowledge base. When the consistency of the N4 reference intervals is insufficient, according to the content of the knowledge base, as well as the kurtosis coefficient and skewness coefficient of the target data set i, determine the reference interval algorithm that is relatively most suitable for calculating the target data set i.

[0076] According to the method for calculating the reference interval of a biomarker provided by an embodiment of the present invention, for a biomarker for which a reference interval needs to be established, by extracting data from multiple databases and combining the processing of multiple optional duplicate removal algorithms and multiple optional exclusion algorithms, multiple target data sets can be obtained, covering as many combinations of databases and algorithms as possible; then, according to the numerical distribution state of the biomarker in the data set, a data set that best conforms to the normal distribution is determined to obtain a good data basis; then all available reference interval algorithms are used to calculate the reference interval based on this data set respectively. If the consistency of all reference intervals is high enough, then an optimal reference interval is further selected according to the confidence interval. If the consistency is not high enough, the reference interval obtained by a recommended algorithm based on a pre-established knowledge base is used as the final result. This solution obtains a reliable reference interval through objective factors such as numerical distribution state and consistency, avoiding subjectively or one-sidedly establishing a reference interval by a certain method, reducing the cost of obtaining the reference interval and having high efficiency.

[0077] As Figure 2 shown, in a preferred embodiment, the selected databases include the following 7: physical examination individual database; outpatient patient database; inpatient patient database; combined database of physical examination individuals and outpatient patients; combined database of physical examination individuals and inpatient patients; combined database of outpatient patients and inpatient patients; combined database of physical examination individuals, outpatient patients and inpatient patients.

[0078] The selected duplicate removal algorithms include the following 4: only retain the first test result for multiple test results of the same individual in the same database (retain the first result); only retain the last test result for multiple test results of the same individual in the same database (retain the last result); randomly retain one test result for multiple test results of the same individual in the same database (randomly retain one result); retain all test results of the same individual in the same database (do not remove duplicates).

[0079] The selected exclusion algorithms include the following 5: use the Tukey method to exclude the value of a single test result; use the Tukey method to exclude the test results of an individual; use the cyclic 4SD method to exclude the value of a single test result; use the cyclic 4SD method to exclude the test results of an individual; do not exclude any test results.

[0080] After the above steps S1 - S2, 7 * 4 * 5 = 140 target data sets can be obtained. Through numerical distribution state analysis, a data set is selected from the 140 target data sets.

[0081] The selected reference interval algorithms include the following 7 types: the combination of the Box-Cox algorithm and the Hoffmann algorithm; the combination of the Box-Cox algorithm and the Bhattacharya algorithm; the combination of the Box-Cox algorithm and the expectation maximization algorithm; the kosmic algorithm; the refineR algorithm; the combination of the Box-Cox algorithm and the parametric method; the non-parametric method.

[0082] For this one target data set, 7 reference intervals can be calculated, and then through steps S5 - S7, the final result is obtained, that is, one reference interval is selected.

[0083] Taking adult free triiodothyronine (a biomarker) as an example, according to the above method, the lower limit of the reference interval is obtained as: 2.85, and its confidence interval is (2.841 - 2.868); the upper limit of the reference interval is: 3.98, and its confidence interval is (3.975 - 3.988). The specific path is to establish the reference interval and its 90% confidence interval of male FT3 based on the physical examination individual database using the transformed Hoffmann method.

[0084] As a comparison, after strict inclusion and exclusion criteria, including but not limited to excluding individuals with abnormal blood pressure, BMI, and thyroid ultrasound, and using the biological test data of a large number of standard individuals, the obtained reference interval is: lower limit: 2.83 (2.805 - 2.855), upper limit: 3.93 (3.895 - 3.976). This reference interval is the "gold standard" in this field. In this field, the "gold standard" generally results from using a large amount of very ideal individual data and can be considered the most accurate reference interval. The problem is that if you want to obtain such a result, a large amount of individual data needs to be collected, and individuals need to be screened to exclude various individuals that may affect the accuracy. The cost of this calculation method is very high. The key is that not all biomarkers have a known "gold standard", and the "gold standard" may not be suitable for all populations. For example, the reference intervals are different for different countries or regions and different ethnic groups. The gold standard cited in this application is only for comparison with this solution.

[0085] Through comparison, it can be found that the reference interval obtained according to this method is very close to this "gold standard", indicating that the method provided in this embodiment has high accuracy.

[0086] The present invention can set a differential optimal algorithm path for establishing a reference interval, including outlier rejection method selection, sample size calculation, reference interval partition analysis, separation of mixed distribution algorithm selection, etc., according to the characteristics of the data distribution parameters of different test analytes, including kurtosis and skewness coefficient and other parameter values. It can guide clinical researchers to establish accurate and reliable reference intervals at low cost, filling the gap in the current such method system.

[0087] In addition, for some biomarkers, the reference intervals are different for different populations, so it is necessary to establish reference intervals for different populations separately. For example, for the free triiodothyronine in adults mentioned in the above example, there are differences in the values between men and women; or for some biomarkers, there are significant differences in the values between infants and adults. In the face of such biomarkers, it is necessary to establish reference intervals separately.

[0088] To this end, the embodiments of the present invention provide a method for calculating the reference interval of a biomarker. First, for each database, the database is divided into multiple regions according to individual age and / or individual gender information. For example, it can be divided into two regions by gender, two regions by age group, or three or four regions by age and gender at the same time, etc. Then, using the method of the previous embodiment, the reference interval of the biomarker is calculated for each partition.

[0089] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0090] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0091] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.

[0093] Obviously, the above-described embodiments are merely examples for clear illustration, and are not intended to limit the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. The obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.

Claims

1. A method for calculating the reference interval of a biomarker, characterized in that, Including: Obtain the detection results of the same biomarker from multiple databases respectively to obtain multiple detection result sets, where each of the detection result sets respectively includes the detection results of the biomarker of a large number of individuals; Process the multiple detection result sets respectively by using a variety of duplicate removal algorithms and a variety of elimination algorithms to obtain multiple target data sets of the biomarker, and the number thereof is the product of the number of the detection result sets, the number of the duplicate removal algorithms and the number of the elimination algorithms; Determine the numerical distribution status of the values of the biomarker in the multiple target data sets respectively, and then determine one of the target data sets that most conforms to the normal distribution; Use a variety of reference interval algorithms respectively to calculate the reference interval of the biomarker based on the target data set; Calculate the consistency data of all the reference interval calculation results, and determine whether the consistency data reaches a predetermined condition; If the consistency data reaches the predetermined condition, then determine one of the reference intervals as the final result according to the confidence interval of each reference interval; If the consistency data does not reach the predetermined condition, then select a reference interval algorithm according to the numerical distribution status of the target data set, and use the corresponding reference interval as the final result.

2. The reference interval calculation method according to claim 1, wherein Determine the normal distribution status of the values of the biomarker in the multiple target data sets respectively, and then determine one of the target data sets with the best normal distribution, specifically including: Calculate the kurtosis coefficient and skewness coefficient of each target data set respectively; Calculate the normal distribution value of each target data set, and the normal distribution value = |skewness coefficient| + |kurtosis coefficient - 3|; Determine the target data set with the smallest normal distribution value.

3. The reference interval calculation method according to claim 1, characterized in that Calculate the consistency data of all the reference interval calculation results, and determine whether the consistency data reaches a predetermined condition, specifically including: Calculate the bias matrix of all the reference intervals, and then calculate the median bias in the bias matrix; The predetermined condition is that the median bias reaches the threshold.

4. The reference interval calculation method according to claim 1, characterized in that, Determine one of the reference intervals as the final result according to the confidence interval of each reference interval, specifically including: Calculate the width of the confidence interval of each reference interval respectively; Use the reference interval with the smallest confidence interval width as the final result.

5. The reference interval calculation method according to any one of claims 1-4, characterized in that, The multiple databases include the following 7: Physical examination individual database; Outpatient patient database; Inpatient patient database; Combined database of physical examination individuals and outpatient patients; Combined database of physical examination individuals and inpatient patients; Combined database of outpatient patients and inpatient patients; Combined database of physical examination individuals, outpatient patients and inpatient patients.

6. The reference interval calculation method according to any one of claims 1-4, characterized in that, The variety of duplicate removal algorithms include the following 4: Only retain the first detection result for multiple detection results of the same individual in the same database; Only retain the last detection result for multiple detection results of the same individual in the same database; Randomly retain one detection result for multiple detection results of the same individual in the same database; Retain all the detection results of the same individual in the same database.

7. The reference interval calculation method according to any one of claims 1-4, characterized in that The variety of elimination algorithms include the following 5: Use the Tukey method to eliminate the value of a single detection result; Use the Tukey method to eliminate the detection results of an individual; The single test result value is removed using the cyclic 4SD method; The test results of individuals are removed using the cyclic 4SD method; Do not remove any test results.

8. The reference interval calculation method according to any one of claims 1-4, characterized in that, The multiple reference interval algorithms include the following seven: The combination of the Box-Cox algorithm and the Hoffmann algorithm; The combination of the Box-Cox algorithm and the Bhattacharya algorithm; The combination of the Box-Cox algorithm and the expectation maximization algorithm; The kosmic algorithm; The refineR algorithm; The combination of the Box-Cox algorithm and the parametric method; The non-parametric method.

9. A method for calculating a reference interval of a biomarker, characterized in that, Include: For each database, the database is divided into multiple regions according to the individual age and / or individual gender information; The reference interval of the biomarker is calculated for each partition using the method described in any one of claims 1-8.

10. A reference interval calculation device for a biomarker, characterized in that Include: A processor and a memory connected to the processor; wherein, the memory stores instructions executable by the processor, and the instructions are executed by the processor to cause the processor to execute the method for calculating the reference interval of the biomarker described in any one of claims 1-9.

Citation Information

Patent Citations

  • Genome copy number variation detection method and device

    CN111916150A

  • Reference interval generation

    CN111936037A