Gender and age-based test item reference interval construction system
By constructing a reference interval system for test items based on gender and age, and utilizing various big data algorithms to group and deduplicate test data, the system solves the problem of inaccurate reference intervals in existing technologies, achieving more efficient and accurate reference interval construction, reducing costs and improving data reliability.
Patent Information
- Application Number
- CN202510377712.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing technologies struggle to accurately establish reference intervals for testing based on gender and age, resulting in inaccurate reference intervals, wasted medical resources, and indirect methods suffer from data bias and computational complexity.
The system constructs reference intervals for test items based on gender and age, utilizing various big data algorithms to group and deduplicate test data, and builds continuous reference intervals. It includes modules for test data collection, deduplication, data filtering, grouping, data distribution description, and reference interval determination. The system is processed using the GAMLSS model and undergoes consistency evaluation and pass rate verification.
It improves the accuracy and appropriateness of reference intervals, reduces costs, minimizes unnecessary diagnoses and treatments, and enhances the reliability and consistency of data.
Smart Images

Figure CN119889561B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of reference interval establishment and verification of test items, and in particular to a test item reference interval construction system based on gender and age. BACKGROUND
[0002] Reference interval is an important standard for clinical interpretation of laboratory test results and judgment of health or disease state. There are significant physiological differences between different age groups and gender groups, which can affect the reference interval of some test items. However, most laboratories have difficulty in establishing reference intervals according to gender and age, resulting in inaccurate reference intervals and unnecessary diagnosis and treatment of patients, wasting medical resources. The establishment of reference interval is divided into direct method and indirect method. The direct method consumes a large amount of time, resources and cost, and it is difficult to obtain the best results for age and gender related items. In clinical practice, it is of great significance to establish a reference interval suitable for the population served by the laboratory based on laboratory real world big data using indirect method for disease diagnosis, treatment and prevention.
[0003] However, the actual algorithm model used at present also has some problems. For example: 1. The reference interval establishment and verification data analysis is complex, time-consuming and has high technical bottlenecks, and it is difficult to realize functions such as data deduplication, threshold exclusion, non-normal data conversion, intelligent grouping, and big data construction algorithm; 2. Due to the difference in data distribution characteristics, there are differences between the reference intervals calculated by different algorithms, making it difficult to choose the appropriate algorithm; 3. The indirect method algorithm has limitations on the proportion of pathological data mixed in the data. When the mixed pathological data exceeds the limit, the established reference interval may deviate.
[0004] Therefore, it is of great urgency to further understand the characteristics and principles of the indirect method and to construct a suitable reference interval establishment scheme. SUMMARY
[0005] An advantage of the present application is to provide a test item reference interval construction system based on gender and age, wherein the test item reference interval construction system based on gender and age groups the test data according to gender and age, and establishes the corresponding reference interval for each age group of different genders using a variety of big data algorithms, which can improve the accuracy and suitability of the reference interval to a certain extent.
[0006] Another advantage of the present application is to provide a gender and age based test item reference interval construction system, wherein the collection of test data of healthy subjects in the gender and age based test item reference interval data processing system comes from real world patient medical records, and the source and use of the medical records are legal and compliant, so that on the one hand, the data source is reliable, and on the other hand, it is not necessary to specially recruit healthy subjects for detection in order to conduct reference interval experiments, which can improve the feasibility and reduce the cost to a certain extent.
[0007] According to an aspect of the present application, a gender and age based test item reference interval construction system is provided, comprising:
[0008] a test data collection module configured to obtain a collection of test data of healthy subjects from a laboratory information system;
[0009] a deduplication module configured to deduplicate the collection of test data of healthy subjects based on age, gender and name or medical record number to obtain a collection of deduplicated test data of healthy subjects;
[0010] a data screening module configured to process the collection of deduplicated test data of healthy subjects based on inclusion and exclusion criteria to obtain a collection of candidate test data of healthy subjects;
[0011] a grouping module configured to group the collection of candidate test data of healthy subjects based on age and gender to obtain a grouping result, the grouping result comprising a plurality of data subsets of candidate test data of healthy subjects;
[0012] a data distribution description module configured to perform outlier processing and data distribution description on each data subset of candidate test data of healthy subjects in the grouping result to obtain a plurality of data distribution description results;
[0013] a reference interval determination module configured to determine a reference interval for each data subset of candidate test data of healthy subjects based on the data distribution description result of the data subset using a plurality of big data algorithms; and
[0014] a continuous reference interval construction module configured to construct a continuous reference interval based on gender and age and the plurality of reference intervals.
[0015] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the continuous reference interval construction module is further configured to process the plurality of reference intervals using a GAMLSS model to obtain the continuous reference interval.
[0016] In an embodiment of the gender and age based test item reference interval construction system according to the present application, further comprising: a consistency evaluation module configured to evaluate the consistency of the results of different construction algorithms of the reference interval of each data subset of the candidate health examinee test data; and a passing rate verification module configured to verify the passing rate of the reference interval of each data subset of the candidate health examinee test data.
[0017] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the deduplication module is further configured to retain only the most recent test data of each health examinee in the set of health examinee test data.
[0018] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the data screening module is further configured to exclude abnormal health examinee test data affecting the test item reference interval from the set of deduplicated health examinee test data based on pre-prepared inclusion and exclusion criteria.
[0019] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the grouping module is further configured to group the set of candidate health examinee test data based on age and gender by Harris & Boyd method or standard deviation ratio method.
[0020] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the data distribution description module is further configured to perform distribution description on the candidate health examinee test data to obtain the plurality of data distribution description results, wherein the manner of performing distribution description on the candidate health examinee test data is selected from at least one of the following manners: probability graph, quantile graph, kernel density graph, Shapiro-Wilk test, Kolmogorov-Smirnov test; and remove outliers of each data subset of the candidate health examinee test data in the grouping result to obtain outlier processed candidate health examinee test data.
[0021] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the reference interval of each data subset of the candidate health examinee test data is determined by using the following big data algorithms: EP28 non-parametric method, EP28 parametric method, refineR algorithm, Kosmic algorithm, TMC algorithm.
[0022] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the consistency evaluation module is further configured to calculate a bias ratio of the reference interval established by the rest of the algorithms based on the reference interval established by the EP28 non-parametric method.
[0023] In an embodiment of the gender and age based test item reference interval construction system according to the present application, the pass rate verification module is further configured to verify the reference interval by obtaining physical examination data of test items in another time period from the laboratory information system as verification data; wherein, when the reference interval pass rate is greater than or equal to 90%, the reference interval passes the verification; otherwise, the reference interval fails the verification.
[0024] The further objects and advantages of the present application will be more readily understood from the following description, with reference to the accompanying drawings, in which:
[0025] These and other objects, features and advantages of the present application will become apparent from the following detailed description of the application, with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0026] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which:
[0027] Figure 1 FIG. 1 illustrates a framework diagram of a gender and age based test item reference interval construction system according to an embodiment of the present application.
[0028] Figure 2 FIG. 2 illustrates a flow diagram of an example of constructing a test item reference interval by a gender and age based test item reference interval construction system according to an embodiment of the present application.
[0029] Figure 3 FIG. 3 illustrates a scenario diagram of data distribution description of a data subset of test data of each alternative healthy examinee in the grouping result by a gender and age based test item reference interval construction system according to an embodiment of the present application.
[0030] Figure 4 FIG. 4 illustrates a scenario diagram of constructing a continuous reference interval by a gender and age based test item reference interval construction system according to an embodiment of the present application.
[0031] Figure 5The illustration shows another scenario diagram of a system for constructing continuous reference intervals for test items based on gender and age, according to an embodiment of this application.
[0032] Figure 6 The illustration shows another framework diagram of a reference interval construction system for gender and age-based test items according to an embodiment of this application. Detailed Implementation
[0033] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0034] It is understood that the term "a" should be understood as "at least one" or "one or more," meaning that in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple. The term "a" should not be construed as a limitation on the quantity. "Multiple" means two or more.
[0035] While ordinal numbers such as “first,” “second,” etc., will be used to describe various components, there is no limitation on which components are used herein. The term is used only to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component, without departing from the teachings of this application. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.
[0036] The terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting. As used herein, the singular form also includes the plural form, unless the context clearly indicates otherwise. It will also be understood that the terms “comprising” and / or “having” as used in this specification specify the presence of the described features, numbers, operations, components, elements or combinations thereof, without excluding the presence or addition of one or more other features, numbers, operations, components, elements or combinations thereof.
[0037] like Figures 1 to 6 As shown, a gender and age-based reference interval construction system 100 for laboratory tests, according to an embodiment of this application, is illustrated. This gender and age-based reference interval data processing system groups test data according to gender and age, and establishes reference intervals corresponding to different age groups for different genders, which can improve the accuracy of the reference intervals to a certain extent.
[0038] In particular, the gender and age based test item reference interval data processing system comprises a test data collection module 10, a deduplication module 20, a data screening module 30, a grouping module 40, a data distribution description module 50, a reference interval determination module 60 and a continuous reference interval construction module 70. The test data collection module 10 is configured to obtain a set of health examination test data from a Laboratory Information System (LIS); the deduplication module 20 is configured to deduplicate the set of health examination test data based on age, gender and name or medical record number to obtain a set of deduplicated health examination test data; the data screening module 30 is configured to process the set of deduplicated health examination test data based on inclusion and exclusion criteria to obtain a set of candidate health examination test data; the grouping module 40 is configured to group the set of candidate health examination test data based on age and gender to obtain a grouping result, the grouping result comprising a plurality of data subsets of candidate health examination test data; the data distribution description module 50 is configured to perform outlier processing and data distribution description on each data subset of candidate health examination test data in the grouping result to obtain a plurality of data distribution description results; the reference interval determination module 60 is configured to determine a reference interval for each data subset of candidate health examination test data based on the data distribution description result of the data subset using a plurality of big data algorithms; and the continuous reference interval construction module 70 is configured to construct a continuous reference interval based on the plurality of reference intervals.
[0039] Correspondingly, in the process of constructing the test item reference interval, the gender and age based test item reference interval data processing system first acquires a set of health examination test data from a laboratory information system through the test data acquisition module 10; then, based on age, gender and name or medical record number, the set of health examination test data is de-duplicated through the de-duplication module 20 to obtain a set of de-duplicated health examination test data; then, based on inclusion and exclusion criteria, the set of de-duplicated health examination test data is processed through the data screening module 30 to obtain a set of candidate health examination test data; then, based on age and gender, the set of candidate health examination test data is grouped through the grouping module 40 to obtain a grouping result, the grouping result including a plurality of data subsets of candidate health examination test data; then, each data subset of candidate health examination test data in the grouping result is processed for outlier and data distribution description through the data distribution description module 50 to obtain a plurality of data distribution description results; thereafter, based on the data distribution description results of each data subset of candidate health examination test data, a plurality of big data algorithms are used to determine the reference interval of each data subset of candidate health examination test data through the reference interval determination module 60; further, based on the plurality of reference intervals, a continuous reference interval is constructed through the continuous reference interval construction module 70.
[0040] In the process of acquiring a set of health examination test data from a laboratory information system through the test data acquisition module 10, the set of health examination test data is a set of real detection data of health examination subjects, which comes from patient medical record data in the real world, i.e., using real patients as health examination subjects, and selecting data in medical record data of a plurality of real patients as the set of health examination test data; in this way, on the one hand, the data source is reliable, and on the other hand, it is not necessary to specially recruit health examination subjects for detection in order to establish a reference interval, which can improve the feasibility and reduce the cost to a certain extent.
[0041] In the process of acquiring a set of health examination test data from a laboratory information system through the test data acquisition module 10, first, database parameters are configured in the laboratory information system, after the test database is communicatively connected with the laboratory information system, at least part of the data of the configured database health examination subject medical record data is selected and exported to obtain preliminary exported data of the database health examination subject medical record data.
[0042] In the process of configuring the laboratory information system database, an existing database or a new database is selected to obtain database basic information and medical record data of a database health examinee; wherein the database basic information includes database name, database type (for example, Mysql, Oracle), database driver, database Internet Protocol Address (i.e., database IP (Internet Protocol Address) address), database port number, database username, and database password; and the medical record data of the database health examinee includes name, medical record number, sample number, gender, age, birth date, detection item name, detection time, detection instrument, and detection result.
[0043] In the process of selecting and exporting at least part of the medical record data of the configured database health examinee, the detection item name, year, or other selection conditions can be selected by writing a structured query language (SQL) to export at least part of the medical record data of the configured database health examinee.
[0044] The preliminary exported data of the database health examinee can be directly used as the set of health examinee test data, or the preliminary exported data of the database health examinee can be managed, for example, added, deleted, or modified, to obtain the managed exported data of the database health examinee.
[0045] Correspondingly, the set of health examinee test data includes test data of multiple health examinees, and the test data of each health examinee is the preliminary exported data of the medical record of the health examinee in the preliminary exported data of the medical record of the database health examinee, or the managed exported data of the medical record of the health examinee in the managed exported data of the medical record of the database health examinee.
[0046] In the process of deduplicating the set of health examinee test data based on age, gender, and name or medical record number by the deduplication module 20 to obtain a set of deduplicated health examinee test data, only the most recent detection data of each health examinee in the set of health examinee test data can be retained.
[0047] In particular, each health checkup examinee data includes at least one health checkup examinee data of at least one detection item. By means of a data chunking algorithm and a fingerprint generation and comparison algorithm, and by adopting a merge sort manner to arrange the detection time in descending order, and based on age, gender and name or medical record number, the multiple health checkup examinee data of the same health checkup examinee are deleted, and only the latest detection data of each health checkup examinee is retained, so as to limit the detection data of each health checkup examinee in the deduplicated health checkup examinee data set to a single health checkup examinee data of each health checkup examinee.
[0048] It should be understood that the deduplication module 20 can also deduplicate the health checkup examinee data set based on other data, for example, based on age, gender and name.
[0049] In the process of processing the deduplicated health checkup examinee data set based on the inclusion and exclusion criteria by the data screening module 30 to obtain a set of candidate health checkup examinee data, based on the pre-prepared inclusion and exclusion criteria, the abnormal health checkup examinee data affecting the reference interval of the detection item in the deduplicated health checkup examinee data set is excluded.
[0050] The abnormal health checkup examinee data includes health checkup examinee data of patients with diabetes, health checkup examinee data of patients with endocrine disorders, health checkup examinee data of patients with tumors, health checkup examinee data of patients with inflammation, health checkup examinee data of abnormal blood test results, health checkup examinee data of abnormal cardiovascular test results, health checkup examinee data of abnormal kidney test results, health checkup examinee data of abnormal liver test results, health checkup examinee data of pregnancy, health checkup examinee data of surgery within a preset time interval, and health checkup examinee data of other abnormal detection results of detection items affecting the reference interval.
[0051] In the process of grouping the set of candidate health checkup examinee data based on age and gender by the grouping module 40 to obtain a grouping result, the set of candidate health checkup examinee data can be grouped based on age and gender by means of the Harris & Boyd method or the standard difference ratio method.
[0052] In an example of the present application, the grouping is established in the following manner.
[0053] First, the set of candidate health checkup examinee data of health checkup examinees of different genders is sorted based on age respectively to obtain sorted candidate health checkup examinee data samples, wherein the sorted candidate health checkup examinee data samples are sequentially recorded as a1、 a 2、 a 3... a n-1 、 a n .
[0054] Then, the sorted candidate health examination data is divided into at least two data subsets, wherein the two data subsets are respectively denoted as a first data subset X 1 and a second data subset X 2, each data subset includes at least two sorted candidate health examination data samples, in an example of the present application, the initial first data subset X 1 is { a 1, a 2}, and the initial second data subset X 2 is { a 3, a 4}.
[0055] Next, it is determined whether the data subsets satisfy a preset grouping condition; when the data subsets satisfy the preset grouping condition, the data subsets are grouped according to a preset grouping iteration manner, and when the data subsets do not satisfy the preset grouping condition, grouping is configured based on the number of sorted candidate health examination data samples included in the at least two data subsets.
[0056] The preset grouping condition is:
[0057]
[0058]
[0059] Z
[0060] an average value of all sorted candidate health examination data samples of the first data subset; an average value of all sorted candidate health examination data samples of the second data subset; S 1 a standard deviation of all sorted candidate health examination data samples of the first data subset; S 2 a standard deviation of all sorted candidate health examination data samples of the second data subset; a variance of all sorted candidate health examination data samples of the first data subset; a variance of all sorted candidate health examination data samples of the second data subset; an absolute value of a difference between a standard deviation of all the ranked candidate health examination data samples of the first data subset and a standard deviation of all the ranked candidate health examination data samples of the second data subset; a maximum value between a standard deviation of all the ranked candidate health examination data samples of the first data subset and a standard deviation of all the ranked candidate health examination data samples of the second data subset; n 1 represents a number of all the ranked candidate health examination data samples of the first data subset; 1 represents a number of all the ranked candidate health examination data samples of the first data subset; k is a preset value, usually 3 to 5.
[0061] The preset grouping iteration mode is: selecting at least two ranked candidate health examination data samples after the ranked candidate health examination data samples in the second data subset X 2 to replace the ranked candidate health examination data samples in the first data subset 1, wherein the at least two ranked candidate health examination data samples used to replace the ranked candidate health examination data samples in the first data subset X 1 are defined as first replacement candidate health examination data samples; selecting at least two ranked candidate health examination data samples after the first replacement candidate health examination data samples to fill the ranked candidate health examination data samples in the second data subset X 2.
[0062] In the process of grouping configuration based on the number of the ranked candidate health examination data samples respectively contained in the at least two groups of data subsets, when the number of the ranked candidate health examination data samples in the first data subset X 1 is equal to the number of the ranked candidate health examination data samples in the second data subset X 2, at least one ranked candidate health examination data sample in the second data subset X 2 is filled into the first data subset X 1, at least one ranked candidate health examination data sample after the ranked candidate health examination data samples in the second data subset X 2 is supplemented to the second data subset X 2 to obtain an updated first data subset X 1 and an updated second data subset X2. For example, the second data subset X One of the sorted candidate healthy physical examination data samples described in 2 Fill into the first data subset X In step 1, the data in the second subset will be ranked. X 2. At least one sample of test data from the healthy individuals who underwent physical examinations. a 5. Add to the second data subset X 2. To obtain the updated first data subset X 1 and the updated second data subset X 2; where the updated first data subset X 1 is { a 1, a 2, a 3}; Updated second data subset X 2 is { a 4, a 5}. When the first subset of data is updated X 1 and the updated second data subset X 2. When the preset grouping conditions are met, the updated first data subset... X 1 and the updated second data subset X 2. Group the data according to the preset grouping iteration method. When the updated first subset of data... X 1 and the updated second data subset X 2. The preset grouping conditions are not met, and the updated second data subset X The number of candidate healthy examinee test data samples after sorting, as described in section 1, is less than the updated first data subset. X When the number of candidate healthy examinee test data samples after sorting as described in section 1 is equal, they will be ranked in the updated second data subset. X At least one of the sorted alternative healthy examinee test data samples following the healthy examinee test data sample in step 2 is added to the second data subset. X 2. To obtain the second data subset after further updates. X 2. For example, it will be ranked in the updated second subset of data. X At least one of the sorted alternative healthy examinee test data samples following the healthy examinee test data sample in 2. a 6. Fill into the second data subset X In step 2, the second data subset is obtained after further updates. X 2 is { a 4, a 5, a 6}. When the first data subset is updated X 1 and the second data subset after the updateX 2 When the preset grouping condition is met, updating the first data subset X 1 and the second data subset again X 2 Grouping according to a preset grouping iteration mode, when the updated first data subset X 1 and the second data subset again X 2 The preset grouping condition is not met, verifying the updated first data subset X 1 and the second data subset again X 2 The number of samples of the sorted candidate health examination data.
[0063] In the process of performing outlier processing and data distribution description on each data subset of candidate health examination data in the grouping result by the data distribution description module 50 (such as shown in Figure 3 Tukey algorithm is used to remove outliers of each data subset of candidate health examination data in the grouping result to obtain outlier-processed candidate health examination data, and the distribution description mode of the outlier-processed candidate health examination data is selected from at least one of the following modes: probability graph, quantile graph, kernel density graph, Shapiro-Wilk test, Kolmogorov-Smirnov test. For the subset data that does not meet the normality, Box-Cox conversion is used to convert the data into normal distribution or approximately normal distribution.
[0064] In the process of removing outliers of each data subset of candidate health examination data in the grouping result by using Tukey algorithm, the quartiles of each data subset of candidate health examination data in the grouping result are calculated, and the upper limit for identifying outliers is calculated by the formula P 25 -1.5*(P 75 -P 25 ) and the lower limit for identifying outliers is calculated by the formula P 75 +1.5*(P 75 -P 25 ) to remove data out of range. P 25 is the 25th percentile of the percentile. P 75 is the 75th percentile of the percentile.
[0065] It should be understood that the outliers of the data subsets of the individual candidate health examination test data in the grouping result can also be removed in other ways to obtain the outlier-processed candidate health examination test data, for example, the outliers of the data subsets of the individual candidate health examination test data in the grouping result are removed by a ±3SD method. In the process of removing the outliers of the data subsets of the individual candidate health examination test data in the grouping result by the ±3SD method, the mean and the standard deviation of the data subsets of the individual candidate health examination test data are calculated, the upper limit and the lower limit for identifying the outliers are determined, and the data out of the range is removed.
[0066] In another technical solution of the present application, the outliers of the data subsets of the individual candidate health examination test data in the grouping result are removed to obtain the outlier-processed candidate health examination test data, comprising the following steps:
[0067] S151: Vectorizing each candidate health examination test data in the data subset of the candidate health examination test data to obtain a set of candidate health examination test data embedding coding vectors. The purpose of vectorizing each candidate health examination test data in the data subset of the candidate health examination test data is to convert the non-numerical data into a numerical vector form that can be processed by a computer, so that mathematical operations can be used to analyze these data.
[0068] S152: Extracting the i th candidate health examination test data embedding coding vector as a query vector from the set of candidate health examination test data embedding coding vectors. Here, the i th selected refers to the data vector of any individual, which is used for comparison and analysis in subsequent steps.
[0069] S153: Inputting the query vector and the set of candidate health examination test data embedding coding vectors into a feature dynamic query analysis network to obtain a candidate test data set within query response coding vector;
[0070] S154: Inputting the candidate test data set within query response coding vector into a support vector machine model-based outlier identifier to obtain an identification result, which is used to indicate whether the i th candidate health examination test data embedding coding vector is an outlier. Finally, the response coding vector obtained in the previous step is input into a support vector machine (SVM) model-based outlier identifier. The support vector machine is a supervised learning model, which is used here to distinguish between normal values and outliers. Through the trained SVM model, it can be determined whether the i th candidate health examination test data embedding coding vector is an outlier. If the model judges it to be an outlier, the data point will be considered as an outlier and be removed.
[0071] In the technical solution of this application, step S153 includes: firstly, inputting the embedded encoding vectors of each candidate healthy examinee's test data from the set of query vectors and the set of embedded encoding vectors of candidate healthy examinee test data into an adaptive decision parser to obtain a set of implicit encoding vectors for the query-candidate test data decision point states. This process can be expressed by the formula:
[0072]
[0073]
[0074]
[0075] in, This represents the set of embedded encoding vectors of the test data of candidate healthy individuals. , , , They represent the first, the second, and the third, respectively. The and the first Embedded encoding vectors of test data from candidate healthy individuals. This represents the total number of embedded coding vectors of the test data of the candidate healthy examinees. This represents the query vector. Indicates the first The first weight matrix, Indicates the first The first bias vector. Represents matrix multiplication. Represents the sigmoid activation function. This represents the set of implicitly encoded vectors representing the state of decision points in the query-alternative test data. , , , They represent the first, the second, and the third, respectively. The and the first Each query-alternative test data decision point state implicit encoding vector.
[0076] It should be understood that the nonlinear decision response unit aims to capture the decision response result of the embedded encoding vector of the candidate healthy examinee's test data, given the query vector. Technically, the nonlinear decision response unit utilizes nonlinear mapping and inter-feature interactions to construct a latent space for the decision response state. Within this latent space, features that were previously difficult to distinguish can be effectively separated, and complex decision boundaries are simplified. The core of the nonlinear decision response unit lies in its use of nonlinear mapping and inter-feature interaction techniques. Nonlinear mapping allows the model to capture complex relationships in the data that are difficult to identify or express in linear models. For example, in health examination data, some potential health risk factors may not be significantly manifested in a single indicator, but require comprehensive consideration of multiple related indicators and their interactions for accurate identification. Inter-feature interactions further enhance this capability. It not only focuses on the performance of individual features but also analyzes the mutual influence between different features. In this way, the nonlinear decision response unit can create a new representation in the latent space, where each point (i.e., the latent encoding vector) reflects the state and characteristics of a candidate healthy examinee relative to the query vector. Within this latent space, features that were previously difficult to distinguish can be effectively separated. This means that individuals that appear similar in the original high-dimensional space but actually have subtle differences can be more clearly identified in the latent space.
[0077] Next, based on the set of implicit encoding vectors of the query-alternative test data decision point states, the neighborhood matrix of the query-alternative test data decision point states is calculated. This process can be expressed by the formula:
[0078]
[0079]
[0080] in, Indicates the first Each query-alternative test data decision point state implicit encoding vector The square of the 1-norm of a vector. Represents the inverse hyperbolic cosine function. This indicates the state class neighborhood matrix of the query-alternative test data decision points. The value of the position, This represents the neighborhood matrix of the state class of the query-alternative test data decision points. , , , These represent the state class neighborhood matrix of the query-alternative test data decision points. , , , value of the position.
[0081] It can be appreciated that the process of computing the class neighborhood matrix is to explicitly model the local structure information between the implicit encoding vectors of the query-alternative examination data decision point states. Specifically, in the context of health examination data, the class neighborhood matrix reflects the relative positions and connection relationships between different health examinees in the feature space. For example, if the implicit encoding vectors of the query-alternative examination data decision point states of two health examinees are very close, the corresponding elements in the class neighborhood matrix of the two health examinees will exhibit a high similarity value, indicating a strong local connection relationship between them. This connection relationship not only depends on numerical proximity, but also considers complex patterns resulting from nonlinear feature interactions. The core significance of constructing the class neighborhood matrix lies in its ability to capture the manifold structure within the data. Manifold structure refers to the inherent geometric shape that data may have in high-dimensional space, which often cannot be described by simple linear methods. By computing the class neighborhood matrix, this complex geometric structure can be converted into a graph form, where each node represents a health examinee, and the weight of the edge reflects the similarity or distance between the nodes. This method is particularly suitable for processing high-dimensional, complex, and possibly nonlinear data sets. In practical applications, the class neighborhood matrix provides a key foundation for subsequent spectral analysis.
[0082] Further, based on the set of query-alternative examination data decision point state implicit encoding vectors, a query-alternative examination data decision point state class degree matrix is computed. This process can be represented by the formula:
[0083]
[0084]
[0085] wherein, denotes the eigenvalue of the first position on the diagonal line in the query-alternative examination data decision point state class degree matrix, denotes the query-alternative examination data decision point state class degree matrix, , , denotes the eigenvalue of the first position on the diagonal line in the query-alternative examination data decision point state class degree matrix.
[0086] It should be appreciated that the process of computing the similarity degree matrix is to quantify the connection between the implicit encoding vectors of the query-candidate examination data decision point state and to evaluate the "importance" and "centrality" of each node (i.e., individual) in the entire network. The similarity degree matrix is a diagonal matrix that records the number of connections of each node, i.e., the connection strength between each individual and other individuals. In the application scenario of health examination data, this means that the relative position and influence of each individual in its group can be quantified. For example, if the health data of an individual is highly similar to that of most other individuals, its value in the similarity degree matrix will be higher, indicating that it is in a relatively "central" position in the network structure. Conversely, if the data of an individual is significantly different from that of other individuals, its degree will be lower, which may mean that it is a potential outlier or outlier. By constructing the similarity degree matrix, the internal structural information of the data can be more clearly seen. For example, in a healthy group, the physiological indicators of most individuals may exhibit certain similarity, forming a tightly connected core network. Those individuals who deviate from the normal range will exhibit a lower degree of connection, which helps to identify individuals who may have health problems. This method not only helps to discover outliers, but also reveals potential sub-group structures, i.e., small groups with similar health characteristics.
[0087] Then, based on the query-candidate examination data decision point state class neighborhood matrix and the query-candidate examination data decision point state similarity degree matrix, a query-candidate examination data decision point state Laplacian matrix is calculated. This process can be represented by the formula as follows:
[0088]
[0089] wherein, represents the query-candidate examination data decision point state Laplacian matrix.
[0090] And the query-candidate examination data decision point state Laplacian matrix is spectrally decomposed to obtain a set of query-candidate examination data decision point core component encoding vectors. This process can be represented by the formula as follows:
[0091]
[0092] wherein, represents spectral decomposition, represents the set of query-candidate examination data decision point core component encoding vectors, , , represents the first, second, and query-candidate examination data decision point core component encoding vector, represents the diagonal matrix after spectral decomposition, denotes a diagonal matrix, 、 、 denotes the first, second, and third eigenvalues on the diagonal of the spectral diagonal matrix.
[0093] It should be appreciated that the query-alternative test data decision point state class neighborhood matrix and the query-alternative test data decision point state class degree matrix are based on the query-alternative test data decision point state class neighborhood matrix and the query-alternative test data decision point state class degree matrix. The Laplacian matrix is a core tool in spectral graph theory, which transforms the structural information of a graph into an algebraic representation. The eigenvalues and eigenvectors of the Laplacian matrix contain important information about the distribution of data, such as frequency components and basis functions. Through further spectral decomposition, the original data can be re-examined from a low-dimensional perspective, the most representative features can be extracted, and the computational complexity can be reduced. In practical applications, the calculation of the Laplacian matrix and the subsequent spectral decomposition process are of great significance. For example, when identifying outliers or outliers, the Laplacian matrix can help determine which individuals are significantly different from the majority of other individuals. In a healthy population, some individuals may be far away from other individuals in the feature space due to underlying health problems or undetected risk factors. Through the Laplacian matrix and its spectral decomposition, these individuals can be more easily identified because their connection relationships show a clear deviation. In addition, the Laplacian matrix provides a basis for constructing low-dimensional embeddings. Through spectral decomposition, the eigenvectors corresponding to the smallest few eigenvalues can be extracted, which are considered to capture the low-dimensional embedding representation of the data manifold, representing the main structural direction of the data. This low-dimensional embedding not only preserves the key structural information of the data, but also significantly reduces the computational complexity, alleviates the curse of dimensionality, and improves the efficiency and robustness of query analysis.
[0094] Finally, the set of query-alternative test data decision point core component encoding vectors is adaptively fused to obtain the query response encoding vector within the alternative test data set. This process can be represented by the formula:
[0095]
[0096] wherein, denotes the i-th query-alternative test data decision point core component encoding vector, denotes the i-th modulation vector, denotes the i-th weight matrix, denotes the i-th bias vector, denotes a normalized exponential function, denotes the i-th query-alternative test data decision point core component encoding vector, denotes the i-th modulation vector, denotes the i-th weight matrix, denotes the i-th bias vector, denotes a normalized exponential function, a normalized query-alternative examination data decision point core component encoding factor, denotes a predetermined threshold value, denotes a mask function, denotes a first normalized decision fusion factor, denotes adaptive fusion, denotes a query response encoding vector within the alternative examination data set.
[0097] It can be understood that the query-alternative examination data decision point core component encoding vector is extracted from the spectral decomposition of the Laplacian matrix. These vectors represent the main structural directions of the data manifold, capturing the key features of the data. However, a single core component encoding vector may not fully reflect all important information, so it is necessary to fuse them in order to better utilize the unique contribution of each vector. The core of adaptive fusion lies in integrating the complementary information of these core component encoding vectors. For example, in health examination data, the physiological indicators of some individuals may be normal in one dimension, but show potential risks in another dimension. Through adaptive fusion, the information in these different dimensions can be considered comprehensively, so as to more comprehensively evaluate the health status of each individual. Specifically, the adaptive fusion method can be realized through attention mechanism or gating mechanism. Attention mechanism allows the model to learn the contribution or importance of each core component encoding vector to a specific query task, and dynamically allocates weights. This means that for each query vector, the system can focus on the most relevant feature components according to its specific needs, suppressing noise or irrelevant information. This flexibility enables the system to generate optimal query response encoding vectors for different query inputs and state changes. Through adaptive fusion, these individual embedding encoding vectors can be compared with the query vector, and the weights can be dynamically adjusted according to the performance of each individual in different feature dimensions. In this way, those individuals who show abnormalities or potential risks in certain key indicators will be given higher weights, so that the finally generated query response encoding vector is more discriminative and robust. In addition, adaptive fusion is not just a simple weighted sum, but a learning of context-dependent feature combination strategies. This method enables the dynamic query response encoding vector to best adapt to different query inputs and state changes.
[0098] In the process of determining the reference interval of each data subset of the alternative health examination data by the reference interval determination module 60 based on the data distribution description results of each data subset of the alternative health examination data, the reference interval of each data subset of the alternative health examination data is determined by using the following big data algorithms to obtain a plurality of reference intervals: EP28 non-parametric method, EP28 parametric method, refineR algorithm, TMC algorithm, Kosmic algorithm.
[0099] In the EP28 nonparametric method, it does not depend on the specific distribution of the data. n The test data of the selected healthy individuals are arranged in ascending order and recorded as follows: x 1. x 2. x 3… x n ; x 1≤ x 2≤ x 3≤…≤ x n ,in, x 1 is n The minimum value among the test data of the selected healthy examinees. x n for n The maximum value among the test data of the selected healthy examinees. n The test data of the selected healthy examinees were divided into 100 equal parts, and the number corresponding to the rank of r% is the rank. r The percentile is denoted by the symbol Pr. The rank of the lower and upper reference limits of the reference interval can be represented by Pr, respectively. 2.5 and P 97.5 express. r =0.025×( n +1), r =0.975×( n +1), if 0.025×( n +1) and 0.975×( n +1) is not an integer, so 0.025 × ( n +1) and 0.975×( n +1) Round to the nearest integer.
[0100] In the EP28 parameter method, if the data subsets of the test data of each candidate healthy examinee follow a normal distribution, or if the data subsets of the test data of each candidate healthy examinee follow a normal distribution after transformation, the mean ±1.96s of the data subsets of the test data of each candidate healthy examinee represents the 95% distribution range, or the mean ±2.58s of the data subsets of the test data of each candidate healthy examinee represents the 99% distribution range.
[0101] The refineR algorithm can be divided into three steps: data preprocessing, model optimization, and reference interval derivation. During the preprocessing of the subset of data from the candidate healthy individuals, the refineR algorithm identifies the dominant peak of the mixed distribution, derives the search region for the Box-Cox transform and normal distribution parameters (λ, μ, σ), and calculates a histogram within a given concentration range around the peak. After preprocessing the subset of data from the candidate healthy individuals, optimal values for λ, μ, and σ are determined through a multi-level grid search to find the optimal model for the subset of data from the candidate healthy individuals. Finally, the optimized model determines the non-pathological distribution, and the reference interval is derived based on the non-pathological distribution.
[0102] The TMC algorithm can be divided into 6 steps: (1) Select an age / gender layer for analysis; (2) Draw a stratified histogram; (3) Obtain the initial estimate of the power normal distribution (PND) parameter λ and derive μ and σ from a series of quantile-quantile (QQ) plots; (4) Obtain the improved estimate of the PND parameter through the TMC program for different cutoff interval candidate schemes; (5) Find the best PND parameter from the candidate parameters considered in step (4); (6) Calculate the reference interval based on the best PND parameter.
[0103] The Kosmic algorithm makes no assumptions about the distribution of pathological samples and minimizes the Kolmogorov-Smirnov distance between the estimated normal distribution and the truncated portion of the observed distribution of the test results after the Box-Cox transformation. The Kolmogorov-Smirnov test is a well-established normality test method, and it numerically optimizes the parameters (µ, σ) of the normal distribution, the Box-Cox transformation parameter (λ), and the truncation interval T.
[0104] like Figure 4 and Figure 5 As shown, in the process of constructing a continuous reference interval based on the multiple reference intervals by the continuous reference interval construction module 70, the multiple reference intervals are processed by the GAMLSS model to obtain the continuous reference interval, thereby achieving a natural, continuous and smooth transition between age groups.
[0105] Specifically, GAMLSS models can flexibly model both the location (mean) and scale (variance) parameters of the data distribution. More specifically, GAMLSS allows the use of smoothing functions to describe how these parameters vary with covariates (e.g., age), rather than simply assuming they are constant or linearly changing.
[0106] In traditional reference interval construction, it is common to establish independent reference intervals for each pre-defined age group (e.g. 0-1 year, 1-5 years, 5-10 years, etc.). The problem with this approach is that the boundaries between adjacent age groups can appear abrupt, lack continuity, and can not accurately reflect the true physiological parameter changes with age. In addition, this segmented approach to reference interval construction can also lead to diagnostic uncertainty at the critical age points.
[0107] The GAMLSS model can capture the non-linear changes in test results as age increases by introducing a smoothing function, such as a spline function, which means that for any given age, an estimated reference interval can be obtained based on all available data points and reflecting the complex relationship between age and test results. In this way, as age increases, the upper and lower limits of the reference interval gradually change, forming a smooth curve rather than a sudden jump.
[0108] In the following, through specific examples, the mechanism of GAMLSS model in constructing continuous reference interval is embodied.
[0109] In the process of using GAMLSS model to process the reference interval of platelet concentration of women in different age groups to obtain the continuous reference interval of platelet concentration of women, a statistical distribution family suitable for describing the distribution of platelet concentration is selected, wherein for platelet count, normal distribution may be suitable, but if the data shows skewness or other non-normal characteristics, other distributions such as Gamma distribution or Log-Normal distribution can also be considered; then, the gamlss package in R language is used for modeling; the trend of platelet concentration changing with age is captured by a smoothing function; other covariates can also be added to the model, such as whether in a specific physiological stage (represented by a binary variable), to further refine the model.
[0110] Through the GAMLSS model, the distribution parameters (such as mean μ and standard deviation σ) of platelet concentration at each age point can be obtained; based on these parameters, the 95% reference interval and 90% confidence interval can be calculated, wherein this reference interval is not fixed, but continuously changes with age, forming a natural, continuous and smooth transition; this means that for any given age, an accurate reference interval can be provided, reflecting the true distribution of platelet concentration of women in that age group.
[0111] In an embodiment of the present application, as Figure 6As shown, the gender and age based test item reference interval construction system 100 further comprises a consistency evaluation module 80 and a passing rate verification module 90. The consistency evaluation module 80 is configured to evaluate the consistency of the results of different construction algorithms of the reference interval of each data subset of the candidate health examination test data; and the passing rate verification module 90 is configured to verify the passing rate of the reference interval of each data subset of the candidate health examination test data.
[0112] Accordingly, the gender and age based test item reference interval construction system 100 can evaluate the consistency of the results of different construction algorithms of the reference interval of each data subset of the candidate health examination test data through the consistency evaluation module 80, and further verify the passing rate of the reference interval of each data subset of the candidate health examination test data through the passing rate verification module 90.
[0113] In the process of evaluating the consistency of the reference interval of each data subset of the candidate health examination test data through the consistency evaluation module 80, the bias degree of the reference interval established through other methods is calculated based on the reference interval established through the EP28 non-parametric method; when the bias degree of the reference interval established through other methods is within the preset range, it is determined that the consistency of the reference interval is good; and when the bias degree of the reference interval established through other methods is not within the preset range, it is determined that the consistency of the reference interval is poor.
[0114] In one specific example of the present application, the bias degree of the reference interval established through other methods is calculated based on the reference interval established through the non-parametric method. Specifically, the reference value is calculated through the following formula:
[0115]
[0116] wherein, SD RI represents the reference value; UL 0 represents the upper limit value of the reference interval established through the EP28 non-parametric method; LL 0 represents the lower limit value of the reference interval established through the EP28 non-parametric method. The bias degree of the lower limit value of the reference interval established through other methods is calculated through the following formula:
[0117]
[0118] wherein, BR LL represents the bias degree of the lower limit value of the reference interval established through other methods; LL represents the lower limit value of the reference interval established through other methods.
[0119] The bias degree of the upper limit value based on the reference interval established by other methods is calculated by the following formula:
[0120]
[0121] Wherein, BR UL The bias degree of the upper limit value based on the reference interval established by other methods; UL The upper limit value based on the reference interval established by other methods. When BR UL Or BR LL Greater than 0.375, it means that the bias degree of the reference interval established by other methods is large, and the consistency between different algorithms is poor; when BR UL And BR LL Less than or equal to 0.375, it means that the bias degree of the reference interval established by other methods is small, and the consistency between different algorithms is good.
[0122] In an embodiment of the present application, the test item reference interval data processing system based on gender and age further comprises the step of: S190, verifying the reference interval. When the reference interval meets the preset condition, the reference interval passes the verification; when the reference interval does not meet the preset condition, the reference interval fails the verification.
[0123] In the process of verifying the pass rate of the reference interval of each data subset of the candidate health examination test data by the pass rate verification module 90, the test data acquisition module is used to obtain the corresponding detection item physical examination data of another time period as verification data, and the grouping verification is carried out according to the grouping of the reference interval. The pass rate of the verification data (i.e. the probability of being consistent with the continuous reference interval) is greater than or equal to 90%, which means that the verification is passed, and the verification result is displayed through a visual picture.
[0124] In summary, the gender and age based test item reference interval data processing system according to the embodiments of the present application is illustrated. The gender and age based test item reference interval data processing system groups the test data according to gender and age, and establishes the reference interval corresponding to each age stage of different genders, which can improve the accuracy and suitability of the reference interval to a certain extent. The collection of the test data of the healthy examinees in the gender and age based test item reference interval data processing system comes from the real world patient medical record data, and the source and use of the medical record data are legal and compliant. In this way, on the one hand, the data source is reliable, and on the other hand, it is not necessary to specially recruit healthy examinees for detection in order to construct the reference interval experiment, which can improve the feasibility and reduce the cost to a certain extent.
[0125] The above describes the present application and its embodiments, which is not restrictive, and the drawings only show one of the embodiments of the present application, and the actual structure is not limited thereto. In summary, if a person skilled in the art is inspired by it, without departing from the purpose of the present application, similar structure and embodiments can be designed without creativity, which should belong to the protection scope of the present application.
Claims
1. A gender and age based test item reference interval construction system, characterized by, The method comprises the following steps: acquiring a set of health examination data of a health examinee from a laboratory information system; performing deduplication on the set of health examination data of the health examinee based on age, gender, and name or medical record number to obtain a set of deduplicated health examination data of the health examinee; processing the set of deduplicated health examination data of the health examinee based on inclusion and exclusion criteria to obtain a set of candidate health examination data of the health examinee; grouping the set of candidate health examination data of the health examinee based on age and gender to obtain a grouping result, wherein the grouping result comprises a plurality of data subsets of candidate health examination data of the health examinee; performing outlier processing and data distribution description on each data subset of candidate health examination data in the grouping result to obtain a plurality of data distribution description results; determining reference intervals for each data subset of candidate health examination data based on the data distribution description results of the data subsets by using a plurality of big data algorithms; constructing continuous reference intervals based on gender, age, and the plurality of reference intervals; wherein the outlier processing on each data subset of candidate health examination data in the grouping result comprises the following steps: vectorizing each candidate health examination data in the data subset of candidate health examination data to obtain a set of candidate health examination data embedding coding vectors; extracting an i-th candidate health examination data embedding coding vector from the set of candidate health examination data embedding coding vectors as a query vector; inputting the query vector and the set of candidate health examination data embedding coding vectors into a feature dynamic query analysis network to obtain an intra-candidate examination data set query response coding vector; inputting the intra-candidate examination data set query response coding vector into an outlier identifier based on a support vector machine model to obtain an identification result, wherein the identification result is used to indicate whether the i-th candidate health examination data embedding coding vector is an outlier. The query vector and the candidate health examination data embedding code vector set are input into a dynamic query analysis network to obtain a query response code vector in the candidate examination data set, including: inputting the query vector and each candidate health examination data embedding code vector in the candidate health examination data embedding code vector set into an adaptive decision analyzer to obtain a query-candidate examination data decision point state hidden code vector set; based on the query-candidate examination data decision point state hidden code vector set, a query-candidate examination data decision point state class neighborhood matrix is calculated; based on the query-candidate examination data decision point state hidden code vector set, a query-candidate examination data decision point state class degree matrix is calculated; based on the query-candidate examination data decision point state class neighborhood matrix and the query-candidate examination data decision point state class degree matrix, a query-candidate examination data decision point state Laplacian matrix is calculated; the query-candidate examination data decision point state Laplacian matrix is spectrally decomposed to obtain a query-candidate examination data decision point core component code vector set; the query-candidate examination data decision point core component code vector set is adaptively fused to obtain the query response code vector in the candidate examination data set.
2. The gender and age based test item reference interval construction system according to claim 1, characterized by, The continuous reference interval construction module is further configured to process the plurality of reference intervals using a GAMLSS model to obtain the continuous reference interval.
3. The gender and age based test item reference interval construction system according to claim 2, wherein, Further comprising: a consistency evaluation module configured to evaluate the consistency of results of different construction algorithms of the reference interval of each data subset of the candidate health examination data; and a pass rate verification module configured to verify the pass rate of the reference interval of each data subset of the candidate health examination data.
4. The gender and age based test item reference interval construction system according to claim 1, wherein, The deduplication module is further configured to retain only the most recent examination data of each health examination person in the set of health examination data.
5. The gender and age based test item reference interval construction system according to claim 1, wherein, The data screening module is further configured to exclude abnormal health examination data of the health examination person affecting the reference interval of the examination item in the set of health examination data after deduplication based on pre-prepared inclusion and exclusion criteria.
6. The gender and age based test item reference interval construction system according to claim 1, wherein, The grouping module is further configured to group the set of candidate health examination data based on age and gender by the Harris & Boyd method or the standard deviation ratio method.
7. The gender and age based test item reference interval construction system according to claim 1, wherein, The data distribution description module is further configured to perform distribution description on the candidate health examination data to obtain a plurality of data distribution description results, wherein the manner of performing distribution description on the candidate health examination data is selected from at least one of the following manners: a probability graph, a quantile graph, a kernel density graph, a Shapiro-Wilk test, a Kolmogorov-Smirnov test; and removing outliers of each data subset of the candidate health examination data in the grouping result to obtain outlier-processed candidate health examination data.
8. The gender and age based test item reference interval construction system according to claim 1, wherein, The reference intervals of the data subsets of the test data of each of the candidate health examinees are determined by using the following big data algorithms to obtain a plurality of reference intervals: EP28 non-parametric method, EP28 parametric method, refineR algorithm, Kosmic algorithm, and TMC algorithm.
9. The gender and age based test item reference interval construction system according to claim 3, wherein, The consistency evaluation module is further configured to calculate bias ratios of the reference intervals established by the rest of the algorithms based on the reference intervals established by the EP28 non-parametric method.
10. The gender and age based test item reference interval construction system according to claim 3, characterized by, The pass rate verification module is further configured to verify the reference intervals by obtaining health examination data of test items in another time period from the laboratory information system as verification data; when the reference interval pass rate is greater than or equal to 90%, the reference intervals pass the verification; otherwise, the reference intervals do not pass the verification.
Citation Information
Patent Citations
Method and equipment for calculating reference interval of biomarker
CN115620819A
Fund management and control method based on grassroots financial management
CN118657618A