ANCA-related vasculitis prognosis auxiliary prediction method and system
By collecting and classifying electronic medical record information for patients with ANCA-related vasculitis (AAV), the replacement treatment parameters and mortality prediction parameters are processed, and the problem of complex and inaccurate AAV prognosis judgment in the prior art is solved, and more accurate prognosis analysis and patient evaluation are achieved.
Patent Information
- Application Number
- CN202411894126.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prognosis judgment process for ANCA-related vasculitis (AAV) in the prior art is complex and inaccurate, making it difficult to accurately analyze the impact of patient clinical data on prognosis, and the model information is not complete enough, so the predictors of adverse outcomes in AAV population have not been fully studied.
A prognostic auxiliary prediction method for ANCA-related vasculitis is provided, by collecting electronic medical record information from prognostic patients, extracting clinical data fields, and classifying them to obtain the first type field associated with alternative treatment and the second type field associated with mortality. Then, alternative treatment parameters and mortality prediction parameters are obtained based on these fields.
Through this method, the prognostic outcomes of AAV patients, including tumor, renal replacement therapy and death, can be more accurately predicted, and relevant reference scores are provided to doctors to help improve the condition evaluation and treatment of AAV patients.
Smart Images

Figure CN120015302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular to a prognosis auxiliary prediction method and system for ANCA-associated vasculitis. Background Art
[0002] ANCA-associated vasculitis (AAV) is an autoimmune disease characterized by small vessel necrotizing inflammation, which causes irreversible organ damage clinically. The kidney is one of the most common sites of involvement, leading to increased mortality. Untreated patients with kidney involvement can rapidly develop end-stage renal disease and require renal replacement therapy, including dialysis or kidney transplantation in some cases. Therefore, the long-term prognosis of AAV patients is not ideal. Although the diagnosis and treatment levels have been continuously improved, making AAV a chronic recurrent disease, a considerable proportion of patients still have acute onset and poor prognosis. A meta-analysis of AAV prognosis confirmed that the risk of death in AAV is at least 2.7 times higher than that in the general population. Therefore, early and correct identification of risk factors is of great significance for assessing their condition and improving AAV.
[0003] In the existing technology, the prediction process for ANCA-associated vasculitis is relatively complicated. In some studies, multivariate logistic regression models were used for a large number of cases to confirm that cardiovascular disease, malignant tumors and renal death may be risk factors for premature death caused by AAV. At the same time, other logistic regression studies have shown that renal function, disease activity and age are important predictors of poor prognosis of AAV.
[0004] However, the above judgment process is mainly based on the multivariate logistic regression model. First, the model is sensitive to multicollinear data and outliers, and the accuracy is not very high; second, the simple form makes it difficult to fit the true distribution of the data; and there is a lack of research on survival, making it difficult to accurately analyze the impact of patient clinical data on prognosis. At present, the prediction model for the outcomes of death, tumors, and renal replacement therapy in AAV patients has problems such as different inclusion criteria, fewer types of prognostic outcomes, and too short follow-up time, which makes the model information incomplete, resulting in the predictive factors for adverse outcomes in the AAV population (including death, malignant tumors, and renal replacement therapy) have not been fully studied. Summary of the invention
[0005] In view of the above problems existing in the prior art, a method for auxiliary prediction of the prognosis of ANCA-associated vasculitis is now provided;
[0006] On the other hand, a prognosis auxiliary prediction system for implementing the prognosis auxiliary prediction method is also provided.
[0007] The specific technical solutions are as follows:
[0008] A method for assisting in predicting the prognosis of ANCA-associated vasculitis, comprising:
[0009] Step S1: collecting electronic medical record information of the prognostic patient and extracting clinical data fields from the medical record information;
[0010] Step S2: classifying the clinical data fields to obtain a first type of fields associated with alternative treatments and a second type of fields associated with mortality rates;
[0011] Step S3: Obtaining alternative treatment parameters according to the first type of fields, and obtaining mortality prediction parameters according to the second type of fields.
[0012] On the other hand, before executing the step S1, a modeling process is also included, and the modeling process includes:
[0013] Step A1: Collect original clinical data of enrolled patients;
[0014] Step A2: searching for a prognosis outcome field from the original clinical data, and grouping the enrolled patients according to the prognosis outcome field to form prognosis groups;
[0015] The prognostic grouping includes a high-risk group, which includes tumors, renal replacement therapy and death;
[0016] Step A3: analyzing the data feature differences of the data features corresponding to each field in the original clinical data for the high-risk group;
[0017] Step A4: construct hypothesis test for the data feature differences;
[0018] Step A5: Based on the test hypothesis, each field in the original clinical data is analyzed for tumor outcome to predict a first field group associated with the tumor outcome, a second field group associated with the renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis screening;
[0019] Step A6: Generate the first type field and the second type field according to the first field group, the second field group and the third field group.
[0020] In another aspect, the clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels;
[0021] The step S1 comprises:
[0022] Step S11: collecting the electronic medical record information for the prognosis patient;
[0023] Step S12: for the electronic medical record information, the clinical data fields are used to search in sequence to obtain medical record fields;
[0024] Step S13: Process the medical record field to obtain the clinical data field.
[0025] On the other hand, the first type of fields includes: heart involvement at first diagnosis, kidney organ involvement at first diagnosis, BVAS score, serum eGFR, and total complement level;
[0026] The second type of fields include serum eGFR and age;
[0027] The step S2 comprises:
[0028] Step S21: searching for an organ involvement condition field from the clinical data field, and screening out the heart involved at the time of the first diagnosis and the kidney involved at the time of the first diagnosis from the organ involvement conditions;
[0029] Step S22: searching the BVAS score, the serum eGFR, the age and the total complement level from the clinical data field;
[0030] Step S23: classify the heart involvement at the time of first diagnosis, the kidney involvement at the time of first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type field, and classify the serum eGFR and the age into the second type field.
[0031] On the other hand, the step S3 comprises:
[0032] Step S31: searching for a first rating scale according to the first type of field, and searching for a second rating scale according to the second type of field;
[0033] Step S32: Processing the first type of fields according to the first scoring scale to obtain the alternative treatment parameters, and processing the second type of fields according to the second scoring scale to obtain the death parameters.
[0034] A prognosis auxiliary prediction system for ANCA-associated vasculitis, used to implement the above-mentioned prognosis auxiliary prediction method;
[0035] include:
[0036] A data extraction module, wherein the data extraction module collects electronic medical record information of the prognosis patient and extracts clinical data fields from the medical record information;
[0037] a data classification module, the data classification module being connected to the data extraction module, the data classification module classifying the clinical data fields to obtain a first type of field associated with alternative treatments and a second type of field associated with mortality;
[0038] A parameter calculation module is connected to the data classification module, and the parameter calculation module obtains alternative treatment parameters according to the first type of fields and obtains mortality prediction parameters according to the second type of fields.
[0039] On the other hand, it also includes modeling modules;
[0040] The modeling module includes:
[0041] A sample collection module, which collects original clinical data from enrolled patients;
[0042] A preprocessing module, the preprocessing module is connected to the sample collection module, the preprocessing module searches for a prognosis outcome field from the original clinical data, and groups the enrolled patients according to the prognosis outcome field to form prognosis groups;
[0043] The prognostic grouping includes a high-risk group, which includes tumors, renal replacement therapy and death;
[0044] A feature extraction module, the feature extraction module is connected to the preprocessing module, and the feature extraction module analyzes the data feature differences corresponding to each field in the original clinical data for the high-risk group;
[0045] A hypothesis building module, the hypothesis building module is connected to the feature extraction module, and the hypothesis framework module constructs a hypothesis test according to the data feature difference;
[0046] A sample analysis module, the sample analysis module is connected to the hypothesis building module, and the sample analysis module analyzes each field in the original clinical data for tumor outcomes based on the test hypothesis to predict a first field group associated with the tumor outcomes, a second field group associated with the renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis screening;
[0047] A sample output module, wherein the sample output module is connected to the sample analysis module, and wherein the sample output module generates the first type field and the second type field according to the first field group, the second field group and the third field group.
[0048] In another aspect, the clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels;
[0049] The data extraction module comprises:
[0050] A medical record collection module, wherein the medical record collection module collects the electronic medical record information for the prognosis patient;
[0051] A field search module, the field search module is connected to the medical record collection module, and the field search module uses the clinical data fields to search for the electronic medical record information in sequence to obtain the medical record fields;
[0052] A field processing module, the field processing module is connected to the field retrieval module, and the field processing module processes the medical record field to obtain the clinical data field.
[0053] On the other hand, the first type of fields includes: heart involvement at first diagnosis, kidney organ involvement at first diagnosis, BVAS score, serum eGFR, and total complement level;
[0054] The second type of fields include serum eGFR and age;
[0055] The data classification module comprises:
[0056] A first search module, wherein the first search module searches for an organ involvement condition field from the clinical data field, and selects the heart involved at the time of the first diagnosis and the kidney involved at the time of the first diagnosis from the organ involvement conditions;
[0057] a second search module, the second search module being connected to the first search module, and the second search module searching the BVAS score, the serum eGFR, the age and the total complement level from the clinical data field;
[0058] A field division module, wherein the field division module is connected to the second search module, and the field division module divides the heart involvement at the first diagnosis, the kidney organ involvement at the first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type of field, and divides the serum eGFR and the age into the second type of field.
[0059] On the other hand, the parameter calculation module includes:
[0060] A scale search module, wherein the scale search module searches for a first rating scale according to the first type of field, and searches for a second rating scale according to the second type of field;
[0061] A parameter calculation module, wherein the parameter calculation module is connected to the scale lookup module, wherein the parameter calculation module processes the first type of field according to the first scoring scale to obtain the alternative treatment parameter, and processes the second type of field according to the second scoring scale to obtain the death parameter.
[0062] The above technical solution has the following advantages or beneficial effects:
[0063] In view of the relatively complex and inaccurate prognosis judgment process of AAV in the existing technology, this solution conducts correlation analysis on different prognostic outcomes (tumor, renal replacement therapy, death) and survival time as dependent variables, thereby determining the relevant indicators associated with the two typical prognostic outcomes of replacement therapy and death. In the actual prognostic analysis process, these indicators can be automatically extracted and processed to provide doctors with relevant reference scores. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The embodiments of the present invention will be described more fully with reference to the attached drawings, which are provided for illustration and description only and are not intended to limit the scope of the present invention.
[0065] Figure 1 It is an overall schematic diagram of an embodiment of the present invention;
[0066] Figure 2 A schematic diagram of a modeling process in an embodiment of the present invention;
[0067] Figure 3 is the P value of each factor of renal survival rate of AAV patients in the embodiment of the present invention;
[0068] Figure 4 is the P value of each factor of the overall survival rate in the embodiment of the present invention;
[0069] Figure 5 This is a schematic diagram of step S1 in an embodiment of the present invention;
[0070] Figure 6 This is a schematic diagram of step S2 in an embodiment of the present invention;
[0071] Figure 7 This is a schematic diagram of step S3 in an embodiment of the present invention;
[0072] Figure 8 A schematic diagram of a system in an embodiment of the present invention;
[0073] Fig. 9 This is a schematic diagram of a modeling module in an embodiment of the present invention;
[0074] Fig.10 Schematic diagram of a data extraction module in an embodiment of the present invention;
[0075] Fig.11 This is a schematic diagram of a data classification module in an embodiment of the present invention;
[0076] Fig.12 Schematic diagram of a parameter calculation module in an embodiment of the present invention. DETAILED DESCRIPTION
[0077] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0078] The terms used in the embodiments of this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. Unless otherwise defined, the technical terms or scientific terms used in the embodiments of this specification should be understood by people with ordinary skills in the field to which this specification belongs. The words "first", "second" and similar words used in this specification and claims do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one" or "one" do not indicate a quantitative limit, but indicate that there is at least one. "Multiple" or "several" means two or more. Unless otherwise specified, words such as "front", "rear", "lower" and / or "upper" are only for the convenience of explanation and are not limited to one position or one spatial orientation. Words such as "include" or "comprise" mean that the elements or objects appearing in front of "include" or "comprise" include the elements or objects listed after "include" or "comprise" and their equivalents, and do not exclude other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, and can include electrical connections, whether direct or indirect. As used in this specification and the appended claims, the singular forms "a", "an", "said", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0079] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0080] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0081] The present invention comprises:
[0082] A method for assisting the prediction of prognosis of ANCA-associated vasculitis, such as Figure 1 As shown, including:
[0083] Step S1: Collecting electronic medical record information of the prognostic patient and extracting clinical data fields from the medical record information;
[0084] Step S2: classifying clinical data fields to obtain a first type of field associated with alternative treatments and a second type of field associated with mortality;
[0085] Step S3: Obtaining alternative treatment parameters according to the first type of field processing, and obtaining mortality prediction parameters according to the second type of field processing.
[0086] Specifically, in view of the relatively complex and inaccurate judgment process of AAV prognosis in the prior art, this solution conducts correlation analysis on different prognostic outcomes (tumor, renal replacement therapy, death) and survival time as dependent variables, thereby determining the relevant indicators associated with the two typical prognostic outcomes of replacement therapy and death. In the actual prognostic analysis process, these indicators can be automatically extracted and processed to provide doctors with relevant reference scores.
[0087] During implementation, the above-mentioned prognosis auxiliary prediction method is mainly configured as a software implementation in a specific computer system, such as a doctor's workstation, for connecting to an external electronic medical record system for data collection and responding to relevant processing requests to feedback relevant data.
[0088] Among them, the electronic medical record information is the electronic medical record data corresponding to the current prognosis of the patient collected from the electronic medical record system. The electronic medical record system is a database system configured in the hospital, which stores the electronic medical record information of a large number of patients and circulates through the relevant data bus.
[0089] Prognostic patients refer to patients who currently need to have their prognosis predicted. The patients should have undergone at least the first diagnosis process and have collected relevant indicators and stored them in the electronic medical record system. When the doctor inputs relevant instructions into the computer, the patient corresponding to the instruction is marked as a prognostic patient for data processing.
[0090] Since the electronic medical record system stores the patient's relevant clinical data based on the database, the corresponding clinical data fields will be obtained during the data collection process. Here, the main purpose is to extract the relevant fields related to the subsequent prediction process.
[0091] The first type of field and the second type of field are two types of field data that are marked based on a classification process, and are used to predict different prognostic outcomes, including alternative treatments and death. According to the analysis results, there may be overlapping fields in the two types of field data.
[0092] The alternative treatment parameters and mortality prediction parameters are evaluation results based on the first type of fields and the second type of fields, respectively, obtained after processing by means of scales, weighted processing, etc. They are used to indicate the risk level of related outcomes, which is convenient for doctors to refer to during the diagnosis process.
[0093] In one embodiment, Figure 2 As shown, before executing step S1, a modeling process is also included, and the modeling process includes:
[0094] Step A1: Collect original clinical data of enrolled patients;
[0095] Step A2: Find the prognostic outcome field from the original clinical data, and group the enrolled patients according to the prognostic outcome field to form prognostic groups;
[0096] Prognostic groups included a high-risk group, which included neoplasms, renal replacement therapy, and death;
[0097] Step A3: analyzing the data feature differences corresponding to each field in the original clinical data for the high-risk group;
[0098] Step A4: Construct hypothesis tests based on data feature differences;
[0099] Step A5: Based on the test hypothesis, each field in the original clinical data is analyzed for tumor outcomes to predict a first field group associated with tumor outcomes, a second field group associated with renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis;
[0100] Step A6: Generate first type fields and second type fields according to the first field group, the second field group and the third field group.
[0101] Specifically, in order to realize the collection of the above fields, in this solution, the above method is used to realize the establishment of the COX regression model for the patient, so as to determine the relevant fields that are helpful in predicting the two prognostic outcomes of alternative treatment and death.
[0102] Specifically, the following clinical data were collected at the time of first diagnosis and treatment of AAV patients: age, sex, smoking and drinking history, general condition (blood pressure, weight), organ involvement (skin, mucosal, ENT, cardiovascular, gastrointestinal, lung, kidney, nervous system), laboratory data including blood routine and urine routine, liver and kidney function, electrolyte levels, inflammatory parameters, immunoglobulin levels, complement, ANCA serology and renal biopsy pathology results. All symptoms and comorbidity diagnoses met the Birmingham Vasculitis Activity Score (BVAS) criteria. The glomerular filtration rate (eGFR) was estimated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equation. BVAS was calculated by clinicians at the time of first diagnosis and was considered active when the score was ≥15 points. Treatment regimens and medication doses were also collected, including glucocorticoid pulse therapy and the cumulative dose of cyclophosphamide at the last follow-up. These data together formed the clinical baseline data.
[0103] Accordingly, in the process of retrospective analysis, different prognostic outcomes can be obtained because historical medical records are used, and the prognostic outcome fields can be found in the original clinical data during the analysis process.
[0104] Then, the patients were grouped according to different prognostic outcome definitions to form prognostic groups. AAV patients who developed tumors, renal replacement therapy, and death were defined as the high-risk group. Renal replacement therapy (RRT) was defined as the improvement of renal disease progression through peritoneal dialysis, hemodialysis, renal transplantation, or combined plasma exchange therapy. Tumor outcomes were defined as the occurrence of hematological tumors and solid tumors after the diagnosis of AAV. The time from the first visit to the occurrence of different prognostic outcomes was also recorded.
[0105] On the basis of grouping, one-way analysis of variance, nonparametric test, and chi-square test for multiple sample comparison were used to analyze the differences in clinical characteristics of AAV patients with different high-risk outcomes to form data feature differences.
[0106] Then, hypothesis tests were constructed for the differences in data characteristics, that is, whether the differences in various data characteristics would have an impact on the relevant disease process.
[0107] Then, different analysis methods were used to analyze each prognostic outcome.
[0108] For example, because the clinical characteristics of the tumor groups were less different, the Logistic regression model was used to predict the risk factors for tumor outcomes in AAV patients, the odds ratio (OR) was calculated to evaluate the risk of each variable in each group, and the ROC curve was used to verify the clinical prediction model. The R language software was used to calculate the hypothesis test related to the Cox regression model, and the multivariate Cox regression analysis after the forward LR method was used to screen the variables to predict the risk factors for renal replacement therapy in AAV patients.
[0109] Kaplan-Meier (KM) survival curves were used to depict the renal survival rate and overall survival rate of patients, and the log-rank test was used to evaluate the survival difference.
[0110] Independent influencing factors of AAV mortality outcome were screened based on univariate and multivariate Cox regression analysis.
[0111] Based on the above process, the first field group, the second field group and the third field group with relatively significant p values can be extracted respectively. Finally, the first type field and the second type field are generated according to the first field group, the second field group and the third field group.
[0112] In one embodiment, the above analysis method is used to process a group of patient data.
[0113] In this example, 89 AAV patients were enrolled. During the median follow-up of 267 months, a total of 46 patients (51.7%) had high-risk events, including 20 patients who received RRT, no patients who received renal transplantation, 12 patients who developed tumors, and 29 patients who died. 21 patients received hormone pulse therapy at the time of first diagnosis, and 56 patients received a cumulative dose of cyclophosphamide of 0.4-16g.
[0114] The median cumulative dose of cyclophosphamide in patients at high risk of events was 0.8 g (IQR 0-9).
[0115] Results 1. Binary logistic regression analysis showed that eGFR and serum complement 3 (C3) levels were independently associated with high-risk outcomes (cancer, renal replacement therapy, and death) in AAV, and higher eGFR and C3 levels reduced the probability of high-risk outcomes.
[0116] Result 2: Subgroup analysis was performed on different prognostic outcomes, and the ROC curve was used to verify the clinical prediction model. It was found that serum potassium had a moderate predictive effect on tumor outcomes in AAV patients.
[0117] Results 3. When the KM curve of renal outcomes was drawn, it was found that the higher the BVAS of AAV patients at the initial diagnosis, the more likely they were to develop end-stage renal disease.
[0118] Patients with AAV that involved the heart and kidneys at first diagnosis were more likely to receive renal replacement therapy in the later stages of the disease. However, there was no significant difference in renal survival between patients with different ANCA types.
[0119] Results 4. Multivariate Cox regression analysis after screening variables using the forward LR method showed that with every 1 unit increase in BVAS, the probability of patients requiring renal replacement therapy increased by 29%; with every 1 unit increase in serum eGFR and total complement levels, the probability of renal replacement therapy decreased by 21.8% and 6.6%, respectively.
[0120] Results 5. The cumulative survival rates of AAV patients at 1, 3, and 5 years were 86.0%, 76.2%, and 68.6%, respectively. During the follow-up period, infection and organ failure caused by AAV remained the main causes of death. Compared with the survival group, AAV patients in the death group were older at the time of initial diagnosis, more likely to have clinical manifestations such as heart and kidney involvement, lower blood calcium levels, and more severe renal impairment.
[0121] Results 6. The KM curve of overall survival of patients showed that the mortality rate of patients aged ≥ 65 years was higher than that of younger patients; in addition, the cumulative survival time of patients with early coagulation abnormalities, heart and kidney organ involvement was also lower than that of normal patients; but whether or not to receive renal replacement therapy had no significant difference in the survival rate of patients.
[0122] Results VII. The results of univariate Cox model analysis indicated that age ≥ 65 years, cardiac involvement, higher fibrinogen, lower blood calcium, and poorer renal function at the time of initial diagnosis increased the probability of death. The multivariate Cox analysis model indicated that only age at diagnosis and eGFR could predict death alone. In conclusion, BVAS and eGFR can be used as important predictors of renal replacement therapy in AAV patients; age and eGFR can independently predict death outcomes. The predictive value of serum potassium level in the development of cancer in AAV patients should be taken seriously.
[0123] Figure 3 Comparison of renal survival in AAV patients according to BVAS score (a), MPO (b), PR3 (c), and complement 3 levels (d) at diagnosis; p values obtained using log-rank analysis.
[0124] Figure 4 Figure 3 Overall survival of AAV patients according to age at diagnosis (a), coagulation abnormality (b), cardiovascular involvement (c), and renal involvement (d); p values obtained using log-rank analysis.
[0125] Table 1 Univariate and multivariate Cox risk model analysis of all-cause mortality variables in AAV patients
[0126]
[0127] Abbreviations: AAV: ANCA-associated vasculitis.
[0128] In one embodiment, clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels;
[0129] like Figure 5 As shown, step S1 includes:
[0130] Step S11: Collecting electronic medical record information for the prognosis patient;
[0131] Step S12: for the electronic medical record information, clinical data fields are used to search in sequence to obtain medical record fields;
[0132] Step S13: Process the medical record fields to obtain clinical data fields.
[0133] Specifically, in order to achieve effective extraction of clinical data fields, in this embodiment, electronic medical record information is first collected for prognosis patients. This process involves a full range of fields. For the electronic medical record information, the clinical data fields are used to perform search operations in turn to obtain the corresponding medical record fields. Finally, the medical record fields are processed, such as assigning null values, formatting the data, etc., to obtain the clinical data fields.
[0134] In one embodiment, the first type of fields include: heart involvement at first diagnosis, kidney organ involvement at first diagnosis, BVAS score, serum eGFR, and total complement level;
[0135] The second type of fields includes serum eGFR, age;
[0136] like Figure 6 As shown, step S2 includes:
[0137] Step S21: searching the organ involvement condition field from the clinical data field, and screening the organ involvement conditions to obtain the heart involved at the first diagnosis and the kidney involved at the first diagnosis;
[0138] Step S22: searching the BVAS score, serum eGFR, age and total complement level from the clinical data fields;
[0139] Step S23: The heart involvement at the time of first diagnosis, the kidney involvement at the time of first diagnosis, the BVAS score, the serum eGFR and the total complement level are classified into the first type field, and the serum eGFR and age are classified into the second type field.
[0140] Specifically, in order to achieve effective division of the two types of fields, in this embodiment, the organ involvement field is first searched from the clinical data field. Since the field records different involved organs in the form of a string, the heart and kidney are matched by regular search to determine whether there is a matching result. If so, the corresponding field is marked as 1, and if not, it is marked as 0, forming the heart involved at the first diagnosis and the kidney involved at the first diagnosis. Then, the BVAS score, serum eGFR, age and total complement level are searched from the clinical data field. The BVAS score is usually calculated by the doctor during the first visit. If not calculated, the associated field is extracted for the empty value, and the score is calculated and filled in.
[0141] Finally, according to the above analysis results, heart involvement at first diagnosis, kidney involvement at first diagnosis, BVAS score, serum eGFR and total complement level were divided into the first type field, and serum eGFR and age were divided into the second type field.
[0142] In one embodiment, Figure 7 As shown, step S3 includes:
[0143] Step S31: searching for a first rating scale according to the first type of field, and searching for a second rating scale according to the second type of field;
[0144] Step S32: Processing the first type of fields according to the first scoring scale to obtain replacement treatment parameters, and processing the second type of fields according to the second scoring scale to obtain death parameters.
[0145] Specifically, in order to achieve effective processing of the above parameters, in this embodiment, after classifying the fields, first, the first rating scale is obtained according to the first type of field, and the second rating scale is obtained according to the second type of field. Then, for each field, the corresponding items are found in the scale, and the score mapping is performed based on the score interval in the corresponding item to obtain the relevant score, and the score is weighted to obtain the total score, and finally the relevant conclusion is matched and output.
[0146] A prognosis auxiliary prediction system for ANCA-associated vasculitis, used to implement the above-mentioned prognosis auxiliary prediction method;
[0147] like Figure 8 As shown, including:
[0148] Data extraction module 1, data extraction module 1 collects electronic medical record information of the prognosis patient and extracts clinical data fields from the medical record information;
[0149] A data classification module 2, the data classification module 2 is connected to the data extraction module 1, and the data classification module 2 classifies the clinical data fields to obtain a first type of field associated with alternative treatments and a second type of field associated with mortality;
[0150] Parameter calculation module 3, parameter calculation module 3 is connected to data classification module 2, parameter calculation module 3 obtains replacement treatment parameters according to first type field processing, and obtains mortality prediction parameters according to second type field processing.
[0151] Specifically, in view of the relatively complex and inaccurate judgment process of AAV prognosis in the prior art, this solution conducts a correlation analysis on different prognostic outcomes (tumor, renal replacement therapy, death) and survival time as dependent variables, thereby determining the relevant indicators associated with the two typical prognostic outcomes of replacement therapy and death. In the actual prognostic analysis process, the data extraction module 1 automatically extracts these indicators and the data classification module 2 classifies them. Finally, the parameter calculation module 3 obtains the replacement therapy parameters according to the first type of field processing, and obtains the mortality prediction parameters according to the second type of field processing, thereby providing the doctor with relevant reference scores.
[0152] In one embodiment, a modeling module is also included;
[0153] like Fig. 9 As shown, the modeling module includes:
[0154] A sample collection module 41, the sample collection module 41 collects original clinical data from the enrolled patients;
[0155] A preprocessing module 42, the preprocessing module 42 is connected to the sample collection module 41, and the preprocessing module 42 searches for a prognosis outcome field from the original clinical data, and groups the enrolled patients according to the prognosis outcome field to form a prognosis group;
[0156] Prognostic groups included a high-risk group, which included neoplasms, renal replacement therapy, and death;
[0157] A feature extraction module 43, the feature extraction module 43 is connected to the preprocessing module 42, and the feature extraction module 43 analyzes the data feature differences corresponding to each field in the original clinical data for the high-risk group;
[0158] A hypothesis building module 44, the hypothesis building module 44 is connected to the feature extraction module 43, and the hypothesis framework module 44 constructs a hypothesis test for the data feature difference;
[0159] A sample analysis module 45, the sample analysis module 45 is connected to the hypothesis building module 44, and the sample analysis module 45 analyzes each field in the original clinical data for tumor outcomes based on the test hypothesis to predict a first field group associated with tumor outcomes, a second field group associated with renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis screening;
[0160] The sample output module 46 is connected to the sample analysis module 45 , and generates a first type field and a second type field according to the first field group, the second field group and the third field group.
[0161] In one embodiment, clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels;
[0162] like Fig.10 As shown, the data extraction module 1 includes:
[0163] Medical record collection module 11, the medical record collection module 11 collects electronic medical record information for the prognosis patient;
[0164] The field search module 12 is connected to the medical record collection module 11. The field search module 12 searches for the electronic medical record information by using clinical data fields in sequence to obtain medical record fields;
[0165] The field processing module 13 is connected to the field retrieval module 12, and processes the medical record fields to obtain clinical data fields.
[0166] In one embodiment, the first type of fields include: heart involvement at first diagnosis, kidney organ involvement at first diagnosis, BVAS score, serum eGFR, and total complement level;
[0167] The second type of fields includes serum eGFR, age;
[0168] like Fig.11 As shown, the data classification module 2 includes:
[0169] A first search module 21, the first search module 21 searches for an organ involvement condition field from the clinical data field, and selects from the organ involvement conditions the heart involved at the first diagnosis and the kidney involved at the first diagnosis;
[0170] A second search module 22, the second search module 22 is connected to the first search module 21, and the second search module 22 searches for BVAS score, serum eGFR, age and total complement level from the clinical data field;
[0171] The field division module 23 is connected to the second search module 22. The field division module 23 divides the heart involved at the time of first diagnosis, the kidney organ involved at the time of first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type of field, and divides the serum eGFR and age into the second type of field.
[0172] Specifically, in order to achieve effective division of the two types of fields, in this embodiment, the first search module 21 first searches for the organ involvement field from the clinical data field. Since the field records different involved organs in the form of a string, the heart and kidney are matched by regular search to determine whether there is a matching result. If there is, the corresponding field is marked as 1, and if not, it is marked as 0, forming the heart involved at the first diagnosis and the kidney involved at the first diagnosis. Then, the second search module 22 searches for the BVAS score, serum eGFR, age and total complement level from the clinical data field. The BVAS score is usually calculated by the doctor during the first visit. If it is not calculated, the associated field is extracted for the empty value, and the score is calculated and filled in.
[0173] Finally, the field division module 23 divides the heart involvement at the first diagnosis, the kidney involvement at the first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type of field, and divides the serum eGFR and age into the second type of field according to the above analysis results.
[0174] In one embodiment, Fig.12 As shown, the parameter calculation module 3 includes:
[0175] A scale search module 31, the scale search module 31 searches for a first scoring scale according to a first type of field, and searches for a second scoring scale according to a second type of field;
[0176] The parameter calculation module 32 is connected to the scale search module 31. The parameter calculation module 32 processes the first type of field according to the first scoring scale to obtain the alternative treatment parameter, and processes the second type of field according to the second scoring scale to obtain the death parameter.
[0177] Specifically, in order to achieve effective processing of the above parameters, in this embodiment, after classifying the fields, the scale search module 31 first searches for the first type of field to obtain the first rating scale, and searches for the second type of field to obtain the second rating scale. Then, for each field, the parameter calculation module 32 finds the corresponding items in the scale, performs score mapping based on the score intervals in the corresponding items to obtain relevant scores, performs weighted calculation on the scores to obtain the total score, and finally matches the relevant conclusion output.
[0178] Those skilled in the art will appreciate that various aspects of the present invention, or possible implementations of various aspects, may be specifically implemented as systems, methods, or computer program products. Therefore, various aspects of the present invention, or possible implementations of various aspects, may take the form of complete hardware embodiments, complete software embodiments (including firmware, resident software, etc.), or embodiments of combined software and hardware aspects, all collectively referred to herein as "circuits," "modules," or "systems." In addition, various aspects of the present invention, or possible implementations of various aspects, may take the form of computer program products, which refer to computer instructions stored in a memory.
[0179] The memory may be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or apparatuses, or any suitable combination of the foregoing, such as random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable read-only memory (CD-ROM).
[0180] The processor in the computer reads the computer instructions stored in the memory, so that the processor can execute the functional actions specified in each step or the combination of steps in the flowchart; and generate a device for implementing the functional actions specified in each block or the combination of blocks in the block diagram.
[0181] It should be understood that the processor in the computer can be understood as one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components implemented to execute the aforementioned computer instructions.
[0182] Computer instructions can be executed completely on the user's local computer, partially on the user's local computer, as a separate software package, partially on the user's local computer and partially on a remote computer, or completely on a remote computer or server. It should also be noted that in some alternative embodiments, the functions noted in each step in the flow chart or each block in the block diagram may not occur in the order noted in the figure. For example, depending on the functions involved, two steps or two blocks shown in succession may actually be executed roughly simultaneously, or these blocks may sometimes be executed in reverse order.
[0183] The above are only preferred embodiments of the present invention, and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for assisting in predicting the prognosis of ANCA-associated vasculitis, characterized in that: include: Step S1: collecting electronic medical record information of the prognostic patient and extracting clinical data fields from the medical record information; Step S2: classifying the clinical data fields to obtain a first type of fields associated with alternative treatments and a second type of fields associated with mortality rates; Step S3: Obtaining alternative treatment parameters according to the first type of fields, and obtaining mortality prediction parameters according to the second type of fields.
2. The prognosis auxiliary prediction method according to claim 1, characterized in that: Before executing step S1, a modeling process is also included, and the modeling process includes: Step A1: Collect original clinical data of enrolled patients; Step A2: searching for a prognosis outcome field from the original clinical data, and grouping the enrolled patients according to the prognosis outcome field to form prognosis groups; The prognostic grouping includes a high-risk group, which includes tumors, renal replacement therapy and death; Step A3: analyzing the data feature differences of the data features corresponding to each field in the original clinical data for the high-risk group; Step A4: construct hypothesis test for the data feature differences; Step A5: Based on the test hypothesis, each field in the original clinical data is analyzed for tumor outcome to predict a first field group associated with the tumor outcome, a second field group associated with the renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis screening; Step A6: Generate the first type field and the second type field according to the first field group, the second field group and the third field group.
3. The prognosis auxiliary prediction method according to claim 1, characterized in that: The clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels; The step S1 comprises: Step S11: collecting the electronic medical record information for the prognosis patient; Step S12: for the electronic medical record information, the clinical data fields are used to search in sequence to obtain medical record fields; Step S13: Process the medical record field to obtain the clinical data field.
4. The prognosis auxiliary prediction method according to claim 1, characterized in that: The first type of fields include: heart involvement at first diagnosis, kidney involvement at first diagnosis, BVAS score, serum eGFR, and total complement level; The second type of fields include serum eGFR and age; The step S2 comprises: Step S21: searching for an organ involvement condition field from the clinical data field, and screening out the heart involved at the time of the first diagnosis and the kidney involved at the time of the first diagnosis from the organ involvement conditions; Step S22: searching the BVAS score, the serum eGFR, the age and the total complement level from the clinical data field; Step S23: classify the heart involvement at the time of first diagnosis, the kidney involvement at the time of first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type field, and classify the serum eGFR and the age into the second type field.
5. The prognosis auxiliary prediction method according to claim 1, characterized in that: The step S3 comprises: Step S31: searching and obtaining a first rating scale according to the first type of field, and searching and obtaining a second rating scale according to the second type of field; Step S32: Processing the first type of fields according to the first scoring scale to obtain the alternative treatment parameters, and processing the second type of fields according to the second scoring scale to obtain the death parameters.
6. A prognosis auxiliary prediction system for ANCA-associated vasculitis, characterized in that: Used to implement the prognosis auxiliary prediction method according to any one of claims 1 to 5; include: A data extraction module, wherein the data extraction module collects electronic medical record information of the prognosis patient and extracts clinical data fields from the medical record information; a data classification module, the data classification module being connected to the data extraction module, the data classification module classifying the clinical data fields to obtain a first type of field associated with alternative treatments and a second type of field associated with mortality; A parameter calculation module is connected to the data classification module, and the parameter calculation module obtains alternative treatment parameters according to the first type of fields and obtains mortality prediction parameters according to the second type of fields.
7. The prognosis auxiliary prediction system according to claim 6, characterized in that: Also includes modeling modules; The modeling module includes: A sample collection module, which collects original clinical data from enrolled patients; A preprocessing module, the preprocessing module is connected to the sample collection module, the preprocessing module searches for a prognosis outcome field from the original clinical data, and groups the enrolled patients according to the prognosis outcome field to form prognosis groups; The prognostic grouping includes a high-risk group, which includes tumors, renal replacement therapy and death; A feature extraction module, the feature extraction module is connected to the preprocessing module, and the feature extraction module analyzes the data feature differences corresponding to each field in the original clinical data for the high-risk group; A hypothesis building module, the hypothesis building module is connected to the feature extraction module, and the hypothesis framework module constructs a hypothesis test according to the data feature difference; A sample analysis module, the sample analysis module is connected to the hypothesis building module, and the sample analysis module analyzes each field in the original clinical data for tumor outcomes based on the test hypothesis to predict a first field group associated with the tumor outcomes, a second field group associated with the renal replacement therapy based on forward screening, and a third field group associated with death based on univariate and multivariate Cox regression analysis screening; A sample output module, wherein the sample output module is connected to the sample analysis module, and wherein the sample output module generates the first type field and the second type field according to the first field group, the second field group and the third field group.
8. The prognosis auxiliary prediction system according to claim 6, characterized in that: The clinical data fields include organ involvement, BVAS score, serum eGFR, age serum and total complement levels; The data extraction module comprises: A medical record collection module, wherein the medical record collection module collects the electronic medical record information for the prognosis patient; A field search module, the field search module is connected to the medical record collection module, and the field search module uses the clinical data fields to search for the electronic medical record information in sequence to obtain the medical record fields; A field processing module, the field processing module is connected to the field retrieval module, and the field processing module processes the medical record field to obtain the clinical data field.
9. The prognosis auxiliary prediction system according to claim 6, characterized in that: The first type of fields include: heart involvement at first diagnosis, kidney involvement at first diagnosis, BVAS score, serum eGFR, and total complement level; The second type of fields include serum eGFR and age; The data classification module comprises: A first search module, wherein the first search module searches for an organ involvement condition field from the clinical data field, and selects the heart involved at the time of the first diagnosis and the kidney involved at the time of the first diagnosis from the organ involvement conditions; a second search module, the second search module being connected to the first search module, and the second search module searching the BVAS score, the serum eGFR, the age and the total complement level from the clinical data field; A field division module, wherein the field division module is connected to the second search module, and the field division module divides the heart involvement at the first diagnosis, the kidney organ involvement at the first diagnosis, the BVAS score, the serum eGFR and the total complement level into the first type of field, and divides the serum eGFR and the age into the second type of field.
10. The prognosis auxiliary prediction system according to claim 6, characterized in that: The parameter calculation module comprises: A scale search module, wherein the scale search module searches for a first rating scale according to the first type of field, and searches for a second rating scale according to the second type of field; A parameter calculation module, wherein the parameter calculation module is connected to the scale lookup module, wherein the parameter calculation module processes the first type of field according to the first scoring scale to obtain the alternative treatment parameter, and processes the second type of field according to the second scoring scale to obtain the death parameter.