Method, device and electronic equipment for predicting probability of disease

By combining the BRF model with GAN and FP-Growth algorithms, the accuracy and stability issues of SLE complication risk assessment were resolved, enabling efficient and accurate prediction of SLE patient complications and supporting individualized treatment.

CN119380978BActive Publication Date: 2026-02-17CAPINFO CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411959084.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2026-02-17
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing technologies are insufficient to comprehensively and accurately assess the risk of complications in patients with systemic lupus erythematosus (SLE), especially in predicting a small number of high-risk cases, where accuracy and stability are lacking, and there is a lack of fine-grained mining and deep learning applications for complex complication associations.

Method used

By employing a Balanced Random Forest (BRF) model combined with Generative Adversarial Networks (GANs) and Frequent Pattern Growth (FP-Growth) algorithms, and through natural language processing and data equalization, keywords associated with the target disease are extracted, labeled, and predicted to improve the comprehensiveness and accuracy of the assessment.

Benefits of technology

It achieves efficient and accurate prediction of complication risks in SLE patients, especially in the identification of a small number of high-risk cases, improving the stability and generalization ability of the model and supporting the implementation of individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380978B_ABST
    Figure CN119380978B_ABST
Patent Text Reader

Abstract

The application provides a disease probability prediction method and device and electronic equipment, at least one target keyword associated with a target disease and disease description information of a target patient are obtained, for each target keyword, the target keyword is marked according to whether the target keyword exists in the disease description information, and a first marking result is obtained, and the disease description information and the first marking result corresponding to each target keyword are input into a pre-trained BRF model to output the probability that the target patient has the target disease, in this way, at least one target keyword associated with the target disease is determined in advance, after each target keyword is marked based on the disease description information of the target patient, the first marking result corresponding to each target keyword and the disease description information are jointly used as the input of the BRF model, which can improve the comprehensiveness and accuracy of evaluating whether the target patient has the target disease.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of disease prediction, in particular to a disease probability prediction method and device and electronic equipment. BACKGROUND

[0002] Systemic lupus erythematosus is a complex autoimmune disease, SLE patients often face the risk of multiple serious complications, including kidney damage, cardiovascular disease, nervous system damage, etc., these complications have a significant impact on the survival rate and quality of life of patients, predicting the risk of systemic lupus erythematosus helps to take individualized treatment programs and intervention measures as soon as possible; the traditional SLE risk assessment method mainly relies on the experience of doctors or is based on a simple statistical model, which is difficult to comprehensively and accurately assess the risk of patients suffering from systemic lupus erythematosus. SUMMARY

[0003] The purpose of the present application is to provide a disease probability prediction method and device and electronic equipment to comprehensively and accurately assess the risk of patients suffering from diseases.

[0004] The disease probability prediction method provided by the present application comprises: acquiring at least one target keyword associated with a target disease and disease description information of a target patient; for each target keyword, marking the target keyword according to whether the target keyword exists in the disease description information to obtain a first marking result corresponding to the target keyword; inputting the disease description information and the first marking result corresponding to each target keyword into a pre-trained BRF model to output the probability of the target patient suffering from the target disease through the BRF model.

[0005] Further, the target keyword is determined by: acquiring an initial medical data set; wherein the initial medical data set includes first diagnosis and treatment data of a plurality of patients; selecting second diagnosis and treatment data corresponding to a preset field from the first diagnosis and treatment data of each patient according to the preset field; wherein the preset field is a field associated with the target disease; inputting the second diagnosis and treatment data of each patient into a preset segmentation model respectively to perform segmentation processing on each second diagnosis and treatment data to obtain a segmentation result corresponding to each second diagnosis and treatment data respectively; and performing screening processing on the segmentation results corresponding to all second diagnosis and treatment data to obtain the target keyword.

[0006] Further, the step of screening the segmentation results corresponding to all the second diagnosis and treatment data to obtain the target keyword includes: screening first segmentation results associated with the target disease from the segmentation results corresponding to all the second diagnosis and treatment data; determining whether there is at least one segmentation combination in the first segmentation results; wherein the plurality of segmentation words in each segmentation combination have the same meaning; if there is at least one segmentation combination, replacing the plurality of segmentation words in each segmentation combination with a specified segmentation word; wherein the specified segmentation word has the same meaning as each segmentation word in the segmentation combination; obtaining the target keyword based on the first segmentation result and each specified segmentation word.

[0007] Further, the BRF model is trained in the following manner: for each second diagnosis and treatment data, if the segmentation result corresponding to the second diagnosis and treatment data contains at least one target keyword, the second diagnosis and treatment data is determined as sample diagnosis and treatment data; for each sample diagnosis and treatment data, each target keyword is labeled respectively according to whether each target keyword exists in the sample diagnosis and treatment data, to obtain a second labeling result corresponding to each target keyword respectively; the plurality of second labeling results respectively corresponding to all the sample diagnosis and treatment data are taken as a first data set; the first data set is subjected to data balancing processing to obtain a second data set; wherein the second data set includes a plurality of sample data; each sample data is labeled with first labeling information; the first labeling information is used to indicate whether the patient corresponding to the sample data has the target disease; based on the second data set, a current sample data is determined, the current sample data is input into the initial model, to output a prediction result corresponding to the current sample data through the initial model; based on the first labeling information and the prediction result corresponding to the current sample data, a loss value is calculated, the initial model is updated based on the loss value, the next group of sample data is taken as a new current sample data, and the step of inputting the current sample data into the initial model is repeatedly executed until the loss value converges, to obtain the trained BRF model.

[0008] Further, the step of performing data balancing processing on the first data set to obtain a second data set includes: processing the first data set through a preset generative adversarial network to obtain a third data set; resampling the third data set to obtain the second data set; wherein in the second data set, the number of sample data of patients with the target disease is the same as the number of sample data of patients without the target disease.

[0009] Further, the method further includes: performing data mining processing on the third data set according to the FP-Growth algorithm to output a correlation result between each target keyword and other target keywords.

[0010] Further, the preset field includes: diagnosis disease, disease description, and disease code.

[0011] The application provides a disease probability prediction device, which comprises an acquisition module, a marking module and an output module.

[0012] The application provides an electronic device, which comprises a processor and a memory.

[0013] The application provides a machine readable storage medium, which stores machine executable instructions.

[0014] The application provides a disease probability prediction method, device and electronic device, which acquire at least one target keyword associated with a target disease and disease description information of a target patient, mark each target keyword according to whether the target keyword exists in the disease description information, obtain a first marking result corresponding to the target keyword, input the disease description information and the first marking result corresponding to each target keyword into a BRF model trained in advance, and output a probability that the target patient has the target disease through the BRF model. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the application or the technical solutions in the prior art, the following will briefly introduce the drawings needed in the specific embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative effort.

[0016] Figure 1 The application provides a disease probability prediction method, device and electronic device, which acquire at least one target keyword associated with a target disease and disease description information of a target patient, mark each target keyword according to whether the target keyword exists in the disease description information, obtain a first marking result corresponding to the target keyword, input the disease description information and the first marking result corresponding to each target keyword into a BRF model trained in advance, and output a probability that the target patient has the target disease through the BRF model.

[0017] Figure 2 FIG. 1 is a flowchart of another method for predicting a disease probability according to an embodiment of the present application;

[0018] Figure 3 FIG. 2 is a structural diagram of a device for predicting a disease probability according to an embodiment of the present application;

[0019] Figure 4 FIG. 3 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] The technical solutions of the present application will be described clearly and completely below with reference to the embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0021] Systemic lupus erythematosus (SLE) is a complex autoimmune disease. Its unknown etiology, diverse and unpredictable symptoms pose many challenges to clinical diagnosis and treatment. SLE patients often face the risk of multiple serious complications, including kidney damage, cardiovascular disease, and nervous system damage, which have a significant impact on patient survival and quality of life.

[0022] SLE symptoms are complex, highly variable, and significantly different between patients, making it challenging to accurately predict complications. Current prediction models are mostly based on single or limited clinical variables, failing to fully utilize complex disease-related features, making it difficult to comprehensively and accurately assess patient complication risks. In addition, data imbalance is a common problem: the amount of data for a few serious complications is small, increasing the difficulty of model prediction, especially the lack of systematic solutions for generating a small number of target samples, resulting in poor model prediction accuracy and stability. Therefore, traditional models still perform poorly in accurately identifying a small number of high-risk cases.

[0023] In summary, the existing technology has the following objective shortcomings:

[0024] Traditional SLE risk assessment methods mainly rely on the experience of doctors or simple statistical models, which are difficult to comprehensively and accurately assess the risk of patients suffering from systemic lupus erythematosus, and cannot efficiently handle the complex correlation of SLE complications. The existing prediction technology lacks an effective data supplement mechanism for a small number of complication samples, resulting in poor prediction effect for a small number of complications. In terms of data processing and model application, the existing technology does not provide a simple and user-friendly input and prediction interface, making it difficult to achieve rapid deployment and clinical application. In terms of complication correlation analysis, the existing technology mainly relies on traditional data mining methods, lacks fine-grained mining of complex complication correlation, and fails to fully utilize the advantages of deep learning and generative models to supplement data.

[0025] Based on this, the embodiment of the present application provides a disease probability prediction method, device and electronic equipment, which can be applied to applications that need to predict the risk of patients suffering from a target disease.

[0026] In order to facilitate the understanding of the present embodiment, first, a disease probability prediction method disclosed by the present embodiment is introduced, as shown in the following formula (I), the method comprises the following steps: Figure 1

[0027] Step S102, at least one target keyword associated with the target disease and disease description information of the target patient are obtained.

[0028] The above-mentioned target disease can be any disease, such as systemic lupus erythematosus, hypertension, etc.; the above-mentioned target keyword usually includes the symptoms, complications, and treatment prescriptions corresponding to the target disease, and the number of target keywords corresponding to different target diseases is usually different; the above-mentioned disease description information can be a natural language form of diagnosis, such as case history, medication history, etc., which can be obtained by a doctor according to the patient's self-reporting; in actual implementation, when the probability of the target patient suffering from the target disease needs to be evaluated, at least one target keyword corresponding to the target disease and the disease description information of the target patient are usually needed.

[0029] Step S104, for each target keyword, whether the target keyword exists in the disease description information is determined, and the target keyword is marked to obtain a first marking result corresponding to the target keyword.

[0030] Each target keyword can be extracted from the above-mentioned disease description information to determine whether each target keyword exists in the disease description information, and each target keyword is marked according to the extraction of each target keyword to obtain a first marking result corresponding to the target keyword, which is used to indicate whether the target keyword exists in the disease description information. ​

[0031] Step S106, input the disease description information and the first marking result corresponding to each target keyword into the pre-trained BRF model to output the probability of the target patient suffering from the target disease through the BRF model.

[0032] The above-mentioned BRF (Balanced Random Forest) model is an ensemble learning method that combines the prediction results of multiple decision trees to improve the accuracy and stability of the model. In actual implementation, the above-mentioned disease description information and the first marking result corresponding to each target keyword can be input into the pre-trained BRF model, and the probability of the target patient suffering from the target disease can be predicted, such as the possibility of suffering from the target disease is 52%.

[0033] The above-mentioned disease probability prediction method obtains at least one target keyword associated with the target disease and disease description information of the target patient. For each target keyword, the target keyword is marked according to whether the target keyword exists in the disease description information to obtain the first marking result corresponding to the target keyword. The disease description information and the first marking result corresponding to each target keyword are input into the pre-trained BRF model to output the probability of the target patient suffering from the target disease through the BRF model. This way, at least one target keyword associated with the target disease is determined in advance. After marking each target keyword based on the disease description information of the target patient, the first marking result corresponding to each target keyword and the disease description information are jointly used as the input of the BRF model, which can improve the comprehensiveness and accuracy of evaluating whether the target patient suffers from the target disease.

[0034] The embodiment of the present application also provides another disease probability prediction method, which is implemented based on the above-mentioned embodiment method. The method comprises the following steps:

[0035] Step one, obtaining at least one target keyword associated with the target disease and disease description information of the target patient;

[0036] The target keyword is determined by the following steps 10 to 13:

[0037] Step 10, obtaining an initial medical data set; wherein the initial medical data set comprises first diagnosis and treatment data of a plurality of patients;

[0038] The first diagnosis and treatment data of the plurality of patients can be medical related data of a plurality of patients in a specified region within a specified time period, such as medical related data in a medical insurance data set of a certain city in the past 5 years. In actual implementation, in order to comprehensively and accurately determine the target keyword associated with the target disease, the first diagnosis and treatment data of the plurality of patients can be obtained as needed.

[0039] Step 11, according to the preset field, the second diagnosis and treatment data corresponding to the preset field is screened out from the first diagnosis and treatment data corresponding to each patient; wherein, the preset field is a field associated with the target disease.

[0040] The above-mentioned preset field is usually a field associated with the target disease, for example, it can include: diagnosed disease, disease description, disease coding, etc.; wherein, the diagnosed disease can be understood as the final disease diagnosis result of the patient, which may be the target disease or other diseases; the above-mentioned disease description can be a natural language form of diagnosis situation arranged by the doctor according to the patient's disease description; the above-mentioned disease coding can refer to the coding in the International Classification of Diseases (ICD) system, which is a standard method that can be used for classification and coding of diseases, injuries and poisoning; in actual implementation, the first diagnosis and treatment data of the plurality of patients obtained usually includes data corresponding to a plurality of different fields, some of which are associated with disease assessment, such as disease description, and some of which are not associated with disease assessment, such as the cost of the patient; in order to reduce the amount of data calculation, the preset field can be set in advance, and the first diagnosis and treatment data of each patient is screened according to the preset field to screen the second diagnosis and treatment data corresponding to the preset field.

[0041] Step 12, input the second diagnosis and treatment data of each patient into the preset word segmentation model respectively, to perform word segmentation processing on each second diagnosis and treatment data, to obtain the word segmentation result corresponding to each second diagnosis and treatment data respectively.

[0042] The above-mentioned preset word segmentation model can be an LTP (Language Technology Platform) model, which can provide rich, efficient and accurate natural language processing including Chinese word segmentation, part-of-speech tagging, named entity recognition, dependency syntax analysis, semantic role labeling, etc.; in actual implementation, the second diagnosis and treatment data of each patient can be input into the preset word segmentation model respectively, and each second diagnosis and treatment data is processed by the preset word segmentation model to obtain the word segmentation result corresponding to each second diagnosis and treatment data respectively.

[0043] Step 13, performing screening processing on the word segmentation results corresponding to all second diagnosis and treatment data to obtain the target keyword.

[0044] This step 13 can be implemented by the following steps 130 to 133:

[0045] Step 130, screening the first word segmentation result associated with the target disease from the word segmentation results corresponding to all second diagnosis and treatment data;

[0046] Step 131, determining whether there is at least one word segmentation combination in the first word segmentation result; wherein the multiple words in each word segmentation combination have the same meaning;

[0047] Step 132, if there is at least one word segmentation combination, replacing the multiple words in each word segmentation combination with a specified word; wherein the specified word has the same meaning as each word in the word segmentation combination;

[0048] Step 133, obtaining the target keyword based on the first word segmentation result and each specified word.

[0049] In actual implementation, for each piece of second diagnosis and treatment data, the word segmentation result corresponding to the second diagnosis and treatment data usually contains multiple words, each of which may or may not be related to the target disease. In order to reduce data calculation and improve prediction accuracy, the first word segmentation result associated with the target disease can be screened from all word segmentation results corresponding to the second diagnosis and treatment data. Considering that in actual application, the same meaning may have different description methods, words with the same meaning can be combined into a word segmentation combination. Therefore, after obtaining the first word segmentation result, it can be determined whether there is at least one word segmentation combination in the first word segmentation result; for example, the first word segmentation result contains headache and headache, although headache and headache are two different description methods, but express the same meaning, therefore, headache and headache can be used as a word segmentation combination. If there is one or more word segmentation combinations, the multiple words in each word segmentation combination can be replaced by a specified word. The specified word can be any one of the multiple words in the word segmentation combination, of course, other words can also be selected for replacement, as long as the specified word has the same meaning as the multiple words in the word segmentation combination. After each word segmentation combination in the first word segmentation result is replaced by a corresponding specified word, the target keyword is obtained; for example, 200 words are collected in the first word segmentation result, after the above replacement processing, 33 target keywords are finally obtained, for example, the target keywords corresponding to systemic lupus erythematosus include: rash, edema, ambrisentan, combined white blood cells, tadalafil, two, methotrexate, etc. Through this way, the uniqueness of each word in the target keyword can be ensured.

[0050] Step two, for each target keyword, marking the target keyword according to whether the target keyword exists in the disease description information, to obtain a first marking result corresponding to the target keyword.

[0051] Step three, input the disease description information and the first mark result corresponding to each target keyword into the pre-trained BRF model to output the probability of the target patient suffering from the target disease through the BRF model.

[0052] For example, the disease description information of the target patient in natural language form is as follows: the patient recently had respiratory tract infection, chest tightness, hypothyroidism, proteinuria, lower limb venous dysfunction, severe osteoporosis, and took Omeprazole, Amlodipine Besylate and Tadalafil, etc., and had thrombocytopenia and leukopenia; the above-mentioned 33 target keywords are extracted and marked as present or not. The pre-trained BRF model is called for prediction, and the prediction result is {False: 51.50, True: 48.50}, i.e. the probability of suffering from SLE is about 48.50%, which should be paid attention to in clinical or other application scenarios.

[0053] The above-mentioned BRF model can be trained through the following steps 30-35:

[0054] Step 30, for each second diagnosis and treatment data, if the second diagnosis and treatment data contains at least one target keyword in the word segmentation result corresponding to the second diagnosis and treatment data, the second diagnosis and treatment data is determined as sample diagnosis and treatment data.

[0055] Step 31, for each sample diagnosis and treatment data, each target keyword is marked respectively according to whether each target keyword exists in the sample diagnosis and treatment data, and the second mark result corresponding to each target keyword is obtained.

[0056] Step 32, the plurality of second mark results respectively corresponding to all sample diagnosis and treatment data are taken as the first data set.

[0057] According to the above-mentioned each target keyword and the word segmentation result respectively corresponding to each second diagnosis and treatment data, the second diagnosis and treatment data containing one or more target keywords are retrieved, the plurality of retrieved second diagnosis and treatment data can be taken as a plurality of sample diagnosis and treatment data, in each sample diagnosis and treatment data, whether each target keyword exists is traversed and the corresponding second mark result is marked, for example, if a target keyword exists, it can be represented as "yes", if it does not exist, it can be represented as "no", etc.; the plurality of second mark results respectively corresponding to all sample diagnosis and treatment data can be taken as the first data set, the form of the first data set is a table with target keywords as features, for example, if the number of target keywords is 33, the first data set is a table with 33 target keywords as features.

[0058] Step 33, data equalization processing is performed on the first data set to obtain a second data set; wherein the second data set includes a plurality of sample data; each sample data is labeled with first labeling information; the first labeling information is used to indicate whether the patient corresponding to the sample data has the target disease.

[0059] This step 33 can be implemented through steps 330 to 332.

[0060] Step 330, the first data set is processed by a preset generative adversarial network to obtain a third data set.

[0061] Step 331, resampling is performed on the third data set to obtain the second data set; wherein in the second data set, the number of sample data of patients with the target disease is the same as the number of sample data of patients without the target disease.

[0062] Step 332, according to the FP-Growth algorithm, the third data set is processed by data mining to output the correlation result between each target keyword and other target keywords.

[0063] The above-mentioned generative adversarial network (GAN) is a deep learning model, which is mainly used to generate new data samples similar to the training data. The above-mentioned FP-Growth (Frequent Pattern Growth) is an algorithm for frequent pattern mining, which is particularly suitable for processing large data sets, mainly used for mining frequent item sets and association rules, and widely used in market basket analysis, recommendation systems and other data mining tasks. In actual implementation, all sample diagnosis and treatment data corresponding to the first data set may only have a small part of patients corresponding to the target disease, and other patients corresponding to other diseases, and the proportion of patients with the target disease is relatively uneven, therefore, more data is generated by the preset generative adversarial network to obtain a third data set; for example, the first data set corresponds to 140,000 data, only 171 data correspond to patients diagnosed with SLE, and the proportion of patients with the disease is uneven, and more data can be generated by the generative adversarial network, for example, 5,000 data diagnosed as SLE can be generated, and finally in the third data set, 5,171 data diagnosed as SLE and 137,157 data diagnosed as SLE are obtained.

[0064] In actual implementation, SMOTE (Synthetic Minority Over-sampling Technique) can be used for resampling. SMOTE is a resampling method for solving the problem of unbalanced data set, especially in classification tasks. The core idea of SMOTE is to enhance the minority class samples by generating synthetic samples, so as to improve the performance of the classifier. Through SMOTE resampling, the characteristics of the minority samples can be highlighted, and data balancing can be achieved. For example, continuing with the above example, the sample distribution in the second data set obtained after resampling is Counter({False: 137157, True: 137157}). After the above steps, the data of patients with SLE and the data of patients without SLE in the second data set have reached the balanced effect. Specifically, the data amount ratio can reach 1:1.

[0065] In the training process of the BRF model, after obtaining the third data set, the correlation between each target keyword and other target keywords can be verified according to the FP-Growth algorithm. For example, the target keywords include symptoms, complications, etc. The correlation between each symptom and complication can be verified. The correlation result is mainly used as an auxiliary conclusion.

[0066] Step 34, based on the second data set, determine the current sample data, input the current sample data into the initial model, and output the prediction result corresponding to the current sample data through the initial model.

[0067] Step 35, based on the first label information and the prediction result corresponding to the current sample data, calculate the loss value, update the initial model based on the loss value, take the next group of sample data as new current sample data, and repeat the step of inputting the current sample data into the initial model until the loss value converges, and obtain the trained BRF model.

[0068] The above BRF is an improved random forest algorithm, which aims to handle unbalanced classification problems. It combines the advantages of random forest and sample balancing techniques to improve the classification performance of minority class samples. It is also widely used in the medical field, such as disease prediction and diagnosis, and identification of rare disease cases. In actual implementation, the current sample data can be determined based on the second data set. For example, taking the number of target keywords as 33 as an example, the second data set includes multiple sample data, each of which includes 33 target keywords corresponding to the label result. The first sample data can be taken as the current sample data and input into the initial model to update the weight parameters in the initial model. The second sample data can be taken as the current sample data, and the above process can be repeated until the loss value converges, and the trained BRF model can be obtained.

[0069] For easy understanding, see Figure 2 Another flowchart of the prediction method of the probability of illness is shown. The diagnosis information is extracted from the medical insurance data set of a city (corresponding to the above-mentioned filtering of the second diagnosis and treatment data corresponding to the preset field from the first diagnosis and treatment data of each patient), the diagnosis information is processed by the LTP model, and the target keywords are obtained by screening. For each second diagnosis and treatment data, if the second diagnosis and treatment data contains at least one target keyword in the word segmentation result, the second diagnosis and treatment data is determined as sample diagnosis and treatment data. For each sample diagnosis and treatment data, each target keyword is marked according to whether each target keyword exists in the sample diagnosis and treatment data, and the second marking result corresponding to each target keyword is obtained. The plurality of second marking results corresponding to all sample diagnosis and treatment data are used as a first data set. If the first data set is extremely unbalanced, the GANs adversarial network is used to make up for the problem of fewer disease data samples, and a third data set is obtained. According to the FP-Growth algorithm, the third data set is subjected to data mining processing, and the correlation result between each target keyword and other target keywords is output as an auxiliary conclusion. The third data set is subjected to SMOTE resampling to balance the data, and a second data set is obtained. The BRF model is trained based on the second data set, and the diagnosis and treatment information of the patient is input into the BRF model, so that the risk probability of the patient suffering from the target disease can be predicted.

[0070] The following introduces a result of actual verification. In the test set, there are 34333 data in total, of which 34287 are patients not diagnosed with SLE, and 46 are patients diagnosed with SLE. In the model prediction result, only 1 of the 34287 patients not suffering from SLE is misdiagnosed as suffering from SLE (or has a higher risk of suffering from SLE); for the 46 patients suffering from SLE, 37 are predicted to have SLE, and 9 are not diagnosed with SLE (or have a lower risk of suffering from SLE). In general, for the rare disease SLE in medicine, the prediction accuracy reaches 80.4%, and the model training effect can be considered good.

[0071] In the use end, the demand end can directly input the disease description information (including case history, medication history) in natural language form to the BRF model to return the prediction result, which is relatively simple and flexible.

[0072] The prediction method of the probability of illness above combines the LTP model (Lexicalized Tree Projection) in natural language processing to automatically analyze the disease description information of the patient, and integrates the target keywords containing the complications and symptom characteristics of the target disease, further enhancing the context accuracy of the data, and can provide more efficient and rich input information sources for training the BRF model.

[0073] The FP-Growth (frequent itemset growth) algorithm is innovatively introduced, which can efficiently discover frequent association rules between the complications of the target disease, capture the association patterns between the complications, and provide multi-dimensional association features for the model, thereby improving the prediction accuracy of the BRF model.

[0074] The generative adversarial network (GANs) is used to generate realistic synthetic data for the minority class samples, and the prediction bias problem caused by the lack of minority samples is innovatively solved through balancing processing, thereby effectively enhancing the generalization ability of the model. Before the model training, the SMOTE (synthetic minority over-sampling technique) is introduced to balance the data set, which avoids over-fitting to the majority class samples, so that the model training can fully capture the feature distribution of different categories, and the stability and accuracy of the prediction are realized. The BRF model is innovatively used to realize the classification task while enhancing the recognition ability of the minority class, so that the high prediction performance can be maintained under the condition of sample imbalance.

[0075] Taking the target disease SLE as an example, the complications of SLE often develop rapidly, and once they occur, the clinical treatment is difficult and the prognosis is poor. For example, lupus nephritis and lupus encephalopathy and other complications can cause irreversible organ damage, which seriously affects the prognosis of patients. The prediction of the risk of systemic lupus erythematosus by the method helps to take individualized treatment programs and intervention measures as soon as possible, thereby delaying or avoiding disease aggravation, and striving for the best treatment window period for patients. The early identification of high-risk patients by the method can also significantly reduce the consumption of medical resources, and effectively improve the treatment effect, reduce the hospitalization rate and the long-term complication cost caused by complications.

[0076] The embodiment of the application provides a prediction device for disease probability, as shown in Figure 3 The device comprises: an acquisition module 30 configured to acquire at least one target keyword associated with a target disease and disease description information of a target patient; a marking module 31 configured to mark each target keyword according to whether the target keyword exists in the disease description information, to obtain a first marking result corresponding to the target keyword; and an output module 32 configured to input the disease description information and the first marking result corresponding to each target keyword into a pre-trained BRF model, to output the probability of the target patient suffering from the target disease by the BRF model.

[0077] The prediction device of the disease probability determines at least one target keyword associated with the target disease in advance, and after marking each target keyword based on the disease description information of the target patient, uses the first marking result corresponding to each target keyword and the disease description information as the input of the BRF model, which can improve the comprehensiveness and accuracy of the evaluation of whether the target patient has the target disease.

[0078] Further, the device further comprises a target keyword determination module, which is configured to: obtain an initial medical data set; wherein the initial medical data set comprises first diagnosis and treatment data of a plurality of patients; according to a preset field, filter second diagnosis and treatment data corresponding to the preset field from the first diagnosis and treatment data of each patient; wherein the preset field is a field associated with the target disease; input the second diagnosis and treatment data of each patient into a preset segmentation model to perform segmentation processing on each second diagnosis and treatment data, and obtain a segmentation result corresponding to each second diagnosis and treatment data; and filter the segmentation results corresponding to all second diagnosis and treatment data to obtain the target keyword.

[0079] Further, the target keyword determination module is further configured to: filter first segmentation results associated with the target disease from the segmentation results corresponding to all second diagnosis and treatment data; determine whether there is at least one segmentation combination in the first segmentation result; wherein the plurality of segments in each segmentation combination have the same meaning; if there is at least one segmentation combination, replace the plurality of segments in each segmentation combination with a specified segment; wherein the specified segment has the same meaning as each segment in the segmentation combination; and obtain the target keyword based on the first segmentation result and each specified segment.

[0080] Further, the device further comprises a BRF model training module, the BRF model training module is used for: for each second diagnosis and treatment data, if the second diagnosis and treatment data corresponding to the word segmentation result contains at least one target keyword, the second diagnosis and treatment data is determined as sample diagnosis and treatment data; for each sample diagnosis and treatment data, whether each target keyword exists in the sample diagnosis and treatment data is determined, and each target keyword is marked respectively to obtain a second marking result corresponding to each target keyword; a plurality of second marking results respectively corresponding to all sample diagnosis and treatment data are used as a first data set; the first data set is subjected to data equalization processing to obtain a second data set; wherein the second data set comprises a plurality of sample data; each sample data is labeled with first labeling information; the first labeling information is used to indicate whether the patient corresponding to the sample data has a target disease; based on the second data set, a current sample data is determined, the current sample data is input into an initial model, and the initial model is used to output a prediction result corresponding to the current sample data; based on the first labeling information and the prediction result corresponding to the current sample data, a loss value is calculated, the initial model is updated based on the loss value, the next group of sample data is used as new current sample data, the step of inputting the current sample data into the initial model is repeatedly executed until the loss value converges, and a trained BRF model is obtained.

[0081] Further, the BRF model training module is used for: processing the first data set by a preset generative adversarial network to obtain a third data set; resampling the third data set to obtain the second data set; wherein in the second data set, the number of sample data of patients with a target disease is the same as the number of sample data of patients without a target disease.

[0082] Further, the BRF model training module is used for: according to the FP-Growth algorithm, the third data set is subjected to data mining processing, and the correlation result between each target keyword and other target keywords is output.

[0083] Further, the preset field comprises: a diagnosed disease, a disease description, and a disease code.

[0084] The prediction device for disease probability provided in the embodiment of the application has the same implementation principle and technical effects as the foregoing prediction method for disease probability, and for brevity of description, the part of the prediction device for disease probability that is not mentioned can refer to the corresponding content in the foregoing prediction method for disease probability.

[0085] The embodiment of the application further provides an electronic device, referring to Figure 4 The electronic device comprises a processor 130 and a memory 131, the memory 131 stores machine executable instructions capable of being executed by the processor 130, and the processor 130 executes the machine executable instructions to implement the foregoing prediction method for disease probability.

[0086] Further, Figure 4 The electronic device also includes a bus 132 and a communication interface 133, and the processor 130, the communication interface 133 and the memory 131 are connected through the bus 132.

[0087] The memory 131 can include a high-speed random access memory (RAM), and can also include a non-volatile memory such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 133 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 132 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0088] The processor 130 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by integrated logic circuits of hardware in the processor 130 or instructions in the form of software. The above processor 130 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiment of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 131, and the processor 130 reads the information in the memory 131, and combines the hardware to complete the steps of the method of the above embodiment.

[0089] The embodiment of the present application further provides a machine readable storage medium which stores machine executable instructions, when the machine executable instructions are invoked and executed by a processor, the machine executable instructions cause the processor to implement the method for predicting the disease probability, and the specific implementation can be referred to the method embodiment, and will not be repeated here.

[0090] The computer program product of the method for predicting the disease probability, the device and the electronic equipment provided by the embodiment of the present application comprises a computer readable storage medium which stores program codes, and the instructions included in the program codes can be used to execute the method described in the foregoing method embodiment, and the specific implementation can be referred to the method embodiment, and will not be repeated here.

[0091] If the functions are realized in the form of software function units and sold or used as independent products, the functions can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or parts of the technical solutions which essentially contribute to the prior art can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media which can store program codes.

[0092] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of predicting a probability of a disease, characterized by, The method comprises: obtaining at least one target keyword associated with a target disease and disease description information of a target patient; for each target keyword, marking the target keyword according to whether the target keyword exists in the disease description information to obtain a first marking result corresponding to the target keyword; wherein the first marking result is used to indicate whether the target keyword exists in the disease description information; inputting the disease description information and the first marking result corresponding to each target keyword into a pre-trained BRF model to output a probability that the target patient has the target disease through the BRF model; The target keyword is determined by the following method: obtaining an initial medical data set; wherein the initial medical data set comprises first diagnosis and treatment data of a plurality of patients; According to a preset field, the second diagnosis and treatment data corresponding to the preset field are screened from the first diagnosis and treatment data of each patient; wherein the preset field is a field associated with the target disease; the preset field includes: diagnosis disease, disease description, disease code; the diagnosis disease is the final disease diagnosis result of the patient; the disease description is a natural language form of the diagnosis obtained by the doctor according to the patient's disease description; the disease code is the code in the international disease classification system; inputting the second diagnosis and treatment data of each patient into a pre-set word segmentation model respectively to perform word segmentation processing on each second diagnosis and treatment data to obtain a word segmentation result corresponding to each second diagnosis and treatment data respectively; filtering the word segmentation results corresponding to all second diagnosis and treatment data to obtain the target keyword; The step of filtering the word segmentation results corresponding to all second diagnosis and treatment data to obtain the target keyword comprises: from the word segmentation results corresponding to all second diagnosis and treatment data, filtering out first word segmentation results associated with the target disease; determining whether at least one word segmentation combination exists in the first word segmentation result; wherein a plurality of words in each word segmentation combination have the same meaning; if there is at least one word segmentation combination, for each word segmentation combination, replace a plurality of words in the word segmentation combination with a specified word; wherein the specified word has the same meaning as each word in the word segmentation combination; the specified word is any one word in the word segmentation combination; based on the first word segmentation result and each specified word, a target keyword is obtained.

2. The method of claim 1, wherein, The BRF model is trained by the following method: for each second diagnosis and treatment data, if the word segmentation result corresponding to the second diagnosis and treatment data contains at least one target keyword, the second diagnosis and treatment data is determined as sample diagnosis and treatment data; for each sample diagnosis and treatment data, according to whether each target keyword exists in the sample diagnosis and treatment data, marking each target keyword to obtain a second marking result corresponding to each target keyword respectively; the plurality of second marking results corresponding to all sample diagnosis and treatment data are used as a first data set; perform data balancing processing on the first data set to obtain a second data set; wherein the second data set includes a plurality of sample data; each of the sample data is labeled with first label information; the first label information is used to indicate whether the sample data corresponds to a patient suffering from the target disease; determine a current sample data based on the second data set, input the current sample data into an initial model, and output a prediction result corresponding to the current sample data through the initial model; based on the first label information corresponding to the current sample data and the prediction result, calculate a loss value, update the initial model based on the loss value, take the next group of sample data as new current sample data, and repeatedly execute the step of inputting the current sample data into the initial model until the loss value converges, to obtain the trained BRF model.

3. The method of claim 2, wherein, The step of performing data balancing processing on the first data set to obtain a second data set includes: processing the first data set through a preset generative adversarial network to obtain a third data set; resampling the third data set to obtain a second data set; wherein in the second data set, the number of sample data of patients suffering from the target disease is the same as the number of sample data of patients not suffering from the target disease.

4. The method of claim 3, wherein, The method further includes: According to the FP-Growth algorithm, performing data mining processing on the third data set to output a correlation result between each target keyword and other target keywords.

5. A device for predicting a probability of a disease, characterized by, The device includes: an acquisition module configured to acquire at least one target keyword associated with a target disease and disease description information of a target patient; a marking module configured to, for each target keyword, mark the target keyword according to whether the target keyword exists in the disease description information to obtain a first marking result corresponding to the target keyword; wherein the first marking result is used to indicate whether the target keyword exists in the disease description information; an output module configured to input the disease description information and the first marking result corresponding to each target keyword into a pre-trained BRF model to output a probability that the target patient suffers from the target disease through the BRF model; The device further includes a target keyword determination module, which is configured to: acquire an initial medical data set; wherein the initial medical data set includes first diagnosis and treatment data of a plurality of patients; According to a preset field, filter second diagnosis and treatment data corresponding to the preset field from the first diagnosis and treatment data corresponding to each patient; wherein the preset field is a field associated with the target disease; the preset field includes: a diagnosed disease, a disease condition description, and a disease code; the diagnosed disease is a final disease diagnosis result of a patient; the disease condition description is a natural language form of a diagnosis condition obtained by a doctor according to a patient's disease description; and the disease code is a code in the international disease classification system. The second diagnosis and treatment data of each of the patients are respectively input into a preset word segmentation model to perform word segmentation processing on each of the second diagnosis and treatment data, to obtain a word segmentation result corresponding to each of the second diagnosis and treatment data; The word segmentation results corresponding to all the second diagnosis and treatment data are subjected to screening processing to obtain the target keyword; The target keyword determination module is further configured to: screen, from the word segmentation results corresponding to all the second diagnosis and treatment data, a first word segmentation result associated with the target disease; determine whether there is at least one word segmentation combination in the first word segmentation result; wherein a plurality of word segments in each of the word segmentation combinations have the same meaning; if there is at least one word segmentation combination, replace, for each of the word segmentation combinations, the plurality of word segments in the word segmentation combination with a specified word segment; wherein the specified word segment has the same meaning as each of the word segments in the word segmentation combination; and the specified word segment is any one of the word segments in the word segmentation combination; obtain the target keyword based on the first word segmentation result and each of the specified word segments.

6. An electronic device, comprising: A processor and a memory are included, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the disease probability prediction method of any one of claims 1-4.

7. A machine-readable storage medium, characterized in that, The machine readable storage medium stores machine executable instructions, and when the machine executable instructions are called and executed by the processor, the machine executable instructions cause the processor to implement the disease probability prediction method of any one of claims 1-4.

Citation Information

Patent Citations

  • Heart-disease intelligent prediction and heart health care information recommendation system and method

    CN108091396A

  • Medical unbalanced data classification method based on biased random forest model

    CN116072302A

  • Keyword extraction method, device, computer equipment and storage medium

    US20240354507A1