Auxiliary diagnosis method and system based on medical data mining
Through an auxiliary diagnostic system based on medical data mining, patient data is automatically processed, high-risk patients are identified and personalized treatment suggestions are provided, and the problem of time-consuming and labor-consuming processing of data by doctors in the prior art is solved, and diagnostic efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510341369.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-01
AI Technical Summary
The existing medical diagnostic system cannot deeply analyze or associate mining patient data, which leads to doctors needing to manually screen and compare data, which is time-consuming and labor-intensive, and diagnosis depends on personal experience.
Through an auxiliary diagnostic system based on medical data mining, including user management, data collection, population division and visual support modules, SMS verification codes are used to verify identity, data preprocessing and prediction models, automatically identify high-risk patients and provide personalized treatment suggestions.
It improves diagnostic efficiency, reduces the time for doctors to process manual data, can identify high-risk patients in advance and issue early warnings, helps doctors to intervene early, provide personalized treatment plans, and improves the convenience and accuracy of diagnosis.
Smart Images

Figure CN120412968A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of requirement analysis, and particularly to an auxiliary diagnosis method and system based on medical data mining. Background Art
[0002] Traditional diagnostic methods usually carry out laboratory tests and imaging examinations, such as blood tests, urine analysis, biochemical analysis and other laboratory testing means, which provide doctors with more information about the metabolic state, immune function, infection situation, etc. inside the patient's body. The emergence of imaging examination technologies such as X-rays, ultrasounds, and electrocardiograms enables doctors to non-invasively observe the internal structure of the human body and diagnose various diseases such as fractures, tumors, and cardiovascular diseases; When dealing with patient data sources, existing medical diagnostic systems only simply organize and record the data, and use static database records to save the basic information, examination results, etc. of patients. They are unable to deeply analyze or associate and mine the data. Doctors still need to manually screen and compare patient data and rely on personal experience to make diagnostic decisions. This processing method is time-consuming and laborious; In view of the above technical deficiencies, a solution is now proposed. Summary of the Invention
[0003] In view of the deficiencies of the prior art, the present invention provides an auxiliary diagnosis method and system based on medical data mining.
[0004] To achieve the above object, the present invention is realized through the following technical solutions: An auxiliary diagnosis system based on medical data mining includes a user management module, a data acquisition module, a population division module, an analysis module, and a visualization support module; The user management module verifies the identity through a SMS verification code, divides users into patients and doctors. When a patient logs in for the first time, they need to upload basic information and health records, providing a case auxiliary diagnosis function for doctors; The data acquisition module obtains patient case data from the associated hospital database, associates the data of the same patient with the ID card as the identifier, and preprocesses the patient data; The population division module classifies and marks different healthy characteristic populations according to information such as the patient's blood type, weight, age, etc. with a preset algorithm, and divides them into different risk disease susceptible populations, different weight category populations, and different physical constitution populations; The analysis module mines similar treatment response groups by associating patients with the same type of labels, trains historical cases with a decision tree algorithm to obtain treatment decision rules, constructs a prediction model to calculate the case process score P and the living habit score D, calculates the health risk score E according to the formula, sets a threshold to compare the risk scores, and recommends routine examinations for low risks, and sends warning information and adjusts the recommended plan for high risks; The visualization support module displays the health trends and treatment effects of patients in the form of charts, and provides a doctor's visualization interface containing the patient's basic information and treatment advice plan.
[0005] The user management module is responsible for user authentication, account management, and role permission control in the system. It conducts identity verification through the SMS verification code method. After successful verification, the system classifies users into patients and doctors, and provides different function interfaces and services respectively. When a patient user first enters the system, they upload their personal basic information and ID number, and provide information such as the patient's health status record and medical history. The doctor user enters the system to provide case analysis functions for doctors. [[ID=u4]]
[0006] The data collection module obtains a large amount of patient case data from the associated hospital database, and associates the patient's data in different medical institutions and hospitals using the patient's ID card as the unique identifier. These case data usually include the patient's basic information, medical records, examination results, treatment plans, and treatment effects. It cleans the outliers and inconsistent records in the case data, removes the incorrect input or invalid duplicate data in the data, fills the missing data with the mean value, and standardizes the parameters with different units.
[0007] The population division module divides patients into different categories according to the patients' basic information. The module divides patients into several different groups and marks them by judging blood type, weight, and age according to a preset algorithm. Each group has different health characteristics. It divides and marks the susceptible population of risk diseases according to blood type, divides and marks different weight categories according to the body mass index (BMI), and divides and marks different physical constitution populations according to gender and age.
[0008] The analysis module associates patients through the same type of label, and discovers patient groups with similar treatment responses or disease courses. These groups show similar treatment effects under the same treatment plan. The analysis module uses the decision tree machine learning algorithm to train the historical case data, thereby obtaining treatment decision rules. Based on the treatment decision rules and the label characteristics of the patient group, a disease treatment prediction model is constructed. The prediction model calculates the case process score P and the lifestyle score D according to the patient's case data. The analysis module will also provide personalized treatment advice plans for the attending doctors using the prediction model based on the patient's label information.
[0009] The analysis module calculates the risk score E according to the formula. The specific formula is as follows: ; Among them, represents the weight coefficient of the case process score, represents the weight coefficient of age, Let \( \omega \) represent the weight coefficient of the living habit score, \( P \) represent the case progression score, \( W \) represent the age in years, and \( D \) represent the living habit score. The above weight coefficient is obtained through training with historical data and automatically estimated by training the data through regression analysis methods. For a preset risk threshold \( T \), when \( E \leq T \), the patient is identified as a low-risk patient, and routine examinations are recommended to the attending physician. When \( E > T \), the patient is identified as a high-risk patient, and a warning message is sent to the attending physician, the treatment recommendation plan is adjusted, and the patient is reminded to undergo examinations.
[0010] The visualization support module is responsible for displaying the health data of the patient, showing the health change trend and treatment effect of the patient through charts and curve graphs, and providing a visualization interface for the doctor. The visualization interface displays the basic information of the patient, the health change trend graph, and the treatment recommendation plan.
[0011] The specific method of the auxiliary diagnosis based on medical data mining is as follows: S1. The user verifies their identity through a text message verification code. After successful login, different function interfaces are assigned according to the identity. The patient interface is used to upload personal basic information. S2. The data acquisition module obtains a large amount of patient case data from associated medical institutions and performs preprocessing. S3. The population division module divides the patients into different categories and marks the patients with labels. S4. The analysis module analyzes the patient case data to generate a disease treatment prediction model, calculates the risk score based on the case progression score, age, and living habit score of the patient, performs risk control according to the preset threshold, and provides a treatment recommendation plan for the doctor. S5. The visualization support module displays the health data of the patient to the doctor and shows the treatment recommendation plan to assist the doctor in making decisions.
[0012] The present invention provides an auxiliary diagnosis method and system based on medical data mining. Compared with the prior art, it has the following beneficial effects: Through the present invention, data collection, cleaning, and preprocessing are carried out by associating with the databases of some hospitals to obtain the associated case data information of the same patient, saving time without manual data collation. The system based on the patient's case data can calculate the risk score, identify high-risk patients in advance and issue warnings, helping doctors to intervene early to reduce the deterioration of the disease. The system helps doctors formulate treatment recommendation plans based on the patient's medical record data and patient label data, accelerating the doctor's understanding of the patient's situation and making the doctor's work more efficient and convenient. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the principle framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] Please refer to Figure 1 , this application provides an auxiliary diagnosis system based on medical data mining, including a user management module, a data collection module, a population classification module, an analysis module, and a visualization support module; The user management module verifies the identity through a SMS verification code, divides users into patients and doctors. When a patient logs in for the first time, they need to upload basic information and health records, providing a case auxiliary diagnosis function for doctors. The data collection module obtains patient case data from the associated hospital database, associates the data of the same patient with the ID card as the identifier, and preprocesses the patient data. The population classification module classifies and marks different healthy characteristic populations according to information such as the patient's blood type, weight, age, etc. using a preset algorithm, and divides them into different risk disease susceptible populations, different weight category populations, and different physical constitution populations. The analysis module mines similar treatment response groups by associating patients with the same type of labels, trains historical cases with a decision tree algorithm to obtain treatment decision rules, constructs a prediction model to calculate the case progress score P and the lifestyle score D, calculates the health risk score E according to the formula, sets a threshold to compare the risk scores, recommends routine examinations for low risks, and sends warning messages and adjusts the recommended plan for high risks. The visualization support module displays the patient's health trends and treatment effects in charts, and provides a doctor's visualization interface containing the patient's basic information and treatment recommended plan.
[0016] The user management module is responsible for user identity verification, account management, and role permission control in the system. It verifies the identity through the SMS verification code. After successful verification, the system divides users into patients and doctors, providing different function interfaces and services respectively. When a patient user enters the system for the first time, they upload personal basic information and ID card number, and provide information such as the patient's health status record and disease history. The doctor user enters the system to provide a case analysis function for doctors.
[0017] The data acquisition module obtains case data of a large number of patients from the associated hospital database, and associates the data of patients in different medical institutions and hospitals through the patient's ID card as the unique identifier. These case data usually include the patient's basic information, medical records, examination results, treatment plans, and treatment effects, etc. Clean the outliers and inconsistent records in the case data, remove the incorrect input or invalid duplicate data in the data, fill in the missing data with the mean value, and standardize the parameters of different units.
[0018] The population division module divides patients into different categories according to the basic information of the patients. The module divides patients into several different groups and marks them by judging blood type, weight, and age through a preset algorithm. Each group of people has different health characteristics. Divide and mark the susceptible population of risk diseases according to blood type, divide into different weight categories and mark according to the body mass index BMI, and divide and mark different physical constitution populations according to gender and age.
[0019] According to the patient's blood type (type A, type B, type AB, type O), the system will identify the groups that may be more susceptible to specific diseases. People with different blood types may show different susceptibilities in certain immune diseases or infectious diseases, and the system will make corresponding marks according to the blood type.
[0020] According to the patient's weight and height, the system calculates the body mass index BMI and divides the patient into different weight categories, such as normal weight, overweight, obesity, or emaciation, etc. The risks of chronic diseases such as cardiovascular diseases, diabetes, and hypertension are different among people in different weight categories, and the system divides into different weight categories through BMI.
[0021] According to the patient's age group, the system divides the patients into different age groups, including children, young people, middle-aged people, and the elderly. There are significant differences in the incidence of diseases, treatment responses, and physical functions among people of different age groups. The elderly population faces more disease risks related to aging.
[0022] The analysis module associates patients through the same type of label, and discovers patient groups with similar treatment responses or disease courses. These groups show similar treatment effects under the same treatment plan. The analysis module uses the decision tree machine learning algorithm to train the historical case data to obtain treatment decision rules, and constructs a disease treatment prediction model based on the treatment decision rules and the label characteristics of the patient group. The prediction model calculates the case process score P and the living habit score D according to the patient's case data. The analysis module will also provide personalized treatment advice plans for the attending physicians using the prediction model according to the patient's label information.
[0023] During the training process, the decision tree algorithm will identify the features that have the most influence on the patient's treatment outcome. The decision tree recursively analyzes the patient data by selecting the optimal features. The patient's living habits are an important feature. By analyzing the impact of living habits on diseases, the decision tree determines which living habits play a key role in the treatment of specific diseases and obtains a series of decision rules through a tree structure.
[0024] The analysis module calculates the risk score E according to the formula. The specific formula is as follows: ; Among them, represents the weight coefficient of the case progress score, represents the weight coefficient of age, represents the weight coefficient of the living habit score. P represents the case progress score, W represents age in years, and D represents the living habit score. The above weight coefficients are obtained through training with historical data and automatically estimated by training the data through regression analysis methods; A preset risk threshold T is set. When E ≤ T, the patient is identified as low-risk, and a routine examination is recommended to the attending doctor. When E > T, the patient is identified as high-risk, and a warning message is sent to the attending doctor and the treatment recommendation plan is adjusted, and the patient is reminded to have an examination.
[0025] The visualization support module is responsible for displaying the patient's health data, showing the patient's health change trend and treatment effect through charts and curve graphs, and providing a visualization interface for doctors. The visualization interface displays the patient's basic information, health change trend graph, and treatment recommendation plan.
[0026] A specific method for auxiliary diagnosis based on medical data mining is as follows: S1. The user verifies their identity through a text message verification code. After successful login, different function interfaces are allocated according to the identity. The patient interface is used to upload personal basic information; S2. The data acquisition module obtains a large amount of patient case data from affiliated medical institutions and performs preprocessing; S3. The population division module divides the patients into different categories and marks the patients with labels; S4. The analysis module analyzes the patient case data to generate a disease treatment prediction model, calculates the risk score based on the patient's case progress score, age, and living habit score, performs risk control according to the preset threshold, and provides a treatment recommendation plan to the doctor; S5. The visualization support module displays the patient's health data to the doctor and shows the treatment recommendation plan to assist the doctor in making decisions.
[0027] Specific working process: The user conducts identity verification through SMS verification code. After successful verification, the user is classified into patients and doctors according to their identity information. When patients log in for the first time, they upload personal basic information, ID number, health records and other information. After doctors enter the system, they can view the case data and relevant health information of patients. The data collection module associates the case data of the same patient in different hospitals through the patient's ID card, integrates a complete data file, and then preprocesses the data file. The population classification module classifies patients using a preset algorithm based on information such as the patient's blood type, weight, age, etc. The analysis module trains a treatment prediction model through the decision tree algorithm on historical case data, calculates the case progress score and lifestyle score of the patient, and calculates the health risk score, providing a treatment advice plan for doctors. A preset risk threshold is used for risk classification: when E ≤ T, routine examinations are recommended; when E > T, the system sends a warning message to the doctor and adjusts the treatment advice plan.
[0028] Furthermore, the present invention collects, cleans and preprocesses data by associating partial hospital databases, obtains the associated case data information of the same patient, saves time without manual data collation. Based on the patient's case data, the system can calculate the risk score, identify high-risk patients in advance and issue warnings, helping doctors intervene early to reduce the deterioration of the disease. The system helps doctors formulate treatment advice plans according to the patient's medical record data and patient label data, speeds up the doctor's understanding of the patient's situation, and makes the doctor's work more efficient and convenient.
[0029] Some of the data in the above formula are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0030] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. An auxiliary diagnosis system based on medical data mining, characterized in that, include: User management module, data collection module, population segmentation module, analysis module and visualization support module; The user management module verifies identity through SMS verification codes and divides users into patients and doctors. Patients need to upload basic information and health records when logging in for the first time, providing doctors with case-assisted diagnosis functions. The data acquisition module obtains patient case data from the associated hospital database, uses the ID card as an identifier to associate the same patient data, and pre-processes the patient data; The population segmentation module uses a preset algorithm to classify and label people with different health characteristics based on the patient's blood type, weight, age and other information, and divides them into people who are susceptible to different risk diseases, people in different weight categories and people with different physical conditions; The analysis module discovers groups with similar treatment responses by associating patients with similar labels. It uses a decision tree algorithm to train historical cases to derive treatment decision rules. It then constructs a predictive model to calculate the case progression score P and lifestyle score D. It then calculates the health risk score E based on a formula. It then sets a threshold to compare risk scores. For low-risk patients, it recommends routine checkups. For high-risk patients, it issues warnings and adjusts the recommended plan. The visualization support module uses charts to display patients' health trends and treatment effects, and provides doctors with a visual interface containing basic patient information and treatment recommendations.
2. The auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that, The user management module is responsible for the identity authentication, account management and role authority control of users in the system. Identity authentication is performed through SMS verification codes. After the verification is passed, the system divides users into patients and doctors, providing different functional interfaces and services respectively. When patient users enter the system for the first time, they upload their personal basic information and ID number, and provide the patient's health status record, medical history and other information. Doctor users enter the system to provide case analysis functions for doctors.
3. An auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that, The data acquisition module obtains a large amount of patient case data from the associated hospital database, and uses the patient's ID card as a unique identifier to associate the patient's data in different medical institutions and hospitals. These case data usually include the patient's basic information, medical records, examination results, treatment plans and treatment effects, etc., cleans up outliers and inconsistent records in the case data, eliminates incorrect input or invalid duplicate data in the data, fills in the missing data with the mean, and standardizes parameters of different units.
4. An auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that, The population classification module divides patients into different categories based on their basic information. The module divides patients into several different groups of people and marks them by judging their blood type, weight and age according to a preset algorithm. Each group of people has different health characteristics. The susceptible group of risk diseases is divided and marked according to blood type, divided into different weight categories according to body mass index (BMI) and marked, and divided into different physical groups according to gender and age and marked.
5. An auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that, The analysis module associates patients through the same type of tags, discovers patient groups with similar treatment responses or disease courses, and these groups show similar treatment effects under the same treatment plan. The analysis module uses the decision tree machine learning algorithm to train historical case data to obtain treatment decision rules, and constructs a disease treatment prediction model based on the treatment decision rules and the label characteristics of the patient group. The prediction model calculates the case progress score P and the lifestyle score D according to the patient's case data. The analysis module also provides a personalized treatment advice plan for the attending physician using the prediction model based on the patient's label information.
6. The auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that, The analysis module calculates the risk score E according to the formula. The specific formula is as follows: ; Among them, represents the weight coefficient of the case progression score, represents the weight coefficient of age, represents the weight coefficient of the living habit score, P represents the case progression score, W represents age in years, D represents the living habit score; the above weight coefficients are obtained through training with historical data and automatically estimated by training the data through the regression analysis method. A preset risk threshold T is set. When E ≤ T, the patient is identified as low-risk, and routine examinations are recommended to the attending physician; when E > T, the patient is identified as high-risk, a warning message is sent to the attending physician, the treatment advice plan is adjusted, and the patient is reminded to undergo examinations.
7. The auxiliary diagnosis system based on medical data mining according to claim 1, characterized in that The visualization support module is responsible for displaying the patient's health data, showing the patient's health change trend and treatment effect through charts and line graphs, and providing a visualization interface for the doctor. The visualization interface displays the patient's basic information, health change trend graph, and treatment advice plan.
8. An auxiliary diagnosis method based on medical data mining, characterized in that, The specific method of an auxiliary diagnosis based on medical data mining is as follows: S1. The user verifies their identity through a text message verification code. After successful login, different function interfaces are assigned according to the identity. The patient interface is used to upload personal basic information. S2. The data collection module obtains a large amount of patient case data from affiliated medical institutions and preprocesses it. S3. The population division module divides patients into different categories and marks patients with tags. S4. The analysis module analyzes the patient case data to generate a disease treatment prediction model, calculates the risk score according to the patient's case progress score, age, and lifestyle score, performs risk control according to the preset threshold, and provides a treatment advice plan for the doctor. S5. The visualization support module displays the patient's health data to the doctor and shows the treatment advice plan to assist the doctor in making decisions.