Construction method, system and equipment of diagnostic model, storage medium and program product

By acquiring and processing historical diagnostic information, performing abnormal analysis and feature extraction, the problems of low data quality and limited generalization capabilities in the existing technology are solved, and an efficient and accurate disease diagnosis model is constructed.

CN120089328APending Publication Date: 2025-06-03CHINA TELECOM YIKANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411952330.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, the low quality of training data, insufficient feature extraction effectiveness and limited model generalization capabilities have resulted in limited accuracy and reliability of disease diagnosis models.

Method used

By obtaining historical diagnostic information, abnormal analysis and repair, extracting features and dividing data sets, training the initial training model based on these data, and fusing model parameters to build an efficient disease diagnosis model.

Benefits of technology

It improves the quality of training data, reduces the impact of noise on model training, and enhances the generalization ability of the model, so that it can diagnose diseases more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089328A_ABST
    Figure CN120089328A_ABST
Patent Text Reader

Abstract

The invention provides a diagnostic model construction method, system and device, a storage medium and a program product. The construction method comprises the following steps: acquiring a plurality of pieces of historical diagnosis information as a first data set; the historical diagnosis information comprises at least one group of historical diagnosis information, and each group of historical diagnosis information comprises the following parameter types: physical symptom data, pathology detection data and treatment scheme data; performing anomaly analysis on the first data set, and repairing abnormal data in the first data set to obtain a second data set; performing feature extraction on the second data set, and dividing the second data set into a plurality of third data sets according to a feature extraction result; and training an initial training model based on the third data to obtain a disease diagnosis model. According to the method and the device, the diagnosis capability of the model facing new and unseen data, namely the generalization capability of the model, is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of medical diagnosis, and particularly to a method, system, device, storage medium, and program product for constructing a diagnosis model. Background Art

[0002] With the rapid growth and increasing complexity of medical data, accurate disease diagnosis has become increasingly challenging. Traditional diagnostic methods rely on doctors' experience and limited detection means, and are easily affected by subjective factors and data incompleteness. In recent years, the application of artificial intelligence and machine learning technologies in the medical field has provided new ways to improve diagnostic accuracy. However, constructing an efficient and accurate disease diagnosis model still faces many technical problems, including data quality issues, the effectiveness of feature extraction, and the generalization ability of the model. Summary of the Invention

[0003] The technical problem to be solved by the present disclosure is to overcome the defects of low-quality training data, insufficient effectiveness of feature extraction, and limited generalization ability of the model in the prior art, and to provide a method, system, device, storage medium, and program product for constructing a diagnosis model.

[0004] The present disclosure solves the above technical problems through the following technical solutions:

[0005] The present disclosure provides a method for constructing a disease diagnosis model, and the construction method includes:

[0006] Obtain a number of historical diagnosis information as a first data set; the historical diagnosis information includes at least one group of the historical diagnosis information, and each group of the historical diagnosis information includes the following parameter types: physical symptom data, pathological test data, and treatment plan data;

[0007] Perform anomaly analysis on the first data set, and repair the abnormal data in the first data set to obtain a second data set;

[0008] Extract features from the second data set, and divide the second data set into a number of third data sets according to the results of feature extraction;

[0009] Train an initial training model based on the third data to obtain a disease diagnosis model.

[0010] Optionally, the performing anomaly analysis on the first data set and repairing the abnormal data in the first data set to obtain a second data set includes:

[0011] Determine the abnormal data and normal data in the first data set;

[0012] Input the abnormal data into a data repair model to obtain the repaired abnormal data; wherein, the data repair model is generated by training with the normal data;

[0013] Use the repaired abnormal data and the normal data as the second data set.

[0014] Optionally, before the step of inputting the abnormal data into a data repair model to obtain the repaired abnormal data, it includes:

[0015] Divide the normal data into several subsets;

[0016] Successively use one of the subsets as the validation set, and the subsets other than the validation set as the training set, and use the training set and the validation set to train and validate the initial data repair model respectively;

[0017] Determine the fitting effect of the initial data repair model according to the result of the validation, and adjust the parameters of the initial data repair model according to the fitting effect until the fitting effect reaches a preset threshold.

[0018] Optionally, the step of extracting features from the second data set and dividing the second data set into several third data sets according to the result of the feature extraction includes:

[0019] Calculate the features of the data corresponding to the target parameter type in each group of the historical diagnosis information in the second data set; the features include central tendency and / or dispersion degree;

[0020] Encode each group of the historical diagnosis information according to the features;

[0021] Divide the second data set into several third data sets according to the encoding corresponding to each group of the historical diagnosis information; each third data set corresponds to one encoding.

[0022] Optionally, the step of training an initial training model based on the third data to obtain a disease diagnosis model includes:

[0023] Use several of the third data sets as training data to train several initial training models respectively;

[0024] Obtain the model parameters of each trained initial training model, and perform a fusion process on the model parameters;

[0025] Construct the disease diagnosis model based on the fused model parameters; the architecture of the disease diagnosis model is consistent with that of the initial training model.

[0026] Optionally, training the initial training model based on the third data to obtain a disease diagnosis model further includes:

[0027] Combining all the third data sets into an overall data set to train an initial training model;

[0028] Using the trained initial training model as the disease diagnosis model.

[0029] The present disclosure provides a system for constructing a disease diagnosis model, the construction system including:

[0030] A data acquisition module, configured to acquire a plurality of historical diagnosis information as a first data set; the historical diagnosis information includes at least one group of the historical diagnosis information, and each group of the historical diagnosis information includes the following parameter types: physical symptom data, pathological test data, and treatment plan data;

[0031] A data analysis module, configured to perform anomaly analysis on the first data set, and repair the abnormal data in the first data set to obtain a second data set;

[0032] A feature extraction module, configured to extract features from the second data set, and divide the second data set into a plurality of third data sets according to the results of the feature extraction;

[0033] A model training module, configured to train an initial training model based on the third data to obtain a disease diagnosis model.

[0034] Optionally, the data analysis module is specifically configured to:

[0035] Determine the abnormal data and normal data in the first data set;

[0036] Input the abnormal data into a data repair model to obtain the repaired abnormal data; wherein, the data repair model is generated by training with the normal data;

[0037] Using the repaired abnormal data and the normal data as the second data set.

[0038] Optionally, the construction system further includes:

[0039] A division module, configured to divide the normal data into a plurality of subsets;

[0040] A training module, configured to sequentially use one of the subsets as a validation set, and the subsets other than the validation set as training sets, and use the training sets and the validation set to train and validate an initial data repair model respectively;

[0041] A fitting module, configured to determine the fitting effect of the initial data repair model according to the result of the verification, and adjust the parameters of the initial data repair model according to the fitting effect until the fitting effect reaches a preset threshold.

[0042] Optionally, the feature extraction module is specifically configured to:

[0043] Calculate the features of the data corresponding to the target parameter type in each group of the historical diagnosis information in the second data set; the features include central tendency and / or dispersion degree;

[0044] Encode each group of the historical diagnosis information according to the features;

[0045] Divide the second data set into several third data sets according to the encodings corresponding to each group of the historical diagnosis information; each third data set corresponds to one encoding.

[0046] Optionally, the model training module is specifically configured to:

[0047] Use several of the third data sets as training data to train several initial training models respectively;

[0048] Obtain the model parameters of each trained initial training model, and perform a fusion process on the model parameters;

[0049] Construct the disease diagnosis model based on the fused model parameters; the architecture of the disease diagnosis model is the same as that of the initial training model.

[0050] Optionally, the model training module is specifically configured to:

[0051] Merge all the third data sets into an overall data set to train an initial training model;

[0052] Use the trained initial training model as the disease diagnosis model.

[0053] The present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor, where the processor implements the method for constructing the diagnosis model according to any one of the above when executing the computer program.

[0054] The present disclosure provides a computer-readable storage medium, on which a computer program is stored, characterized in that the computer program implements the method for constructing the diagnosis model according to any one of the above when being executed by a processor.

[0055] The present disclosure provides a computer program product, including a computer program, characterized in that the computer program implements the method for constructing the diagnosis model according to any one of the above when being executed by a processor.

[0056] Based on the common knowledge in this field, the above preferred conditions can be combined arbitrarily to obtain the preferred embodiments of the present disclosure.

[0057] The positive and progressive effects of the present disclosure are as follows: By detecting and repairing abnormal data, the quality of the training data is ensured, and the influence of noise on model training is reduced. The diverse dataset partitioning and model training methods increase the diversity of the model training data, enabling the model to better learn the distribution and patterns of the data. This method improves the diagnostic ability of the model when facing new and unseen data, that is, the generalization ability of the model. Description of the Drawings

[0058] Figure 1 It is a flowchart of a method for constructing a diagnostic model provided by an exemplary embodiment of the present disclosure;

[0059] Figure 2 It is a flowchart of step 102 provided by an exemplary embodiment of the present disclosure;

[0060] Figure 3 It is a flowchart of step 103 provided by an exemplary embodiment of the present disclosure;

[0061] Figure 4 It is a flowchart of step 104 provided by an exemplary embodiment of the present disclosure;

[0062] Figure 5 It is another flowchart of step 104 provided by an exemplary embodiment of the present disclosure;

[0063] Figure 6 It is a schematic diagram of the modules of a diagnostic model construction system provided by an exemplary embodiment of the present disclosure;

[0064] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed Embodiments

[0065] The present disclosure will be further described below by way of examples, but the present disclosure is not limited to the scope of the described examples.

[0066] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. The use of prefix words such as ordinal numbers for distinguishing described objects in the embodiments of the present disclosure does not constitute a restriction on the described objects. For the statements of the described objects, refer to the descriptions in the claims or the context of the embodiments, and no redundant restrictions should be formed due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0067] In the embodiments of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0068] Embodiment 1

[0069] Figure 1 It is a flowchart of a method for constructing a diagnosis model provided for an exemplary embodiment of the present disclosure.

[0070] The present disclosure provides a method for constructing a disease diagnosis model, and the construction method includes:

[0071] Step 101: Obtain a number of historical diagnosis information as the first data set; the historical diagnosis information includes at least one set of historical diagnosis information, and each set of historical diagnosis information includes the following parameter types: physical symptom data, pathological test data, and treatment plan data.

[0072] The core of this step 101 is data collection. The "historical diagnosis information" mentioned here refers to the past patient diagnosis records, which usually contain information about the patient's physical condition, disease manifestations, laboratory test results, and treatment measures taken by doctors. These data are crucial for training machine learning models because the models need to learn the patterns in these historical data to predict new cases.

[0073] Here is a specific example: Suppose we want to build a data set for a cardiovascular disease diagnosis model. We can obtain historical diagnosis information from the following sources:

[0074] 1. Electronic medical record system, including:

[0075] Physical symptom data of patient A: Chest pain, shortness of breath.

[0076] Pathological test data: Electrocardiogram shows ST segment elevation, and blood test shows elevated myocardial enzymes.

[0077] Treatment plan data: Administer anticoagulants, nitroglycerin, and arrange coronary angiography.

[0078] 2. Hospital database, including:

[0079] Physical symptom data of Patient B: Fatigue, palpitations.

[0080] Pathological test data: Echocardiogram shows reduced left ventricular ejection fraction.

[0081] Treatment plan data: Initiate ACE inhibitor treatment and recommend cardiac rehabilitation training.

[0082] 3. Clinic records, including:

[0083] Physical symptom data of Patient C: Dizziness, sweating.

[0084] Pathological test data: Blood test shows low blood pressure and high blood sugar.

[0085] Treatment plan data: Supplement fluids and electrolytes, adjust the dosage of hypoglycemic drugs.

[0086] 4. Public health database, including:

[0087] Physical symptom data of Patient D: History of heart attack in the family.

[0088] Pathological test data: Genetic screening shows the presence of certain gene mutations related to heart disease.

[0089] Treatment plan data: Regularly monitor cardiovascular health status and provide advice on lifestyle improvement.

[0090] After collecting the above information, it is necessary to integrate these scattered data into a structured first dataset. For example, a table can be created where each row represents the diagnostic record of a patient, and the columns correspond to physical symptom data, pathological test data, and treatment plan data respectively. And this process usually involves data encoding. For categorical data (such as physical symptoms, treatment plans), one-hot encoding or label encoding can be used. For continuous data (such as blood test indicators), the numerical values can be directly retained or standardized.

[0091] In this way, a dataset containing historical diagnostic information of multiple parameter types can be constructed, laying a foundation for subsequent data processing and model training.

[0092] It should be understood that when collecting data, it is necessary to ensure the privacy and security of the data and comply with relevant medical data protection regulations. The data should be as comprehensive and accurate as possible to reduce biases and errors in model training. For missing or incomplete data, it needs to be processed in the subsequent data preprocessing stage.

[0093] Step 102: Perform anomaly analysis on the first dataset and repair the abnormal data in the first dataset to obtain the second dataset.

[0094] The purpose of this step 102 is to identify and repair the abnormal data in the first dataset through anomaly analysis, obtaining a cleaner and more accurate second dataset. Abnormal data may be caused by measurement errors, recording errors, or other reasons, and they may have a negative impact on the performance of the model.

[0095] Optionally, referring to Figure 2 it can be known that step 102 specifically includes:

[0096] Step 1021: Determine the abnormal data and normal data in the first dataset.

[0097] The purpose of this step 1021 is to identify the abnormal data that does not conform to the conventional pattern or has errors in the dataset, as well as the normal data that conforms to the expected pattern, for subsequent data cleaning and processing. Abnormal data refers to data that is significantly inconsistent with other data and may be data that deviates from the normal range due to measurement errors, recording errors, or other reasons. Abnormal data may have a negative impact on the training and prediction of the model. Normal data refers to data that conforms to the distribution law of most data points and is within a reasonable range. These data are the basis for model training.

[0098] Here, based on step 1021, a specific example is given: Suppose a dataset containing cardiovascular and cerebrovascular disease diagnosis information has been collected and sorted out, including the following fields: patient age, blood pressure, cholesterol level, blood sugar level, whether smoking, family medical history, etc. Among them, some examples of abnormal data are:

[0099] 1. Patient age:

[0100] Normal range: Usually for adults, such as 18 - 100 years old.

[0101] Abnormal data: Recorded as 3 years old or 150 years old, which are obviously incorrect records.

[0102] 2. Blood pressure:

[0103] Normal range: Systolic blood pressure 90 - 140 mmHg, diastolic blood pressure 60 - 90 mmHg.

[0104] Abnormal data: Systolic blood pressure recorded as 250 mmHg or diastolic blood pressure recorded as 15 mmHg, these values are far beyond the normal range and may be measurement errors.

[0105] 3. Cholesterol level:

[0106] Normal range: Generally considered that LDL cholesterol below 130 mg / dL is normal.

[0107] Abnormal data: Recorded as 500 mg / dL, far higher than the normal upper limit, which may be a laboratory testing error.

[0108] 4. Blood glucose level:

[0109] Normal range: Fasting blood glucose 70 - 100 mg / dL.

[0110] Abnormal data: Recorded as 500 mg / dL, which may be an extreme case of diabetic ketoacidosis or a data entry error.

[0111] Therefore, an example of normal data can be:

[0112] Patient age: 55 years old, within the age range of normal adults.

[0113] Blood pressure: Systolic blood pressure 120 mmHg, diastolic blood pressure 80 mmHg, within the normal blood pressure range.

[0114] Cholesterol level: LDL cholesterol 100 mg / dL, within the normal range.

[0115] Blood glucose level: Fasting blood glucose 85 mg / dL, also within the normal range.

[0116] In addition, for the methods of handling abnormal data, the following ways can be included:

[0117] 1. Deletion: If the abnormal data is caused by an obvious error and cannot be corrected, deleting this record can be considered.

[0118] 2. Correction: If the abnormal data can be verified and corrected through other information sources, corresponding modifications should be made. This is also one of the technical problems that the present invention focuses on solving.

[0119] 3. Marking: For abnormal data whose error status cannot be determined, marking can be carried out, and these data should be particularly concerned or excluded in subsequent analysis.

[0120] By determining abnormal data and normal data, the quality of the data set can be ensured, and the accuracy and reliability of model training can be improved.

[0121] Step 1022: Input the abnormal data into the data repair model to obtain the repaired abnormal data; wherein, the data repair model is trained and generated by normal data.

[0122] Optionally, step 1022 specifically includes: dividing the normal data into several subsets; sequentially using one of the subsets as the validation set and the subsets other than the validation set as the training set, and training and validating the initial data repair model using the training set and the validation set respectively; determining the fitting effect of the initial data repair model according to the validation result, and adjusting the parameters of the initial data repair model according to the fitting effect until the fitting effect reaches a preset threshold.

[0123] This step 1022 is a key step in data preprocessing, aiming to repair or correct outliers in the dataset through a specific model. The following is an interpretation of this step and specific examples:

[0124] 1. The role of the data repair model: The data repair model is trained based on normal data, and it can learn the distribution rules and feature patterns of normal data. When abnormal data is input, the model will try to "repair" these outliers according to the learned normal patterns to make them closer to the distribution of normal data.

[0125] 2. The training and validation process of the model: The optional step of dividing the normal data into several subsets and sequentially using them as the validation set and the training set is a cross-validation method. In this way, the fitting effect of the data repair model can be evaluated, and the model parameters can be adjusted according to the validation result until the preset fitting threshold is reached.

[0126] Here, based on step 1022, a specific example is given: Suppose there is a dataset containing patients' blood pressure data, and some of the blood pressure values are wrongly recorded as extremely high or extremely low. To repair these outliers, a data repair model can be constructed. The specific steps are as follows:

[0127] 1. Prepare normal data: Screen out the records with blood pressure values within the normal range from the dataset as normal data.

[0128] 2. Divide the dataset: Randomly divide the normal data into 5 subsets, such as A, B, C, D, and E.

[0129] 3. Training and validation: First, use subset A as the training set and subset B as the validation set to train the initial data repair model. Evaluate the fitting effect of the model on validation set B, for example, by calculating the mean squared error (MSE). If the fitting effect is not ideal, adjust the model parameters and retrain. Then, sequentially use other subset combinations (such as C as the training set and D as the validation set) to repeat the above process.

[0130] 4. Determine the best model: After multiple rounds of training and validation, select the model with the best fitting effect as the final data repair model.

[0131] 5. Repair abnormal data: Input the abnormal blood pressure values identified in the dataset into the finally determined data repair model. The model will output the repaired blood pressure values, which should be closer to the distribution range of normal data.

[0132] Example results:

[0133] Original abnormal data: The systolic blood pressure of patient X was recorded as 250 mmHg (abnormally high).

[0134] Repaired data: After being processed by the data repair model, the systolic blood pressure of patient X was corrected to 130 mmHg (within the normal range).

[0135] It should be understood that the effect of the data repair model depends on the representativeness and quality of the training data. When using the repaired data for further analysis or modeling, its reliability and impact should be carefully evaluated. In some cases, if the abnormal data reflects real extreme situations or rare events, over-repair may lead to loss of information; therefore, the integrity and accuracy of the data should be balanced according to the specific application scenarios and requirements.

[0136] Step 1023: Use the repaired abnormal data and normal data as the second dataset.

[0137] This step 1023 is an important link in the data processing flow. The purpose is to form a more complete and higher-quality dataset by integrating the repaired abnormal data and the originally normal data without abnormalities. This new dataset (the second dataset) will be used for subsequent analysis, modeling, or other data processing tasks. Repairing abnormal data helps improve the overall quality and accuracy of the data. Combining with normal data can more comprehensively reflect the true distribution and characteristics of the data, thereby enhancing the generalization ability and prediction accuracy of the model.

[0138] Here, based on step 1023, a specific example is given: Suppose a dataset on patient health indicators is being processed, including multiple physiological parameters such as blood pressure and blood sugar. Some abnormal values (such as too high or too low blood pressure) were found in the preliminary analysis and corresponding repairs were made. The example process is as follows:

[0139] 1. Original dataset: Contains 1000 patient records, and some of the blood pressure values in the records are significantly abnormal (such as the systolic blood pressure reaching 300 mmHg).

[0140] 2. Abnormal data processing: Repair these abnormal blood pressure values through the data repair model. For example, correct the unreasonable 300 mmHg to a more reasonable 140 mmHg.

[0141] 3. Forming a second data set: The repaired blood pressure values ​​are combined with other normal data that are not abnormal to form a new data set. At this time, the second data set contains 1,000 records, all of which are within a reasonable range (such as systolic blood pressure 90-140 mmHg).

[0142] It should be understood that before forming the second data set, it should be ensured that the repair process of abnormal data is reasonable and effective. The repaired data should maintain the distribution characteristics and authenticity of the original data as much as possible to avoid introducing new deviations or errors. When using the second data set for analysis, it should be clearly stated that the data has been repaired so that others can understand the data processing process and potential impact.

[0143] Based on the above steps 1021 to 1023, a specific example is given here: suppose that the following abnormal data is encountered in the data set for building a heart disease diagnosis model:

[0144] Patient ID: F008;

[0145] Chest pain severity (1-10 points): 9 points;

[0146] Heart rate (beats / minute): 120 beats / minute;

[0147] Myocardial enzyme level (U / L): 5U / L.

[0148] In this example, patient F008 has a high degree of chest pain (9 points) and a fast heart rate (120 beats / minute), but an abnormally low myocardial enzyme level (only 5 U / L), which is obviously inconsistent with the clinical experience reflected by his symptoms and heart rate, and may be an abnormal data point. Therefore, based on steps 1021 to 1023, it is as follows:

[0149] Step 1021: Determine abnormal data and normal data. Through comparative analysis, it can be determined that the myocardial enzyme level of patient F is abnormal data.

[0150] Step 1022: Repair abnormal data. A data repair model generated by training normal data can be used to predict and repair this abnormal value. For example, the following two key steps:

[0151] Data repair model training: Assume that there is a certain relationship between the myocardial enzyme level in normal data and the degree of chest pain and heart rate. A linear regression model can be used to fit this relationship. Assume that the model is: myocardial enzyme level = a*chest pain degree + b*heart rate + c. Use normal data to train this model and obtain the values ​​of parameters a, b, and c.

[0152] Abnormal data repair: Substitute the chest pain level (9) and heart rate (120) of patient F into the trained model. The predicted myocardial enzyme level by the model is: Myocardial enzyme level = a * 9 + b * 120 + c = 200 U / L (assuming the predicted value is 200 U / L).

[0153] Step 1023: Form the second data set. Combine the repaired data of patient F with other normal data to form the second data set. As shown in the following table:

[0154] Patient ID: F008;

[0155] Chest pain level (1 - 10 points): 9 points;

[0156] Heart rate (beats / minute): 120 beats / minute;

[0157] Myocardial enzyme level (U / L): 200 U / L.

[0158] It should be understood that in actual application scenarios, an appropriate repair model should be selected according to the characteristics of the data, and the performance of the model should be verified through methods such as cross - validation. If the repaired data still has abnormalities, the repair process may need to be iterated multiple times until the data quality meets the requirements. In some cases, domain experts may need to review the abnormal data and repair results to ensure the accuracy and reasonableness of the data.

[0159] Through these specific numerical examples and detailed steps, the process and importance of abnormal data analysis and repair can be more clearly understood.

[0160] Step 103: Extract features from the second data set and divide the second data set into several third data sets according to the results of feature extraction.

[0161] This step 103 is a key step in the data processing process. Feature extraction is the process of extracting attributes or variables from data that are helpful for describing, classifying, or predicting the data. In the field of healthcare, features may include age, gender, blood pressure, blood glucose level, genetic markers, etc. It can also be mathematical processing based on the above - mentioned features, such as mean, variance, standard deviation, and quartiles, etc. These mathematical processes help to extract more meaningful features from the original data, thereby enhancing the prediction ability and analysis accuracy of the model. Data set division is based on the results of feature extraction, and the data set can be divided into multiple subsets, which may be based on certain specific feature values or feature combinations. The purpose of dividing the data set is to better understand the structure of the data, simplify the analysis process, or prepare specific data subsets for different analysis tasks.

[0162] Optionally, refer to Figure 3 It can be seen that step 103 specifically includes:

[0163] Step 1031: Calculate the characteristics of the data corresponding to the target parameter type in each group of historical diagnostic information in the second dataset. The characteristics include central tendency and / or dispersion.

[0164] This step 1031 is mainly to conduct a more in-depth analysis of the data characteristics of the repaired second dataset.

[0165] 1. Target parameter type: This refers to analyzing a specific category of diagnostic information (such as physical symptom data, pathological test data, or treatment plan data) in the second dataset. For example, in physical symptom data, the parameter "blood pressure" may be of concern; in pathological test data, the "blood glucose level" etc. may be of concern.

[0166] 2. Feature calculation:

[0167] Central tendency: Describes the degree to which data converges towards the central value. Common measures include the mean (average), median, etc.

[0168] Dispersion: Describes the dispersion or variation of data. Common measures include variance, standard deviation, range (the difference between the maximum and minimum values), etc.

[0169] 3. Purpose:

[0170] By calculating these characteristics, we can better understand the data distribution, identify potential outliers or trends, and provide a basis for subsequent data analysis and model construction.

[0171] Here, based on step 1031, a specific example is given: Suppose the second dataset contains the following patient information: age, gender, systolic blood pressure, diastolic blood pressure, and blood glucose level.

[0172] If the analysis is conducted for the parameter type of "systolic blood pressure":

[0173] 1. Central tendency characteristics: Calculate the average (mean) of the systolic blood pressure of all patients. For example, the average systolic blood pressure is obtained as 130 mmHg. Calculate the median, assumed to be 128 mmHg (the value in the middle after sorting all systolic blood pressure values from smallest to largest).

[0174] 2. Dispersion characteristics: Calculate the variance, which reflects the degree of deviation of the systolic blood pressure values from their mean. Calculate the standard deviation, assumed to be 15 mmHg, which means that most patients' systolic blood pressure values fluctuate within the range of 15 mmHg above and below the mean. Calculate the range, that is, the difference between the maximum and minimum systolic blood pressures. Assume the maximum value is 160 mmHg and the minimum value is 100 mmHg, then the range is 60 mmHg.

[0175] If the analysis is conducted for "blood glucose level":

[0176] Central tendency: The average blood glucose level may be 95 mg / dL, and the median may be 93 mg / dL.

[0177] Dispersion: The variance may reflect the fluctuations in blood glucose levels. The standard deviation may be 10 mg / dL, indicating that the blood glucose values of most patients vary within the range of 10 mg / dL above and below the average blood glucose. The range can give the highest and lowest limits of blood glucose.

[0178] Practical significance: These calculated characteristics can help doctors or researchers quickly understand the overall condition of the systolic blood pressure or blood glucose of the patient group and the degree of difference among individuals. When constructing a disease diagnosis model, these characteristics can serve as important input variables to help the model more accurately identify disease patterns or predict disease risks. At the same time, by comparing the differences between different characteristics, potential health problems or disease trends can also be discovered, providing a basis for further clinical decision-making.

[0179] In summary, step 1031 is to deeply analyze and understand the data by calculating the central tendency and dispersion of the characteristics of specific parameter types in the second dataset, laying a foundation for subsequent data processing and analysis.

[0180] Step 1032: Encode each group of historical diagnosis information according to the characteristics.

[0181] This step 1032 is to transform these characteristics into a coding form that can be processed by a machine learning model or other analysis tools after feature extraction and analysis of the historical diagnosis information in the second dataset.

[0182] 1. Purpose of encoding: Transform the extracted characteristics (such as central tendency and dispersion) into a standardized format for easy computer processing and analysis. Encoding can reduce the complexity of the data, improve the efficiency of data processing, and help the model better understand and utilize these characteristics.

[0183] 2. Methods of encoding: Methods such as binary encoding, one-hot encoding, and label encoding can be used. In this example, the central tendency and dispersion are transformed into binary representations and corresponding identification codes are assigned.

[0184] Here, a specific example is given based on step 1032: Suppose the characteristics of "systolic blood pressure" in the second dataset have been calculated:

[0185] Central tendency: The average systolic blood pressure is 130 mmHg, and the variance is 150.

[0186] Dispersion: The discrete value obtained through paired calculation is 20.

[0187] Coding process:

[0188] 1. Central tendency coding: Calculate the trend identification code for the average systolic blood pressure. Assume a threshold is set, such as 125 mmHg. Values above this are "high" and values below are "low". The average systolic blood pressure is 130 mmHg > 125 mmHg, so the trend identification code "high" is assigned, with a binary representation of 1.

[0189] 2. Dispersion coding:

[0190] Calculate the dispersion identification code for the degree of dispersion. Assume a threshold is set, such as 15. Values above this are "high dispersion" and values below are "low dispersion". The dispersion value is 20 > 15, so the dispersion identification code "high dispersion" is assigned, with a binary representation of 1.

[0191] Coding result: 11.

[0192] Applied to historical diagnostic information: Assume there are the following patient records:

[0193] The systolic blood pressure of Patient 1 is 140, that of Patient 2 is 110, and that of Patient 3 is 135.

[0194] Patient 1:

[0195] The systolic blood pressure is 140 mmHg > 125 mmHg, and the central tendency coding is 1.

[0196] According to specific calculations (assuming the dispersion value is also above the threshold), the dispersion coding may be 1.

[0197] Then the coded data is 11.

[0198] Patient 2:

[0199] The systolic blood pressure is 110 mmHg < 125 mmHg, and the central tendency coding is 0.

[0200] According to specific calculations (assuming the dispersion value is below the threshold), the dispersion coding may be 0.

[0201] Then the coded data is 00.

[0202] Patient 3:

[0203] The systolic blood pressure is 135 mmHg > 125 mmHg, and the central tendency coding is 1.

[0204] According to specific calculations (assuming the dispersion value is above the threshold), the dispersion coding may be 1.

[0205] Then the coded data is 11.

[0206] It should be understood that the specific coding method and threshold setting need to be determined according to the actual data and business requirements. During the coding process, the consistency and interpretability of the coding should be ensured to facilitate subsequent data analysis and model interpretation. The coded data should be saved together with the original data for verification and traceability.

[0207] Through the above steps, the features in the historical diagnostic information can be transformed into a standardized coding form, facilitating subsequent data analysis and model construction.

[0208] Step 1033: Divide the second data set into several third data sets according to the codes corresponding to each group of historical diagnostic information; each third data set corresponds to one code.

[0209] This step 1033 is a process of further subdividing the data according to these codes after coding the historical diagnostic information.

[0210] 1. Purpose: By coding to group the data, data subsets with similar features or patterns can be more clearly identified. This helps to perform more refined analysis or modeling for data subsets with different features, improving the accuracy and generalization ability of the model.

[0211] 2. Operation method: According to the coding results of each historical diagnostic information, divide the data into the corresponding third data sets. Each third data set contains data records with the same coding features.

[0212] Here, based on step 1033, a specific example is given: Continuing with the previous example, there is already a set of coding results for the systolic blood pressure data of patients:

[0213] Patient 1, systolic blood pressure is 140, central tendency code is 1, dispersion degree code is 1;

[0214] Patient 2, systolic blood pressure is 110, central tendency code is 0, dispersion degree code is 0;

[0215] Patient 3, systolic blood pressure is 135, central tendency code is 1, dispersion degree code is 1;

[0216] Patient 4, systolic blood pressure is 120, central tendency code is 0, dispersion degree code is 0;

[0217] Patient 5, systolic blood pressure is 145, central tendency code is 1, dispersion degree code is 1;

[0218] Based on the above coding results, divide the third data sets:

[0219] 1. Third data set 1 (central tendency code is 1, dispersion degree code is 1):

[0220] Contains records with patient IDs 1, 3, and 5. The systolic blood pressure data of these patients shows a high degree of central tendency and high dispersion.

[0221] 2. The third dataset 2 (central tendency coded as 0, dispersion coded as 0):

[0222] Contains records with patient IDs 2 and 4. The systolic blood pressure data of these patients shows a low degree of central tendency and low dispersion.

[0223] Practical significance: Through this partitioning method, researchers can conduct in-depth analysis on data subsets with different characteristics, such as exploring the reasons or influencing factors behind systolic blood pressure data with high central tendency and high dispersion. In machine learning modeling, these third datasets can be used as training sets or validation sets respectively to improve the pertinence and prediction performance of the model. In addition, this grouping method helps to discover potential data patterns or anomalies, providing support for clinical decision-making.

[0224] In summary, step 1033 is a process of further subdividing the second dataset into several third datasets with similar characteristics according to the coding results of historical diagnosis information, which helps to improve the fineness and depth of data analysis.

[0225] Step 104: Train the initial training model based on the third data to obtain a disease diagnosis model.

[0226] The purpose of this step 104 is to use the processed and partitioned third dataset to optimize the initial training model so that it can accurately diagnose diseases. The third dataset is the data after screening, sorting, and feature engineering in the previous steps, with better data quality and pertinence.

[0227] Optionally, step 104 specifically includes:

[0228] Step 1041: Use several third datasets as training data to train several initial training models respectively.

[0229] This step 1041 mainly conducts model training based on the third datasets with different characteristics divided by operations such as calculation, extraction, and coding of data features in the previous step 103. In step 103, feature extraction (such as calculating features such as central tendency and dispersion of systolic blood pressure) was performed on the second dataset and encoded, and then the data was divided into different third datasets according to the encoding. Each of these third datasets has a specific combination of features. In step 1041, these third datasets with different feature combinations are used to train the initial training models respectively. The advantage of doing this is that each initial model can focus on learning the data patterns under specific feature combinations, thus more accurately capturing the cardiovascular disease diagnosis patterns related to these features.

[0230] Since different third data sets have different characteristics, the trained initial models will also vary. For example, one third data set may consist of patient data with high and fluctuating systolic blood pressure (high degree of dispersion). The initial model trained for this data set will focus on learning the relationship between such blood pressure characteristics and cardiovascular diseases. Another third data set may be composed of patient data with older age and high blood sugar levels. The corresponding initial model will then pay attention to the role of age and blood sugar factors in the diagnosis of cardiovascular diseases.

[0231] Here, a specific example is given based on step 1041: Suppose a cardiovascular disease diagnosis model is being built, and in step 103, the calculations and encodings of characteristics such as patient age, blood pressure (systolic and diastolic), and blood sugar level have been obtained, and the following third data sets are divided:

[0232] Third data set 1

[0233] Feature combination: Age between 40 - 50 years old, normal systolic blood pressure (central tendency within the normal range and low degree of dispersion), normal blood sugar.

[0234] Training model 1: Use the data of third data set 1 to train an initial logistic regression model. This model will learn the relationship pattern between the risk of cardiovascular diseases and other factors (such as lifestyle habits and other variables that may exist in the data set) under this specific age range, normal blood pressure, and normal blood sugar conditions. For example, it may be found that among people with normal blood pressure and blood sugar in this age group, long-term smoking will increase the risk ratio of cardiovascular diseases.

[0235] Third data set 2

[0236] Feature combination: Age greater than 60 years old, high systolic blood pressure (high central tendency), high blood sugar.

[0237] Training model 2: Use the data of third data set 2 to train an initial decision tree model. This model will focus on learning the diagnosis pattern of cardiovascular diseases in the patient group with older age, high blood pressure, and high blood sugar. For example, it may be found that for such patients, if there is also dyslipidemia (also a possible factor in the data set) at the same time, the incidence risk of cardiovascular diseases will increase significantly.

[0238] Third data set 3

[0239] Feature combination: Female, normal systolic blood pressure but high diastolic blood pressure (relatively high degree of dispersion), no family history of cardiovascular diseases.

[0240] Training Model 3: Use the data of the third dataset 3 to train an initial model of a support vector machine. This model will explore the association patterns between cardiovascular diseases and other factors (such as hormone levels and other potentially relevant factors) in the case of normal systolic blood pressure but high diastolic blood pressure and no family history in women. For example, it may be found that in this case, the obesity index is a relatively key influencing factor.

[0241] By training the initial model separately for the third dataset with different feature combinations, multiple targeted models can be obtained. Subsequently, these models can be combined through methods such as ensemble learning to construct a more comprehensive and accurate cardiovascular disease diagnosis model.

[0242] Step 1042: Obtain the model parameters of each trained initial training model and perform a fusion process on the model parameters.

[0243] In this step 1042, multiple initial training models are trained using different third datasets. Each model has learned the data patterns under specific feature combinations. The purpose of step 1042 is to integrate the "knowledge" of these models and create a more powerful and comprehensive disease diagnosis model by fusing their model parameters. Model parameters are the "weights" or "rules" learned by the model during the learning process, which determine how the model makes predictions based on input features. Different initial models will have different learned model parameters due to the different characteristics of their training data. By fusing the parameters of multiple models, the advantages of each model can be combined, improving the prediction accuracy and generalization ability of the overall model. The fusion process can also reduce the bias and variance of a single model, making the final diagnosis model more robust.

[0244] Here, based on step 1042, a specific example is given: Suppose in the previous steps, three initial training models are trained: Model A, Model B, and Model C.

[0245] Model parameter example:

[0246] 1. Model A (trained based on the third dataset 1):

[0247] Suppose Model A is a logistic regression model, and its model parameters include the weights of each feature, such as the age weight is 0.5, the systolic blood pressure weight is 0.3, the blood sugar weight is 0.2, etc.

[0248] 2. Model B (trained based on the third dataset 2):

[0249] Suppose Model B is a decision tree model, and its model parameters can be represented as the splitting thresholds of each node, such as the age splitting threshold is 55 years old, the systolic blood pressure splitting threshold is 140 mmHg, etc.

[0250] 3. Model C (trained based on the third dataset 3):

[0251] Assume that Model C is a support vector machine model, and its model parameters include the coefficients of support vectors and bias terms, etc.

[0252] Example of fusion processing:

[0253] 1. Weight averaging method:

[0254] For the logistic regression model A and other linearly additive models, the weights of each feature can be directly averaged. For example, if the age weight of another model D is 0.4, then the fused age weight is (0.5 + 0.4) / 2 = 0.45.

[0255] 2. Voting method or stacking method:

[0256] For classification models such as the decision tree model B, the voting method or stacking method can be used for fusion. For example, when predicting whether a certain patient has cardiovascular disease, models A, B, and C respectively give prediction results, and the final result is determined by majority voting through the voting method.

[0257] 3. Model integration:

[0258] More complex fusion methods include model integration techniques such as Boosting or Bagging. These methods can optimize the overall performance by adjusting the contribution degrees of each model.

[0259] It should be understood that the selection of the fusion processing method should be determined according to the specific task and model type. During the fusion process, it is necessary to ensure that the data preprocessing and feature engineering steps of each model are consistent to avoid introducing additional biases. The fused model needs to be fully verified and evaluated to ensure that its performance is better than that of a single initial model.

[0260] Through the fusion processing in step 1042, the advantages of multiple initial models can be fully utilized to construct a more powerful and accurate disease diagnosis model.

[0261] Step 1043, construct a disease diagnosis model based on the fused model parameters; the architecture of the disease diagnosis model is consistent with that of the initial training model.

[0262] This step 1043 is the final crucial step in constructing a disease diagnosis model. The aim is to build a comprehensive disease diagnosis model by fusing the parameters of multiple initial training models. The architecture consistent with the initial training models is maintained to ensure the interpretability and compatibility of the new model. - The initial training models may have adopted specific algorithms and structures (such as logistic regression, decision tree, support vector machine, etc.). Based on these algorithms and structures, the fused model improves its performance by updating the parameters rather than changing the basic architecture of the model. The parameters of the fused model contain the knowledge obtained by each initial model during the learning process. These parameters act together on the new model, enabling it to capture the patterns and features in the data more comprehensively.

[0263] Here, a specific example is given based on step 1043: Suppose three initial models have been trained before: Model A (logistic regression), Model B (decision tree), and Model C (support vector machine), and parameter fusion has been performed through step 1042.

[0264] Constructing a disease diagnosis model:

[0265] 1. Architecture of the logistic regression model A:

[0266] Input layer: Features such as the patient's age, blood pressure, blood sugar, etc.

[0267] Output layer: The risk probability of cardiovascular disease.

[0268] Parameters: Weights of each feature (such as age weight 0.5, systolic blood pressure weight 0.3, etc.).

[0269] 2. Architecture of the decision tree model B:

[0270] Input layer: The same patient features.

[0271] Output layer: Classification result of cardiovascular disease (diseased or not diseased).

[0272] Parameters: Splitting thresholds of each node (such as age splitting threshold 55 years old, systolic blood pressure splitting threshold 140 mmHg, etc.).

[0273] 3. Architecture of the support vector machine model C:

[0274] Input layer: The same patient features.

[0275] Output layer: Classification result of cardiovascular disease.

[0276] Parameters: Coefficients and bias terms of the support vectors.

[0277] Construction of the fused model:

[0278] Parameter fusion: Assume that through a certain fusion method (such as weighted averaging), new feature weights are obtained: age weight 0.45, systolic blood pressure weight 0.35, etc. For the parameters of decision trees and support vector machines, corresponding fusion processing is also carried out.

[0279] Construct a new model: The new disease diagnosis model adopts the same architecture as logistic regression model A (because logistic regression models have good interpretability and linear feature processing capabilities). Apply the fused parameters to the corresponding positions of the new model.

[0280] In this way, when the new model inputs patient features, it will calculate according to the fused parameters and output the risk probability of cardiovascular disease or the diagnosis and treatment plan.

[0281] For example, when inputting new patient data, such as age 60, systolic blood pressure 150 mmHg, and blood sugar 110 mg / dL, the new disease diagnosis model will calculate the risk probability of cardiovascular disease for this patient according to the fused parameters. For example, the model may output that the risk probability of cardiovascular disease for this patient is 70%, which means that this patient has a relatively high risk of getting the disease and requires further examinations and interventions.

[0282] It should be understood that the method of fusing parameters should ensure that the performance of the new model is better than that of a single initial model. When constructing a new model, sufficient verification and evaluation are required to ensure its accuracy and reliability. Maintaining the consistency of the model architecture helps with the interpretability of the new model and its compatibility with other systems.

[0283] Through step 1043, a disease diagnosis model with better performance and consistent structure can be constructed using the fused model parameters, providing strong support for clinical decision-making.

[0284] Optionally, step 104 can also specifically include:

[0285] Step 1044: Merge all the third data sets into an overall data set to train an initial training model.

[0286] This step 1044 is another method for constructing a disease diagnosis model, which is different from the way of training multiple initial models separately using different third datasets and then fusing them in step 1041. In the previous steps, the third datasets were divided according to different feature combinations. Merging them into an overall dataset can provide a more comprehensive data view. Doing so helps to discover potential relationships between different features and avoid global pattern information that may be lost due to data segmentation. Using this merged overall dataset to train an initial training model, this model will attempt to learn the patterns in the data as a whole to diagnose diseases. This initial model may comprehensively consider the complex interactions between multiple factors, rather than learning separately for specific feature subsets as in step 1041.

[0287] Here, a specific example is given based on step 1044: Suppose still taking the diagnosis of cardiovascular diseases as an example, and several third datasets have been obtained in the previous steps:

[0288] Situation of the third datasets:

[0289] 1. Third dataset 1:

[0290] Features: Age between 40 - 50 years old, normal systolic blood pressure (central tendency within the normal range and low degree of dispersion), normal blood sugar. Patient data containing such features.

[0291] 2. Third dataset 2:

[0292] Features: Age greater than 60 years old, high systolic blood pressure (high central tendency), high blood sugar. Patient data containing such features.

[0293] 3. Third dataset 3:

[0294] Features: Female, normal systolic blood pressure but high diastolic blood pressure (relatively high degree of dispersion), no family history of cardiovascular diseases. Patient data containing such features.

[0295] Merging and training process:

[0296] 1. Data merging: Merge the third dataset 1, the third dataset 2, and the third dataset 3 into an overall dataset. This overall dataset contains patient data under various factor combinations such as different age ranges, blood pressure and blood sugar conditions, and gender.

[0297] 2. Initial training model training: Select a machine learning algorithm to build an initial training model, such as the random forest algorithm. Input features such as patient age, blood pressure (systolic and diastolic), blood glucose level, and gender in the merged overall dataset into the random forest model for training. The model will learn the relationships between these features and the occurrence of cardiovascular diseases in this comprehensive dataset. For example, it may find that although age and blood pressure are individually related to cardiovascular diseases, when considering gender and blood glucose level simultaneously, there are some specific interaction relationships. For instance, in a specific age range, women with high blood glucose may have a higher risk of cardiovascular diseases even if their blood pressure is normal.

[0298] Compared with separately training multiple initial models in step 1041 and then fusing them, this method pays more attention to mining feature relationships from the overall data, while the former focuses more on constructing models for different feature subsets respectively and then integrating their advantages. Both methods have their own advantages and disadvantages, and can be selected according to specific research purposes and data characteristics.

[0299] Step 1045: Use the trained initial training model as the disease diagnosis model.

[0300] This step 1045 is the last step in the entire model construction process, marking the transition from data preparation to model application. Before this, through a series of steps (including data preprocessing, feature extraction, model training, etc.), the training of the initial training model has been completed. This initial training model has learned the patterns and features in the data and can make predictions on new input data. Designating the trained initial model as the disease diagnosis model officially means that it can be used for actual clinical or research scenarios to diagnose diseases. This model represents the final result of the entire data processing and model construction process.

[0301] Here, based on step 1043, a specific example is given: Suppose a model for diagnosing hypertension has been under construction and all necessary steps have been completed, including data collection, preprocessing, feature extraction, model training, etc.

[0302] Review of the model training process:

[0303] 1. Data preparation: A dataset containing information such as patient age, blood pressure (systolic and diastolic), blood glucose level, and family medical history was collected. Preprocessing operations such as data cleaning, feature extraction, and encoding were performed on the data.

[0304] 2. Model training: Multiple initial training models (such as logistic regression, decision tree, support vector machine, etc.) were trained using different third datasets separately. Or an initial training model (such as random forest) was trained by merging all third datasets into an overall dataset.

[0305] Determined as a disease diagnosis model:

[0306] After verification and evaluation, it is found that one of the initial training models (such as a random forest-based model) performs best on the test set, with high accuracy and generalization ability.

[0307] Therefore, it is decided to officially use this trained random forest model as the diagnosis model for hypertension disease.

[0308] Application example: When a new patient comes to the clinic, the doctor collects information such as the patient's age (50 years old), systolic blood pressure (140 mmHg), diastolic blood pressure (90 mmHg), blood glucose level (100 mg / dL), and family history (no family history of cardiovascular disease). These information are passed as input features to the random forest model that has been determined as the diagnosis model for hypertension disease. The model analyzes based on the patterns and features learned before and outputs a prediction result, such as the patient has a 70% probability of having hypertension disease.

[0309] It should be understood that before determining the initial training model as a disease diagnosis model, sufficient verification and evaluation must be carried out to ensure its accuracy and reliability in actual applications. The model should be updated and maintained regularly to adapt to new data and clinical needs. When using the model for diagnosis, doctors still need to make a comprehensive judgment by combining professional knowledge and clinical experience.

[0310] Through step 1045, the trained initial model is officially applied to disease diagnosis, providing strong support for clinical decision-making.

[0311] In summary, in this embodiment, by collecting the basic information data, physical symptom data, pathological test data, and treatment plan data of historical patients, abnormal analysis is performed on these data to ensure the integrity and availability of the data, thereby improving the overall value of the data. During the data processing process, the system uses data fusion technology to integrate the same type of data, effectively reducing the data volume and accelerating the speed of subsequent data processing and analysis.

[0312] Embodiment 2

[0313] Corresponding to the foregoing embodiment of the method for constructing a diagnosis model, the present disclosure also provides an embodiment of a system for constructing a disease diagnosis model.

[0314] Figure 6 It is a schematic diagram of the modules of a system for constructing a disease diagnosis model provided by an exemplary embodiment of the present disclosure. The system includes:

[0315] The present disclosure provides a system for constructing a disease diagnosis model. The construction system includes:

[0316] A data acquisition module 21, configured to acquire a plurality of historical diagnostic information as a first data set; the historical diagnostic information includes at least one set of historical diagnostic information, and each set of historical diagnostic information includes the following parameter types: physical symptom data, pathological test data, and treatment plan data;

[0317] A data analysis module 22, configured to perform anomaly analysis on the first data set, and repair the abnormal data in the first data set to obtain a second data set;

[0318] A feature extraction module 23, configured to extract features from the second data set, and divide the second data set into a plurality of third data sets according to the results of the feature extraction;

[0319] A model training module 24, configured to train an initial training model based on the third data to obtain a disease diagnosis model.

[0320] Optionally, the data analysis module 22 is specifically configured to:

[0321] Determine the abnormal data and normal data in the first data set;

[0322] Input the abnormal data into a data repair model to obtain the repaired abnormal data; wherein, the data repair model is generated by training with the normal data;

[0323] Use the repaired abnormal data and the normal data as the second data set.

[0324] Optionally, the construction system further includes:

[0325] A division module, configured to divide the normal data into a plurality of subsets;

[0326] A training module, configured to sequentially use one of the subsets as a validation set, and the subsets other than the validation set as a training set, and use the training set and the validation set to train and validate the initial data repair model respectively;

[0327] A fitting module, configured to determine the fitting effect of the initial data repair model according to the verification results, and adjust the parameters of the initial data repair model according to the fitting effect until the fitting effect reaches a preset threshold.

[0328] Optionally, the feature extraction module 23 is specifically configured to:

[0329] Calculate the features of the data corresponding to the target parameter type in each group of historical diagnostic information in the second data set; the features include central tendency and / or dispersion degree;

[0330] Encode each group of historical diagnostic information according to the features;

[0331] According to the codes corresponding to each group of historical diagnosis information, divide the second data set into several third data sets; each third data set corresponds to one code.

[0332] Optionally, the model training module 24 is specifically configured to:

[0333] Use several third data sets as training data to train several initial training models respectively;

[0334] Obtain the model parameters of each trained initial training model, and perform fusion processing on the model parameters;

[0335] Construct a disease diagnosis model based on the fused model parameters; the architecture of the disease diagnosis model is the same as that of the initial training model.

[0336] Optionally, the model training module 24 is specifically configured to:

[0337] Merge all third data sets into an overall data set to train an initial training model;

[0338] Use the trained initial training model as the disease diagnosis model.

[0339] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only illustrative, where the units described as separate components may or may not be physically separated, and the components as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0340] In summary, in this embodiment, by collecting the basic information data, physical symptom data, pathological test data, and treatment plan data of historical patients, abnormal analysis is performed on these data to ensure the integrity and availability of the data, thereby improving the overall value of the data. In the data processing process, the system uses data fusion technology to integrate the same type of data, effectively reducing the data volume and accelerating the speed of subsequent data processing and analysis.

[0341] Embodiment 3

[0342] Figure 7 The structure diagram of an electronic device shown in an exemplary embodiment of the present disclosure, the electronic device includes a memory, a processor, and a computer program stored on the memory and for running on the processor, and when the processor executes the computer program, it implements the diagnostic model construction method described in any of the above embodiments. Figure 7The illustrated electronic device 90 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0343] As Figure 7 shown, the electronic device 90 may be presented in the form of a general-purpose computing device. For example, it may be a server device. The components of the electronic device 90 may include, but are not limited to: at least one of the above-mentioned processors 91, at least one of the above-mentioned memories 92, and a bus 93 that connects different system components (including the memory 92 and the processor 91).

[0344] The bus 93 includes a data bus, an address bus, and a control bus.

[0345] The memory 92 may include volatile memory, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.

[0346] The memory 92 may also include a program tool 925 (or utility) having a set of (at least one) program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0347] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92, such as the method for constructing a diagnostic model provided in any of the above embodiments.

[0348] The electronic device 90 may also communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication may be carried out through an input / output (I / O) interface 95. Moreover, the electronic device 90 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 96. As shown in the figure, the network adapter 96 communicates with other modules of the electronic device 90 through the bus 93. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0349] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0350] Embodiment 4

[0351] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for constructing the diagnostic model provided in any one of the above embodiments is implemented.

[0352] Among them, the more specific computer-readable storage medium that can be adopted may include but is not limited to: portable disk, hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical storage device, magnetic storage device or any suitable combination of the above.

[0353] Embodiment 5

[0354] The embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, the method for constructing the diagnostic model described in any one of the above is implemented.

[0355] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0356] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A method for constructing a disease diagnosis model, characterized in that: The construction method comprises: Acquire a number of historical diagnosis information as a first data set; the historical diagnosis information includes at least one group of the historical diagnosis information, each group of the historical diagnosis information includes the following parameter types: physical symptom data, pathological test data and treatment plan data; Performing anomaly analysis on the first data set, and repairing the abnormal data in the first data set to obtain a second data set; performing feature extraction on the second data set, and dividing the second data set into a plurality of third data sets according to the result of the feature extraction; The initial training model is trained based on the third data to obtain a disease diagnosis model.

2. The construction method according to claim 1, characterized in that: The performing anomaly analysis on the first data set and repairing the abnormal data in the first data set to obtain a second data set includes: Determining abnormal data and normal data in the first data set; Inputting the abnormal data into a data repair model to obtain the repaired abnormal data; wherein the data repair model is generated by training the normal data; The repaired abnormal data and normal data are used as the second data set.

3. The construction method according to claim 2, characterized in that: Before the step of inputting the abnormal data into a data repair model to obtain the repaired abnormal data, the method includes: Dividing the normal data into a plurality of subsets; One of the subsets is used as a validation set, and the subsets other than the validation set are used as training sets, and the initial data repair model is trained and validated using the training set and the validation set respectively; The fitting effect of the initial data repair model is determined according to the verification result, and the parameters of the initial data repair model are adjusted according to the fitting effect until the fitting effect reaches a preset threshold.

4. The construction method according to claim 1, characterized in that: The step of extracting features from the second data set and dividing the second data set into a plurality of third data sets according to the result of the feature extraction includes: Calculating the characteristics of the data corresponding to the target parameter type in each group of the historical diagnostic information in the second data set; the characteristics include central tendency and / or dispersion degree; encoding each group of the historical diagnostic information according to the characteristics; According to the codes corresponding to each group of the historical diagnostic information, the second data set is divided into a plurality of third data sets; each third data set corresponds to a code.

5. The construction method according to claim 1, characterized in that: The step of training the initial training model based on the third data to obtain a disease diagnosis model comprises: Using the plurality of third data sets as training data to train a plurality of initial training models; Obtaining model parameters of the initial training models after each training is completed, and fusing the model parameters; The disease diagnosis model is constructed based on the fused model parameters; the disease diagnosis model is consistent with the architecture of the initial training model.

6. The construction method according to claim 1, characterized in that: The step of training the initial training model based on the third data to obtain a disease diagnosis model further includes: Merging all the third data sets into an overall data set to train an initial training model; The trained initial training model is used as the disease diagnosis model.

7. A system for constructing a disease diagnosis model, characterized in that: The build system includes: A data acquisition module, used to acquire a number of historical diagnosis information as a first data set; the historical diagnosis information includes at least one group of the historical diagnosis information, each group of the historical diagnosis information includes the following parameter types: physical symptom data, pathological test data and treatment plan data; A data analysis module, configured to perform an abnormality analysis on the first data set, and repair abnormal data in the first data set to obtain a second data set; a feature extraction module, configured to extract features from the second data set, and divide the second data set into a plurality of third data sets according to the result of the feature extraction; A model training module is used to train an initial training model based on the third data to obtain a disease diagnosis model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the method for constructing a diagnosis model according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a diagnosis model according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing a diagnosis model according to any one of claims 1 to 6 is implemented.