Mining method and device of medical data, medium and equipment
Patent Information
- Application Number
- CN202211090935.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-09-07
AI Technical Summary
[0004]本公开的目的在于提供一种医疗数据的挖掘方法、医疗数据的挖掘装置、计算机可读存储介质以及电子设备,进而至少在一定程度上克服由于相关技术的限制和缺陷而导致的风险规则的精确度较低的问题
[0019] This disclosure discloses a method for mining medical data. On one hand, it acquires a medical data prediction model and original medical data features, and calculates the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. Then, it filters the original medical data features based on multiple feature contribution rates to obtain target medical data features and constructs a correlation dependency graph between the target medical data features. Next, it determines the feature association relationship between the target medical data features based on the correlation dependency graph and determines the feature threshold of the target medical data features based on the feature association relationship. Finally, it generates risk rules affecting diseases corresponding to the original medical data features based on the feature thresholds and the target medical data features. This solves the problem in the prior art where the accuracy of the obtained risk rules is low due to the inability to achieve cross-calculation of cross-associations between multiple features, thus improving the accuracy of the risk rules. On the other hand, it solves the problem in the prior art where only associations between features can be established, but feature thresholds cannot be obtained.
Smart Images

Figure CN115620914B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of machine learning technology, and more specifically, to a method for mining medical data, a device for mining medical data, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Existing self-interpretable models, such as logistic regression and tree-based models commonly used in medical research, often only yield the relationship between a single feature and the outcome, failing to perform cross-calculation of cross-correlation between multiple features, thus resulting in low accuracy of the risk rules obtained.
[0003] It should be noted that the information in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this disclosure is to provide a method for mining medical data, a device for mining medical data, a computer-readable storage medium, and an electronic device, thereby overcoming, at least to some extent, the problem of low accuracy of risk rules due to limitations and defects in related technologies.
[0005] According to one aspect of this disclosure, a method for mining medical data is provided, comprising: Obtain the medical data prediction model and the original medical data features, and calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model; The original medical data features are filtered based on the contribution rates of multiple features to obtain target medical data features, and a dependency graph corresponding to the target medical data features is constructed. The feature relationships between the target medical data features are determined based on the dependency graph, and the feature thresholds of the target medical data features are determined based on the feature relationships. Based on the feature threshold and the target medical data features, risk rules that affect the diseases corresponding to the original medical data features are generated.
[0006] In one exemplary embodiment of this disclosure, calculating the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model includes: Based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model.
[0007] In one exemplary embodiment of this disclosure, based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model, including: Step S10: Randomly select one feature from the original medical data features as the target object, and construct a feature set based on the other features in the original medical data features excluding the target object; Step S20: Construct multiple feature subsets based on the feature set, and calculate the first expected value of the first predicted value obtained by inputting the original medical data features into the medical data prediction model, and the second expected value of the second predicted value obtained by inputting the feature subsets into the medical data prediction model, based on the preset tree model interpreter. Step S30: Calculate the feature contribution rate of the target object during parameter adjustment in the medical data prediction model based on the number of features included in the feature subset, the number of feature subsets, the number of target objects, the first expected value, and the second expected value. Step S40: Repeat steps S10-S30 sequentially to obtain the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model.
[0008] In one exemplary embodiment of this disclosure, filtering the original medical data features based on the feature contribution rate to obtain target medical data features includes: The original medical data features are sorted according to the value of the feature contribution rate, and the original medical data features with a feature contribution rate greater than a preset threshold are selected from the sorted original medical data features as target medical data features.
[0009] In one exemplary embodiment of this disclosure, constructing a dependency graph corresponding to the target medical data features includes: Step S10': Select any feature from the target medical data features as the first feature object, select a feature from the original medical data features other than the first feature object as the second feature object, and obtain the first current contribution rate of the first feature object and the second current contribution rate of the second feature object. Step S20': Construct a coordinate system with the feature contribution rate as the first ordinate, the first current feature value of the first feature object as the abscissa, and the second current feature value of the second feature object as the second ordinate. Step S30': Determine a first coordinate pair based on the first current feature value and the first current contribution rate, and determine a second coordinate pair based on the second current feature value and the second current contribution rate; Step S40': Determine a first coordinate point based on the first coordinate position of the first coordinate pair in the coordinate system, and determine a second coordinate point based on the second coordinate position of the second coordinate pair in the coordinate system; Step S50': Based on the first coordinate point and the second coordinate point, generate a dependency graph between the first feature object and the second feature object, and repeat steps S10'-S40' to obtain a dependency graph corresponding to each feature in the target medical data features.
[0010] In one exemplary embodiment of this disclosure, determining the feature association relationships between the target medical data features based on the dependency graph includes: Obtain the target coordinates points that have overlapping relationships in the association dependency graph, and obtain the first target feature value, the second target feature value, the first target contribution rate corresponding to the first target feature, and the second target contribution rate corresponding to the second target feature corresponding to the second target feature; Linear fitting is performed on the first target feature value, the second target feature value, the first target contribution rate, and the second target contribution rate, and the feature correlation relationship between the target medical data features is determined based on the linear fitting result.
[0011] In one exemplary embodiment of this disclosure, determining the feature threshold of the target medical data feature based on the feature association includes: Based on the aforementioned feature relationships, the trend categories of the correlation dependency graph corresponding to the target medical data features are determined; wherein, the trend categories include an upward trend category and / or a downward trend category; Based on the trend category of the dependency graph, a data fitting trend is determined, and based on the data fitting trend, a feature threshold for the target medical data features included in the dependency graph is determined.
[0012] In one exemplary embodiment of this disclosure, the association dependency graph includes two different target medical data features, namely a first feature object and a second feature object; Wherein, when the trend category includes an upward trend category, the feature thresholds of the target medical data features included in the correlation dependency graph are determined based on the data fitting trend, including: The first target feature value of the first feature object and the second target feature value of the second feature object included in the association dependency graph are sorted, and the first maximum feature value and the first minimum feature value corresponding to the first target feature value, and the second maximum feature value and the second minimum feature value corresponding to the second target feature value are selected from the first target feature value and the second target feature value according to the sorting result. A first value range interval is constructed based on the first maximum eigenvalue, the first minimum eigenvalue, and other eigenvalues in the first target eigenvalues excluding the first maximum eigenvalue and the first minimum eigenvalue; and a second value range interval is constructed based on the second maximum eigenvalue, the second minimum eigenvalue, and other eigenvalues in the second target eigenvalues excluding the second maximum eigenvalue and the second minimum eigenvalue. Select a first set of points from the first value range interval that makes the sum of the first current contribution rates of the first feature object reach the maximum value, and select a second set of points from the second value range interval that makes the sum of the second current contribution rates of the second feature object reach the maximum value. The first minimum value is selected from the first set of points as the first minimum feature threshold of the first feature object, and the second minimum value is selected from the second set of points as the second minimum feature threshold of the second feature object.
[0013] In one exemplary embodiment of this disclosure, when the trend category includes a downward trend category, determining the feature threshold of the target medical data feature included in the correlation dependency graph based on the data fitting trend includes: The first maximum value is selected from the first set of points as the first maximum feature threshold of the first feature object, and the second maximum value is selected from the second set of points as the second maximum feature threshold of the second feature object.
[0014] In an exemplary embodiment of this disclosure, when the trend category includes an upward trend category and a downward trend category, determining the feature threshold of the target medical data feature included in the correlation dependency graph based on the data fitting trend includes: Select a first minimum value from the first set of points as the first minimum feature threshold of the first feature object, and select a second minimum value from the second set of points as the second minimum feature threshold of the second feature object; and Select the first maximum value from the first set of points as the first maximum feature threshold of the first feature object, and select the second maximum value from the second set of points as the second maximum feature threshold of the second feature object; A first threshold interval for the first feature object is constructed based on the first minimum feature threshold and the first maximum feature threshold, and a second threshold interval for the second feature object is constructed based on the second minimum feature value and the second maximum feature value.
[0015] According to one aspect of this disclosure, a method for processing medical data is provided, comprising: Acquire the medical data to be processed and extract the medical features to be processed included in the medical data to be processed; Obtain the risk rules associated with the diseases corresponding to the medical data to be processed; wherein the risk rules are generated based on any of the medical data mining methods described above; The medical characteristics to be processed are judged according to the risk rules to obtain the judgment result, which is used by medical personnel to obtain the diagnosis and treatment results of the disease based on the judgment result.
[0016] According to one aspect of this disclosure, a medical data mining apparatus is provided, comprising: The feature contribution rate calculation module is used to acquire the medical data prediction model and the original medical data features, and to calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. The dependency graph construction module is used to filter the original medical data features according to the contribution rates of multiple features to obtain target medical data features, and construct a dependency graph corresponding to the target medical data features. The feature threshold determination module is used to determine the feature association relationship between the target medical data features based on the association dependency graph, and to determine the feature threshold of the target medical data features based on the feature association relationship; The risk rule generation module is used to generate risk rules that affect diseases corresponding to the original medical data features based on the feature threshold and the target medical data features.
[0017] According to one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the medical data mining method and the medical data processing method described in any one of the preceding claims.
[0018] According to one aspect of this disclosure, an electronic device is provided, comprising: Processor; and Memory for storing the executable instructions of the processor; The processor is configured to execute the medical data mining method and the medical data processing method described above by executing the executable instructions.
[0019] This disclosure discloses a method for mining medical data. On one hand, it acquires a medical data prediction model and original medical data features, and calculates the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. Then, it filters the original medical data features based on multiple feature contribution rates to obtain target medical data features and constructs a correlation dependency graph between the target medical data features. Next, it determines the feature association relationship between the target medical data features based on the correlation dependency graph and determines the feature threshold of the target medical data features based on the feature association relationship. Finally, it generates risk rules affecting diseases corresponding to the original medical data features based on the feature thresholds and the target medical data features. This solves the problem in the prior art where the accuracy of the obtained risk rules is low due to the inability to achieve cross-calculation of cross-associations between multiple features, thus improving the accuracy of the risk rules. On the other hand, it solves the problem in the prior art where only associations between features can be established, but feature thresholds cannot be obtained.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0022] Figure 1 The flowchart illustrates a method for mining medical data according to an example embodiment of the present disclosure.
[0023] Figure 2 The diagram schematically illustrates a structural example of a gradient boosting tree (GBDT) according to an exemplary embodiment of the present disclosure.
[0024] Figure 3 An example diagram illustrating the feature contribution rate of a raw medical data feature according to an exemplary embodiment of the present disclosure is shown.
[0025] Figure 4 A schematic diagram illustrating the correlation dependency between age and neutrophil-lymphocyte ratio (NLR) according to an example embodiment of the present disclosure is shown.
[0026] Figure 5 This illustration schematically shows a correlation dependency graph between mammography (CC) and serum uric acid (UA) according to an example embodiment of the present disclosure.
[0027] Figure 6 This illustration schematically shows a correlation dependency graph between a brain natriuretic peptide (BNP) and mammography (CC) according to an example embodiment of the present disclosure.
[0028] Figure 7 This illustration schematically depicts a correlation dependency graph between fasting plasma glucose (FPG) and age according to an example embodiment of the present disclosure.
[0029] Figure 8 This illustration schematically depicts a correlation dependency graph between triglycerides (TG) and age according to an example embodiment of this disclosure.
[0030] Figure 9 A schematic diagram illustrating the correlation dependency between an aldosterone-renin ratio level (ARRL) and triglycerides (TG) according to an example embodiment of the present disclosure is shown.
[0031] Figure 10 A characteristic threshold map of age is schematically shown according to an example embodiment of the present disclosure.
[0032] Figure 11 The illustration schematically shows a characteristic threshold map of a mammogram (CC) according to an exemplary embodiment of the present disclosure.
[0033] Figure 12 A characteristic threshold map of a brain natriuretic peptide (BNP) according to an example embodiment of the present disclosure is schematically shown.
[0034] Figure 13 A characteristic threshold map of fasting plasma glucose (FPG) is schematically shown according to an example embodiment of the present disclosure.
[0035] Figure 14 A characteristic threshold map of a triglyceride (TG) according to an example embodiment of the present disclosure is schematically shown.
[0036] Figure 15 A characteristic threshold graph of an aldosterone-renin ratio (ARRL) level according to an example embodiment of the present disclosure is illustrated schematically.
[0037] Figure 16 The flowchart illustrates a method for processing medical data according to an example embodiment of the present disclosure.
[0038] Figure 17 The diagram schematically illustrates a block diagram of a medical data mining apparatus according to an exemplary embodiment of the present disclosure.
[0039] Figure 18 An electronic device for implementing the above-described method for mining and processing medical data, according to an example embodiment of the present disclosure, is illustrated. Detailed Implementation
[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0041] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0042] In some problems involving interpretability, the non-mathematical definition of interpretability can be understood as the degree to which humans can understand the reasons behind a decision. In machine learning research, highly interpretable methods can help humans understand the decision-making process and prediction results of machine learning models. Interpretable methods can be divided into two categories: self-interpretable methods and model-agnostic interpretable methods. In practical applications, self-interpretable models, such as tree models or linear models, allow humans to easily and intuitively understand the model's decision-making process. Model-agnostic interpretable methods, such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-Agnostic Explanations), can explain the model locally or globally without considering the model's own decision-making process.
[0043] In the context of medical risk rule mining within medical big data processing, the aforementioned model-interpretable scenarios are required. Medical risk rule mining falls under the category of medical data mining, generating specific medical risk rules by establishing predictive models driven by medical data. Simultaneously, in medical outcome prediction tasks, predictive models are built, and associations between key medical variables are generated using model-independent methods (such as SHAP). Based on these associations of medical features, they are visualized, and risk thresholds (cut-offs) for the variables are identified. This method further determines risk rules related to outcomes, thus achieving medical risk rule mining.
[0044] In some medical risk rule mining schemes, this can be achieved in the following two ways: One approach is to use supervised learning to build a self-interpretable model. Through this interpretability, the correlation between outcomes and predictors can be discovered, thus identifying risk factors. However, self-interpretable models, such as logistic regression and tree-based models commonly used in medical research, often only reveal the relationship between a single feature and the outcome. Cross-correlation between multiple features has greater value in understanding the impact of a single variable on the outcome.
[0045] Another approach is to use supervised learning, employing the model-independent interpretation method SHAP to discover the correlation between results and predictors, thereby identifying risk factors. However, in current medical research, many researchers use SHAP to interpret predictive models. While SHAP can establish correlations between features, it cannot obtain the feature's risk threshold (cut-off). In medical research, however, defining the feature's risk threshold provides a more intuitive way to define the range of risk factors. Therefore, how to obtain the feature's risk threshold (i.e., feature threshold) has become a pressing issue.
[0046] Based on this, this exemplary embodiment first provides a method for mining medical data, which can run on servers, server clusters, or cloud servers, etc. Of course, those skilled in the art can also run the method disclosed herein on other platforms as needed, and this exemplary embodiment does not impose any special limitations on this. Specifically, refer to... Figure 1 As shown, the method for mining this medical data may include the following steps: Step S110. Obtain the medical data prediction model and the original medical data features, and calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model; Step S120. Filter the original medical data features according to the contribution rates of multiple features to obtain target medical data features, and construct an association dependency graph corresponding to the target medical data features; Step S130. Determine the feature association relationships between the target medical data features based on the dependency graph, and determine the feature thresholds of the target medical data features based on the feature association relationships; Step S140. Generate risk rules that affect the diseases corresponding to the original medical data features based on the feature threshold and the target medical data features.
[0047] The aforementioned method for mining medical data, on the one hand, involves acquiring a medical data prediction model and original medical data features, and calculating the feature contribution rate of each medical data feature in the parameter adjustment process of the medical data prediction model. Then, the original medical data features are filtered based on multiple feature contribution rates to obtain target medical data features, and a dependency graph is constructed between these target medical data features. Next, the feature relationships between the target medical data features are determined based on the dependency graph, and feature thresholds for the target medical data features are determined based on these relationships. Finally, risk rules affecting diseases corresponding to the original medical data features are generated based on the feature thresholds and the target medical data features. This solves the problem in existing technologies where the accuracy of risk rules is low due to the inability to perform cross-calculation of cross-correlation between multiple features, thus improving the accuracy of risk rules. On the other hand, it also solves the problem in existing technologies where only relationships between features can be established, but feature thresholds cannot be obtained.
[0048] The following will provide a detailed explanation and description of the medical data mining method of the present disclosure in conjunction with the accompanying drawings.
[0049] First, the terms used in the example embodiments of this disclosure will be explained and described.
[0050] Real-world research refers to research data derived from real medical environments, reflecting actual diagnosis and treatment processes and patients' health status under real conditions.
[0051] Supervised learning vs. unsupervised learning: Whether something is supervised depends on whether the input data has labels. If the input data has labels, it is supervised learning; if it does not, it is unsupervised learning. Supervised learning refers to the process of adjusting the parameters of a classifier using a set of samples of known categories to achieve the required performance. It is also called supervised training or teacher-led learning. Unsupervised learning refers to the attempt to find hidden structures in unlabeled data.
[0052] Interpretable method: The non-mathematical definition of interpretability is "the degree to which humans can understand the reasons for a decision"; in the field of machine learning research, a highly interpretable method can help humans understand the decision-making process and prediction results of machine learning models.
[0053] Medical data mining is a step in the discovery of medical knowledge. It consists of specific data mining algorithms that, under certain acceptable computational efficiency constraints, generate specific medical patterns (such as medical risk rules).
[0054] SHAP (SHapley Additive exPlanations): Based on the optimal Shapley value in game theory, SHAP provides many global interpretation methods based on Shapley value aggregation. At the same time, SHAP interprets the model by calculating the contribution value of each feature to the outcome. For tree models, the TreeSHAP method can be used to discover the correlation between features.
[0055] Secondly, the inventive purpose of the exemplary embodiments of this disclosure will be explained and described. Specifically, the medical data mining method described in the exemplary embodiments of this disclosure can discover feature associations from medical prediction tasks based on the interpretable method SHAP, and find the risk thresholds of relevant risk factors through feature associations to form medical risk rule mining; the specific technical problems to be solved may include the following aspects: on the one hand, establishing a prediction model and using the model-independent interpretable method SHAP to interpret the prediction model, i.e., generating feature contributions; on the other hand, establishing feature association dependency graphs for features with high SHAP contribution values to find feature associations; furthermore, based on feature associations, using the optimal SHAP set principle to find the risk thresholds of features, thereby generating medical risk rules.
[0056] Furthermore, the medical data prediction model involved in the exemplary embodiments of this disclosure will be explained and described. Specifically, the medical data prediction model can be a Gradient Boosting Decision Tree (GBDT) model, or it can be an XGBoost (eXtreme Gradient Boosting) model; this example does not impose any special limitations on it. A specific structural example diagram of the Gradient Boosting Decision Tree model can be found in [reference needed]. Figure 2 As shown.
[0057] Specifically, in the GBDT model, each sample no longer strictly falls on a single branch, but rather on two or more branches with a certain probability. Each node in the GBDT model can be computed independently, thus enabling parallelization. In some example implementations, the GBDT model can consist of multiple base learners, where the base learners are decision trees. Local loss calculations are performed on the residuals of each tree based on the previous tree, and multiple trees are concatenated to form the GBDT model. Furthermore, since the computation of each tree is independent and can be parallelized, and the loss function of each tree is consistent with the GBDT calculation, a global loss function is ultimately used, thereby improving the final computational efficiency. During training, the entire model can be trained by minimizing the global loss function L. Although the result of each tree fits the residual of the previous tree, due to the independence between the nodes and trees of the decision tree, training only requires summing the local loss functions of all trees at the end, minimizing the sum of all local loss functions. This fully utilizes parallel computing to significantly improve the training speed of the model.
[0058] The medical data prediction model can be obtained by training the GBDT model using the features of the original medical data, which can be achieved in the following way: First, the original medical data features are input into the request type prediction model to be trained to obtain the medical prediction result. Specifically, this can be achieved as follows: First, the first linear regression component of the original medical data features in the internal nodes of the GBDT model is calculated using the GBDT model; second, the first linear regression component is normalized using the normalization layer containing the leaf nodes of the GBDT model to obtain the output value of the first linear regression component in the leaf node of the current decision tree in the GBDT model; finally, the medical prediction result is obtained based on the output value of the first linear regression component. Next, after obtaining the request type prediction result, a loss function needs to be constructed based on the request type prediction result and the feature labels of the historical access requests. Specifically, this can be achieved as follows: First, the local loss function of the current decision tree included in the GBDT model is calculated based on the medical prediction result and the feature labels of the original medical data features; second, the global loss function of the GBDT model is calculated based on the local loss function; finally, the model parameters included in the request type prediction model to be trained are adjusted based on the global loss function to obtain the medical data prediction model.
[0059] In some example embodiments, references Figure 2As shown, the first linear regression component (Logistic Regression) of the standard requested features can be calculated using the internal nodes of XGBoost. Specifically, this internal node can include multiple layers, such as layer zero, layer one, layer two, etc. The specific number of layers can be set according to actual needs; this example does not impose any special restrictions on this. Then, the first linear regression component is normalized using the normalization layer (Softmax) where the leaf node is located, obtaining the output value of the first linear regression component at the leaf node where the current decision tree is located. Finally, the medical prediction result is obtained based on this output value.
[0060] In some example embodiments, the local function can be implemented as follows: First, calculate the sum of the output values at the leaf nodes of all decision trees preceding the current decision tree; then, normalize the sum of the output values at the leaf nodes of all decision trees preceding the current decision tree to obtain the first prediction result of the standard request feature among all decision trees preceding the current decision tree; finally, construct the local loss function of the current decision tree based on the first prediction result, the output value of the first linear regression part at the leaf node of the current decision tree, and the feature label; wherein the local loss function also includes a regularization term, which can be used to represent the complexity of the current decision tree. Furthermore, after obtaining the local loss function, a global loss function can be constructed based on the local loss function and the regularization term.
[0061] In some example embodiments, adjusting the model parameters included in the GBDT model based on the global loss function to obtain a medical data prediction model can be achieved as follows: First, calculate the first gradient of the internal nodes of the GBDT model according to the global loss function, and then update the parameters of the internal nodes according to the first gradient; second, calculate the second gradient of the leaf nodes of the GBDT model according to the global loss function, and update the parameters of the leaf nodes according to the second gradient; finally, obtain the medical data prediction model based on the updated parameters of the internal nodes and the updated parameters of the leaf nodes. It should be noted that during the process of updating the parameters of the internal nodes and the leaf nodes, it is necessary to perform a second derivative of the global loss function, and then update the parameters of the internal nodes and the leaf nodes according to the gradient obtained from the second derivative. This method can improve the accuracy of the parameters of the internal nodes and the leaf nodes, thereby improving the accuracy of the obtained parameter-adjusted request type prediction model.
[0062] The following will combine Figure 2 right Figure 1 The steps involved in the medical data mining method shown are explained and illustrated in detail. Specifically: In step S110, the medical data prediction model and the original medical data features are obtained, and the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model.
[0063] In this example embodiment, firstly, a medical data prediction model and raw medical data features can be obtained. Here, the medical data prediction model refers to the model obtained by training the GBDT model using the raw medical data features. Specifically, taking the evaluation of the role of machine learning classification methods based on routine clinical biochemical test results in assessing early atrial function reduction in hypertension in a real-world study as an example, the raw medical data features involved in this example embodiment will be explained and illustrated. Specifically, the evaluation target can be the Left Atrial Stiffness Index (LASI), with LASI serving as the label (medical outcome) for this supervised learning task. A classification node of 0.29 is used, indicating LASI values of 0.29 or higher as positive samples and values less than 0.29 as negative samples. After feature selection, 20 variables (original medical data features) are used as the input matrix X of the prediction model, with LASI as the label y, to establish a classification prediction model for whether the left atrial stiffness index of hypertensive patients is too high. Furthermore, in practical applications, the gradient descent-based GBDT (Gradient Boosting Decision Tree) algorithm can be selected to build the prediction model, i.e., the medical data prediction model.
[0064] In one example embodiment, in the scenario of predicting the left atrial stiffness index, the raw medical data features involved may specifically include the following features: Age; CC: "cephalothorax" (also known as "axial" or "superior-inferior") view of mammography, where X-rays are projected from top to bottom during the examination; BNP: Brain Natriuretic Peptide; FPG: Fasting Plasma Glucose; TG: Triglyceride; ARRL: Aldosterone-Renin Ratio Level; BMI: Body Mass Index; HsCRP: Hypersensitive C-Reactive Protein; AldL: Aldosterone Level; NLR: Neutrophil-Lymphocyte Ratio; ApoA: Apolipoprotein-A; RDW_SD: Red Blood Cell Distribution Width. Width; HbAIc: Glycosylated Hemoglobin; eGFR: Estimated Glomerular Filtration Rate; LPa: Lipoprotein Converterase; ApoB: Apolipoprotein-B; UA: Blood Uric Acid; HT-duration: Hypertension-duration; LDLC: Low-Density Lipoprotein Cholesterol; RenL: Renin Level.
[0065] It should be further noted that the original medical data features involved in the exemplary embodiments of this disclosure may represent different specific features under different medical backgrounds and / or different medical scenarios. For example, in the scenario of predicting the left atrial stiffness index, the specific features included in the original medical data features are as shown above. As another example, in scenarios such as cancer cell metastasis or leukemia deterioration, other original medical data features are used, and this example does not impose any special restrictions on them. At the same time, the main application principle of the aforementioned medical data prediction model is: inputting the original medical data features into the medical data prediction model, and then obtaining the corresponding prediction results. The prediction results may be the prediction index of left atrial stiffness, the probability of cancer cell metastasis, the probability of leukemia deterioration or cure, etc., and this example does not impose any special restrictions on them.
[0066] Furthermore, once the medical data prediction model and the original medical data features are obtained, the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model can be calculated. In calculating the feature contribution rate, a preset tree model interpreter can be used to calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. It should be noted that other models can also be selected when choosing a medical data prediction model, such as convolutional neural network models, recurrent neural network models, deep neural network models, etc. This example does not impose any special restrictions on this. Of course, different types of medical data prediction models require different SHAP interpretation methods, and the corresponding SHAP interpretation method can be selected according to the specific category of the medical data prediction model. In this example embodiment, the medical data prediction model selected is the GBDT model; therefore, TreeSHAP is selected as the specific interpretation method.
[0067] In one example embodiment, the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model is calculated based on a preset tree model interpreter (TreeSHAP). This can be achieved as follows: Step S10, arbitrarily select a feature from the original medical data features as the target object, and construct a feature set based on the other features in the original medical data features excluding the target object; Step S20, construct multiple feature subsets based on the feature set, and calculate the input of the original medical data features into the medical data prediction model based on the preset tree model interpreter. Step S30: Calculate the feature contribution rate of the target object during parameter adjustment in the medical data prediction model based on the number of features included in the feature subset, the number of feature subsets, the number of target objects, the first expected value, and the second expected value; Step S40: Repeat steps S10-S30 sequentially to obtain the feature contribution rate of each medical data feature in the original medical data features during parameter adjustment in the medical data prediction model.
[0068] The following will explain in detail the specific calculation process of the feature contribution rate. Specifically, in application, firstly, the features of the original medical data are abstracted to obtain a specific feature vector x; where, Let k be the number of original medical data features (i.e., 20). For each original medical data feature, its feature contribution rate is calculated, which is the Shapley value of each original medical data feature calculated using TreeSHAP. Further, set... Let be the Shapley value of the i-th original medical data feature, and its calculation formula can be as follows: ; Where F is the set of subscripts for the features. ,but It is the set of indices after removing the i-th feature, which is S is All subsets, Let S be the number of elements in set S. for factorial; Where X is a vector of feature variables in the dataset, and it is k-dimensional. ; Let S be the set of feature variables extracted from the elements in set S. dimension, For the corresponding features in the original medical data The feature data; f is the medical data prediction model; This represents the medical prediction result obtained by the medical data prediction model when the feature variable data corresponding to S is input.
[0069] Meanwhile, the TreeShap model calculates... The algorithm is as follows: Input: a raw medical data feature x, a subset S of the i-th feature variable set, and parameters {v, a, b, t, r, d} of the medical data prediction model; where v is a q-dimensional vector, q is the number of nodes in the tree included in the medical data prediction model, containing the values of all tree nodes. If the node is a leaf node, it is assigned the output value of the leaf node; if the node is an internal node, it is assigned the value of internal. a is a vector containing the left node index of each internal node; b is a vector containing the right node index of each internal node; t is a vector containing the threshold in each internal node; d is a vector containing the indices of the feature variables used when splitting all internal nodes; r is a vector containing the samples that fall into the next subtree in each internal node; j∈{1,2,…,q}, define the function G(j): Check if vj is a leaf node; if it is a leaf node, output vj directly. If not, then check if dj exists in set S; if it exists in set S and Then output If it exists in set S and Then output ; If dj does not exist in set S, then output ; Calculate G(1), which is... .
[0070] It should be further explained here that, in step S20 above, the first predicted value refers to the specific disease prediction result obtained by inputting the original medical data features into the medical data prediction model, such as the prediction result of the left atrial stiffness index; the first expected value refers to the specific expectation of the first predicted value calculated by TreeSHAP, that is, the expected value of the first predicted value; the second predicted value refers to the specific disease prediction result obtained by inputting the feature subset other than the target object into the medical data prediction model, and the second expected value refers to the expected value of the second predicted value; at the same time, in the process of calculating the second expected value, when inputting the feature subset into the medical data prediction model, if there are multiple feature subsets corresponding to a certain target object, multiple feature subsets can be input into the medical data prediction model respectively to obtain the second expected value, or a feature subset can be selected to be input into the medical data prediction model. This example does not impose any special restrictions on this.
[0071] At this point, the calculation process for the feature contribution rate of the original medical data features has been completed; the specific values of the feature contribution rate for each original medical data feature can be found in [reference needed]. Figure 3 As shown. Among them, in Figure 3 In the graph, the horizontal axis represents the SHAP value (feature contribution rate), and the larger the value, the greater the impact on the outcome. The left vertical axis represents the feature ranking, and the right vertical axis represents the feature value (red indicates a higher feature value, and blue indicates a lower feature value). The graph also ranks the features on the outcome from highest to lowest.
[0072] In step S120, the original medical data features are filtered according to the feature contribution rate to obtain target medical data features, and a dependency graph corresponding to the target medical data features is constructed.
[0073] In this example embodiment, firstly, the original medical data features are filtered according to the feature contribution rate to obtain target medical data features. Specifically, this can be achieved as follows: the original medical data features are sorted according to the value of the feature contribution rate, and original medical data features with a feature contribution rate greater than a preset threshold are selected from the sorted original medical data features as target medical data features. Continuing with the example of the original medical data features involved in the prediction scenario of the left atrial stiffness index, the selected target medical data features may include: Age; CC: mammography; BNP: brain natriuretic peptide; FPG: fasting plasma glucose; TG: triglycerides; ARRL: aldosterone-renin ratio level. It should be noted that the target medical data features selected here are only for illustrative purposes and will not have any effect on diagnosis or treatment in actual medical scenarios.
[0074] Secondly, after obtaining the target medical data features, a dependency graph corresponding to the target medical data features can be constructed. Specifically, this can be achieved as follows: Step S10', arbitrarily select one feature from the target medical data features as the first feature object, select one feature from the original medical data features other than the first feature object as the second feature object, and obtain the first current contribution rate of the first feature object and the second current contribution rate of the second feature object; Step S20', construct a coordinate system with the feature contribution rate as the first vertical axis, the first current feature value of the first feature object as the horizontal axis, and the second current feature value of the second feature object as the second vertical axis; Step S30', according to the first... Step S10': Determine a first coordinate pair based on a current feature value and a first current contribution rate, and determine a second coordinate pair based on a second current feature value and a second current contribution rate; Step S40': Determine a first coordinate point based on the first coordinate position of the first coordinate pair in the coordinate system, and determine a second coordinate point based on the second coordinate position of the second coordinate pair in the coordinate system; Step S50': Generate a dependency graph between the first feature object and the second feature object based on the first coordinate point and the second coordinate point, and repeat steps S10'-S40' to obtain a dependency graph corresponding to each feature in the target medical data features.
[0075] In some example embodiments, the correlation dependency graph between the obtained target medical data features can be specifically referred to as Figure 4- Figure 9 As shown; where, Figure 4 A plot showing the association between age and the neutrophil-to-lymphocyte ratio (NLR); Figure 5 A correlation dependency plot between mammography (CC) and serum uric acid (UA); Figure 6 This is a correlation dependency plot between brain natriuretic peptide (BNP) and mammography (CC). Figure 7 This is a graph showing the association between fasting plasma glucose (FPG) and age. Figure 8 This is a dependency plot showing the relationship between triglycerides (TG) and age. Figure 9 This is a dependency graph showing the association between the aldosterone-renin ratio (ARRL) and triglycerides (TG). It should be noted that when drawing this dependency graph, one can draw dependency graphs between target medical data features, or between target medical data features and other non-target medical data features; this example does not impose any special restrictions on this. In some possible example embodiments, dependency graphs can also be drawn between all original medical data features; this example also does not impose any special restrictions on this.
[0076] It is important to further explain here that the purpose of drawing the dependency graph is to perform feature association analysis. That is, by performing feature association analysis on the features with high influence (target medical data features), we can find the joint features (feature interactions) that have a significant impact on the outcome. In the specific drawing process, the dependency plot function in SHAP can be used to draw the joint feature graph of each feature. For specific drawing details, please refer to the method described above, which will not be elaborated further here. At the same time, for each feature (which can be the target medical data feature or other features in the original medical data besides the target medical data feature), its corresponding associated features can be obtained, which have a significant impact on the outcome. The combinations that have a significant impact on the outcome (medical prediction result) can include, but are not limited to: age & NLR, CC & UA, BNP & CC, FPG & age, TG & age, ARRL & TG, etc.
[0077] In step S130, the feature association relationship between the target medical data features is determined according to the association dependency graph, and the feature threshold of the target medical data features is determined according to the feature association relationship.
[0078] In this example embodiment, firstly, the feature association relationships between target medical data features are determined based on the dependency graph. Specifically, this can be achieved as follows: First, target coordinate points with overlapping relationships are obtained in the dependency graph, and a first target feature value, a second target feature value, a first target contribution rate, and a second target contribution rate corresponding to each target coordinate point are obtained; secondly, linear fitting is performed on the first target feature value, the second target feature value, the first target contribution rate, and the second target contribution rate, and the feature association relationships between the target medical data features are determined based on the linear fitting result. It should be noted that in obtaining target coordinate points with overlapping relationships in the dependency graph, completely overlapping target coordinate points or target coordinate points within a certain error range can be selected; this example does not impose any special restrictions on this. Furthermore, the specific value range of the error range can be determined according to actual needs; this example does not impose any special restrictions on this.
[0079] It is worth noting that, in determining the target coordinates, the horizontal axis can be used as a reference, and all coordinates in the dependency graph can be traversed sequentially according to a window of a preset step size. Coordinates with overlapping relationships can then be used as target coordinates. Simultaneously, in calculating feature relationships, linear fitting can be performed on the first target feature value, the second target feature value, the first target contribution rate, and the second target contribution rate. The feature relationships between the target medical data features can then be obtained based on the linear fitting results. These feature relationships can include simultaneous increases, simultaneous decreases, or one feature increasing while another decreases, etc. This example does not impose any special restrictions on this.
[0080] Furthermore, once the feature association relationship is obtained, the feature threshold of the target medical data feature can be determined based on the feature association relationship. Specifically, this can be achieved as follows: First, obtain the trend category of the association dependency graph corresponding to the target medical data feature based on the feature association relationship; wherein, the trend category includes an upward trend category and / or a downward trend category; second, determine the data fitting trend based on the trend category of the association dependency graph, and determine the feature threshold of the target medical data feature included in the association dependency graph based on the data fitting trend; wherein, the association dependency graph includes two different target medical data features, the two different target medical data features being a first feature object and a second feature object.
[0081] In practical applications, when determining the feature thresholds of the target medical data features included in the correlation dependency graph based on data fitting trends, this can be achieved through the following methods: One approach is to determine the feature thresholds of the target medical data features included in the dependency graph based on the data fitting trend when the trend category includes an upward trend category. This can be achieved as follows: First, sort the first target feature values of the first feature objects and the second target feature values of the second target objects included in the dependency graph, and select the first maximum feature value, the first minimum feature value, the second maximum feature value, and the second minimum feature value from the first and second target feature values according to the sorting results; Second, select the first maximum feature value, the first minimum feature value, and the other feature values from the first target feature values excluding the first maximum and the first minimum feature values. A first value range is constructed based on the feature values, and a second value range is constructed based on the second maximum feature value, the second minimum feature value, and other feature values of the second target feature value excluding the second maximum and second minimum feature values. Then, a first set of points is selected from the first value range that makes the sum of the first current contribution rates of the first feature object reach its maximum value, and a second set of points is selected from the second value range that makes the sum of the second current contribution rates of the second feature object reach its maximum value. Finally, a first minimum value is selected from the first set of points as the first minimum feature threshold of the first feature object, and a second minimum value is selected from the second set of points as the second minimum feature threshold of the second feature object.
[0082] Another approach is to determine the feature thresholds of the target medical data features included in the dependency graph based on the data fitting trend when the trend category includes a downward trend category. This can be achieved as follows: First, sort the first target feature values of the first feature objects and the second target feature values of the second target objects included in the dependency graph, and select the first maximum feature value, the first minimum feature value, and the feature value corresponding to the second target feature value from the first target feature value and the second target feature value according to the sorting result; Second, construct the first maximum feature value, the first minimum feature value, and other feature values in the first target feature value excluding the first maximum feature value and the first minimum feature value. First, a value range interval is defined, and a second value range interval is constructed based on the second maximum eigenvalue, the second minimum eigenvalue, and other eigenvalues of the second target eigenvalue, excluding the second maximum eigenvalue and the second minimum eigenvalue. Then, a first set of points is selected from the first value range interval that makes the sum of the first current contribution rates of the first feature object reach its maximum value, and a second set of points is selected from the second value range interval that makes the sum of the second current contribution rates of the second feature object reach its maximum value. Finally, a first maximum value is selected from the first set of points as the first maximum feature threshold of the first feature object, and a second maximum value is selected from the second set of points as the second maximum feature threshold of the second feature object.
[0083] Another approach is to determine the feature thresholds of the target medical data features included in the dependency graph based on the data fitting trend when the trend categories include both upward and downward trends. This can be achieved as follows: First, sort the first target feature values of the first feature objects and the second target feature values of the second target objects included in the dependency graph. Then, based on the sorting results, select the first maximum feature value, the first minimum feature value, and the feature value corresponding to the second target feature value from the first and second target feature values. Second, construct a first value range interval based on the first maximum feature value, the first minimum feature value, and other feature values in the first target feature values excluding the first maximum and first minimum feature values. Construct a second value range interval based on the second maximum feature value, the second minimum feature value, and other feature values in the second target feature values excluding the second maximum and second minimum feature values. Then, from the... A first set of points is selected from the first value range interval, such that the sum of the first current contribution rates of the first feature object reaches its maximum value; and a second set of points is selected from the second value range interval, such that the sum of the second current contribution rates of the second feature object reaches its maximum value. Further, a first minimum value is selected from the first set of points as the first minimum feature threshold of the first feature object, and a second minimum value is selected from the second set of points as the second minimum feature threshold of the second feature object. A first maximum value is selected from the first set of points as the first maximum feature threshold of the first feature object, and a second maximum value is selected from the second set of points as the second maximum feature threshold of the second feature object. Finally, a first threshold range for the first feature object is constructed based on the first minimum feature threshold and the first maximum feature threshold, and a second threshold range for the second feature object is constructed based on the second minimum feature value and the second maximum feature value.
[0084] The following will further explain and illustrate the specific calculation process of the feature threshold of the target medical data. Specifically, the medical data mining method described in the example embodiment of this disclosure proposes an optimal SHAP set principle to find the feature risk threshold (feature threshold) for the feature association generated by SHAP, thereby generating medical risk rules; wherein, the calculation method of the feature threshold is shown in the following formula (1): ; Formula (1) Where k represents the data fitting trend, which can be obtained by linearly fitting the feature association data and the SHAP value; where k > 0 if the feature association dependency graph shows an upward trend (e.g., ...). Figure 4shown in age & NLR), if the feature association dependency graph shows a downward trend, then k≤0 (as in Figure 9 ARRL & TG). The risk threshold (feature threshold) c (which may include c1 and / or c2) is selected from the minimum value i=min(f) to the maximum value I=max(f) of the feature f. Since there are different orders of magnitude among different features (for example, the order of magnitude of age is 1~100, and the order of magnitude of BNP is 1~400), in order to ensure that the step size is sufficiently small, all features can be amplified by n times, where n is the maximum order of magnitude in the data set; of course, amplification treatment may not be performed, and this example does not impose special restrictions herewith; meanwhile, some features have associated trends with k>0 and k≤0, and two effective risk thresholds c1 & c2 can be obtained in this case. In practical applications, there are several situations as follows: (1) When k>0, when i reaches c(k>0), so that the SHAP value set corresponding to the portion of i>c(k>0 in this feature f reaches the maximum value, then the risk threshold c1=c(k>0) / n can be obtained (n is the amplification factor, and there is no need to divide by n here if no amplification treatment is performed), and the obtained risk rule is f>c1; that is, the minimum feature threshold can be c1; (2) When k≤0, when i reaches c(k≤0), so that the SHAP value set corresponding to the portion of i≤c(k≤0) in this feature f reaches the maximum value, then the risk threshold c2=c(k≤0) / n can be obtained (n is the amplification factor, and there is no need to divide by n here if no amplification treatment is performed), and the obtained risk rule is f<c2; that is, the maximum feature threshold can be c2; (3) When both k>0 and k≤0 exist, the minimum feature threshold c1 may be determined first, then the maximum feature threshold c2 is determined, and thus the risk rule c1<f<c2 is obtained.
[0085] So far, the specific determination process of the feature thresholds of the target medical data features has been fully completed. Wherein, for the specific feature thresholds of each target medical data feature, reference may be made to Figures 10-15 shown. Wherein, Figure 10 is a characteristic threshold diagram of age (Age); Figure 11 is a characteristic threshold diagram of mammography (CC); Figure 12 is a characteristic threshold diagram of brain natriuretic peptide (BNP); Figure 13 is a characteristic threshold diagram of fasting plasma glucose (FPG); Figure 14 is a characteristic threshold diagram of triglyceride (TG); Figure 15 is a characteristic threshold diagram of aldosterone-renin ratio level (ARRL).
[0086] In step S140, according to the feature thresholds and the target medical data features, a risk rule affecting the disease corresponding to the original medical data features is generated.
[0087] Specifically, once the feature thresholds for each target medical data feature are obtained, corresponding risk rules can be generated based on the feature thresholds and the target medical data features. For example, taking a real-world study, the risk thresholds (feature thresholds) for the aforementioned target medical data features can yield the following risk rules: age>52, CC>0.90, BNP>35.70, FPG>5.40, TG>1.10. Simultaneously, since the ARRL & TG feature correlation dependency graph exhibits both upward and downward trends, two risk thresholds for ARRL, c1=28.70 (maximum feature threshold) and c2=4.90 (minimum feature threshold), can be obtained. Therefore, the risk rule for ARRL can be ARRL<4.90 & ARRL>28.70.
[0088] In specific applications, taking age and NLR as an example, one of the medical risk rules that can be found is that age contributes the most to a higher left atrial stiffness index, that is, it has a greater effect on early atrial function reduction in hypertension. Furthermore, NLR has the greatest impact on the relationship between age and outcome, with the risk threshold for age being 52 years.
[0089] It should be further clarified here that the first maximum eigenvalue, second maximum eigenvalue, first minimum eigenvalue, and second minimum eigenvalue mentioned above refer to the first maximum target eigenvalue and first minimum target eigenvalue corresponding to the first target eigenvalue, and the second minimum target eigenvalue and second directly reachable target eigenvalue corresponding to the second target eigenvalue. At this point, the specific generation process of the risk rules has been completed. Based on the foregoing description, it can be understood that the medical data mining method described in the example embodiments of this disclosure has at least the following advantages: on the one hand, it can improve the interpretability of prediction models in real-world medical outcome prediction tasks, making it easier for medical workers to better understand the performance of prediction models; on the other hand, in real-world medical outcome prediction tasks, it can automatically discover medical risk rules related to outcomes, improving the research efficiency of researchers.
[0090] This disclosure also provides an example embodiment of a method for processing medical data. Specifically, refer to... Figure 16 As shown, the method for processing this medical data may include the following steps: Step S1610: Obtain the medical data to be processed and extract the medical features to be processed included in the medical data to be processed; Step S1620: Obtain the risk rules for the diseases corresponding to the medical data to be processed; wherein the risk rules are generated based on any of the above-described medical data mining methods; Step S1630: The medical characteristics to be processed are judged according to the risk rules to obtain the judgment result, which is used by medical personnel to obtain the diagnosis and treatment results of the disease based on the judgment result.
[0091] Figure 16 The illustrated method for processing medical data addresses two key issues. First, it involves obtaining a medical data prediction model and original medical data features, calculating the feature contribution rate of these features during parameter adjustment within the prediction model, filtering the original features based on their contribution rate to obtain target medical data features, and constructing a dependency graph between these features. The dependency graph then determines the feature relationships between these target features, and the feature thresholds are determined based on these relationships. Finally, the method generates risk rules based on these thresholds and target medical data features. This approach solves the problem of low accuracy in existing technologies due to the inability to perform cross-calculation of multiple features, thus improving the accuracy of risk rules and consequently, the accuracy of diagnostic results. Second, it addresses the issue of existing technologies only establishing relationships between features but failing to obtain feature thresholds. This allows for direct discrimination based on feature thresholds of the medical features to be processed, yielding corresponding diagnostic results and improving the efficiency of diagnostic result generation.
[0092] This disclosure also provides an example embodiment of a medical data mining apparatus. Specifically, refer to... Figure 17 As shown, the medical data mining device may include a feature contribution rate calculation module 1710, an association dependency graph construction module 1720, a feature threshold determination module 1730, and a risk rule generation module 1740. Wherein: The feature contribution rate calculation module 1710 can be used to acquire the medical data prediction model and the original medical data features, and calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. The dependency graph construction module 1720 can be used to filter the original medical data features according to the contribution rates of multiple features to obtain target medical data features, and construct a dependency graph corresponding to the target medical data features. The feature threshold determination module 1730 can be used to determine the feature association relationship between the target medical data features based on the association dependency graph, and to determine the feature threshold of the target medical data features based on the feature association relationship; The risk rule generation module 1740 can be used to generate risk rules that affect diseases corresponding to the original medical data features based on the feature threshold and the target medical data features.
[0093] In one exemplary embodiment of this disclosure, calculating the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model includes: Based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model.
[0094] In one exemplary embodiment of this disclosure, based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model, including: Step S10: Randomly select one feature from the original medical data features as the target object, and construct a feature set based on the other features in the original medical data features excluding the target object; Step S20: Construct multiple feature subsets based on the feature set, and calculate the first expected value of the first predicted value obtained by inputting the original medical data features into the medical data prediction model, and the second expected value of the second predicted value obtained by inputting the feature subsets into the medical data prediction model, based on the preset tree model interpreter. Step S30: Calculate the feature contribution rate of the target object during parameter adjustment in the medical data prediction model based on the number of features included in the feature subset, the number of feature subsets, the number of target objects, the first expected value, and the second expected value. Step S40: Repeat steps S10-S30 sequentially to obtain the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model.
[0095] In one exemplary embodiment of this disclosure, filtering the original medical data features based on the feature contribution rate to obtain target medical data features includes: The original medical data features are sorted according to the value of the feature contribution rate, and the original medical data features with a feature contribution rate greater than a preset threshold are selected from the sorted original medical data features as target medical data features.
[0096] In one exemplary embodiment of this disclosure, constructing a dependency graph corresponding to the target medical data features includes: Step S10': Select any feature from the target medical data features as the first feature object, select a feature from the original medical data features other than the first feature object as the second feature object, and obtain the first current contribution rate of the first feature object and the second current contribution rate of the second feature object. Step S20': Construct a coordinate system with the feature contribution rate as the first ordinate, the first current feature value of the first feature object as the abscissa, and the second current feature value of the second feature object as the second ordinate. Step S30': Determine a first coordinate pair based on the first current feature value and the first current contribution rate, and determine a second coordinate pair based on the second current feature value and the second current contribution rate; Step S40': Determine a first coordinate point based on the first coordinate position of the first coordinate pair in the coordinate system, and determine a second coordinate point based on the second coordinate position of the second coordinate pair in the coordinate system; Step S50': Based on the first coordinate point and the second coordinate point, generate a dependency graph between the first feature object and the second feature object, and repeat steps S10'-S40' to obtain a dependency graph corresponding to each feature in the target medical data features.
[0097] In one exemplary embodiment of this disclosure, determining the feature association relationships between the target medical data features based on the dependency graph includes: Obtain the target coordinates points that have overlapping relationships in the association dependency graph, and obtain the first target feature value, the second target feature value, the first target contribution rate corresponding to the first target feature, and the second target contribution rate corresponding to the second target feature corresponding to the second target feature; Linear fitting is performed on the first target feature value, the second target feature value, the first target contribution rate, and the second target contribution rate, and the feature correlation relationship between the target medical data features is determined based on the linear fitting result.
[0098] In one exemplary embodiment of this disclosure, determining the feature threshold of the target medical data feature based on the feature association includes: The trend category of the association dependency graph corresponding to the target medical data feature is determined based on the feature association relationship; wherein, the trend category includes an upward trend category and / or a downward trend category; Based on the trend category of the dependency graph, a data fitting trend is determined, and based on the data fitting trend, a feature threshold for the target medical data features included in the dependency graph is determined.
[0099] In one exemplary embodiment of this disclosure, the association dependency graph includes two different target medical data features, namely a first feature object and a second feature object; Wherein, when the trend category includes an upward trend category, the feature thresholds of the target medical data features included in the correlation dependency graph are determined based on the data fitting trend, including: The first target feature value of the first feature object and the second target feature value of the second feature object included in the association dependency graph are sorted, and the first maximum feature value and the first minimum feature value corresponding to the first target feature value, and the second maximum feature value and the second minimum feature value corresponding to the second target feature value are selected from the first target feature value and the second target feature value according to the sorting result. A first value range interval is constructed based on the first maximum eigenvalue, the first minimum eigenvalue, and other eigenvalues in the first target eigenvalues excluding the first maximum eigenvalue and the first minimum eigenvalue; and a second value range interval is constructed based on the second maximum eigenvalue, the second minimum eigenvalue, and other eigenvalues in the second target eigenvalues excluding the second maximum eigenvalue and the second minimum eigenvalue. Select a first set of points from the first value range interval that makes the sum of the first current contribution rates of the first feature object reach the maximum value, and select a second set of points from the second value range interval that makes the sum of the second current contribution rates of the second feature object reach the maximum value. The first minimum value is selected from the first set of points as the first minimum feature threshold of the first feature object, and the second minimum value is selected from the second set of points as the second minimum feature threshold of the second feature object.
[0100] In one exemplary embodiment of this disclosure, when the trend category includes a downward trend category, determining the feature threshold of the target medical data feature included in the correlation dependency graph based on the data fitting trend includes: The first maximum value is selected from the first set of points as the first maximum feature threshold of the first feature object, and the second maximum value is selected from the second set of points as the second maximum feature threshold of the second feature object.
[0101] In an exemplary embodiment of this disclosure, when the trend category includes an upward trend category and a downward trend category, determining the feature threshold of the target medical data feature included in the correlation dependency graph based on the data fitting trend includes: Select a first minimum value from the first set of points as the first minimum feature threshold of the first feature object, and select a second minimum value from the second set of points as the second minimum feature threshold of the second feature object; and Select the first maximum value from the first set of points as the first maximum feature threshold of the first feature object, and select the second maximum value from the second set of points as the second maximum feature threshold of the second feature object; A first threshold interval for the first feature object is constructed based on the first minimum feature threshold and the first maximum feature threshold, and a second threshold interval for the second feature object is constructed based on the second minimum feature value and the second maximum feature value.
[0102] This disclosure also provides an example embodiment of a medical data processing apparatus. Specifically, the medical data processing apparatus may include a medical feature extraction module, a risk rule acquisition module, and a diagnosis and treatment result generation module. Wherein: The medical feature extraction module can be used to acquire medical data to be processed and extract the medical features to be processed included in the medical data to be processed. The risk rule acquisition module can be used to acquire the risk rules of the disease corresponding to the medical data to be processed; wherein, the risk rules are generated based on any of the above-described medical data mining methods; The diagnosis and treatment result generation module can be used to identify the medical characteristics to be processed according to the risk rules, and obtain the identification results, which can be used by medical staff to obtain the diagnosis and treatment results of the disease.
[0103] The specific details of each module in the aforementioned medical data mining device and medical data processing device have been described in detail in the corresponding medical data mining method and medical data processing method, so they will not be repeated here.
[0104] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0105] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0106] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0107] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0108] The following reference Figure 18 To describe an electronic device 1800 according to such an embodiment of the present disclosure. Figure 18 The electronic device 1800 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0109] like Figure 18 As shown, the electronic device 1800 is manifested in the form of a general-purpose computing device. The components of the electronic device 1800 may include, but are not limited to: at least one processing unit 1810, at least one storage unit 1820, a bus 1830 connecting different system components (including storage unit 1820 and processing unit 1810), and a display unit 1840.
[0110] The storage unit stores program code that can be executed by the processing unit 1810, causing the processing unit 1810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1810 can perform actions such as... Figure 1Step S110: Obtain the medical data prediction model and the original medical data features, and calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model; Step S120: Filter the original medical data features according to the multiple feature contribution rates to obtain the target medical data features, and construct the association dependency graph corresponding to the target medical data features; Step S130: Determine the feature association relationship between the target medical data features according to the association dependency graph, and determine the feature threshold of the target medical data features according to the feature association relationship; Step S140: Generate risk rules affecting the diseases corresponding to the original medical data features according to the feature threshold and the target medical data features.
[0111] For example, the processing unit 1810 can perform actions such as Figure 16 Step S1610: Obtain medical data to be processed and extract the medical features to be processed included in the medical data to be processed; Step S1620: Obtain the risk rules of the disease corresponding to the medical data to be processed; wherein, the risk rules are generated based on any of the above-described medical data mining methods; Step S1630: Judge the medical features to be processed according to the risk rules to obtain the judgment result, so that medical staff can obtain the diagnosis and treatment results of the disease based on the judgment result.
[0112] Storage unit 1820 may include readable media in the form of volatile storage units, such as random access memory (RAM) 18201 and / or cache memory 18202, and may further include read-only memory (ROM) 18203.
[0113] Storage unit 1820 may also include a program / utility 18204 having a set (at least one) program module 18205, such program module 18205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0114] Bus 1830 can represent one or more of several types of bus structures, including memory cell bus or memory cell controller, peripheral bus, graphics acceleration port, processing unit, or local bus using any of the various bus structures.
[0115] Electronic device 1800 can also communicate with one or more external devices 1900 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1800, and / or any device that enables electronic device 1800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1850. Furthermore, electronic device 1800 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1860. As shown, network adapter 1860 communicates with other modules of electronic device 1800 via bus 1830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0116] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0117] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.
[0118] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0119] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0120] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0121] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0122] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0123] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0124] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention described herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not invented by this disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A method for mining medical data, characterized in that, include: Obtain the medical data prediction model and the original medical data features, and calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model; The original medical data features are filtered based on the contribution rates of multiple features to obtain target medical data features, and a dependency graph corresponding to the target medical data features is constructed. The feature relationships between the target medical data features are determined based on the dependency graph, and the trend category of the dependency graph corresponding to the target medical data features is determined based on the feature relationships; the data fitting trend is determined based on the trend category of the dependency graph, and the feature threshold of the target medical data features included in the dependency graph is determined based on the data fitting trend. The trend category includes an upward trend category and / or a downward trend category; the feature threshold is determined as follows: when the trend category is an upward trend category, a set of points is selected that makes the sum of the feature contribution rates of the target medical data feature reach the maximum value, and the minimum value in the set of points is determined as the feature threshold of the target medical data feature; When the trend category is a downward trend category, select the set of points that makes the sum of the feature contribution rates of the target medical data feature reach the maximum value, and determine the maximum value in the set of points as the feature threshold of the target medical data. Based on the feature threshold and the target medical data features, risk rules that affect the diseases corresponding to the original medical data features are generated.
2. The method for mining medical data according to claim 1, characterized in that, Calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model, including: Based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model.
3. The method for mining medical data according to claim 2, characterized in that, Based on a preset tree model interpreter, the feature contribution rate of each medical data feature in the original medical data features is calculated during the parameter adjustment process of the medical data prediction model, including: Step S10: Randomly select one feature from the original medical data features as the target object, and construct a feature set based on the other features in the original medical data features excluding the target object; Step S20: Construct multiple feature subsets based on the feature set, and calculate the first expected value of the first predicted value obtained by inputting the original medical data features into the medical data prediction model, and the second expected value of the second predicted value obtained by inputting the feature subsets into the medical data prediction model, based on the preset tree model interpreter. Step S30: Calculate the feature contribution rate of the target object during parameter adjustment in the medical data prediction model based on the number of features included in the feature subset, the number of feature subsets, the number of target objects, the first expected value, and the second expected value. Step S40: Repeat steps S10-S30 sequentially to obtain the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model.
4. The method for mining medical data according to claim 1, characterized in that, The original medical data features are filtered based on the feature contribution rate to obtain target medical data features, including: The original medical data features are sorted according to the value of the feature contribution rate, and the original medical data features with a feature contribution rate greater than a preset threshold are selected from the sorted original medical data features as target medical data features.
5. The method for mining medical data according to claim 1, characterized in that, Constructing a dependency graph corresponding to the target medical data features, including: Step S10': Select any feature from the target medical data features as the first feature object, select a feature from the original medical data features other than the first feature object as the second feature object, and obtain the first current contribution rate of the first feature object and the second current contribution rate of the second feature object. Step S20': Construct a coordinate system with the feature contribution rate as the first ordinate, the first current feature value of the first feature object as the abscissa, and the second current feature value of the second feature object as the second ordinate. Step S30': Determine a first coordinate pair based on the first current feature value and the first current contribution rate, and determine a second coordinate pair based on the second current feature value and the second current contribution rate; Step S40': Determine a first coordinate point based on the first coordinate position of the first coordinate pair in the coordinate system, and determine a second coordinate point based on the second coordinate position of the second coordinate pair in the coordinate system; Step S50': Based on the first coordinate point and the second coordinate point, generate a dependency graph between the first feature object and the second feature object, and repeat steps S10'-S40' to obtain a dependency graph corresponding to each feature in the target medical data features.
6. The method for mining medical data according to claim 5, characterized in that, Determining the feature associations between the target medical data features based on the dependency graph includes: Obtain the target coordinates that have overlapping relationships in the dependency graph, and obtain the first target feature value, the second target feature value, the first target contribution rate corresponding to the first target feature, and the second target contribution rate corresponding to the second target feature corresponding to the second target feature; Linear fitting is performed on the first target feature value, the second target feature value, the first target contribution rate, and the second target contribution rate, and the feature correlation relationship between the target medical data features is determined based on the linear fitting result.
7. The method for mining medical data according to claim 1, characterized in that, The association dependency graph includes two different target medical data features, namely a first feature object and a second feature object; Wherein, when the trend category includes an upward trend category, the feature thresholds of the target medical data features included in the correlation dependency graph are determined based on the data fitting trend, including: The first target feature value of the first feature object and the second target feature value of the second feature object included in the association dependency graph are sorted, and the first maximum feature value and the first minimum feature value corresponding to the first target feature value, and the second maximum feature value and the second minimum feature value corresponding to the second target feature value are selected from the first target feature value and the second target feature value according to the sorting result. A first value range interval is constructed based on the first maximum eigenvalue, the first minimum eigenvalue, and other eigenvalues in the first target eigenvalues excluding the first maximum eigenvalue and the first minimum eigenvalue; and a second value range interval is constructed based on the second maximum eigenvalue, the second minimum eigenvalue, and other eigenvalues in the second target eigenvalues excluding the second maximum eigenvalue and the second minimum eigenvalue. Select a first set of points from the first value range interval that makes the sum of the first current contribution rates of the first feature object reach the maximum value, and select a second set of points from the second value range interval that makes the sum of the second current contribution rates of the second feature object reach the maximum value; The first minimum value is selected from the first set of points as the first minimum feature threshold of the first feature object, and the second minimum value is selected from the second set of points as the second minimum feature threshold of the second feature object.
8. The method for mining medical data according to claim 7, characterized in that, When the trend category includes a downward trend category, the feature thresholds of the target medical data features included in the correlation dependency graph are determined based on the data fitting trend, including: The first maximum value is selected from the first set of points as the first maximum feature threshold of the first feature object, and the second maximum value is selected from the second set of points as the second maximum feature threshold of the second feature object.
9. The method for mining medical data according to claim 7, characterized in that, When the trend category includes both an upward trend category and a downward trend category, the feature thresholds of the target medical data features included in the correlation dependency graph are determined based on the data fitting trend, including: Select a first minimum value from the first set of points as the first minimum feature threshold of the first feature object, and select a second minimum value from the second set of points as the second minimum feature threshold of the second feature object; and Select the first maximum value from the first set of points as the first maximum feature threshold of the first feature object, and select the second maximum value from the second set of points as the second maximum feature threshold of the second feature object; A first threshold interval for the first feature object is constructed based on the first minimum feature threshold and the first maximum feature threshold, and a second threshold interval for the second feature object is constructed based on the second minimum feature value and the second maximum feature value.
10. A medical data mining device, characterized in that, include: The feature contribution rate calculation module is used to acquire the medical data prediction model and the original medical data features, and to calculate the feature contribution rate of each medical data feature in the original medical data features during the parameter adjustment process of the medical data prediction model. The dependency graph construction module is used to filter the original medical data features according to the contribution rates of multiple features to obtain target medical data features, and construct a dependency graph corresponding to the target medical data features. The feature threshold determination module is used to determine the feature association relationship between the target medical data features based on the association dependency graph, and to determine the trend category of the association dependency graph corresponding to the target medical data features based on the feature association relationship; to determine the data fitting trend based on the trend category of the association dependency graph, and to determine the feature threshold of the target medical data features included in the association dependency graph based on the data fitting trend. The trend category includes an upward trend category and / or a downward trend category; the feature threshold is determined as follows: when the trend category is an upward trend category, a set of points is selected that makes the sum of the feature contribution rates of the target medical data feature reach the maximum value, and the minimum value in the set of points is determined as the feature threshold of the target medical data feature; When the trend category is a downward trend category, select the set of points that makes the sum of the feature contribution rates of the target medical data feature reach the maximum value, and determine the maximum value in the set of points as the feature threshold of the target medical data. The risk rule generation module is used to generate risk rules that affect diseases corresponding to the original medical data features based on the feature threshold and the target medical data features.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the medical data mining method according to any one of claims 1-9.
12. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the medical data mining method according to any one of claims 1-9 via the executable instructions.
Citation Information
Patent Citations
Medical data mining method and device, storage medium and electronic equipment
CN114334167A
Disease prognosis risk prediction model training method and device and electronic equipment
CN114566284A