Early-stage Parkinson's disease interpretable three-branch intelligent evaluation method under multi-task remote data
Through multitasking remote data fusion and three interpretable decision-making mechanisms, a fuzzy decision tree is constructed, which solves the problems of insufficient generalization ability and lack of interpretability in the existing technology, and achieves efficient and accurate diagnosis of early Parkinson's disease, reducing the risk of misdiagnosis.
Patent Information
- Application Number
- CN202510735373.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing multitasking remote data evaluation model has limited generalization ability in early diagnosis of Parkinson's disease, making it difficult to sensitively identify changes in patient status, and the traditional model lacks interpretability, resulting in a high risk of misjudgment.
A multi-task remote data fusion strategy is adopted, combined with three interpretable decision mechanisms, a clear decision tree is built through the C4.5 algorithm, and learnable fine-tuning parameters and offset parameters are introduced to convert them into a fuzzy decision tree to process uncertain data, and realize sensitive identification and interpretable evaluation of early Parkinson's symptoms.
It improves the accuracy and robustness of early Parkinson's disease diagnosis, reduces the risk of misdiagnosis, is suitable for telemedicine scenarios, reduces the dependence on face-to-face medical treatment, and improves diagnostic efficiency and interpretability.
Smart Images

Figure CN120277616A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of assisted medical evaluation, and specifically discloses an interpretable three-branch intelligent evaluation method for early Parkinson's disease under multi-task remote data. Background Art
[0002] Parkinson's disease is a chronic progressive neurodegenerative disease, and its early stage is usually accompanied by mild but critical motor and non-motor symptoms. Research shows that compared with the irreversible nerve damage in the middle and late stages, the early stage is the window period with the most significant intervention effect. Therefore, achieving accurate evaluation early not only helps delay the progression of the disease but also significantly improves the quality of life of patients.
[0003] With the development of remote acquisition devices and data technologies, it has become more feasible and efficient to obtain and analyze the symptom characteristics of early Parkinson's disease, such as gait changes, voice fluctuations, and writing abnormalities, providing an important basis for the early intelligent evaluation of Parkinson's disease.
[0004] At the data collection level, although remote acquisition devices have broadened the data sources, the acquisition process is still easily interfered by factors such as the environment, patient status, and device methods. Traditional evaluation models constructed based on a single task often have limited generalization ability and are often difficult to sensitively identify changes in the patient's status. Therefore, obtaining data by designing multiple task scenarios has become an effective means. For example, using insoles with pressure sensors to collect gait data of patients walking normally, walking while listening to music, and walking while communicating; using professional sound pickup devices to record voice tasks such as pronouncing vowels and consonants; or using intelligent tablets to collect handwriting trajectories of writing spiral lines, meandering lines, and graphics. Such multi-task remote data contains multi-dimensional early manifestations of Parkinson's disease and has important application potential.
[0005] At the data processing level, existing methods for modeling evaluation models using multi-task remote data simply splice different task data and then process them uniformly. They neither fully eliminate the interference of individual physiological factors (such as height, age, weight) on features nor take into account that the feature distribution differences between different tasks may weaken the model's ability to extract general features across tasks. Therefore, integrating the commonalities and differences of multi-task remote data and effectively avoiding the interference of physiological differences are the basis for constructing accurate and robust evaluation models.
[0006] At the model construction level, although deep learning methods have strong feature extraction capabilities, due to their "black box" characteristics, they lack interpretability and it is difficult to gain the trust of doctors. In addition, traditional binary classification models usually rigidly divide patients into "diseased" or "healthy". When faced with patients with uncertain conditions, it is easy to cause misjudgment risks. Therefore, in the context of telemedicine, there is an urgent need to construct an early intelligent assessment method for Parkinson's disease that can handle uncertainty and has good interpretability, providing more reliable and transparent auxiliary diagnosis support for doctors.
[0007] The present invention provides an interpretable three-way intelligent assessment method for early Parkinson's under multi-task remote data to solve the above problems. Summary of the Invention
[0008] The purpose of the present invention is to construct an interpretable three-way intelligent assessment method for early Parkinson's under multi-task remote data. By integrating multi-task remote data, introducing a fuzzy partitioning strategy for effectively processing uncertain data, and an interpretable three-way decision-making mechanism, it realizes the processing of uncertainty and sensitive recognition of early Parkinson's symptoms, thereby improving the assessment accuracy and clinical usability, and supporting the early screening and intervention of diseases.
[0009] To achieve the above purpose, the basic solution of the present invention provides an interpretable three-way intelligent assessment method for early Parkinson's under multi-task remote data, including the following steps: Collect multi-task remote data to be measured, input it into the interpretable three-way intelligent assessment model for processing and analysis, and obtain an assessment result and an interpretability report; When the assessment result determines whether the patient is diseased, directly output healthy or diseased. Otherwise, output uncertain and further evaluate the disease risk; The construction method of the interpretable three-way intelligent assessment model includes the following steps: Step S1: Collect and process multi-task remote data to obtain a multi-task remote data set; Step S2: Use the multi-task remote data set as a sample set and input it into the C4.5 algorithm to construct a clear decision tree; Step S3: Take several non-leaf nodes in the clear decision tree as division points, introduce learnable fine-tuning parameters for each division point according to its position, and introduce learnable offset parameters for the left and right sides of each division point according to its position; Step S4: Input each sample in the sample set into the clear decision tree after step S3. Taking the minimum classification error function as the loss function, iteratively optimize the fine-tuning parameters of each division point position and the offset parameters of the left and right sides of each division point position. By constructing a triangular fuzzy membership function for all optimized division points and their offset parameters on the left and right sides, realize the transformation from the clear decision tree to the fuzzy decision tree; Step S5: According to the obtained fuzzy decision tree, aggregate the membership degrees of each sample with multiple splitting point paths, normalize them as the weights of the samples under this path, multiply the path weights by the probability distribution of the samples in the leaf nodes for different classes, and then sum them up for the same class, so as to determine the weighted probability vector of the samples for different classes; Step S6: Introduce a three-way decision mechanism to determine the interval into which the weighted probability vector falls as the evaluation result, use the size of the weighted probability vector corresponding to the evaluation result as the disease risk value, and construct an interpretable three-way intelligent evaluation model.
[0010] Furthermore, during the implementation of step S1, the following steps are included: Step S11: Collection of multi-task remote data: Collect multi-task remote data of early Parkinson's patients through remote data sources or intelligent devices; Step S12: Processing of multi-task remote data: including manual feature extraction, physiological feature de-correlation, cross-task feature distribution alignment, and data normalization carried out in sequence, and finally obtain a multi-task remote data set with unified structure and coordinated features.
[0011] Furthermore, during the implementation of step S2, the following steps are included: Step S21: Perform dichotomy processing on all features of the nodes of the clear decision tree to be established; Step S22: Use the multi-task remote data set as the sample set input, and construct a clear decision tree by running the C4.5 algorithm; Step S23: Use the post-pruning strategy to prevent overfitting of the clear decision tree.
[0012] Furthermore, in step S3, the expression for adjusting the splitting point position based on the fine-tuning parameter is as follows: ; where q = 1, 2, …, Q, and Q is the number of non-leaf nodes, is the value of the splitting point of the original clear decision tree, is the splitting point after fine-tuning, is the learnable fine-tuning parameter; The offset parameter for the left position of the splitting point and the offset parameter for the right position of the splitting point are respectively represented in the following ways: ; ; where is the learnable parameter for controlling the offset parameter of the left position of the splitting point, is the learnable parameter for controlling the offset parameter of the right position of the splitting point, The initial values of are both 0.
[0013] Furthermore, during the implementation of step S4, the following steps are included: Step S41: Input each sample in the multi-task remote dataset into the clear decision tree after step S3 for forward propagation. Pass through each splitting point in sequence. According to each splitting point and the offset parameters on its left and right sides, construct a triangular fuzzy membership function to obtain the membership degree of each sample to different splitting points; Step S42: Aggregate the membership degrees of each sample containing multiple splitting point paths. After normalization, use it as the weight of the sample under this path. Multiply the path weight by the class distribution probability of the sample in the leaf node, and then sum them under the same class to determine the weighted probability vector of the sample to different classes, and select the class with the maximum weighted probability value as the predicted label of the sample; Step S43: Use the error between the true label and the predicted label of the sample to construct a classification error loss function, and optimize three parameters through backpropagation and gradient update: the fine-tuning parameter of the splitting point, the offset parameter of the position on the left side of the splitting point, and the offset parameter of the position on the right side of the splitting point; Step S44: Through multiple rounds of training, iteratively optimize all parameters to minimize the loss function, and finally obtain the fine-tuning parameter of each splitting point and the offset parameters of its left and right positions; Step S45: Based on the finally optimized fine-tuning parameter and the offset parameters of its left and right positions, obtain the fine-tuned splitting point and the offset parameters of its left and right positions. Construct a triangular fuzzy membership function for each splitting point to convert each splitting point into a fuzzy splitting point, and finally obtain a fuzzy decision tree.
[0014] Furthermore, in step S41, the triple representing the construction of the triangular fuzzy membership function is , so the triangular fuzzy membership function corresponding to each fuzzy splitting point is: ; In the formula, q = 1, 2,..., Q, where Q is the number of non-leaf nodes, is the fine-tuned splitting point, is the learnable offset parameter of the position on the left side of the splitting point, is the learnable offset parameter of the position on the right side of the splitting point, and x is a sample in the multi-task remote dataset; When , ; When , .
[0015] Furthermore, in step S42, the membership degrees of the aggregated sample x to all splitting points on each path are expressed as: ; where w m (x) represents the weight of sample x under the m-th path, d is the path depth, and μ mj (x) is the membership degree of sample x to the j-th splitting point in path m; The normalized path weight W m is expressed as: ; where M is the number of paths, and W m is the normalized weight of sample x under the m-th (m = 1, 2,..., M) path; The calculation formula for the weighted probability vector (P0, P1) of sample x for different classes is as follows: ; where M is the number of paths, t represents two classes of healthy or diseased, t = 0 or 1, represents the class distribution probability of samples in the leaf nodes under the m-th path, that is, the class distribution probability of samples in the leaf nodes of the corresponding clear decision tree; The predicted label t of sample x is calculated by the maximum weighted probability formula: Predicted label t = max t P t ; where t = 0 or 1 represents two classes of healthy and diseased.
[0016] Furthermore, during the implementation of step S5, it includes the following steps: Step S51: After the sample is input into the fuzzy decision tree, calculate the fuzzy splitting point membership degree: According to the obtained fuzzy decision tree, for any fuzzy splitting point, use the triangular fuzzy membership function to calculate the fuzzy membership degree of the sample to this fuzzy splitting point; Step S52: Calculate the path weight: Aggregate the comprehensive membership degree values of the sample on each path, and after normalization, obtain the weight value of the sample on each path; Step S53: Output the weighted class probability vector: Multiply the path weight and the class distribution probability of the sample in the leaf node, and then add them up under the same class to determine the probability vector of the sample for different classes , where represents the healthy probability, represents the diseased probability.
[0017] Furthermore, during the implementation of step S6, the three-way decision mechanism is: When it is the case, a determined result is output, that is, when and it is the case, the output result is: diseased. When and it is the case, the output result is: healthy; When it is the case, the output result is: uncertain, as well as a diseased evaluation value ; Among them, the value range of the confidence threshold ξ is: 0.5 < ξ < 1.
[0018] The principle and effect of this solution are as follows: 1. Compared with the prior art, the present invention is more applicable to the telemedicine scenario: obtaining multi-task behavior data of patients through a remote data acquisition module, and combining with an intelligent classification model to achieve timely and accurate diagnostic support, effectively reducing patients' dependence on face-to-face consultations, and significantly improving the practicability and diagnostic efficiency of telemedicine.
[0019] 2. Compared with the prior art, the cross-task generality and difference joint modeling strategy provided by the present invention can utilize any one of multi-task remote data such as gait, speech, writing, etc. to achieve joint modeling of multi-dimensional symptom features, which can not only effectively enhance the comprehensiveness and robustness of the model for early assessment of Parkinson's disease, but also improve the stability and adaptability of the model in different tasks and different populations.
[0020] 3. Compared with the prior art, the present invention uses a learnable fuzzy partitioning mechanism, with the decision tree partitioning point as the center, introducing learnable fine-tuning parameters and offset parameters, and fine-tuning the partitioning point and optimizing the fuzzy region boundary through forward propagation and backward propagation to achieve adaptive fuzzification. Therefore, the present invention has strong flexibility, self-adaptability, and the ability to process and interpret sample uncertainty.
[0021] 4. Compared with the prior art, the present invention has an interpretable evaluation framework, constructs a fuzzy decision tree based on a clear decision tree, and introducing fuzzy partitioning enhances the ability to process uncertain data. By introducing a three-way decision mechanism, flexible discrimination of "healthy - uncertain - diseased" is achieved. Therefore, the present invention conforms to medical diagnosis logic and can effectively reduce the risk of misdiagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It shows the flowchart of the interpretable three - branch intelligent evaluation method for early Parkinson's disease under multi - task remote data proposed in the embodiments of the present application; Figure 2 It shows the technical roadmap of multi - task remote data collection and multi - task remote data processing proposed in the embodiments of the present application; Figure 3 It shows the triangular fuzzy membership function obtained after fuzzification for any division point b in the interpretable three - branch intelligent evaluation method for early Parkinson's disease under multi - task remote data of the embodiments of the present application q obtained through fuzzification; Figure 4 It shows the visual - aided display of using multi - task remote gait data in step S5 of the interpretable three - branch intelligent evaluation method for early Parkinson's disease under multi - task remote data of the embodiments of the present application; Figure 5 It shows the schematic diagram of the interpretable three - branch intelligent evaluation device for early Parkinson's disease under multi - task remote data proposed in the embodiments of the present application. Detailed implementation manners
[0024] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and their effects of the present invention as follows.
[0025] The interpretable three - branch intelligent evaluation method for early Parkinson's disease under multi - task remote data, as an example of the embodiment Figures 1 to 4 is shown as follows: It includes the following steps: Collect the multi - task remote data to be measured, input it into the interpretable three - branch intelligent evaluation model for processing and analysis, and obtain the evaluation result and the interpretability report; When the evaluation result determines whether the patient is ill, directly output healthy or ill; otherwise, output uncertain and further evaluate the risk of illness.
[0026] Among them, the construction method of the interpretable three - branch intelligent evaluation model includes the following steps: Step S1: Collect and process the multi - task remote data to obtain the multi - task remote data set.
[0027] As Figure 2 shown, it specifically includes the following steps: Step S11: Collection of multi-task remote data: Collect multi-task remote data of early Parkinson's patients through remote data sources or intelligent devices. The multi-task remote data includes any one of the following: Obtain multi-task gait data of the patient walking in a normal state, walking while listening to rhythmic music, and walking during a communication task using insoles with pressure sensors; Obtain multi-task speech data of the patient during tasks such as pronouncing vowels and consonants using a microphone (such as a mobile phone or a medical-grade sound pickup device); Collect multi-task hand-drawn image data of the patient with handwritten spiral lines, meandering lines, and simple graphics as templates using an intelligent tablet.
[0028] Step S12: Processing of multi-task remote data: Include operations such as manual feature extraction, de-correlation of physiological features, cross-task feature distribution alignment, and data normalization, and finally obtain a multi-task remote data set with unified structure and coordinated features. It includes: Processing of multi-task gait data, specifically, process the multi-task gait data of the patient walking normally, walking while listening to music, and walking during communication collected by the insole with a pressure sensor, including the following four steps: Step a121: Manually extract time-domain features, spatial-domain features, gait cycle variability, symmetry index, step length, single / double support time ratio, etc. of multi-task gait data; Step a122: De-correlation of physiological features: Use multiple regression methods to remove the influence of physiological features such as height, weight, and age on multi-task gait data; Step a123: Cross-task feature distribution alignment: Use the principal component analysis (PCA) method or other feature selection methods to reduce the dimension of the gait data of each task, extract the shared principal components between tasks for constructing a unified feature representation, and at the same time retain some task-specific components to enhance individual discriminability, so as to achieve cross-task feature distribution alignment; Step a124: Data normalization: Normalize all features, such as using Z-score standardization or Min-Max normalization.
[0029] Processing of multi-task speech data, specifically, process the multi-task speech data such as vowels and consonants emitted by the patient using a professional sound pickup device: including the following four steps: Step b121: Manually extract frequency-domain features, fundamental frequency features, harmonic features, and glottal dynamic features, etc. of multi-task speech data; Step b122: De-correlation of physiological features: Use multiple regression methods to remove the influence of physiological features such as height, weight, and age on multi-task speech data; Step b123: Cross-task feature distribution alignment: Use the principal component analysis (PCA) method or other feature selection methods to reduce the dimension of the speech data for each task, extract the shared principal components between tasks for constructing a unified feature representation, and at the same time retain some task-specific components to enhance individual discriminability, so as to achieve cross-task feature distribution alignment; Step b124: Data normalization: Normalize all features, such as using Z-score standardization or Min-Max normalization.
[0030] The processing of multi-task hand-drawn image data specifically refers to the processing of multi-task hand-drawn images of the handwriting trajectories of patients collecting spiral lines, meandering lines and standard graphics with the help of a smart tablet, including the following four steps: Step c121: Manually extract the trajectory offset, tremor amplitude, line regularity and morphological statistical features of the multi-task hand-drawn image; Step c122: Physiological feature de-correlation: Use the multiple regression method to remove the influence of physiological features such as height, weight and age on the multi-task hand-drawn image data; Step c123: Cross-task feature distribution alignment: Use the principal component analysis (PCA) method or other feature selection methods to reduce the dimension of the hand-drawn image data for each task, extract the shared principal components between tasks for constructing a unified feature representation, and at the same time retain some task-specific components to enhance individual discriminability, so as to achieve cross-task feature distribution alignment; Step c124: Data normalization: Normalize all features, such as using Z-score standardization or Min-Max normalization.
[0031] Process different multi-task remote datasets respectively to obtain multi-task remote datasets with unified structure and coordinated features: gait multi-task remote dataset, speech multi-task remote dataset and hand-drawn image multi-task remote dataset.
[0032] Step S2: Use the multi-task remote dataset as the sample set and input it into the C4.5 algorithm to construct a clear decision tree.
[0033] It includes the following specific steps: Step S21: Perform dichotomy processing on all features of the nodes of the clear decision tree to be established: For the discrete attributes of the nodes, use the frequency to replace the discrete attribute values to obtain continuous attribute values; then, use dichotomy (such as threshold-based, equal-interval or equal-frequency segmentation) to transform the discrete attributes and the original continuous attributes for subsequent evaluation model modeling.
[0034] Step S22: Using the multi-task remote dataset as the sample set input, construct a clear decision tree by running the C4.5 algorithm: By running the C4.5 algorithm, calculate the partitioning ability of each discrete feature using the information gain ratio, and select the discrete feature with the maximum information gain ratio for partitioning. This process is carried out recursively until one of the following termination conditions is met: (i) the information gain ratio of the best partitioning discrete feature is lower than the set threshold; (ii) the number of samples in the subset is lower than the set threshold; (iii) all samples in the subset belong to the same category.
[0035] Step S23: Use the post-pruning strategy to prevent overfitting of the clear decision tree. Prune some redundant branches through cross-validation or minimizing the validation set error, thereby improving the generalization ability of the model.
[0036] Step S3: Using several non-leaf nodes in the clear decision tree as partitioning points, introduce learnable fine-tuning parameters for each partitioning point according to its position, and introduce learnable offset parameters for the left and right sides of each partitioning point according to its position.
[0037] Step S31: Extract several non-leaf nodes from the clear decision tree as partitioning points: From the clear decision tree constructed in Step S2, extract the partitioning attributes and corresponding partitioning points b q (q = 1, 2,..., Q) of Q non-leaf nodes as the central parameters of the subsequent triangular fuzzy membership function.
[0038] Step S32: Introduce learnable fine-tuning parameters for each partitioning point according to its position, and introduce learnable offset parameters for the left and right sides of each partitioning point according to its position: The offset parameter on the left side of the partitioning point and the offset parameter on the right side of the partitioning point , and allow fine-tuning of the partitioning point b q . Specifically as follows: ; where q = 1, 2,..., Q, Q is the number of non-leaf nodes, is the value of the partitioning point of the original clear decision tree, is the partitioning point after fine-tuning, is the learnable fine-tuning parameter; The offset parameter at the left position of the partitioning point and the offset parameter at the right position of the partitioning point are respectively represented by the following methods: ; ; where is the learnable parameter that controls the offset parameter at the left position of the partitioning point, is a learnable parameter for the offset parameter of the position on the right side of the control division point, and both have an initial value of 0 (i.e., consistent with the clear decision tree).
[0039] In this example, the softplus function has the following effects: (1) ensuring that > 0 and > 0, that is, the offset parameter is always the effective width; (2) remaining continuously differentiable for easy gradient descent; (3) having a nonlinearity similar to ReLU with a stable gradient range.
[0040] Step S4: Input each sample in the sample set into the clear decision tree after Step S3, use the minimum classification error function as the loss function, iteratively optimize the fine-tuning parameter of each division point position and the offset parameters of the positions on both the left and right sides of each division point, and construct a triangular fuzzy membership function for all the optimized division points and their offset parameters on both sides to achieve the transformation from the clear decision tree to the fuzzy decision tree. Specifically, it includes the following steps: Step S41: Input each sample in the multi-task remote dataset into the clear decision tree after Step S3 for forward propagation, pass through each division point in sequence, and construct a triangular fuzzy membership function according to each division point and its offset parameters on both the left and right sides to obtain the membership degree of each sample to different division points.
[0041] According to Step S32, the parameter triple of the triangular fuzzy membership function is obtained as , so, as Figure 3 shown, in the forward propagation stage, based on the triple the triangular fuzzy membership function is: ; In the formula, q = 1, 2,..., Q, where Q is the number of non-leaf nodes, is the fine-tuned division point, is the learnable offset parameter of the position on the left side of the division point, is the learnable offset parameter of the position on the right side of the division point, and x is the input data: a sample in the multi-task remote dataset.
[0042] Specifically, when , ; When , .
[0043] According to the above triangular fuzzy membership function, calculate the membership degree values of the sample \(x\) for each fuzzy partition node. This step uses the triangular fuzzy membership function to fuzzify the partition points in the clear decision tree, thereby constructing a soft partition area with continuous transition characteristics at each partition node.
[0044] Step S42: Aggregate the membership degrees of each sample containing multiple partition point paths, normalize them as the weights of the sample under this path, multiply the path weights by the class distribution probabilities of the samples in the leaf nodes, and then sum them under the same class, so as to determine the weighted probability vectors of the sample for different classes, and select the class with the largest weighted probability value as the predicted label of the sample.
[0045] Specifically, aggregate the membership degrees of the sample \(x\) at all partition points on each path through T-norm, which is expressed as: ; In the formula, \(w\) m (x) represents the weight of the sample \(x\) under the \(m\)-th path, \(d\) is the path depth, and \(\mu\) mj (x) is the membership degree of the sample \(x\) for the \(j\)-th partition point in the path \(m\); The normalized path weight \(W\) m is expressed as: ; In the formula, \(M\) is the number of paths, and \(W\) m is the normalized weight of the sample \(x\) under the \(m\) ( \(m = 1, 2, \ldots, M\))-th path; The formula for the weighted probability vector \((P_0, P_1)\) of the sample \(x\) for different classes is as follows: ; In the formula, \(M\) is the number of paths, \(t = 0\) or \(1\) represents the two classes of healthy or diseased, represents the class distribution probability of the samples in the leaf node under the \(m\)-th path, that is, the class distribution probability of the samples in the leaf node in the corresponding clear decision tree; The predicted label \(t\) of the sample \(x\) is calculated through the maximum weighted probability formula: The predicted label \(t=\max\) t P t ; In the formula, \(t = 0\) or \(1\) represents the two classes of healthy and diseased.
[0046] The weights of this weighted fusion come from the membership degree values of the paths, ensuring that paths with higher membership degrees contribute more to the final classification result.
[0047] Select the class with the maximum membership degree and weighted probability value as the predicted class of sample x according to the above formula. The weights of this weighted fusion come from the membership degree values of the paths, ensuring that paths with higher membership degrees contribute more to the final classification result.
[0048] Step S43: Construct a classification error loss function using the error between the true label and the predicted label of the sample, and optimize three parameters through backpropagation and gradient update: the fine-tuning parameter of the splitting point, the offset parameter of the position to the left of the splitting point, and the offset parameter of the position to the right of the splitting point.
[0049] During the training process, use the cross-entropy loss function to measure the error between the prediction and the true label. The cross-entropy loss function is: ; In the formula, is the probability distribution of the prediction where the model fuses the outputs of multiple paths, y i is the one-hot encoding of the true label, and N represents the number of samples in the input dataset.
[0050] To improve the generalization performance and robustness of the model, set the regularization term of the offset parameter to constrain the width of the offset parameter, prevent the fuzzy region from being too wide or too narrow, and maintain boundary stability. The regularization term is expressed as: ; In the formula, λ is the regularization parameter.
[0051] Therefore, the total loss function is: ; Backpropagation and gradient update: During the backpropagation process, the model automatically calculates the gradients of the loss function with respect to all learnable parameters , and using the chain rule: ; Through layer-by-layer backpropagation, the gradients are passed from the output layer to each node in turn. Subsequently, use the gradient descent algorithm or its variants (such as Adam, RMSprop, etc.) to update the model parameters: ; ; ; In the formula, η is the learning rate, and n is the current iteration number.
[0052] Through this optimization process, the model continuously adjusts the values of each parameter to minimize the loss function and improve the evaluation performance.
[0053] Step S44: Through multiple rounds of training, iteratively optimize all parameters to minimize the loss function, and finally obtain the fine-tuned parameters for each splitting point and the offset parameters for the positions on its left and right sides.
[0054] Multiple rounds of training and tuning: Through multiple rounds of training (epochs), the model gradually adjusts all parameters to minimize the loss function. In each round of training, an optimization algorithm (such as Adam) is used to continuously adjust the three parameters until the change in the loss function tends to zero or reaches a preset stopping condition. Finally, three optimal parameters are obtained: ; ; ; where, Based on the splitting point b q obtained after fine-tuning, is the learnable parameter for optimizing the offset parameter on the left side of the splitting point, is the learnable parameter for optimizing the offset parameter on the right side of the splitting point.
[0055] Step S45: Based on the finally optimized fine-tuned parameters and the offset parameters for the positions on their left and right sides, obtain the fine-tuned splitting points and the offset parameters for the positions on their left and right sides, construct a triangular fuzzy membership function for each splitting point to convert each splitting point into a fuzzy splitting point, and finally obtain a fuzzy decision tree.
[0056] The decision tree fuzzification method based on learnable offset parameters can effectively avoid problems such as being prone to falling into local optima, difficult to parallelize training, and poor generalization ability caused by relying on heuristic search in traditional fuzzy decision trees. By implementing differentiable and learnable modeling of the left and right offsets of the splitting points, it can significantly improve the training efficiency, robustness, and generalization ability of the model.
[0057] Step S5: According to the obtained fuzzy decision tree, aggregate the membership degrees of each sample containing multiple splitting point paths, normalize them as the weights of the sample under this path, multiply the path weights by the class distribution probabilities of the samples in the leaf nodes, and then sum them for the same class, so as to determine the weighted probability vector of the sample for different classes.
[0058] It includes the following specific steps: Step S51: After the sample is input into the fuzzy decision tree, calculate the membership degree of the fuzzy splitting point: According to the obtained fuzzy decision tree, for any fuzzy splitting point, use the triangular fuzzy membership function to calculate the fuzzy membership degree of the sample on this fuzzy splitting point. According to the obtained optimal parameters, calculate the fine-tuned splitting point , the offset parameter for its left position and the offset parameter for its right position : ; ; ; For each sample x, starting from the root node, traverse all paths with non-zero membership degrees for this sample (the paths are fuzzy, so a sample may belong to multiple paths). For any fine-tuned partitioning point , calculate the fuzzy membership degree of this sample at this node based on the following triangular fuzzy membership function: ; where q = 1, 2,..., Q, and Q is the number of non-leaf nodes, is the triangular fuzzy membership function corresponding to each non-leaf node q in the final fuzzy decision tree, obtained based on the fine-tuned partitioning point b, is the offset parameter for the left position of the learnable partitioning point after optimization, is the offset parameter for the right position of the learnable partitioning point after optimization.
[0059] Specifically, when , ; When , ; Step S52: Path weight calculation: Aggregate the comprehensive membership degree values of the sample on each path, and after normalization, obtain the weight value of the sample on each path: Use T-norm to aggregate the membership degrees of all nodes on the path. The comprehensive membership degree of the sample x for each path from the root node to the leaf node is expressed as: ; where represents the weight of the sample x under the m-th path, d is the path depth, is the membership degree of the sample x at the j-th partitioning point under the m-th path in the final fuzzy decision tree; The normalized path weight is expressed as: ; where M is the number of paths, is the normalized weight of the sample x on the m-th (m = 1, 2,..., M) path in the final fuzzy decision tree.
[0060] Step S53: Output the weighted category probability vector: Multiply the path weight by the category distribution probability of the samples in the leaf node, and then sum them up for the same category to determine the probability vector of the sample for different categories. , where represents the health probability, represents the disease probability: The weighted classification probability formula for sample x is: ; In the formula, t represents two categories of health or disease, t = 0 or 1, represents the category distribution probability of the samples in the leaf node under the m-th path, that is, the category distribution probability of the samples in the leaf node in the corresponding clear decision tree.
[0061] Step S6: Introduce a three-way decision mechanism to determine the interval into which the weighted probability vector falls as the evaluation result, use the size of the weighted probability vector corresponding to the evaluation result as the disease risk value, and construct an interpretable three-way intelligent evaluation model.
[0062] Specifically, first set the confidence threshold ξ (0.5 < ξ < 1) in combination with expert opinions to define the boundary between "definite" and "uncertain". The three-way decision is as follows: When , output a definite result, specifically including: when and , output the result "diseased", when and , output the result "healthy"; When , output the result "uncertain & diseased evaluation value ".
[0063] In this embodiment, the interpretable three-way intelligent evaluation method for early Parkinson's under multi-task remote data also includes generating a model decision path report and visualizing auxiliary display, and the model decision path report is the interpretability report.
[0064] Model decision path report: Extract the weights of the samples (a tester) in all paths, and select the top-K paths with the largest weights for the diseased category as the key paths for the model to make a judgment. Integrate all key paths, node contributions, and membership degrees to generate a structured interpretable evaluation report.
[0065] The report content includes: the confidence level and the final diagnosis conclusion determined according to expert opinions; the direction and membership degree of the sample in each path; the action interval of each key feature in the judgment.
[0066] Visualization auxiliary display: Such as Figure 4As shown, the positions of the patient's features in the fuzzy division intervals, the path weights, and the final judgment results are graphically and visually displayed, facilitating doctors to quickly understand the model behavior and decision-making basis.
[0067] In this embodiment, an early Parkinson's interpretable three-branch intelligent evaluation device under multi-task remote data is also provided. It should be understood that this device is used to execute the construction method of the interpretable three-branch intelligent evaluation model and is used to execute the early Parkinson's interpretable three-branch intelligent evaluation method under multi-task remote data.
[0068] As Figure 5 shown, it specifically includes a multi-task remote data acquisition module, a multi-task remote data processing module, a model construction module for executing the construction method of the interpretable three-branch intelligent evaluation model, and a model training module and a disease evaluation module for executing the early Parkinson's interpretable three-branch intelligent evaluation method under multi-task remote data.
[0069] The multi-task remote data acquisition module refers to collecting multi-task remote data of early Parkinson's patients through remote data sources or intelligent devices. The multi-task remote data includes any one of the following: obtaining multi-task gait data of patients walking in a normal state, walking while listening to rhythmic music, and walking during a communication task with people by using insoles with pressure sensors; obtaining multi-task voice data of patients during tasks such as pronouncing vowels and consonants by using a microphone (such as a mobile phone or a medical-grade sound pickup device); collecting multi-task hand-drawn image data of patients using handwritten spiral lines, meandering lines, and simple graphics as templates by using an intelligent tablet. Then, the multi-task remote data processing module performs operations such as manual feature extraction, physiological feature de-correlation, cross-task feature distribution alignment, and data normalization on the multi-task remote data, and finally obtains a multi-task remote data set with unified structure and coordinated features.
[0070] The model construction module is specifically divided into: The first construction module: constructing a clear decision tree. Using the C4.5 algorithm to construct an initial clear decision tree for the processed multi-task remote data set; The second construction module: decision tree fuzzification. Taking each division point in the clear decision tree as the initial center of the triangular fuzzy membership function, introducing learnable offset parameters for its position on both its left and right sides. Through the forward propagation and backward propagation mechanisms, jointly fine-tuning the position of the division point and the offset parameters of the positions on both its left and right sides, enabling the model to adaptively adjust the fuzzy range of the division point. Constructing a triangular fuzzy membership function according to the optimized parameters, thereby realizing the fuzzification of the clear decision tree. The third construction module: constructing an intelligent evaluation model based on a three-branch fuzzy decision tree. According to the obtained fuzzy decision tree, calculating the probability vector of the sample for each category. By introducing a three-branch decision mechanism, constructing an intelligent evaluation model based on a three-branch fuzzy decision tree.
[0071] The model training module refers to training the interpretable three-branch intelligent evaluation model through a multi-task remote dataset to meet the usage standards.
[0072] The disease evaluation module refers to inputting the multi-task remote data of a new patient and using the intelligent evaluation model based on the three-branch fuzzy decision tree to output the evaluation result, the model decision path report, and the visual auxiliary display.
[0073] As mentioned above, it is only the preferred embodiment of the present invention, and there is no any form of limitation to the present invention. Although the present invention has been disclosed above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An interpretable three-branch intelligent evaluation method for early Parkinson's disease under multi-task remote data, characterized in that, It includes the following steps: Collect multi-task remote data to be measured, input it into an interpretable three-way intelligent evaluation model for processing and analysis, and obtain an evaluation result and an interpretability report; When the evaluation result determines whether the patient is ill, directly output healthy or ill; otherwise, output uncertain and further evaluate the disease risk; The construction method of the interpretable three-way intelligent evaluation model includes the following steps: Step S1: Collect and process multi-task remote data to obtain a multi-task remote data set; Step S2: Use the multi-task remote data set as a sample set and input it into the C4.5 algorithm to construct a clear decision tree; Step S3: Take several non-leaf nodes in the clear decision tree as dividing points, introduce learnable fine-tuning parameters for each dividing point according to its position, and introduce learnable offset parameters for the left and right sides of each dividing point according to its position; Step S4: Input each sample in the sample set into the clear decision tree after step S3. Taking the minimum classification error function as the loss function, iteratively optimize the fine-tuning parameters at each dividing point position and the offset parameters at the left and right positions of each dividing point. By constructing a triangular fuzzy membership function for all optimized dividing points and their offset parameters on the left and right sides, realize the transformation from the clear decision tree to the fuzzy decision tree; Step S5: According to the obtained fuzzy decision tree, aggregate the membership degrees of each sample containing multiple dividing point paths, normalize them as the weights of the sample under this path, multiply the path weights by the class distribution probabilities of the samples in the leaf nodes, and then sum them under the same class to determine the weighted probability vector of the sample for different classes; Step S6: Introduce a three-way decision-making mechanism to determine the interval into which the weighted probability vector falls as the evaluation result, take the size of the weighted probability vector corresponding to the evaluation result as the disease risk value, and construct an interpretable three-way intelligent evaluation model.
2. The early Parkinson's interpretable three-branch intelligent evaluation method for multi-task remote data according to claim 1, characterized in that, During the implementation of step S1, it includes the following steps: Step S11: Collection of multi-task remote data: Collect multi-task remote data of early Parkinson's patients through remote data sources or intelligent devices; Step S12: Processing of multi-task remote data: It includes manual feature extraction, physiological feature de-correlation, cross-task feature distribution alignment, and data normalization in sequence, and finally obtains a multi-task remote data set with unified structure and coordinated features.
3. The early Parkinson's interpretable three-branch intelligent evaluation method for multitask remote data according to claim 1, characterized in that, During the implementation of step S2, it includes the following steps: Step S21: Perform dichotomy processing on all features of the nodes of the clear decision tree to be established; Step S22: Use the multi-task remote data set as a sample set and input it. By running the C4.5 algorithm, construct a clear decision tree; Step S23: Use a post-pruning strategy to prevent overfitting of the clear decision tree.
4. The early Parkinson's interpretable three-branch intelligent evaluation method for multitask remote data according to claim 1, wherein In step S3, the expression for adjusting the dividing point position based on the fine-tuning parameter is as follows: ; where \(q = 1, 2, \ldots, Q\), and \(Q\) is the number of non-leaf nodes, is the value of the splitting point of the original clear decision tree, is the splitting point after fine-tuning, is the learnable fine-tuning parameter; Offset parameter for the position to the left of the division point and the offset parameter for the position to the right of the division point are respectively represented as follows: ; ; In the formula, is a learnable parameter for the offset parameter controlling the position on the left side of the division point, is a learnable parameter for the offset parameter controlling the position on the right side of the division point, and both have an initial value of 0.
5. The early Parkinson's interpretable three-branch intelligent evaluation method for multi-task remote data according to claim 1, wherein During the implementation of step S4, it includes the following steps: Step S41: Input each sample in the multi-task remote data set into the clear decision tree after step S3 for forward propagation, pass through each dividing point in sequence, and construct a triangular fuzzy membership function according to each dividing point and its offset parameters on the left and right sides to obtain the membership degree of each sample for different dividing points; Step S42: Aggregate the membership degrees of each sample containing multiple dividing point paths, normalize them as the weights of the sample under this path, multiply the path weights by the probability distribution of the samples in the leaf nodes for the same category, and then sum them up for the same category, so as to determine the weighted probability vector of the sample for different categories, and select the category with the largest weighted probability value as the predicted label of the sample; Step S43: Use the error between the true label and the predicted label of the sample to construct a classification error loss function, and optimize three parameters through backpropagation and gradient update: the fine-tuning parameter of the dividing point, the offset parameter of the position on the left side of the dividing point, and the offset parameter of the position on the right side of the dividing point; Step S44: Through multiple rounds of training, iteratively optimize all parameters to minimize the loss function, and finally obtain the fine-tuning parameter of each dividing point and the offset parameters of its left and right positions; Step S45: Based on the finally optimized fine-tuning parameter and the offset parameters of its left and right positions, obtain the fine-tuned dividing point and the offset parameters of its left and right positions, construct a triangular fuzzy membership function for each dividing point to convert each dividing point into a fuzzy dividing point, and finally obtain a fuzzy decision tree.
6. The early Parkinson's interpretable three-branch intelligent evaluation method for multi-task remote data according to claim 5, characterized in that, In step S41, the triple of the triangular fuzzy membership function is expressed as , so the triangular fuzzy membership function corresponding to each fuzzy division point is as follows: ; where q = 1, 2, …, Q, and Q is the number of non-leaf nodes, is the partition point after fine-tuning, is the offset parameter for the position to the left of the learnable partition point, is the offset parameter for the position to the right of the learnable partition point, and x is a sample in the multi-task remote dataset; When then ; When then 。 7. The multi-task remote data-based early Parkinson's interpretable three-branch intelligent evaluation method according to claim 5, characterized in that In step S42, aggregate the membership degrees of sample x for all dividing points on each path, which is expressed as: ; where w m (x) represents the weight of sample x under the m-th path, d is the path depth, and μ mj (x) is the membership degree of sample x to the j-th splitting point in path m; Normalized path weight W m It is expressed as: ; where M is the number of paths, and W m is the normalized weight of the sample x on the m-th (m = 1, 2, …, M) path; The arithmetic formula for the weighted probability vector (P0, P1) of sample x for different categories is as follows: ; Where M is the number of paths, t represents two categories of healthy or diseased, t = 0 or 1, represents the class distribution probability of samples in the leaf nodes under the m-th path, that is, the class distribution probability of samples in the leaf nodes in the corresponding clear decision tree; The predicted label t of sample x is calculated by the maximum weighted probability formula: Predicted label \(t = \max\) t P t ; In the formula, t = 0 or 1 represents the two categories of healthy and diseased.
8. The early Parkinson's interpretable three-branch intelligent evaluation method for multi-task remote data according to claim 1, characterized in that During the implementation of step S5, it includes the following steps: Step S51: After the sample is input into the fuzzy decision tree, calculate the fuzzy dividing point membership degree: According to the obtained fuzzy decision tree, for any fuzzy dividing point, use the triangular fuzzy membership function to calculate the fuzzy membership degree of the sample on this fuzzy dividing point; Step S52: Calculate the path weight: Aggregate the comprehensive membership degree values of the sample on each path, and normalize them to obtain the weight value of the sample on each path; Step S53: Output the weighted category probability vector: Multiply the path weights by the category distribution probabilities of the samples in the leaf nodes, and then sum them up for the same category, so as to determine the probability vector of the samples for different categories. , where represents the healthy probability, represents the diseased probability.
9. The early Parkinson's interpretable three-branch intelligent evaluation method for multi-task remote data according to claim 1, wherein During the implementation of step S6, the three-way decision mechanism is: When then the determined result is output, that is, when and then the output result is: diseased, when and then the output result is: healthy; When the output result is: uncertain, and the disease evaluation value ; Among them, the value range of the confidence threshold ξ is: 0.5 < ξ < 1.
Citation Information
Patent Citations
Sports equipment safety assessment system based on machine learning
CN117689219A
Three-branch diagnosis and evaluation method and device based on multi-modal remote medical data
CN119337166A
Complex concept-based interpretable data representation learning method
CN119378598A
Method and apparatus for constructing decision model, computer device and storage device
WO2017215370A1
KR20240171709A