A device for predicting multiple chronic diseases based on dynamic task-attention balance

By introducing a dynamic task attention mechanism into the multi-chronic disease prediction model, the problems of task conflict and information limitation in multi-task learning are solved, feature balance and accuracy are improved, and the effect of multi-task learning is enhanced.

CN120108738BActive Publication Date: 2025-09-05ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510584840.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-05
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing multi-task learning in chronic disease prediction has conflicts between tasks and information limitations, resulting in insufficient accuracy. In particular, the hard parameter sharing and soft parameter sharing methods are limited in computing resources and storage overhead, and the expressive power of Embedding & MLP is limited.

Method used

The dynamic task attention mechanism (DTA) is introduced to perform soft selection of input features through attention activation units, calculate task relevance weights, dynamically adjust feature representation, and optimize the feature requirements of each chronic disease prediction task.

Benefits of technology

It achieves feature balance in multiple chronic disease prediction tasks, improves the accuracy of prediction for each type of chronic disease and the expressiveness of the model, and simultaneously improves the effect of multi-task learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108738B_ABST
    Figure CN120108738B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-chronic disease prediction device based on dynamic task attention balance, which belongs to the field of smart medical technology, including: obtaining electronic medical case data and initializing multi-information features in the data; introducing a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model, which can perform soft selection of input multi-information features related to a single chronic disease prediction target task through an attention activation unit, focus on features related to the target task, and calculate corresponding activation weights, so that features with higher correlation obtain greater weights, and dominate the feature representation of each sub-expert network in the input model, thereby realizing input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation, which can more effectively adapt to the feature requirements of different chronic disease prediction tasks, thereby balancing all chronic disease prediction tasks and simultaneously improving prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of chronic disease medical prevention, and in particular relates to a multi-chronic disease prediction device based on dynamic task-attention balance. Background Art

[0002] In chronic disease prediction, multi-task learning (MTL) models can be used to predict a patient's likelihood of developing different chronic diseases within a specific timeframe. Compared to single-task approaches that train independent models for each task, MTL offers several advantages. First, due to the presence of shared network layers, memory usage is significantly reduced. Second, it avoids recalculating features for each task in the shared layers, significantly improving inference speed. Most importantly, if multiple tasks share complementary information or act as regularizers, MTL can potentially improve the predictive performance of the overall model.

[0003] Multi-task learning can be implemented in three forms: hard parameter sharing, soft parameter sharing, and expert parameter sharing. Hard parameter sharing is a classic multi-task learning strategy. Its core idea is that all tasks share underlying network parameters, while top-level parameters are optimized independently to suit their respective task requirements. However, due to the complete sharing of underlying parameters, this approach can lead to potential conflicts between tasks, especially for tasks with weak correlations. This conflict limits the performance gains of hard parameter sharing in practical applications.

[0004] Unlike hard parameter sharing, soft parameter sharing does not force all tasks to use the same underlying parameters. Instead, it assigns a separate model to each task while allowing the models to share and learn information. Because soft parameter sharing does not rely on inter-task correlation, it performs better when optimizing low-correlation or unrelated tasks. However, the main drawback of this approach is that it requires additional computing resources for online inference, and storing the additional network parameters of multiple models incurs significant storage overhead.

[0005] To address the limitations of hard and soft parameter sharing, a mixture of experts (MoE) network was proposed. It combines multiple expert sub-networks with a gating mechanism to dynamically select the most appropriate expert for task optimization. Further improvements, such as the multi-gate mixture of experts (MMoE), introduce independent gating mechanisms for each task, allowing tasks to be learned more accurately from multiple expert networks.

[0006] All three forms of multi-task learning employ the "embedding and multi-layer perceptron (MLP)" paradigm: large-scale, sparse features are first mapped into a low-dimensional embedding space, and these embedding vectors are then concatenated and fed into the MLP to learn the nonlinear relationships between features. However, the expressive power of the embedding and MLP method is limited by the dimensionality of the input vectors, which has become a bottleneck in multi-task learning.

[0007] Furthermore, the complex and competitive relationship between multiple tasks may further exacerbate the limitations of input information. If all tasks are simply trained using the same input features, a seesaw problem may occur, whereby improved performance in one task may lead to decreased performance in other tasks. During training, each task extracts patterns from the input features and makes predictions, with the loss function measuring the difference between the predicted value and the true value. The loss of each task generates a corresponding gradient for the shared parameters, and these gradients compete with each other in the multi-task network. Due to the uniqueness of the input features, competition between different tasks is amplified: if the gradient of a task is much larger than that of other tasks, it dominates the update of the shared parameters, and the model tends to optimize that task, resulting in the neglect of important features of other tasks and affecting overall performance. Conversely, if the gradient of a task is too small, its influence is insufficient, making it difficult to fully learn. Summary of the Invention

[0008] In view of the above, the purpose of the present invention is to provide a multi-chronic disease prediction device based on dynamic task-attention balance, which introduces a new attention mechanism, namely the dynamic task-attention mechanism (DTA). The input multi-information features can more effectively adapt to the feature requirements of different chronic disease prediction tasks, thereby balancing all chronic disease prediction tasks and simultaneously improving the accuracy of prediction for each type of chronic disease.

[0009] To achieve the above-mentioned purpose of the invention, an embodiment provides a device for predicting multiple chronic diseases based on dynamic task-attention balance, comprising:

[0010] A data processing module, which is used to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing;

[0011] A model building module is used to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in a multi-task prediction model for multi-chronic disease prediction. The attention activation unit can softly select the input multi-information features related to a single chronic disease prediction target task, focusing on features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation;

[0012] A prediction module is applied, which is used to perform multi-chronic disease prediction tasks based on a multi-task prediction model that introduces a dynamic task attention mechanism layer.

[0013] Preferably, multiple information features in the electronic medical record data are concatenated and used as input features of the model.

[0014] Preferably, the multi-task prediction model adopts a hybrid expert model, a multi-gated hybrid expert model, a cross-task collaborative model, or a step-by-step hierarchical extraction model.

[0015] Preferably, when a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task is introduced into a hybrid expert model, a multi-gated hybrid expert model, a cross-task collaborative model, or a hierarchical extraction model, the task covariate corresponding to each target task is calculated based on the input multi-information features by each gate in the model. ;

[0016] In the attention activation unit of each dynamic task attention mechanism layer, the feature sequence K as the key vector and the feature sequence V as the value vector are constructed based on the multi-information features of the input, and the task covariates With characteristic sequence Perform matrix operations to generate attention activation weights ( ), and then use the attention activation weight ( ) The feature sequence The corresponding features of are weighted to obtain feature representations rich in task attention information.

[0017] The feature representation output by each dynamic task attention mechanism layer is input into each corresponding sub-expert network to predict the probability of each chronic disease, and then multiple sub-expert networks running simultaneously realize the synchronous prediction of the probability of multiple chronic diseases.

[0018] Preferably, each gate in the model uses a multi-layer perceptron to calculate the task covariate corresponding to each target task based on the multi-information features of the input .

[0019] Preferably, constructing a feature sequence K as a key vector and a feature sequence V as a value vector based on the input multi-information features includes:

[0020] The input multi-information features are mapped by feature operations with different parameters to obtain the feature sequence K as the key vector and the feature sequence V as the value vector. The dimension of the feature sequence K is different from that of the feature sequence V. The dimension of the feature sequence K is , is the number of features, is the feature dimension, and the dimension of the feature sequence V is consistent with the input multi-information feature.

[0021] Preferably, the input multi-information features are mapped by feature operations with different parameters to obtain a feature sequence K as a key vector and a feature sequence V as a value vector, including:

[0022] ;

[0023] ;

[0024] in, express Class information features, 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. Represents a splicing operation.

[0025] Preferably, the multi-task prediction model is trained to learn the optimal parameters before being applied. During training, a multi-label based cross entropy loss function and back propagation are used to update the parameters.

[0026] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the device is used to implement the multi-type chronic disease prediction steps of the multi-chronic disease prediction device:

[0027] S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing;

[0028] S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation;

[0029] S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

[0030] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the program implements the steps of predicting multiple chronic diseases of the above-mentioned multi-chronic disease prediction device:

[0031] S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing;

[0032] S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation;

[0033] S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] A dynamic task attention mechanism layer corresponding to each chronic disease prediction target task is introduced into the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to a single chronic disease prediction target task, focus on the features related to the target task, and calculate the corresponding attention activation weights, so that the features with higher correlation obtain greater weights, and dominate the feature representation of each sub-expert network in the input model to achieve input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation, so that the input multi-information features can more effectively adapt to the feature requirements of different chronic disease prediction tasks, thereby balancing all chronic disease prediction tasks and simultaneously improving the accuracy of prediction for each type of chronic disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 2 is a schematic structural diagram of a device for predicting multiple chronic diseases based on dynamic task-attention balance provided by an embodiment;

[0038] Figure 2 1 is a schematic diagram of a structure of introducing a dynamic task attention mechanism layer into a Mixed Expert Model (MOE) provided in an embodiment;

[0039] Figure 3 This is a schematic diagram of the structure and process of introducing a dynamic task attention mechanism layer into a multi-gate mixture of experts (MMOE) model provided by an embodiment;

[0040] Figure 4 This is a schematic diagram of the structure and process of introducing a dynamic task attention mechanism layer into the cross-task collaboration model (CGC) provided in the embodiment;

[0041] Figure 5 This is a schematic diagram of the structure and flow of introducing a dynamic task attention mechanism layer into a progressive layered extraction model (PLE) provided by an embodiment;

[0042] Figure 6 Schematic diagram of the structure and flow of the dynamic task attention mechanism layer provided by the embodiment;

[0043] Figure 7 This is a heat map of the output of the dynamic task attention mechanism layer for different multi-chronic disease prediction tasks provided by the embodiment;

[0044] Figure 8 This is the ratio of task covariates used as weights for each sub-expert network under different tasks for the mixed expert model and the gated mixed expert model provided in the embodiment. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0046] The inventive concept of the present invention is: in existing multi-task learning, the complex and competitive relationship between multiple tasks may further aggravate the limitations of input features, resulting in insufficient accuracy of multi-task learning when applied to multiple types of chronic disease prediction tasks. The embodiment of the present invention provides a multi-chronic disease prediction device based on dynamic task attention balance. The dynamic task attention mechanism layer introduced by the device can adaptively calculate the attention weights of different tasks according to the correlation between the tasks and the input features, thereby generating a dynamic input feature representation for each task.

[0047] like Figure 1 As shown, the multi-chronic disease prediction device 10 based on dynamic task attention balance provided by the embodiment includes a data processing module 11, a model building module 12, and an application prediction module 13. Through the mutual cooperation of these three modules, the accuracy of the prediction of each type of chronic disease is simultaneously improved on the basis of balancing the feature input representation of multiple types of chronic disease prediction tasks from the input features.

[0048] In this embodiment, the data processing module 11 is used to obtain electronic medical record data and initialize the multi-information features in the electronic medical record data through preprocessing. The specific electronic medical record data can be obtained from the multi-chronic disease variable joint dataset ChronicDataset, which contains 180,000 samples extracted from the medical system. The obtained electronic medical record data is preprocessed by screening and cleaning, and then initialized to obtain multi-information features including biometric features, clinical features, and demographic features. This data is for five chronic diseases: diabetes, hypertension, kidney disease, coronary heart disease, and cerebrovascular disease. Based on this data, five joint modeling multi-chronic disease prediction tasks are established.

[0049] In an embodiment, the model building module 12 is used to introduce a dynamic task attention mechanism (DTA) layer corresponding to each chronic disease prediction target task in a multi-task prediction model for multi-chronic disease prediction. The dynamic task attention mechanism layer can soft-select the input multi-information features related to a single chronic disease prediction target task through the attention activation unit, focusing on the features related to the target task, and calculating the corresponding attention activation weights, so that the features with higher relevance obtain greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation. Through the dynamic task attention mechanism, the input multi-information features can change with the task, thereby enhancing the expressive power of the model within a limited dimension and improving the effect of multi-task learning.

[0050] In the embodiment, the multi-task prediction model adopts a mixture of experts model (MOE), a multi-gated mixture of experts model (MMOE), a cross-task collaboration model (CGC), or a step-by-step layered extraction model (PLE), and introduces a DTA layer corresponding to each chronic disease prediction target task into these models, such as Figure 2 ,3,4,5. After the introduction of the DTA layer, the principle process of each model is as follows:

[0051] First, multiple information features are concatenated to form input features Input into the model, ,in, express Class information features, Represents a splicing operation;

[0052] Then each gate in the model is based on the input features Calculate the task covariates corresponding to each target task , specifically a multi-layer perceptron Calculating task covariates ,Right now ,in The dimension is , to represent the attention information of the task.

[0053] It should be noted that for the MOE model, Figure 2 As shown, since there is only one gate, the multilayer perceptron with different parameters in one gate is To calculate the task covariate for each sub-expert network , that is, the task covariates for the five sub-expert networks 、 、 、 、 For the MMOE model, CGC model, and PLE model, each sub-expert network contains a corresponding gate. At this time, each gate calculates the task covariate of each sub-expert network separately. ,like Figure 3 , Figure 4 , Figure 5 shown.

[0054] Next, if Figure 6 As shown, in the attention activation unit of the dynamic task attention mechanism layer corresponding to each target task, a feature sequence K as a key vector and a feature sequence V as a value vector are constructed based on the input multi-information features. Specifically, the input multi-information features are mapped by feature operations with different parameters to obtain the feature sequence K as a key vector and the feature sequence V as a value vector. The dimension of the feature sequence K is different from that of the feature sequence V. The dimension of the feature sequence K is , is the number of features, is the feature dimension. The dimension of the feature sequence V is consistent with the input multi-information feature and is expressed as:

[0055] ;

[0056] ;

[0057] in, express Class information features, 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. Represents a splicing operation.

[0058] Then the task covariate With characteristic sequence Perform matrix operations to generate attention activation weights ( ), and then use the attention activation weight ( ) The feature sequence The corresponding feature weights are added to obtain feature representations rich in task attention information, which can be expressed as follows:

[0059] ;

[0060] in, Indicates that the attention activation weight will be The corresponding feature weights, Represents a feature sequence Dimensions, Represents a feature representation rich in task attention information, which makes the feature representation input to each sub-expert network task-biased.

[0061] Finally, the feature representation output by each dynamic task attention mechanism layer is input into each corresponding sub-expert network to predict the probability of each chronic disease, and then multiple sub-expert networks running simultaneously realize the synchronous prediction of the probability of multiple chronic diseases.

[0062] By allowing different sub-expert networks to obtain different DTA outputs, they are biased towards different tasks, alleviating the problem of a small number of expert networks dominating the majority of tasks. Furthermore, because the DTA structure focuses on building an attention relationship between tasks and input features, it effectively compensates for the insensitivity of existing multi-task prediction models to feature differences.

[0063] The above multi-task prediction model is trained to learn the optimal parameters before being applied. During training, a multi-label based cross entropy loss function is used. and backpropagation to update the parameters, ;

[0064] in, Represents the chronic disease category index, represents the total number of chronic disease categories, represents the predicted disease risk probability output by the model, represents the true classification label of chronic diseases, Represents the sigmoid function.

[0065] The embodiment also uses the chronic disease database Chronic Dataset to compare the AUCs of the original MOE model, the MOE (DTA) model with DTA introduced into the original MOE model, the original MMOE model, the MMOE (DTA) model with DTA introduced into the original MMOE model, and the original PLE model, the PLE (DTA) model with DTA introduced into the original PLE model. Table 1 shows the AUCs of the multi-chronic disease prediction tasks.

[0066] Table 1 AUC improvement of different models using DTA layer

[0067]

[0068] It can be clearly seen from Table 1 that introducing DTA into any multi-task prediction model can achieve an AUC improvement of about 0.5%.

[0069] In the DTA layer, a weight matrix can be generated by calculating the attention activation weights of each chronic disease prediction task on the input features. Figure 7 As shown, this weight matrix can be represented as a heatmap, where the color depth indicates the magnitude of the weight. For example, darker colors (such as red) indicate larger weights and a higher relevance between the feature and the task; lighter colors (such as blue) indicate smaller weights and a lower relevance between the feature and the task. When predicting the probability of different chronic diseases, DTA can effectively capture the relationship between different tasks and input features. For example, the correlation between heart disease and blood pressure features is often greater than that between kidney disease and blood pressure features.

[0070] DTA softly selects features through an attention mechanism, retaining only those relevant to the current chronic disease prediction task while suppressing the influence of irrelevant features. This generates dynamic inputs for each chronic disease prediction task. For example, for a diabetes prediction task, blood sugar-related features may be given a higher weight, while in a kidney disease prediction task, features related to renal function may be more important. This dynamic adjustment enables the model to flexibly adjust the input representation based on task requirements, making the input vector better suited to the current task. This adaptability enables the model to more effectively utilize features within a limited dimension, avoiding the information redundancy or insufficiency that may result from a fixed input representation, thereby improving the effectiveness of multi-task learning.

[0071] Figure 8 Given the task covariate ratio, the MMOE model can experience a situation where a small number of sub-expert networks monopolize the majority of tasks (load imbalance). In this case, the model degenerates from an MMOE to a single-expert network structure. However, adding the DTA layer can significantly alleviate this monopolization issue. For PLE, the weights of different tasks on different expert networks vary significantly. However, due to its structure, a single expert network cannot provide information to other tasks, resulting in a certain amount of information loss. This explains why the PLE(DTA) model achieves better training results than the PLE model.

[0072] In summary, a DTA layer is introduced into a multi-task prediction model for multi-chronic disease prediction. Training is performed based on a joint dataset of multiple chronic disease variables, integrating the biological, clinical, and demographic characteristics of multiple chronic diseases. The model is optimized using approximate gradient descent, and model performance is evaluated using multiple metrics such as consistency index and area under curve (AUC). By defining multiple risk indicators (such as baseline risk and relative risk), the model can comprehensively assess a patient's risk of disease. This implementation not only fully leverages the advantages of multi-task learning but also ensures model accuracy and generalization through optimization strategies and evaluation metrics. This approach enables efficient and accurate predictions in chronic disease prediction scenarios, providing strong support for clinical decision-making.

[0073] Based on the same inventive concept, an embodiment further provides a computing device comprising a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, the device is configured to implement the multi-type chronic disease prediction steps of the multi-chronic disease prediction apparatus described above:

[0074] S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing;

[0075] S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation;

[0076] S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

[0077] The computing device provided in the embodiment, at the hardware level, includes not only a processor and memory, but also hardware required for other services such as an internal bus, a network interface, and memory. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the multiple chronic disease prediction steps described in S1-S3 above. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0078] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the program implements the multi-type chronic disease prediction steps of the multi-chronic disease prediction device described above:

[0079] S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing;

[0080] S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation;

[0081] S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

[0082] In the embodiment, computer-readable media includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.

[0083] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multi-chronic disease prediction device based on dynamic task attention balance, characterized in that: include: A data processing module, which is used to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing; A model building module is used to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in a multi-task prediction model for multi-chronic disease prediction. The attention activation unit can softly select the input multi-information features related to a single chronic disease prediction target task, focusing on features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation; Among them, the multi-task prediction model adopts a hybrid expert model, a multi-gated hybrid expert model, a cross-task collaborative model, or a step-by-step hierarchical extraction model. In these models, when the dynamic task attention mechanism layer corresponding to each chronic disease prediction target task is introduced, each gate in the model calculates the task covariate corresponding to each target task based on the input multi-information features. In the attention activation unit of each dynamic task attention mechanism layer, the feature sequence K as the key vector and the feature sequence V as the value vector are constructed based on the input multi-information features, and the task covariate With characteristic sequence Perform matrix operations to generate attention activation weights ( ), and then use the attention activation weight ( ) The feature sequence The corresponding features are weighted to obtain a feature representation rich in task attention information; the feature representation output by each dynamic task attention mechanism layer is input into each corresponding sub-expert network to predict the probability of each chronic disease, and then multiple sub-expert networks running simultaneously achieve the simultaneous prediction of the probability of multiple chronic diseases; A prediction module is applied, which is used to perform multi-chronic disease prediction tasks based on a multi-task prediction model that introduces a dynamic task attention mechanism layer.

2. The multi-chronic disease prediction device based on dynamic task-attention balance according to claim 1, characterized in that: The multiple information features in the electronic medical record data are concatenated and used as the input features of the model.

3. The multi-chronic disease prediction device based on dynamic task-attention balance according to claim 1, characterized in that: Each gate in the model uses a multi-layer perceptron to calculate the task covariate corresponding to each target task based on the multi-information features of the input .

4. The multi-chronic disease prediction device based on dynamic task-attention balance according to claim 1, characterized in that: Based on the input multi-information features, a feature sequence K as a key vector and a feature sequence V as a value vector are constructed, including: The input multi-information features are mapped by feature operations with different parameters to obtain the feature sequence K as the key vector and the feature sequence V as the value vector. The dimension of the feature sequence K is different from that of the feature sequence V. The dimension of the feature sequence K is , is the number of features, is the feature dimension, and the dimension of the feature sequence V is consistent with the input multi-information feature.

5. The multi-chronic disease prediction device based on dynamic task-attention balance according to claim 4 is characterized in that: The input multi-information features are mapped by feature operations with different parameters to obtain the feature sequence K as the key vector and the feature sequence V as the value vector, including: ; ; in, express Class information features, 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. 、 、 、 Indicates that in calculation Time Target The class information features correspond to the multi-layer perceptron mappings with different parameters. Represents a splicing operation.

6. The multi-chronic disease prediction device based on dynamic task-attention balance according to claim 3, characterized in that: The multi-task prediction model is trained to learn the optimal parameters before being applied. During training, a multi-label based cross entropy loss function is used. and backpropagation to update the parameters.

7. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, the processors are used to implement the steps of predicting multiple types of chronic diseases using the multi-chronic disease prediction device according to any one of claims 1 to 6: S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing; S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation; Among them, the multi-task prediction model adopts a hybrid expert model, a multi-gated hybrid expert model, a cross-task collaborative model, or a step-by-step hierarchical extraction model. In these models, when the dynamic task attention mechanism layer corresponding to each chronic disease prediction target task is introduced, each gate in the model calculates the task covariate corresponding to each target task based on the input multi-information features. In the attention activation unit of each dynamic task attention mechanism layer, the feature sequence K as the key vector and the feature sequence V as the value vector are constructed based on the input multi-information features, and the task covariate With characteristic sequence Perform matrix operations to generate attention activation weights ( ), and then use the attention activation weight ( ) The feature sequence The corresponding features are weighted to obtain a feature representation rich in task attention information; the feature representation output by each dynamic task attention mechanism layer is input into each corresponding sub-expert network to predict the probability of each chronic disease, and then multiple sub-expert networks running simultaneously achieve the synchronous prediction of the probability of multiple chronic diseases; S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

8. A computer-readable storage medium, characterized in that A program is stored thereon, which, when executed by a processor, is used to implement the steps of predicting multiple types of chronic diseases using the multi-chronic disease prediction device according to any one of claims 1 to 6: S1, using the data processing module to obtain electronic medical data and initialize multiple information features in the electronic medical data through preprocessing; S2 uses the model building module to introduce a dynamic task attention mechanism layer corresponding to each chronic disease prediction target task in the multi-task prediction model for multi-chronic disease prediction. Through the attention activation unit, it can soft-select the input multi-information features related to the single chronic disease prediction target task, focusing on the features related to the target task and calculating the corresponding attention activation weights. Features with higher relevance receive greater weights and dominate the feature representation of each sub-expert network in the input model, achieving input feature balance at the task level. The model predicts the probability of multiple chronic diseases based on the input feature representation; Among them, the multi-task prediction model adopts a hybrid expert model, a multi-gated hybrid expert model, a cross-task collaborative model, or a step-by-step hierarchical extraction model. In these models, when the dynamic task attention mechanism layer corresponding to each chronic disease prediction target task is introduced, each gate in the model calculates the task covariate corresponding to each target task based on the input multi-information features. In the attention activation unit of each dynamic task attention mechanism layer, the feature sequence K as the key vector and the feature sequence V as the value vector are constructed based on the input multi-information features, and the task covariate With characteristic sequence Perform matrix operations to generate attention activation weights ( ), and then use the attention activation weight ( ) The feature sequence The corresponding features are weighted to obtain a feature representation rich in task attention information; the feature representation output by each dynamic task attention mechanism layer is input into each corresponding sub-expert network to predict the probability of each chronic disease, and then multiple sub-expert networks running simultaneously achieve the synchronous prediction of the probability of multiple chronic diseases; S3, using the application prediction module to perform multi-chronic disease prediction tasks based on a hybrid expert model with a dynamic task attention mechanism layer, a multi-gated hybrid expert model, or a step-by-step hierarchical extraction model.

Citation Information

Patent Citations

  • Risk prediction method for chronic diseases and related equipment

    CN115862842A

  • Method and system for evaluating condition of sepsis patient based on multi-modal data fusion

    CN118280579A