Diabetes prediction system and method fusing multiple models

By fusing multiple models of diabetes prediction system, multi-scale convolutional neural network and Transformer encoder extract features, and dynamically selecting features through TabNet, the problem of difficult feature interaction relationships in traditional methods is solved, achieving higher prediction accuracy and model interpretability, supporting early diabetes diagnosis and intervention.

CN119943405APending Publication Date: 2025-05-06ANHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510092121.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional diabetes prediction methods rely on statistical models and traditional machine learning algorithms, and it is difficult to effectively capture the complex interaction between features, resulting in feature selection relying on manual and insufficient prediction accuracy.

Method used

A diabetes prediction system that integrates multiple models is adopted, including data preprocessing, feature extraction, feature transformation and splicing, and TabNet dynamic feature selection module. Local and global features are extracted using multi-scale convolutional neural network and Transformer encoder, and important features are automatically selected through TabNet, and finally classified through Softmax layer.

Benefits of technology

It improves the accuracy and generalization ability of diabetes prediction, can more accurately predict the risk of diabetes in individuals, enhances the interpretability and robustness of the model, and provides a basis for early intervention and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943405A_ABST
    Figure CN119943405A_ABST
Patent Text Reader

Abstract

The invention discloses a diabetes prediction system and method fusing multiple models. The method comprises the following steps: inputting preprocessed two-dimensional structured physical examination data into the diabetes prediction system fusing multiple models; extracting local features through convolution kernels of different sizes of a plurality of convolution layers in the CNN module; a self-attention mechanism in a Transform module is utilized to capture a long-distance dependency relationship between the features; the features extracted by the CNN module and the Transform module are spliced and transformed; and carrying out feature selection and model interpretation. According to the method, the features can be automatically extracted, the complex relation between the features is effectively captured, the generalization ability is high, and the accuracy of diabetes prediction is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data processing, and in particular to a diabetes prediction system and method integrating multiple models. Background Art

[0002] Diabetes is a global chronic disease, and early prediction and intervention are crucial to improving patients' quality of life and reducing medical burdens. Traditional diabetes prediction methods rely on statistical models and traditional machine learning algorithms, which have limitations such as manual feature selection and difficulty in capturing complex interactions between features. Summary of the invention

[0003] The purpose of the present invention is to provide a diabetes prediction system and method that integrates multiple models, which can automatically extract features, effectively capture the complex relationships between features, and have strong generalization capabilities, thereby greatly improving the accuracy of diabetes prediction.

[0004] In order to achieve the above object, the present invention provides a diabetes prediction system integrating multiple models, the system comprising:

[0005] A data preprocessing module, used for receiving and processing structured physical examination data;

[0006] Feature extraction module, used for local feature extraction and global feature extraction;

[0007] Feature transformation and splicing module, used to unify feature dimensions and perform splicing;

[0008] TabNet dynamic feature selection module, used to automatically select important features;

[0009] Feature accumulation and classification module, used to output diabetes prediction results.

[0010] Preferably, the feature extraction module includes a multi-scale convolutional neural network CNN and a Transformer encoder structure, wherein the multi-scale convolutional neural network CNN is used for local feature extraction and the Transformer encoder structure is used for global feature extraction.

[0011] Preferably, the feature transformation and concatenation module includes linear transformation and feature concatenation, wherein the linear transformation is used to unify the feature dimension and the feature concatenation is used to form a comprehensive feature representation.

[0012] Preferably, the TabNet dynamic feature selection module includes Attentive Transformer and Feature Transformer, wherein Attentive Transformer is used for feature selection and Feature Transformer is used for feature transformation.

[0013] Preferably, the feature accumulation and classification module comprises a feature accumulation layer and a Softmax layer, wherein the feature accumulation layer is used for accumulating feature outputs, and the Softmax layer is used for classifying outputs.

[0014] Another aspect of the present invention provides a diabetes prediction method integrating multiple models. The diabetes prediction method integrating multiple models utilizes the diabetes prediction system integrating multiple models as described above to perform diabetes prediction.

[0015] Preferably, the diabetes prediction method integrating multiple models comprises:

[0016] The pre-processed two-dimensional structured physical examination data is input into the diabetes prediction system integrating multiple models;

[0017] Extract local features through convolution kernels of different sizes in multiple convolution layers in the CNN module;

[0018] Use the self-attention mechanism in the Transformer module to capture long-distance dependencies between features;

[0019] Concatenate and transform the features extracted by the CNN module and the Transformer module;

[0020] Perform feature selection and model interpretation.

[0021] Preferably, the local feature extraction includes: the CNN module uses a parallel structure of multi-scale convolution kernels to capture local patterns of different scales, and at the same time, uses a ReLU activation function and a pooling layer to process feature maps and unify the dimensions of feature maps of different convolution layers through linear transformation;

[0022] Use the convolution layer to convolve the input features to obtain a feature map, and process the feature map through the ReLU activation function so that the elements in the feature map greater than 0 remain unchanged, and the elements less than 0 are set to 0;

[0023] Use maximum pooling to process the feature map after the convolution layer, slide a pooling window on the feature map and take the maximum value in each window as the output of the window;

[0024] The linear transformation layer includes a learnable weight matrix and bias term, which can map the feature maps corresponding to convolution kernels of different sizes to the same dimensional space;

[0025] Global feature extraction includes: using the Transformer module to capture the dependencies between features through self-attention mechanism and multi-head self-attention mechanism, and capturing complex nonlinear relationships in the feedforward neural network.

[0026] Preferably, the features are concatenated to form a new feature vector, which integrates local and global information; at the same time, a linear transformation is used in the TabNet module to map the input features to a new space.

[0027] Preferably, feature selection includes: using Attentive Transformer to generate a sparse attention mask to determine which features should be retained in the current step; using Feature Transformer to further nonlinearly transform and process the selected features to extract higher-level feature representations, wherein:

[0028] The Feature Transformer consists of a feedforward neural network FFN, which includes a shared part across decision steps and an independent part related to a specific decision step. The fully connected layer FC and gated linear unit GLU of the shared part are used to perform linear transformation and gating processing on the input features. At the same time, the independent part is used to provide a unique feature transformation for each decision step.

[0029] According to the above technical scheme, the present invention first performs data preprocessing, including data standardization and missing value processing, to ensure the consistency and integrity of the input data; secondly, the data is input into the CNN and Transformer modules in parallel to extract local features and global features respectively; then, the extracted features are linearly transformed to unify their dimensions, and the local and global features are spliced ​​to form a comprehensive feature representation; then, TabNet automatically selects the most important features for diabetes risk prediction through a dynamic feature selection module, while retaining the interpretability of the model; finally, the output after feature transformation is accumulated through a feature accumulation layer and passed to the next decision step, and the accumulated features are classified through a Softmax layer to output the results of diabetes prediction. This method can more accurately predict the risk of an individual suffering from diabetes, thereby helping healthcare professionals to intervene and treat early.

[0030] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present invention but do not constitute a limitation of the present invention. In the accompanying drawings:

[0032] Figure 1 This is an overall architecture diagram of a diabetes prediction system integrating multiple models provided by the present invention;

[0033] Figure 2It is a flowchart of local feature extraction in the diabetes prediction method integrating multiple models provided by the present invention;

[0034] Figure 3 The present invention provides a flowchart of global feature extraction in a diabetes prediction method integrating multiple models. DETAILED DESCRIPTION

[0035] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0036] See also Figure 1 On the one hand, the present invention provides a diabetes prediction system integrating multiple models, the system comprising:

[0037] A data preprocessing module, used for receiving and processing structured physical examination data;

[0038] Feature extraction module, used for local feature extraction and global feature extraction;

[0039] Feature transformation and splicing module, used to unify feature dimensions and perform splicing;

[0040] TabNet dynamic feature selection module, used to automatically select important features;

[0041] Feature accumulation and classification module, used to output diabetes prediction results.

[0042] Specifically, in this embodiment, preferably, the above-mentioned feature extraction module includes a multi-scale convolutional neural network CNN and a Transformer encoder structure, wherein the multi-scale convolutional neural network CNN is used for local feature extraction, and the Transformer encoder structure is used for global feature extraction.

[0043] The above-mentioned feature transformation and splicing module includes linear transformation and feature splicing, wherein linear transformation is used to unify feature dimensions, and feature splicing is used to form a comprehensive feature representation.

[0044] The TabNet dynamic feature selection module includes Attentive Transformer and Feature Transformer, where Attentive Transformer is used for feature selection and Feature Transformer is used for feature transformation.

[0045] In this embodiment, the feature accumulation and classification module includes a feature accumulation layer and a Softmax layer, wherein the feature accumulation layer is used to accumulate feature outputs, and the Softmax layer is used to classify outputs.

[0046] In addition, another aspect of the present invention provides a diabetes prediction method integrating multiple models. The diabetes prediction method integrating multiple models utilizes the above-mentioned diabetes prediction system integrating multiple models to perform diabetes prediction.

[0047] Specifically, the diabetes prediction method integrating multiple models includes:

[0048] The pre-processed two-dimensional structured physical examination data is input into the diabetes prediction system integrating multiple models;

[0049] Extract local features through convolution kernels of different sizes in multiple convolution layers in the CNN module;

[0050] Use the self-attention mechanism in the Transformer module to capture long-distance dependencies between features;

[0051] Concatenate and transform the features extracted by the CNN module and the Transformer module;

[0052] Perform feature selection and model interpretation.

[0053] Among them, the above-mentioned local feature extraction includes: the CNN module uses a parallel structure of multi-scale convolution kernels to capture local patterns of different scales, and at the same time, uses the ReLU activation function and the pooling layer to process the feature map and unify the feature map dimensions of different convolution layers through linear transformation. Specifically, after the convolution layer performs a convolution operation on the input feature, a convolutional feature map is obtained. These feature maps are processed by the ReLU activation function so that the elements greater than 0 in the feature map remain unchanged and the elements less than 0 are set to 0. In this way, some unimportant information can be removed, useful features are retained, and the distribution of the feature map is made sparser, which is conducive to subsequent feature learning and model training; the feature map after the convolution layer will be processed by the pooling layer, which uses the maximum pooling. The basic principle is to slide a pooling window (such as 2×2) on the feature map and take the maximum value in each window as the output of the window. For example, for a feature map of size n×n, after being processed by a pooling window of size 2×2, the size of the feature map will become n / 2×n / 2. In this way, the pooling layer can effectively reduce the spatial dimension of the feature map, reduce the amount of subsequent calculations, and retain the main feature information in the feature map, such as edges, corners, etc., which helps to improve the robustness of the model to the input data; the feature maps after convolution operations and pooling processing with convolution kernels of different sizes (1×3, 1×5, 1×7) will be unified through linear transformation to unify their dimensions. Specifically, for each feature map output by the convolution layer, it will pass through a linear transformation layer, which contains a learnable weight matrix and bias terms. Through this linear transformation, the feature maps corresponding to convolution kernels of different sizes are mapped to the same dimensional space. For example, assuming that the dimensions of the feature maps after convolution and pooling are d1, d2, and d3 respectively, after linear transformation, they are uniformly mapped to a space of dimension d, so that these feature maps can be spliced ​​in the feature dimension to obtain a comprehensive feature representation. In this way, the model can effectively integrate features from convolution layers of different scales and provide more comprehensive feature input for subsequent global feature extraction and classification tasks.

[0054] Global feature extraction includes: using the Transformer module, capturing the dependencies between features through the self-attention mechanism and the multi-head self-attention mechanism, and capturing complex nonlinear relationships in the feedforward neural network. Specifically, the input features are first processed by the embedding layer, mapped to a high-dimensional space, and position encoding is added to retain the relative position information between the features. Then, through the self-attention mechanism, the model calculates the dot product between the query, key, and value matrices to obtain the attention weights. These weights reflect the importance relationship between the input features, that is, the influence of certain features on other features; the input features are divided into multiple heads, each with a dimension of dk / h, where h is the number of heads. Each head independently calculates self-attention to obtain its own weighted feature representation. Then, the outputs of these heads are concatenated and integrated through a linear transformation layer to obtain the final multi-head self-attention output; the output of the multi-head self-attention mechanism is processed by a feedforward neural network. Specifically, the feedforward neural network contains two linear transformation layers, and the ReLU activation function is used in the middle to introduce nonlinearity.

[0055] In the above process, features are concatenated to form a new feature vector that integrates local and global information; at the same time, linear transformation is used in the TabNet module to map the input features to a new space.

[0056] In this embodiment, the preferred feature selection includes: using Attentive Transformer to generate a sparse attention mask to determine which features should be retained in the current step; using Feature Transformer to further nonlinearly transform and process the selected features. Specifically, Feature Transformer is composed of a feed-forward neural network (FFN), which is specially designed to include a part shared across decision steps and an independent part related to a specific decision step. The shared part linearly transforms and gates the input features through a fully connected layer (FC) and a gated linear unit (GLU) to enhance the expressive power of the features and selectively retain important features. The independent part provides a unique feature transformation for each decision step to further refine the feature representation. These transformed features are accumulated through a feature accumulation layer, combined with the feature output of the previous step, to form a gradually enhanced feature representation, and finally provide input for the classification decision of the model. This dynamic feature selection and transformation mechanism not only improves the efficiency and interpretability of the model, but also enhances the flexibility and prediction performance of the model, providing a richer and more accurate feature representation for the diabetes prediction task.

[0057] Through the above technical solution, data preprocessing is first performed, including data standardization and missing value processing, to ensure the consistency and integrity of the input data; secondly, the data is input into the CNN and Transformer modules in parallel to extract local features and global features respectively; then, the extracted features are linearly transformed to unify their dimensions, and the local and global features are spliced ​​to form a comprehensive feature representation; then, TabNet automatically selects the most important features for diabetes risk prediction through the dynamic feature selection module, while retaining the interpretability of the model; finally, the output after feature transformation is accumulated through the feature accumulation layer and passed to the next decision step, and the accumulated features are classified through the Softmax layer to output the results of diabetes prediction. This method can more accurately predict the risk of an individual developing diabetes, thereby helping healthcare professionals to intervene and treat early.

[0058] A specific embodiment is provided below to illustrate the present invention:

[0059] 1. Data preparation:

[0060] The dataset contains the following characteristics of patients: age, gender, height, weight, blood pressure, blood sugar, cholesterol, triglycerides, uric acid, liver function indicators (such as ALT, AST), kidney function indicators (such as creatinine, urea nitrogen), etc., a total of 20 features. In addition, the dataset also contains a label to indicate whether the patient has diabetes (1 means the patient has the disease, 0 means the patient does not have the disease).

[0061] 2. Data preprocessing:

[0062] 1. Missing value processing: In the data set, some indicators of some patients (such as triglycerides and uric acid) have missing values. For numerical features, the median filling method is used, that is, the median of the feature is used to fill the missing values ​​to reduce the impact of extreme values ​​on the interpolation results; for categorical features (such as gender), the mode filling method is used for filling;

[0063] 2. Data standardization: Since the dimensions and value ranges of different features vary greatly, in order to enable the model to better learn the relationship between features, all numerical features are subjected to Z-score standardization, that is, the mean of the feature is subtracted and divided by its standard deviation, so that the mean of the feature is 0 and the standard deviation is 1;

[0064] 3. Outlier processing: Through box plot analysis, it was found that some patients had abnormal values ​​in indicators such as blood pressure and blood sugar. The Z-score method was used to regard points with an absolute value of Z-score greater than 3 as outliers and replace them with the mean of the feature.

[0065] 3. Model training and prediction:

[0066] 1. Model architecture construction: building a multi-module integrated diabetes prediction model based on CNN, Transformer and TabNet:

[0067] CNN module: Input preprocessed physical examination data, extract local features through three convolutional layers using 1x3, 1x5, and 1x7 convolution kernels respectively, and each convolutional layer is followed by a ReLU activation function and a maximum pooling layer to reduce feature dimensions and enhance feature expression capabilities. Finally, features of different scales are mapped to the same dimension through linear transformation.

[0068] Transformer module: embeds input features into a high-dimensional space and adds positional encoding to consider the relative positional relationship between features; captures long-range dependencies between features through the self-attention mechanism. The multi-head self-attention mechanism models feature dependencies from multiple angles. After residual connection and layer normalization, the features are further processed through a feedforward neural network.

[0069] TabNet module: It consists of multiple decision steps, each of which includes feature selection (AttentiveTransformer) and feature transformation (Feature Transformer). Attentive Transformer generates sparse attention masks and selects key features; Feature Transformer performs nonlinear transformation on the selected features, and finally aggregates the features through the feature accumulation layer to provide input for classification.

[0070] 2. Model training: The preprocessed dataset is divided into a training set (80%) and a test set (20%). The AdamW optimizer is used, the initial learning rate is 0.0005, the number of training rounds is 50, and the batch size is 64. The early stopping strategy is adopted, and the training is stopped when the validation set loss no longer decreases for 5 consecutive epochs to prevent overfitting.

[0071] 3. Model prediction: After training, the test set data is used for prediction. The test set data is input into the model. After being processed by the CNN, Transformer and TabNet modules, the probability of each sample suffering from diabetes is output through the Softmax layer.

[0072] 4. Technical effects and improvements:

[0073] 1. Performance improvement: On the test set, the model's accuracy reached 0.8817 and AUC was 0.8845. Compared with traditional statistical models and single machine learning algorithms, such as logistic regression (0.7754) and support vector machine (0.8453), the performance was significantly improved, indicating that the model can more accurately predict the occurrence of diabetes.

[0074] 2. Feature extraction and selection: The CNN module automatically extracts local features from physical examination data through multi-scale convolution kernels, such as subtle change patterns captured from indicators such as blood sugar and blood pressure; the Transformer module captures long-range dependencies between features, such as the potential correlation between age and renal function indicators; the TabNet module dynamically selects key features through a sparse attention mechanism, reducing the interference of redundant features and improving the generalization ability of the model.

[0075] 3. Enhanced interpretability: The sparse attention mask of the TabNet module provides interpretability for the model, which can clearly indicate which features play an important role in the prediction process. For example, when predicting whether a patient has diabetes, the model may find that features such as blood sugar, BMI, and family history have higher weights, thus providing doctors with valuable diagnostic evidence.

[0076] 4. Improved robustness: The model shows good performance on different data sets (such as the Pima data set), and when faced with noise and outliers in the data, it can maintain a high prediction accuracy through data preprocessing and model structure design, showing strong robustness.

[0077] In summary, through the above specific examples, it can be seen that the model has not only achieved excellent performance in the diabetes prediction task, but also made significant improvements in feature extraction, selection and model interpretability, providing strong technical support for the early diagnosis and prevention of diabetes.

[0078] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0079] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0080] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0082] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0083] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0084] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0085] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0086] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A diabetes prediction system integrating multiple models, characterized in that: The system comprises: A data preprocessing module, used for receiving and processing structured physical examination data; Feature extraction module, used for local feature extraction and global feature extraction; Feature transformation and splicing module, used to unify feature dimensions and perform splicing; TabNet dynamic feature selection module, used to automatically select important features; Feature accumulation and classification module, used to output diabetes prediction results.

2. The diabetes prediction system integrating multiple models according to claim 1, characterized in that: The feature extraction module includes a multi-scale convolutional neural network (CNN) and a Transformer encoder structure, wherein the multi-scale convolutional neural network (CNN) is used for local feature extraction and the Transformer encoder structure is used for global feature extraction.

3. The diabetes prediction system integrating multiple models according to claim 2, characterized in that: The feature transformation and splicing module includes linear transformation and feature splicing, where linear transformation is used to unify feature dimensions and feature splicing is used to form a comprehensive feature representation.

4. The diabetes prediction system integrating multiple models according to claim 3, characterized in that: The TabNet dynamic feature selection module includes Attentive Transformer and Feature Transformer, where Attentive Transformer is used for feature selection and Feature Transformer is used for feature transformation.

5. The diabetes prediction system integrating multiple models according to claim 4, characterized in that: The feature accumulation and classification module includes a feature accumulation layer and a Softmax layer, wherein the feature accumulation layer is used to accumulate feature outputs, and the Softmax layer is used for classification outputs.

6. A diabetes prediction method integrating multiple models, characterized in that: The diabetes prediction method integrating multiple models utilizes the diabetes prediction system integrating multiple models as described in claim 5 to perform diabetes prediction.

7. The diabetes prediction method integrating multiple models according to claim 6, characterized in that: The diabetes prediction method integrating multiple models includes: The pre-processed two-dimensional structured physical examination data is input into the diabetes prediction system integrating multiple models; Extract local features through convolution kernels of different sizes in multiple convolution layers in the CNN module; Use the self-attention mechanism in the Transformer module to capture long-distance dependencies between features; Concatenate and transform the features extracted by the CNN module and the Transformer module; Perform feature selection and model interpretation.

8. The diabetes prediction method integrating multiple models according to claim 7, characterized in that: Local feature extraction includes: the CNN module uses a parallel structure of multi-scale convolution kernels to capture local patterns of different scales. At the same time, the ReLU activation function and pooling layer are used to process feature maps and unify the dimensions of feature maps of different convolution layers through linear transformation: Use the convolution layer to convolve the input features to obtain a feature map, and process the feature map through the ReLU activation function so that the elements in the feature map greater than 0 remain unchanged, and the elements less than 0 are set to 0; Use maximum pooling to process the feature map after the convolution layer, slide a pooling window on the feature map and take the maximum value in each window as the output of the window; The linear transformation layer includes a learnable weight matrix and bias term, which can map the feature maps corresponding to convolution kernels of different sizes to the same dimensional space; Global feature extraction includes: using the Transformer module to capture the dependencies between features through self-attention mechanism and multi-head self-attention mechanism, and capturing complex nonlinear relationships in the feedforward neural network.

9. The diabetes prediction method integrating multiple models according to claim 7, characterized in that: The features are concatenated to form a new feature vector, which integrates local and global information; at the same time, a linear transformation is used in the TabNet module to map the input features to a new space.

10. The diabetes prediction method integrating multiple models according to claim 7, characterized in that: Feature selection includes: using Attentive Transformer to generate sparse attention masks to determine which features should be retained in the current step; using Feature Transformer to further nonlinearly transform and process the selected features to extract higher-level feature representations, where: The Feature Transformer consists of a feedforward neural network FFN, which includes a shared part across decision steps and an independent part related to a specific decision step. The fully connected layer FC and gated linear unit GLU of the shared part are used to perform linear transformation and gating processing on the input features. At the same time, the independent part is used to provide a unique feature transformation for each decision step.