Disease risk prediction method and device, training method and device, electronic equipment and medium
Through multimodal data fusion and knowledge distillation technology, global feature representations are generated, which are used to train regression prediction models, solving the problems of one-sided, low accuracy and inefficiency in traditional disease diagnosis and prediction, and achieving a more comprehensive and accurate disease risk assessment.
Patent Information
- Application Number
- CN202510150168.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional disease diagnosis and prediction have problems such as one-sided, low accuracy and inefficiency in disease risk prediction. Especially when doctors need to quickly process multiple data, timeliness cannot meet the needs of rapid disease risk prediction.
A disease risk prediction method for multimodal data fusion is proposed. Through the feature extraction and knowledge distillation technology of medical imaging, physiological signals and gene sequence data, global feature representation is generated, which is used to train a regression prediction model to achieve comprehensive, accurate and efficient prediction of multimodal data.
Through multimodal data fusion, the accuracy and efficiency of disease risk prediction are improved, the robustness of the model to data noise is enhanced, the problem of insufficient single-modal data is compensated, and a more comprehensive disease risk assessment is achieved.
Smart Images

Figure CN120089363A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular, to a method and device for predicting disease risk, a training method, an electronic device, and a medium. Background Art
[0002] Traditional disease diagnosis and prediction mainly rely on doctors to judge the lesions, disease causes, etc. of patients through the patient's medical imaging data (such as X-ray, CT, MRI and other image data), gene sequence data (such as RNA-seq, DNA-seq and other sequencing data), physiological signal data (such as electrocardiogram ECG, electroencephalogram EEG, etc.), and issue corresponding diagnosis and treatment plans. On the one hand, due to experience guidance or deviation, doctors are prone to focus on one type of data and ignore the influence of other data, resulting in low accuracy in disease risk prediction and diagnosis. On the other hand, when doctors comprehensively judge by combining medical imaging data, gene sequence data and physiological signal data, the timeliness cannot meet the needs of rapid disease risk prediction and diagnosis. Summary of the Invention
[0003] The present application provides a method and device for predicting disease risk, a training method, an electronic device, and a medium, which can solve the problems of one-sided, low accuracy, and inefficient prediction of disease risk.
[0004] To achieve the above object, the present application adopts the following technical solutions:
[0005] In a first aspect, a training method for a disease risk prediction model is provided. The disease risk prediction model is a regression prediction model, and the training method includes:
[0006] Obtaining training data; and
[0007] Training the regression prediction model by using the training data.
[0008] Wherein, obtaining training data specifically includes:
[0009] Obtaining data to be processed, where the data to be processed includes medical imaging data, physiological signal data, and gene sequence data. Among them, the medical imaging data includes first data and a first annotation representing the true value of the disease from the medical imaging aspect, the gene sequence data includes second data and a second annotation representing the true value of the disease from the gene sequence aspect, and the physiological signal includes third data and a third annotation representing the true value of the disease from the physiological signal aspect;
[0010] Processing to obtain a first feature vector corresponding to the first data, a second feature vector corresponding to the second data, and a third feature vector corresponding to the third data;
[0011] Using the first feature vector as the knowledge to guide the second and third feature vectors, the knowledge distillation technique is adopted to make the expressions of the second and third feature vectors consistent with that of the first feature vector; and
[0012] Process the first, second, and third feature vectors with consistent expressions to obtain a globally feature representation after feature weighted fusion,
[0013] The training data includes: the globally feature representation as the input, the disease prediction value corresponding to the globally feature representation as the output, and the true disease value.
[0014] Based on the above technical solution, in the process of obtaining training data, using the annotation of the multi-modal data to be processed as the true disease value, processing the multi-modal data to be processed to obtain the corresponding feature vectors, using the feature vector corresponding to the medical image as the knowledge to guide the feature vectors corresponding to other modalities, adopting the knowledge distillation technique to make the expressions of the feature vectors consistent, and then obtaining a globally feature representation after feature weighted fusion as the input of the regression prediction model. The regression prediction model outputs the disease prediction value. In this way, using the trained regression prediction model, it is possible to perform comprehensive, accurate, and efficient disease risk prediction based on multi-modal data; using the medical image with rich features as the knowledge to guide other modalities, making the prediction accuracy of the model higher.
[0015] In a possible design of the first aspect, processing the first, second, and third feature vectors with consistent expressions to obtain a globally feature representation after feature weighted fusion specifically includes:
[0016] Using the first feature vector as the query, the second feature vector with the same expression as the first feature vector as the key, and the third feature vector with the same expression as the first feature vector as the value, and adopting the scaled dot-product attention mechanism to obtain a fusion feature representation containing the weighted information of the first, second, and third feature vectors; and
[0017] Processing the fusion feature representation to obtain the globally feature representation.
[0018] In a possible design of the first aspect, processing the fusion feature representation to obtain the globally feature representation specifically includes:
[0019] Inputting the fusion feature representation into a large language model for processing to obtain an encoded output; and
[0020] Performing average pooling on the encoded output to obtain the globally feature representation.
[0021] In a possible design of the first aspect, obtaining the first feature vector corresponding to the first data, the second feature vector corresponding to the second data, and the third feature vector corresponding to the third data specifically includes:
[0022] Extracting the low-frequency components and high-frequency components from the first data, the second data, and the third data respectively; and
[0023] Concatenating the low-frequency components and high-frequency components to form the first feature vector corresponding to the first data, the second feature vector corresponding to the second data, and the third feature vector corresponding to the third data.
[0024] In a possible design of the first aspect, the training method further includes:
[0025] Successively performing unified formatting, data cleaning, annotation, synchronization, and data augmentation processing on the initial medical image data, the initial physiological signal data, and the initial gene sequence data to correspondingly obtain the medical image data, the physiological signal data, and the gene sequence data.
[0026] In a second aspect, a method for predicting disease risk is provided. The prediction method includes:
[0027] Obtaining the current medical image data, the current physiological signal data, and the current gene sequence data;
[0028] Processing the current medical image data, the current physiological signal data, and the current gene sequence data through the same process as obtaining the training data above to obtain the current global feature representation; and
[0029] Using the current global feature representation as the input of the disease prediction model as above to obtain the current disease prediction value.
[0030] In a third aspect, a training device for a disease risk prediction model is provided. The training device includes:
[0031] A first acquisition unit for obtaining training data, specifically including: obtaining data to be processed, where the data to be processed includes: medical image data, physiological signal data, and gene sequence data. Among them, the gene sequence data includes: first data and a first annotation related to a disease in terms of gene sequence, the medical image data includes: second data and a second annotation related to a disease in terms of image, and the physiological signal includes: third data and a third annotation related to a disease in terms of physiological signal; processing to obtain a first feature vector corresponding to the first data, a second feature vector corresponding to the second data, and a third feature vector corresponding to the third data; using the first feature vector as knowledge to guide the second feature vector and the third feature vector, and adopting a knowledge distillation technique to make the expressions of the second feature vector and the third feature vector consistent with that of the first feature vector; and processing the first feature vector, the second feature vector, and the third feature vector with consistent expressions to obtain a globally feature-represented after weighted fusion; and
[0032] A training unit for training the regression prediction model by using the training data,
[0033] where the training data includes: the globally feature-represented as input, the disease prediction value corresponding to the globally feature-represented as output, and the true disease value.
[0034] In a fourth aspect, a device for predicting disease risk is provided, and the prediction device includes:
[0035] A second acquisition unit for obtaining current medical image data, current physiological signal data, and current gene sequence data;
[0036] A data preprocessing unit for processing the current medical image data, the current physiological signal data, and the current gene sequence data through the same processing as obtaining the training data above to obtain a current globally feature-represented; and
[0037] A data postprocessing unit for using the current globally feature-represented as the input of the disease prediction model above to obtain a current disease prediction value.
[0038] In a fifth aspect, an electronic device is provided, and the electronic device includes: a processor, and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the training method in any possible implementation manner in the first aspect, or executes the prediction method in any possible implementation manner in the second aspect.
[0039] In a sixth aspect, there is provided a computer-readable storage medium including a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the training method according to any possible implementation manner in the first aspect, or execute the prediction method according to any possible implementation manner in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the related art descriptions. Obviously, the drawings in the following descriptions are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 is the overall flowchart of multi-modal data processing provided by the embodiments of the present application;
[0042] Figure 2 is the data cleaning flowchart provided by the embodiments of the present application;
[0043] Figure 3 is the overall annotation flowchart provided by the embodiments of the present application;
[0044] Figure 4 is the medical image annotation flowchart provided by the embodiments of the present application;
[0045] Figure 5 is the gene data annotation flowchart provided by the embodiments of the present application;
[0046] Figure 6 is the annotation flowchart of physiological signals provided by the embodiments of the present application;
[0047] Figure 7 is the synchronization flowchart of medical images and physiological signals provided by the embodiments of the present application;
[0048] Figure 8 is the training flowchart provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0052] The following will specifically describe the disease risk prediction method and training method of the embodiments of this application with reference to the accompanying drawings.
[0053] (1) Training data processing part:
[0054] The processing of step S1 is as Figure 1 shown:
[0055] S1: Obtain input data from the intelligent health system. The input data includes medical images I (such as X-rays, MRIs, CTs), gene sequencing data Tgene (such as RNA-seq, DNA-seq), and physiological signals Tphysio (such as ECG, EEG). For these data, they need to be processed into a unified format respectively.
[0056] S11: The medical image data comes from various medical imaging devices, such as X-ray, MRI, CT scanners, etc. The input part includes the original medical image data I. The original medical image data refers to the continuous scanned images obtained by medical imaging devices. Each image contains information such as the patient's body part and lesion area. For medical image data, the unified DICOM format is adopted. Write a batch conversion script to make it automated. The processing process includes: modifying the image size to 512×512, checking the original data format. If the original data is not in DICOM format, use the dcmconv command of the DCMTK tool for format conversion, and automatically add metadata, such as patient information, scanning parameters, device information, etc. to the newly converted DICOM files. For the image data among them, perform normalization processing, using logarithmic Log normalization and Min-Max normalization methods. Log normalization solves the problem of long-tail distribution, and Min-Max normalization further maps the data to a unified range. The processing flow is shown in the following formula. Through the above steps, it is ensured that the data from different devices and scanning methods can be in a unified format.
[0057]
[0058]
[0059] where x i is the original value, ∈ is a constant to prevent mathematical errors when the data contains zero values, is the data after Log normalization. max(x log ), min(x log ) are the minimum and maximum values of the data after Log normalization respectively. The data value after Min - Max normalization. Var(x i ) is the variance of x i , representing the degree of dispersion of this feature among all samples. n is the number of samples. x ij is the value of x i in sample j. is the mean of x i , representing the average value among all samples. Output medical image data Is1 in a unified format.
[0060] S12: The gene data Tgene comes from a gene sequencing platform, including sequencing data of types such as RNA - seq and DNA - seq. The original sequencing data is usually stored in the FASTQ format, containing gene sequences and sequencing quality information; the data after alignment and processing is usually stored in the BAM format, representing the aligned gene sequences and their positions in the reference genome. For gene data, first, it is necessary to use the BWA tool to create an index for the reference genome. After the index creation is completed, use the BWAmem command to align the FASTQ file with the reference genome to generate a SAM file. Next, use the SAMtools view command to convert the SAM file to a BAM file, and use the SAMtools sort command to sort the BAM file. To better process and analyze the data, it is also necessary to create an index for the sorted BAM file. Finally, use SAMtools flagstat to view the basic statistical information of the BAM file to ensure the quality of the alignment results.
[0061] S13: The physiological signal data Tphysio includes electrophysiological signals such as ECG (electrocardiogram) and EEG (electroencephalogram). These signals are acquired through physiological sensors and record continuous electrophysiological activities. To ensure data consistency, all physiological signals need to have a unified sampling rate and be saved in CSV format. First, unify the sampling rates of different signals. Select 1000Hz as the target sampling rate. EEG signals usually contain higher-frequency components, while ECG signals have lower frequencies. 1000Hz can generally meet the needs of most signal processing. First, unify the sampling rate through resampling, with the target sampling rate being 1000HZ, and then process it through the Fourier dual-channel filtering method. The specific steps are as follows. First, perform the Fourier transform,
[0062]
[0063] where X(f) is the frequency-domain representation, x(t) is the time-domain signal, and f is the frequency.
[0064] Next, combine the low-pass and high-pass filters, which is achieved through cascaded filtering. First, use the low-pass filter, which is achieved by setting the high-frequency components to zero,
[0065]
[0066] where f cutoff is the cut-off frequency of the low-pass filter.
[0067] Next, use the high-pass filter, which is achieved by setting the low-frequency components to zero,
[0068]
[0069] where f cutoff is the cut-off frequency of the high-pass filter.
[0070] Then combine the above processing results. Finally, perform the inverse Fourier transform,
[0071]
[0072] where X(f) is the processed frequency-domain signal and x(t) is the restored time-domain signal.
[0073] The processing of step S2 is as Figure 2 shown:
[0074] S2: Clean the data in the unified format processed in the S1 stage to improve the quality and consistency of the data. The main goals of this process are to remove noise, repair missing values, and detect abnormal data.
[0075] S21: Input the DICOM format medical image data Is1 after being processed by S1, and use Gaussian filtering to remove the noise in the image to ensure image smoothness and reduce interference. First, select a filtering window with a size of 3×3, calculate the corresponding Gaussian function values, and generate a Gaussian filter kernel by selecting an appropriate standard deviation. Then, convolve each pixel in the image with the filter kernel to calculate the new value of each pixel after filtering. For boundary pixels, use the mirror padding method to reduce boundary distortion. Use a method combining bilinear interpolation and GAN network to repair it. The specific steps are as follows. First, for the initial two-dimensional image, if the missing value is at the pixel block coordinate (x, y) point, perform simple repair through the following formula:
[0076]
[0077] where x 0 , y 0 represents the pixel block coordinate of the upper left corner, and x 1 , y 1 represents the pixel block coordinate of the lower right corner.
[0078] Then use the GAN network to perform deep repair on it. The objective of GAN is expressed as:
[0079]
[0080] where I is the real image, D is the discriminator, G is the generator, I M is the missing area, G(IM) is the generated image after repair, P data represents the distribution of the real image, represents the distribution of the complete images sampled from the dataset, P gen represents the distribution of the generated image, represents the distribution of the repaired image generated by the generator G, and E I~Pdata is the logarithmic expectation of the discriminator's discrimination result for the real image I, and E IM~Pgen is the logarithmic expectation of the discriminator's discrimination result for the image G(IM) repaired by the generator.
[0081] Judge whether the image is abnormal by calculating the signal-to-noise ratio SNR and the structural similarity index SSIM. For abnormal images with SSIM lower than the threshold of 0.85 or SNR lower than 20 dB, it is necessary to re-acquire and perform S11 operation.
[0082] S22: Input the gene sequencing data Tgenes1 processed by S1. First, use the quality control tool FastQC to perform quality detection on the original FASTQ file, check the quality scores of each sequencing read, and remove sequences with quality scores lower than 30. Then use the tool Trimmomatic to remove the residual adapter sequences during the sequencing process. Missing values commonly found in RNA-seq and DNA-seq data usually appear in low-expression genes or missing parts of gene regions. For these missing values, a multiple imputation algorithm is used for filling. Then use the BWA tool to align the FASTQ file, align it with the reference genome, and generate a BAM file. Next, use SAMtools flagstat to perform statistics on the alignment results to identify genomic regions with low alignment rates or abnormal alignments. Use the genomic data analysis platform GATK to further evaluate the data quality, detect abnormal data, and repair it. For abnormal data that cannot be repaired, consider deleting it to ensure the reliability of the data quality in subsequent analyses.
[0083] S23: Input the physiological signal data Tphysios1 processed by S1. First, perform a filtering operation to remove high-frequency and low-frequency noise. For low-frequency noise (such as baseline drift), use a high-pass filter with a cut-off frequency ranging from 0.5 Hz to 1 Hz for filtering. For high-frequency noise (such as electrical interference and electromyographic noise, etc.), use a low-pass filter with a cut-off frequency of 30 Hz for filtering. After filtering, view the waveform of the signal through visualization means to confirm whether the noise has been effectively removed. During the filtering process, the edges of the signal may be affected. To avoid phase delay, use filtfilt for filtering. However, boundary effects still need to be noted. It is recommended to perform mirror padding on both ends of the signal to reduce the influence of boundary effects. Next, for the missing values in the signal, use the KGM clustering method for repair. First, perform K-nearest neighbor calculation. For each data point xi, calculate its distance from other points in the dataset, and select the nearest K points as neighbors. The nearest neighbors of X i are {x i1 , x i2 , …, x ik}}. Then extract the nearby information and calculate the local structure of each point:
[0084]
[0085] where is the preliminary label based on local calculation, and KNN is the K-nearest neighbor algorithm.
[0086] Next, GMM is used to model the global distribution. For the entire dataset, GMM assumes that the data points come from a mixture of multiple Gaussian distributions. Each cluster is represented by a Gaussian distribution. GMM estimates the parameters of each cluster through the Expectation-Maximization (EM) algorithm. Specifically, the GMM model is as follows:
[0087]
[0088] where N(x i |μ k ,Σ k ) is the probability density function of the k-th Gaussian distribution, π k is the mixing coefficient of the k-th Gaussian distribution, representing the proportion of the k-th cluster. μ k is the mean of the k-th Gaussian distribution, and Σ k is the covariance matrix of the k-th Gaussian distribution. θ is the set of model parameters of GMM, including the weights, means, and covariance matrices of each Gaussian distribution.
[0089] EM algorithm: First, calculate the posterior probability of each data point belonging to each cluster:
[0090]
[0091] where is the probability density of the k-th Gaussian distribution for the data point x i , and π k is the mixing coefficient of the k-th cluster.
[0092] Then, update the parameters (means, covariances, mixing coefficients) of GMM
[0093]
[0094] where γ ik is the posterior probability that the data point x i belongs to the k-th cluster, and N is the total number of data points.
[0095] After training GMM is completed, for each data point x i , GMM will output the probability that this point belongs to each cluster. According to the probability γik, select the cluster with the highest probability as the belonging cluster of the data point:
[0096]
[0097] where argmax represents the index of the maximum number.
[0098] Finally, the random forest method is used for anomaly detection to automatically identify abnormal data in the signal. This process trains a model to learn the patterns of normal signals, thereby accurately identifying outliers and marking or processing them.
[0099] The processing in Step 3 is as Figure 3 shown below:
[0100] S3: Label the clean data processed in the S2 stage.
[0101] S31: During the process of labeling medical image data, this solution uses the pre-trained model U-Net. This model has been trained on a large-scale dataset and can quickly and accurately segment medical images, especially for common medical image regions (such as tumors, lesion regions, etc.). In the improved U-net model, Gaussian filtering is added to both the hidden layer and the decoder. For the output feature map F of each layer l , perform Gaussian filtering, and the formula is:
[0102] F' l = F l * G σ
[0103] where G σ is the Gaussian filter kernel, and σ is the standard deviation.
[0104] The formula for the Gaussian filter kernel:
[0105]
[0106] Apply the convolution kernel Gσ to smooth the feature map, reduce noise and improve image quality, where x and y are the coordinates of the pixel points.
[0107] To solve the problem of over-focusing on unimportant regions in U-Net, this solution adds a channel attention mechanism after the convolutional layer. First, perform global average pooling on each channel to extract global context information:
[0108]
[0109] where H and W are the height and width of the image, and z c is the channel-level descriptor.
[0110] Then calculate the importance of each channel through a fully connected layer:
[0111] w c = σ(W 1 z c + b 1 )
[0112] w' c = σ(W 2 w c + b 2 )
[0113] Among them, W1 and W2 are the weight matrices of the fully connected layers, σ is the activation function (usually the Sigmoid function is used), and b1 and b2 are the bias terms.
[0114] Next, the feature map is weighted according to the weights of each channel:
[0115] F' l = F l ⊙ w' c
[0116] Among them, ⊙ represents element-wise multiplication.
[0117] To enhance the model's segmentation ability for targets of different scales, a multi-scale feature fusion method is introduced. First, different-scale features are extracted through pooling operations of multiple scales:
[0118] F ppm = [Pool 1 (F l ), Pool 2 (F l ), …, Pool n (F l )]
[0119] Next, the feature maps after multi-scale pooling are fused:
[0120] F fusion = concat(F ppm )
[0121] Then, the fused feature map is processed through a convolutional layer to finally obtain the segmentation result. The improved U-Net model is used to load the pre-trained weights so that the model can segment medical images. Secondly, the medical image data Is2 processed in the S2 stage is input, and the U-Net model is used to predict each frame of the image to obtain a binary segmentation mask. First, noise and small holes are removed through opening (removing small objects) and closing (filling small holes) operations. Secondly, connected component labeling is performed. If there are multiple regions in the segmentation result, connected component labeling is used for region separation to label each target region. Then, the annotation results of each frame of medical image are saved in the COCO format, and the file save path needs to be the same as the original image to associate the annotation data with the image file. Finally, a script is written to process the entire medical image data to ensure that it has all been predicted by the U-Net model to generate a complete image annotation dataset Is3. As Figure 4 shown.
[0122] S32: For the annotation of gene data Tgenes2, this solution uses an improved BioBERT model for automatic annotation.
[0123] 1) Input layer: The gene sequence data is converted into vectors using k-mer encoding. Then, not only the gene sequence data but also the gene expression data and mutation information are input into the model as additional inputs.
[0124] X fusion = α @ X seq + β · X meta
[0125] where X seq is the gene sequence, X meta is the additional input, and α and β are weighting coefficients.
[0126] 2) Add the combined multi-scale attention mechanism and hierarchical attention mechanism in the encoder layer.
[0127] The multi-scale attention mechanism and hierarchical attention mechanism are combined to aggregate context information from multiple perspectives.
[0128] For each scale and each level, calculate local and global attention:
[0129]
[0130] where, Q local,s represents the query matrix (Query) in the local range at the s-th scale; Q global,s represents the query matrix in the global range at the s-th scale; d k is the dimension of the key matrix, the local key matrix K local,s is the feature representation in the local range at the s-th scale, which is usually the feature of the local region extracted from the input features and used to capture local context information; the global key matrix K global,s is the feature representation in the global range at the s-th scale, which is used to capture global context information and usually covers the entire input feature; V local,s represents the value matrix (Value) in the local range at the s-th scale, which represents the feature representation of the local region and is used to weight and aggregate local context information.
[0131] Then, the information of multiple scales and levels is weighted and fused:
[0132]
[0133] where, α s represents the weighting coefficient at the s-th scale. This coefficient is learned through training and used to represent the relative importance of each scale.
[0134] By combining the multi-scale attention mechanism and the hierarchical attention mechanism, the model can better capture the long-range dependencies and local and global features in gene sequences, thereby improving the model's ability to understand and annotate complex gene data.
[0135] For each input segment, the model will predict the label for each position to identify whether it belongs to a functional gene region (such as exons, introns, etc.), and annotate whether there are mutations and the mutation types. The annotation result Tgenes3 is saved in VCF format. For the above steps, write a script to enable automatic execution. Gene data annotation is as Figure 5 shown.
[0136] S33: For the annotation of physiological signal Tphysios2, this solution needs to annotate the ECG and EEG signal data separately. For the ECG physiological signal, this solution uses the Transformer-based ECG-FM model, which is an open basic model for ECG analysis [4]. Input the ECG data processed by S2, and split the signal into non-overlapping 5s segments to generate the model input. Use ECG-FM to generate a label for each signal and record the timestamp of the event occurrence. For EEG data, this solution uses the Transformer-based EEG Conformer model for automatic annotation. Input the EEG data processed by S2, and split the input data using the time window method, with each segment being 2 seconds. Different types of annotations are performed for different tasks. For the data with event timestamps, map these events to the EEG signals to create labels. Then let EEGConformer load the pre-trained file to enable automatic annotation. Finally, output the annotated data Tphysios3 in CSV format. Each EEG signal within a time window will correspond to a label and be saved in the annotation file. As Figure 6 shown.
[0137] S34: Manually check the medical image annotation file Is3 generated in step S31. First, load the original medical image and the segmentation result generated by the U-Net model through OpenCV, overlay and display the two, and check whether the segmentation result is consistent with the actual target area to ensure that the target area is accurately annotated. If it is found that the segmentation boundary generated by the U-Net model is inconsistent with the actual target area, manually adjust the segmentation area to correct the mis-segmented part. In addition, for the target objects not detected by the improved U-Net, manually add class labels and manually mark the positions of the objects in the unsegmented areas. Through these manual checks and corrections, ensure that the final medical image annotation result is more accurate and reliable and suitable for subsequent analysis and research.
[0138] S35: Manually review the gene data annotation file Tgenes3 generated in step S32. First, load and visualize the BAM file using the gene data analysis tool IGV, carefully check the overall quality of the file, especially the accuracy of the alignment and the fluctuations in the alignment quality. Use the Samtools flagstat command to obtain alignment statistics, analyze the alignment quality, success rate, and coverage of different genomic regions to ensure the integrity and accuracy of the data. During the review process, focus on checking the variant sites marked in the BAM file, verify whether the mutation types (such as synonymous mutations, missense mutations, nonsense mutations) and the gene regions of the variants are consistent with the reference genome annotation, especially whether the variant is located in a functional region (such as exons, introns, etc.) and whether this mutation may affect gene function. At the same time, conduct a detailed review of regions with low alignment quality or small coverage, and correct the coordinates of the variant positions if necessary. For mislabeled mutation types (such as confusion between missense and synonymous mutations), make manual corrections in a timely manner to ensure the accuracy of the annotation. If it is found that some mutations or gene regions are not detected by the annotation tool, manually add the missing variant sites. Finally, after completing the manual correction, update and save the annotation data in the BAM file to ensure the comprehensiveness and accuracy of the variant annotation.
[0139] S36: First, load the electrocardiogram data processed in S2 and the annotation file Tphysios3 generated by the ECG-FM model through a dedicated electrocardiogram viewer, and check whether the division of each 5-second signal segment is accurate to ensure that the signals are not lost or misclassified, especially in areas where the signal amplitude increases or the fluctuations are more frequent. Focus on checking whether the event times generated by the model match the actual physiological events (such as arrhythmia, tachycardia, etc.), and verify the accuracy of the model labels by manually backtracking and comparing the event labels. If there are any unlabeled or mislabeled situations (such as incorrect category labels for arrhythmia events), then it is necessary to manually add the missing labels or modify the incorrect labels. The event categories of all labels should ensure consistency and error-freeness, and cross-validate the label results, especially the consistency among different labelers, to prevent data inconsistencies caused by label biases. For EEG data, use the EEG analysis tool EEGLAB to load the EEG signal data processed in S2 and the annotation file generated by the EEGConformer model, check whether the division of each 2-second time window is reasonable, and ensure that the details of brain electrical activities can be effectively captured, especially for changing brain electrical activities (such as brain wave mutations or rapid eye movement stages). By comparing with the expected times of brain electrical events (such as changes in sleep stages, appearance of specific brain wave activities, etc.), ensure that the labels generated by the model are consistent with the actual physiological phenomena. If there are obvious errors or omissions, they should be manually modified, especially in brain regions where the signals are blurred or the boundaries are not clear. For areas of brain electrical activities not labeled by the model (such as certain series of activities), supplementation should be carried out to ensure the harmlessness of the supplemented file. Finally, to ensure the consistency and accuracy of the labels, check the deviations between the labels through cross-validation by different labelers, and ensure that the finally generated label file (in CSV format) meets the standards and has high consistency. Ensure that the EEG signals within each window have accurate labels and the label file formats are consistent for subsequent analysis and applications. S37: After completing the data annotation, conduct a comprehensive secondary quality inspection to ensure that the annotation information of all medical images, gene sequences, and physiological signal data is not only consistent and accurate in content, but also error-free in key details such as format, annotation category, and timestamp. In addition, make a detailed backup of all manually modified and reviewed data, and strictly follow the data integrity management process to ensure the traceability and security of the data throughout its life cycle.
[0140] The processing of step S4 is as Figure 7 shown:
[0141] S4: Synchronize the data after the annotation in stage S3 is completed.
[0142] S41: To achieve the timestamp alignment between the medical image Is3 and the physiological signal data Tphysios3, it is first necessary to extract the timestamp of each frame of the image, which is obtained from the "AcquisitionTime" field in the DICOM file of the image. If the image timestamp is missing, the time interval can be calculated based on the sequence number of the image. Subsequently, the timestamp is extracted from the sampling file of the physiological signal data (such as ECG, EEG). If there is partial or complete loss of the signal time series, but the sampling rate of the signal remains consistent, the timestamp is calculated based on the sampling frequency of the signal to ensure that the time of each data point is consistent with the actual sampling time. Then, alignment is directly performed based on the timestamps. Finally, it is ensured that the timestamps of the image and the signal are consistent so that the corresponding data can be accurately matched in the multi-modal data analysis. The synchronized data is D1 = {(t1, Is41, Tphysios41), (t2, Is42, Tphysios42), …, (tn, Is4n, Tphysios4n)}, where Tphysios represents the synchronized data of the physiological signal at time t. The synchronization process is as Figure 7 shown.
[0143] S42: To achieve the linkage between the gene data Tgenes3 and the image data Is3, it is first necessary to associate the gene data with the corresponding medical image data through the unique identifier of the patient (such as patient ID or other medical record numbers). First, the identifier information of the patient is obtained from the BAM of the gene data and the DICOM file of the image data to ensure information consistency. Then, the gene data is matched with the corresponding image data based on the patient ID. Next, for specific diseases (such as hereditary tumors), data linkage and correlation analysis are performed by comparing the mutation information detected in the gene data with the lesion location in the image data. Finally, more accurate disease feature analysis and clinical research are carried out using the linked multi-modal data (gene information and image data) to ensure that the integration and analysis of gene information and image data can provide comprehensive support for the diagnosis and treatment of patients. The synchronized data is D2 = {(t1, Is41, Tgenes41), (t2, Is42, Tgenes42), …, (tn, Is4n, Tgenes4n)}, where Tgenes represents the synchronized data of the gene data at time t.
[0144] Processing of step S5: To improve the robustness and generalization ability of the model, data augmentation is performed on the data synchronized in S4.
[0145] S51: First, the medical image Is4i is randomly rotated at a random angle (such as between -30 and +30 degrees) to enhance the robustness of the model to rotation changes. Next, the medical image is randomly translated horizontally and vertically, and the translation range is set to 10-30 pixels to help the model learn the different positions of objects in the image. Then, the medical image is randomly scaled to change the image size (such as enlarging or reducing 10%-20%) to simulate organs or lesion areas of different sizes. A random shear transformation is applied to adjust the image shape so that the model can adapt to various geometric deformations of the image. Horizontal or vertical flipping is particularly suitable for most medical images (such as CT, MRI), especially when the tumor or organ is bilaterally symmetrical. By randomly adjusting the brightness, contrast and saturation of the image, different imaging conditions are simulated, image diversity is enhanced, and the model is helped to cope with the influence of different equipment and scanning settings. In addition, the image is randomly cropped, a part of the image is intercepted and resized to enhance the model's perception of the local area and simulate the local lesion area. Finally, the above process is written into a script file to achieve the random enhancement effect of the medical image. It is worth noting that when enhancing the medical image, the synchronized physiological signal needs to be modified accordingly. The enhanced medical image Is5 is output.
[0146] S52: Input the synchronized gene data Tgenes4i, and create different mutation versions by simulating base mutations in the gene sequence (such as randomly replacing A with T, G or C, etc.), thereby simulating the natural mutation process. By setting the mutation probability, some bases are randomly replaced with other bases to further enhance the diversity of the sequence. By randomly inserting or deleting gene fragments, insertion / deletion mutations (Indels) in the genome are simulated. This method generates mutation data by randomly selecting a position to insert or delete base pairs or sequence fragments. In addition, reverse complementation of the sequence is performed to generate the reverse complementary chain of the original sequence, which can simulate the different directions of the gene chain and help the model learn the gene sequence in both directions. By randomly replacing part of the sequence from other gene sequences, it can help the model learn the diversity of gene structure and simulate different genetic variations. At the same time, random noise is introduced into the gene sequence, such as inserting misplaced base pairs or irrelevant sequences at specific positions to simulate noise or errors in sequencing. These methods help to improve the robustness of the model to gene sequence variation, technical noise, etc. Output the enhanced Tgenes5.
[0147] S53: Use a generative adversarial network (GAN) to generate medical image samples of rare diseases, aiming to alleviate the long-tail effect in medical image data, improve the performance of minority-class samples, and achieve a balanced data distribution through data augmentation. For rare cases or complex pathological situations (such as rare lesions, special pathological features, heterogeneous lesions, etc.), first divide the dataset into a training set and a test set, and label it for specific pathological types. Use the generator and discriminator in the GAN framework for adversarial training to optimize the generator to generate medical image samples that meet clinical requirements, and ensure that the generated images are highly consistent with real medical images in terms of structure, texture, and semantics. To further improve the generalization ability of the model and the recognition ability of minority-class samples, finally optimize the generated samples through image post-processing techniques (such as denoising, enhancing details, contrast adjustment, etc.), and then add them to the original dataset. This not only enriches the diversity of the dataset, but also effectively improves the learning and reasoning ability of the deep learning model for rare pathological samples, promoting the accuracy and robustness of medical image analysis.
[0148] Processing of step S6: Establish an efficient hierarchical storage structure.
[0149] Specifically, first classify the data according to disease categories and data modalities (such as medical images, gene data, physiological signals, etc.). First, create a top-level folder named "Disease Data", and then further create sub-folders according to the types of diseases (such as heart diseases, cancers, neurological diseases, etc.). Under each disease category, further divide it into different sub-directories according to data modalities, such as "Image Data", "Gene Data", "Physiological Signals", etc. Under each data modality, create folders according to the specific types of data (such as CT, MRI, ECG, EEG, etc.). Inside each folder, further classify according to information such as collection date, patient ID, or experiment number to ensure that the hierarchical structure of data storage is clear and convenient for subsequent retrieval. When naming files, it is recommended to use a consistent naming convention that includes key information (such as disease category, data type, collection date, etc.) to ensure quick positioning during retrieval. Through this hierarchical structure, efficient storage management is achieved, data redundancy is reduced, and it is convenient for quick searching, analysis, and sharing of data. At the same time, combined with a database management system or a data lake system, the flexibility and efficiency of data access are further improved.
[0150] This solution completes data collection, formatting, cleaning, annotation, synchronization, and augmentation in an automated manner. The entire process achieves high efficiency and automation, and only requires manual intervention in a few key steps, which not only reduces the computational resource consumption of generating the dataset, but also effectively alleviates the imbalance problem of real-world datasets.
[0151] (II) Training part:
[0152] The processing of step S7 is asFigure 8 As shown below:
[0153] S7: Train the data stored in S6.
[0154] S71: (1) Input the processed medical image Is6, gene sequence data Tgenes6, and physiological signal text data Tphysios6. Preprocess and encode these data. First, use discrete wavelet transform to preprocess the medical image to obtain low-frequency and high-frequency components.
[0155] I low ,I high = DWT(I s6 )
[0156] where I is the medical image; I low is the low-frequency component, which contains the global structure information of the image; I high is the high-frequency component, which contains the detailed information of the image, and DWT represents discrete wavelet transform.
[0157] Then, for each frequency band, independently use ResNet to encode the medical image to obtain the feature representations of different frequency bands.
[0158]
[0159] where f image is the image encoder.
[0160] Then, perform feature fusion. First, map the encoded features of the wavelet components to the shared feature space through an adapter, and finally use the attention mechanism for weighted fusion:
[0161]
[0162] where A image is the fully connected layer.
[0163] Then, concatenate the features of different frequency bands (concat).
[0164]
[0165] (2) Immediately perform discrete wavelet transform DWT on the gene data.
[0166]
[0167] where T gene is the gene data, is the low-frequency component, is the high-frequency component.
[0168] Then, the BERT encoder is applied to process the low-frequency and high-frequency components respectively.
[0169]
[0170] where f gene is the encoder.
[0171] Then, the features of the low-frequency and high-frequency are mapped to a unified feature space through an adapter and fused:
[0172]
[0173] where A gene is the fully connected layer.
[0174] Next, the features of different frequency bands are concatenated.
[0175]
[0176] (3) For the physiological signal text data, first perform wavelet transform on it to obtain low-frequency and high-frequency components:
[0177]
[0178] where T physio is the physiological signal data, are the low-frequency and high-frequency components respectively. Then use BERT to encode the low-frequency and high-frequency components.
[0179]
[0180] where f physio is the encoder.
[0181] Next, map the features of the low-frequency and high-frequency to the shared feature space respectively:
[0182]
[0183] where A physio is the fully connected layer.
[0184] Next, concatenate these features.
[0185]
[0186] S72: For the data processed in S71, perform fusion, and adopt the attention mechanism for fusion. For the input modal features, perform linear transformation on the features of each modality to make their dimensions consistent.
[0187]
[0188] where Wv , W g and W p are the weight matrices of linear transformations, and the output V q , are all eigenvectors after linear transformation.
[0189] Then, a cross-modal distillation module is added. Taking Vq as the key feature of the strong modality, it serves as the knowledge to guide the weak modality.
[0190] V s = W s ·V q
[0191] where W s is the dimensionality reduction matrix of the strong modality features, used to extract key information. V s is the knowledge feature of the strong modality after extraction.
[0192] The physiological signals and gene data are adjusted as the features of the weak modality to the same space as the strong modality for calculating similarity.
[0193]
[0194] where W′ p and W′ g are projection matrices, used to adjust the weak modality feature space. is the aligned weak modality feature.
[0195] Through the distillation loss function, the expressions of the strong modality features and the weak modality features are made similar.
[0196]
[0197] where L distill multi is the distillation loss, used to measure the difference between the strong modality and weak modality features; MSE is the mean square error function, calculating the squared difference between two eigenvectors; m, n represent different modalities, m is the physiological signal (physio), n is the gene data (gene); modalities represents the set of all modalities.
[0198] Then, these features are input into the attention mechanism. The Scaled Dot-ProductAttention is used to calculate the attention weights, taking V s , as the input of query Q, key K, and value V respectively
[0199]
[0200] where W Qis the weight matrix for the query (Query), used to map the input feature V s to the query space, W K is the weight matrix for the key (Key), used to map the gene feature to the key space, W V is the weight matrix for the value (Value), used to map the physiological feature to the value space.
[0201] Then, by calculating the dot product between the query and the key and using a scaling factor to avoid large numerical values, the attention scores are calculated,
[0202]
[0203] where d is the dimension of the feature and T is the transpose.
[0204] Next, the softmax activation function is used to calculate the weights for each modality:
[0205] α i = Softmax(Attention Score(Q, K))
[0206] The final fused feature is obtained by applying the attention weight α i to the corresponding value V:
[0207]
[0208] where, X f is the fused feature representation, which contains the weighted information of all modalities.
[0209] S74: The fused feature X f will be fed into the large language model ViLT for further processing, with the function of: learning the global and local relationships between the two by jointly modeling the visual and text modalities.
[0210] First, the fused feature undergoes a non-linear transformation through a feed-forward neural network:
[0211] X ffn = ReLU(X f @W 1 + b 1 )·W 2 + b 2
[0212] where, W 1 , W 2 are the weight matrices of the feed-forward network. b 1 , b 2 are the biases of the feed-forward network. X ffnFeatures processed by the feedforward network.
[0213] Then, residual connection normalization is performed.
[0214] X norm1 = LayerNorm(X ViLT + X a )
[0215] X norm2 = LayerNorm(X norm1 + X ffn )
[0216] Among them, X ViLT is the input feature, the feature extracted by the Vision-Language Model ViLT module. Xa is the output feature after passing through the attention mechanism.
[0217] Next, after stacking multiple Transformer encoders, the final output is generated:
[0218] X final = Transformer(X norm2 )
[0219] S75: Input the data processed by S74, and use average pooling to generate a global feature representation X from N feature vectors g
[0220]
[0221] final,i where X g is the i-th feature vector. X g
[0222] is the global representation, with a dimension of [1, d]. Input the global representation X g into the fully connected layer, and generate the regression prediction value through linear transformation
[0223]
[0224] o and b o are the weights and biases of the regression layer, is the regression output.
[0225] S76: Train the model. The loss function during the training process is
[0226]
[0227] i where N is the total number of samples, and y is the true label of sample i (i.e., the above annotation).It is the predicted value of the model for sample i (the risk score of the disease).
[0228] (2) Prediction part:
[0229] Based on the trained model, through the processing of multi-modal data output, the predicted value of the disease risk can be finally obtained.
[0230] Optionally, when using a Gaussian filter for denoising during image cleaning, zero-padding can be used instead of mirror-padding, which is more convenient and simpler in operation; when annotating ECG, a CNN-based model can be selected for annotation, which can focus on more local information.
[0231] Applying the technical solution of this application:
[0232] 1. Through the guidance of medical images, the features of physiological signals and gene sequences can be better combined with image features to provide more comprehensive disease information.
[0233] 2. Through knowledge distillation, the reliable knowledge of medical images can be transferred to physiological signals and gene sequences to improve their feature expression ability, especially when the data quality is poor.
[0234] 3. The guidance of medical images can reduce the influence of noise in physiological signals and gene sequences and enhance the robustness of the model to data noise.
[0235] 4. Through knowledge distillation, the rich information of medical images can make up for the problem of insufficient data in physiological signals and gene sequences and improve the performance of the model on small-sample data.
[0236] 5. The guidance of medical images can optimize the fusion process of multi-modal features to ensure that the contributions of each modal feature are balanced during fusion and avoid one modal dominating.
[0237] 6. The guidance of medical images can enhance the interpretability of the model and help doctors understand the role of physiological signal and gene sequence features in disease prediction.
[0238] 7. Through the guidance of medical images, the trained model can be more easily transferred to other diseases or fields to improve generality.
[0239] The embodiment of this solution also provides a training device for a disease risk prediction model, and the training device includes:
[0240] A first acquisition unit for acquiring training data, specifically including: acquiring data to be processed, where the data to be processed includes medical image data, physiological signal data, and gene sequence data. Among them, the gene sequence data includes: first data and first annotations related to diseases in terms of gene sequences, the medical image data includes: second data and second annotations related to diseases in terms of images, and the physiological signals include: third data and third annotations related to diseases in terms of physiological signals; processing to obtain a first feature vector corresponding to the first data, a second feature vector corresponding to the second data, and a third feature vector corresponding to the third data; using the first feature vector as knowledge to guide the second feature vector and the third feature vector, and adopting knowledge distillation technology to make the expressions of the second feature vector and the third feature vector consistent with the first feature vector; and processing the first feature vector, the second feature vector, and the third feature vector with consistent expressions to obtain a globally feature-represented global feature after weighted fusion; and
[0241] A training unit for training the regression prediction model using the training data,
[0242] where the training data includes: the globally feature-represented global feature as input, the disease prediction value corresponding to the globally feature-represented global feature as output, and the true disease value.
[0243] An embodiment of this solution also provides a device for predicting disease risk, and the prediction device includes:
[0244] A second acquisition unit for acquiring current medical image data, current physiological signal data, and current gene sequence data;
[0245] A data preprocessing unit for processing the current medical image data, the current physiological signal data, and the current gene sequence data through the same processing as that for acquiring the training data to obtain a current globally feature-represented global feature; and
[0246] A data postprocessing unit for using the current globally feature-represented global feature as the input of the disease prediction model as described above to obtain a current disease prediction value.
[0247] An embodiment of this solution also provides an electronic device, including: a processor, and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.
[0248] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.
[0249] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device, and uses various interfaces and circuits to connect all parts of the device.
[0250] The memory can be used to store the computer program. The processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.
[0251] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMCs), secure digital (SD) cards, flash cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0252] The embodiment of this solution also provides a storage medium. The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0253] An embodiment of this solution also provides a computer program product, including: a computer program or instruction, which, when running on a computer, causes the computer to execute the method of any one of the above possible implementation manners.
[0254] The above is the preferred implementation manner of this solution. It should be noted that for those of ordinary skill in the art, without departing from the principle of this solution, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of this solution.
Claims
1. A method for training a disease risk prediction model, characterized in that: The disease risk prediction model is a regression prediction model, and the training method includes: obtaining training data; and Using the training data, the regression prediction model is trained. Among them, obtaining training data specifically includes: Obtaining data to be processed, the data to be processed comprising: medical image data, physiological signal data and gene sequence data, wherein the medical image data comprises: first data and a first annotation representing the true value of the disease from the aspect of the medical image, the gene sequence data comprises: second data and a second annotation representing the true value of the disease from the aspect of the gene sequence, and the physiological signal comprises: third data and a third annotation representing the true value of the disease from the aspect of the physiological signal; Processing to obtain a first eigenvector corresponding to the first data, a second eigenvector corresponding to the second data, and a third eigenvector corresponding to the third data; Using the first feature vector as knowledge to guide the second feature vector and the third feature vector, and using knowledge distillation technology to make the second feature vector and the third feature vector consistent with the expression of the first feature vector; and Processing the first feature vector, the second feature vector and the third feature vector that have consistent expressions to obtain a global feature representation after feature weighted fusion, The training data includes: the global feature representation as input, the disease prediction value corresponding to the global feature representation as output, and the true value of the disease.
2. The training method according to claim 1, characterized in that: Processing the first feature vector, the second feature vector, and the third feature vector that have consistent expressions to obtain a global feature representation after weighted fusion of features specifically includes: Using the first feature vector as a query, the second feature vector consistent with the first feature vector as a key, and the third feature vector consistent with the first feature vector as a value, a scaled dot product attention mechanism is used to obtain a fused feature representation including weighted information of the first feature vector, the second feature vector, and the third feature vector; and The fused feature representation is processed to obtain a global feature representation.
3. The training method according to claim 2, characterized in that: The fused feature representation is processed to obtain a global feature representation, specifically including: Inputting the fused feature representation into a large language model for processing to obtain an encoded output; and The encoded output is average pooled to obtain the global feature representation.
4. The training method according to claim 1, characterized in that: Processing to obtain a first feature vector corresponding to the first data, a second feature vector corresponding to the second data, and a third feature vector corresponding to the third data specifically includes: extracting low-frequency components and high-frequency components from the first data, the second data, and the third data, respectively; and The low-frequency component and the high-frequency component are concatenated to form a first eigenvector corresponding to the first data, a second eigenvector corresponding to the second data, and a third eigenvector corresponding to the third data.
5. The training method according to claim 1, characterized in that: The training method further comprises: The initial medical image data, the initial physiological signal data and the initial gene sequence data are sequentially subjected to unified formatting, data cleaning, labeling, synchronization and data enhancement processing to obtain the medical image data, the physiological signal data and the gene sequence data accordingly.
6. A method for predicting disease risk, characterized in that: The prediction method comprises: Obtain current medical imaging data, current physiological signal data, and current gene sequence data; The current medical image data, the current physiological signal data and the current gene sequence data are processed by obtaining training data as described in any one of claims 1 to 5 to obtain a current global feature representation; and The current global feature representation is used as an input of the disease prediction model as claimed in any one of claims 1 to 5 to obtain a current disease prediction value.
7. A training device for a disease risk prediction model, characterized in that: The training device comprises: The first acquisition unit is used to obtain training data, specifically including: obtaining data to be processed, the data to be processed includes: medical image data, physiological signal data and gene sequence data, wherein the gene sequence data includes: first data and a first annotation related to the disease in terms of gene sequence, the medical image data includes: second data and a second annotation related to the disease in terms of image, and the physiological signal includes: third data and a third annotation related to the disease in terms of physiological signal; processing to obtain a first feature vector corresponding to the first data, a second feature vector corresponding to the second data and a third feature vector corresponding to the third data; using the first feature vector as the knowledge guiding the second feature vector and the third feature vector, using the knowledge distillation technology to make the second feature vector and the third feature vector consistent with the expression of the first feature vector; and processing the first feature vector, the second feature vector and the third feature vector with consistent expression to obtain a global feature representation after feature weighted fusion; and A training unit, used to train the regression prediction model using the training data, The training data includes: the global feature representation as input, the disease prediction value corresponding to the global feature representation as output, and the true value of the disease.
8. A device for predicting disease risk, characterized in that: The prediction device comprises: A second acquisition unit, used to obtain current medical image data, current physiological signal data and current gene sequence data; a data pre-processing unit, configured to obtain a current global feature representation by processing the current medical image data, the current physiological signal data, and the current gene sequence data through the training data obtained according to any one of claims 1 to 5; and A data post-processing unit is used to use the current global feature representation as an input of the disease prediction model as claimed in any one of claims 1 to 5 to obtain a current disease prediction value.
9. An electronic device, characterized in that: The electronic device comprises: a processor, and a memory coupled to the processor, The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the electronic device performs the training method as described in any one of claims 1 to 5, or performs the prediction method as described in claim 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer executes the training method according to any one of claims 1 to 5, or executes the prediction method according to claim 6.
Citation Information
Patent Citations
Multi-modal data-based cancer prediction method, system and equipment
CN117422704A
Chronic disease risk prediction model training method and system
CN119069127A
Intelligent decision-making method and device based on knowledge retrieval multi-modal large model
CN119357315A
Liver transplantation medicine dosage prediction method based on multiple modes
CN119397492A
Method for establishing robust prediction model, prediction system, and prognostic system for alzheimer's disease
US20230377746A1
Cited By
Complex cardiovascular risk prediction model training method, prediction method, equipment and medium
CN122455375A