A Parkinson's Disease Evolution Prediction Modeling Method Based on Higher-Order Gene Communication Tensors

By constructing a high-order gene communication tensor and an adaptive Transformer architecture, and integrating information from Parkinson's disease-specific gene co-expression networks, the problem of insufficient accuracy in existing early diagnosis models of Parkinson's disease is solved, and higher-precision disease evolution prediction is achieved.

CN121011340BActive Publication Date: 2026-01-30ZHEJIANG SCI-TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511543502.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-01-30
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Existing technologies struggle to fully capture the multi-level, dynamic characteristics of the molecular system of Parkinson's disease using high-throughput transcriptome data, resulting in insufficient accuracy of early diagnostic models for Parkinson's disease.

Method used

We employ a method based on high-order gene communication tensors, combined with multi-scale embedded gene co-expression network analysis and adaptive Transformer architecture, to extract temporal dynamic features and spatial topological features of gene expression. We then integrate the information through a multi-perspective collaborative fusion mechanism to construct a Parkinson's disease evolution prediction model.

Benefits of technology

This improved the predictive accuracy and biological interpretability of early diagnosis models for Parkinson's disease, enabling more precise predictions of disease evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011340B_ABST
    Figure CN121011340B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of disease evolution trajectory prediction in precision medicine, and discloses a method for predicting and modeling the evolution of Parkinson's disease based on high-order gene communication tensors. First, based on intercellular communication mechanisms, a high-order gene module communication tensor is constructed as spatial topological data to characterize the information interaction patterns between gene modules. Second, an adaptive Transformer module is designed to simultaneously extract gene expression features and high-order spatial topological features. Finally, a multi-perspective collaborative fusion mechanism is used to integrate the multi-perspective interaction information between the high-order gene module communication tensor and whole blood transcriptome data, achieving effective feature fusion and utilization. Experimental results show that, compared with traditional gene maps, the high-order spatial representation proposed in this invention has more refined resolution and significantly improves multiple evaluation indicators, providing a new technical means for the accurate staging of Parkinson's disease.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of disease evolution trajectory prediction in precision medicine, and particularly relates to a Parkinson's disease evolution prediction modeling method based on a high-order gene communication tensor. BACKGROUND

[0002] The research on the pathogenesis of neurodegenerative diseases is a key research direction in the field of precision medicine. Such diseases have complex pathological characteristics, involving dynamic synergistic effects of multiple factors such as genetic susceptibility, age-related degenerative changes, and environmental exposure. Notably, these pathogenic factors leave unique molecular imprints at the transcriptional regulation level, making it possible to assess the disease progression stage through high-throughput gene expression profiling. In recent years, the breakthrough progress of single-cell RNA sequencing (scRNA-seq) technology has brought new opportunities for the precision diagnosis of Parkinson's disease (PD), Alzheimer's disease (AD), and other neurodegenerative diseases. This technology can analyze the cellular heterogeneity characteristics in the central nervous system at single-cell resolution and reveal abnormal activation or inhibition states of key regulatory networks.

[0003] Deep mining research based on human whole blood RNA sequencing data has successfully identified multiple transcriptional features closely related to disease development stages. These features include but are not limited to: differentially expressed genes (DEGs), abnormally activated signaling pathways, alternative splicing variants, and regulatory abnormalities of non-coding RNAs. These findings not only significantly deepen our understanding of the molecular mechanisms of neurodegenerative diseases, but more importantly, these detectable transcriptional features in peripheral blood are becoming highly clinically valuable biomarkers, providing new technical approaches for early warning, molecular subtype identification, and treatment response monitoring of brain diseases.

[0004] In terms of data analysis and pattern recognition, graph neural networks (GNNs) have shown significant advantages in high-throughput transcriptome data analysis in recent years due to their excellent data mining capabilities. Although various deep learning methods have been developed to identify abnormal patterns in gene expression data, direct application of traditional deep neural networks to RNA sequencing data often fails to achieve ideal results. This limitation is mainly due to: on the one hand, RNA sequencing data has high dimensionality and high noise characteristics; on the other hand, a single perspective analysis method is difficult to fully capture the multi-level and dynamic changes in the molecular system.

[0005] Recent research progress shows that extracting spatial pattern features from single-cell transcriptome data can not only significantly improve the prediction accuracy of computational models, but also effectively enhance the biological interpretability of the models. In particular, the system biology analysis method based on disease-specific gene networks provides a new research perspective for revealing the molecular mechanisms of disease occurrence and development and discovering potential drug targets. Taking Parkinson's disease as an example, researchers have been able to preliminarily reveal the structural disorder characteristics of its gene regulatory network by constructing a Parkinson's disease-specific gene co-expression network (PD-GCN). However, it is still difficult to establish a high-precision diagnostic model with clinical practical value by relying only on such coarse-grained network information.

[0006] In order to more accurately capture the molecular dynamic change characteristics of the early stage of Parkinson's disease, it is necessary to deeply mine the fine-grained information in single-cell RNA sequencing data, so as to construct a more comprehensive and more predictive computational model. It is worth noting that the PD-GCN network naturally contains topological information of different scales, and by systematically exploring the transmission mode of information flow in the network, it is helpful to understand the structural disorder mechanism of the gene network at high resolution.

[0007] The present application innovatively draws on the analysis paradigm of intercellular communication, and conceptualizes the dynamic communication mode between gene modules as information flow and functional hierarchical features in the disease-specific gene atlas. Based on this idea, a four-dimensional high-order gene module communication tensor, also known as a four-dimensional communication tensor (4D-GMC), is invented, which can effectively capture high-order spatial pattern features in the gene regulatory network.

[0008] The present application creatively applies the high-order 4D-GMC tensor to the structural disorder analysis of the Parkinson's disease-specific gene co-expression network; then uses an adaptive Transformer architecture to extract the temporal dynamic features of gene expression and the spatial topological features of the 4D-GMC tensor, respectively; finally designs a multi-view collaborative fusion mechanism to integrate the multi-view interaction information between the 4D-GMC tensor and the gene expression level through attention weighting, realizing the optimal fusion and collaborative use of features of different scales and different perspectives. SUMMARY

[0009] Based on the problems existing in the background art, the present application proposes a Parkinson's disease evolution prediction modeling method based on high-order gene communication tensors. By deeply studying the whole blood transcriptome data of Parkinson's disease, the two perspectives of data are collaboratively fused and trained, providing a new framework for modeling the evolution of Parkinson's disease.

[0010] To solve the above problems, the application adopts the following technical solutions:

[0011] A Parkinson's disease evolution prediction modeling method based on high-order gene communication tensor, comprising the following steps:

[0012] S1: Assuming that the whole blood transcriptome data at each time is X, containing m patients and n genes; a multi-scale embedded gene co-expression network analysis algorithm (MEGENA) is used to perform co-expression network analysis on the whole blood transcriptome data X at each time to obtain a Parkinson's disease-specific gene atlas G containing the correlation between genes; then a high-order gene communication tensor, i.e. a four-dimensional communication tensor 4D-GMC, is constructed based on the intercellular communication mode;

[0013] S2: A multi-scale adaptive Transformer model is used to extract the time sequence features of the whole blood transcriptome data and the spatial features of the high-order gene communication tensor ;

[0014] S3: A multi-view collaborative fusion scheme is used to integrate the multi-view interaction information between the high-order gene communication tensor and the whole blood transcriptome data, and realize effective fusion and utilization of the features;

[0015] S4: A full connection layer is used as a classifier for Parkinson's disease evolution modeling, and based on the data of the previous time, the motor dysfunction and non-motor dysfunction scores of the next time are predicted.

[0016] Further, the whole blood transcriptome data in step S1 specifically includes but is not limited to:

[0017] S11: Blood transcriptome (sncRNA) data of the Parkinson's Progression Markers Initiative (PPMI);

[0018] S12: Blood transcriptome data of the Parkinson's Disease Biomarker Program (PDBP).

[0019] Further, the basic steps of the multi-scale embedded gene co-expression network analysis algorithm in step S1 include:

[0020] S13: Calculate the correlation coefficient matrix between genes;

[0021] S14: Introduce parallelization, early termination and prior quality control fast planar filtering network construction-FPFNC, which is part of the MEGENA algorithm in the application;

[0022] S15: A multi-scale clustering analysis method is used to divide k gene modules (M1, M2, …, M K), whose formula is:

[0023] ;

[0024] wherein is expressed as a gene module compactness score, used to divide gene modules, k represents the number of gene modules; represents the number of edges within the i-th gene module, represents the total number of edges, represents a resolution parameter.

[0025] Further, the method for calculating the correlation coefficient of the gene pair in step S13 includes but is not limited to

[0026] The first method is to use the Pearson correlation coefficient method to calculate the correlation between genes and genes , and the specific formula is as follows:

[0027] ;

[0028] Assuming that there are m patients in total, the correlation between genes i and j is , wherein the dimension of A is n*n, n is the number of genes, and respectively represent the expression value of the i-th and j-th gene of the k-th patient, and represent the average expression of the i-th and j-th genes;

[0029] The second method is to use the Spearman rank correlation coefficient to calculate the correlation between genes and genes , and the specific formula is as follows:

[0030] ;

[0031] Assuming that there are m patients in total, the correlation between genes i and j is , wherein the dimension of A is n*n, n is the number of genes, and respectively represent the ranking of the expression value of the i-th and j-th gene of the k-th patient;

[0032] The third method is to use the Kendall tau correlation coefficient to calculate the correlation between genes and genes , and the specific formula is as follows:

[0033] ;

[0034] Assuming that there are m patients in total, the correlation between genes i and j is , wherein the dimension of A is n*n, n is the number of genes, and respectively represent the expression values of the i-th and j-th genes of the k-th patient, and respectively represent the expression values of the i-th and j-th genes of the l-th patient, C represents the number of consistent pairs, i.e. the number of inconsistent pairs, i.e. the number of inconsistent pairs, represents the number of i-th gene expression values, represents the number of j-th gene expression values.

[0035] Further, the basic steps of constructing the high-order gene communication tensor in step S1 include:

[0036] S16: Construct a two-dimensional tensor for the k gene modules (M1, M2, …, MK) divided in S15, specifically, for gene a from gene module A, if there is a gene b in gene module B connected to it, it means that there is information interaction between module A and module B, and the communication volume increases by 1; for gene a, the final two-dimensional communication tensor is represented as , and the formula is as follows:

[0037] ;

[0038] where i and j represent the indices of the gene modules;

[0039] S17: After calculating the two-dimensional communication tensor, for the i-th patient, construct a three-dimensional communication tensor using all genes, G represents the total number of genes; the specific construction formula is as follows:

[0040] ;

[0041] where g represents the g-th gene, G represents the total number of genes, represents the two-dimensional communication tensor constructed using the g-th gene;

[0042] S18: Finally, combine the three-dimensional communication tensors of all samples to form the final four-dimensional communication tensor 4D-GMC ; the specific construction formula is as follows:

[0043] ;

[0044] where n represents the n-th patient, N represents the total number of patients, represents the three-dimensional communication tensor constructed using the n-th patient.

[0045] Further, the basic steps of using the multi-scale adaptive Transformer model to extract the space-time features in step S2 include:

[0046] S21: In order to process the whole blood transcriptome data X, the present application designs a linear layer to linearly project X, thereby generating a feature vector of the whole blood transcriptome data ; The formula is as follows:

[0047] ;

[0048] Wherein, Liner represents a linear layer in a neural network model;

[0049] S22: In order to process the whole blood transcriptome data , the present application uses the convolutional neural network and KAN model in combination to extract the feature tensor, and the final result is ;

[0050] S23: Using multi-scale adaptive Transformer to extract features from the whole blood transcriptome scale and high-order communication tensor scale respectively, the two scale features obtained are and , and the formula is:

[0051] ;

[0052] Wherein, the query vector Q, the key vector K, and the value vector V are and respectively obtained by a linear layer, T represents a transpose operation, softmax represents an activation function, represents a parameter.

[0053] Further, the basic steps of using convolution and KAN model in step S22 include:

[0054] S221: The four-dimensional gene module communication tensor is first input to a convolution layer, which has 2000 input channels and 1024 output channels; then, an activation function is applied to the result output , and the formula is:

[0055] ;

[0056] Wherein, ReLU represents an activation function, and Conv represents a convolution layer;

[0057] S222: Then, an average pooling layer is used to reduce the dimension of to , and then the obtained tensor is flattened into two dimensions, and the formula is as follows:

[0058] ;

[0059] wherein, view denotes the flatten operation, AvgPool denotes the average pooling layer operation, denotes the feature representation after the average pooling layer extraction;

[0060] S223: Finally, the KAN layer is used to extract the nonlinear relationship existing in the high-dimensional data, and the formula is as follows:

[0061] ;

[0062] wherein, KAN denotes the KAN layer, denotes the result after the KAN model feature extraction.

[0063] Further, the basic steps of multi-view collaborative fusion in step S3 include:

[0064] S31: The core of multi-view collaborative fusion is cross-attention mechanism, and the formula is as follows:

[0065] ;

[0066] ;

[0067] ;

[0068] wherein, respectively denote the query (Q1), key (K1) and value (V1) derived from the first view; ; similarly, ; similarly, respectively denote the query (Q2), key (K2) and value (V2) derived from the second view; ; and respectively denote the feature matrix generated after the first cross-attention and the second cross-attention, and the final result obtained after the two cross-attentions is ;

[0069] S32: Then, the feature fusion is performed on and to obtain an enhanced matrix containing more information, and the formula is as follows:

[0070] ; ​​​​​

[0071] where Concat denotes the feature concatenation operation.

[0072] Further, the basic steps of using the fully connected layer to identify the disease stage in the S4 step include:

[0073] S41: using a fully connected layer with a softmax activation function to perform classification prediction, whose formula is as follows:

[0074] ;

[0075] where the softmax function converts the output of the fully connected layer into a class probability, and FC denotes the fully connected layer.

[0076] S42: using a cross-entropy loss function to extract the characteristics and commonalities between cross-view angles, whose mathematical formula is as follows:

[0077] ;

[0078] where L denotes the loss function, which quantifies the difference between the predicted probability distribution and the true label distribution output by the model, and reflects the prediction accuracy of the model, is a hyperparameter, is and the cross-entropy loss function between is the cross-entropy loss function between and is the cross-entropy loss function between and

[0079] Further, in the step S4, the accuracy (ACC), F1-score, recall (Recall), and precision (Precision) are used to evaluate the accuracy of the model; among them, the accuracy refers to the proportion of the number of samples correctly predicted by the classifier to the total number of samples, which can directly evaluate the overall effect of the model; the recall rate, also known as the total rate, indicates the ratio of the number of positive examples correctly predicted by the classifier to the total number of actual positive examples, which evaluates the ability of the model to find all relevant instances; the F1-score is the harmonic mean of the accuracy and the integrity of the model; the precision indicates the proportion of the number of samples correctly predicted as positive examples by the classifier to the number of samples predicted as positive examples, which focuses on the accuracy of the model predicting positive examples. Its formula is as follows:

[0080] ;

[0081] ;

[0082] ;

[0083] ;

[0084] Wherein TP is true positive, TN is true negative, FP is false positive, and FN is false negative, the greater the indicators, the more accurate the predicted results.

[0085] Compared with the prior art, the present application has the beneficial effects that:

[0086] (1) The present application uses a multi-scale embedded gene co-expression network analysis method to generate a silver co-expression network. Then based on intercellular communication mode analysis, a high-order gene communication tensor is constructed for capturing high-order spatial patterns in the gene network, i.e. information interaction patterns between gene modules. The constructed communication tensor contains more information than the traditional gene co-expression network and has higher resolution.

[0087] (2) The present application first designs an adaptive Transformer module, introduces a learnable adaptive weight mechanism, and can simultaneously extract gene expression features and high-order spatial topology features. Finally, a collaborative fusion module based on cross-attention mechanism is constructed to realize deep integration and complementary enhancement of gene expression features and spatial information. BRIEF DESCRIPTION OF DRAWINGS

[0088] Figure 1 is the flowchart of Parkinson's disease evolution prediction based on high-order gene communication tensor in the specific embodiments of the present application;

[0089] Figure 2 is the multi-view multi-scale Transformer framework in the specific embodiments of the present application;

[0090] Figure 3 is the box plot of comparison between different baseline methods in the Motor PPMI case in the embodiments of the present application;

[0091] Figure 4 is the communication mode visualization of the Motor PPMI case at the initial treatment time in the embodiments of the present application;

[0092] Figure 5 is the communication mode visualization of the Motor PDBP case at the six-month follow-up time in the embodiments of the present application;

[0093] Figure 6 is the gene co-expression network visualization of the Motor PPMI case in the embodiments of the present application;

[0094] Figure 7is the gene co-expression network visualization of the Motor PDBP case in the embodiments of the present application. DETAILED DESCRIPTION

[0095] The specific embodiments of the present application are described below to facilitate the understanding of the present application by any person skilled in the art, but the protection scope of the present application is not limited to the range of the specific embodiments, and any person skilled in the art should cover the protection scope of the present application according to the technical solutions and inventive concept of the present application and equivalent replacement or change.

[0096] The technical solutions of the present application are further specifically described below by embodiments and in conjunction with the drawings.

[0097] Embodiment 1

[0098] As shown in Figure 1 , a Parkinson's disease evolution prediction modeling method based on high-order gene communication tensor includes the following steps:

[0099] S1: Assuming that the whole blood transcriptome data at each time is X, containing m patients and n genes; using a multi-scale embedded gene co-expression network analysis algorithm (MEGENA) to perform co-expression network analysis on the whole blood transcriptome data X at each time, to obtain a Parkinson-specific gene graph G containing the correlation between genes, and then using a method based on intercellular communication mode to construct a high-order gene communication tensor (4D-GMC);

[0100] S2: using a multi-scale adaptive Transformer model to extract the time sequence characteristics of the whole blood transcriptome data and the spatial characteristics of the high-order gene communication tensor ;

[0101] S3: using a multi-view collaborative fusion scheme to integrate the multi-view interaction information between the high-order gene communication tensor and the whole blood transcriptome data, to realize effective fusion and utilization of the characteristics;

[0102] S4: using a fully connected layer as a classifier for Parkinson's disease evolution modeling, to predict the next time disease motor dysfunction and non-motor dysfunction scores based on the data of the previous time.

[0103] The whole blood transcriptome data in step S1 specifically includes but is not limited to:

[0104] S11: blood transcriptome (sncRNA) data of Parkinson's disease progression marker initiative (PPMI);

[0105] S12: blood transcriptome data of Parkinson's disease data and biomarker program (PDBP).

[0106] The basic steps of the multi-scale embedded gene co-expression network analysis algorithm in step S1 include:

[0107] S13: Calculate the correlation coefficient matrix between genes;

[0108] S14: Introduce parallelization, early termination and prior quality control of fast planar filtering network construction-FPFNC, which is part of the MEGENA algorithm;

[0109] S15: Use multi-scale clustering analysis method to divide k gene modules (M1, M2, …, Mk); K ), the formula is:

[0110] ;

[0111] Wherein is the gene module compactness score, which is used to divide gene modules, k represents the number of gene modules, is the number of edges in the i-th gene module, is the total number of edges, is the resolution parameter.

[0112] The method for calculating the correlation coefficient of gene pairs in step S13 includes but is not limited to

[0113] S131: The first method is to use Pearson correlation coefficient method to calculate the correlation between genes , the specific formula is as follows:

[0114] ;

[0115] Assuming there are m patients in total, the correlation between gene i and gene j is , wherein the dimension of A is n*n, n is the number of genes, and represent the expression value of the i-th and j-th gene of the k-th patient, and represent the average expression of the i-th and j-th genes;

[0116] S132: The second method is to use Spearman rank correlation coefficient to calculate the correlation between genes , the specific formula is as follows:

[0117] ;

[0118] Assuming there are m patients in total, the correlation between gene i and gene j is where A is n*n, n is the number of genes, and respectively represent the rank of the i and j gene expression value of the kth patient;

[0119] S133: The third way is to use Kendall tau correlation coefficient to calculate the correlation between genes , the specific formula is as follows:

[0120] ;

[0121] Assuming that there are m patients in total, the correlation between gene i and gene j is where A is n*n, n is the number of genes, and respectively represent the expression value of the i and j gene of the kth patient, and respectively represent the expression value of the i and j gene of the lth patient, C represents the number of consistent pairs, that is the number of, D represents the number of inconsistent pairs, that is the number of, represents the number of i gene expression values, represents the number of j gene expression values.

[0122] Further, the basic steps of constructing the high-order gene communication tensor in step S1 include:

[0123] S16: Constructing a two-dimensional tensor for the k gene modules (M1, M2, …, MK) divided in S15; specifically, for gene a from gene module A, if there is gene b in gene module B connected to it, it means that there is information interaction between module A and module B, and the communication amount increases by 1. For gene a, the final two-dimensional communication tensor is represented as , the formula is as follows:

[0124] ;

[0125] where i and j represent the index of the gene module;

[0126] S17: After calculating the two-dimensional communication tensor, for the ith patient, use all genes to construct a three-dimensional communication tensor , G represents the total number of genes; the specific construction formula is as follows:

[0127] ;

[0128] where g represents the gth gene, G represents the total number of genes, denotes a two-dimensional communication tensor constructed using the gth gene;

[0129] S18: Finally, the three-dimensional communication tensors of all samples are combined to form a final four-dimensional communication tensor (4D-GMC) ; The specific construction formula is as follows:

[0130] ;

[0131] wherein n represents the nth patient, N represents the total number of patients, denotes a three-dimensional communication tensor constructed using the nth patient.

[0132] Further, the basic steps of using a multi-scale adaptive Transformer model to extract spatio-temporal features in step S2 include:

[0133] S21: In order to process the whole blood transcriptome data X, the present application designs a linear layer to linearly project X, thereby generating a feature vector of the whole blood transcriptome data , whose formula is as follows:

[0134] ;

[0135] wherein Liner represents a linear layer in a neural network model;

[0136] S22: In order to process the whole blood transcriptome data , the present application uses a convolutional neural network and a KAN model in combination to extract feature tensors, and the final result is ;

[0137] S23: Using a multi-scale adaptive Transformer, the whole blood transcriptome scale and the high-order communication tensor scale are respectively extracted, and the two scale features finally obtained are and , whose formula is as follows:

[0138] ;

[0139] wherein the query vector Q, the key vector K, and the value vector V are respectively and obtained through a linear layer, T represents a transposition operation, softmax represents an activation function, denotes a parameter.

[0140] The basic steps of using a convolution and a KAN model in step S22 include:

[0141] S221: Four-dimensional gene module communication tensor is first input to a convolutional layer with 2000 input channels and 1024 output channels, and then an activation function is applied to the resulting output , where

[0142] ;

[0143] , where ReLU denotes the activation function and Conv denotes the convolutional layer

[0144] S222: Then, an average pooling layer is used to reduce the dimension of to , and then the resulting tensor is flattened into two dimensions, as follows:

[0145] ;

[0146] , where view denotes the flattening operation, AvgPool denotes the average pooling layer operation, denotes the feature representation after the average pooling layer

[0147] S223: Finally, a KAN layer is used to extract the nonlinear relationship present in the high-dimensional data, as follows:

[0148] ;

[0149] , where KAN denotes the KAN layer, denotes the result after KAN model feature extraction.

[0150] Further, the basic steps of multi-view collaborative fusion in step S3 include:

[0151] S31: The core of multi-view collaborative fusion is the cross-attention mechanism, as follows:

[0152] ;

[0153] ;

[0154] ;

[0155] , where denote the query ( ), key ( ) and value ( ) derived from , respectively; similarly, denote the query ( ), key ( ) and value ( ) derived from , respectively; and are the feature matrices generated after the first cross-attention and the second cross-attention, respectively; the final result after the two cross-attentions is ;

[0156] S32: Then, the and are fused to obtain an enhanced matrix containing more information , whose formula is as follows:

[0157] ;

[0158] where Concat represents the feature concatenation operation.

[0159] The basic steps of using the full connection layer to perform disease stage recognition in S4 include:

[0160] S41: The full connection layer is used in combination with the softmax activation function to perform classification prediction on , and its formula is as follows:

[0161] ;

[0162] where the softmax function converts the output of the full connection layer into a class probability, and FC represents the full connection layer.

[0163] S42: The cross-entropy loss function is used to extract the characteristics and commonalities between the cross-view angles, and its mathematical formula is as follows:

[0164] ;

[0165] where L represents the loss function, which quantifies the difference between the predicted probability distribution and the true label distribution of the model output, reflecting the prediction accuracy of the model, is a hyperparameter, is the cross-entropy loss function between and , is the cross-entropy loss function of , is the cross-entropy loss function of , and is the cross-entropy loss function of .

[0166] ​In the step S4, the accuracy of the model is evaluated by using the four indexes of accuracy (ACC), F1-score, recall and precision. The accuracy is the proportion of the number of samples correctly predicted by the classifier to the total number of samples, which can directly evaluate the overall effect of the model. The recall, also known as the recall rate, is the proportion of the number of positive examples correctly predicted by the classifier to the total number of positive examples, which evaluates the ability of the model to find all relevant instances. The F1-score is the harmonic mean of precision and recall, which considers the accuracy and integrity of the model. The precision is the proportion of the number of samples correctly predicted as positive by the classifier to the number of samples predicted as positive, which focuses on the accuracy of the model in predicting positive examples. The formula is as follows:

[0167] ;

[0168] ;

[0169] ;

[0170] ;

[0171] where TP is true positive, TN is true negative, FP is false positive, and FN is false negative. The larger these indexes are, the more accurate the prediction result is.

[0172] Example 2

[0173] The hardware environment for running this example includes one notebook computer, CPU: Xeon(R) Gold 6430, RAM: 24.0GB; the software environment includes Python 3.6 and R 3.3.3; and the operating platform is Win10. The communication mode visualization of the Motor PPMI case at the initial treatment time is as shown in Figure 4 , the communication mode visualization of the Motor PDBP case at the six-month follow-up time is as shown in Figure 5 , the gene co-expression network visualization of the Motor PPMI case is as shown in Figure 6 , and the gene co-expression network visualization of the Motor PDBP case is as shown in Figure 7 .

[0174] The present embodiment tests the effect of the proposed method on the whole blood transcriptome dataset, which includes: PPMI dataset and PDBP data, and analyzes from two aspects of motor function assessment (Motor PPMI and Motor PDBP) and non-motor function assessment (Non-motor PPMI and Non-motor PDBP); in the aspect of motor function assessment, the Hoehn and Yahr scale label is used for evaluation; this scoring scale is based on the severity of motor symptoms and functional impairment, and is a commonly used disease staging system for assessing the severity and progression of the disease, as follows:

[0175] Stage 0: No symptoms;

[0176] Stage 1: Only unilateral / side of the body is affected, but balance is not affected;

[0177] Stage 2: Only bilateral / side of the body is affected, but balance is not affected;

[0178] Stage 3: Balance is affected, mild to moderate disease; but the patient can live independently;

[0179] Stage 4: Severe immobility. But the patient can walk and stand by himself;

[0180] Stage 5: Can only lie in bed or sit in a wheelchair without the help of others.

[0181] In the aspect of non-motor assessment, the cognitive dysfunction label in UPDRS-I is adopted, which mainly includes cognitive function (such as memory, attention, executive function, etc.), emotional state (depression, anxiety) and autonomic nervous symptoms (constipation, urinary dysfunction) and the like; this label is divided into five categories in total, as follows:

[0182] Stage 0: No cognitive impairment;

[0183] Stage 1: Mild forgetfulness, no impact on daily life;

[0184] Stage 2: Moderate impairment of memory or executive function, with mild impact on daily life (such as forgetting appointments, difficulty making decisions);

[0185] Stage 3: Severe cognitive impairment, requiring assistance to complete complex tasks (such as financial management, medication);

[0186] Stage 4: Dementia (unable to live independently, requiring full-time care).

[0187] The top 2000 high variable genes were used as features in all four cases, and the top 10 important gene modules were selected to construct the 4D-GMC. In terms of data division, the training set used the gene expression data of the patients at the baseline (BL) stage, i.e., the initial stage of treatment, and the test set used the time series data at the sixth month (M6) of follow-up. The constructed 4D-GMC information is shown in Table 1.

[0188] Table 1 High-order gene communication tensor information

[0189] .

[0190] The Parkinson's disease evolution prediction method based on the high-order gene communication tensor proposed in this embodiment is shown in the flowchart as Figure 1 shown. The implementation process of this method is illustrated by taking the motor evaluation (Motor PPMI) of the whole blood transcriptome data PPMI dataset as an example, which specifically includes the following steps:

[0191] S1: First, the original PPMI dataset is preprocessed, the whole blood transcriptome data of multiple disease stages of the patient is separated to obtain the whole blood transcriptome data X of two disease time points of the same batch of patients, the number of patients is 4425, then the top 2000 high variable genes are selected as features using the scanpy package; Finally, the whole blood transcriptome data at the two time points is analyzed using the multi-scale embedded gene co-expression network analysis algorithm (MEGENA) to obtain the relevant adjacency matrix representation , the size is 2000 2000, and the calculation formula is:

[0192] ;

[0193] Where m represents the number of patients, the size is 4425; the correlation between gene i and gene j is , n represents the number of genes, the size is 2000; and represent the expression values of the i-th and j-th genes of the k-th patient; and represent the average expression of the i-th and j-th genes; In addition, the gene co-expression network contains k gene module information, and the calculation formula is:

[0194] ;

[0195] Where represents the gene module compactness score, which is used to divide the gene module; k represents the number of gene modules; represents the number of edges in the i-th gene module. represents the total number of edges; represents the resolution parameter.

[0196] S2: Based on the intercellular communication mode analysis method, a high-order gene communication tensor is constructed , Figure 1 The specific technical scheme is shown in the flowchart:

[0197] S21: According to the gene importance score obtained by the MEGENA algorithm, the top ten gene modules are selected to obtain a two-dimensional interaction matrix between gene modules with a size of 10*10; if a gene in module A and a gene b in module B have an edge connected, it indicates that module A and module B have information interaction, and the value of the corresponding bit in the interaction matrix increases by one; the calculation formula is:

[0198] ;

[0199] In the above formula, i and j represent the index of the gene module, represents the generated interaction matrix between gene modules;

[0200] S22: An interaction matrix between gene modules is generated for each gene, thereby forming a three-dimensional matrix with a size of 2000 10 The calculation formula is:

[0201] ;

[0202] In the above formula, g represents the index of the gene, represents the interaction matrix of the gth gene, represents the generated three-dimensional matrix;

[0203] S23: The three-dimensional matrix generated in step S22 is for a single patient, therefore, a three-dimensional matrix is generated for each patient according to the previous steps, thereby generating a four-dimensional matrix, i.e., a high-order gene communication tensor, with a size of 4425 2000 10 The calculation formula is:

[0204] ;

[0205] In the above formula, n represents the index of the patient, represents the three-dimensional matrix of the nth patient, represents the generated high-order gene communication tensor.

[0206] S3: Whole blood transcriptome data and high-order gene communication tensor As two scale information, the application adopts a multi-scale adaptive Transformer model to extract the feature information of the two, as shown in Figure 2 ; The specific technical scheme is:

[0207] S31: For whole blood transcriptome data X, linear projection is performed using a linear layer to obtain its feature vector , and the calculation formula is:

[0208] ;

[0209] Wherein, Liner represents a linear layer in a neural network model;

[0210] S32: For whole blood transcriptome data , first use a convolutional neural network for dimension reduction, and the calculation formula is:

[0211] ;

[0212] Wherein, ReLU represents an activation function, and Conv represents a convolutional layer with 2000 input channels and 1024 output channels, and then use an average pooling layer to flatten to two dimensions, and the calculation formula is:

[0213] ;

[0214] Wherein, view represents a flattening operation, AvgPool represents an average pooling layer operation, represents the feature representation extracted after the average pooling layer. Finally, use KAN layer to extract the nonlinear relationship existing in , and the calculation formula is:

[0215] ;

[0216] Wherein, KAN represents a KAN layer, represents the result after KAN model feature extraction;

[0217] S33: Use a multi-scale adaptive Transformer to extract features of whole blood transcriptome scale and high-order communication tensor scale respectively, and the final two scale features are and , and the calculation formula is:

[0218] ;

[0219] Wherein, the query vector Q, the key vector K, and the value vector V are and ​are obtained by linear layers respectively, T represents the transpose operation, and softmax represents the activation function, represent parameters.

[0220] S4: A multi-view collaborative fusion scheme is adopted to integrate the multi-view interaction information between the high-order gene communication tensor and the whole blood transcriptome data, and the core principle is the cross attention mechanism; the specific technical scheme is as follows:

[0221] S41: First, a plurality of cross attention layers are used to extract feature information in each view, and the calculation formula is as follows:

[0222] ;

[0223] ;

[0224] ;

[0225] wherein, respectively represent the query (Q), key (K) and value (V) derived from the first view. Similarly, respectively represent the query (Q), key (K) and value (V) derived from the second view. and respectively represent the feature matrix generated after the first cross attention and the second cross attention, and T represents the transpose operation for the matrix vector. The final result obtained after the two cross attentions is as follows: ; S42: Then, the feature splicing is performed on and to obtain an enhanced matrix containing more information, and the calculation formula is as follows:

[0226] ; wherein, Concat represents the feature splicing operation; S5: The full connection layer is used in combination with the softmax activation function to perform classification prediction on

[0227] , and the calculation formula is as follows:

[0228] ;

[0229] wherein

[0230] ;

[0231] ​​​​​​The softmax function, representing the final prediction result of the model, transforms the output of the fully connected layer into class probabilities. FC stands for Fully Connected Layer. The model uses the cross-entropy loss function to extract characteristics and commonalities across viewpoints, and its calculation formula is as follows:

[0232] ;

[0233] Where L represents the loss function of the neural network model, which reflects the model's prediction accuracy by quantifying the difference between the predicted probability distribution output by the model and the true label distribution; the smaller the difference, the lower the loss value. It is a hyperparameter. for and The cross-entropy loss function between them For gene embedding The cross-entropy loss function, For higher-order communication tensors The cross-entropy loss function.

[0234] Finally, the actual motor / non-motor function assessment labels y and the model-predicted labels are compared. The model's accuracy was evaluated using four metrics: accuracy (ACC), F1-score, recall, and precision. Accuracy refers to the proportion of correctly predicted samples out of the total number of samples, providing a direct assessment of the model's overall performance. Recall, also known as the full detection rate, represents the ratio of correctly predicted positive instances to the actual number of positive instances, evaluating the model's ability to find all relevant instances. The F1-score is the harmonic mean of precision and recall, comprehensively considering both accuracy and completeness. Precision represents the proportion of correctly predicted positive samples out of the total number of predicted positive samples, focusing on the accuracy of the model's positive predictions. Its calculation formula is as follows:

[0235] ;

[0236] ;

[0237] ;

[0238] ;

[0239] TP represents a true positive, TN represents a true negative, FP represents a false positive, and FN represents a false negative. The higher these indicators are, the more accurate the prediction results.

[0240] Example 3

[0241] The main contribution of the present application is to construct a high-order gene communication tensor, and use an adaptive Transformer model to extract feature information at two scales, and finally construct a collaborative fusion scheme to fuse the characteristics within the perspective and the commonality between the perspectives, and extract information from past period data to predict future disease status. In the present application, the M2GMC model (see Figure 1 The M2GMC model is a M2GMC method based on spatiotemporal joint perspective, which integrates the mixed perspective of temporal expression pattern and high-order communication tensor (corresponding to spatial perspective). In order to test the effectiveness of the model, six baseline methods are selected as comparative models:

[0242] (1) CNN: Convolutional Neural Network is a typical deep learning architecture mainly applied in visual data processing. The network automatically extracts spatial features of images through multi-layer convolution operation, combines with down-sampling operation to compress feature dimension, and finally realizes mapping of high-level semantic information and task prediction by using full connection layer.

[0243] (2) Transformer: Its core architecture is based on multi-head self-attention mechanism and feed-forward neural network, which realizes context modeling and sequence generation of input data through stacking of multi-layer encoder and decoder.

[0244] (3) ACmix: The model reveals the strong potential relationship between Self-Attention and convolution, provides a new perspective for understanding the relationship between the two modules, and provides inspiration for designing new learning paradigms.

[0245] (4) MAET: The model is a novel multi-modal adaptive emotion Transformer, which can adaptively process single modal and multi-modal input.

[0246] (5) Meta-Transformer: The model consists of three main components: a unified data tokenizer, a modality-shared encoder and a task-specific head for downstream tasks. Meta-Transformer is the first framework that can perform unified learning on 12 modalities and use unpaired data.

[0247] (6) Perceiver: The model adopts a single Transformer-based architecture to manipulate inconsistent arrangements of different modalities.

[0248] The experimental results of the present application in Motor PPMI, Motor PDBP, Non-motor PPMI and Non-motor PDBP cases are shown in Tables 2-5:

[0249] Table 2 Motor evaluation case of PPMI data set

[0250] .

[0251] Table 3. Motor assessment cases for PDBP dataset

[0252] .

[0253] Table 4. Non-motor assessment cases for PPMI dataset

[0254] .

[0255] Table 5. Non-motor assessment cases for PDBP dataset

[0256] .

[0257] The M2GMC model proposed in this embodiment achieves good results on the four cases. First, single-view models (including CNN, Transformer, and ACmix) are used to independently test high-order communication tensors and whole blood transcriptome data. Subsequently, dual-view models (including MAET, Meta-Transformer, and Perceiver) are used to jointly analyze and compare high-order communication tensors and whole blood transcriptome data. The specific experimental results are shown in Tables 2-5 and Figure 3 According to the experimental results, in the single-view model, the model using only high-order communication tensors has a slightly lower prediction effect than the model using only whole blood transcriptome data, which indicates the effectiveness of high-order communication tensors. The experimental results show that in the single-view model, the model using only high-order communication tensors has a slightly lower prediction effect than the model using only whole blood transcriptome data, which preliminarily verifies the effectiveness of high-order communication tensors. In dual-view learning, the model performance is significantly better than the single-view model using only high-order communication tensors or gene expression patterns, which indicates that there is complementarity between the two data types, and through multi-view joint analysis, the feature representation capability can be effectively enhanced, thereby improving the prediction performance of the model. For example, on the Motor PDBP dataset, the accuracy of M2GMC reaches 0.732, which is 7.9% higher than the MAET method, and significantly better than other methods.

[0258] Although the candidate PD modeling methods based on deep learning (such as ACmix and MAET) show certain competitiveness, there are still deficiencies in effectively capturing complex patterns in the dataset. The excellent performance of M2GMC is due to its advanced design, which integrates the relationship between views and perspectives into a unified framework, thereby enhancing its ability to more comprehensively model the underlying data structure. The evaluation indicators obtained by M2GMC and the candidate methods prove the effectiveness of M2GMC in various applications, and emphasize the positive role of inter-module communication patterns in disease modeling.

[0259] To further explore the role of high-order communication tensors, multi-scale adaptive modules, and collaborative fusion modules, the M2GMC model underwent ablation analysis on four sets of whole blood transcriptome datasets. The experimental design used a strict control method, with each module undergoing individual and combined ablation to evaluate the contribution of each module to the model. The specific experimental results are shown in Tables 6 and 7.

[0260] Table 6 Ablation experiment on motor function evaluation dataset

[0261] .

[0262] Table 7 Ablation experiment on non-motor function evaluation dataset

[0263] .

[0264] The experimental data showed that after integrating the communication tensor, the performance indicators of the model in three different application scenarios were significantly improved, which fully confirmed the effectiveness of the high-order communication tensor in improving the efficiency of the model.

[0265] In addition, the experiment also analyzed the influence of the multi-scale adaptive module and the multi-view collaborative module on the performance of the model. The comparison results showed that both had a positive enhancement effect on the model.

[0266] To further verify the irreplaceability of the four-dimensional communication tensor, a comparative experiment was designed: using a three-dimensional communication tensor (patient x sending module x receiving module) to replace the original four-dimensional communication tensor (patient x gene x sending module x receiving module). The experimental results are shown in Table 8:

[0267] Table 8 Performance comparison of four-dimensional and three-dimensional communication tensors

[0268] .

[0269] From Table 8, it can be found that in all cases, the four-dimensional communication tensor is superior to the three-dimensional communication tensor in all four evaluation indicators, which confirms the independent contribution of gene-specific features to the performance of the model.

[0270] The above examples are used to explain and illustrate the present application, but not to limit the present application. Any modifications and changes made to the present application within the spirit and protection scope of the claims fall within the protection scope of the present application.

Claims

1. A Parkinson evolution prediction modeling method based on high-order gene communication tensors, characterized in that, The method comprises the following steps: S1: the whole blood transcriptome data at each time point is X, containing m patients and n genes; A multi-scale embedded gene co-expression network analysis algorithm MEGENA is used to analyze the whole blood transcriptome data X at each time point to obtain a Parkinson-specific gene atlas G containing the correlation between genes and genes; Then a high-order gene communication tensor, i.e. a four-dimensional communication tensor 4D-GMC, is constructed based on a cell intercommunication mode; The steps of the multi-scale embedded gene co-expression network analysis algorithm include: Step S11: calculating the correlation coefficient matrix between genes; Step S12: parallelization, early termination and fast plane filtering network construction-FPFNC with prior quality control; Step S13: divide the k gene modules (M1, M2, …, M K ) by using a multi-scale clustering analysis method, and the formula is: ; wherein is denoted as the gene module compactness score, k denotes the number of gene modules; denotes the number of edges within the i-th gene module, denotes the total number of edges, denotes the resolution parameter; The basic steps of constructing the high-order gene communication tensor include: S14: Constructing two-dimensional tensor for k gene modules M1, M2, …, MK divided in S13; for gene a from gene module A, if there is gene b in gene module B connected with it, it means that there is information interaction between module A and module B, and the communication volume increases by 1; for gene a, the finally constructed two-dimensional communication tensor is represented as The formula is as follows: ; Where i and j represent the indexes of the gene modules, respectively; S15: After the two-dimensional communication tensor is calculated, for the i-th patient, a three-dimensional communication tensor is constructed using all genes The construction of the three-dimensional communication tensor G represents the total number of genes; the construction formula is as follows: ; wherein g denotes the gth gene, G denotes the total number of genes, denotes the two-dimensional gene communication tensor constructed using the gth gene; S16: Finally, the three-dimensional communication tensors of all samples are combined to form the final four-dimensional communication tensor 4D-GMC ; the formula is constructed as follows: ; where n denotes the nth patient, and N denotes the total number of patients, denotes the three-dimensional communication tensor constructed using the nth patient; S2: adopt multi-scale adaptive Transformer model to extract the time sequence features of whole blood transcriptome data and the spatial features of high-order gene communication tensor, respectively ​​ S3: a multi-view collaborative fusion scheme is used to integrate the multi-view interaction information between the high-order gene communication tensor and the whole blood transcriptome data, so as to realize effective fusion and utilization of features; S4: a full connection layer is used as a classifier for Parkinson disease evolution modeling, and the data at the previous time point is used to predict the motor dysfunction and non-motor dysfunction scores at the next time point.

2. The Parkinson evolution prediction modeling method based on high-order genetic communication tensor according to claim 1, characterized in that, The types of the whole blood transcriptome data in step S1 include but are not limited to: The blood transcriptome sncRNA data of the Parkinson's Disease Progression Markers Initiative PPMI and the blood transcriptome data of the Parkinson's Disease Biomarkers Program PDBP.

3. The Parkinson evolution prediction modeling method based on high-order genetic communication tensor according to claim 1, characterized in that, The calculation method of the correlation coefficient matrix between genes in step S11 is any one of the following ways: The first way is to calculate the correlation between genes and genes using Pearson correlation coefficient method The formula is as follows: ; m represents the number of patients, and the correlation between gene i and gene j is where A has dimension n*n, n is the number of genes, and Eikand Ejrepresent the expression value of the ith and jth gene of the kth patient, respectively, and Eiand Ejrepresent the average expression of the ith and jth gene, respectively. The second way is to calculate the correlation between genes and genes using Spearman rank correlation coefficient The specific formula is as follows: ; m represents the number of patients, and the correlation between gene i and gene j is where A has dimension n*n, n is the number of genes, and ranked expression values of the ith and jth genes of the kth patient, respectively. The third method is to calculate the correlation between genes and genes using the Kendall tau correlation coefficient The specific formula is as follows: ; m represents the number of patients, and the correlation between gene i and gene j is where A has dimension n*n, n is the number of genes, and Xi k represents the expression value of the ith gene of the kth patient, and Xi l represents the expression value of the ith gene of the 1th patient, C represents the number of consistent pairs, i.e. the number of consistent pairs, D represents the number of inconsistent pairs, i.e. the number of inconsistent pairs, Xi represents the number of the same expression values of the ith gene, Xj represents the number of the same expression values of the jth gene.

4. The Parkinson evolution prediction modeling method based on high-order genetic communication tensor according to claim 1, characterized in that, The basic steps of using the multi-scale adaptive Transformer model in step S2 include: S21: To process the whole blood transcriptome data X, a linear layer is utilized to linearly project X, resulting in a feature vector of the whole blood transcriptome data ; the formula is as follows: ; Where Liner represents a linear layer in the neural network model; S22: To process whole blood transcriptome data , the feature tensor extraction is carried out in the way of convolutional neural network and KAN model combination, and the final result is ; S23: Use multi-scale adaptive Transformer to extract features on the whole blood transcriptome scale and high-order communication tensor scale respectively, and the final two scale features are and The formula is: ; where the query vector Q, the key vector K, and the value vector V are obtained respectively by and linear layers; T represents the transpose operation, and softmax represents the activation function, denote parameters.

5. The Parkinson evolution prediction modeling method based on high-order genetic communication tensor according to claim 4, characterized in that, The basic steps of using the convolutional neural network and KAN model combination to extract the feature tensor in step S22 include: S221: four-dimensional genetic module communication tensor First input to a convolutional layer with 2000 input channels, 1024 output channels; subsequently, an activation function is applied to the resulting output The formula is: ; Where ReLU represents an activation function, and Conv represents a convolutional layer; S222: Subsequently, an average pooling layer is used to reduce the dimension of from to ; then the resulting tensor is flattened to two dimensions, whose formula is as follows: ; wherein, view denotes a flattening operation, AvgPool denotes an average pooling layer operation, denotes the feature representation after the average pooling layer extraction; S223: the KAN layer is used to extract the nonlinear relationship existing in the high-dimensional data, and the formula is as follows: ; wherein KAN represents the KAN layer, represents Results after KAN model feature extraction.

6. The Parkinson evolution prediction modeling method based on high-order genetic communication tensor according to claim 1, characterized in that, The steps of using the multi-view collaborative scheme to perform feature fusion in step S3 include: S31: a plurality of cross-attention layers are used to extract features from two views, and the formula is as follows: ; ; ; wherein, respectively represent queries, keys and values derived from ; , keys and values derived from ; ; similarly, respectively represent queries, keys and values derived from ; , keys and values derived from ; ; and respectively are feature matrices generated after the first cross-attention and the second cross-attention, T represents a transpose operation for matrix vectors; the final result obtained after the two cross-attentions is ; S32: Subsequently, the following is performed and feature fusion is performed to obtain an enhanced matrix containing more information ; the formula is as follows: ; Where Concat represents a feature concatenation operation.

7. The high-order gene communication tensor based Parkinson's evolution prediction modeling method according to claim 1, characterized in that, The basic steps of using the full connection layer to perform disease stage recognition in step S4 include: S41: Using a fully connected layer with a softmax activation function The formula for classification prediction is as follows: ; where the softmax function converts the output of the fully connected layer into class probabilities, FC denotes the fully connected layer, denotes the final prediction of the model; S42: a cross-entropy loss function is used to extract the characteristics and commonalities between cross-views, and the mathematical formula is as follows: ; wherein L represents a loss function reflecting the prediction accuracy of the model by quantifying the difference between the predicted probability distribution output by the model and the true label distribution, is a hyperparameter, is and a cross-entropy loss function between is a cross-entropy loss function between is a cross-entropy loss function between Finally, the real motion / non-motion function evaluation label y and the model predicted label The accuracy of the model is evaluated by comparing the accuracy ACC, F1-score, recall, and precision. The F1-score is the harmonic mean of precision and recall. The precision indicates the proportion of correctly predicted positive samples in the predicted positive samples, and the formula is as follows: ; ; ; ; Where TP is true positive, TN is true negative, FP is false positive, and FN is false negative.

Citation Information

Patent Citations

  • IncRNA-disease association prediction method based on high-order proximity and matrix completion algorithm

    CN113160880A

  • Microorganism-disease relevance prediction method based on multi-order similarity fusion learning

    CN117219173A