A multi-modal omics feature fusion cervical cancer risk prediction system

By constructing a cervical cancer-specific heterogeneity map and a spatial micro-domain latent distribution dictionary, and combining it with a cross-modal gating network, the biological noise and modality loss problems of existing cervical cancer multimodal omics prediction systems were solved, achieving high-precision and robust risk prediction.

CN122637902APending Publication Date: 2026-08-25HANGZHOU FUYANG DISTRICT FIRST PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610802626.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing multimodal omics prognostic risk prediction systems for cervical cancer lack cervical cancer-specific biological topological prior constraints, and the fusion process generates biologically meaningless noise. They cannot robustly address the lack of spatial heterogeneity information in the tumor microenvironment and the lack of modality in multimodal omics data in routine clinical Bulk omics data.

Method used

A multimodal omics feature fusion system is adopted, which uses HPV integration breakpoint coordinates as anchor points for heterogeneous graphs. The heterogeneous graphs are constructed by combining chromosome physical distance and cis-regulation priors. Spatial micro-domain latent distribution dictionary and Sinkhorn optimal transport algorithm are used to inject spatial micro-environment priors. Cross-modal gating network is combined to handle modality missing, so as to achieve cross-modal adaptive completion and risk prediction.

Benefits of technology

It improves the accuracy of risk stratification, significantly enhances the biological interpretability and robustness of the model, and can maintain high predictive performance even with a modality missing rate of up to 50%, breaking through the performance bottleneck of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637902A_ABST
    Figure CN122637902A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal omics feature fusion's cervical cancer risk prediction system, belong to biomedical engineering and artificial intelligence cross field.The system is composed of four core modules: multi-modal omics data preprocessing and coding module, ST-HGAT topological graph attention fusion module, based on optimal transmission's spatial microenvironment priori knowledge injection module and cross-modal meta-learning adaptive completion and risk prediction module.ST-HGAT module takes HPV integration breakpoint coordinate as the anchor point node of graph, takes DNA methylation site and gene expression node as other two kinds of nodes, constructs heterogeneous graph with physical distance attenuation edge and cis-regulation priori edge, realizes the multi-omics fusion of compliance with the specific biological causal logic of cervical cancer, and the application significantly improves the precision of cervical cancer prognosis risk stratification and clinical applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of biomedical engineering, smart healthcare and artificial intelligence, and specifically relates to a cervical cancer risk prediction system based on the fusion of multimodal omics features. Background Technology

[0002] Cervical cancer is one of the most common malignant tumors in women worldwide, and persistent infection with high-risk human papillomavirus (HPV) is a necessary condition for its development. After the HPV genome integrates into the host chromosome, it can disrupt chromosomal topology, topological association domains, and TADs in the region adjacent to the integration site. This leads to abnormal activation or silencing of cis-regulatory elements of host tumor suppressor genes or proto-oncogenes, such as MYC, TERT, and CCND1. Consequently, through epigenetic remodeling, abnormal DNA methylation, and dysregulation of gene expression, it drives the malignant progression of cervical cancer.

[0003] Currently, multi-omics data integration has become a major direction in cervical cancer prognostic prediction research, and existing representative protocols can be divided into the following three categories:

[0004] The first type is a multi-omics fully connected fusion network based on feature splicing. This approach independently reduces the dimensionality of each omics data and then splices them horizontally, inputting them into a multi-layer fully connected network to output a survival risk score. End-to-end training is performed using the partial likelihood function of the Cox proportional hazards model as the loss function. This approach treats each omics feature as an independent and identically distributed vector element, neglecting the specific cis-regulatory effect of HPV integration sites on adjacent host genome regions. That is, there is a strong physical distance-dependent correlation between the coordinates of the integration site and the methylation status and expression level of surrounding genes. This leads to the core cervical cancer driving event of HPV integration being diluted in high-dimensional noise, resulting in insufficient ability of the model to capture key pathogenic pathways, and the risk stratification consistency index is significantly lower than the optimal level.

[0005] The second category is multi-omics association modeling based on graph neural networks. This approach constructs a gene interaction graph using public databases, performs neighborhood aggregation through graph convolutional networks or graph attention networks to obtain node embeddings that consider network topology, and then outputs a patient-level risk score through full graph pooling. This approach uses general protein-protein interaction priors and does not model the specific biological topology of HPV integration in cervical cancer, thus failing to effectively capture the cervical cancer-specific causal chain of "viral invasion—epigmoid remodeling—abnormal gene expression."

[0006] The third category is Transformer-based multimodal self-attention fusion. This approach treats different omics data as different labeled sequences, utilizes the Transformer self-attention mechanism for cross-modal fusion, and takes the classification label output and passes it through a linear layer to obtain a risk score. However, this approach also lacks cervical cancer-specific biological topological prior constraints and is not adaptable to scenarios with missing modalities.

[0007] The three approaches mentioned above share three core limitations: First, none of them model the specific biological topology of HPV integration in cervical cancer, resulting in a large amount of biologically meaningless noise during the fusion process. Second, the spatial distribution pattern of immune cells and stromal cells in the tumor microenvironment is a key independent factor determining patient prognosis, but spatial transcriptomics technology, which can resolve this information, cannot be incorporated into routine clinical testing procedures due to its high cost per test and stringent requirements for fresh frozen samples; existing Bulk omics-based models completely ignore this. Third, multimodal omics data from patients in real clinical settings generally suffer from modality loss; existing models experience a sharp decline in performance or even complete failure when any omics dimension is missing, making robust deployment in real clinical scenarios impossible.

[0008] Therefore, there is an urgent need for a multimodal omics risk prediction system for cervical cancer that can simultaneously solve the above three core technical problems. Summary of the Invention

[0009] This invention aims to address three core technical challenges faced by existing cervical cancer multimodal omics prognostic risk prediction systems in practical clinical applications:

[0010] First, existing multi-omics fusion models lack cervical cancer-specific biological topological prior constraints, and the fusion process generates a large amount of biologically meaningless noise, resulting in low accuracy of risk stratification.

[0011] Second, routine clinical Bulk omics data completely loses information on the spatial heterogeneity of the tumor microenvironment, resulting in a systematic lack of key prognostic information;

[0012] Third, in real-world clinical settings, patient multimodal omics data often suffers from modality loss, which existing models cannot robustly address.

[0013] To address the aforementioned technical problems, this invention provides a cervical cancer risk prediction system based on multimodal omics feature fusion, which consists of the following four core modules:

[0014] (I) Module A: Multimodal omics data preprocessing and encoding module

[0015] Module A is responsible for standardizing and initial feature encoding of the input omics data. The types of input data processed include HPV integration breakpoint coordinates, DNA methylation profiles, RNA-seq transcriptome data, and clinical baseline data.

[0016] For HPV integration breakpoint coordinates, HPV integration sites are identified through whole-genome sequencing or targeted high-throughput sequencing. The chromosome number and base pair position coordinates of the integration breakpoints on the human reference genome are recorded. Each integration site is represented as a binary tuple, which is then mapped to a D-dimensional breakpoint coordinate vector after linear normalization.

[0017] For DNA methylation profiles, Illumina 450K or EPIC methylation chip data were used to extract the Beta values ​​of the corresponding CpG sites, with a value range of [0,1], representing the degree of methylation. Missing CpG sites were filled using the K-nearest neighbor interpolation method, and the top N_m CpG sites with the highest variance were retained to form a methylation feature matrix.

[0018] For RNA-seq transcriptome data, after TPM normalization, log1p transformation is performed to retain the set of differentially expressed genes associated with cervical cancer to form a transcriptome feature matrix.

[0019] For clinical baseline data, HPV genotyping is one-hot encoded, viral load is log normalized, and cytology results are ordinal encoded. These results are then mapped to a global condition vector through a linear embedding layer, which is used as the global condition input for the subsequent meta-learning module.

[0020] (II) Module B: ST-HGAT Topology Graph Attention Fusion Module

[0021] Module B, based on the topological constraints of HPV integration sites, performs heterogeneous graph attention fusion on the host DNA methylation profile and transcriptome, and outputs a patient-level topological fusion characterization.

[0022] The rules for constructing heterogeneous graphs are as follows:

[0023] In terms of node definition, the graph contains three types of nodes: HPV integration anchor nodes, with one node corresponding to each integration site, and the initial feature being the corresponding normalized breakpoint coordinate vector; methylation site nodes, with one node corresponding to each retained CpG site, and the initial feature being the corresponding Beta value; and gene expression nodes, with one node corresponding to each retained gene, and the initial feature being the corresponding log-TPM value.

[0024] Regarding edge definition and weight calculation, the graph contains three types of edges:

[0025] Physical distance decay edges are established from HPV integration anchors to neighboring CpG site nodes or gene expression nodes. A directed edge is created when the chromosomal physical distance between the HPV integration anchor and a CpG site or gene is less than a preset threshold. In this system, this is set to 1Mb, i.e., 1 million base pairs. The edge weights are calculated using a Gaussian distance decay function.

[0026]

[0027] in, For HPV integration anchor nodes With the target node Edge weights between them; The physical distance between the two on the chromosome, in bp; The distance attenuation bandwidth parameter is set to 200000, or 200kb, in this system, to control the rate at which the weights decrease as the distance increases.

[0028] Cis-regulatory prior edges are established from CpG site nodes to target gene nodes. Directed edges are established based on the regulatory relationships between promoter CpG sites and corresponding target genes annotated in the ENCODE and RoadmapEpigenomics databases. The initial weight is set to 1, and subsequent weights are adaptively adjusted by the graph attention mechanism.

[0029] Co-expression related edges are established between gene nodes. Based on the gene co-expression matrix pre-computed by the TCGA-CESC queue, undirected edges are established for gene pairs whose absolute Pearson correlation coefficient exceeds the threshold of 0.7, and the edge weight is the corresponding absolute value of the correlation coefficient.

[0030] The process of attention propagation in the diagram is as follows:

[0031] First, the initial features of the three types of nodes are mapped to a unified dimension through a linear transformation. The implicit vectors are denoted as follows: , , .

[0032] For each directed edge in the graph The attention coefficient is calculated as follows: : The latent vector of the source node and target node latent vector After concatenation, the learnable attention vector Inner product with Leaky ReLU activation, then combined with the corresponding prior structure weights. Multiply and apply to the target node. Softmax normalization is performed on all source neighbor nodes to obtain the final attention coefficient.

[0033] For each target node, the latent vectors of all its source neighbors are weighted and summed according to the attention coefficient. The latent vectors of the target node are then updated through the ELU activation function and residual connections, thus completing one layer of graph attention propagation.

[0034] This system employs a multi-head attention mechanism, and the number of heads... ,parallel computing Group of independent attention coefficients, The aggregated results were spliced ​​together and then compressed back by linear transformation. Dimension, repeated execution The attention propagation of the layer graph yields the final hidden vectors of each node.

[0035] Graph-level readout is performed on the final latent vectors of all nodes in the graph, employing a joint strategy of mean pooling and attention pooling: the mean pooling vector is obtained by calculating the mean of the latent vectors of all nodes, and the attention pooling vector is obtained by weighted summation using a learnable node importance scoring function. The two are concatenated and then linearly transformed to obtain a patient-level topological fusion representation. .

[0036] (III) Module C: Module for injecting prior knowledge of the spatial microenvironment based on optimal transmission

[0037] Module C is divided into an offline pre-training stage and an online inference stage, which realizes the injection of the spatial micro-domain latent distribution dictionary built offline into the main network features through Wasserstein distance regularization.

[0038] The offline spatial micro-domain distribution dictionary construction process is as follows:

[0039] We collected publicly available cervical cancer spatial transcriptome datasets, such as cervical cancer slice data from the 10x Visium platform in the GEO database. We performed Scran normalization on the gene expression data of each spot to remove low-quality spots and detect spots with fewer than 200 genes.

[0040] The gene expression vectors of each spot in the preprocessed spatial transcriptome data are input into a pre-trained variational autoencoder to extract the latent vector representation of each spot in the latent space. This system .

[0041] The latent vectors of each spot point After concatenating with its spatial coordinate information, all spot points are clustered into groups using a graph-based spatial clustering algorithm (Graph-SAGE combined with Leiden community detection algorithm). Each spatial functional microdomain (this system) Spatial clustering algorithm: Graph-SAGE combined with Leiden community discovery algorithm. Each microdomain corresponds to a typical spatial functional region in cervical cancer tissue, including T cell infiltration microdomain, CAF enrichment microdomain, tumor core microdomain, angiogenesis microdomain, etc.

[0042] For each micro-domain ( Collect the latent vectors of all spot points within the micro-domain and calculate the micro-domain prototype vector. (The mean of the latent vectors of all spot points within this microdomain) and the covariance matrix (Diagonal approximation), store it as a dictionary of hidden distributions in a spatial micro-domain. Freeze all parameters in the dictionary.

[0043] The optimal transmission alignment process for online Sinkhorn is as follows:

[0044] Topological fusion characterization of patients As a single-point sample of the source distribution, in the spatial micro-domain dictionary The mean vector of the prototype of each micro-domain As the target distribution Each support point is assigned a uniform weight to the target distribution. Source distribution weights .

[0045] Calculate the cost matrix , of which The elements are With the Micro-domain prototype The square Euclidean distance between them, i.e. .

[0046] Introducing entropy regularization coefficient This system The optimal transmission matrix is ​​obtained by solving the regularized optimal transmission problem using the Sinkhorn iterative algorithm. The specific execution method of Sinkhorn iteration is as follows: initialize dual variables. , Then, row normalization and column normalization operations are performed alternately until convergence (the maximum number of iterations is set to 50).

[0047] With optimal transmission matrix The element As weight, for We obtain the spatial prior alignment vector by weighted summation of the prototype vectors of each micro-domain:

[0048]

[0049] in, represent Transmitted to the The probability mass of a spatial micro-domain For the first The prototype mean vector of a spatial micro-domain.

[0050] Topology fusion representation Alignment vector with spatial prior By performing residual superposition, a spatial prior enhancement representation is obtained. Simultaneously calculate the Wasserstein regularization loss. That is, the total cost of optimal transmission, used as an auxiliary training loss to guide... Approximate a reasonable spatial micro-domain distribution in the feature space:

[0051]

[0052]

[0053] in, The spatial prior fusion intensity hyperparameter is set to 0.3 in this system; For the cost matrix, the first Each element, namely With the The squared Euclidean distance of each micro-domain prototype.

[0054] (iv) Module D: Cross-modal meta-learning adaptive completion and risk prediction module

[0055] Module D processes modality-deficient scenarios and outputs the final risk score, consisting of a cross-modal gating network and a risk prediction head.

[0056] The cross-modal adaptive completion process is as follows:

[0057] Spatial prior augmentation representation output by receiver module C Global conditional vector obtained by encoding clinical baseline data And read the existence mask vector of each omics mode. ,in This indicates the presence of DNA methylation data. This indicates that RNA-seq data exists. This indicates that the HPV integration breakpoint coordinate data exists, while a value of 0 indicates that the modality data is missing.

[0058] With global condition vector A query vector for multi-head cross-attention is generated through linear transformation; key vectors and value vectors are generated respectively using the encoded features of each existing modality through linear transformation; multi-head cross-attention calculation is performed to obtain a context-aware modality importance weight vector. , The Middle The element represents the first The importance score of each modality under the current patient condition.

[0059] For each missing mode, use the global condition vector and context weight vector The concatenated vector is used as input, and a two-layer MLP decoder is used to infer the high-dimensional latent representation corresponding to the missing mode. The parameters of the MLP decoder are trained using a meta-learning paradigm, enabling it to quickly adapt to different modality-missing scenarios with a small number of samples.

[0060] For all modes, including the encoded representation of existing modes. Inference representation of missing modes Based on the existence of a modal mask The fusion weights of each modality representation are adaptively adjusted: the weights of modes that exist are adjusted through... The importance score corresponding to the missing mode is directly determined; the weight of the missing mode is determined in... The corresponding importance score is multiplied by a penalty coefficient. This system This is to reflect the uncertainty of the inferred representation of missing modes relative to the actual observed representation.

[0061] The final fused representation is obtained by summing all modal representations according to their adjusted weights. and combine it with spatial prior enhancement representation The risk prediction head is input through residual connection fusion.

[0062] The risk prediction head consists of a three-layer fully connected network, with each layer followed by BatchNorm and Dropout. The Dropout rate is 0.3, and the final output is a scalar risk score. The larger the continuous value, the higher the risk. High-risk / low-risk stratification is performed based on the median cutoff point.

[0063] The system uses the negative partial likelihood function of the Cox proportional hazards model. The main training loss is combined with the optimal transport regularization loss. And mode reconstruction consistency loss The total loss for end-to-end joint training is:

[0064]

[0065] in, and The system provides the balancing hyperparameters for each auxiliary loss. , ; The mean squared error loss between the missing modality inferred representation and the true encoded representation on complete modality samples is used to supervise the learning of the cross-modal MLP decoder.

[0066] The system training consists of three stages: The first stage is offline spatial micro-domain dictionary pre-training, which involves constructing a dictionary based on open-source cervical cancer spatial transcriptome data and then freezing the parameters; the second stage is end-to-end joint training of the main network, which simulates clinical modality missing scenarios through random modality masking data augmentation strategy, optimizes it by combining three loss functions, uses the AdamW optimizer with an initial learning rate of 1e-4, and uses cosine annealing learning rate scheduling, with a maximum of 200 training epochs; the third stage is meta-learning fine-tuning, which involves executing MAML inner and outer loops on the modality missing task set to further enhance the cross-modal completion network's ability to quickly adapt to multiple missing modes. After the three stages of training are completed, the model parameters with the optimal C-index on the validation set are saved for clinical deployment.

[0067] The beneficial effects of this invention are:

[0068] 1. By using the HPV integration breakpoint coordinates as anchor nodes of the heterogeneous graph and constructing the edges of the heterogeneous graph using the chromosomal physical distance decay function and cis-regulatory biological priors, the multi-omics fusion process follows the cervical cancer-specific biological causal logic of "viral invasion - epigenetic remodeling - abnormal gene expression". This eliminates meaningless fusion noise, improves the model's biological interpretability and risk stratification accuracy. In the standard cervical cancer multi-omics cohort, the risk stratification consistency index can be improved by 5% to 10% compared with the existing best baseline method.

[0069] 2. By constructing a spatial micro-domain latent distribution dictionary offline and using the Sinkhorn optimal transfer algorithm for feature alignment during the online inference stage, the system can implicitly perceive the spatial heterogeneity information of the tumor microenvironment based solely on conventional Bulk omics data, breaking through the performance limit of Bulk omics models. Under the condition of only inputting Bulk transcriptome data, the prognostic stratification performance surpasses existing comparative methods that do not contain spatial information.

[0070] 3. By using a meta-learning-based cross-modal gated completion network, the system dynamically infers the potential representation of missing modalities using clinical baseline data as a global condition, and adaptively adjusts the fusion weights of each modality based on the modality presence mask. This ensures that the system's prediction performance does not decrease by more than 3% even when the random missing modality rate reaches 50%, which is significantly better than the precipitous performance drop of existing methods, enabling truly feasible clinical deployment. Attached Figure Description

[0071] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1This is a flowchart illustrating the overall architecture of the system of the present invention;

[0073] Figure 2 Complete workflow diagram for injecting prior knowledge of spatial microenvironment based on optimal transmission;

[0074] Figure 3 This is a flowchart illustrating the complete three-stage training process of the system of the present invention. Detailed Implementation

[0075] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0076] As attached Figure 1-3 As shown;

[0077] Figure 1 This is a flowchart of the overall architecture of the system of the present invention, describing the entire process from multimodal data input to risk score output, and showing the data transmission relationship from module A to module D and the core processing steps of each module.

[0078] Figure 2 A complete workflow diagram for the module injecting prior knowledge of spatial micro-environments based on optimal transmission is presented, showing the process of constructing the spatial micro-domain latent distribution dictionary in the offline pre-training stage and the process of aligning Sinkhorn optimal transmission features in the online inference stage.

[0079] Figure 3 This is a complete three-stage training flowchart of the system of the present invention, which shows the training steps, data input and parameter update methods of the three stages: offline dictionary pre-training, end-to-end joint training of the main network and meta-learning fine-tuning.

[0080] Example 1

[0081] This embodiment uses the TCGA-CESC cervical cancer multi-omics cohort as the training dataset to demonstrate the complete training and validation process of the system of the present invention.

[0082] During the data preparation phase, patient samples containing RNA-seq transcriptome data, DNA methylation 450K microarray data, HPV integration breakpoint coordinates obtained from whole-genome sequencing (WGS), and clinical follow-up data were acquired from the TCGA-CESC cohort. For spatial transcriptome data, publicly available cervical cancer 10x Visium slice datasets were obtained from GEO databases (such as GSE210616) for offline dictionary construction.

[0083] In the first training phase, using the aforementioned open-source spatial transcriptome data of cervical cancer, Algorithm 2 of module C was executed to construct a spatial microdomain latent distribution dictionary containing eight spatial functional microdomains (T cell infiltration microdomain, CAF enrichment microdomain, tumor core microdomain, angiogenesis microdomain, and four other types of functional microdomains). The dictionary parameters are frozen and will not participate in subsequent main network training.

[0084] In the second training phase, for each training batch of samples, one of the three omics modalities is randomly masked with a probability of 0.3 to simulate a clinical modality missing scenario. The samples are then processed sequentially through Module A: preprocessing and encoding; Module B: ST-HGAT topology graph attention fusion; Module C: Sinkhorn optimal transport feature alignment; and Module D: forward propagation: cross-modal adaptive completion and risk scoring, to calculate the total loss. The AdamW optimizer is used to update parameters via backpropagation. The initial learning rate is 1e-4, and the learning rate is scheduled using cosine annealing. Training continues until the maximum number of epochs is 200 or the C-index on the validation set no longer increases.

[0085] In the third training phase, meta-learning fine-tuning is performed. On a task set containing multiple modal missing modes, MAML inner and outer loops are executed to perform meta-learning fine-tuning on the cross-modal MLP decoder parameters in module D, enhancing its rapid adaptability to few samples and multiple missing modes. After the three-phase training is completed, the model parameters with the optimal C-index on the validation set are saved for clinical inference deployment.

[0086] Example 2

[0087] This embodiment demonstrates the complete predictive workflow for patients who simultaneously possess HPV integrated breakpoint coordinates, DNA methylation profiles, RNA-seq transcriptome data, and clinical baseline data.

[0088] Taking a patient's data as an example, whole-genome sequencing identified three HPV16 integration sites, with breakpoint coordinates located at chr8:128,748,315 (approximately 80kb upstream of the MYC gene), chr13:115,251,489, and chr3:186,741,023, respectively. After linear normalization, each of the three integration sites was mapped to a 128-dimensional breakpoint coordinate vector, forming... .

[0089] DNA methylation 450K microarray data, after KNN imputation, retains the top 5000 CpG sites with the highest variance to form a methylation feature matrix. RNA-seq data, after TPM normalization and log1p transformation, retained 2000 differentially expressed genes selected based on prior literature on cervical cancer, forming a transcriptome feature matrix. Clinical baseline data is encoded and mapped to a 64-dimensional global conditional vector. .

[0090] In Module B, a heterogeneous graph was constructed based on the above data: three HPV integration anchor nodes, where the chr8 nodes (128, 748, 315) are approximately 80 kb physically distant from the MYC gene region, meeting the 1 Mb threshold, with a Gaussian decay weight of approximately 0.817; 5000 methylation site nodes; and 2000 gene expression nodes, totaling 7003 nodes. Based on the physical distance decay threshold, approximately 380 physical distance decay edges were established within the 1 Mb range of the chr8 integration site; approximately 1200 cis-regulation prior edges were established based on ENCODE database annotations; and approximately 4300 co-expression related edges were established based on the co-expression matrix, with gene pairs having an absolute Pearson correlation coefficient exceeding 0.7. After attention propagation through a 3-layer, 8-head graph, a patient-level topological fusion representation was obtained using a combined mean pooling and attention pooling strategy. .

[0091] In module C, with For a single-point sample from the source distribution, calculate the cost matrix by plotting its squared Euclidean distance to the eight spatial micro-domain prototypes in the dictionary. The optimal transfer matrix is ​​solved through Sinkhorn iteration. Assuming The patient's topological characterization was shown to be primarily transmitted to the T-cell infiltration microdomain and the tumor core microdomain, and a weighted sum was obtained from these. Spatial prior enhancement representation is obtained by residual superposition. .

[0092] In module D, the modal masks exist. All modalities exist, and the importance weight of each modality is calculated through multi-head cross-attention. After weighted fusion, a risk score is output through a three-layer fully connected network. The value is higher than the median cutoff point of 1.87, indicating a high risk.

[0093] Example 3

[0094] This embodiment demonstrates a robust prediction process for the system when only RNA-seq transcriptome and clinical baseline data are available.

[0095] Modal Existence Mask Set to That is, only RNA-seq data ( ) exists, DNA methylation ( ) and HPV integration breakpoint coordinates ( All are missing.

[0096] In Module B, due to the lack of HPV integration breakpoint coordinate data, a complete heterogeneous graph cannot be directly constructed. Therefore, only RNA-seq transcriptome features are used to initialize gene expression nodes, skipping the construction of HPV integration anchor nodes and related physical distance decay edges. Graph attention propagation is then performed on the reduced graph structure to obtain a degraded topological fusion representation. .

[0097] In module D, the cross-modal gating network uses a global condition vector. Based on RNA-seq modality weights alone, the latent characterization of missing methylation modalities is inferred through a two-layer MLP decoder. Potential characterization of HPV integration modality The fusion weights for the two missing modes are multiplied by a penalty coefficient based on their corresponding importance scores. The weighted fusion is then processed by a risk prediction head to output a risk score. Experiments show that, in this bimodal missing scenario, the C-index of our system decreases by only about 2.1% compared to the full-modal input, which is significantly better than the precipitous performance drop of existing methods in the same missing scenario.

[0098] Example 4

[0099] This embodiment demonstrates a method for performing biological interpretability analysis by extracting graph attention weights from the ST-HGAT module.

[0100] After completing model training, for the validation set of patients, the attention coefficients of all physical distance decay edges originating from the HPV integration anchor node are extracted. And combined with prior structural weights The actual information contribution weight distribution of each HPV integration site to its surrounding methylation sites and gene expression nodes can be obtained. For the chr8:128,748,315 integration site, its attention weight to the MYC gene expression node is significantly higher than its weight to other unrelated gene nodes, averaging 3.7 times higher. This verifies the biological hypothesis that this integration event activates MYC through a cis-regulatory mechanism, demonstrating the biological interpretability advantage of the model in this invention.

[0101] Meanwhile, by analyzing the optimal transmission matrix in module C... The transport probability quality of each spatial microdomain can determine the dominant type of the patient's tumor microenvironment. In high-risk patients, the proportion of patients transported mainly to the "tumor core microdomain" and "CAF-enriched microdomain" is significantly higher than that in low-risk patients, which is consistent with the known pathological mechanisms by which immune escape and stromal remodeling promote the malignant progression of cervical cancer.

[0102] In the description of this specification, the references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0103] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in the claims, they should all fall within the protection scope of the present invention.

Claims

1. A cervical cancer risk prediction system based on multimodal omics feature fusion, characterized in that, The system includes: The multimodal omics data preprocessing and encoding module receives HPV integration breakpoint coordinates, DNA methylation profiles, RNA-seq transcriptome data, and clinical baseline data. These data are then standardized and initially encoded. Specifically, HPV integration breakpoint coordinates are mapped to breakpoint coordinate vectors after linear normalization. DNA methylation profiles extract CpG site Beta values ​​and construct a methylation feature matrix after K-nearest neighbor interpolation. RNA-seq data are normalized by TPM and transformed by log1p to construct a transcriptome feature matrix. Clinical baseline data are mapped to a global conditional vector after one-hot encoding, log normalization, and ordinal encoding. The ST-HGAT topology graph attention fusion module is used to construct a heterogeneous graph with HPV integration breakpoint coordinates as anchor nodes, DNA methylation sites and gene expression as the other two types of nodes, and physical distance decay directed edges with weights calculated based on the Gaussian decay function of chromosome physical distance, cis-regulation prior directed edges based on cis-regulation database annotations, and co-expression related undirected edges based on gene co-expression matrices. Multi-layer multi-head graph attention propagation is performed on this heterogeneous graph, and patient-level topology fusion representation is obtained by joint reading of graph-level mean pooling and attention pooling. The spatial microenvironment prior knowledge injection module based on optimal transmission is used to construct a spatial microdomain latent distribution dictionary using cervical cancer spatial transcriptome data in the offline stage. In the online inference stage, the optimal transmission matrix between the patient's topological fusion representation and each spatial microdomain prototype vector in the dictionary is calculated by the Sinkhorn iterative algorithm. The microdomain prototype vectors are weighted and summed with each element of the optimal transmission matrix as weights to obtain the spatial prior alignment vector. Then, the residual is superimposed with the topological fusion representation to obtain the spatial prior enhanced representation. The cross-modal meta-learning adaptive completion and risk prediction module receives spatial prior augmentation representations, global condition vectors, and existence mask vectors for each omics modality. It calculates the importance weights of each modality through a multi-head cross-attention mechanism with the global condition vector as the query vector. For missing modalities with an existence mask value of 0, it infers their latent representations through a two-layer MLP decoder. Based on the existence mask, it adaptively adjusts the fusion weights of each modality and then performs weighted fusion. Finally, it outputs risk scores and risk stratification results through a three-layer fully connected network risk prediction head.

2. The system according to claim 1, characterized in that, In the ST-HGAT topology graph attention fusion module, the edge weight of the directed edge with physical distance decay is calculated according to the Gaussian distance decay function. When the physical distance between the HPV integration anchor point and the target node on the chromosome is less than the preset distance threshold, the directed edge is established. The input of the Gaussian distance decay function is the physical distance between the HPV integration anchor point node and the target node on the chromosome, and the output is the edge weight. The distance decay bandwidth parameter controls the rate at which the weight decays as the distance increases. During the graph attention propagation process, the attention coefficient of each directed edge in the graph is obtained by concatenating the latent vectors of the source node and the target node, performing the inner product of the learnable attention vectors, and then taking LeakyReLU activation. The original score is then multiplied by the corresponding prior structure weights, and Softmax normalization is performed on all source neighbor nodes of the target node to obtain the final attention coefficient. For each target node, the latent vectors of all its source neighbors are weighted and summed according to the attention coefficient, and then the target node's latent vector is updated through the ELU activation function and residual connection. The multi-head attention mechanism calculates multiple independent attention coefficients in parallel, concatenates the aggregated results, compresses them to a unified dimension through linear transformation, repeats the multi-layer graph attention propagation, and then concatenates the mean pooling vector and the attention pooling vector through linear transformation to obtain the patient-level topological fusion representation.

3. The system according to claim 1, characterized in that, The offline phase of the spatial micro-domain latent distribution dictionary construction by the optimal transmission-based spatial micro-environment prior knowledge injection module specifically includes: We collected a spatial transcriptome dataset of cervical cancer, standardized the gene expression data of each spatial spot, and removed low-quality spot points. The preprocessed gene expression vectors of each spatial spot point are input into a variational autoencoder to extract the latent vector representation of each spot point in the latent space. The latent vector representation of each spot point is concatenated with its spatial coordinate information, and all spot points are clustered into K spatial functional micro-domains by a graph-based spatial clustering algorithm. Each micro-domain corresponds to a typical spatial functional region in cervical cancer tissue. For each spatial microdomain, calculate the mean of the latent vectors of all spot points in the microdomain as the prototype vector of the microdomain, calculate the covariance matrix using a diagonal approximation, store the prototype vectors and covariance matrices of each microdomain as a spatial microdomain latent distribution dictionary, and freeze all parameters in the dictionary.

4. The system according to claim 1, characterized in that, In the cross-modal meta-learning adaptive completion and risk prediction module, the parameters of the two-layer MLP decoder for the missing modality are trained on a task set containing multiple modality missing modes using the Model-Agnostic Meta-Learning paradigm, enabling it to quickly adapt to different modality missing scenarios under conditions of a small number of samples. For fusion weights with modalities, the corresponding importance scores obtained from multi-head cross-attention calculation are directly used to determine them; The fusion weights for missing modes are multiplied by the missing mode weight penalty coefficient based on the corresponding importance score to reflect the uncertainty of the inferred representation of missing modes relative to the actual observed representation. All modal representations are weighted and summed according to the adjusted weights to obtain the final fused representation, which is then fused with the spatial prior enhancement representation through residual connection and input into the risk prediction head.

5. The system according to claim 1, characterized in that, The training loss function of the system consists of a weighted sum of three parts: the negative biased likelihood loss of the Cox proportional hazards model, the optimal transmission regularization loss, and the modality reconstruction consistency loss. The optimal transmission regularization loss is the optimal total transmission cost, which is used to guide the patient topology fusion representation to approach a reasonable spatial micro-domain distribution in the feature space. The modality reconstruction consistency loss is the mean square error between the missing modality inferred representation and the true encoded representation on complete modality samples, which is used to supervise the learning of the cross-modal MLP decoder.

6. The system according to claim 1, characterized in that, The training of the system is divided into three stages: the first stage is the offline spatial micro-domain dictionary pre-training stage, which involves constructing a spatial micro-domain hidden distribution dictionary based on open-source cervical cancer spatial transcriptome data and then freezing the dictionary parameters. The second stage is the end-to-end joint training stage of the main network. A random modality masking data augmentation strategy is used to simulate clinical modality missing scenarios, and the parameters of the main network are optimized by combining three loss functions. The third stage is the meta-learning fine-tuning stage, which performs Model-Agnostic Meta-Learning inner and outer loops on a task set containing multiple modality missing modes to further enhance the rapid adaptability of the cross-modal completion network.

7. The system according to any one of claims 1 to 6, characterized in that, In the construction of the heterogeneous graph in the ST-HGAT topology graph attention fusion module, the cis-regulatory prior directed edges are established based on the regulatory relationship between promoter CpG sites and corresponding target genes annotated in the ENCODE database and the RoadmapEpigenomics database; the co-expression related undirected edges are established based on the gene co-expression matrix pre-computed in the cervical cancer cohort and gene pairs whose absolute Pearson correlation coefficient exceeds a preset threshold, with the edge weight being the absolute value of the corresponding correlation coefficient.

8. The system according to any one of claims 1 to 6, characterized in that, In the online inference phase, the spatial microenvironment prior knowledge injection module based on optimal transmission uses the patient's topological fusion representation as a single-point sample of the source distribution, and uses the prototype vectors of each micro-domain in the spatial micro-domain dictionary as support points of the target distribution and uniformly assigns weights to the target distribution. The cost matrix is ​​constructed by calculating the squared Euclidean distance between the topological fusion representation and each prototype vector of the micro-domain. An entropy regularization coefficient is introduced and the Sinkhorn iterative algorithm is used to solve the regularized optimal transmission problem to obtain the optimal transmission matrix. The spatial prior alignment vector is obtained by weighting each element of the optimal transmission matrix and summing the prototype vectors of each micro-domain. The spatial prior enhancement representation is obtained by residual superposition of the topological fusion representation and the spatial prior alignment vector.

9. The system according to any one of claims 1 to 6, characterized in that, In the multimodal omics data preprocessing and encoding module, the clinical baseline data includes HPV genotyping, viral load, and cytology results. HPV genotyping is one-hot encoded, viral load is log-normalized, and cytology results are ordinal encoded. The three are mapped to a global condition vector through a linear embedding layer. This global condition vector also serves as the global condition input of the cross-modal meta-learning adaptive completion and risk prediction module for the cross-modal gating network, used to infer the potential representation of missing modalities in modality-deficient scenarios.