An exosome multi-dimensional health status evaluation method and system based on artificial intelligence (general platform and tumor screening application)

CN122531704APending Publication Date: 2026-08-07LIANWU JINNUO (SHANGHAI) BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIANWU JINNUO (SHANGHAI) BIOTECHNOLOGY CO LTD
Filing Date
2026-06-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]然而,现有技术在外泌体数据的临床转化应用中仍面临五大核心瓶颈:

Benefits of technology

[0046](1)建立了首个外泌体特异性的三维质量控制体系。 区别于通用多组学质控工具,本发明从标志蛋白检出一致性、RNA长度分布完整性和膜蛋白纯度三个维度构建了专门针对外泌体生物学特性的质控框架。该三维质控体系能够有效过滤低质量数据,确保下游分析的生物学可信度,预期可显著提升诊断模型在不同批次和不同实验室数据上的泛化性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based method and system for multidimensional health status assessment of exosomes, belonging to the interdisciplinary field of biomedical informatics and artificial intelligence. The method includes: acquiring multi-omics data of exosome RNA and proteins; performing exosome-specific three-dimensional quality control to assess the consistency of marker protein detection, the integrity of RNA length distribution (KL divergence < 0.5), and the purity of membrane proteins (EMII > 0.6); extracting 64-dimensional RNA and 64-dimensional protein modal features respectively through a cross-modal feature extraction network with quality control anchor constraints; performing adaptive weighted fusion based on exosome origin type, employing a dynamic modal weight decay strategy; inputting the 128-dimensional fused features into a graph attention network analysis engine (2-5 layers of graph attention convolution) based on an exosome domain knowledge graph for association reasoning; and verifying the interpretation through multi-method interpretability analysis, outputting the assessment conclusion and a marker contribution interpretation report when the intersection of the top-10 markers from the three methods is non-empty and they are in the top 50% quantile. This invention constructs an intelligent health status assessment scheme for exosomes that covers the entire chain of "data quality control - feature extraction - fusion analysis - knowledge reasoning - interpretable output". It solves the technical problems of lack of quality control, difficulty in fusion, black box diagnosis, insufficient performance in small sample scenarios, and lack of systematic multi-omics integration methods for aging assessment in the existing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of biomedical informatics and artificial intelligence, specifically relating to an AI-based method and system for multidimensional health status assessment of exosomes. This invention comprehensively utilizes exosome multi-omics data acquisition, exosome-specific quality control, cross-modal deep feature fusion, graph neural network (GNN) correlation reasoning, and explainable artificial intelligence (XAI) technology to construct an intelligent exosome health status assessment solution covering the entire chain of "data input—quality control—feature extraction—fusion analysis—knowledge reasoning—explainable assessment output." It is applicable to scenarios such as aging status assessment, monitoring of anti-aging intervention effects, early tumor screening, and auxiliary diagnosis and prognostic assessment of diseases. Background Technology

[0002] Exosomes are extracellular vesicles (EVs) with a diameter of approximately 30-150 nanometers, widely distributed in various bodily fluids such as blood, urine, and saliva. Exosomes carry a variety of bioactive molecules from their source cells, including proteins, RNA (ribonucleic acid, including mRNA and miRNA), and lipids, reflecting the physiological and pathological state of these cells. Therefore, they are considered a highly promising source of biomarkers for liquid biopsy. In recent years, with the rapid development of high-throughput sequencing and mass spectrometry technologies, the cost of acquiring exosomal RNA and proteomics data has significantly decreased, while the amount of data has grown exponentially, providing an unprecedented data foundation for exosome-based disease screening and diagnosis.

[0003] However, existing technologies still face five major bottlenecks in the clinical translation of exosome data:

[0004] First, there are limitations and fragmented information in single-omics analysis. Traditional exosome diagnostic methods are mostly based on a single omics dimension, or rely only on a few known miRNA (microRNA) markers, or only detect the expression levels of individual membrane proteins. While RNomics can capture regulatory information at the transcriptional level, it lacks functional execution information at the protein level; while proteomics can reflect the post-translational functional state, it is difficult to reveal upstream regulatory mechanisms. A single omics perspective cannot depict the complete biological picture of exosomes as intercellular communication carriers, resulting in limited diagnostic sensitivity and specificity, and making it difficult to reveal the systemic molecular mechanisms of disease development.

[0005] Second, there are structural difficulties in multimodal data fusion. Exosome RNA expression data and protein expression data differ fundamentally in data type, distribution characteristics, dimensionality, and biological semantics. RNA data is typically high-dimensional, sparse, count-based data, while protein data is mostly continuous abundance values; the data acquisition platforms, preprocessing procedures, and sources of batch effects differ between the two modalities. Existing general multimodal fusion methods (such as early splicing fusion or late decision fusion) fail to fully consider the specificities of the exosome domain—exosomes from different sources (such as plasma, cell culture supernatant, and urine) exhibit significant differences in RNA and protein expression patterns and correlation strengths. Simply applying general fusion strategies can easily lead to information conflicts or loss of complementary information between modalities.

[0006] Third, there is a lack of expertise in exosome data quality control. Existing quality control methods are mostly simple transplants of general multi-omics quality control tools, lacking specific quality control dimensions tailored to the biological characteristics of exosomes. The Minimum Information Requirement (MISEV) for exosome identification explicitly recommends members of the four-transmembrane protein family, such as CD63, CD81, and CD9, as exosome marker proteins; however, existing quality control procedures do not systematically assess the consistency of detection and the reliability of quantification for these marker proteins. Furthermore, exosome RNA has a specific length distribution characteristic (mature miRNAs are concentrated in the approximately 22 nt length range), a key quality control indicator often overlooked in existing RNA quality control procedures. Simultaneously, the purity of exosome membrane proteins directly affects the biological reliability of proteomics data, but quantitative purity assessment methods are lacking. Weaknesses in quality control significantly reduce the reproducibility and clinical translational value of downstream analyses.

[0007] Fourth, the "black box" problem of AI models and the clinical trust gap. Deep learning models have demonstrated powerful predictive capabilities in exosome diagnostic tasks, but their internal decision-making mechanisms are opaque, making it difficult for clinicians and review agencies to understand and trust the diagnostic basis of the models. When a model gives a positive diagnosis, it cannot answer key questions such as "which exosome biomarkers drove this judgment" and "which disease pathways are these biomarkers associated with?" Existing interpretation methods are mostly post-hoc general interpretation techniques (such as Grad-CAM, LIME, etc.), which are not organically integrated with the biomarker-pathway-disease knowledge system in the exosome field. The interpretation results lack domain consistency and clinical verifiability, which seriously restricts the clinical implementation of AI diagnostic systems.

[0008] Fifth, the knowledge transfer capability is insufficient in scenarios with few samples. The accumulation of exosome data for novel or rare diseases is limited, and the number of training samples is far from sufficient to support adequate training of deep models. Existing transfer learning methods mostly involve fine-tuning parameters in general domains, failing to utilize the accumulated biomarkers—disease-related knowledge—in the exosome field. When faced with diagnostic tasks for new diseases or indications, the model struggles to effectively utilize shared biological knowledge across diseases, resulting in poor cold-start performance and a large demand for labeled data.

[0009] Sixth, aging assessment lacks a systematic multi-omics integrated approach. Aging is a complex biological process involving multiple systems and levels of physiological changes. The Hallmarks of Aging framework proposed by the International Consortium for Research on Aging systematically summarizes several core dimensions of aging, including genomic instability, telomere loss, epigenetic alterations, loss of protein homeostasis, dystrophic nutrient sensing, mitochondrial dysfunction, cellular senescence, stem cell depletion, altered intercellular communication, and chronic inflammation. This framework provides scientific theoretical guidance for aging assessment, but it also places higher demands on the comprehensiveness of detection methods.

[0010] Exosomes, as important carriers of intercellular communication, carry molecular information that can reflect changes in multiple dimensions of aging biomarkers. For example, specific miRNAs and proteins in exosomes can reflect genomic instability and DNA damage repair capabilities; changes in the expression of exosome-related genes are closely related to the cellular aging process; mitochondrial components carried by exosomes can indicate the degree of mitochondrial dysfunction; and the role of exosomes in regulating inflammatory responses and tissue repair reflects the aging dimension of altered intercellular communication. Therefore, exosomes naturally possess the molecular basis to serve as a tool for systemic aging assessment.

[0011] However, existing aging assessment methods have limitations in terms of comprehensiveness and practicality. Epigenetic assessment methods, such as DNA methylation clocks (e.g., Horvath clocks), while providing relatively accurate predictions of biological age, only reflect the single dimension of aging—epigenetic alterations—and cannot cover other key dimensions such as mitochondrial function, protein homeostasis, and intercellular communication. Telomere length detection, another classic aging assessment tool, also only reflects one aspect of cellular replication aging, and its standardization is insufficient, limiting its clinical applicability. Furthermore, existing aging assessment methods generally use single-omics data types, lacking the ability to integrate and analyze multi-omics information, making it difficult to construct a comprehensive assessment system that can fully reflect the multidimensional state of aging.

[0012] Therefore, how to fully utilize the multi-omics information of exosomal RNA and proteins, combined with the known associations of aging biomarkers in the knowledge graph of the aging field, to construct a systematic assessment model covering multiple dimensions of aging is an important technical problem that urgently needs to be solved in the current field of aging research.

[0013] In summary, there is an urgent need in this field for a systematic solution that can cover the entire chain of exosome screening and diagnosis, and demonstrate specific innovations in the field of exosomes in key aspects such as data quality control, cross-modal fusion, knowledge-driven reasoning, and interpretable output, thereby overcoming the six major technical bottlenecks mentioned above. Summary of the Invention

[0014] I. The technical problem to be solved by the invention

[0015] To address the shortcomings of existing technologies, this invention provides a method and system for multidimensional health status assessment of exosomes based on artificial intelligence, aiming to solve the following technical problems:

[0016] (1) How to establish a specific quality control system covering both exosomal RNA and proteomics to ensure the biological reliability and reproducibility of input data;

[0017] (2) How to construct a cross-modal adaptive fusion mechanism that adapts to the heterogeneity of exosome origin types and fully explore the complementary diagnostic information between RNA and protein modalities;

[0018] (3) How to use the knowledge graph of the exosome domain to drive the graph neural network for association reasoning, so as to realize systematic analysis from the marker level to the pathway level and then to the disease or physiological function system level;

[0019] (4) How to construct a marker-pathway-system interpretation chain through multi-method interpretability verification and domain knowledge integration to achieve multi-dimensional health status assessment from molecular markers to physiological functional system level;

[0020] (5) How to establish a longitudinal tracking and evaluation mechanism, and evaluate the effect of anti-aging intervention measures and the trend of individual aging rate changes by comparing and analyzing the exosome multi-omics data of the same subject at different time points.

[0021] II. Technical Solution

[0022] To achieve the above objectives, the present invention adopts the following technical solution:

[0023] A multi-dimensional health status assessment method for exosomes based on artificial intelligence, characterized by the following steps:

[0024] S1: Multi-omics Data Acquisition. Acquire multi-omics data of exosomes from the test samples. This multi-omics data includes exosome RNA expression data and exosome protein expression data. The exosome RNA expression data includes expression levels of miRNA, mRNA, and long non-coding RNA (lncRNA); the exosome protein expression data includes quantitative abundance information of membrane proteins and intraluminal proteins. The test samples may be sourced from plasma, serum, cell culture supernatant, urine, saliva, dried blood spot (DBS), cerebrospinal fluid, or feces.

[0025] S2: Exosome-Specific Three-Dimensional Quality Control. Exosome-specific quality control is performed on the exosome multi-omics data. This quality control includes: assessing the consistency of exosome marker protein detection, calculating a quality control score based on the detection rate and quantitative correlation of three tetraspanic membrane protein families (CD63, CD81, and CD9); assessing the integrity of exosome RNA length distribution, evaluating RNA quality based on the enrichment of miRNA reads in the approximately 22 nt length range; and assessing the purity of exosome membrane proteins, calculating a purity index based on the abundance ratio of membrane-bound proteins to soluble proteins. Data that pass all three quality control measures are marked as "quality control qualified," otherwise marked as "quality control unqualified," triggering a data retest or rejection process.

[0026] S3: Cross-modal feature extraction. The quality-controlled exosome multi-omics data is input into a cross-modal feature extraction network to extract RNA and protein modal feature representations, respectively. The cross-modal feature extraction network includes an RNA modality encoder and a protein modality encoder. The RNA modality encoder employs a hierarchical dimensionality reduction structure adapted to the sparse and high-dimensional characteristics of RNA, while the protein modality encoder employs a fully connected structure adapted to the continuous abundance characteristics of proteins.

[0027] S4: Adaptive Weighted Fusion. The RNA modality feature representation and the protein modality feature representation are adaptively weighted and fused based on the exosome source type to obtain a fused feature representation. The adaptive weighted fusion dynamically adjusts the fusion weights of the RNA and protein modalities according to the source type (plasma, serum, cell culture supernatant, urine, saliva, dried blood spots, cerebrospinal fluid, or feces), enabling the fusion network to adapt to the differences in the strength of RNA-protein association under different source types.

[0028] S5: Graph Neural Network Association Reasoning. The fused feature representation is input into a graph neural network analysis engine, which performs message passing and association reasoning based on an exosome domain knowledge graph. The knowledge graph uses RNA, proteins, and metabolites as nodes, aging biomarkers (including aging-related miRNAs, aging-related secretory phenotype proteins, and mitochondrial function-related metabolites) and physiological functional systems (including the cardiovascular system, nervous system, metabolic system, immune system, skeletal system, skin system, endocrine system, digestive system, and respiratory system) as extension nodes, and known molecular interactions, regulatory relationships, disease associations, and aging-related pathway associations as edges. Through a graph attention mechanism, the fused features are propagated and enhanced on the knowledge graph, outputting a probability ranking of candidate disease or health status assessment results.

[0029] S6: Interpretable Diagnostic Output. The diagnostic results are interpreted and verified based on multi-method interpretability analysis, outputting evaluation conclusions and a biomarker contribution interpretation report. The multi-method interpretability analysis includes biomarker importance scoring based on attention weights, feature contribution calculation based on gradient backpropagation, and biomarker-pathway-system interpretation chain construction based on knowledge graph path reasoning. The results of these three methods must meet consistency verification conditions before the final evaluation report can be output.

[0030] Preferably, the assessment of the consistency of exosome marker protein detection in step S2 further includes: calculating the quantitative abundance Pearson correlation coefficient of any two of the three proteins CD63, CD81, and CD9, constructing a protein detection consistency triangle, and comprehensively assessing the overall consistency level of marker protein detection based on the area index and centroid position of the triangle.

[0031] Preferably, the assessment of exosomal RNA length distribution integrity in step S2 further includes: statistically analyzing the distribution curves of miRNA reads in the 18-25 nt length range, calculating the Kullback-Leibler divergence between the measured distribution curves and the standard exosomal miRNA length distribution template, and determining that the RNA length distribution integrity is qualified when the Kullback-Leibler divergence is less than 0.5.

[0032] Preferably, the evaluation of exosome membrane protein purity in step S2 further includes: calculating the EMII (Exosomal Membrane Integrity Index), whereby the EMII is defined as the ratio of the total abundance of membrane-bound proteins to the total abundance of membrane-bound proteins plus soluble proteins. When the EMII value is higher than 0.6, the membrane protein purity is deemed acceptable.

[0033] Preferably, the cross-modal feature extraction network described in step S3 introduces exosome quality control anchor point constraints during the training phase: the detection status of CD63, CD81, and CD9 is used as auxiliary supervision signals to constrain the intermediate layer features of the RNA modality encoder and the protein modality encoder to maintain consistent representation in the quality control anchor point dimension.

[0034] Preferably, the adaptive weighted fusion in step S4 adopts a dynamic modality weight decay strategy: during the training process, the fusion weights are dynamically adjusted according to the rate of decrease of the loss function of each modality, so that the modality with a slower loss decreases can obtain a higher fusion weight, thereby promoting the balanced utilization of information between modalities.

[0035] Preferably, the method for constructing the exosome domain knowledge graph in step S5 includes: automatically extracting RNA-protein interaction pairs, protein-pathway annotation pairs, and pathway-disease association pairs from exosome databases and biomedical literature, constructing nodes and edges of the knowledge graph using the extracted triples, and embedding knowledge through a graph convolutional network.

[0036] Preferably, step S5 further includes a multi-task dynamic routing mechanism: when diagnosing multiple diseases or performing disease classification at the same time, the graph neural network analysis engine dynamically adjusts the message passing path according to the characteristic activation mode of each disease, so that different diagnostic tasks activate different subgraph regions in the knowledge graph.

[0037] Preferably, step S5 further includes a multi-system aging assessment mechanism: the output of the graph neural network analysis engine is modularly grouped according to physiological functional systems, and aging scores for nine major systems (cardiovascular, nervous, metabolic, immune, skeletal, skin, endocrine, digestive, and respiratory) are calculated to generate a multi-system aging profile. The score for each system is calculated using a system-specific scoring function based on the inference results of the biomarker-pathway sub-map related to that system.

[0038] Preferably, step S5 further includes a few-sample knowledge transfer mechanism: when the number of training samples for the target disease is lower than a preset threshold, the source disease knowledge subgraph that shares pathways or markers with the target disease is retrieved from the knowledge graph, and cross-disease knowledge sharing is achieved through parameter transfer of the graph neural network.

[0039] Preferably, the consistency verification condition in step S6 is defined as follows: the intersection of the Top-K important marker sets obtained by the three interpretability methods is not empty, and the elements of the intersection are all in the top 50 percentile in the importance ranking of the three methods.

[0040] Preferably, the method for constructing the biomarker-pathway-system explanatory chain in step S6 includes: starting from the important biomarkers selected in the evaluation results, locating the relevant protein nodes along the RNA-protein interaction edge of the knowledge graph, locating the enriched signaling pathway nodes along the protein-pathway annotation edge, and locating the target disease or system node along the pathway-disease or physiological function system association edge, thus forming a multi-hop explanatory path from biomarker to pathway and then to disease or system.

[0041] Preferably, the output format of the explanation report in step S6 includes: evaluation conclusion, confidence score, Top-N key marker list, contribution score of each marker, annotation of the signaling pathways involved, and consistency verification results of the three methods. The report is presented in two forms: structured data and visual charts.

[0042] Preferably, step S6 further includes longitudinal intervention effect assessment: when there is historical exosome multi-omics data of the same subject at multiple time points, the exosome characteristic representation and health status assessment results at each time point are compared, the rate of change of aging score is calculated, and the effect of anti-aging intervention measures is evaluated.

[0043] Preferably, step S6 further includes peer percentile benchmarking: comparing the individual's health status assessment results with the reference distribution of the peer population, calculating the percentile position of the individual's score in the peer population, and providing a relative health level assessment.

[0044] III. Beneficial Effects

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] (1) The first exosome-specific three-dimensional quality control system was established. Unlike general multi-omics quality control tools, this invention constructs a quality control framework specifically for the biological characteristics of exosomes from three dimensions: consistency of marker protein detection, integrity of RNA length distribution, and purity of membrane proteins. This three-dimensional quality control system can effectively filter low-quality data, ensure the biological reliability of downstream analysis, and is expected to significantly improve the generalization performance of diagnostic models on data from different batches and different laboratories.

[0047] (2) An adaptive cross-modal fusion mechanism based on source type was achieved. This invention addresses the heterogeneity of exosomes from different sources (plasma, serum, cell culture supernatant, urine, saliva, dried blood spots, cerebrospinal fluid, and feces) by designing an adaptive weighted fusion strategy and a dynamic modal weight decay mechanism. This mechanism can dynamically adjust the fusion weights of RNA and protein modalities according to the source type, avoiding the inter-modal information conflict problem in general fusion methods. It is expected to fully explore the complementary diagnostic information of dual-omics data, improving the sensitivity and specificity of the assessment.

[0048] (3) A knowledge-driven graph neural network-based reasoning engine was constructed. This invention deeply integrates the exosome domain knowledge graph with the graph neural network, guiding the message passing process through molecular interactions and disease association information in the knowledge graph. Unlike purely data-driven black-box models, the reasoning process of this invention embeds domain knowledge constraints, which is expected to enhance the model's diagnostic capabilities in low-sample scenarios and provide a biologically interpretable reasoning path.

[0049] (4) An interpretable evaluation framework with multi-method consistency verification was established. This invention integrates three interpretation methods: attention weight analysis, gradient backpropagation, and knowledge graph path reasoning, and ensures the reliability of the interpretation results through a consistency verification mechanism. The output biomarker-pathway-system interpretation chain organically connects AI evaluation decisions with the biological knowledge system of exosomes, which is expected to bridge the trust gap between AI models and clinicians and promote the clinical translation of AI evaluation systems.

[0050] (5) A systematic solution covering the entire chain has been formed. This invention integrates data quality control, feature extraction, cross-modal fusion, knowledge reasoning, and interpretable output into a unified technical framework, with clear input-output interfaces and progressive relationships between each link. This full-chain design ensures standardization and automation of the entire process from raw multi-omics data to clinical evaluation reports, and is expected to lower the clinical translation threshold of exosome diagnostic technology and accelerate the popularization and application of liquid biopsy technology.

[0051] (6) Expanding the application scenarios of exosome multi-omics analysis to the field of aging assessment. This invention extends exosome multi-omics analysis from traditional disease diagnosis to the fields of health status assessment and aging monitoring. By constructing a knowledge graph linking aging biomarkers and physiological functions, it achieves multi-dimensional aging assessment from the molecular to the systemic level. This expansion enables exosome liquid biopsy technology not only for early disease detection but also for monitoring the aging status of healthy individuals and evaluating the effectiveness of anti-aging interventions, significantly broadening the applicable population and commercial application scenarios of the technology. Attached Figure Description

[0052] Figure 1 The present invention provides an overall flowchart of an artificial intelligence-based multidimensional health status assessment method for exosomes, illustrating the entire processing flow from S1 multi-omics data acquisition to S6 interpretability assessment output, as well as the data transfer relationships between each step.

[0053] Figure 2This is a schematic diagram of the exosome-specific three-dimensional quality control system provided in an embodiment of the present invention, illustrating the evaluation principles and joint judgment logic of three quality control dimensions: the consistency triangle diagram for marker protein detection, the RNA length distribution integrity assessment curve, and the EMII index for membrane protein purity.

[0054] Figure 3 This is a schematic diagram of the cross-modal adaptive fusion network provided in an embodiment of the present invention, showing the connection relationship between the RNA modality encoder, the protein modality encoder, the exosome origin type embedding layer, the adaptive weighted fusion layer, and the quality control anchor constraint signal.

[0055] Figure 4 This is a schematic diagram of the architecture for constructing an exosome domain knowledge graph and using a graph neural network for inference, provided in an embodiment of the present invention. It illustrates the node types (RNA, protein, pathway, disease, aging biomarkers, physiological functional systems), edge types (interaction, regulation, association, aging-related pathway association), and graph attention message passing mechanism of the knowledge graph.

[0056] Figure 5 The flowchart of multi-method interpretability analysis and consistency verification provided in the embodiments of the present invention illustrates the parallel execution logic of three methods: attention weight analysis, gradient backpropagation interpretation, and knowledge graph path reasoning, as well as the consistency verification and interpretation chain construction process.

[0057] Figure 6 The overall architecture diagram of the AI-based multidimensional health status assessment system for exosomes provided in this embodiment of the invention shows the modular structure and data flow of the data acquisition module, quality control module, feature extraction module, fusion analysis module, knowledge reasoning module, and assessment output module. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0059] Example 1: Full-process implementation of the multidimensional health status assessment method for exosomes

[0060] In an exemplary implementation scenario, the complete implementation process of the method described in this invention in the context of early tumor screening is demonstrated.

[0061] S1: Acquisition of multi-omics data.

[0062] In one exemplary implementation scenario, it is assumed that an exosome sample derived from plasma needs to be analyzed. In other exemplary scenarios, the sample may also be derived from saliva (collected via a non-invasive saliva collection tube), dried blood spots (collected via fingertip blood onto a dedicated filter paper card), venous blood (collected via standard venipuncture), cerebrospinal fluid (collected via lumbar puncture, suitable for neurological assessment scenarios), or feces (collected via a fecal sampling tube, suitable for gut microbiota-related assessment scenarios). Exosomes are isolated and purified using ultracentrifugation combined with size exclusion chromatography. Small RNA-seq technology is used to sequence the RNA of the extracted exosomes, obtaining miRNA expression profile data, in the format of RPM (Reads Per Million) normalized expression values ​​for each miRNA. Data-independent acquisition (DIA) mass spectrometry is used to quantitatively analyze the exosome proteome, obtaining a protein abundance matrix. Based on the technical characteristics of small RNA-seq and DIA mass spectrometry, detectable exosome multi-omics data typically includes expression values ​​of hundreds to thousands of miRNAs and quantitative abundance values ​​of hundreds to thousands of proteins, the specific number depending on sequencing depth and mass spectrometry resolution.

[0063] S2: Exosome-specific quality control.

[0064] (2a) Consistency assessment of marker protein detection: The detection of three exosomal marker proteins, CD63, CD81, and CD9, in mass spectrometry data was detected. The quantitative abundance Pearson correlation coefficients between each pair of the three proteins were calculated, resulting in three pairs of correlation coefficients: CD63-CD81, CD81-CD9, and CD63-CD9. A quality control triangular diagram with the three proteins as vertices was constructed, and the area index and centroid position index of the triangular diagram were calculated. Thresholds for the area index and the centroid offset were set. When both indices met the threshold requirements, the consistency of marker protein detection was deemed acceptable.

[0065] (2b) RNA length distribution integrity assessment: Statistical analysis of the length distribution of raw small RNA-seq data was performed, and read distribution curves for the 18-30 nt interval were plotted. A template curve for the length distribution of known standard exosomal miRNAs was extracted, and the KL divergence value between the measured distribution and the template distribution was calculated. A KL divergence threshold of 0.5 was set; RNA length distribution integrity was considered acceptable when the measured KL divergence was below this threshold.

[0066] (2c) Membrane protein purity assessment: In the proteomic data, proteins are classified into membrane-bound proteins and soluble proteins according to the cellular component annotation of Gene Ontology (GO). The EMII index is calculated as: EMII = Total abundance of membrane-bound proteins / (Total abundance of membrane-bound proteins + Total abundance of soluble proteins). An EMII threshold of 0.6 is set; membrane protein purity is considered acceptable when the EMII value is higher than this threshold.

[0067] Sample data will only proceed to downstream analysis if all three quality control indicators meet the requirements. If any one of them fails, the system will automatically mark the sample as "quality control unqualified," triggering a sample resampling or experimental retesting process.

[0068] S3: Cross-modal feature extraction.

[0069] Quality-controlled RNA expression data is input into an RNA modality encoder. This encoder consists of three fully connected layers. The first layer reduces the input dimension from hundreds to thousands of miRNAs (the exact number depends on sequencing depth) to 512. The second layer reduces it from 512 to 128, and the third layer outputs a 64-dimensional RNA modality feature representation from 128. ReLU activation and batch normalization are applied between layers, along with Dropout regularization (inactivation probability of 0.3) to prevent overfitting. To address the high-dimensional sparsity of the RNA data, a logarithmic transformation and sparse coding preprocessing step is added before the first layer.

[0070] Protein abundance data is input into a protein modality encoder. The protein modality encoder consists of two fully connected layers. The first layer reduces the input dimension from hundreds to thousands of proteins (the exact number depends on the mass spectrometry resolution) to 256. The second layer outputs a 64-dimensional protein modality feature representation from 256. To accommodate the continuous abundance characteristics of the protein data, Z-score normalization is performed before the input layer.

[0071] During the training phase, the detection states of CD63, CD81, and CD9 (represented as binary vectors) are used as quality control anchor constraint signals and applied to the intermediate layer outputs of the RNA modality encoder and the protein modality encoder, respectively. By minimizing the representational differences between the two modalities in the quality control anchor dimension, the consistency and biological interpretability of cross-modal features are enhanced.

[0072] S4: Adaptive weighted fusion.

[0073] The RNA and protein modality feature representations are input into an adaptive weighted fusion layer. First, the source type embedding table is queried based on the sample's source type (assumed to be plasma in this exemplary implementation scenario) to obtain the corresponding type embedding vector. The type embedding vector is then concatenated with the RNA and protein modality features respectively, and input into the fusion weight generation network to output the RNA modality weight w_RNA and the protein modality weight w_prot, satisfying w_RNA + w_prot = 1.

[0074] The fusion feature F_fusion is obtained by weighted summation: F_fusion = w_RNA × F_RNA + w_prot × F_prot, where F_RNA and F_prot are the RNA modality feature and the protein modality feature, respectively.

[0075] During model training, a dynamic modality weight decay strategy is employed: the rate of decrease in validation loss for both the RNA and protein modalities is recorded in each training cycle. If the rate of decrease in RNA modality loss is lower than that of the protein modality, the initial bias value of w_RNA is increased in the next cycle, and vice versa. This strategy ensures that information from both modalities is utilized in a balanced manner, avoiding a single modality dominating the fusion result.

[0076] S5: Graph Neural Network Association Reasoning.

[0077] The fused feature representations are input into a graph neural network analysis engine. The analysis engine performs inference based on a pre-built exosome domain knowledge graph.

[0078] The knowledge graph is constructed as follows: nodes are exosome-related miRNAs, proteins, signaling pathways, and disease types; edges are miRNA-protein interaction pairs, protein-pathway annotation pairs, and pathway-disease association pairs reported in the literature, all collected in the exosome database. The initial node embeddings are obtained by pre-training a graph convolutional network on the graph structure to capture the topological relationships between nodes.

[0079] The graph neural network analysis engine employs a Graph Attention Network (GAT) architecture. Fused feature representations serve as the initial input for miRNA and protein nodes, and message passing occurs across the knowledge graph through multiple layers of graph attention convolutions. In each layer of graph attention convolution, each node aggregates information from its neighbors and weights this neighbor information according to attention coefficients. After three layers of graph attention convolutions, each disease node in the knowledge graph receives an updated embedding representation, and the probability ranking of each candidate disease is output through fully connected layers and a softmax activation function.

[0080] When multi-task diagnosis is required (such as simultaneous screening for multiple tumor types), a multi-task dynamic routing mechanism is activated. The routing network corresponding to each diagnostic task generates a task-specific attention mask based on the activation mode of the fusion features, dynamically adjusting the effective weights of each edge during message passing, so that different tasks focus on different subgraph regions in the knowledge graph.

[0081] S6: Interpretable diagnostic output.

[0082] Once the graph neural network outputs a ranking of candidate diseases, a multi-method interpretability analysis process is initiated.

[0083] (6a) Attention weight analysis: Extract the attention coefficients of each miRNA and protein node to the target disease node in the last layer of the graph attention convolution of the graph neural network, and use them as biomarker importance scores.

[0084] (6b) Gradient backpropagation interpretation: Calculate the absolute value of the gradient of each dimension in the fusion feature representation relative to the probability of the target disease, backmap it to the original miRNA and protein input space, and obtain the feature contribution score of each marker.

[0085] (6c) Knowledge Graph Path Reasoning: Starting from the target disease node, backtrack along the edges of the knowledge graph to search for multi-hop paths from disease to pathway to marker. The weight of the path takes into account both the confidence of the association strength on the edge and the graph attention coefficient.

[0086] Consistency verification: Set K=10, take the top-10 important marker sets of the sorting methods for each of the three methods, and calculate the intersection of the three sets. The consistency verification is considered successful when the intersection is not empty and each element in the intersection is in the top 50 percentile in the sorting methods.

[0087] After the consistency verification is passed, a biomarker-pathway-system interpretation chain is constructed. Starting with the key biomarkers in the intersection, a multi-hop interpretation path from biomarker to pathway and then to disease or system is constructed along the RNA-protein interaction edge, protein-pathway annotation edge and pathway-to-disease or physiological function system association edge of the knowledge graph.

[0088] The final evaluation report includes: target disease assessment conclusions, model confidence score, Top-10 key biomarkers and their contribution scores, names of involved signaling pathways and enrichment significance evaluation, consistency verification results of the three methods, and a visualization of the explanatory chain. The report is output simultaneously in both structured JSON data format and PDF visualization charts.

[0089] Example 2: Adaptive Diagnosis of Exosome Samples from Multiple Sources

[0090] In an exemplary implementation scenario, the adaptive fusion capability of the method of the present invention is demonstrated when processing exosome samples from multiple sources.

[0091] In one exemplary implementation scenario, it is assumed that four types of exosome samples from different sources need to be processed simultaneously: plasma samples, serum samples, cell culture supernatant samples, and urine samples. The number of each type of sample is determined according to actual research needs, for example, ranging from dozens to hundreds of samples per type. All samples undergo multi-omics data acquisition and quality control according to the S1-S2 process described in Example 1.

[0092] In the S4 adaptive weighted fusion step, a source type embedding matrix is ​​constructed for different source types. The dimension of the embedding matrix is ​​4 (number of source types) × 16 (embedding dimension), and the embedding vectors are jointly optimized with other network parameters during model training.

[0093] Comparative experimental setup: A fixed fusion weight strategy (w_RNA = w_prot = 0.5) was used as the baseline method and compared with the adaptive weighted fusion strategy of this invention. The adaptive fusion strategy is expected to achieve superior diagnostic results compared to the fixed weight strategy on samples from various sources, especially on cell culture supernatants and urine samples where the correlation strength between RNA and protein differs significantly. The targeted adjustment capability of the adaptive strategy is expected to bring more significant performance advantages. The specific performance improvement needs to be verified through actual experimental data.

[0094] The specific implementation of dynamic modality weight decay is as follows: A recording window of 5 training epochs is set, and the average validation loss descent slope for each modality within this window is calculated. If the descent slope of the RNA modality is less than 70% of that of the protein modality, the initial weight bias of the RNA modality in the next epoch is increased by 0.1; conversely, if the descent slope of the protein modality is less than 70% of that of the RNA modality, the initial weight bias of the protein modality is increased by 0.1. The upper limit of the weight bias is 0.3 to prevent excessive shift.

[0095] Example 3: Knowledge Transfer Diagnosis in Few-Sample Scenarios

[0096] In an exemplary implementation scenario, the knowledge transfer capability of the method described in this invention is demonstrated in a scenario with few samples of a novel disease.

[0097] Suppose the target disease D is a rare type of tumor, with only 20 labeled samples (far fewer than the number of samples required to fully train a deep learning model). Enable a few-shot knowledge transfer mechanism.

[0098] (1) Retrieve knowledge subgraphs related to disease D from the exosome domain knowledge graph: Using known associated markers and pathways of disease D as seed nodes, perform a breadth-first search in the knowledge graph to expand neighboring nodes within a 2-hop range to form a local knowledge subgraph of disease D.

[0099] (2) Identify source diseases: Search the knowledge graph for other diseases that share at least 2 pathways or 3 biomarkers with disease D, and use them as the source disease set. Suppose that source diseases D1 (common tumor type with 500 labeled samples) and D2 (related chronic disease with 300 labeled samples) are identified.

[0100] (3) Parameter transfer: The graph neural network analysis engine was trained using samples from source diseases D1 and D2 to obtain the pre-trained model parameters. The parameters of the knowledge graph embedding layer and the first two graph attention convolutional layers were frozen, and only the third graph attention convolutional layer and the output classification layer were fine-tuned. During fine-tuning, 20 labeled samples of disease D were used.

[0101] (4) Knowledge-enhanced reasoning: In the reasoning stage, the sample features of disease D and the D-D1 and D-D2 association path information retrieved from the knowledge graph are input into the model. The expression of shared markers and pathways is enhanced through the attention mechanism, thereby improving the diagnostic reliability in scenarios with few samples.

[0102] It is expected that this few-shot knowledge transfer mechanism can achieve better diagnostic performance than the de novo trained model under the condition of only a small number of labeled samples, and significantly improve the predictive stability on shared markers and pathways.

[0103] Example 4: Boundary Case Handling for Multi-Method Consistency Verification

[0104] In an exemplary implementation scenario, the handling strategy for multi-method consistency verification in step S6 under boundary conditions is demonstrated.

[0105] The standard consistency verification condition is set as follows: the intersection of the three methods of the Top-10 marker set is not empty and the intersection element is located in the top 50 percentile of each method ranking.

[0106] Case 1 (Complete Consistency): The Top-10 marker sets obtained by the three methods are completely identical, the consistency verification passes directly, and the explanatory chain is constructed using this set to output an evaluation report.

[0107] Scenario 2 (Partial Consistency): The top-10 sets of the three methods have pairwise non-empty intersections, but the intersection of all three sets is empty. In this case, the consistency verification standard is lowered to: the intersection of the top-10 sets of at least two methods is non-empty, and the intersection elements meet the top 50% quantile requirement. An explanatory chain is constructed using the intersection of the two verified methods, and the report notes the discrepancy in the results of the third method and records a list of discrepancy markers.

[0108] Scenario 3 (Inconsistency): The top-10 sets of the three methods have pairwise empty intersections. In this case, the consistency verification is deemed to have failed, triggering the model retraining process: increasing the number of training epochs, adjusting the learning rate, or expanding the training samples. A preliminary evaluation conclusion is output, but the report prominently displays the warning message "Consistency verification failed; results are for reference only."

[0109] Through the above-described hierarchical processing strategy, this invention provides a degradation output mechanism in boundary cases while ensuring the reliability of the evaluation report, thereby enhancing the robustness and practicality of the system.

[0110] Example 5: Multisystem Aging Status Assessment Based on Exosome Multi-omics

[0111] In an exemplary implementation scenario, the application of the method described in this invention in the fields of health status assessment and aging monitoring is demonstrated.

[0112] (1) Sample collection and data acquisition: Assume that a subject's saliva and venous blood samples are collected. Saliva samples are collected using standard saliva collection tubes, and venous blood samples are collected into EDTA anticoagulant tubes using standard venipuncture. Both types of samples are purified by exosome separation and then subjected to small RNA-seq and DIA mass spectrometry analysis to obtain miRNA expression profiles and proteomic data.

[0113] (2) Three-dimensional quality control: The multi-omics data of the two groups of samples were subjected to quality control according to the S2 procedure. After the consistency assessment of marker protein detection, the integrity assessment of RNA length distribution and the purity assessment of membrane protein were all qualified, the data were entered into downstream analysis.

[0114] (3) Cross-modal fusion and knowledge reasoning: After obtaining the fused feature representation through steps S3-S4, it is input into the graph neural network analysis engine. The knowledge graph pre-constructs aging biomarker nodes (including reported aging-related miRNA and protein biomarkers) and physiological function system nodes (nine major systems: cardiovascular, nervous, metabolic, immune, skeletal, skin, endocrine, digestive, and respiratory). The graph neural network propagates the fused features on the aging knowledge graph through a message passing mechanism, and calculates the aging scores of the nine major systems respectively.

[0115] (4) Output of multi-system aging profile: Step S6 outputs a multi-system aging profile report, including independent scores for each system, inter-system correlation analysis, and percentile comparison results with the reference distribution of the same age group. The report also explains the key driving biomarkers for each system score and their biological function annotations through a biomarker-pathway-system interpretation chain.

[0116] (5) Longitudinal tracking example: Suppose the subject is sampled again after 6 months. The system will compare the two test results longitudinally, calculate the trend of aging scores for each system, and assess the rate of aging. If the subject has undergone specific anti-aging interventions (such as dietary adjustments, exercise interventions, or supplement regimens) between the two samplings, the system can further assess the impact of the interventions on the aging status of each system.

Claims

1. A multi-dimensional health status assessment method for exosomes based on artificial intelligence, characterized in that, Includes the following steps: S1: Obtain multi-omics data of exosomes from the sample to be tested. The multi-omics data includes exosome RNA expression data and exosome protein expression data. The exosome RNA expression data includes miRNA expression level information, and the exosome protein expression data includes quantitative abundance information of membrane proteins and luminal proteins. S2: Perform exosome-specific three-dimensional quality control on the exosome multi-omics data, including: evaluating the detection consistency of marker proteins based on the detection rates and quantitative correlations of three members of the tetraspanin family, CD63, CD81, and CD9; evaluating the integrity of RNA length distribution based on the KL divergence between the distribution curve of miRNA reads in the 18-25 nt length interval and the standard exosome miRNA length distribution template; and calculating the exosome membrane integrity index EMII based on the abundance ratio of membrane-bound proteins to soluble proteins to evaluate the purity of membrane proteins. When the KL divergence is less than 0.5 and the EMII is greater than 0.6, the quality control is determined to be qualified. S3: Input the exosome multi-omics data that has passed the quality control into a cross-modal feature extraction network, which includes an RNA modality encoder and a protein modality encoder. The RNA modality encoder adopts a three-layer fully connected hierarchical dimensionality reduction structure (input layer → 512 dimensions → 128 dimensions → 64-dimensional output), and the protein modality encoder adopts a two-layer fully connected structure (input layer → 256 dimensions → 64-dimensional output) to extract 64-dimensional RNA modality feature representations and 64-dimensional protein modality feature representations respectively. During the training phase, introduce quality control anchor point constraints: use the detection status of CD63, CD81, and CD9 as auxiliary supervision signals and apply them to the middle layers of the two encoders to constrain the consistency of the two modalities in the middle layer representations. S4: Perform adaptive weighted fusion on the RNA modality feature representations and the protein modality feature representations based on the exosome source type to obtain 128-dimensional fusion feature representations. The adaptive weighted fusion queries a preset source type embedding table according to the source type to obtain a type embedding vector, concatenates the type embedding vector with the RNA modality features and protein modality features, and inputs them into a fusion weight generation network to dynamically output RNA modality weights and protein modality weights that satisfy the normalization constraints. Adopt a dynamic modality weight decay strategy: During the training process, record the validation loss decrease rate of each modality within a 5-training cycle window, and increase the fusion weight bias of the modality whose loss decrease rate is lower than 70% of the other modality by 0.1 in the next training cycle (upper limit 0.3) S5: The fused feature representation is input into a graph neural network analysis engine. This engine performs message passing and association reasoning based on an exosome domain knowledge graph, outputting a probability ranking of candidate diseases or health states. The graph neural network adopts a graph attention network architecture, containing 2-5 graph attention convolutional layers. The knowledge graph uses RNA, proteins, signaling pathways, and diseases as nodes, and molecular interactions, regulatory relationships, and disease associations as edges, performing knowledge embedding learning through a graph convolutional network. S6: The evaluation results are interpreted and verified based on multi-method interpretability analysis, outputting evaluation conclusions and a biomarker contribution interpretation report. The multi-method interpretability analysis includes biomarker importance scoring based on attention weights, feature contribution calculation based on gradient backpropagation, and biomarker-pathway-system interpretation chain construction based on knowledge graph path backtracking. The results of the three methods must meet consistency verification conditions before a final evaluation report can be output. The consistency verification condition is defined as a non-empty intersection of the Top-10 important biomarker sets obtained by each of the three methods, and the intersection elements are all in the top 50% percentile in the importance ranking of the three methods.

2. The method according to claim 1, characterized in that, The assessment of the consistency of exosome marker protein detection in step S2 further includes: calculating the quantitative abundance Pearson correlation coefficient of any two of the three proteins CD63, CD81, and CD9; constructing a protein detection consistency triangle diagram with the three proteins as vertices; and comprehensively assessing the overall consistency level of marker protein detection based on the area index and centroid position index of the triangle diagram; when the area index is higher than a preset area threshold and the centroid offset distance is lower than a preset offset threshold, the consistency of marker protein detection is deemed qualified.

3. The method according to claim 1, characterized in that, The method for constructing the exosome domain knowledge graph in step S5 includes: automatically extracting RNA-protein interaction pairs, protein-pathway annotation pairs, and pathway-disease association pairs from exosome databases and biomedical literature; constructing nodes and edges of the knowledge graph using the extracted triples; and performing knowledge embedding learning through a graph convolutional network to obtain the topological association embedding vector of each node.

4. The method according to claim 1, characterized in that, Step S5 further includes a multi-task dynamic routing mechanism: when diagnosing multiple diseases or performing disease classification at the same time, the graph neural network analysis engine generates task-specific attention masks based on the characteristic activation patterns of each disease, dynamically adjusts the effective weights of message passing paths in the knowledge graph, and enables different evaluation tasks to activate different knowledge subgraph regions.

5. The method according to claim 1, characterized in that, Step S5 further includes a few-shot knowledge transfer mechanism: when the number of training samples for the target disease is less than 50, the source disease knowledge subgraph that shares at least two pathways or three markers with the target disease is retrieved from the knowledge graph, the parameters of the knowledge graph embedding layer and the first two graph attention convolutional layers of the graph neural network are frozen, and only the third graph attention convolutional layer and the output classification layer are fine-tuned to achieve cross-disease knowledge sharing.

6. The method according to claim 1, characterized in that, Step S5 further includes a multi-system aging assessment mechanism: the output of the graph neural network analysis engine is modularly grouped according to physiological functional systems, and aging scores of nine major systems (cardiovascular, nervous, metabolic, immune, skeletal, skin, endocrine, digestive, and respiratory) are calculated to generate a multi-system aging profile; Each system score is calculated using a system-specific scoring function based on the inference results of the biomarker-pathway sub-map related to that system.

7. The method according to claim 1, characterized in that, The method for constructing the biomarker-pathway-system explanatory chain in step S6 includes: starting from important biomarker nodes, locating relevant protein nodes along the RNA-protein interaction edge of the knowledge graph, locating signaling pathway nodes along the protein-pathway annotation edge, and locating target disease or system nodes along the pathway-disease or physiological function system association edge, forming a multi-hop explanatory path from biomarker to pathway and then to disease or system; the final output evaluation report includes evaluation conclusions, confidence scores, a list of key biomarkers, contribution scores of each biomarker, annotations of the involved signaling pathways, and consistency verification results of the three methods.

8. The method according to claim 1, characterized in that, Step S6 further includes longitudinal intervention effect assessment: when there is historical exosome multi-omics data of the same subject at multiple time points, the exosome characteristics and health status assessment results at each time point are compared, the rate of change of aging score is calculated, and the effect of anti-aging intervention measures is evaluated.

9. The method according to claim 1, characterized in that, Step S6 further includes peer percentile benchmarking: comparing the individual's health status assessment results with the reference distribution of the peer population, calculating the percentile position of the individual's score in the peer population, and providing a relative health level assessment.

10. A multi-dimensional health status assessment system for exosomes based on artificial intelligence, characterized in that, include: The data acquisition module is used to acquire multi-omics data of exosomes from the sample to be tested, including exosome RNA expression data and exosome protein expression data. A quality control module, connected to the data acquisition module, is used to perform exosome-specific three-dimensional quality control on the exosome multi-omics data. This three-dimensional quality control includes assessing the consistency of marker protein detection based on the detection rate and quantitative correlation of CD63, CD81, and CD9; assessing the integrity of RNA length distribution based on the distribution of miRNA reads in the 18-25 nt length range and the KL divergence of the standard template; and assessing membrane protein purity by calculating the EMII based on the abundance ratio of membrane-bound proteins and soluble proteins. The KL divergence threshold is set to 0.5, and the EMII threshold is set to 0.

6. A feature extraction module, also connected to the quality control module, is used to extract 64-dimensional RNA modality feature representations and 64-dimensional protein modality feature representations from the quality-controlled exosome multi-omics data. This feature extraction module includes a three-layer fully connected RNA modality encoder and a two-layer fully connected protein modality encoder, and incorporates CD63 and CD81 during the training phase.

81. CD9 detection status serves as a quality control anchor constraint; the fusion analysis module, connected to the feature extraction module, is used to adaptively weight and fuse the RNA modality feature representation and the protein modality feature representation based on the exosome origin type to obtain a 128-dimensional fusion feature representation; the knowledge reasoning module, connected to the fusion analysis module, has a built-in graph attention network analysis engine, which contains 2-5 graph attention convolutional layers, and performs message passing and association reasoning on the fusion feature representation based on the exosome domain knowledge graph, outputting a probability ranking of candidate diseases or health states; The diagnostic output module, connected to the knowledge reasoning module, is used to interpret and verify the probability ranking of the candidate diseases or health states based on multi-method interpretability analysis, and output the evaluation conclusion and biomarker contribution interpretation report.

11. The method or system according to claim 1 or 10, characterized in that, The method is implemented by a computer program stored in a computer-readable storage medium. When the computer program is executed by a processor, it implements all the steps of the artificial intelligence-based multidimensional health status assessment method for exosomes.