Methods for subtyping acute respiratory distress syndrome biological subtypes
By integrating multi-omics data through the SNF-MCIA-DIABLO fusion framework, the problems of single and inaccurate ARDS classification systems have been solved, enabling efficient classification of biological subtypes and optimization of personalized treatment, thereby reducing mortality.
Patent Information
- Application Number
- CN202511040791.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-28
AI Technical Summary
Existing ARDS classification systems are too simplistic and cannot effectively utilize the correlation information of multi-omics data, resulting in insufficient classification accuracy and reliability. Furthermore, traditional methods are unable to discover new subtypes, which affects treatment efficacy and leads to high mortality rates.
A multi-omics similarity network was constructed using similarity fusion network (SNF). Combined with multiple synergistic principal component analysis (MCIA) and data integration analysis (DIABLO), multi-omics data dimensionality reduction and clustering were performed to identify key feature variables and achieve biological subtype classification.
It improves the accuracy of biological subtyping of ARDS patients, can identify new subtypes, optimize treatment strategies, reduce mortality, enhance the effectiveness of personalized treatment, and adapt to new subtypes and sample changes through model update mechanisms.
Smart Images

Figure CN121034430A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and in particular to a method for biological subtype typing of acute respiratory distress syndrome. BACKGROUND
[0002] Acute respiratory distress syndrome (ARDS) is a common severe respiratory failure in ICU, with persistent hypoxemia as the main clinical manifestation. ARDS has many causes and complex pathogenesis, and is highly heterogeneous, with a mortality rate of about 50%. The pathogenesis of ARDS is complex, and there is a lack of early diagnostic biomarkers. With the continuous in-depth exploration of the pathogenesis of the disease, more and more evidence suggests that the great heterogeneity of the disease is an important reason for the poor treatment effect and high mortality of patients.
[0003] However, the existing typing system does not meet the needs of clinical diagnosis and treatment, and the existing typing system still has the following problems: 1. Traditional single clinical typing: Traditional clinical typing only classifies the clinical phenotype of patients from a single dimension, and cannot predict the diagnosis and treatment effect of patients and optimize the diagnosis and treatment plan.
[0004] 2. Single omics data typing based on supervised learning algorithm: Based on supervised learning method, the known types are accurately classified, but new subtypes cannot be found, and single omics data analysis often cannot fully capture the association information between different data types, resulting in insufficient information utilization, affecting the accuracy and reliability of ARDS patient typing.
[0005] 3. Traditional unsupervised learning method: K-means clustering or hierarchical clustering method, although has the potential to discover new categories, but cannot preserve the non-linear structure of high-dimensional data, is sensitive to omics data noise, and cannot fully utilize the local similarity relationship between samples, resulting in high randomness, poor stability and poor repeatability between different data sets.
[0006] The method of the present application can predict the prognosis and complications of ARDS patients, optimize the treatment strategy of patients, and improve the prognosis of patients. SUMMARY
[0007] The present application covers the following technical solutions: One aspect of the present application relates to a method for biological subtype typing of acute respiratory distress syndrome, comprising: a) obtaining multi-omics data and performing standardization preprocessing; The multi-omics data includes transcriptomic data, proteomic data and metabolomic data from biological samples; b) constructing a similarity network of each omics by using a similarity fusion network (SNF), and obtaining a unified sample similarity matrix through multi-omics network fusion and iteration; Adopting multiple collaborative principal component analysis (MCIA) to reduce the dimension of multi-omics data to realize the visualization of clustering results; Using a data integration analysis (DIABLO) method to perform multi-omics joint discriminant modeling under the guidance of SNF clustering to identify key feature variables; Classifying biological subtypes based on the clustering results of the above steps.
[0008] Still another aspect of the present application relates to a computer readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when running on a computer, enable the computer to execute the method according to any one of the above aspects.
[0009] Still another aspect of the present application relates to an electronic device, comprising: one or more processors; and a computer readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when running on a computer, enable the one or more processors to implement the method according to any one of the above aspects.
[0010] Advantages of the present application: 1. The method of the present application can effectively solve the problem of multi-omics data heterogeneity, improve the accuracy of biological classification of patients with acute respiratory distress syndrome, and help improve the effect of personalized treatment.
[0011] 2. The fusion framework model of the present application also includes a model updating mechanism, which updates the model through incremental learning for new subtypes and samples, so that the model can complete self-optimization and iterative updating, has sustainable use, and meets all classification and new subtype detection. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0013] Figure 1 The first aspect of the present application provides a flowchart of the acute respiratory distress syndrome subtype classification method based on SNF-MCIA-DIABLO fusion framework and multi-omics data.
[0014] Figure 2Multi-omics fusion network heat map.
[0015] Figure 3 MCIA integrated scatter plot (subtype spatial distribution).
[0016] Figure 4 DIABLO integrated vector arrow plot (subtype spatial distribution).
[0017] Figure 5 Subtype verification AUC curve plot.
[0018] Figure 6 Subtype Cluster1 group-specific pathway enrichment network plot.
[0019] Figure 7 Subtype Cluster2 group-specific pathway enrichment network plot.
[0020] Figure 8 Subtype Cluster3 group-specific pathway enrichment network plot. Figure 9 Structure diagram of an acute respiratory distress syndrome biological subtype classification model adopted by an embodiment of the present application. DETAILED DESCRIPTION
[0021] Reference will now be made in detail to embodiments of the present application, one or more examples of which are illustrated below. Each example is provided by way of explanation of the present application, not limitation thereof. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made in the present application without departing from the scope or spirit of the present application. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment.
[0022] Unless otherwise defined, all terms (including technical and scientific terms) used in disclosing the present application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. By further guidance, the following definitions are set forth to better define the present teachings. The terminology used in the description of the present application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application.
[0023] In the present application, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by a person of ordinary skill in the art. Also, the terms related to protein and nucleic acid chemistry, molecular biology, microbiology, immunology and laboratory operation procedures used herein are the terms and conventional procedures widely used in the corresponding fields. At the same time, in order to better understand the present application, the definitions and explanations of the related terms are provided as follows.
[0024] The choice of the conjunctive range of the terms "and / or", "or / and", "and / or" used herein includes any one of the two or more related listed items, and also includes any and all combinations of the related listed items, including a combination of any two related listed items, a combination of any more related listed items, or a combination of all related listed items. It should be noted that when at least two conjunctions selected from "and / or", "or / and", "and / or" are combined to connect at least three items, it should be understood that in the present application, the technical solution undoubtedly includes the technical solution connected by "logical and", and also undoubtedly includes the technical solution connected by "logical or". For example, "A and / or B" includes three parallel solutions of A, B and A+B. For another example, the technical solution of "A, and / or, B, and / or, C, and / or, D" includes any one of A, B, C and D (i.e. the technical solution connected by "logical or"), and also includes any and all combinations of A, B, C and D, i.e. includes a combination of any two or any three of A, B, C and D, and also includes a four-item combination of A, B, C and D (i.e. the technical solution connected by "logical and").
[0025] The terms "containing", "including" and "comprising" used in the present application are synonymous and are inclusive or open-ended and do not exclude additional, unrecited members, elements or method steps.
[0026] The numerical ranges used in the present application include all the values and fractions included in the range and the recited endpoints.
[0027] The term "about" or "approximately" as used herein means within 20%, preferably within 10%, and more preferably within 5% of a given value or range. It is also intended to cover a specific number, e.g. about 20 includes 20.
[0028] Further, in describing representative embodiments of the present application, the specification can have presented the method and / or process of the present application as a particular sequence of steps. However, to the extent that the method or process depends on the performance of such steps, the method or process should not be limited to the
[0029] In the present application, the concentration values are intended to include fluctuations within a certain range. For example, they can fluctuate within a corresponding accuracy range. For example, 2% can fluctuate within a range of ±0.1%. For values that are relatively large or do not need to be controlled too precisely, the values are also intended to include larger fluctuations. For example, 100 mM can fluctuate within a range of ±1%, ±2%, ±5%, etc. For molecular weights, the values are intended to include fluctuations of ±10%.
[0030] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise.
[0031] In the present application, the descriptions such as "a plurality of", "a plurality of kinds", etc. refer to a number greater than or equal to 2 unless otherwise specified.
[0032] In the present application, the technical features described in an open form include both a closed technical solution consisting of listed features and an open technical solution including the listed features.
[0033] In the present application, "preferably", "more preferably", "even more preferably", "suitably" only describe embodiments or examples with better effects, and should be understood as not constituting a limitation on the protection scope of the present application. In the present application, "optionally", "optional" and "may" mean that something can or can not exist, i.e. it means that either of the two parallel schemes "yes" or "no" is selected. If multiple "optionally" appear in a technical solution, and there is no special description, and there is no contradictory relationship or mutual restriction, each "optionally" is independent.
[0034] The first aspect of the present application relates to a method for acute respiratory distress syndrome biological subtype typing, comprising: a) obtaining multi-omics data and performing standardization preprocessing; The multi-omics data includes transcriptomics data, proteomics data and metabolomics data from biological samples; b) constructing a similarity network for each omics using a similarity fusion network (SNF), and obtaining a unified sample similarity matrix through multi-omics network fusion and iteration; Using multiple collaborative principal component analysis (MCIA) to reduce the dimension of multi-omics data to realize the visualization of clustering results; Using the data integration analysis (DIABLO) method to perform multi-omics joint discriminant modeling under the guidance of SNF clustering to identify key feature variables; Based on the clustering results of the above steps, biological subtype classification is performed.
[0035] The application first integrates three multi-omics fusion algorithms SNF (Similarity Network Fusion) + MCIA (Multiple Co-Inertia Analysis) + DIABLO (Data Integration Analysis for Biomarker discovery using Latent components) in series, realizes the collaborative optimization of heterogeneous data structures by SNF clustering → MCIA dimensionality reduction visualization → DIABLO constructing cross-omics discriminant model, and applies the method to the biological subtype typing of ARDS patients, and verifies the accuracy and prediction ability of the model in clinical samples.
[0036] If DIABLO is used without SNF label, high AUC cannot be achieved; if MCIA is not aligned with the principal direction according to SNF clustering, the subtype boundary will be blurred. In the comparative experiment, if DIABLO is used to train the set of artificial preset labels instead of SNF clustering results, the AUC of the typing model on the test set decreases from 0.93 to 0.78, which shows that the label collaborative path has a significant impact on the stability of the results.
[0037] In addition, the application adopts multi-omics data to reveal complex pathophysiological processes and pathogenesis. The existing model is difficult to solve the problems of data heterogeneity and computational efficiency in the process of integrating multi-omics data. Data heterogeneity refers to the fact that different omics data have different data characteristics and noise levels due to differences in source, measurement method and scale, which makes data integration complex and challenging.
[0038] In some embodiments, the standardization preprocessing includes: The FPKM method is used for RNA-seq data normalization; Intra-sample standardization (Z-score or Min-max) is used for proteomics data; Log2 conversion and centralization are performed on metabolomics data.
[0039] The preprocessing process can improve the stability and consistency of subsequent fusion analysis. Of course, those skilled in the art can also select the publicly disclosed mature methods such as TPM normalization, Z-score standardization, Pareto scaling, PQN correction, ComBat batch effect correction, etc. according to the actual data type, to realize similar data consistency processing without changing the overall process of the application.
[0040] In specific embodiments, if the missing value of a protein in the proteomic data is more than 50%, the protein is removed; if the missing rate is between 10-50%, KNN imputation is used; in addition, for samples deviating from the main distribution by more than 2 times the standard deviation after batch correction, a weight degradation strategy is used to avoid affecting the training of the main model.
[0041] In some embodiments, the principal axis projection direction of the MCIA in the method is spatially aligned based on the SNF clustering result to enhance the consistency of clustering visualization.
[0042] This way guides the MCIA principal axis direction selection with the SNF unsupervised clustering structure, forms a logically interpretable low-dimensional visualization coordinate system, and enhances the clarity and intuitiveness of the subtype structure display.
[0043] In some embodiments, the DIABLO in the method performs cross-omics modeling with the clustering label obtained by SNF as a supervision signal.
[0044] The supervision label used by the DIABLO module is derived from the SNF fusion clustering result, not a manually set label. This step first realizes the reverse linkage of the process from "unsupervised clustering result → pseudo-supervised discriminant modeling", and in the process of displaying the principal components, the discriminant axis extracted by DIABLO is projected in the MCIA three-dimensional principal axis space, so that the low-dimensional expression results generated by different methods have consistent visual directions, enhancing the interpretability of the model.
[0045] In some embodiments, the latent components extracted by the DIABLO in the method are nested in the dimension consistent with the projection direction of the MCIA principal axis.
[0046] Further enhance the consistency of MCIA and DIABLO output space, so that the key features and typing structure are fused at the visualization level, enhancing the interpretability of multi-omics discriminant features.
[0047] In some embodiments, the standardization preprocessing includes: For RNA-seq data, FPKM method is used for normalization; For proteomic data, intra-sample standardization is used; For metabolomic data, log2 transformation and centering processing are performed.
[0048] In some embodiments, the method further uses a batch effect correction method for multi-omics data, and the method is ComBat correction.
[0049] In some embodiments, the biological sample in the method is a peripheral venous blood sample, a BALF sample, a bronchial brushing sample or a lung alveolar epithelial fluid, preferably the biological sample is a peripheral venous blood sample. Using peripheral blood sample can significantly improve the availability and popularization of the method in clinic, avoiding invasive operation (such as BALF acquisition), and having good practical availability and popularization potential.
[0050] In some embodiments, the biological sample in the method is derived from an ARDS patient receiving mechanical ventilation. Defining the suitable population as typical severe ARDS patients helps to enhance the applicability of the method in clinical monitoring and risk warning, and meets the needs of precise identification in clinical practice.
[0051] The present application also relates to a computer readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when executed on a computer, cause the computer to perform the method as described above.
[0052] Any combination of one or more computer readable medium can be employed. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, the computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0053] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for use by or in connection with an instruction execution system, apparatus, or device. The computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0054] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0055] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, Swift, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0056] The present application also relates to an electronic device comprising: one or more processors; and a computer readable storage medium storing computer instructions, programs, sets or instructions, which, when executed on a computer, cause the one or more processors to implement the method as described above.
[0057] In some embodiments, the electronic device can further comprise a transceiver. The processor and the transceiver are connected, such as through a bus. It should be noted that the transceiver is not limited to one in actual application, and the structure of the electronic device does not constitute a limitation on the embodiments of the present application.
[0058] The processor can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure. The processor can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of DSP and microprocessor, etc.
[0059] The bus can include a path for transmitting information between the above-mentioned components. The bus can be a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0060] The embodiments of the present application will be described in detail below with reference to the embodiments. It should be understood that these embodiments are only used to illustrate but not to limit the scope of the present application. The experimental methods in the following examples without specific conditions are preferred to refer to the guidance given in the present application, but also can be according to the experimental manual or conventional conditions in the art, or can refer to other experimental methods known in the art, or according to the conditions suggested by the manufacturer.
[0061] In the following specific examples, the measurement parameters of the raw material components may have slight deviations within the weighing accuracy range without specific instructions. The temperature and time parameters allow for acceptable deviations caused by instrument testing accuracy or operating accuracy.
[0062] Example 1 Reference Figure 1 As shown, the embodiment of the present application discloses a method for acute respiratory distress syndrome biological subtype typing based on SNF-MCIA-DIABLO fusion framework, which can be executed by a multi-omics acute respiratory distress syndrome data typing device based on SNF-MCIA-DIABLO fusion framework (hereinafter referred to as a typing device), in particular, by one or more processors of the typing device to implement the following method: S1, acquiring multi-omics data: According to the multi-omics data, a test data set is constructed, wherein the multi-omics data includes transcriptomics data, proteomics data and metabolomics data. 120 cases of acute respiratory failure patients are collected from 5 medical centers, and the transcriptomics data, proteomics data and metabolomics data of the patient's peripheral venous blood are obtained, wherein the transcriptomics data is an RNA-seq gene expression matrix, the data dimension is N × 36,184 genes, the proteomics data is a mass spectrometry quantitative protein matrix, the data dimension is N × 8,411 proteins, and the metabolomics data is an LC-MS metabolite matrix, the data dimension is N × 2587 metabolites.
[0063] S2, preprocessing the multi-omics data of the patient: The transcriptomics, proteomics and metabolomics data of the patient are preprocessed, wherein the data preprocessing includes data standardization and batch processing to ensure the consistency and stability of the data in the training process.
[0064] (1) Transcriptomic gene expression matrix preprocessing: input the reads count matrix generated by RNA-seq (calculated by featureCounts software) and transform it, steps include: using FPKM (Fragments Per Kilobase of transcript per Million fragments mapped) to normalize the original reads count matrix of genes, as an indicator to measure the expression level of transcripts or genes, the formula is , R ij : the original reads / fragments value of the jth gene in the ith sample; L j : the length of the jth gene (unit: kb); T i = : the total number of reads of sample i. Thus, the standardized matrix is obtained (N: sample number, G: gene number).
[0065] (2) Proteomic mass spectrometry quantitative matrix preprocessing: input the LFQ (Label-Free Quantification) protein quantitative matrix obtained by mass spectrometry (calculated by DIA-NN software) and transform it, steps include: using intra-sample standardization to process the abundance quantification of each protein in different samples, the formula is , where i represents the sample and j represents the protein. Thus, the standardized matrix is obtained (N: sample number, P: protein number).
[0066] (3) Metabolomic LC-MS matrix preprocessing: when OPLS-DA analysis, input the metabolite peak area matrix detected by LC-MS and transform it, steps include: log2 transformation of the original peak area to reduce the dominance of high abundance metabolites; then center scaling, the formula is , = metabolite mean, (N: sample number, M: metabolite number).
[0067] S3, construct a single omics similarity network: For each omics data : (1) Distance matrix calculation:
[0068] (2) K-nearest neighbor similarity matrix: set
[0069] Hot kernel weight calculation: , where is the sample Distance to its Kth neighbor, = 0.5 (3) Normalized transition matrix:
[0070] S3, Fusion and iteration of multi-omics network: (1) Initialization of fusion network:
[0071] (2) Iterative update (T = 15 times):
[0072] (3) Output fusion similarity matrix:
[0073] S4, Feature extraction and visualization of multi-omics fusion expression (1) Multi-omics component dimension reduction: Use Multiple Co-Inertia Analysis (MCIA) to perform joint dimension reduction on the original expression matrix, extract comprehensive variables representing the covariant structure of different omics, steps include: first input the data matrix: , where is the number of features of the mth omics, N is the number of samples, and each omics is independently centered and standardized, the formula is ; find the maximum co-structure in the sample dimension of the standardized omics matrix, extract the first three principal axes (var1, var2, var3), the formula is . Then is used as a three-dimensional projection to show the spatial distribution of SNF clusters in each sample.
[0074] (2) Key feature identification and classification modeling: Use Data Integration Analysis for Biomarker discovery using Latent components (DIABLO) to extract discriminant features and jointly model multi-omics under the guidance of SNF cluster labels, steps include: input the data matrix: , where is the cluster label defined by SNF; find the collaborative latent discriminant components between omics, maximize the inter-group discrimination, and select the most representative feature subset, finally the main discriminant components of each omics and the corresponding variable weight vector are selected, which are used for subsequent biological interpretation and subtype feature marker extraction.
[0075] S5, ARDS biological subtype classification and verification: (1) Density clustering classification. Specifically, in this embodiment, the multi-omics data of real acute respiratory distress syndrome patients is used, and the data comes from 5 different medical centers, a total of 120 samples. For all experiments, training and testing are performed in the same hardware environment. Three kinds of omics information of acute respiratory failure patients, including transcriptomic data, proteomic data and metabolomic data. Divided into 3 categories, the category and the corresponding sample number are respectively: Cluster1: 29, Cluster2: 44, Cluster3: 47. The stability of the model is verified, and the AUC of the three classification methods is above 0.88 (Cluster1: 0.927, Cluster2: 0.968, Cluster3: 0.889) (2) Correlation analysis of clinical indicators of three groups of patients: oxygenation index PaO2 / FiO2, SOFA score, APACHE II score, ICU mortality, 3-month follow-up mortality, 6-month follow-up mortality, 12-month follow-up mortality, and new-onset acute renal failure after ARDS diagnosis of three groups of patients were analyzed, wherein the categorical variable data uses chi-square test, the continuous variable data belongs to normal distribution data ANOVA test, and the skewed data uses Kruskal-waills test.
[0076] Table 1 Clinical correlation analysis of ARDS biological subtype classification
[0077] Compared with the traditional Berlin standard, which classifies mild, moderate and severe according to simple oxygenation level, the new classification is superior to the Berlin standard classification in predicting various indicators, and can better identify patients with poor prognosis and complications affecting patient prognosis, providing important reference basis for early intervention and optimization of treatment strategies.
[0078] In terms of ICU mortality and 3-month follow-up results, the mortality rate of Cluster1 (severe condition group) based on the classification method of the application (62.07%, 65.52%) is higher than that of the severe patient group of the Berlin standard (56.67%, 63.33%); the mortality rate of Cluster3 (treatment recovery group) based on the classification method of the application (21.28%, 25.53%) is lower than that of the mild patient group of the Berlin standard (29.41%, 35.29%). The classification method of the application can well identify and predict the prognosis and survival status of patients, and provide an important reference for the optimization of clinical diagnosis and treatment.
[0079] In terms of complication prediction, the prediction rate of new AKI of Cluster1 (severe group) based on the subtyping method of the application (41.38%) is higher than that of the severe patient group of the Berlin standard (26.67%); the prediction rate of new AKI of Cluster3 (treatment recovery group) based on the subtyping method of the application (6.38%) is lower than that of the mild patient group of the Berlin standard (23.53%). The subtyping method of the application can well predict the incidence of complications AKI, early warning and prediction of patients, and improve the prognosis of patients.
[0080] Table 2 Berlin standard grading method-clinical correlation analysis
[0081] S6, omics-specific differential analysis: (1) Transcriptional omics differential gene screening: DESeq2 algorithm based on negative binomial distribution model is used for differential analysis of non-standardized RNA-seq original count matrix, the key steps include: estimation of sample size factor (SizeFactor) standardization; fitting dispersion-mean relationship; Wald test to calculate difference significance, where the threshold is set to (Benjamini-Hochberg correction), get the differential gene list D RNA .
[0082] (2) Proteomics differential protein screening: LIMMA algorithm based on linear model and empirical Bayes shrinkage is used for differential analysis of log2 transformed LFQ intensity matrix, the key steps include: design matrix (including subtype grouping variables) construction; fitting linear model, formula is eBayes adjustment variance estimation, where the threshold is set to (equivalent to FC>1.5), get the differential protein list D Protein .
[0083] (3) Metabolomics differential metabolite screening: MASS algorithm is used for differential analysis of Pareto scaled metabolite matrix, the key steps include: calculate mean and variance according to subtype grouping; perform t test, ; FDR correction of multiple hypotheses, where the threshold is set to (VIP is variable importance projection), get the differential protein list D Metabolite .
[0084] S7, multi-level pathway enrichment analysis: (1) Single omics pathway enrichment analysis: The algorithm, database and test method used are shown in Table 3, and the significance threshold is set to FDR < 0.05 & Count ≥ 5 (number of different molecules in the pathway) Table 3 Single omics pathway enrichment analysis method
[0085] (2) Multi-omics common pathway integration analysis: Based on STRING (gene-protein) and STITCH (protein-metabolite) databases, the interaction of is extracted, and a molecular interaction network is constructed to obtain an integrated network ; The KEGG pathway topology structure is superimposed on G, and the pathways covered by ≥ 2 types of molecules are identified (such as pathway reporter gene , protein , metabolite ); Joint enrichment significance calculation, Fisher joint probability method, formula:
[0086] Among them, is the single enrichment p value of the molecule in the pathway, and the final output multi-omics synergistically regulated pathway ( ) In summary, compared with the existing traditional Berlin standard grading method, the acute respiratory distress syndrome patient biological subtype classification method based on the SNF-MCIA-DIABLO fusion framework is used for integrating and analyzing multi-omics data of acute respiratory distress syndrome patients, and is particularly suitable for biological subtype classification. The method captures complementary biological signals by using a similarity fusion network SNF model to fuse transcriptomic, proteomic and metabolomic data. The subtypes obtained by the classification method are significantly associated with the prognosis of patients in the prospective cohort of multiple medical centers. Experimental results show that the proposed method is superior to the traditional Berlin standard grading in terms of classification accuracy and stability, and can better capture the correlation information between different omics data. The method realizes efficient integration of multi-omics data and significantly improves the accuracy of acute respiratory distress syndrome patient classification. Through experimental verification, the method shows excellent performance and performs better than the traditional Berlin standard grading in the classification task.
[0087] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the patent scope of the present application. It should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims, and the description and drawings can be used to explain the content of the claims.
Claims
1. A method for acute respiratory distress syndrome (ARDS) biological subtype typing, comprising: a) obtaining and standardizing preprocessing of multi-omics data; the multi-omics data comprising transcriptomics data, proteomics data and metabolomics data derived from biological samples; b) constructing a similarity network for each omics using a similarity fusion network (SNF), and obtaining a unified sample similarity matrix through multi-omics network fusion and iteration; using multiple collaborative principal component analysis (MCIA) to reduce the dimensionality of multi-omics data to realize the visualization of clustering results; using a data integration analysis (DIABLO) method to perform multi-omics joint discriminant modeling under the guidance of SNF clustering to identify key feature variables; based on the clustering results of the above steps, biological subtype classification is performed.
2. The method of claim 1, wherein the standardization preprocessing comprises: using FPKM method to normalize RNA-seq data; using intra-sample standardization for proteomics data; log2 transformation and centering processing for metabolomics data.
3. The method of claim 2, wherein the principal axis projection direction of the MCIA is based on the spatial alignment of the SNF clustering results to enhance the consistency of clustering visualization.
4. The method of claim 2, wherein the DIABLO uses the clustering labels obtained by SNF as a supervisory signal for cross-omics modeling.
5. The method of claim 2, wherein the latent components extracted by the DIABLO are nested in the dimension consistent with the principal axis projection direction of the MCIA.
6. The method of any one of claims 1-5, wherein the standardization preprocessing comprises: using FPKM method to normalize RNA-seq data; using intra-sample standardization for proteomics data; log2 transformation and centering processing for metabolomics data.
7. The method of claim 6, wherein the multi-omics data is further corrected for batch effect using a ComBat correction method.
8. The method of any one of claims 1-5, 7, wherein the biological samples are peripheral venous blood samples, BALF samples, bronchial brushing samples, and alveolar epithelial fluid; optionally, wherein the biological samples are derived from ARDS patients receiving mechanical ventilation.
9. A computer-readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when executed on a computer, cause the computer to perform the method of any one of claims 1-8.
10. An electronic device, comprising: one or more processors; and a computer-readable storage medium for storing computer instructions, programs, code sets or instruction sets, which, when executed on a computer, cause the one or more processors to implement the method of any one of claims 1-8.
Citation Information
Cited By
Cancer prediction model construction method based on causal network and adaptive feature selection
CN122050850A