Factor analysis system based on long-chain non-coding RNA multi-omics integration analysis

By employing a factor analysis system based on multi-omics integration analysis of long non-coding RNAs, the problem of integrating and interpreting multimodal LncRNA data was solved, enabling accurate prediction of LncRNA function and efficient identification of molecular markers, thereby improving the biological interpretability and accuracy of data analysis.

CN120913641AInactive Publication Date: 2025-11-07BOCE BIOMEDICAL (TIANJIN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511037505.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate multimodal lncRNA data, making it difficult to analyze their complex regulatory networks. They lack targeted analysis of lncRNA lineage specificity, classification-related regulatory mechanisms, and tumor pathological phenotypic variations. Traditional factor analysis methods have insufficient biological interpretability, and missing value estimation affects data integrity and analytical accuracy.

Method used

A factor analysis system based on multi-omics integration analysis of long non-coding RNA was adopted, including modules for multi-omics data acquisition, data preprocessing, unsupervised factor analysis, heterogeneity analysis, and factor annotation. Through multimodal data integration, heterogeneity analysis, and functional annotation, potential factors were identified and mapping relationships were constructed, missing value estimation was optimized, and the efficiency of LncRNA functional prediction and molecular marker identification was improved.

Benefits of technology

This method enables precise functional annotation of LncRNA multi-omics data, improves the efficiency of LncRNA functional prediction and molecular marker identification, overcomes the technical bottlenecks of traditional methods, and enhances data integrity and analytical accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913641A_ABST
    Figure CN120913641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics and molecular biology, and discloses a factor analysis system based on long-chain non-coding RNA multi-omics integration analysis. The system comprises: a multi-omics data acquisition module configured to acquire multi-modal omics data related to long-chain non-coding RNA; the data pre-processing module is configured to pre-process the multi-modal omics data to generate a data set in a unified format; the unsupervised factor analysis module is configured to integrate and analyze the data set in the unified format and identify potential factors; the heterogeneity analysis module is configured to construct a mapping relation between the potential factors and multiple omics data features; and the factor annotation and expansion analysis module is configured to perform function annotation on the analyzed potential factors based on biological function enrichment analysis, regulation and control network inference and cross-modal data association, perform missing value estimation on the multi-omics data in combination with the mapping relationship, and output an analysis result containing factor annotation information and complete data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics and molecular biology, and particularly relates to a factor analysis system based on long non-coding RNA multi-omics integrated analysis. BACKGROUND

[0002] As an important regulatory molecule, the lineage specificity, spatiotemporal expression characteristics and key role in tumor development of long non-coding RNA (LncRNA) have been widely proven. LncRNA is involved in gene regulation through various mechanisms, including acting as a molecular scaffold, regulating alternative splicing, and absorbing microRNAs (miRNAs), and its abnormal expression is closely related to tumor malignant transformation, cell proliferation and invasion and other phenotypes. However, the functional analysis of LncRNA faces multiple challenges: the classification (antisense type, enhancer type, intergenic type, etc.) of LncRNA has significant differences in regulatory mechanisms, and the heterogeneity of multi-omics data (such as RNA-protein interaction data, expression profile data, and clinical phenotype data) makes it difficult for traditional single-omics analysis to reveal the complex regulatory network.

[0003] Existing multi-omics integration methods usually adopt simple data merging or statistical correlation-based analysis strategies, which cannot effectively separate common variations and unique variations in different modal data, especially lacking targeted analysis of core heterogeneity factors such as LncRNA lineage specificity, classification-related regulatory mechanisms and tumor pathological phenotype variations. For example, traditional factor analysis methods do not combine the biological characteristics (such as subcellular localization and functional classification) of LncRNA for model optimization, resulting in insufficient biological interpretability of the identified latent factors; the missing value estimation process does not associate the LncRNA regulatory network information, affecting data integrity and analysis accuracy. In addition, existing technologies lack a systematic functional annotation system for latent factors, making it difficult to establish a hierarchical association of "factor-regulatory molecule-function pathway", limiting the efficiency of LncRNA function prediction and molecular marker identification.

[0004] Therefore, in view of the complexity and specificity of LncRNA multi-omics data, there is an urgent need for an analysis system that can integrate multi-modal data, analyze heterogeneous regulatory factors and achieve accurate functional annotation, to break through the technical bottleneck of traditional methods in LncRNA function analysis. SUMMARY

[0005] The present application provides a factor analysis system based on long non-coding RNA multi-omics integrated analysis, which aims to address the complexity and specificity of LncRNA multi-omics data, and to develop an analysis system that can integrate multi-modal data, analyze heterogeneous regulatory factors and achieve accurate functional annotation, to break through the technical bottleneck of traditional methods in LncRNA function analysis.

[0006] In a first aspect, the application provides a factor analysis system based on long non-coding RNA multi-omics integrated analysis, comprising:

[0007] A multi-omics data acquisition module configured to acquire multi-modal omics data related to long non-coding RNA, the multi-modal omics data at least including RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data, and associated phenotype or disease state data;

[0008] A data preprocessing module configured to standardize, preliminarily process missing values, and select features of the multi-modal omics data, and generate a preprocessed unified format data set;

[0009] An unsupervised factor analysis module configured to use an unsupervised multi-omics factor analysis algorithm to integrate and analyze the unified format data set, identify potential factors affecting the main variation sources of the multi-omics data modalities, and the potential factors include hidden variables reflecting the lineage specificity, spatiotemporal expression characteristics, and disease correlation of long non-coding RNA;

[0010] A heterogeneity analysis module configured to analyze the heterogeneity factors behind different omics data from the perspective of factor analysis, and construct a mapping relationship between the potential factors and the multi-omics data characteristics, and the heterogeneity factors include differences in regulatory mechanisms related to long non-coding RNA classification and pathological phenotype variations related to tumors;

[0011] A factor annotation and extension analysis module configured to perform functional annotation on the identified potential factors based on biological function enrichment analysis, regulatory network inference, and cross-modality data association, and estimate missing values of the multi-omics data based on the mapping relationship, and output analysis results containing factor annotation information and complete data, so as to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

[0012] In some embodiments, the acquisition of multi-modal omics data related to long non-coding RNA includes: acquiring sequence feature data, subcellular localization data, and spatiotemporal expression data of multiple classified long non-coding RNAs; when RNA-protein interaction data is obtained based on RNA immunoprecipitation technology, the interaction information of chromatin modification proteins, RNA binding proteins, and microRNAs combined with long non-coding RNA is synchronously captured to obtain the multi-modal omics data; the associated phenotype or disease state data includes pathological grading, cell proliferation index, invasion ability parameter, and clinical prognosis data of tumor samples, and the data covers paired cancer tissue and paracancer tissue samples of at least two or more tumor types; the classification at least includes antisense type, enhancer type, intergenic type, bidirectional type, and intron type.

[0013] In some embodiments, the standardization, preliminary treatment of missing values and feature screening of the multi-modal omics data to generate a pre-processed uniform format dataset comprises: for RNA-protein interaction data, using quantile normalization to eliminate batch effects, and for long non-coding RNA expression profile data, using variance stabilization transformation to adapt to non-normal distribution characteristics; using K-nearest neighbor interpolation based on a latent low-dimensional factor space to preliminarily fill in missing values, wherein the interpolation process combines the lineage-specific expression pattern weight of long non-coding RNA; through a coefficient of variation to filter features with a variation higher than a threshold in cross-modal data, and based on a single variable test to filter redundant features that have no significant association with a disease phenotype, to generate the uniform format dataset.

[0014] In some embodiments, the integrated analysis of the uniform format dataset using an unsupervised multi-omics factor analysis algorithm to identify potential factors affecting the main variation sources of multi-omics data modalities comprises: constructing a hierarchical factor model containing shared factors and unique factors of multi-modal data, wherein the shared factors capture cross-modal common variation, and the unique factors retain unique characteristics of each modality; through sparse regularization to constrain the factor loading matrix, focusing the potential factors on key regulatory modules related to long non-coding RNA classification; using the Bayesian information criterion to determine the optimal number of factors, wherein the factor number inference process combines long non-coding RNA functional module enrichment verification to ensure that the identified potential factors correspond to biologically interpretable regulatory programs.

[0015] In some embodiments, the analysis of the heterogeneity factors affecting different omics data from the perspective of factor analysis to construct a mapping relationship between potential factors and multi-omics data characteristics comprises: for regulatory mechanism differences related to long non-coding RNA classification, identifying significantly associated regulatory pathways of each factor through enrichment analysis; for tumor-related pathological phenotype variation, establishing a partial least squares regression model of potential factors and cancer cell proliferation, invasion phenotype to quantify the contribution of factors to the phenotype; through structural equation modeling to verify the causal relationship between potential factors and multi-omics characteristics, forming a mapping network containing regulatory levels as the mapping relationship.

[0016] In some embodiments, the functional annotation of the identified potential factors based on the biological function enrichment analysis, regulatory network inference and cross-modality data correlation comprises: performing gene ontology function enrichment and Kyoto Encyclopedia of Genes and Genomes pathway analysis on each potential factor, screening specific enrichment items related to nuclear scaffold function or cytoplasmic translation regulation based on the subcellular localization information of long non-coding RNA; constructing a regulatory module of the potential factor and a coding gene through co-expression network analysis to identify key mRNA nodes regulated by the factor; and inferring the core regulatory protein corresponding to the potential factor by using the cross-modality correlation of RNA-protein interaction data and expression profile data to form a three-level annotation system of factor-regulatory molecule-function pathway for functional annotation of the potential factor.

[0017] In some embodiments, the missing value estimation of the multi-omics data based on the mapping relationship comprises: constructing a conditional probability model based on the mapping relationship between the potential factor and the multi-modality data, predicting the mapping coordinates of the missing data in the low-dimensional space through the factor loading matrix, and then back-projecting to the original data space; for the missing values of the RNA-protein interaction data, the regulatory protein information annotated by the factor is used for priority weighted filling to preferentially repair the interaction items related to known tumor drivers; and the output analysis result comprises a biological function annotation report of each potential factor, a cross-modality data completion matrix and a factor-phenotype correlation heat map, which is used for visualizing the contribution degree sorting of the potential factor to the long non-coding RNA function prediction and molecular marker identification.

[0018] In a second aspect, the present application provides a factor analysis method based on long non-coding RNA multi-omics integrated analysis, which is applied to the factor analysis system based on long non-coding RNA multi-omics integrated analysis provided in any of the embodiments of the present application; the method comprises:

[0019] obtaining multi-modality omics data related to long non-coding RNA, the multi-modality omics data at least including RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data and associated phenotype or disease state data;

[0020] performing standardization, preliminary processing of missing values and feature screening on the multi-modality omics data to generate a pre-processed unified format data set;

[0021] performing integrated analysis on the unified format data set by using an unsupervised multi-omics factor analysis algorithm to identify potential factors affecting the main variation sources of the multi-omics data modalities, the potential factors including hidden variables reflecting the lineage specificity, spatiotemporal expression characteristics and disease correlation of long non-coding RNA;

[0022] From the perspective of factor analysis, the heterogeneity factors behind different omics data are analyzed, and the mapping relationship between potential factors and multi-omics data characteristics is constructed. The heterogeneity factors include differences in regulatory mechanisms related to LncRNA classification and variations in tumor-related pathological phenotypes.

[0023] Based on biological function enrichment analysis, regulatory network inference and cross-modal data association, the potential factors are functionally annotated, and the mapping relationship is used to estimate the missing values of multi-omics data. The analysis results containing factor annotation information and complete data are outputted to improve the efficiency of LncRNA function prediction and molecular marker identification.

[0024] In a third aspect, a computer device is provided, which includes a memory and a processor. The memory is configured to store a computer program. The processor is configured to execute the computer program and implement the method provided in any of the embodiments of the present application when executing the computer program.

[0025] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program. The computer readable instructions are executed by the processor to cause one or more processors to execute the method provided in any of the embodiments of the present application.

[0026] The factor analysis system based on LncRNA multi-omics integration analysis provided in the present application includes the following core modules: a multi-omics data acquisition module: acquiring LncRNA related multi-modal data, including RNA-protein interaction data (based on RNA immunoprecipitation technology), LncRNA expression profile data and tumor pathological phenotype data, covering various LncRNA classifications (antisense type, enhancer type, etc.) and spatiotemporal expression characteristics. A data preprocessing module: standardizing, preliminarily processing missing values and feature selection of multi-modal data, generating a unified format data set, and solving the problem of data heterogeneity. An unsupervised factor analysis module: using an improved factor analysis algorithm to identify potential factors reflecting LncRNA lineage specificity and disease correlation, and separating cross-modal common variation and modal specific variation. A heterogeneity analysis module: analyzing differences in LncRNA classification related regulatory mechanisms and variations in tumor pathological phenotypes, and constructing the mapping relationship between potential factors and multi-omics characteristics. A factor annotation and extended analysis module: based on functional enrichment, regulatory network inference and cross-modal association, the potential factors are functionally annotated, and the mapping relationship is used to optimize the estimation of missing values, and the complete analysis results containing annotation information are outputted.

[0027] The system integrates the classification characteristics of LncRNA (such as antisense type and enhancer type) and tumor phenotype data, solves the problem of insufficient analysis of specific regulation mechanism of LncRNA by traditional methods, separates common and unique variations by unsupervised factor analysis, combines sparse regularization and Bayesian criterion to optimize the factor model, and ensures that the latent factors correspond to explainable biological regulation programs (such as core pathways of tumor malignant transformation). A three-level annotation system of "factor-regulatory molecule-function pathway" is established to associate LncRNA subcellular localization, regulatory proteins and functional pathways, and improve the biological explanation accuracy of the factors. Combined with the latent factors and regulatory network information, the missing values are filled, the key regulatory interaction items are repaired preferentially, and the data integrity and analysis reliability are improved. Through integrated analysis and accurate annotation, the efficiency of LncRNA function prediction and tumor molecular marker identification is significantly improved, and data support is provided for individualized medical treatment.

[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0030] Figure 1 is a structural schematic block diagram of a factor analysis system based on long-chain non-coding RNA multi-omics integrated analysis provided by an embodiment of the present application;

[0031] Figure 2 is a step schematic flow chart of a factor analysis method based on long-chain non-coding RNA multi-omics integrated analysis provided by an embodiment of the present application;

[0032] Figure 3 is a structural schematic block diagram of a computer device provided by an embodiment of the present application.

[0033] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0035] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order can be changed according to actual conditions.

[0036] It should be understood that, in order to facilitate clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. also do not necessarily mean different.

[0037] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in the specification and the appended claims of the present application, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0038] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0039] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0040] Long non-coding RNAs (LncRNAs) are a class of important regulatory molecules, whose lineage specificity, spatiotemporal expression characteristics, and key roles in tumor development have been widely demonstrated. LncRNAs participate in gene regulation through various mechanisms, including acting as molecular scaffolds, regulating alternative splicing, and absorbing microRNAs (miRNAs), and their abnormal expression is closely related to tumor malignant transformation, cell proliferation, and invasion phenotypes. However, the functional analysis of LncRNAs faces multiple challenges: the classification of LncRNAs (antisense, enhancer, intergenic, etc.) corresponds to significantly different regulatory mechanisms, and the heterogeneity of multi-omics data (such as RNA-protein interaction data, expression profile data, and clinical phenotype data) makes it difficult for traditional single-omics analysis to reveal their complex regulatory networks.

[0041] The existing multi-omics integration method usually adopts simple data merging or statistical correlation-based analysis strategy, which cannot effectively separate the common variation and specific variation in different modal data, especially lacks specific analysis of core heterogeneity factors such as LncRNA lineage specificity, classification-related regulatory mechanism and tumor pathological phenotype variation. For example, the traditional factor analysis method does not combine the biological characteristics (such as subcellular localization, functional classification) of LncRNA for model optimization, resulting in insufficient biological interpretability of the identified latent factors; the missing value estimation process is not associated with the regulatory network information of LncRNA, affecting the data integrity and analysis accuracy. In addition, the existing technology lacks a systematic functional annotation system for latent factors, making it difficult to establish a hierarchical association of “factor-regulatory molecule-functional pathway”, which limits the efficiency of LncRNA function prediction and molecular marker identification.

[0042] Therefore, in view of the complexity and specificity of LncRNA multi-omics data, an analysis system capable of integrating multi-modal data, analyzing heterogeneous regulatory factors and realizing accurate functional annotation is needed to break through the technical bottleneck of traditional methods in LncRNA function analysis.

[0043] To solve the above problems, please refer to Figure 1 The present application provides a factor analysis system based on long non-coding RNA multi-omics integration analysis, comprising: a multi-omics data acquisition module configured to acquire multi-modal omics data related to long non-coding RNA, the multi-modal omics data at least including RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data and associated phenotype or disease state data; a data preprocessing module configured to standardize, preliminarily process missing values and select features of the multi-modal omics data, and generate a preprocessed unified format data set; an unsupervised factor analysis module configured to use an unsupervised multi-omics factor analysis algorithm to integrate and analyze the unified format data set, identify latent factors affecting the main variation sources of the multi-omics data modal, and the latent factors include hidden variables reflecting the lineage specificity, spatiotemporal expression characteristics and disease correlation of long non-coding RNA; a heterogeneity analysis module configured to analyze the heterogeneity factors affecting different omics data from the perspective of factor analysis, and construct the mapping relationship between the latent factors and the multi-omics data characteristics, the heterogeneity factors including the regulatory mechanism difference related to the classification of long non-coding RNA and the pathological phenotype variation related to tumor; a factor annotation and extension analysis module configured to perform functional annotation on the identified latent factors based on biological function enrichment analysis, regulatory network inference and cross-modal data association, and estimate the missing values of the multi-omics data based on the mapping relationship, and output the analysis results containing factor annotation information and complete data, so as to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

[0044] The application provides a factor analysis system based on long non-coding RNA (LncRNA) multi-omics integrated analysis, which realizes integrated analysis of the whole process from data acquisition to functional annotation through modular design in view of the complexity and heterogeneity of LncRNA multi-modal data, and integrates three types of core data: LncRNA basic characteristic data: sequence characteristics, subcellular localization (nuclear / cytoplasm) and space-time expression data (expression profile in different tissues, development stages or disease states) of different classified LncRNA such as antisense type, enhancer type and intergenic type; regulation interaction data: RNA-protein interaction data (including chromatin modification proteins, RNA binding proteins, miRNA and other interaction molecules) obtained through RNA immunoprecipitation (RIP) technology; phenotype correlation data: pathological grading of tumor samples, cell proliferation index (such as Ki-67), invasion ability parameters (such as matrix metalloproteinase expression) and clinical prognosis data (survival time, recurrence rate), covering at least two or more tumor types of paired cancer tissue and paracancer tissue samples.

[0045] The data acquisition module is used for collecting multi-modal data, ensuring that the LncRNA classification characteristics, regulation interaction relationship and disease phenotype correlation are covered; the pretreatment module is used for solving the data heterogeneity problem, generating a unified format data set through standardization, missing value processing and feature selection; the factor analysis module is used for identifying potential factors based on an improved unsupervised algorithm, separating cross-modal common variation (such as tumor core regulation program) and modal specific variation (such as tissue-specific interaction network); the heterogeneity analysis module is used for analyzing the differences in LncRNA classification related regulation mechanisms (such as histone modification regions enriched by enhancer type LncRNA) and tumor pathological phenotype variation, and constructing a “factor-feature” mapping relationship; the annotation and expansion module is used for establishing a hierarchical functional annotation system, and combining missing value estimation to improve data integrity and analysis accuracy.

[0046] The data type refinement is performed by collecting sequence length, genomic localization (position of adjacent genes), subcellular localization (nuclear molecular scaffold type / cytoplasmic miRNA adsorption type) and space-time expression data (expression profile of different sample types obtained through RNA-seq) of LncRNA according to classification (antisense type, enhancer type, etc.); when LncRNA binding proteins are captured through the RIP technology, the cell substructure (such as chromosomal regions in the nucleus or ribosomes in the cytoplasm) and interaction molecule types (protein domains, miRNA seed sequences, etc.) where the interaction occurs are recorded synchronously; the phenotype data acquisition includes pathological section score (such as WHO grading) of tumor tissue, cell function experiment data (Transwell invasion experiment result) and clinical follow-up data (progression-free survival time, overall survival time).

[0047] Standardization strategy includes quantile normalization for RNA-protein interaction data to eliminate technical bias across different batches of RIP experiments; variance stabilizing transformation (e.g. DESeq2 algorithm) for LncRNA expression profile data to fit the over-dispersed feature of count data.

[0048] Missing value treatment adopts K-Nearest Neighbors imputation based on lineage-specific weights: according to the classification of LncRNAs (e.g. enhancer-type LncRNAs are highly expressed in specific tissues), set the sample similarity weight, and fill in the missing values with the same lineage sample data first.

[0049] Feature selection filters the top 30% of features with high variability in cross-modality data through coefficient of variation (CV), retains LncRNAs and interacting molecules with high dynamic regulation; reduces the data dimension by filtering redundant features with no significant association (p-value > 0.05) with disease phenotypes based on univariate tests (e.g. Wilcoxon rank-sum test).

[0050] Hierarchical factor model construction establishes a mixed model containing shared factors (capturing common variability across modalities, such as core regulatory programs related to tumor malignant transformation) and unique factors (retaining unique features of each modality, such as tissue specificity of RNA-protein interaction networks); focuses potential factors on key regulatory modules related to LncRNA classification (e.g. alternative splicing events associated with intron-type LncRNAs) by sparse regularization (e.g. L1 penalty term) to constrain factor loading matrix. Factor number optimization: determine the number of factors using Bayesian Information Criterion (BIC), and verify the enrichment degree of LncRNA functional modules (e.g. significance of GO pathway enrichment) to ensure that each factor corresponds to a biologically interpretable regulatory program (e.g. cell cycle regulation, EMT pathway).

[0051] Regulatory mechanism analysis identifies factor-associated regulatory pathways (e.g. enhancer-type LncRNA factors significantly enrich H3K27ac histone modification regions, suggesting their involvement in enhancer activity regulation) through enrichment analysis for LncRNA classification differences; establishes partial least squares regression (PLS-R) models of potential factors and cancer cell proliferation, invasion phenotypes to quantify the contribution of factors to phenotypes (e.g. factor 1 explains 60% of the variation in invasion ability).

[0052] Mapping relationship construction: verify the causal relationship between potential factors and multi-omics features through structural equation modeling (SEM), and construct hierarchical networks (e.g. "genomic mutations -> LncRNA expression abnormalities -> regulatory protein interaction changes -> pathological phenotypes").

[0053] Primary annotation: based on GO / KEGG enrichment analysis, combined with LncRNA subcellular localization to screen specific functional items (such as nuclear factor enrichment "chromosome organization" pathway, cytoplasmic factor enrichment "mRNA degradation" pathway); secondary annotation: through co-expression network analysis, identify key mRNA nodes regulated by factors (such as oncogene MYC or tumor suppressor PTEN), and construct "factor-LncRNA-mRNA" regulatory module; tertiary annotation: using the cross-modal association of RIP data and expression profiles, infer the core regulatory proteins corresponding to the factors (such as RNA polymerase II complex or splicing proteins), and form a "factor-regulatory molecule-function pathway" three-level association.

[0054] Construct conditional probability model: map missing data to low-dimensional factor space through factor loading matrix, and predict the coordinates of missing values in the original space using the correlation between factors; for missing entries in RNA-protein interaction data, prioritize filling based on factor-annotated regulatory protein information (such as preferentially repairing interactions related to known tumor driver factor p53).

[0055] Address LncRNA heterogeneity: Unlike general multi-omics tools, the system focuses on LncRNA classification features (antisense / enhancer type, etc.) and lineage specificity, separates common and unique variations through hierarchical factor models, and avoids biological signal loss caused by "one-size-fits-all" analysis.

[0056] Improve the biological interpretability of factors: through sparse regularization and functional module enrichment verification, ensure that the identified latent factors correspond to real regulatory programs (such as only retaining factors that are enriched in tumor-related pathways), rather than meaningless statistical noise.

[0057] Hierarchical annotation improves application value: the three-level annotation system (factor → regulatory molecule → pathway) converts abstract statistical factors into verifiable biological hypotheses (such as "factor A affects tumor cell cycle by regulating splicing proteins"), providing a clear direction for experimental verification. Optimize data integrity and analysis accuracy: based on missing value filling strategy based on regulatory network information, preferentially repair key regulatory interactions, avoid ignoring biological background by traditional interpolation method, and improve the reliability of subsequent analysis (such as marker screening).

[0058] By integrating RIP interaction data, expression profiles, and phenotype data, the system can identify regulatory hubs that cannot be discovered by traditional single-omics analysis (such as LncRNAs that participate in both RNA splicing and miRNA adsorption), providing multi-dimensional evidence for tumor molecular marker identification. Incorporate tumor pathological grading, prognosis data, and other clinical phenotypes to directly associate factors with disease phenotype variations, accelerating the research process from mechanism analysis to clinical translation (such as screening LncRNA factors related to prognosis as candidate markers).

[0059] The application breaks through the limitations of traditional multi-omics analysis on LncRNA specific regulation mechanism by biological feature driven algorithm optimization (such as LncRNA classification oriented factor modeling) and hierarchical data integration strategy, and constructs a closed-loop system from data acquisition to functional annotation, which provides an efficient tool for LncRNA function research and tumor precision medicine.

[0060] In some embodiments, the obtaining includes obtaining multi-modal omics data related to long non-coding RNA, including: obtaining sequence feature data, subcellular localization data and spatiotemporal expression data of a plurality of classified long non-coding RNAs; when RNA immunoprecipitation technology is used to obtain RNA-protein interaction data, the interaction information of chromatin modification proteins, RNA binding proteins and microRNAs combined with long non-coding RNA is captured synchronously, and the multi-modal omics data is obtained; the associated phenotype or disease state data includes pathological grading, cell proliferation index, invasion ability parameter and clinical prognosis data of tumor samples, and the data covers paired cancer tissue and pericancer tissue samples of at least two or more tumor types; the classification includes at least antisense type, enhancer type, intergenic type, bidirectional type and intron type.

[0061] The embodiments specify the specific composition of multi-modal omics data, including LncRNA classification feature data (sequence features, subcellular localization, spatiotemporal expression), RNA-protein interaction data (interaction information of chromatin modification proteins, RNA binding proteins, miRNAs) and tumor phenotype data (paired cancer / cancer-adjacent tissue samples covering ≥2 tumor types, including pathological grading, proliferation / invasion parameters, clinical prognosis). The classification covers five core LncRNA types: antisense type, enhancer type, intergenic type, bidirectional type and intron type, ensuring data integrity and biological specificity.

[0062] LncRNA classification feature collection: sequence features: extract the length, genomic coordinates (such as the position information of antisense LncRNA adjacent to the target gene), open reading frame (ORF) integrity of each classified LncRNA; subcellular localization: determine the localization of LncRNA in the nucleus (such as enhancer type enriched in chromatin region) or cytoplasm (such as intron type enriched near ribosome) by fluorescence in situ hybridization (FISH) experiment or subcellular component separation combined with RNA-seq; spatiotemporal expression: obtain the expression profile by RNA-seq for cancer / cancer-adjacent tissue of different tumor types (such as lung cancer, liver cancer), different developmental stage cell lines, and mark the tissue source and pathological stage (such as TNM stage) of the sample.

[0063] Interaction data capture employs RNA immunoprecipitation (RIP) combined with mass spectrometry (MS) or high-throughput sequencing to capture proteins (such as chromatin-modifying protein EZH2 and RNA-binding protein HNRNPA1) and miRNAs (recording miRNA seed sequences and LncRNA binding sites), while simultaneously recording subcellular regions (nuclear / cytoplasmic) where interactions occur.

[0064] Phenotypic data standardization includes: pathological grading: scoring tumor samples according to WHO standards (e.g., adenocarcinoma G1-G4 grade); functional parameters: measuring cell proliferation index using the CCK-8 assay and invasive ability (number of cells that can penetrate the membrane) using the Transwell assay; clinical prognosis: collecting patients' overall survival (OS), disease-free survival (DFS), and related treatment plans such as surgery / chemotherapy.

[0065] It covers 5 LncRNA classifications and ≥2 tumor types, avoiding the limitations of single classification or disease type analysis, and ensuring that factor analysis can capture common regulatory mechanisms across lineages and tumor-specific differences; it directly links LncRNA molecular characteristics (sequence, location) with functional interactions (protein / miRNA binding) and clinical phenotypes, providing a data foundation for subsequent cross-level analysis of "molecular characteristics → regulatory mechanisms → disease phenotypes"; it clarifies classification boundaries (e.g., bidirectional LncRNAs are adjacent to bidirectional promoter regions), enabling subsequent analyses to specifically focus on the regulatory characteristics of different classifications (e.g., enhancer-type LncRNAs are enriched with enhancer histone modifications).

[0066] In some embodiments, the standardization, preliminary processing of missing values, and feature screening of multimodal omics data to generate a preprocessed unified format dataset includes: using quantile standardization to eliminate batch effects for RNA-protein interaction data; performing variance stabilization transformation on long non-coding RNA expression profile data to adapt to non-normal distribution characteristics; using K-nearest neighbor interpolation based on a potential low-dimensional factor space to initially fill missing values, wherein the interpolation process incorporates lineage-specific expression pattern weights for long non-coding RNAs; screening features with variability higher than a threshold in cross-modal data using the coefficient of variation, and filtering redundant features that are not significantly associated with disease phenotypes based on univariate tests, thereby generating the unified format dataset.

[0067] To address data heterogeneity, we propose a modal standardization strategy (quantile standardization and variance stabilization transformation), lineage-specific missing value imputation (based on K-nearest neighbor interpolation in low-dimensional factor space, combined with lineage weights), and feature selection (coefficient of variation + univariate test) to generate a unified format dataset, thereby eliminating technical noise and redundant information for subsequent modeling.

[0068] The standardized method includes RNA-protein interaction data: quantile normalization is performed on different batches of RIP experimental data to make the distribution of each batch of data consistent, and to eliminate the bias caused by antibody efficiency and sequencing depth difference; LncRNA expression profile: DESeq2 or EdgeR is used for variance stabilization transformation, and the count data (reads number) is converted into approximate normal distribution, which is suitable for the assumption of data distribution of factor analysis. Missing value filling: constructing LncRNA lineage-specific weight matrix: according to LncRNA classification (such as intron-type LncRNA mainly expressed in specific tissues), the similarity of lineages between samples is calculated (the same classification sample weight is greater than or equal to 0.8, and the different classification is less than or equal to 0.5); K-nearest neighbor interpolation based on low-dimensional factor space: dimensionality reduction is performed on the data by principal component analysis (PCA), and K most similar samples (weight weighted) are found in the factor space, and the missing values are filled by weighted average, and the information of the same lineage sample is preferentially used. Feature selection: coefficient of variation (CV) screening: features with CV greater than 0.5 are retained (high cross-sample variability, excluding low dynamic regulation molecules); univariate test: Wilcoxon rank-sum test is performed on each feature, and features with no significant association (p value greater than 0.05) with tumor phenotype (such as pathological grade) are filtered out, and the data dimension is compressed to 30%-50% of the original data.

[0069] The multi-modal standardization specifically solves the technical noise of different data types (such as batch effect of RIP and overdispersion of RNA-seq), avoids the interference of experimental errors on the analysis results; the missing value filling of lineage-specific weight avoids the damage of traditional interpolation methods (such as mean filling) to LncRNA classification characteristics, and retains key signals such as “enhancer-type LncRNA is only highly expressed in specific cell types”; through double screening of coefficient of variation and univariate test, the core features with high variability and phenotype association are focused, the computational complexity is reduced, and the efficiency and accuracy of subsequent factor analysis are improved.

[0070] In some embodiments, the unsupervised multi-omics factor analysis algorithm is used to integrate and analyze the unified format data set, and to identify potential factors affecting the main variation source of multi-omics data modalities, including: constructing a hierarchical factor model containing shared factors and unique factors of multi-modal data, wherein the shared factors capture the common variation across modalities, and the unique factors retain the unique characteristics of each modality; through sparse regularization constraint factor loading matrix, the potential factors are focused on the key regulation modules related to LncRNA classification; the Bayesian information criterion is used to determine the optimal number of factors, and the factor number inference process is combined with the functional module enrichment degree verification of LncRNA to ensure that the identified potential factors correspond to biologically interpretable regulation programs.

[0071] By constructing hierarchical factor model (shared factor + unique factor), through sparse regularization constraint factor loading matrix, combined with Bayesian information criterion (BIC) and functional module enrichment degree verification, identify the potential factors reflecting the classification regulation mechanism of LncRNA and the variation of tumor phenotype, ensure the biological interpretability of the factors.

[0072] Hierarchical model construction: mathematical model: let the multi-modal data be X m = Λ m F + Ψ m E m, wherein F is a shared factor (cross-modal common variation, such as tumor driving pathway), and Ψ m E m is a modality-specific factor (such as tissue-specific expression program specific to expression profile); sparse constraint: L1 regularization is applied on the factor loading matrix Λ m, so that each factor is strongly associated with only a few key features (such as the expression amount of specific classification LncRNA, key interacting protein), avoiding the biological meaning blurred by the dispersion of factor load.

[0073] Factor number determination: initial screening: select the factor number with the minimum BIC value through the BIC formula BIC = -2lnL + klnn (L is the likelihood function, k is the number of parameters, and n is the sample number); biological verification: perform GO pathway enrichment analysis on each candidate factor, and only keep the factors with significant enrichment p < 0.01, to ensure that each factor corresponds to a real regulation program (such as cell apoptosis pathway factor, EMT pathway factor).

[0074] Shared factors capture cross-modal common mechanisms (such as LncRNA proliferation-promoting pathways common to all tumor types), and unique factors retain modality specificity (such as the chromatin remodeling network specific to RNA-protein interaction data), avoiding the mixed masking of heterogeneous signals by traditional single-factor models; sparse regularization forces the model to focus on classification-related key features (such as only enhancing the strong association between the expression amount of LncRNA and a factor), so that the factor loading matrix can be directly mapped to specific regulation modules (such as “factor 2 = enhancer LncRNA-histone modification protein interaction module”); combined with BIC and functional enrichment double verification, eliminate meaningless statistical factors (such as factors that only reflect technical variation), and ensure that the final factors have clear biological direction (such as factors directly related to tumor invasion).

[0075] In some embodiments, the analysis of the heterogeneous factors affecting different omics data from the perspective of factor analysis, and the mapping relationship between the potential factors and the multi-omics data features, include: for the differences in classification-related regulation mechanisms of long non-coding RNA, identify the regulation pathways significantly associated with each factor through enrichment analysis; for tumor-related pathological phenotype variation, establish a partial least squares regression model of potential factors and cancer cell proliferation, invasion phenotype, and quantify the contribution of factors to the phenotype; through structural equation model, verify the causal relationship between potential factors and multi-omics features, and form a mapping network containing regulation levels as the mapping relationship.

[0076] By enrichment analysis, partial least squares regression (PLS-R), structural equation model (SEM), the differences in classification-related regulatory mechanisms of LncRNA and the variation of tumor phenotype are analyzed, a causal mapping network of "potential factors → multi-omics characteristics → disease phenotypes" is constructed, and the hierarchical regulation relationship of heterogeneous factors is revealed. The classification difference analysis adds classification screening conditions (such as only analyzing enhancer-type LncRNA-related characteristics) when performing GO enrichment on each factor, identifies the pathways significantly enriched by the factor (such as enhancer-type factors enriched in "transcription regulator binding" GO entries); the PLS-R model is used to establish the regression relationship between the factor score and the cancer cell proliferation / invasion parameters, and the contribution of the factor to the phenotype is calculated (such as factor 3 explaining 70% of the Ki-67 index variation). The SEM model is constructed: the path assumption of "LncRNA classification characteristics → regulatory interaction changes → potential factor activation → phenotype variation" is set, the model fitting degree is verified by chi-square test (p>0.05), RMSEA (<0.08), and the regulation hierarchy is determined (such as "high expression of antisense LncRNA → binding HNRNPA1 → activating factor 1 → promoting cell invasion"). Abstract problems such as "classification-related regulatory mechanism differences" and "tumor phenotype variation" are converted into quantifiable factor contribution (such as determining that factor A mainly reflects the regulatory mechanism of enhancer-type LncRNA, and factor B is mainly associated with the invasion phenotype of liver cancer), avoiding the ambiguous handling of heterogeneity by traditional analysis; the regulation network is hierarchical: the SEM verifies the causal relationship, and a multi-level mapping network containing molecular characteristics (LncRNA classification), regulatory interactions (protein binding), and phenotype effects is constructed, providing a clear hypothesis path for mechanism research (such as "the activation of factor C depends on the intron-type LncRNA-spliceosome protein interaction, which further affects the cell cycle"); the PLS-R quantifies the contribution of the factor to the phenotype, which can focus on high-contribution factors (such as factors with a contribution degree of >50%) and accelerate the screening of key regulatory factors (such as functional verification of the core factor for lung cancer invasion).

[0077] In some embodiments, the biological function enrichment analysis, regulatory network inference, and cross-modal data association are used to perform functional annotation on the potential factors, including: performing gene ontology function enrichment and Kyoto Encyclopedia of Genes and Genomes pathway analysis on each potential factor, and screening specific enrichment entries related to nuclear scaffold function or cytoplasmic translation regulation based on the subcellular localization information of long non-coding RNA; constructing a regulatory module of potential factors and coding genes through co-expression network analysis, and identifying key mRNA nodes regulated by the factors; using the cross-modal association of RNA-protein interaction data and expression profile data to infer the core regulatory proteins corresponding to the potential factors, and forming a three-level annotation system of factor-regulatory molecule-function pathway to perform functional annotation on the potential factors.

[0078] Establish a three-level annotation system: the first level is based on GO / KEGG enrichment (combined with subcellular localization), the second level identifies key mRNA nodes through co-expression networks, and the third level infers core regulatory proteins using cross-modal data to form a hierarchical annotation of "factor → regulatory molecule → functional pathway" and convert statistical factors into verifiable biological hypotheses.

[0079] First-level annotation (pathway enrichment): Subcellular localization screening: For nuclear localization factors, preferentially retain "chromatin organization" and "transcription regulation" related GO entries; for cytoplasmic localization factors, focus on "mRNA metabolism" and "translation regulation" pathways; KEGG pathway analysis: use the clusterProfiler package for enrichment, set FDR < 0.05, and retain tumor-related pathways (such as PI3K-AKT and MAPK signaling pathways). Second-level annotation (co-expression network): Calculate the Spearman correlation coefficient between factor score and mRNA expression, and screen mRNA nodes with |p| > 0.6 and p < 0.01; Construct regulatory modules: visualize through Cytoscape, identify hub genes (degree > 10, such as cancer gene MYC or tumor suppressor gene TP53), and label "factor → LncRNA → mRNA" regulatory relationships (such as factor 4 up-regulating LncRNA-ATB expression, which in turn promotes VEGFA transcription). Third-level annotation (regulatory protein inference): Cross-modal association: integrate RNA-protein interaction data and factor load, and screen proteins with strong association with factors (absolute value of load > 0.5) (such as chromatin modification protein EZH2 and splicing factor SF3B1); Function verification suggestion: combine co-expression mRNA nodes and pathway enrichment results to form a complete regulatory chain (such as "factor 5 → binds EZH2 → inhibits tumor suppressor gene promoter H3K27me3 modification → activates Wnt pathway → promotes tumor proliferation").

[0080] From pathway enrichment (macro function) to key molecules (specific targets) to regulatory proteins (action mechanism), gradually refine the biological significance of factors, avoiding the ambiguity of single enrichment analysis (such as "factor X enriches cancer pathways" upgrading to "factor X regulates Wnt pathway to promote invasion by binding EZH2"); three-level annotation directly provides verifiable hypotheses (such as "can LncRNA-protein interaction associated with factor knockout affect pathway activity"), shortening the cycle from basic analysis to wet experiment verification; Combine expression profiles, interaction data, and localization information to ensure that the annotation results have multidimensional support (such as nuclear localization factors associated with chromatin modification proteins, consistent with the biological common sense of subcellular functional partitioning), improving the credibility of the annotation.

[0081] In some embodiments, the binding mapping relationship performs missing value estimation on multi-omics data, and outputs an analysis result containing factor annotation information and complete data, including constructing a conditional probability model based on the mapping relationship between latent factors and multi-modal data, predicting the mapping coordinates of missing data in a low-dimensional space through a factor loading matrix, and then back-projecting to the original data space; for missing values of RNA-protein interaction data, the priority weighted filling is performed using the regulatory protein information annotated by the factors, and the interaction entries related to known tumor drivers are preferentially repaired; the output analysis result contains biological function annotation reports of each latent factor, cross-modal data completion matrix, and factor-phenotype association heat map, which is used for visualizing the contribution degree of latent factors to LncRNA function prediction and molecular marker identification.

[0082] By proposing a conditional probability model driven missing value filling (based on factor-feature mapping relationship), key regulatory interaction entries are preferentially repaired, and the analysis result containing function annotation report, complete data matrix, and factor-phenotype heat map is output, which supports the visual analysis of LncRNA function prediction and marker identification.

[0083] Missing value estimation: conditional probability model: in the factor space, it is assumed that the missing data point xij satisfies xij ~ N(Λjfi, Ψj), where fi is the factor score of sample i, and the posterior probability is calculated through Bayesian inference to predict the coordinates of the missing value in the original space; for missing entries of RNA-protein interaction data, if the interaction protein belongs to known tumor drivers (such as p53, Ras family), the weight coefficient is increased by 50%, and such key interactions are preferentially filled.

[0084] The function annotation report contains factor number, classification association (such as “factor 2 = enhancer type LncRNA regulatory factor”), enriched pathway, key regulatory protein list, and phenotype contribution degree; the complete data matrix is used for marking filled data (such as red marked interpolation entries), and the filling method is explained (such as “interpolation based on factor 3 of adjacent samples”); the factor-phenotype heat map includes visualizing the Spearman correlation coefficient of factor score and pathological grade, survival time, etc. (color depth represents correlation, asterisk marks significance), and highlights high contribution factors (such as factors negatively correlated with OS highlighted in blue).

[0085] By combining the regulatory protein information annotated by factors, the missing values of interactions related to tumor drivers are filled preferentially, avoiding the loss of key signals caused by the equal treatment of all missing values in traditional methods (such as missing the interaction between LncRNA and p53, which may mask the anticancer mechanism); the visualization heat map supports quick screening of factors with strong association with phenotypes (such as directly locating to factor 6 related to prognosis), the functional report provides clear biological interpretation, and the use threshold of non-professional users is reduced; the conditional probability model uses the correlation between factors to infer missing values, which can better preserve the potential structure of the data (such as the distribution pattern of samples in the factor space) compared to the simple interpolation method (such as mean filling), and improve the accuracy of subsequent machine learning modeling (such as marker screening).

[0086] The present application solves the problem of insufficient analysis of specific regulatory mechanisms of LncRNA by the traditional method by integrating the classification characteristics of LncRNA (such as antisense type, enhancer type) and tumor phenotype data through the system. By separating common and unique variations through unsupervised factor analysis, combining sparse regularization and Bayesian criteria to optimize the factor model, the potential factors are ensured to correspond to explainable biological regulatory programs (such as core pathways of tumor malignant transformation). A three-level annotation system of "factor-regulatory molecule-function pathway" is established to associate LncRNA subcellular localization, regulatory proteins and functional pathways, and improve the biological interpretation accuracy of factors. Combined with the missing value filling of potential factors and regulatory network information, key regulatory interaction items are repaired preferentially, and the data integrity and analysis reliability are improved. Through integrated analysis and accurate annotation, the efficiency of LncRNA function prediction and tumor molecular marker identification is significantly improved, providing data support for individualized medical treatment.

[0087] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a factor analysis method based on long non-coding RNA multi-omics integrated analysis provided by an embodiment of the present application. The execution device of the method is the computer device of the factor analysis system based on long non-coding RNA multi-omics integrated analysis provided by any embodiment of the present application.

[0088] As Figure 2 shown, the provided method includes steps S101 to S105. The computer device can be a handheld terminal, a notebook computer, a wearable device, or a robot, etc. The embodiments for implementing steps S101 to S105 and their corresponding embodiments are used.

[0089] Step S101. Obtain multi-modal omics data related to long non-coding RNA, which at least includes RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data, and associated phenotype or disease state data;

[0090] Step S102. Standardize, preliminary handle missing values and feature selection on multi-modal omics data, and generate a pre-processed uniform format data set;

[0091] Step S103. Integrate and analyze the uniform format data set by using unsupervised multi-omics factor analysis algorithm, identify potential factors affecting the main variation sources of multi-omics data modalities, and the potential factors include hidden variables reflecting the lineage specificity, spatio-temporal expression characteristics and disease correlation of long non-coding RNA;

[0092] Step S104. Analyze the heterogeneity factors behind different omics data from the perspective of factor analysis, and construct the mapping relationship between potential factors and multi-omics data characteristics, and the heterogeneity factors include the differences in regulatory mechanisms related to long non-coding RNA classification and the pathological phenotype variations related to tumors;

[0093] Step S105. Based on biological function enrichment analysis, regulatory network inference and cross-modality data association, functionally annotate the resolved potential factors, and estimate the missing values of multi-omics data based on the mapping relationship, and output the analysis results containing factor annotation information and complete data, to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

[0094] In some embodiments, the multi-modal omics data related to long non-coding RNA is obtained, including: obtaining sequence feature data, subcellular localization data and spatio-temporal expression data of multiple classified long non-coding RNAs; when RNA immunoprecipitation technology is used to obtain RNA-protein interaction data, the interaction information of chromatin modification proteins, RNA binding proteins and microRNAs combined with long non-coding RNAs is captured synchronously, and the multi-modal omics data is obtained; the associated phenotype or disease state data includes pathological grading, cell proliferation index, invasion ability parameter and clinical prognosis data of tumor samples, and the data covers paired cancer tissue and paracancer tissue samples of at least two or more tumor types; the classification at least includes antisense type, enhancer type, intergenic type, bidirectional type and intron type.

[0095] In some embodiments, the multi-modal omics data is standardized, the missing values are preliminarily handled and the features are selected to generate a pre-processed uniform format data set, including: using quantile standardization to eliminate batch effects for RNA-protein interaction data, and using variance stabilization transformation for long non-coding RNA expression profile data to adapt to non-normal distribution characteristics; using K-nearest neighbor interpolation method based on latent low-dimensional factor space to preliminarily fill the missing values, and the interpolation process combines the expression mode weight of long non-coding RNA lineage specificity; through coefficient of variation, the features with variation higher than a threshold in cross-modality data are screened, and based on univariate test, the redundant features with no significant association with disease phenotype are filtered to generate the uniform format data set.

[0096] In some embodiments, the unsupervised multi-omics factor analysis algorithm is used to perform integrated analysis on the unified format dataset, and identify potential factors affecting the main variation sources of multi-omics data modalities, including: constructing a hierarchical factor model containing shared factors and unique factors of multi-modal data, wherein the shared factors capture cross-modality common variation, and the unique factors retain the unique characteristics of each modality; focusing the potential factors on the key regulatory modules related to long non-coding RNA classification by sparse regularization constraint factor loading matrix; determining the optimal number of factors by Bayesian information criterion, and the factor number inference process combines the enrichment degree of functional modules of long non-coding RNA to ensure that the identified potential factors correspond to biologically interpretable regulatory programs.

[0097] In some embodiments, the heterogeneity factors behind the different omics data are analyzed from the perspective of factor analysis, and a mapping relationship between potential factors and multi-omics data characteristics is constructed, including: for the differences in regulatory mechanisms related to long non-coding RNA classification, identifying the regulatory pathways significantly associated with each factor through enrichment analysis; for tumor-related pathological phenotype variation, establishing a partial least squares regression model of potential factors and cancer cell proliferation and invasion phenotype to quantify the contribution of factors to the phenotype; verifying the causal relationship between potential factors and multi-omics characteristics through structural equation modeling to form a mapping network containing regulatory levels as the mapping relationship.

[0098] In some embodiments, the potential factors are functionally annotated based on biological function enrichment analysis, regulatory network inference, and cross-modality data association, including: performing gene ontology function enrichment and Kyoto Encyclopedia of Genes and Genomes pathway analysis on each potential factor, and screening specific enrichment items related to nuclear scaffold function or cytoplasmic translation regulation based on the subcellular localization information of long non-coding RNA; constructing a regulatory module of potential factors and coding genes through co-expression network analysis to identify key mRNA nodes regulated by factors; using the cross-modality association of RNA-protein interaction data and expression profile data to infer the core regulatory proteins corresponding to the potential factors, and forming a three-level annotation system of factor-regulatory molecule-function pathway to functionally annotate the potential factors.

[0099] In some embodiments, the binding mapping relationship performs missing value estimation on multi-omics data, and outputs an analysis result containing factor annotation information and completed data, including constructing a conditional probability model based on the mapping relationship between the latent factors and the multi-modal data, predicting the mapping coordinates of the missing data in the low-dimensional space through the factor loading matrix, and then back-projecting to the original data space; for the missing values of RNA-protein interaction data, the priority weighted filling is performed using the regulatory protein information annotated by the factors, and the interaction entries related to the known tumor driver factors are preferentially repaired; the output analysis result contains a biological function annotation report of each latent factor, a cross-modal data completion matrix, and a factor-phenotype association heat map, which is used to visualize the contribution degree of the latent factors to the long non-coding RNA function prediction and molecular marker identification.

[0100] It should be noted that, for the convenience and brevity of description, the specific working processes of the above-described factor analysis method based on long non-coding RNA multi-omics integrated analysis and each step can be clearly understood by those skilled in the art, and the corresponding processes in the above-described factor analysis system embodiments based on long non-coding RNA multi-omics integrated analysis can be referred to, which will not be described here.

[0101] Please refer to Figure 3 , Figure 3 is a structural schematic block diagram of a computer device provided by the embodiment of the present application. The computer device includes a processor, a memory and a network interface connected through a device bus, wherein the memory can include a storage medium and an internal memory.

[0102] The storage medium can store an operating device and a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute any one of the embodiments of the factor analysis method based on long non-coding RNA multi-omics integrated analysis.

[0103] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0104] The internal memory provides an environment for the execution of the computer program in the non-volatile storage medium, which, when executed by the processor, can cause the processor to execute any one of the factor analysis system methods based on long non-coding RNA multi-omics integrated analysis.

[0105] The network interface is used for network communication, such as sending assigned tasks. Those skilled in the art can understand that Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0106] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0107] In one embodiment, the processor is configured to run a computer program stored in the memory to implement the following steps:

[0108] Obtaining multi-omics data related to long non-coding RNA, the multi-omics data at least including RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data and associated phenotype or disease state data;

[0109] Standardizing, performing preliminary processing of missing values and performing feature screening on the multi-omics data to generate a pre-processed unified format data set;

[0110] Integrating and analyzing the unified format data set by using unsupervised multi-omics factor analysis algorithm to identify potential factors affecting the main variation sources of the multi-omics data modalities, the potential factors including hidden variables reflecting long non-coding RNA lineage specificity, spatio-temporal expression characteristics and disease correlation;

[0111] Analyzing the heterogeneity factors behind different omics data from the perspective of factor analysis, and constructing the mapping relationship between the potential factors and the multi-omics data characteristics, the heterogeneity factors including differences in regulatory mechanisms related to long non-coding RNA classification and pathological phenotype variations related to tumors;

[0112] Based on biological function enrichment analysis, regulatory network inference and cross-modality data correlation, the potential factors are functionally annotated, and the multi-omics data are estimated for missing values, and the analysis results including factor annotation information and complete data are outputted, so as to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

[0113] In some embodiments, the obtaining the multi-omics data associated with long non-coding RNA comprises: obtaining sequence feature data, subcellular localization data, and spatiotemporal expression data of a plurality of classified long non-coding RNAs; obtaining RNA-protein interaction data based on RNA immunoprecipitation technology, and simultaneously capturing interaction information of chromatin modification proteins, RNA binding proteins, and microRNAs combined with the long non-coding RNA to obtain the multi-omics data; the associated phenotype or disease state data comprises pathological grading, cell proliferation index, invasion ability parameter, and clinical prognosis data of tumor samples, and the data covers paired cancer tissue and paracancer tissue samples of at least two tumor types; and the classification comprises antisense type, enhancer type, intergenic type, bidirectional type, and intron type.

[0114] In some embodiments, the standardizing, preliminary processing of missing values, and feature screening of the multi-omics data to generate a preprocessed unified format data set comprises: eliminating batch effects by using quantile normalization for RNA-protein interaction data, and performing variance stabilization transformation on long non-coding RNA expression profile data to adapt to non-normal distribution characteristics; performing preliminary filling of missing values by using K-nearest neighbor interpolation based on a latent low-dimensional factor space, and the interpolation process combines lineage-specific expression pattern weights of long non-coding RNAs; filtering features with a variation degree higher than a threshold in cross-modality data by using a coefficient of variation, and filtering redundant features that have no significant association with disease phenotypes based on univariate testing to generate the unified format data set.

[0115] In some embodiments, the integrating analysis of the unified format data set by using an unsupervised multi-omics factor analysis algorithm to identify latent factors of main variation sources affecting multi-omics data modalities comprises: constructing a hierarchical factor model containing shared factors and unique factors of multi-modality data, wherein the shared factors capture cross-modality common variation, and the unique factors retain unique characteristics of each modality; focusing latent factors on key regulatory modules related to long non-coding RNA classification by using sparse regularization to constrain factor loading matrices; determining an optimal number of factors by using Bayesian information criterion, and the factor number inference process combines verification of enrichment degrees of functional modules of long non-coding RNAs to ensure that the identified latent factors correspond to biologically interpretable regulatory programs.

[0116] In some embodiments, the factor analysis is used to analyze the heterogeneity behind different omics data, and a mapping relationship between potential factors and omics data features is constructed, including: for the differences in regulatory mechanisms related to long non-coding RNA classification, significant regulatory pathways associated with each factor are identified through enrichment analysis; for tumor-related pathological phenotype variations, a partial least squares regression model of potential factors and cancer cell proliferation and invasion phenotypes is established to quantify the contribution of factors to the phenotypes; the causal relationship between potential factors and multi-omics features is verified by a structural equation model to form a mapping network containing a regulatory hierarchy as the mapping relationship.

[0117] In some embodiments, the potential factors are functionally annotated based on biological function enrichment analysis, regulatory network inference, and cross-modal data association, including: gene ontology function enrichment and Kyoto Encyclopedia of Genes and Genomes pathway analysis are performed on each potential factor, and specific enrichment items related to nuclear scaffold function or cytoplasmic translation regulation are screened based on the subcellular localization information of long non-coding RNA; a regulatory module of potential factors and coding genes is constructed through co-expression network analysis to identify key mRNA nodes regulated by the factors; core regulatory proteins corresponding to the potential factors are inferred using cross-modal association of RNA-protein interaction data and expression profile data, and a three-level annotation system of factor-regulatory molecule-function pathway is formed to functionally annotate the potential factors.

[0118] In some embodiments, the multi-omics data is subjected to missing value estimation based on the mapping relationship, and an analysis result containing factor annotation information and complete data is output, including: a conditional probability model is constructed based on the mapping relationship between potential factors and multi-modal data, the missing data is predicted in the low-dimensional space based on the factor loading matrix, and then is back-projected to the original data space; for missing values of RNA-protein interaction data, the regulatory protein information annotated by the factors is used for priority weighted filling, and the interaction items related to known tumor drivers are preferentially repaired; the output analysis result includes biological function annotation reports of each potential factor, cross-modal data completion matrix, and factor-phenotype association heat map, which is used to visualize the contribution degree of potential factors to long non-coding RNA function prediction and molecular marker identification.

[0119] It should be noted that, for the convenience and brevity of description, the specific working process of the processor described above can refer to the corresponding process in the method embodiments described in the above embodiments, which will not be described here.

[0120] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program comprises program instructions. The processor executes the program instructions to realize the steps of the factor analysis method based on long-chain non-coding RNA multi-omics integrated analysis provided by each embodiment of the present application.

[0121] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like.

[0122] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A factor analysis system based on long-chain non-coding RNA multi-omics integration analysis, characterized in that, The method comprises the following steps: A multi-omics data acquisition module is configured to acquire multi-modal omics data related to long non-coding RNA, which at least includes RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data, and associated phenotype or disease state data; A data preprocessing module is configured to standardize, preliminarily process missing values, and filter features of the multi-modal omics data, and generate a preprocessed unified format data set; An unsupervised factor analysis module is configured to use an unsupervised multi-omics factor analysis algorithm to integrate and analyze the unified format data set, identify potential factors that affect the main variation sources of the multi-omics data modalities, and the potential factors include hidden variables reflecting the lineage specificity, spatiotemporal expression characteristics, and disease correlation of long non-coding RNA; A heterogeneity analysis module is configured to analyze the heterogeneity factors behind different omics data from the perspective of factor analysis, and construct a mapping relationship between the potential factors and the multi-omics data features, and the heterogeneity factors include differences in regulatory mechanisms related to long non-coding RNA classification and pathological phenotype variations related to tumors; A factor annotation and extension analysis module is configured to perform functional annotation on the resolved potential factors based on biological function enrichment analysis, regulatory network inference, and cross-modality data association, and estimate missing values of the multi-omics data based on the mapping relationship, and output analysis results containing factor annotation information and complete data to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

2. The system of claim 1, wherein, The method for acquiring multi-modal omics data related to long non-coding RNA comprises the following steps: Obtain sequence feature data, subcellular localization data, and spatiotemporal expression data of multiple classified long non-coding RNAs; When obtaining RNA-protein interaction data based on RNA immunoprecipitation technology, capture the interaction information of chromatin modification proteins, RNA binding proteins, and microRNAs combined with long non-coding RNA at the same time, and obtain the multi-modal omics data; the associated phenotype or disease state data includes pathological grading, cell proliferation index, invasion ability parameter, and clinical prognosis data of tumor samples, and the data covers paired cancer tissue and paracancer tissue samples of at least two or more tumor types; the classification at least includes antisense type, enhancer type, intergenic type, bidirectional type, and intron type.

3. The system of claim 1, wherein, The method for standardizing, preliminarily processing missing values, and filtering features of the multi-modal omics data to generate a preprocessed unified format data set comprises the following steps: For RNA-protein interaction data, use quantile standardization to eliminate batch effects, and perform variance stabilization transformation on long non-coding RNA expression profile data to adapt to non-normal distribution characteristics; Use K-nearest neighbor interpolation method based on a low-dimensional latent factor space to preliminarily fill the missing values, and the interpolation process combines the expression mode weight of the lineage specificity of long non-coding RNA; Filter features with a variation degree higher than a threshold in cross-modality data through coefficient of variation, and filter redundant features with no significant association with disease phenotypes based on univariate test, and generate the unified format data set.

4. The system of claim 1, wherein, The unsupervised multi-omics factor analysis algorithm is used to perform integrated analysis on the unified format data set, identify potential factors affecting the main variation sources of multi-omics data modalities, and the potential factors include: A hierarchical factor model containing multi-modal data shared factors and unique factors is constructed, wherein the shared factors capture the common variation across modalities, and the unique factors retain the unique characteristics of each modality; By sparse regularization constraint factor loading matrix, the potential factors are focused on the key regulatory modules related to long non-coding RNA classification; the optimal number of factors is determined by the Bayesian information criterion, and the factor number inference process is combined with the enrichment degree verification of the functional modules of long non-coding RNA, to ensure that the identified potential factors correspond to biologically interpretable regulatory programs.

5. The system of claim 1, wherein, The heterogeneity factors behind the different omics data are analyzed from the perspective of factor analysis, and the mapping relationship between the potential factors and the multi-omics data characteristics is constructed, including: For the differences in regulatory mechanisms related to long non-coding RNA classification, the regulatory pathways significantly associated with each factor are identified through enrichment analysis; For tumor-related pathological phenotype variations, a partial least squares regression model of potential factors and cancer cell proliferation and invasion phenotypes is established to quantify the contribution of factors to the phenotype; The causal relationship between the potential factors and the multi-omics characteristics is verified by structural equation modeling, and a mapping network containing regulatory levels is formed as the mapping relationship.

6. The system of claim 1, wherein, The potential factors are functionally annotated based on biological function enrichment analysis, regulatory network inference, and cross-modality data association, including: Gene ontology function enrichment and Kyoto Encyclopedia of Genes and Genomes pathway analysis are performed on each potential factor, and the nuclear scaffold function or cytoplasmic translation regulation related specific enrichment items are screened combined with the subcellular localization information of long non-coding RNA; The regulatory modules of potential factors and coding genes are constructed through co-expression network analysis to identify key mRNA nodes regulated by factors; The core regulatory proteins corresponding to the potential factors are inferred using the cross-modality association of RNA-protein interaction data and expression profile data, and a three-level annotation system of factor-regulatory molecule-function pathway is formed to functionally annotate the potential factors.

7. The system of claim 1, wherein, The missing values of multi-omics data are estimated based on the mapping relationship between potential factors and multi-modal data, and the analysis results containing factor annotation information and complete data are output, including: Based on the mapping relationship between potential factors and multi-modal data, a conditional probability model is constructed, the missing data is predicted in the low-dimensional space through the factor loading matrix, and then it is back-projected to the original data space; For the missing values of RNA-protein interaction data, the regulatory protein information annotated by factors is used for priority weighted filling, and the interaction items related to known tumor drivers are preferentially repaired; The output analysis results include biological function annotation reports of each potential factor, cross-modality data complete matrix, and factor-phenotype association heat map, which is used to visualize the contribution degree sorting of potential factors to long non-coding RNA function prediction and molecular marker identification.

8. A factor analysis method based on long-chain non-coding RNA multi-omics integration analysis, characterized in that, The method is applied to the long non-coding RNA multi-omics integrated analysis factor analysis system of any one of claims 1-7, and the method comprises: Obtaining multi-modal omics data related to long non-coding RNA, the multi-modal omics data at least including RNA-protein interaction data obtained based on RNA immunoprecipitation technology, long non-coding RNA expression profile data and associated phenotype or disease state data; Standardizing, performing preliminary processing of missing values and performing feature screening on the multi-modal omics data to generate a pre-processed unified format data set; Integrating and analyzing the unified format data set by using unsupervised multi-omics factor analysis algorithm, identifying potential factors of main variation sources affecting the modal of multi-omics data, the potential factors including hidden variables reflecting long non-coding RNA lineage specificity, space-time expression characteristics and disease correlation; Analyzing the heterogeneity factors behind different omics data from the perspective of factor analysis, constructing the mapping relationship between the potential factors and the multi-omics data features, and the heterogeneity factors including differences in regulatory mechanisms related to long non-coding RNA classification and pathological phenotype variation related to tumors; Based on biological function enrichment analysis, regulatory network inference and cross-modal data association, the potential factors are functionally annotated, and the multi-omics data are estimated for missing values, and the analysis results including factor annotation information and complete data are outputted, so as to improve the efficiency of long non-coding RNA function prediction and molecular marker identification.

9. A computer device, comprising: The computer device comprises a memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program and realize the method of claim 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the method of claim 8.

Citation Information

Cited By

  • Multi-omics data three-dimensional adjustment network construction method based on hierarchical Bayesian model

    CN121191568A