Methods, devices, equipment, and storage media for single-sample ncRNA regulatory network inference.

By calculating and fusing the regulatory relationship matrix between ncRNAs and target genes, the single-sample ncRNA regulatory network is inferred, solving the problem that existing technologies cannot identify ncRNA regulation at the single-sample level. This enables the construction of a sample-specific network, providing technical support for disease diagnosis and treatment.

CN115527615BActive Publication Date: 2026-05-26DALI UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALI UNIV
Filing Date
2022-10-17
Publication Date
2026-05-26

Smart Images

  • Figure CN115527615B_ABST
    Figure CN115527615B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and storage medium for inferring a single-sample ncRNA regulatory network, relating to the field of gene identification technology. The specific implementation includes: acquiring ncRNA and target gene transcriptome data of a matched sample; for each target sample in the matched sample, calculating a first regulatory relationship strength matrix and a second regulatory relationship strength matrix of ncRNA and target gene before and after removing the target sample using a preset statistical algorithm; obtaining the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample based on the first and second regulatory relationship strength matrices; and inferring the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample. This invention can reflect the regulatory relationship strength between ncRNA and target gene at the single-sample level by constructing statistical correlation values, thereby inferring the single-sample ncRNA regulatory network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene recognition technology, and more specifically, to a method, apparatus, device, and storage medium for single-sample ncRNA regulatory network inference. Background Technology

[0002] Non-coding ribonucleic acid (ncRNA) is a class of RNA molecules that do not encode proteins. There are many types of ncRNA, including microRNA (miRNA), long non-coding RNA (lncRNA), circular RNA (circRNA), pseudogenes, and other common ncRNA types.

[0003] At the genomic and chromosomal levels, ncRNAs can regulate the expression levels of target genes. A series of studies have shown that ncRNAs can influence chromosome structure, participate in RNA processing and modification, participate in the stability and translational regulation of messenger RNA (mRNA), regulate cell development and differentiation, and are closely related to the occurrence and development of complex human diseases. Previous studies have found that selective gene expression is the basis of sample (e.g., tissue and cell) specificity, and its expression level is regulated by ncRNAs. Therefore, dysregulation of ncRNAs is closely related to many complex human diseases. Previous research suggests that the gene regulatory mechanisms involved in ncRNAs will provide new breakthroughs for disease diagnosis and treatment.

[0004] Current research methods can identify ncRNA regulatory networks at multiple sample levels based on dual transcriptome data of ncRNAs and target genes. However, due to the heterogeneity of each sample, current methods cannot construct ncRNA regulatory networks at the single-sample level. In other words, current methods cannot be used to study ncRNA regulatory networks at the single-sample level. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of the prior art by providing a single-sample ncRNA regulatory network inference method, apparatus, device, and storage medium, which can identify and infer ncRNA regulatory networks at the single-sample level.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention provides a method for inferring a single-sample ncRNA regulatory network, comprising: acquiring ncRNA and target gene transcriptome data of a matching sample, wherein the number of matching samples is multiple; for each target sample in the matching samples, calculating a first regulatory relationship strength matrix of ncRNA and target gene before removing the target sample and a second regulatory relationship strength matrix of ncRNA and target gene after removing the target sample according to a preset statistical algorithm; obtaining a fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample based on the first and second regulatory relationship strength matrices, wherein the fusion regulatory relationship strength matrix is ​​used to indicate the regulatory relationship strength of ncRNA and target gene in the target sample; and inferring the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample.

[0008] In some implementations, the preset statistical algorithm includes any one of the following: correlation method, distance method, information method, and regression method.

[0009] In some implementations, when the preset statistical algorithm is a distance algorithm, before obtaining the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, the method further includes: summing the first regulatory relationship strength matrix with a preset double-precision accuracy to obtain a first summation result; taking the reciprocal of the first summation result as the transformed first regulatory relationship strength matrix; summing the second regulatory relationship strength matrix with the double-precision accuracy to obtain a second summation result; taking the reciprocal of the second summation result as the transformed second regulatory relationship strength matrix.

[0010] The step of obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the first regulatory strength matrix and the second regulatory strength matrix includes: obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the transformed first regulatory strength matrix and the transformed second regulatory strength matrix.

[0011] In some implementations, obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the first regulatory strength matrix and the second regulatory strength matrix includes: calculating a first product between the first regulatory strength matrix and the number of matching samples; calculating a second product between the second regulatory strength matrix and the difference between the number of matching samples and 1; and calculating the difference between the first product and the second product to obtain the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample.

[0012] In some implementations, the step of inferring the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample includes: normalizing the fusion regulatory relationship strength matrix; obtaining the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; determining that there is a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strengths whose significance p-values ​​meet preset conditions in the normalized fusion regulatory relationship strength matrix; and inferring the single-sample ncRNA regulatory network of the target sample based on the regulatory relationship between the ncRNA and the target gene.

[0013] In some implementations, the normalization of the fusion regulation relationship strength matrix includes: obtaining the mean and standard deviation of the fusion regulation relationship strength matrix; calculating the ratio of the difference between each regulation relationship strength and the mean to the standard deviation in the fusion regulation relationship strength matrix to obtain the ratio corresponding to each regulation relationship strength; and replacing the regulation relationship strength with the ratio corresponding to each regulation relationship strength as the normalized regulation relationship strength.

[0014] In some implementations, the ncRNA is any one of microRNA, long non-coding RNA, circular RNA, and pseudogene, and the target gene is messenger RNA.

[0015] Secondly, the present invention also provides a single-sample ncRNA regulatory network inference device, comprising: an acquisition module for acquiring ncRNA and target gene transcriptome data of a matching sample, wherein the number of matching samples is multiple; a statistics module for calculating, according to a preset statistical algorithm, a first regulatory relationship strength matrix of ncRNA and target gene before removing the target sample and a second regulatory relationship strength matrix of ncRNA and target gene after removing the target sample for each target sample in the matching sample; a fusion module for acquiring, based on the first and second regulatory relationship strength matrices, a fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample, wherein the fusion regulatory relationship strength matrix is ​​used to indicate the regulatory relationship strength of ncRNA and target gene in the target sample; and an inference module for inferring the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample.

[0016] In some implementations, the preset statistical algorithm includes any one of the following: correlation method, distance method, information method, and regression method.

[0017] In some implementations, when the preset statistical algorithm is a distance algorithm, before the fusion module obtains the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, it is further used to sum the first regulatory relationship strength matrix with a preset double precision accuracy to obtain a first summation result; take the reciprocal of the first summation result as the transformed first regulatory relationship strength matrix; sum the second regulatory relationship strength matrix with the double precision accuracy to obtain a second summation result; take the reciprocal of the second summation result as the transformed second regulatory relationship strength matrix.

[0018] The fusion module is specifically used to obtain the fusion regulatory relationship strength matrix between the target RNA and the target gene corresponding to the target sample based on the transformed first regulatory relationship strength matrix and the transformed second regulatory relationship strength matrix.

[0019] In some implementations, the fusion module is specifically used to calculate a first product between the first regulatory relationship strength matrix and the number of matching samples; calculate a second product between the second regulatory relationship strength matrix and the difference between the number of matching samples and 1; and calculate the difference between the first product and the second product to obtain the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample.

[0020] In some implementations, the inference module is specifically used to normalize the fusion regulatory relationship strength matrix; obtain the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; determine that there is a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strength with a significance p-value that meets the preset conditions in the normalized fusion regulatory relationship strength matrix; and infer the single-sample ncRNA regulatory network of the target sample based on the regulatory relationship between the ncRNA and the target gene.

[0021] In some implementations, the inference module is specifically used to obtain the mean and standard deviation of the fusion regulation relationship strength matrix; calculate the ratio of the difference between each regulation relationship strength and the mean in the fusion regulation relationship strength matrix to the standard deviation, and obtain the ratio corresponding to each regulation relationship strength; replace the regulation relationship strength with the ratio corresponding to each regulation relationship strength as the normalized regulation relationship strength.

[0022] In some implementations, the ncRNA is any one of microRNA, long non-coding RNA, circular RNA, and pseudogene, and the target gene is messenger RNA.

[0023] Thirdly, the present invention provides an electronic device, comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of any of the methods described in the first aspect.

[0024] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of any of the methods described in the first aspect above.

[0025] The beneficial effects of this invention are as follows: This invention can reflect the strength of the regulatory relationship between ncRNAs and target genes at the single-sample level by constructing statistical correlation values, and then infer the single-sample ncRNA regulatory network. Considering the heterogeneity of each sample, this invention can construct a specific ncRNA regulatory network for each sample, that is, one ncRNA regulatory network corresponds to one sample. This invention can be applied to disease transcriptome data to identify ncRNA regulatory networks in disease samples, providing technical support for the clinical diagnosis and treatment of complex human diseases. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A flowchart illustrating the single-sample ncRNA regulatory network inference method provided in this embodiment of the invention;

[0028] Figure 2 A schematic diagram illustrating the number of miRNA regulatory relationships predicted based on four different analysis methods, provided in an embodiment of the present invention;

[0029] Figure 3 A schematic diagram illustrating the percentage of miRNA regulatory relationship verification based on four different analysis methods provided in this embodiment of the invention;

[0030] Figure 4 A schematic diagram illustrating the number of lncRNA regulatory relationships predicted based on four different analysis methods, provided in an embodiment of the present invention;

[0031] Figure 5 A schematic diagram illustrating the percentage of lncRNA regulatory relationship verification based on four different analysis methods provided in this embodiment of the invention;

[0032] Figure 6 This is a schematic diagram of the structure of a single-sample ncRNA regulatory network inference device provided in an embodiment of the present invention;

[0033] Figure 7 This is a schematic diagram of the electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0035] Current research methods can identify ncRNA regulatory networks at multiple sample levels based on dual transcriptome data of ncRNAs and target genes. However, due to the heterogeneity of each sample, current methods cannot construct ncRNA regulatory networks at the single-sample level. In other words, current methods cannot be used to study ncRNA regulatory networks at the single-sample level.

[0036] Based on this, the present invention provides a method for inferring single-sample ncRNA regulatory networks. This method can construct statistical correlation values ​​to reflect the strength of the regulatory relationship between ncRNA and target genes at the single-sample level, and then infer the single-sample ncRNA regulatory network. It can combine the unique characteristics of each biological sample to reveal the gene regulation mechanism of a single sample and realize the identification of ncRNA regulatory networks at the single-sample level.

[0037] In other words, the technical solution of this invention can be based on dual transcriptome data of ncRNA and target genes, taking into account the heterogeneity of each biological sample, to construct an ncRNA regulatory network for each sample (one ncRNA regulatory network per sample). This method can be applied to disease transcriptome data, aiming to identify ncRNA regulatory networks in disease samples, and provide technical support for the clinical diagnosis and treatment of complex human diseases.

[0038] For example, the execution subject of this single-sample ncRNA regulatory network inference method can be a desktop computer, laptop computer, server, cloud server, smart terminal, tablet computer, or other device with data processing capabilities, and there are no restrictions on this.

[0039] Optionally, this single-sample ncRNA regulatory network inference method can be implemented using multiple programming languages ​​such as R, Python, and C++. During implementation, only a runtime environment for R, Python, or C++ is required.

[0040] Figure 1This is a flowchart illustrating the single-sample ncRNA regulatory network inference method provided in an embodiment of the present invention. Figure 1 As shown, this single-sample ncRNA regulatory network inference method may include:

[0041] S110. Obtain ncRNA and target gene transcriptome data of the matching samples. There are multiple matching samples.

[0042] For example, in a given dual transcriptome dataset of ncRNA and target genes, expression profiles of ncRNA and target genes from the same sample can be extracted to obtain ncRNA and target gene transcriptome data of the matching sample.

[0043] In some implementations, based on the biological characteristics of the matched sample, expression profiles of ncRNAs and target genes with the same biological characteristics can be extracted from the given dual transcriptome data of ncRNAs and target genes to obtain the ncRNA and target gene transcriptome data of the matched sample.

[0044] It should be noted that the given ncRNA and target gene dual transcriptome data can come from databases that provide gene expression profile data (i.e., transcriptome matrices), such as the Gene Expression Omnibus (GEO).

[0045] Optionally, ncRNA can be microRNA (miRNA), long non-coding RNA (lncRNA), circular RNA (circRNA), pseudogene, etc., and the target gene can be messenger RNA (mRNA), without any restrictions.

[0046] S120. For each target sample in the matching samples, calculate the first regulatory relationship strength matrix between ncRNA and target gene before removing the target sample and the second regulatory relationship strength matrix between ncRNA and target gene after removing the target sample, according to the preset statistical algorithm.

[0047] For example, the number of matching samples obtained in S110 can be c, where c is an integer greater than 1; ncRNA can be represented as R = {R1, R2, ..., R...} q}∈R c×q Target gene transcriptome data can be represented as T = {T1, T2, ..., T} p}∈R c×p Where q represents the number of ncRNAs in each matched sample; p represents the number of target genes in each matched sample.

[0048] In some implementations, the preset statistical algorithm may include: correlation methods (e.g., Pearson correlation method), distance methods (e.g., Euclidean distance method), information methods (e.g., mutual information (MI) method), regression methods (e.g., Lasso regression method), etc. This invention does not limit the specific type of preset statistical algorithm.

[0049] Taking the target sample as sample k as an example, according to the preset statistical algorithm, the first regulatory relationship strength matrix between ncRNA and target gene before removing sample k can be represented as X. (k) After removing sample k, the second regulatory relationship strength matrix between ncRNA and target gene can be represented as Y. (k) .

[0050] X (k) and Y (k) They are as follows:

[0051]

[0052]

[0053] in, This indicates the strength of the regulatory relationship between ncRNA(j) and target gene(i) before sample k was removed; This indicates the strength of the regulatory relationship between ncRNA(j) and target gene(i) after removing sample k.

[0054] For related methods, information methods, and regression methods, and The larger the absolute value, the stronger the regulatory relationship between ncRNA(j) and target gene(i); for distance methods, and The smaller the absolute value, the stronger the regulatory relationship between ncRNA(j) and target gene(i).

[0055] To standardize the calculation, the first regulatory relationship strength matrix X between ncRNA and target gene before removing sample k was obtained using a distance method. (k) And the second regulatory relationship strength matrix Y between ncRNA and target gene after removing sample k. (k) After that, it is also possible to target X (k) and Y (k) Further transformation yields the transformed first regulatory relationship strength matrix X′. (k) And the transformed second regulation relationship strength matrix Y′ (k) , such that X′ (k) and Y′ (k)The larger the absolute value of the element (ij) in the expression, the stronger the regulatory relationship between ncRNA (j) and target gene (i).

[0056] Among them, the first regulation relationship strength matrix X (k) The second regulation relationship strength matrix Y (k) Further transformation yields the transformed first regulatory relationship strength matrix X′. (k) and the transformed second regulation relationship strength matrix f′ (k) This can include: the first regulatory relationship strength matrix X (k) The summation is performed with the preset double-precision accuracy to obtain the first summation result; the reciprocal of the first summation result is taken as the transformed first regulation relationship strength matrix X′. (k) The second regulatory relationship strength matrix Y (k) The result is summed with the preset double-precision accuracy to obtain a second summation result; the reciprocal of the second summation result is taken as the transformed second control relationship strength matrix Y′. (k) .

[0057] That is, when the preset statistical algorithm is a distance algorithm, before obtaining the fusion regulatory relationship strength matrix of the target sample corresponding to ncRNA and target gene based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, the method further includes: summing the first regulatory relationship strength matrix with a preset double precision accuracy to obtain a first summation result; taking the reciprocal of the first summation result as the transformed first regulatory relationship strength matrix; summing the second regulatory relationship strength matrix with the double precision accuracy to obtain a second summation result; taking the reciprocal of the second summation result as the transformed second regulatory relationship strength matrix.

[0058] The step of obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the first regulatory strength matrix and the second regulatory strength matrix includes: obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the transformed first regulatory strength matrix and the transformed second regulatory strength matrix.

[0059] For example, X' (k) and Y' (k) They are as follows:

[0060]

[0061]

[0062] Where eps represents double precision accuracy (typically 2.22E-16 by default). For distance methods, and The larger the absolute value, the stronger the regulatory relationship between ncRNA(j) and target gene(i).

[0063] S130. Based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, obtain the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample. The fusion regulatory relationship strength matrix is ​​used to indicate the regulatory relationship strength between ncRNA and target gene in the target sample.

[0064] Taking the target sample as a single sample k as an example, in single sample k, for correlation methods, information methods, and regression methods, the strength Z of the regulatory relationship between ncRNA(j) and target gene(i) is... (k) The definition is as follows:

[0065]

[0066] In a single sample k, for the distance method, the strength Z of the regulatory relationship between ncRNA(j) and target gene(i) is... (k) The definition is as follows:

[0067]

[0068] Where c is the number of matched samples. For correlation methods, distance methods, information methods, and regression methods, The larger the absolute value, the stronger the regulatory relationship between ncRNA(j) and target gene(i) in a single sample k.

[0069] That is, obtaining the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample based on the first regulatory strength matrix and the second regulatory strength matrix includes: calculating the first product between the first regulatory strength matrix and the number of matching samples; calculating the second product between the second regulatory strength matrix and the difference between the number of matching samples and 1; and calculating the difference between the first product and the second product to obtain the fusion regulatory strength matrix of ncRNA and target gene corresponding to the target sample.

[0070] S140. Based on the strength matrix of the fusion regulatory relationship between the target sample and the target gene, the single-sample ncRNA regulatory network of the target sample is deduced.

[0071] In this invention, reasoning can also be understood as identification. The step of reasoning to obtain the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample may include: normalizing the fusion regulatory relationship strength matrix; obtaining the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; determining that there is a regulatory relationship between the ncRNA and target gene corresponding to the regulatory relationship strengths whose significance p-values ​​meet preset conditions in the normalized fusion regulatory relationship strength matrix; and reasoning to obtain the single-sample ncRNA regulatory network of the target sample based on the regulatory relationship between the ncRNA and target gene.

[0072] For example, Z ij (k Since the fusion regulation relationship strength matrix basically follows a normal distribution, normalizing it can be understood as normalizing the strength of each regulation relationship in the matrix. The steps for normalizing the strength of each regulation relationship in the fusion regulation relationship strength matrix include: obtaining the mean and standard deviation of the matrix; calculating the ratio of the difference between each regulation relationship strength and the mean to the standard deviation, thus obtaining the ratio corresponding to each regulation relationship strength; and replacing the regulation relationship strength with the ratio corresponding to each regulation relationship strength as the normalized regulation relationship strength.

[0073] For example, the normalized intensity of the regulatory relationship for:

[0074]

[0075] Where, μ (k) Z represents the strength matrix of the fusion regulation relationship. (k) The mean; σ (k) Z represents the strength matrix of the fusion regulation relationship. (k) The standard deviation.

[0076] Normalized regulation relationship strength matrix Z' (k) for:

[0077]

[0078] For example, the strength of each normalized regulatory relationship Corresponding to a significance p-value, The corresponding significance p-value can be expressed as: The calculation can be performed as follows:

[0079]

[0080] in, express The absolute value of a random number from a standard normal distribution is calculated using the pnorm() function. The probability p-value; The smaller the value, the more likely a regulatory relationship exists between ncRNA(j) and target gene(i) in a single sample k.

[0081] In this invention, a significant p-value meeting a preset condition may include: a significant p-value less than a preset threshold, where the preset threshold can be 0.05, and the magnitude of the preset threshold is not limited. In determining the normalized fusion regulatory relationship strength matrix, the existence of a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strengths where the significant p-value meets the preset condition means that: when... corresponding When the threshold is less than a preset threshold, a regulatory relationship is determined between ncRNA(j) and target gene(i) in a single sample k.

[0082] Following the above method, the regulatory relationships between all ncRNAs and target genes in a single sample k can be obtained. Based on the regulatory relationships between ncRNAs and target genes, the single-sample ncRNA regulatory network of the target sample can be deduced.

[0083] As described above, the single-sample ncRNA regulatory network inference method provided in this invention can identify ncRNA regulatory networks at the single-sample level based on dual transcriptome data of ncRNAs and target genes. Considering the heterogeneity of each sample, this method can construct a specific ncRNA regulatory network for each sample, meaning one sample corresponds to one ncRNA regulatory network. This method can be applied to disease transcriptome data to identify ncRNA regulatory networks in disease samples, providing technical support for the clinical diagnosis and treatment of complex human diseases.

[0084] For example, this method can be applied to malignant tumor transcriptome data to screen out malignant tumor sample-specific ceRNA networks, providing technical support for the clinical diagnosis and treatment of human malignant tumors, which has important biological significance.

[0085] Optionally, the single-sample ncRNA regulatory network inference method provided in this embodiment of the invention can be applied to inferring and identifying the regulatory networks of miRNA, lncRNA, circRNA, and pseudogene and target genes.

[0086] The following examples illustrate how to infer and identify the regulatory networks of miRNA and lncRNA with their target genes.

[0087] Taking miRNA as an example, according to the single-sample ncRNA regulatory network inference method provided in this embodiment of the invention, the process of identifying the single-sample miRNA regulatory network can be as follows:

[0088] 1) Obtain matched sample miRNA and target gene transcriptome data;

[0089] For example, miRNA and mRNA transcriptome data of K562 leukemia cells can be collected from the public database (Gene Expression Omnibus, GEO) (dataset storage number GSE114071). Through preprocessing (removal of duplicate genes, logarithmic processing, and removal of genes consistently expressed in all single cells), expression profiles of 212 miRNAs and 15361 mRNAs from 19 matched K562 leukemia cell samples were obtained. To reduce computational complexity, feature extraction was performed (retaining the top 20% of miRNAs and the top 10% of mRNAs by expression variance), ultimately yielding expression profiles of 43 miRNAs and 1536 mRNAs from the 19 matched K562 leukemia cell samples. Therefore, in this embodiment, R = {R1, R2, ..., R...} q}∈R 19×43 and T = {T1,T2,…,T} p}∈R 19×1536 .

[0090] The experimentally validated miRNA-mRNA regulatory relationships were obtained from two databases, miRTarBase v9.0 and TarBase v8.0, ultimately yielding 762,540 miRNA-mRNA regulatory relationship pairs.

[0091] 2) Identify single-sample miRNA regulatory networks

[0092] Given miRNA and mRNA transcriptome data (R and T) matching K562 leukemia cells, the single-sample miRNA-mRNA regulatory network was identified. In this embodiment, the correlation method Pearson, the distance method Euclidean, the information method MI, and the regression method Lasso were applied to calculate the strength of the miRNA-mRNA regulatory relationship before and after removing leukemia cell k. The significance p-threshold for the miRNA-mRNA regulatory relationship strength was set to 0.05 for each leukemia cell. For each leukemia cell, all miRNA-mRNA regulatory relationships were fused to obtain the miRNA-mRNA regulatory network for each leukemia cell.

[0093] 3) Single-sample miRNA regulatory network analysis

[0094] The identified single-sample miRNA regulatory networks were analyzed from the following two aspects:

[0095] (3.1) Similarity analysis of single-sample miRNA-mRNA regulatory networks

[0096] Given the miRNA-mRNA regulatory network N for sample i and sample j i and N j The similarity S(N) between the two single-sample miRNA-mRNA regulatory networks i N j The calculation is as follows:

[0097]

[0098] Among them, overlap(N) i N j ) represents the logarithm of the identical miRNA-mRNA regulatory relationships in the two miRNA regulatory networks, min(N) i N j S(N) represents the logarithm of the miRNA-mRNA regulatory relationship in the smallest network of the two miRNA regulatory networks. i N j The value of ) ranges from [0, 1], and the larger the value, the more similar the two single-sample miRNA regulatory networks are.

[0099] (3.2) Validation of the single-sample miRNA-mRNA regulatory relationship

[0100] Based on the given experimentally validated miRNA-mRNA regulatory relationships, the overlap rate between the miRNA-mRNA regulatory relationships and the experimentally validated miRNA-mRNA regulatory relationships in a single-sample miRNA regulatory network is obtained. A larger overlap rate indicates more experimentally validated miRNA-mRNA regulatory relationships in that sample.

[0101] Figure 2 This diagram illustrates the number of miRNA regulatory relationships predicted by four different analysis methods, as provided in embodiments of the present invention. Figure 2As shown in this embodiment, the number of single-sample miRNA regulatory relationships predicted by four different statistical methods (Pearson, Euclidean, MI, and Lasso) varied. In 19 leukemia single cells, the median number of single-sample miRNA regulatory relationships predicted by the four methods (Pearson, Euclidean, MI, and Lasso) were 5732, 2219, 5805, and 281, respectively. This result indicates that the MI information method can generally predict more miRNA regulatory relationships in single-sample miRNA regulatory relationship prediction. Furthermore, based on the four different statistical methods (Pearson, Euclidean, MI, and Lasso), the similarity among the 19 leukemia single-sample miRNA-mRNA regulatory networks was generally less than 1. This result indicates that even samples with the same phenotype generally exhibit differences in single-sample miRNA-mRNA regulatory networks.

[0102] Figure 3 This is a schematic diagram illustrating the percentage of miRNA regulatory relationship verification based on four different analytical methods, provided for embodiments of the present invention. Figure 3 As shown, the percentage of validation of single-sample miRNA regulatory relationships using four different statistical methods (Pearson, Euclidean, MI, and Lasso) also differed. In 19 leukemia single cells, the median percentages of validation of single-sample miRNA regulatory relationships using the Pearson, Euclidean, MI, and Lasso methods were 5.80%, 4.00%, 4.46%, and 10.94%, respectively. This result indicates that the Lasso regression method performed best in terms of the percentage of validation of single-sample miRNA regulatory relationships.

[0103] Taking lncRNA as an example, according to the single-sample ncRNA regulatory network inference method provided in this embodiment of the invention, the process of identifying the single-sample lncRNA regulatory network can be as follows:

[0104] 1) Obtain matched sample lncRNA and target gene transcriptome data

[0105] lncRNA and mRNA expression profiles of autism-matched samples were collected from the Gene Expression Omnibus (GEO) database (dataset storage number GSE18123). After preprocessing (removing duplicates and lncRNAs and mRNAs without gene names), expression profiles of 595 lncRNAs and 18114 mRNAs from 104 autism-matched and 82 normalized samples were obtained. Gene differential expression table analysis revealed 119 lncRNAs and 1811 mRNAs differentially expressed between the autism-matched and normalized samples. In this embodiment, only the differentially expressed genes from the 104 autism-matched samples were considered; therefore, R = {R1, R2, ..., R...} q}∈R 104×119 and T = {T1,T2,…,T} p}∈R 104×1811 .

[0106] High-confidence lncRNA-mRNA regulatory relationships were obtained from the ENCORI database, ultimately yielding 44,230 lncRNA-mRNA regulatory pairs for validation.

[0107] 2) Identify single-sample lncRNA regulatory networks

[0108] Given lncRNA and mRNA transcriptome data R and T for matched autism samples, the lncRNA-mRNA regulatory network of a single sample is identified. In this embodiment, the correlation method Pearson, the distance method Euclidean, the information method MI, and the regression method Lasso are applied to calculate the strength of the regulatory relationship between lncRNA and mRNA before and after removing autism sample k. The significance p-threshold for the strength of the lncRNA-mRNA regulatory relationship is set to 0.05 for each autism sample. Within each autism sample, all lncRNA-mRNA regulatory relationships are fused to obtain the lncRNA-mRNA regulatory network for each autism sample.

[0109] 3) Single-sample lncRNA regulatory network analysis

[0110] The identified single-sample lncRNA regulatory network was analyzed from the following two aspects:

[0111] (3.1) Similarity analysis of single-sample lncRNA-mRNA regulatory networks

[0112] Given the lncRNA-mRNA regulatory network M of sample i and sample j i and M j Similarity S(M) between two single-sample lncRNA-mRNA regulatory networks i M jThe calculation is as follows:

[0113]

[0114] Among them, overlap (M i M j ) represents the logarithm of the same lncRNA-mRNA regulatory relationship in the two lncRNA regulatory networks, min(M i M j S(M) represents the logarithm of the lncRNA-mRNA regulatory relationship in the smallest of the two lncRNA regulatory networks. i M j The value range of ) is

[01] , and the larger the value, the more similar the two single-sample lncRNA regulatory networks are.

[0115] (3.2) Validation of the single-sample lncRNA-mRNA regulatory relationship

[0116] Based on a given high-confidence lncRNA-mRNA regulatory relationship, the overlap rate between the lncRNA-mRNA regulatory relationship and the high-confidence miRNA-mRNA regulatory relationship in a single-sample lncRNA regulatory network is obtained. A larger overlap rate indicates more validated lncRNA-mRNA regulatory relationships in that sample.

[0117] Figure 4 This diagram illustrates the number of lncRNA regulatory relationships predicted based on four different analysis methods, as provided in embodiments of the present invention. Figure 4 As shown in this embodiment, the number of single-sample lncRNA regulatory relationships predicted by four different statistical methods (Pearson, Euclidean, MI, and Lasso) varies (e.g., Figure 4 In 104 autism samples, the median number of single-sample lncRNA regulatory relationships predicted by the four methods—Pearson, Euclidean, MI, and Lasso—were 16234.5, 12856, 18602.5, and 2563, respectively. This result indicates that the MI method typically predicts more lncRNA regulatory relationships in single-sample prediction. Furthermore, based on the four different statistical methods (Pearson, Euclidean, MI, and Lasso), the similarity between the lncRNA-mRNA regulatory networks of the 104 autism samples was less than 1. This result suggests that even samples with the same phenotype exhibit differences in their single-sample lncRNA-mRNA regulatory networks.

[0118] Figure 5 This is a schematic diagram illustrating the percentage of lncRNA regulatory relationship verification based on four different analytical methods, provided for embodiments of the present invention. Figure 5 As shown, the percentage of validation of single-sample lncRNA regulatory relationships differed among the four different statistical methods (Pearson, Euclidean, MI, and Lasso). In 104 autism samples, the median percentages of validation of single-sample lncRNA regulatory relationships using the Pearson, Euclidean, MI, and Lasso methods were 0.17%, 0.02%, 0.21%, and 0.48%, respectively. This result indicates that the Lasso regression method performed best in terms of the percentage of validation of single-sample lncRNA regulatory relationships.

[0119] The results from the two examples above regarding the identification of single-sample miRNA regulatory networks and single-sample lncRNA regulatory networks demonstrate that even samples with the same phenotype exhibit differences in their single-sample ncRNA regulatory networks. To study the gene regulatory mechanisms involved by ncRNAs in a single sample, it is essential to identify ncRNA regulatory networks at the single-sample level. Furthermore, single-sample ncRNA regulatory networks can be used to calculate the similarity between samples, providing a new method for sample classification. In summary, the single-sample ncRNA regulatory network inference method proposed in this invention can identify disease-specific ncRNA regulatory networks, providing technical support for the clinical diagnosis and treatment of complex human diseases, and has significant biological implications.

[0120] This invention also provides a single-sample ncRNA regulatory network inference device. Figure 6 This is a schematic diagram of the structure of a single-sample ncRNA regulatory network inference device provided in an embodiment of the present invention. Figure 6 As shown, the device may include: an acquisition module 61, a statistics module 62, a fusion module 63, and an inference module 64.

[0121] The acquisition module 61 is used to acquire ncRNA and target gene transcriptome data of multiple matching samples. The statistics module 62 is used to calculate, for each target sample in the matching samples, a first regulatory relationship strength matrix of ncRNA and target genes before target sample removal, and a second regulatory relationship strength matrix of ncRNA and target genes after target sample removal, according to a preset statistical algorithm. The fusion module 63 is used to obtain the fusion regulatory relationship strength matrix of ncRNA and target genes corresponding to the target sample based on the first and second regulatory relationship strength matrices. The fusion regulatory relationship strength matrix indicates the regulatory relationship strength between ncRNA and target genes in the target sample. The inference module 64 is used to infer the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix of ncRNA and target genes corresponding to the target sample.

[0122] In some implementations, the preset statistical algorithm includes any one of the following: correlation method, distance method, information method, and regression method.

[0123] In some implementations, when the preset statistical algorithm is a distance algorithm, before obtaining the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, the fusion module 63 is further used to sum the first regulatory relationship strength matrix with a preset double precision accuracy to obtain a first summation result; take the reciprocal of the first summation result as the transformed first regulatory relationship strength matrix; sum the second regulatory relationship strength matrix with the double precision accuracy to obtain a second summation result; take the reciprocal of the second summation result as the transformed second regulatory relationship strength matrix.

[0124] The fusion module 63 is specifically used to obtain the fusion regulatory relationship strength matrix of the target sample and the target gene corresponding to the target sample based on the transformed first regulatory relationship strength matrix and the transformed second regulatory relationship strength matrix.

[0125] In some implementations, the fusion module 63 is specifically used to calculate a first product between the first regulatory relationship strength matrix and the number of matching samples; calculate a second product between the second regulatory relationship strength matrix and the difference between the number of matching samples and 1; and calculate the difference between the first product and the second product to obtain the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample.

[0126] In some implementations, the inference module 64 is specifically used to normalize the fusion regulatory relationship strength matrix; obtain the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; determine that there is a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strength with a significance p-value that meets the preset conditions in the normalized fusion regulatory relationship strength matrix; and infer the single-sample ncRNA regulatory network of the target sample based on the regulatory relationship between the ncRNA and the target gene.

[0127] In some implementations, the inference module 64 is specifically used to obtain the mean and standard deviation of the fusion regulation relationship strength matrix; calculate the ratio of the difference between each regulation relationship strength and the mean in the fusion regulation relationship strength matrix to the standard deviation, and obtain the ratio corresponding to each regulation relationship strength; replace the regulation relationship strength with the ratio corresponding to each regulation relationship strength as the normalized regulation relationship strength.

[0128] In some implementations, the ncRNA is any one of microRNA, long non-coding RNA, circular RNA, and pseudogene, and the target gene is messenger RNA.

[0129] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0130] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0131] This invention also provides an electronic device. Figure 7 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Figure 7 As shown, the electronic device includes a processor 71, a computer-readable storage medium 72, and a bus 73. The electronic device may include one or more processors 71, the storage medium 72 is used to store machine-readable instructions, the processor 71 is communicatively connected to the storage medium 72 via the bus 73, and the processor 71 executes the machine-readable instructions stored in the storage medium 72 to perform the steps of the method described in the above method embodiment.

[0132] The electronic device can be a general-purpose computer, server, or mobile terminal, etc., and is not limited thereto. The electronic device is used to implement the methods described in the above-described embodiments of the present invention.

[0133] It should be noted that processor 71 may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, processor may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC) computer, or a microprocessor, or any combination thereof.

[0134] Storage medium 72 may include: mass storage, removable storage, volatile read-write storage, or read-only memory (ROM), or any combination thereof. For example, mass storage may include disks, optical disks, solid-state drives, etc.; removable storage may include flash drives, floppy disks, optical disks, memory cards, zip disks, magnetic tapes, etc.; volatile read-write storage may include random access memory (RAM); RAM may include dynamic random access memory (DRAM), double data rate synchronous dynamic RAM (DDR SDRAM), static random-access memory (SRAM), thyristor-based random access memory (T-RAM), and zero-capacitor RAM, etc. For example, ROMs can include mask read-only memory (MROM), programmable read-only memory (PROM), programmable read-only memory (PEROM), electrically erasable programmable read-only memory (EEPROM), optical disc ROM (CD-ROM), and digital universal disk ROM, etc.

[0135] For ease of explanation, only one processor 71 is described in the electronic device. However, it should be noted that the electronic device of the present invention may also include multiple processors 71, and therefore the steps performed by one processor as described in the present invention may also be performed jointly or individually by multiple processors. For example, if the processor 71 of the electronic device performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.

[0136] Optionally, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method described above.

[0137] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0138] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0139] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0140] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0141] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for inferring a single-sample non-coding ribonucleic acid (ncRNA) regulatory network, characterized in that, include: Obtain ncRNA and target gene transcriptome data of matching samples, wherein there are multiple matching samples; For each target sample in the matched samples, the first regulatory relationship strength matrix between ncRNA and target gene before removing the target sample and the second regulatory relationship strength matrix between ncRNA and target gene after removing the target sample are calculated according to a preset statistical algorithm. Based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample is obtained, and the fusion regulatory relationship strength matrix is ​​used to indicate the regulatory relationship strength between ncRNA and target gene in the target sample; Calculate the first product between the first regulation relationship strength matrix and the number of matched samples; Calculate the second product between the second regulation relationship strength matrix and the difference between the number of matched samples and 1; Calculate the difference between the first product and the second product to obtain the strength matrix of the fusion regulatory relationship between the ncRNA and the target gene corresponding to the target sample; Based on the fusion regulatory relationship strength matrix between the ncRNA and the target gene corresponding to the target sample, the single-sample ncRNA regulatory network of the target sample is deduced. The intensity matrix of the fusion regulation relationship is normalized; Obtain the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; In the normalized fusion regulatory relationship strength matrix, it is determined that there is a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strengths whose significance p-values ​​meet the preset conditions; Based on the regulatory relationship between ncRNAs and target genes, the single-sample ncRNA regulatory network of the target sample is deduced.

2. The method according to claim 1, characterized in that, The preset statistical algorithm includes any one of the following: correlation method, distance method, information method, and regression method, wherein the correlation method includes at least the Pearson correlation method, the distance method includes at least the Euclidean distance method, the information method includes at least the mutual information method, and the regression method includes at least the Lasso regression method.

3. The method according to claim 2, characterized in that, When the preset statistical algorithm is a distance algorithm, before obtaining the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, the method further includes: The first regulation relationship strength matrix is ​​summed with the preset double precision accuracy to obtain the first summation result, wherein the default value of the preset double precision accuracy is generally 2.22E-16; Take the reciprocal of the first summation result as the transformed first regulation relationship strength matrix; The second regulation relationship strength matrix is ​​summed with the double precision accuracy to obtain the second summation result; Take the reciprocal of the second summation result as the transformed second regulatory relationship strength matrix; The step of obtaining the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix includes: Based on the transformed first regulatory relationship strength matrix and the transformed second regulatory relationship strength matrix, the fusion regulatory relationship strength matrix of the ncRNA and target gene corresponding to the target sample is obtained.

4. The method according to claim 1, characterized in that, The normalization of the fusion regulation relationship strength matrix includes: Obtain the mean and standard deviation of the fusion regulation relationship strength matrix; Calculate the ratio of the difference between each regulation relationship strength and the mean in the fusion regulation relationship strength matrix to the standard deviation, and obtain the ratio corresponding to each regulation relationship strength; The normalized regulatory relationship strength is obtained by replacing the ratio corresponding to each regulatory relationship strength with the normalized ratio.

5. The method according to any one of claims 1-4, characterized in that, The ncRNA is any one of microRNA, long non-coding RNA, circular RNA, and pseudogene, and the target gene is messenger RNA.

6. A single-sample ncRNA regulatory network inference device, characterized in that, include: The acquisition module is used to acquire ncRNA and target gene transcriptome data of the matching samples, wherein the number of matching samples is multiple; The statistics module is used to calculate, according to a preset statistical algorithm, the first regulatory relationship strength matrix between ncRNA and target gene before removing the target sample and the second regulatory relationship strength matrix between ncRNA and target gene after removing the target sample for each target sample in the matching sample. The fusion module is used to obtain the fusion regulatory relationship strength matrix of ncRNA and target gene corresponding to the target sample based on the first regulatory relationship strength matrix and the second regulatory relationship strength matrix, wherein the fusion regulatory relationship strength matrix is ​​used to indicate the regulatory relationship strength between ncRNA and target gene in the target sample; The inference module is used to infer the single-sample ncRNA regulatory network of the target sample based on the fusion regulatory relationship strength matrix between the ncRNA and the target gene corresponding to the target sample. The fusion module is specifically used to calculate the first product between the first regulation relationship strength matrix and the number of matching samples; Calculate the second product between the second regulation relationship strength matrix and the difference between the number of matched samples and 1; Calculate the difference between the first product and the second product to obtain the strength matrix of the fusion regulatory relationship between the ncRNA and the target gene corresponding to the target sample; The reasoning module is specifically used to normalize the fusion regulation relationship strength matrix; Obtain the significance p-value of each regulatory relationship strength in the normalized fusion regulatory relationship strength matrix; In the normalized fusion regulatory relationship strength matrix, it is determined that there is a regulatory relationship between the ncRNA and the target gene corresponding to the regulatory relationship strengths whose significance p-values ​​meet the preset conditions; Based on the regulatory relationship between ncRNAs and target genes, the single-sample ncRNA regulatory network of the target sample is deduced.

7. An electronic device, characterized in that, It includes a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions that can be executed by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the method of any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and the computer program, when executed by a processor, performs the method of any one of claims 1-5.