Omics Data Capture and Reconstruction Method Based on JDLGR Algorithm
By applying the Omics data capture and reconstruction method based on JDLGR algorithm in the biological field, and using the joint dictionary Dc to recover the Omics data, the problem of poor accuracy in the reconstruction of Omics data in the prior art is solved, and fast, accurate and low-cost Omics data acquisition is achieved.
Patent Information
- Application Number
- CN202411423831.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The existing omics data reconstruction methods have the problem of poor accuracy in recovery results, especially in the high randomness environment in the biological field, which makes it difficult to ensure the robustness of the algorithm.
Omics data capture and reconstruction method based on JDLGR algorithm is used to restore the omics data by training the joint dictionary Dc, and the combination of high and low-dimensional dictionaries is used to perform sparse representation and recovery of the data.
It realizes rapid, accurate and low-cost acquisition of omics data in the biological field, reduces experimental costs and time, and improves the accuracy of recovery results, and is suitable for practical application scenarios.
Smart Images

Figure CN119311476B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of life sciences, and particularly relates to a method for capturing and reconstructing multi-omics data of different cells. Background Art
[0002] With the progress of life science technology, studying cell states based on multi-omics data has become a more accurate research method. By integrating different omics data such as genomics, transcriptomics, proteomics, and metabolomics, a more comprehensive description of cell states can be provided. In terms of disease treatment, by regulating the expression levels of genes, proteins, and metabolites in cells with drugs, multi-level regulation can be achieved, thereby achieving the treatment purpose. Therefore, quickly, low-costly, and accurately obtaining multi-omics data and sequencing it is the key technology to achieve this goal.
[0003] Common methods for obtaining multi-omics data include gene sequencing, mass spectrometry analysis, and nuclear magnetic resonance spectroscopy, etc. For example, gene expression levels are statistically analyzed for all RNA sequences in cells through gene sequencing technology (such as NGS), and the expression levels of genes are calculated through data processing; proteomics uses mass spectrometry technology to identify and quantify the structural information of proteins and the detailed composition of peptide chains; while metabolomics analyzes the types and concentrations of all metabolites through mass spectrometry or nuclear magnetic resonance spectroscopy.
[0004] Current omics data sequencing methods mainly include traditional sequencing methods and dimensionality reduction observation methods. Traditional methods require sequencing all omics data, and this method needs to be sent to a professional company for inspection. The process is numerous, the workload is huge, and it often takes several weeks, which is extremely inconvenient. To improve the sequencing efficiency, dimensionality reduction observation methods have been proposed. Dimensionality reduction observation reduces the detection amount by downsampling data and shortens the detection time. However, it is difficult to recover the original data sequencing results from the sequencing results of a small part of data. Currently, such methods generally have problems such as low data credibility, complex design of experimental schemes, inability to simultaneously retain the linear and non-linear information of gene data, and high costs, which limit the application scenarios of this method in practice.
[0005] Traditional methods are time-consuming and costly, but the single sequencing method is mature and has high accuracy; dimensionality reduction observation methods are time-consuming, but have low accuracy and complex algorithms. To reduce costs and improve efficiency in sequencing, it is necessary to reduce the sequencing amount, and dimensionality reduction observation methods have achieved remarkable results in the field of image processing. Therefore, the urgent task is to find the essence of the inadaptability of dimensionality reduction observation methods in the biological field and its solutions.
[0006] By summarizing a large amount of experimental experience, the many problems faced by the above-mentioned dimensionality reduction observation methods in the biological field are essentially caused by the high randomness of the biological environment, which is mainly reflected in the following two aspects:
[0007] (1) The problem of biological randomness. The problem of biological randomness refers to the inherent randomness and uncertainty in biological systems, resulting in differences in individuals' gene expression, cell behavior, and disease responses compared to the expected values.
[0008] (2) The problem of high off-target rate. When using biosensors to detect substances in the biological environment through biochemical reactions, signal crosstalk may occur, such as the detection results of non-specific sensors containing substances that are not expected to be detected, as well as the mutual interference between sensors, etc. These problems bring high uncertainty to sequencing.
[0009] The high randomness of biological problems requires higher algorithm robustness than other fields. Blindly applying mature algorithms from other fields without considering the particularity of biological problems is difficult to achieve the required accuracy of the recovery results. Summary of the Invention
[0010] The present invention aims to solve the problem of poor accuracy of the recovery results in the reconstruction of existing omics data.
[0011] An omics data capture and reconstruction method based on the JDLGR algorithm, which uses the trained joint dictionary D c to recover omics data;
[0012] The training process of the joint dictionary D c includes the following steps:
[0013] Use the high-dimensional data X h and the low-dimensional data X l to form the joint data X c The high-dimensional data X h is a matrix, and each column vector in the matrix represents the omics data contained in a sample, and each row vector represents the data of the same omics feature in different samples; the low-dimensional data X l is obtained by randomly sampling the total number of rows of X h ;
[0014] Use the high-dimensional dictionary D h and the low-dimensional dictionary D l to form the joint dictionary D c and train the joint dictionary D c The objective during the training process is as follows:
[0015]
[0016] Among them, is the Lagrange operator; α is the sparse representation vector;
[0017] The use of the trained joint dictionary D cThe process of recovering omics data is based on the trained joint dictionary D c to recover the omics information X of the high-dimensional data from the low-dimensional data Y of the whole sequencing * In this process, first, based on formula (2), use D l and Y to solve for α * ; After obtaining α * , then use α * and D h to calculate X h * ;
[0018]
[0019] where α * is the sparse representation vector;
[0020] The low-dimensional data Y of the whole sequencing is the low-dimensional omics data measured by experiments. The specific omics features included in Y are the same as the type and size of the low-dimensional data X c used when training the joint dictionary D l .
[0021] Furthermore, the high-dimensional data X h is obtained based on the original omics data X, that is, the original omics data X is normalized and then decentralized. The decentralized operation means that each sample in X h subtracts the mean of the samples.
[0022] Furthermore, the joint data X c , the joint dictionary D c are as follows:
[0023]
[0024] where N and M are the number of rows of the X h , X l matrices respectively.
[0025] Furthermore, the process of training the joint dictionary D c includes the following steps:
[0026] S121. Initialize D c and α;
[0027] where D c is randomly assigned any non-zero value and then needs to be decentralized and normalized; α is all assigned zero;
[0028] S122. Substitute the existing D c into formula (1) and optimize this formula with α as the variable to obtain the trained α;
[0029] S123. Substitute the existing α into Equation (1), and optimize this equation with D c as the variable to obtain the trained D c ;
[0030] S124. Repeat S121 - S123 until the number of loop iterations reaches a preset value, and finally obtain the trained joint dictionary D c .
[0031] Furthermore, the process of recovering the omics information X of the high - dimensional data from the low - dimensional data Y of the whole - genome sequencing * includes the following steps:
[0032] S201. Substitute the low - dimensional dictionary component D l and the low - dimensional data Y into Equation (2), and optimize Equation (2) with α * as the variable to obtain the trained sparse representation vector α * ;
[0033] S202. According to the relationship in Equation (3), use α * and D h to solve and recover the data texture matrix X h ;
[0034] X h = D h α * (3)
[0035] S203. According to Equation (4), centralize the recovery of X h to obtain the recovered data X h * :
[0036] X h * = X h + mean(Y) (4)
[0037] where mean(Y) is the sample mean of the whole - genome sequencing low - dimensional data Y
[0038] An omics data capture and reconstruction method based on the JDLGR algorithm, which uses the trained joint dictionary D c to recover omics data;
[0039] The training process of the joint dictionary D c includes the following steps:
[0040] Step 101. Obtain the high - dimensional data X h , X his a matrix, where each column vector in the matrix represents the omics data contained in a sample, and each row vector represents the data of the same omics feature in different samples;
[0041] For the high-dimensional data matrix X h for each column, perform clustering by row, and arrange the results in descending order of correlation;
[0042] Step 102: Divide each sample in X h , that is, each column, into n non-overlapping m×1 blocks of the same size. For each block, perform downsampling at the same position to obtain X l1 , X l2 ,......, X ln . The same-position sampling means that the sampling position is the same relative position in each block. Vertically splice the blocks formed after these samplings as a column in the low-dimensional data block X l , thereby obtaining X l ;
[0043] Step 103: Use the high-dimensional dictionary D h and the low-dimensional dictionary D l to form a combined dictionary D c , and train the combined dictionary D c . The objective during the training process is as follows:
[0044]
[0045] where, is the Lagrange operator; α is the sparse representation vector;
[0046] The process of using the trained combined dictionary D c to recover the omics data includes the following steps:
[0047] For Y, starting from the upper left corner, take non-overlapping m×1 blocks, denoted as y; the low-dimensional data Y of whole-genome sequencing is the low-dimensional omics data measured by experiments, and the specific omics features contained in Y are the same as the type and size of the low-dimensional data X c used when training the combined dictionary D l ; For each y, perform the following processing:
[0048] Step 201: Given the basis D l and y, solve the optimization equation where α * is the sparse representation vector, is the Lagrange operator;
[0049] Step 202: Generate the high-dimensional data texture block x = D h α * , where α* is the result corresponding to the sparse representation vector obtained from the optimization equation in step 201, and x is the recovered high-dimensional data block after decentralization;
[0050] Step 203: Calculate the high-dimensional data block x* = x + mean(y), where mean represents taking the mean;
[0051] Step 204: Take each generated data block x * as a column vector of the high-dimensional data matrix X * and finally obtain X * .
[0052] Furthermore, the high-dimensional data X h is obtained based on the original omics data X, that is, the original omics data X is normalized and then decentralized, where the decentralization operation means that each sample in X h is subtracted by the mean of the samples.
[0053] Furthermore, the joint data X c , the joint dictionary D c is as follows:
[0054]
[0055] where N and M are the number of rows of the X h , X l matrix, respectively.
[0056] Furthermore, the process of training the joint dictionary D c includes the following steps:
[0057] Step 121: Initialize D c and α;
[0058] where D c needs to be decentralized and normalized after being randomly assigned any non-zero value; α is all assigned zero;
[0059] Step 122: Substitute the existing D c into Equation (5), optimize this equation with α as the variable, and obtain the trained α;
[0060] Step 123: Substitute the existing α into Equation (5), optimize this equation with D c as the variable, and obtain the trained D c ;
[0061] Step 124: Repeat steps 121 - 123 until the number of loop iterations reaches the preset value, and finally obtain the trained joint dictionary D c .
[0062] Further, in step 101, for each column of the high-dimensional data matrix X h during the process of clustering by rows, the K-Means clustering algorithm is used for clustering.
[0063] Beneficial effects:
[0064] 1. The present invention combines the theoretical method of reducing the dimension of omics data with the actual biochemical reaction process. For example, in transcriptomics, this method uses PCR reactions to reduce the dimension to obtain gene data, and recovers high-dimensional gene information from it to characterize the state of the organism.
[0065] 2. The biochemical reaction used in the present invention is a traditional amplification method (such as PCR amplification), and no special instrument is required for operation. For example, in transcriptomics, only common PCR amplification reactions are needed during the experiment, and only a small number of genes need to be measured once, which greatly reduces the experimental difficulty and ensures the accuracy of the measurement results.
[0066] 3. The present invention provides a method for obtaining omics data that is short-term, accurate, and low-cost. For example, in transcriptomics, the present invention sparsely samples and measures genes to obtain the gene expression levels of cells. The cost of a single experiment is about 4%-8% of that of next-generation gene sequencing (NGS); the experimental time is 2 hours, which is about 1% of that of NGS. The measurement accuracy meets the accuracy requirements for actual application scenarios such as tumor typing and cell state observation. Description of the drawings
[0067] Figure 1 It is a schematic diagram for partitioning and downsampling each sample.
[0068] Figure 2 It is a schematic diagram of the recovery process in Embodiment 2.
[0069] Figure 3 It is a schematic diagram of a PCR 96-well plate.
[0070] Figure 4 It is a schematic diagram of the data reconstruction and recovery process of the embodiment. Detailed implementation manners
[0071] The key technology to solve the problems in the background art is to reduce the dimension of omics data, obtain omics data through conventional amplification methods, and establish an experimental process that meets the actual application requirements. Taking transcriptomics as an example (according to the central dogma, the methods applicable to transcriptomics can actually be applied to the processing of other omics data), the present invention has established a fast way to obtain the gene expression level of cells. Based on the modular expression of genes and the property of being compressible, first, a high-low dictionary pair for gene data is trained, the modular information of genes is recorded in the low-dimensional dictionary, and the gene data structure in each low-dimensional module is recorded at the corresponding position in the high-dimensional dictionary. Then, through a small number of single-PCR measurements of certain genes, the current cell state is perceived by dimension reduction. The obtained cell state (such as the low-dimensional data Y in transcriptomics) is combined with the high-low dictionary to decode the state of the organism to be measured (high-dimensional gene information). The obtained data is evaluated using the single-gene recovery accuracy, gene enrichment analysis, and information transfer entropy as indicators to verify the recovery effect. The verification results show that the multi-omics data measured and detected by this method can meet the perception requirements of the current organism state.
[0072] The present invention is directed to transcriptomics data. By using the method of training a high-low dictionary pair of genes, a small number of individual measurements are performed on the sparse sampling of high-dimensional genes through single-PCR reactions, and the high-dimensional gene expression level is decoded and obtained. Specifically, the present invention first uses open-source data to establish a high-low dictionary pair for genes, then determines the genes to be measured through a subsampling matrix, uses single-PCR reactions (real-time PCR, qPCR, RNA-seq, microarray, etc.) to obtain gene observations, and finally quickly obtains the gene expression level of cells, establishes the current cell state, and provides data guidance at the gene level for the treatment, prognosis, and pathological analysis of major diseases such as cancer, chronic diseases, and major infectious diseases. The following further describes the present invention in conjunction with specific embodiments. Specific Embodiment 1:
[0074] This embodiment is an omics data capture and reconstruction method based on the JDLGR algorithm. Taking transcriptomics as an example, it can realize an accurate, efficient, and low-cost method for measuring the gene expression level of cells. Among them, we use the method of random sampling to separately measure about 4% of all genes, use the PCR reaction process to realize the observation process of the dimension-reduced data, and use the JDLGR algorithm (Joint Dictionaries Learning for Genomic data Recovery, the joint dictionary learning algorithm for genomic data recovery) to train the gene dictionary through a publicly available gene sequencing dataset, perform dictionary training and reconstruct the gene expression level of cells, and finally meet the requirements of reducing experimental costs, shortening the measurement time, and improving the recovery accuracy. According to the central dogma, the methods applicable to transcriptomics can also be applied to the processing of other omics data.
[0075] According to the sparse coding theory, data can be represented as the product of an overcomplete dictionary and a sparse representation vector. The basic principle of data recovery based on the JDLGR algorithm (Joint Dictionaries Learning for Genomic data Recovery) is to train a joint dictionary with a small amount of high-dimensional and low-dimensional data, and then recover all the high-dimensional data from all the sequenced low-dimensional data through the joint dictionary. The following will elaborate on the basic dictionary training method, the basic data recovery method, the implementation and improvement of the algorithm, and its beneficial results.
[0076] An omics data capture and reconstruction method based on the JDLGR algorithm includes:
[0077] S100. Training the joint dictionary:
[0078] The dictionary training in the present invention is implemented based on an optimization method, using the high-dimensional data X h , the low-dimensional data X l to form the joint data X c to train the joint dictionary D h composed of the high-dimensional dictionary D l and the low-dimensional dictionary D c . Before explaining the training process, the meanings of physical quantities will be explained first:
[0079] X h , X l respectively represent high-dimensional and low-dimensional data matrices. Each column vector in the high-dimensional and low-dimensional data matrices represents the omics data contained in a sample, and each row vector represents the data of the same omics feature in different samples. For example, in transcriptomics, the omics data is the gene expression intensity, which is generally characterized by the Ct value amplified by qPCR. Each column represents the data in a sample, and each row represents the data of a gene in different samples.
[0080] X h , X l as the data used for training is preprocessed based on the original omics data X: X h is obtained by normalizing and decentralizing X, where the decentralizing operation means that each sample in X h subtracts the mean of the sample; X l is obtained by randomly extracting about 4% of the total number of rows of X h .
[0081] The joint data X c , the joint dictionary D c are obtained by vertically concatenating the previous quantities, specifically as follows:
[0082]
[0083] Among them, N and M are X h and X l The number of rows of the matrix. In transcriptomics, its physical meaning is the number of genes contained in X h and X l .
[0084] The specific process of training the dictionary is as follows:
[0085] The ideal joint dictionary D c and the corresponding sparse representation vector α should satisfy X c = D c α. And according to the sparse coding theory, α should be sparse enough to obtain a better recovery effect. Applying the Lagrange dual problem, the above objective can be transformed into the following optimization problem:
[0086]
[0087] Among them, is the Lagrange operator, which is obtained by transforming the Lagrange dual problem, and this hyperparameter is set artificially.
[0088] The solution of this optimization problem is the dictionary training process. The training adopts the strategy of alternately training D c and α, which is as follows:
[0089] S121. Initialize D c and α;
[0090] Among them, after randomly assigning non-zero values to D c , it is necessary to perform decentralization and normalization operations; α is all assigned zero.
[0091] S122. Substitute the existing D c into Equation (1), optimize this equation with α as the variable, and obtain the trained α.
[0092] S123. Substitute the existing α into Equation (1), optimize this equation with D c as the variable, and obtain the trained D c .
[0093] S124. Repeat S121 - S123 until the number of loop rounds reaches the preset value, and finally obtain the trained joint dictionary D c .
[0094] S200. Use the trained joint dictionary to recover omics data:
[0095] Data recovery is based on the trained joint dictionary Dc Recovering the omics information X of high-dimensional data from the low-dimensional data Y of whole sequencing * The core of this process is to first solve for α through D l and Y * and then use α * and D h to calculate X h * . Before explaining the recovery process, first explain the meaning of physical quantities:
[0096] The low-dimensional data Y of whole sequencing is the low-dimensional omics data obtained through experiments (such as the Ct value obtained by qPCR, the mass-to-charge ratio obtained by mass spectrometry, etc.). The specific omics features included in Y are the same as those of the low-dimensional data X c used when training the joint dictionary D l in terms of type and size. Similar to formula (1), use D l and Y to solve for α * to satisfy the relationship shown in formula (2):
[0097]
[0098] The specific process of recovering omics data is as follows:
[0099] S201. Substitute the low-dimensional dictionary component D c in the known joint dictionary D l and the low-dimensional data Y into formula (2), and optimize this formula with α * as the variable to obtain the trained sparse representation vector α * , and α * is the same size as α
[0100] S202. According to the relationship in formula (3), use α * and D h to solve for the recovered data texture matrix X h .
[0101] X h = D h α * (3)
[0102] S203. According to formula (4), centralize the recovery of X h to obtain the recovered data X h * :
[0103] X h * = X h + mean(Y) (4)
[0104] Among them, mean(Y) is the sample mean of the full-sequencing low-dimensional data Y.
[0105] After the above steps, the restored data X is obtained h * . Specific Embodiment 2:
[0107] This embodiment is an omics data capture and reconstruction method based on the JDLGR algorithm. Based on Specific Embodiment 1, this embodiment proposes an optimization and improvement scheme:
[0108] Omics data has the characteristics of high throughput, that is, a large amount of data, and the amount of data to be restored at one time can reach millions. The iterative optimization algorithm used in the algorithm has high requirements for computing power and memory. If calculated directly without processing, the running speed is often slow, and it may even be unable to run to completion due to insufficient running memory.
[0109] In addition, if no operations other than data normalization and decentralization are performed when training the dictionary, it may lead to weak connections between the rows of the high- and low-dimensional data matrices X h 、X l . That is, each row of X l as a feature may represent several rows that are far apart in X h , and the truly adjacent rows in X h may have very different corresponding features, which will result in the sparse representation vector α obtained by training not being sparse enough, not meeting the requirements of sparse coding theory. Therefore, the correlation between the restored data obtained using the dictionary trained by the original algorithm and the original data is relatively low. Sometimes, the correlation of the experimental results is only about 33%, far from meeting the requirements of sequencing accuracy.
[0110] To solve the above problems, an improved JDLGR algorithm based on consensus clustering and block training is proposed on the basis of this Embodiment 1. This algorithm is mainly improved in the following aspects:
[0111] 1. Perform consensus clustering on the training data, and arrange the results in descending order according to the correlation (evaluated by Pearson coefficient or Spearman coefficient). This can ensure high correlation between adjacent rows and sufficient feature extraction;
[0112] 2. Perform block training on the joint dictionary, group and amplify Y, and perform block recovery on the data. This can not only reduce the occupied running memory but also avoid interference to the results caused by too large a dimension crossing.
[0113] The improved JDLGR algorithm involves two major processes: dictionary training and data recovery. The specific processing process is as follows:
[0114] I. The specific process of training the dictionary is as follows:
[0115] Step 101: For each column of the high-dimensional data matrix X h , perform consensus clustering on the rows using the K-Means clustering algorithm, and arrange the results in descending order according to the correlation (evaluated by Pearson coefficient or Spearman coefficient);
[0116] Step 102: Divide each sample in X h , that is, each column, into n blocks of the same size m×1 (the blocks are Group1, Group2, etc. in Figure 1 ). For each block, perform downsampling at the same position to obtain X l1 , X l2 ,......, X ln . The same-position sampling means that the sampling positions are the same relative positions in each block. For example, if the 2nd, 4th, 9th, 10th, 11th, 13th, and 14th rows in the first block are selected for sampling, then the 2nd, 4th, 9th, 10th, 11th, 13th, and 14th rows should also be sampled for other blocks. Vertically splice the blocks formed after these samplings as a column in the low-dimensional data block X l , that is, a sample, as shown in Figure 1 ;
[0117] Step 103: Train the joint dictionary D c according to the dictionary training method in S100. The specific steps are as follows: Preprocessing of training data
[0118] Normalize X h and X l , and decentralize X h . Then splice X h and X l into the joint data X c according to the following formula; At the same time, generate D h and D l of the corresponding size according to the training data size and the set dictionary feature number, and splice and process them into D c
[0119]
[0120] Optimal training of the joint dictionary
[0121] Alternately train D c and α according to Equation (5). The specific steps are as follows:
[0122]
[0123] Step 121: Initialize D c and α;
[0124] Among them, D c needs to be decentralized and normalized after being randomly assigned any non-zero value; α is all assigned zero.
[0125] Step 122: Substitute the existing D c into Equation (5), optimize this equation with α as the variable, and obtain the trained α.
[0126] Step 123: Substitute the existing α into Equation (5), optimize this equation with D c as the variable, and obtain the trained D c .
[0127] Step 124: Repeat Step 121 - Step 123 until the number of loop iterations reaches the preset value, and finally obtain the trained joint dictionary D c .
[0128] It should be noted that: in this solution, the training process is a holistic training, and the restoration is performed in blocks. The restoration process is as Figure 2 shown.
[0129] II. Using the trained joint dictionary to restore omics data:
[0130] For Y, starting from the upper left corner, take non-overlapping m×1 blocks, denoted as y, and perform the following operations for each y:
[0131] Step 201: Based on the low-dimensional dictionary D l and y, solve the optimization equation where α * is the sparse representation vector, is the Lagrange operator;
[0132] Step 202: Generate the high-dimensional data texture block x = D h α * , where α * is the result of the sparse representation vector obtained from the optimization equation in Step 201, and x is the restored decentralized high-dimensional data block;
[0133] Step 203: Calculate the high-dimensional gene data block x * = x + mean(y), where the function of the mean method is to calculate the mean value;
[0134] Step 204: Take each generated data block x * as a column vector of the high-dimensional data matrix X * , and finally return X * .
[0135] Example
[0136] The omics data is captured and reconstructed in the manner of the second specific implementation method. The specific process is as follows:
[0137] Take transcriptomics data qPCR as an example to design the reaction experiment process:
[0138] Design PCR primers for the genes corresponding to the designed observation matrix, and customize the designed upstream and downstream PCR primers to be assembled into the same 96-well plate, such as Figure 3 As shown, the primer concentration in a single well is 10 nmol / ul. At this point, the primer preparation process is completed.
[0139] Through qPCR (PCR) reaction, a qPCR instrument (PCR instrument) is used for amplification process. The number of amplification rounds is generally selected as 30 or 40 rounds, and the entry type data of the genetic data can be obtained, which is recorded as the reduced dimension observation value of the genetic data and recorded in the observation result matrix y (m, 1).
[0140] Then, the omics data is captured and reconstructed based on the JDLGR algorithm, such as Figure 4 The specific process is shown as follows:
[0141] (1) Dictionary training:
[0142] Use the JDLGR algorithm to perform dictionary training. During the training process, for a set of genetic data X(M*N), M is the number of genes and N is the number of samples. First, according to step 101 in the JDLGR dictionary training algorithm, the genes are arranged according to the strength of the correlation through the consensus clustering method based on K-Means to obtain the genetic data set X_sort(M*N). Then, according to step 102 in the JDLGR dictionary training algorithm, the data X_sort is divided into x1, x2..., and x1, x2... are downsampled in the same way (same relative position) to obtain the observation values y1, y2..., and then the dictionary module is trained to obtain the joint dictionary D c The effect of dictionary training is evaluated by Pearson correlation and Spearman correlation, and the number of blocks is adjusted according to the evaluation results.
[0143] (2) Data recovery:
[0144] According to step 201 of the JDLGR data recovery algorithm, the gene Y to be measured is obtained, single-plex PCR primers are designed and the observed value Y is obtained through single-plex PCR experiment, and then according to step 202 of the JDLGR data recovery algorithm, the recovery algorithm is used to solve the recovered gene set X.
[0145] The main process of data recovery is as follows: divide each sample y in Y into k blocks of equal size y1, y2, ..., y i ,......,yk , for each block y i the corresponding sparse representation vector can be obtained according to step 201 the restored decentralized high-dimensional gene data block x can be obtained through step 202 i , and then the high-dimensional gene data block is obtained through step 203 Let x1, x2,......, x i ,......, x k vertically concatenate to obtain the restored high-dimensional omics data x * . For each restored omics data x * horizontally concatenate according to step 204 to obtain the restored omics feature set X.
[0146] Taking transcriptomics and proteomics as examples, the situations and effects of the following specific embodiments are described.
[0147] Example 1:
[0148] In the present invention, the gene amplification method selects a PCR instrument or a qPCR instrument produced by ThermoFisher Company for gene amplification, which includes a supporting 96-well plate, 384-well plate, and fluorescence value reading function. Use the traditional cell RNA extraction method for extraction, and then perform reverse transcription of the gene through a configured 20ul reverse transcription system and a PCR instrument to obtain the corresponding cDNA system. Then configure a 20ul qPCR reaction system. After configuration, add the sampled genes to the reaction systems at the corresponding positions of the 96-well plate or 384-well plate respectively, and start the qPCR amplification process.
[0149] Use publicly available data on websites such as TCGA and microarray to establish a high-low dictionary pair for genes in advance. The present invention can establish a dictionary for all cells with publicly available datasets. During the establishment process, use the JDLGR algorithm (Joint Dictionaries Learning for Genomic data Recovery, the joint dictionary learning algorithm for genomic data recovery) for dictionary training. The trained dictionary can be classified according to different types of cells for subsequent use.
[0150] After the qPCR reaction is completed, obtain the Ct value in the well and record the observed value according to the recording method. Then use the JDLGR algorithm (Joint Dictionaries Learning for Genomic data Recovery, the joint dictionary learning algorithm for genomic data recovery) to reproduce the gene expression profile.
[0151] In this embodiment, the qPCR amplification process can give relatively accurate Ct values. On the premise that the compressive sensing sampling rate is about 6%-10%, the correlation between the gene expression level reproduced by the present invention and the original data is 75%-95% in theory. For different samples, the Pearson correlation coefficient is above 70%, and the Spearman correlation coefficient is above 65%. The effect of reproducing the gene expression profile is good.
[0152] For common PCR and qPCR instruments, 96-well plates and 384-well plates can be used to simultaneously perform amplification experiments on the corresponding number of reaction wells. The number of genes in the reproducible gene expression profile is more than 10,000, and the overall experimental time is within 2-3 hours, which can meet the requirements for the dimension and real-time performance of gene expression profile reproduction in most practical applications.
[0153] Example 2:
[0154] In this embodiment, the amplification process of the PCR and qPCR instruments can be changed to the amplification process of magnetic bead-linked primers. Only the gene primers in the reaction wells need to be correspondingly linked to the magnetic beads, and the rest can be carried out according to the above experimental scheme.
[0155] Example 3:
[0156] In the present invention, the amplification process of the PCR and qPCR instruments can be changed to the amplification process using microarray droplets. Only the reaction system containing gene primers in the reaction wells needs to be correspondingly dropped on the glass slide in the form of a microarray, and the rest can be carried out according to the above experimental scheme.
[0157] Example 4:
[0158] In the present invention, the amplification process of the PCR and qPCR instruments can be changed to the amplification process using the T7 promoter. Only the front end or the back end of the gene primers in the reaction wells needs to be linked with the T7 promoter, and then the amplification is carried out by the way of linear amplification. The rest can be carried out according to the above experimental scheme.
[0159] Example 5:
[0160] In the present invention, the amplification process of the PCR and qPCR instruments can be changed to the amplification process using NGS or microarray. Only the genes to be measured need to be sent to a gene sequencing company, and then the corresponding instrument is used to measure the gene expression level. The rest can be carried out according to the above experimental scheme.
[0161] Example 6:
[0162] According to the central dogma, there is a corresponding relationship between various RNAs in transcriptomics and various proteins in proteomics. Therefore, the methods applicable to transcriptomics can also be applied to proteomics data analysis.
[0163] The above calculation examples of the present invention are only for illustrating in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solution of the present invention still fall within the protection scope of the present invention.
Claims
1. A method for capturing and reconstructing omics data based on a joint dictionary learning (JDLGR) algorithm for genomic data recovery, characterized in that: Using the trained joint dictionary D c Perform omics data recovery; The joint dictionary D c The training process includes the following steps: Using high-dimensional data X h , low-dimensional data X l Composition of joint data X c , high-dimensional data X h is a matrix, each column vector in the matrix represents the omics data contained in a sample, and each row vector represents the data of the same omics feature in different samples; low-dimensional data X l X is selected by random sampling h The total number of rows is extracted; Using a high-dimensional dictionary D h , low-dimensional dictionary D l Composition of joint dictionary D c , for the joint dictionary D c Conduct training. The goals of the training process are as follows: in, is the Lagrange operator; α is the sparse representation vector; The trained joint dictionary D c The process of recovering omics data is based on the trained joint dictionary D c , recover the omics information X of high-dimensional data from the low-dimensional data Y of full sequencing * In this process, firstly based on formula (2), using D l Solve for α with Y * ; get α * Then use α * and D h Calculate X h * ; Among them, α * is a sparse representation vector; Recovering high-dimensional omics information X from low-dimensional data Y of full sequencing * The process includes the following steps: S201, the low-dimensional dictionary component D l Substitute the low-dimensional data Y into formula (2), and use α * For the variable optimization formula (2), we can obtain the sparse representation vector α after training. * ; S202, according to the relationship in formula (3), use α * and D h Solve and restore the data texture matrix X h ; X h =D h a * (3) S203, according to formula (4), X h Recovery centralization, obtain recovery data X h * : X h * =X h +mean(Y) (4) Among them, mean(Y) is the sample mean of the full sequencing low-dimensional data Y; The low-dimensional data Y of the whole sequencing is the low-dimensional omics data measured by experiments. The specific omics features contained in Y are combined with the training joint dictionary D c The low-dimensional data X used l of the same type and size.
2. The method for capturing and reconstructing omics data based on the joint dictionary learning (JDLGR) algorithm for genome data recovery according to claim 1, characterized in that: High dimensional dataX h Based on the original omics data X, that is, the original omics data X is normalized and then decentralized, where the decentralized operation refers to X h The sample mean is subtracted from each sample.
3. The omics data capture and reconstruction method based on the joint dictionary learning JDLGR algorithm for genome data recovery according to claim 1 or 2, characterized in that: Joint DataX c , joint dictionary D c as follows: Among them, N and M are respectively h , Z l The number of rows in the matrix.
4. The method for capturing and reconstructing omics data based on the joint dictionary learning (JDLGR) algorithm for genome data recovery according to claim 3, characterized in that: For the joint dictionary D c The training process includes the following steps: S121, Initialization D c and α; Among them, D c After randomly assigning any non-zero value, decentralization and normalization operations are required; α is assigned to zero; S122, the existing D c Substitute into formula (1), optimize the formula with α as the variable, and obtain the trained α; S123, Substitute the existing α into formula (1), and use D c Optimize the formula for the variable and obtain the trained D c ; S124, repeat S121-S123 until the number of cycles reaches the preset value, and finally obtain the trained joint dictionary D c .
5. A method for capturing and reconstructing omics data based on a joint dictionary learning (JDLGR) algorithm for genomic data recovery, characterized in that: Using the trained joint dictionary D c Perform omics data recovery; The joint dictionary D c The training process includes the following steps: Step 101: Obtain high-dimensional data X h , X h is a matrix, each column vector in the matrix represents the omics data contained in a sample, and each row vector represents the data of the same omics feature in different samples; For a high-dimensional data matrix X h For each column, cluster them by rows and arrange the results from high to low in terms of relevance; Step 102: X h Each sample in the matrix, i.e. each column, is divided into n m×1 blocks of the same size, and each block is downsampled at the same position to obtain X l1 ,X l2 ,……,X ln , the same position sampling means that the relative position of the sampling in each block is the same; these blocks formed after sampling are vertically spliced as low-dimensional data blocks X l One column in , so that X l ; Step 103: Use high-dimensional dictionary D h , low-dimensional dictionary D l Composition of joint dictionary D c , for the joint dictionary D c Conduct training. The goals of the training process are as follows: in, is the Lagrange operator; α is the sparse representation vector; The trained joint dictionary D c The process of recovering omics data includes the following steps: For Y, we start from the upper left corner and take m×1 blocks without overlap, which are recorded as y. The low-dimensional data Y of the whole sequencing is the low-dimensional omics data measured by experiments. The specific omics features contained in Y are consistent with the training joint dictionary D. c The low-dimensional data X used l The type and size are the same; for each y, the following processing is performed: Step 201: Base D l and y, solve the optimization equation where α * is a sparse representation vector, is the Lagrange operator; Step 202: Generate a high-dimensional data texture block x=D h α * , where α * is the result corresponding to the sparse representation vector obtained by the optimization equation in step 201, and x is the restored decentralized high-dimensional data block; Step 203: Calculate the high-dimensional data block x * =x+mean(y), where mean means the mean; Step 204: generate each data block x * As a high-dimensional data matrix X * A column vector of X * .
6. The method for capturing and reconstructing omics data based on the joint dictionary learning of genomic data recovery and the JDLGR algorithm according to claim 5, characterized in that: High dimensional dataX h Based on the original omics data X, that is, the original omics data X is normalized and then decentralized, where the decentralized operation refers to X h The sample mean is subtracted from each sample.
7. The method for capturing and reconstructing omics data based on the joint dictionary learning (JDLGR) algorithm for genome data recovery according to claim 5 or 6, characterized in that: Joint DataX c , joint dictionary D c as follows: Among them, N and M are respectively h , X l The number of rows in the matrix.
8. The method for capturing and reconstructing omics data based on the joint dictionary learning (JDLGR) algorithm for genome data recovery according to claim 7, characterized in that: For the joint dictionary D c The training process includes the following steps: Step 121: Initialize D c and α; Among them, D c After randomly assigning any non-zero value in, decentralization and normalization operations are required; α is all assigned zero; Step 122: Change the existing D c Substitute into formula (5), optimize the formula with α as the variable, and obtain the trained α; Step 123: Substitute the existing α into equation (5) and use D c Optimize the formula for the variable and obtain the trained D c ; Step 124: Repeat steps 121 to 123 until the number of cycles reaches a preset value, and finally obtain the trained joint dictionary D c .
9. The method for capturing and reconstructing omics data based on the joint dictionary learning (JDLGR) algorithm for genome data recovery according to claim 8, characterized in that: In step 101, the high-dimensional data matrix X h In the process of clustering rows in each column, K-Means clustering algorithm is used for clustering.
Citation Information
Patent Citations
Robust sparse recovery STAP method and system based on alternating direction method of multipliers
CN106501785A
RGB image spectral information reconstruction method based on dictionary atom embedding
CN114240756A