Unknown cellular network expression observation method and system based on multiple random encoding
By training the dictionary through multiple random encodings and blind compressed sensing algorithms, the high cost and time constraints of measuring gene expression levels in unknown cells are solved, achieving efficient and accurate observation of gene expression levels, which is suitable for observing gene network expression in unknown cells.
Patent Information
- Application Number
- CN202310117288.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-02-15
AI Technical Summary
Existing technologies for obtaining gene expression levels in unknown cells are costly, time-consuming, and cannot simultaneously measure gene expression profiles in a single well plate. In particular, additional discrimination and manipulation steps are required for unknown cells, leading to increased costs.
A method based on multiple random coding is adopted, which determines the random measurement matrix and gene name group through finite isochronous conditions, performs multiple PCR reactions, trains a dictionary by combining blind compressed sensing algorithm, calculates gene expression levels, and realizes sparse coding and dimensionality reduction data observation.
Without prior knowledge, this method efficiently and accurately observes gene expression levels in unknown cells, shortens measurement time to 2 hours, reduces costs, and achieves a Pearson correlation coefficient accuracy of 85%.
Smart Images

Figure CN116469460B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cell gene expression measurement technology in the field of life science, and particularly relates to an unknown cell network expression observation method and system based on multiple random encoding. BACKGROUND
[0002] With the progress of life science technology, the research on cell state based on gene expression can obtain more accurate data results. The common method for obtaining gene expression at present is to match all RNA sequences in cells through gene sequencing, and then count the same gene fragments to obtain the expression of genes through data processing. In addition, the commonly used methods are PCR, qPCR and digital PCR, which measure all RNA in cells to obtain relative quantitative or absolute quantitative gene expression.
[0003] In the process of obtaining gene expression by using gene sequencing (NGS, etc.), specific instruments are needed or samples are sent to companies for measurement. Moreover, the sequencing data obtained needs to be converted into corresponding gene expression by algorithm. In the above process, it will take several weeks to obtain gene expression, and the cost is high. When using common PCR and qPCR instruments, only 96-well plates or 384-well plates can be equipped. Traditional gene sequencing and PCR methods cannot be used in single well plates for experiments, that is, the gene data required by gene expression profile cannot be obtained in a PCR instrument at the same time, which greatly increases the time and cost of experiments.
[0004] In addition, there are some theoretical ways for gene dimension reduction observation at present, which can quickly obtain the expression of part of the genes in the cell, such as neural network dimension reduction method, which combines early analysis data and biological relationship through part of gene measurement to infer the remaining gene expression profile. However, the existing technology is more applied to known cells. For unknown cells, it is necessary to first distinguish the cells and then reduce the dimension, which needs to consume more time and operation steps, further leading to cost increase. SUMMARY
[0005] The purpose of the present application is to provide an unknown cell network expression observation method and system based on multiple random encoding, which directly measures unknown cells to improve the measurement efficiency.
[0006] To achieve the above purpose, the present application provides the following scheme:
[0007] An unknown cell network expression observation method based on multiple random encoding, comprising:
[0008] determining a plurality of random measurement matrices based on a limited isometry condition and a gene dimension of the unknown cell to be measured; each of the random measurement matrices comprises 0 values, positive values and negative values;
[0009] for each of the random measurement matrices, determining a plurality of gene name groups according to the random measurement matrix and the gene dimension of the unknown cell to be measured; the number of the gene name groups is the same as the number of rows of the random measurement matrix;
[0010] performing multiple PCR reactions on the unknown cell to be measured based on the plurality of gene name groups to determine a plurality of Ct data sets of the unknown cell to be measured; each of the Ct data sets comprises Ct positive values and Ct negative values;
[0011] determining gene observations according to relative quantity calculation of the Ct positive values and the Ct negative values; a plurality of the gene observations constitute a gene observation matrix;
[0012] performing dictionary training on a plurality of gene observation matrices corresponding to the plurality of random measurement matrices based on a blind compressed sensing algorithm to obtain an unknown cell gene dictionary;
[0013] calculating gene expression of the unknown cell to be measured according to the unknown cell gene dictionary and the plurality of gene observation matrices corresponding to the plurality of random measurement matrices.
[0014] Optionally, the determining of the random measurement matrix based on the limited isometry condition and the gene dimension of the unknown cell to be measured specifically comprises:
[0015] calculating a gene overall sampling rate and a gene single-row sampling rate according to a preset sparsity and the gene dimension of the unknown cell to be measured based on the limited isometry condition;
[0016] determining the random measurement matrix according to the gene overall sampling rate and the gene single-row sampling rate.
[0017] Optionally, the performing of the dictionary training on the plurality of gene observation matrices corresponding to the plurality of random measurement matrices based on the blind compressed sensing algorithm to obtain the unknown cell gene dictionary specifically comprises:
[0018] generating an initial module dictionary matrix and an initial module activity matrix randomly according to the plurality of gene observation matrices corresponding to the plurality of random measurement matrices;
[0019] performing iterative calculation on the initial module activity matrix corresponding to each of the gene observation matrices by using an orthogonal matching pursuit algorithm;
[0020] When the iteration number of the initial module activity matrix reaches a first preset iteration number, the initial module activity matrices corresponding to the gene observation matrices after iteration are connected to obtain a secondary activity matrix;
[0021] The initial module dictionary matrix corresponding to each gene observation matrix is iteratively calculated using a proximal gradient algorithm;
[0022] When the iteration number of the initial module dictionary matrix reaches a second preset iteration number, the initial module dictionary matrices corresponding to the gene observation matrices after iteration are connected to obtain a secondary dictionary matrix;
[0023] The secondary activity matrix is iteratively calculated using an orthogonal matching pursuit algorithm, and the secondary dictionary matrix is iteratively calculated using a proximal gradient algorithm;
[0024] When the iteration number of the secondary activity matrix and the iteration number of the secondary dictionary matrix both reach a third preset iteration number, the secondary activity matrix after iteration and the secondary dictionary matrix after iteration are respectively output; the secondary activity matrix after iteration and the secondary dictionary matrix after iteration constitute an unknown cell gene dictionary.
[0025] Optionally, each of the gene name groups includes a positive value name subgroup and a negative value name subgroup;
[0026] Based on the multiple gene name groups, multiple PCR reactions are performed on the measured unknown cell to determine multiple Ct data sets of the measured unknown cell, specifically including:
[0027] For each of the gene name groups, a first PCR primer group is determined according to the positive value name subgroup, and a second PCR primer group is determined according to the negative value name subgroup;
[0028] Based on the first PCR primer group and the second PCR primer group, PCR reactions are respectively performed on the measured unknown cell to determine the Ct data set of the measured unknown cell.
[0029] Optionally, a relative amount calculation is performed according to the Ct positive value and the Ct negative value to determine a gene observation value, specifically including:
[0030] The Ct positive value is subtracted from the Ct negative value to obtain a gene observation value.
[0031] Optionally, a gene expression amount of the measured unknown cell is calculated according to the unknown cell gene dictionary and the gene observation matrices corresponding to the multiple random measurement matrices, specifically including:
[0032] According to the gene observation matrix corresponding to the plurality of random measurement matrices, an unknown cell activity matrix is determined;
[0033] According to the unknown cell activity matrix and the unknown cell gene dictionary, the gene expression quantity of the measured unknown cell is calculated.
[0034] Optionally, according to the unknown cell activity matrix and the unknown cell gene dictionary, the gene expression quantity of the measured unknown cell is calculated, and specifically includes:
[0035] According to the formula X'=Uxs, the gene expression quantity of the unknown cell dictionary is calculated.
[0036] Wherein, X' represents the gene expression quantity of the unknown cell dictionary, U represents the unknown cell gene dictionary, and s represents the unknown cell activity matrix.
[0037] Optionally, the unknown cell network expression observation method further includes:
[0038] The gene data corresponding to the measured unknown cell is reproduced by using a compressed sensing algorithm to obtain actual reproduction data.
[0039] The gene data corresponding to the random measurement matrix is reproduced by using a compressed sensing algorithm to obtain predicted reproduction data.
[0040] According to the actual reproduction data and the predicted reproduction data, a Pearson correlation coefficient is calculated.
[0041] When the Pearson correlation coefficient is not within a preset deviation range, the gene single-row sampling rate is adjusted, and the random measurement matrix is re-determined according to the gene overall sampling rate and the adjusted gene single-row sampling rate.
[0042] To achieve the above purpose, the present application also provides the following technical solutions:
[0043] An unknown cell network expression observation system based on multiple random encodings, comprising:
[0044] A random matrix determination module is configured to determine a plurality of random measurement matrices based on a limited equidistance condition and a gene dimension of a measured unknown cell; each random measurement matrix includes 0 values, positive values and negative values.
[0045] A gene name determination module is configured to determine a plurality of gene name groups for each random measurement matrix according to the random measurement matrix and the gene dimension of the measured unknown cell; the number of gene name groups is the same as the number of rows of the random measurement matrix.
[0046] a PCR reaction module configured to perform multiplex PCR reactions on the unknown cell to be tested based on multiple sets of gene name groups to determine multiple Ct data sets of the unknown cell to be tested; each of the Ct data sets includes Ct positive values and Ct negative values;
[0047] an observation value determination module configured to determine gene observation values according to relative quantity calculation based on the Ct positive values and the Ct negative values; multiple gene observation values constitute a gene observation matrix;
[0048] a dictionary training module configured to perform dictionary training on gene observation matrices corresponding to the multiple random measurement matrices based on a blind compressed sensing algorithm to obtain an unknown cell gene dictionary;
[0049] a gene expression quantity calculation module configured to calculate gene expression quantities of the unknown cell to be tested according to the unknown cell gene dictionary and the gene observation matrices corresponding to the multiple random measurement matrices.
[0050] According to the specific embodiments of the present application, the following technical effects are provided:
[0051] The present application provides a kind of unknown cell network expression observation method and system based on multiple random encoding, based on the gene dimension of measured unknown cell and limited equidistant condition, determine multiple random measurement matrices, realize the sparse coding of gene.According to random measurement matrix and the gene dimension of measured unknown cell, determine multiple gene name groups, and then based on multiple gene name groups, the multiple PCR reactions of unknown cell to be tested are carried out to determine multiple Ct data sets of unknown cell to be tested;That is, the observation process of dimensionality reduction data is realized by PCR reaction process.According to Ct positive values and Ct negative values in Ct data set, calculate gene observation value;Based on blind compressed sensing algorithm, the gene observation matrix corresponding to the multiple random measurement matrices is trained to obtain an unknown cell gene dictionary, so that the gene network expression is reproduced by blind dictionary algorithm without any prior knowledge, and the gene dictionary is trained.Finally, according to the unknown cell gene dictionary and the gene observation matrix corresponding to the multiple random measurement matrices, the gene expression quantity of the unknown cell to be tested is calculated.The unknown cell state is observed, and the correlation between the cell reproduction data and the original data can reach 85% (Pearson correlation coefficient), the observation task of cell state can be completed within 2h, high-precision and high-efficiency observation calculation is realized. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only constitute some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0053] Figure 1 A flowchart of the unknown cell network expression observation method based on multiple random encoding of the present application;
[0054] Figure 2 A structural diagram of the unknown cell network expression observation system based on multiple random encoding of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0056] The present application provides an unknown cell network expression observation method and system based on multiple random encoding. The method encodes genes based on RIP conditions, uses PCR reaction process to realize the observation process of dimension reduction data, trains the dictionary without prior knowledge through multiple encoding measurement processes, uses blind compressed sensing algorithm to reconstruct the gene expression of cells, and finally achieves the requirements of reducing experimental cost and shortening measurement time.
[0057] In order to make the purpose, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0058] Embodiment one
[0059] As shown in the figure, the present embodiment provides an unknown cell network expression observation method based on multiple random encoding, which comprises: Figure 1 Step 100, based on the limited equidistant condition and the gene dimension of the unknown cell to be measured, a plurality of random measurement matrices are determined; each random measurement matrix comprises 0 value, positive value and negative value. Specifically, 4-7 random measurement matrices are determined, and generally, 5 random measurement matrices are sufficient for subsequent data processing.
[0060]
[0061] The elements in the matrix in the prior art common Gaussian random matrix are mostly decimals. In the digital channel or signal path, since there is a value itself, the Gaussian matrix can better perform weight operation; however, in biochemical reactions, such weight sampling method is limited by biochemical measurement method and measurement accuracy and cannot be actually used.
[0062] At the same time, in the current biochemical reaction, most of them still use specific primers for gene amplification experiments, and the common primers still have a small application range in the current research due to the problems of poor specificity and high mismatch probability. Even if the barcode experiment method is designed, the overall experimental process is more complex than the traditional PCR, and the cost is higher, and how to design and verify the specificity of the barcode fragment is a big problem. Therefore, if the elements in the matrix are partially sampled, and the sampling method is selected, it is more suitable for actual application in terms of time cost and economic cost.
[0063] Based on this, the application sets the measurement matrix corresponding to the gene dimension of the measured cell as A(m, n), m represents the number of rows of the matrix, and n represents the gene dimension of the measured cell; the sampled elements in the measurement matrix are defined as 1 / -1, or defined as , represents any weight, and s represents a weight matrix composed of a plurality of arbitrary weights. In order to facilitate calculation and reduce calculation difficulty, positive and negative values are generally set and determined by +1 and -1. In the process of finally iterating the gene expression amount of the measured unknown cell, the weight of the sampled element has little effect on the final result:
[0064] ;
[0065] Wherein, y represents the gene observation value, γ represents the penalty term, the value of γ can be selected by itself, || ||0 represents the L0 norm, represents the square of the L2 norm.
[0066] If there is an accurate weight matrix s, then the final result is , will not affect the effect of gene expression amount reproduction or reconstruction, that is, when the primer is excessive, it will not affect the measurement result, so that this measurement method has very strong anti-interference ability.
[0067] The traditional RIP condition is described as:
[0068] ;
[0069] ;
[0070] Wherein, x represents the measured vector, that is, the gene expression of the measured cell, A represents the measurement matrix.
[0071] Or can be described as: A and Ψ are not related (y=AΨs).
[0072] ;
[0073] Wherein, Ψ represents the dictionary, The kth row of the measurement matrix, The jth row of the dictionary.
[0074] Or described as:
[0075] ;
[0076] Wherein, A i The ith row of the measurement matrix, A j The jth row of the measurement matrix, || represents the modulus of the matrix.
[0077] Based on the above theory, the present application designs the measured basis in the measurement matrix to be positive and negative 1, the single-row sampled position is random, the single-row sampling rate of the gene is about 1%-2%, and the measurement gene entry, that is, the number of rows m of the measurement matrix is:
[0078] ;
[0079] Generally, the ratio is about 5%-10%, that is, m / n.
[0080] The following formula is used to verify whether the preliminary random matrix generated by the above method satisfies the RIP condition:
[0081] ;
[0082] In summary, step 100 specifically includes:
[0083] 1) Based on the limited equidistance condition, the overall gene sampling rate and the single-row gene sampling rate are calculated according to the preset sparsity and the gene dimension of the measured unknown cell.
[0084] 2) Determine the random measurement matrix according to the overall gene sampling rate and the single-row gene sampling rate.
[0085] In one specific application, the unknown cell network expression observation method further includes:
[0086] 1) The gene data corresponding to the measured unknown cell is reproduced by using the compressed sensing algorithm to obtain the actual reproduced data.
[0087] 2) The gene data corresponding to the random measurement matrix is reproduced by using the compressed sensing algorithm to obtain the predicted reproduced data.
[0088] 3) calculating a Pearson correlation coefficient according to the actual recurrence data and the predicted recurrence data; when the Pearson correlation coefficient is not within a preset deviation range, adjusting the gene single-line sampling rate, and re-determining a random measurement matrix according to the gene overall sampling rate and the adjusted gene single-line sampling rate; when the Pearson correlation coefficient is within the preset deviation range, then performing subsequent steps according to the random measurement matrix. Specifically, the preset deviation range is that the Pearson correlation coefficient is greater than 0.85.
[0089] Generally, under the conditions of the gene overall sampling rate and the gene single-line sampling rate, the obtained usable random measurement matrix is 99% of the total randomly generated matrix.
[0090] Step 200, for each random measurement matrix, obtaining a plurality of gene name groups corresponding to the random measurement matrix according to the random measurement matrix and the gene dimension of the unknown cell to be measured based on the principle of matrix multiplication; the number of gene name groups is the same as the number of rows of the random measurement matrix. Each gene name group includes a positive value name subgroup and a negative value name subgroup.
[0091] Step 300, based on a plurality of gene name groups, performing multiplex PCR reaction on the unknown cell to be measured to determine a plurality of Ct data sets of the unknown cell to be measured; each Ct data set includes a Ct positive value and a Ct negative value.
[0092] Step 300 specifically includes: for each gene name group, determining a first PCR primer group according to the positive value name subgroup and a second PCR primer group according to the negative value name subgroup; based on the first PCR primer group and the second PCR primer group, extracting RNA of the unknown cell to be measured, and performing PCR reaction on the unknown cell to be measured respectively to determine a Ct data set of the unknown cell to be measured.
[0093] In one specific example, the genes corresponding to each row in the designed random measurement matrix are subjected to multiplex PCR primer design, and the designed multiplex PCR primers are customized and assembled into the same 96-well plate, and the concentration of each primer contained in each 96-well multiplex primer is diluted to 10 nmol / ul, thereby completing the primer preparation process.
[0094] The Ct value of gene data can be obtained by a qPCR (PCR) reaction, using a qPCR instrument (PCR instrument) for the amplification process, generally 30 rounds or 40 rounds of amplification, and a marker gene is needed as a baseline in the production process, and the gapdh gene is usually used for normalization (the expression amount of the gapdh gene in the cell is basically fixed), and the Ct value corresponding to the gapdh gene is compared with the expression amount to obtain a standard line; in order to prevent errors in the operation process, two duplicate wells need to be prepared for each experimental well. If the Ct values of three experiments differ by no more than 0.2, the final Ct value is the average of the three wells.
[0095] In step 400, the relative amount is calculated according to the Ct positive value and the Ct negative value to determine the gene observation value; and a plurality of gene observation values constitute a gene observation matrix. Specifically, the Ct positive value is subtracted from the Ct negative value to obtain the gene observation value. And the Ct positive value and the Ct negative value are recorded as counts type data. Then, the gene observation value is recorded in the observation matrix. The Ct data of the counts type data obtained is calculated based on this, so that the data reproduction accuracy is better, and the measurement value of this method can be less when the same reproduction effect is achieved.
[0096] In addition, in order to reduce the reaction time, a plurality of parallel PCR or qPCR reaction methods are used to measure the sampled genes, so as to obtain 4-7 gene observation matrices, and the number of gene observation matrices is the same as the number of random matrix generations.
[0097] In step 500, based on the blind compressed sensing algorithm, the gene observation matrix corresponding to each of the plurality of random measurement matrices is subjected to dictionary training to obtain an unknown cell gene dictionary. Step 500 includes:
[0098] 1) An initial module dictionary matrix and an initial module activity matrix are randomly generated according to the gene observation matrix corresponding to each of the plurality of random measurement matrices.
[0099] 2) An orthogonal matching pursuit algorithm is used to iteratively calculate the initial module activity matrix corresponding to each of the gene observation matrices.
[0100] 3) When the iteration number of the initial module activity matrix reaches a first preset iteration number, the initial module activity matrices corresponding to the plurality of gene observation matrices after iteration are connected to obtain a secondary activity matrix.
[0101] 4) A proximal gradient algorithm is used to iteratively calculate the initial module dictionary matrix corresponding to each of the gene observation matrices.
[0102] 5) When the number of iterations of the initial module dictionary matrix reaches a second preset number of iterations, the initial module dictionary matrix corresponding to the gene observation matrix after iteration is connected to obtain a secondary dictionary matrix.
[0103] 6) The secondary activity matrix is iterated using an orthogonal matching pursuit algorithm, and the secondary dictionary matrix is iterated using a proximal gradient algorithm.
[0104] 7) When the number of iterations of the secondary activity matrix and the number of iterations of the secondary dictionary matrix both reach a third preset number of iterations, the secondary activity matrix after iteration and the secondary dictionary matrix after iteration are respectively output; the secondary activity matrix after iteration and the secondary dictionary matrix after iteration constitute an unknown cell gene dictionary.
[0105] In an actual application, in order to further improve the training accuracy, a clustering algorithm is added in the training process of the unknown cell gene dictionary, and the specific process is as follows:
[0106] A) Initialize the parameters of spectral clustering.
[0107] B) Obtain a preset spectral clustering, and determine the number of clusters L = max(5, min(20, (n / 50))). Each cluster includes multiple samples, and the samples in each cluster are random; the samples are the gene observation matrix corresponding to any random measurement matrix.
[0108] C) For each cluster, randomly generate an initial module dictionary matrix and an initial module activity matrix. The number of initial module dictionary matrices is dc = max(5, (|c| / 20)), where |c| represents the number of samples in the cluster.
[0109] D) Perform 5 iterations:
[0110] 1) If a sample is in the cluster, update the initial module activity matrix corresponding to the sample using an orthogonal matching pursuit algorithm (OMP), otherwise set it to 0.
[0111] 2) If a sample is in the cluster, update the initial module dictionary matrix corresponding to the sample using a proximal gradient algorithm (Proximal Gradient Descent, ProxGrad).
[0112] E) Connect the multiple initial module activity matrices after 5 iterations to obtain Connect the multiple initial module dictionary matrices after 5 iterations to obtain .
[0113] F) performing 5 iterations on the two connected matrices respectively, and performing iteration on the activity matrix by using an orthogonal matching pursuit algorithm, and performing iteration on the dictionary matrix by using a proximal gradient algorithm, and finally returning the activity matrix and the dictionary matrix after 5 iterations.
[0114] G) splicing the plurality of iteration activity matrices obtained after clustering, and splicing the plurality of iteration dictionary matrices obtained after clustering, so as to obtain a final activity matrix and a final dictionary matrix.
[0115] Step 600: calculating the gene expression quantity of the measured unknown cell according to the unknown cell gene dictionary and the gene observation matrix corresponding to the plurality of random measurement matrices.
[0116] Step 600 includes: determining an unknown cell activity matrix according to the gene observation matrix corresponding to the plurality of random measurement matrices; and calculating the gene expression quantity of the measured unknown cell according to the unknown cell activity matrix and the unknown cell gene dictionary.
[0117] The specific process of calculating the gene expression quantity of the measured unknown cell is: calculating the gene expression quantity of the unknown cell dictionary according to the formula X' = Uxs. Wherein, X' represents the gene expression quantity of the unknown cell dictionary, U represents the unknown cell gene dictionary, and s represents the unknown cell activity matrix.
[0118] In addition, through research, it is found that for the same type of disease, the interaction between genes has a strong correlation with the expression quantity, and the use of multiple encoding methods can obtain observation values through actual reactions, and then the gene network expression state of the unknown cell can be observed. Based on this, the present application provides a method that can be applied to any unknown cell (without any prior knowledge), which sparsely encodes genes based on the RIP condition, uses the PCR reaction process to realize the observation process of the reduced dimension data, trains the dictionary through multiple encoding measurement processes without prior knowledge, uses the blind compression sensing algorithm to reconstruct the gene expression quantity of the cell, and finally meets the requirements of reducing the experimental cost and shortening the measurement time. And when the PCR (qPCR) reaction is carried out in parallel, the overall operation can obtain the gene expression quantity of the current cell within 2h.
[0119] Embodiment Two
[0120] As Figure 2 shown, in order to perform the method corresponding to the above-mentioned embodiment one to realize the corresponding functions and technical effects, the present embodiment provides an unknown cell network expression observation system based on multiple random encoding, which comprises:
[0121] The random matrix determination module 101 is configured to determine a plurality of random measurement matrices based on a limited isometric condition and a gene dimension of a measured unknown cell; each of the random measurement matrices comprises 0 values, positive values and negative values.
[0122] The gene name determination module 201 is configured to determine, for each of the random measurement matrices, a plurality of gene name groups according to the random measurement matrix and the gene dimension of the measured unknown cell; the number of the gene name groups is the same as the number of rows of the random measurement matrix.
[0123] The PCR reaction module 301 is configured to perform multiplex PCR reaction on the measured unknown cell based on the plurality of gene name groups to determine a plurality of Ct data sets of the measured unknown cell; each of the Ct data sets comprises Ct positive values and Ct negative values.
[0124] The observation value determination module 401 is configured to determine gene observation values according to relative quantity calculation of the Ct positive values and the Ct negative values; the plurality of gene observation values constitute a gene observation matrix.
[0125] The dictionary training module 501 is configured to perform dictionary training on gene observation matrices corresponding to the plurality of random measurement matrices based on a blind compressed sensing algorithm to obtain an unknown cell gene dictionary.
[0126] The gene expression quantity calculation module 601 is configured to calculate gene expression quantities of the measured unknown cell according to the unknown cell gene dictionary and the gene observation matrices corresponding to the plurality of random measurement matrices.
[0127] In addition, the present application has the following advantages in addition to the prior art:
[0128] (1) The present application realizes the sparse coding measurement process of genes and the gene decoding process based on the blind dictionary iterative algorithm through the multiplex PCR reaction guided by multiple random coding. It does not need to distinguish the cells in advance or train the dictionary of the cell type, and can determine the measured gene combination without prior knowledge, obtain the relative quantitative observation values of genes by using the multiplex sparse coding measurement matrix, and finally observe the expression of the unknown gene network, establish the current cell state, and provide data guidance at the gene level for the treatment, prognosis and pathological analysis of major diseases such as cancer, chronic diseases and major infectious diseases.
[0129] (2) The present application combines the theoretical way of gene data dimension reduction with the actual biochemical reaction. The RIP condition is combined with the PCR reaction, the theoretical dimension reduction method is realized through the biochemical reaction, and the high-dimensional information of genes is stored through the multiplex PCR.
[0130] (3) The biochemical reaction used in the present application is a traditional gene amplification method (PCR amplification method), and no special instrument is needed for operation. Only a PCR instrument or a qPCR instrument or a digital PCR instrument is needed in the experiment process, thereby providing a stable and universal method for obtaining gene expression.
[0131] (4) The present application provides a short, accurate and low-cost method for obtaining cell gene expression. The present application obtains cell gene expression by sparse coding of genes and combining the dictionary training method of compressed sensing, the experimental time is 2h, about 1% of NGS, and the measurement accuracy is higher (Pearson correlation coefficient is about 85%, and Spearman correlation coefficient is more than 65%).
[0132] The principles and implementation modes of the present application are described by applying specific examples in the present article, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation modes and application ranges will be changed. In summary, the content of the present specification should not be understood as a limitation of the present application.
Claims
1. A method for observing unknown cellular network expression based on multiple random encoding, characterized in that, The unknown cell network expression observation method comprises: Determine a plurality of random measurement matrices based on the limited isometric condition and the gene dimension of the measured unknown cell; each of the random measurement matrices comprises 0 values, positive values and negative values; For each of the random measurement matrices, determine a plurality of gene name groups according to the random measurement matrix and the gene dimension of the measured unknown cell; the number of gene name groups is the same as the number of rows of the random measurement matrix; Based on the plurality of gene name groups, perform multiple PCR reactions on the measured unknown cell to determine a plurality of Ct data sets of the measured unknown cell; each of the Ct data sets comprises Ct positive values and Ct negative values; According to the relative amount calculation of the Ct positive values and the Ct negative values, determine gene observation values; a plurality of the gene observation values constitute a gene observation matrix; Based on a blind compressed sensing algorithm, perform dictionary training on the gene observation matrices corresponding to the plurality of random measurement matrices to obtain an unknown cell gene dictionary; specifically comprising: According to the gene observation matrices corresponding to the plurality of random measurement matrices, randomly generate an initial module dictionary matrix and an initial module activity matrix; Using an orthogonal matching pursuit algorithm, perform iterative calculation on the initial module activity matrix corresponding to each of the gene observation matrices; When the number of iterations of the initial module activity matrix reaches a first preset iteration number, connect the initial module activity matrices corresponding to the plurality of gene observation matrices after iteration to obtain a secondary activity matrix; Using a proximal gradient algorithm, perform iterative calculation on the initial module dictionary matrix corresponding to each of the gene observation matrices; When the number of iterations of the initial module dictionary matrix reaches a second preset iteration number, connect the initial module dictionary matrices corresponding to the gene observation matrices after iteration to obtain a secondary dictionary matrix; Iterate the secondary activity matrix using an orthogonal matching pursuit algorithm and iterate the secondary dictionary matrix using a proximal gradient algorithm; When the number of iterations of the secondary activity matrix and the number of iterations of the secondary dictionary matrix both reach a third preset iteration number, respectively output the secondary activity matrix after iteration and the secondary dictionary matrix after iteration; the secondary activity matrix after iteration and the secondary dictionary matrix after iteration constitute an unknown cell gene dictionary; According to the unknown cell gene dictionary and the gene observation matrices corresponding to the plurality of random measurement matrices, calculate the gene expression amount of the measured unknown cell.
2. The method of claim 1, wherein, The determination of the random measurement matrix based on the limited isometric condition and the gene dimension of the measured unknown cell specifically comprises: Based on the limited isometric condition, calculate the gene overall sampling rate and the gene single-row sampling rate according to the preset sparsity and the gene dimension of the measured unknown cell; Determine the random measurement matrix according to the gene overall sampling rate and the gene single-row sampling rate.
3. The method of claim 1, wherein, In the training process of the unknown cell gene dictionary, a clustering algorithm is added, and the specific process is as follows: Initialize the parameters of spectral clustering; Obtaining a preset spectrum clustering, and determining a number of clusters; each cluster includes a plurality of samples, and the samples in each cluster are random; the samples are gene observation matrices corresponding to any random measurement matrix; For each cluster, an initial module dictionary matrix and an initial module activity matrix are randomly generated; the number of initial module dictionary matrices is dc =max(5, (|c| / 20)), where |c| represents the number of samples in the cluster; 5 iterations are performed: if a sample is in the cluster, the initial module activity matrix corresponding to the sample is updated using an orthogonal matching pursuit algorithm OMP, otherwise it is set to 0; if a sample is in the cluster, the initial module dictionary matrix corresponding to the sample is updated using a proximal gradient algorithm; For the connection of the plurality of initial module activity matrices after 5 iterations, the following is obtained For the connection of the plurality of initial module dictionary matrices after 5 iterations, the following is obtained ; The two matrices connected above are iterated for 5 times respectively; and the activity matrix is iterated using an orthogonal matching pursuit algorithm, and the dictionary matrix is iterated using a proximal gradient algorithm, and finally the activity matrix and the dictionary matrix after 5 iterations are returned; The plurality of iterated activity matrices obtained after clustering are spliced, and the plurality of iterated dictionary matrices obtained after clustering are spliced, so as to obtain a final activity matrix and a final dictionary matrix.
4. The method of claim 1, wherein, Each of the gene name groups includes a positive value name group and a negative value name group; Based on a plurality of the gene name groups, a multiplex PCR reaction is performed on the measured unknown cell to determine a plurality of Ct data sets of the measured unknown cell, specifically including: For each of the gene name groups, a first PCR primer group is determined according to the positive value name group, and a second PCR primer group is determined according to the negative value name group; Based on the first PCR primer group and the second PCR primer group, a PCR reaction is performed on the measured unknown cell respectively to determine a Ct data set of the measured unknown cell.
5. The method of claim 1, wherein, According to the Ct positive value and the Ct negative value, a relative amount calculation is performed to determine a gene observation value, specifically including: The Ct positive value is subtracted from the Ct negative value to obtain a gene observation value.
6. The method of claim 1, wherein, According to the unknown cell gene dictionary and the gene observation matrices corresponding to a plurality of the random measurement matrices, the gene expression amount of the measured unknown cell is calculated, specifically including: According to the gene observation matrices corresponding to a plurality of the random measurement matrices, an unknown cell activity matrix is determined; According to the unknown cell activity matrix and the unknown cell gene dictionary, the gene expression amount of the measured unknown cell is calculated.
7. The method of claim 6, wherein the method further comprises: According to the unknown cell activity matrix and the unknown cell gene dictionary, the gene expression amount of the measured unknown cell is calculated, specifically including: According to the formula X' = Ux s, the gene expression amount of the unknown cell dictionary is calculated; Wherein, X' represents the gene expression amount of the unknown cell dictionary, U represents the unknown cell gene dictionary, and s represents the unknown cell activity matrix.
8. The method of claim 2, wherein, The unknown cell network expression observation method further includes: The gene data corresponding to the measured unknown cell is reproduced using a compressed sensing algorithm to obtain actual reproduction data; The compressed sensing algorithm is used to reproduce gene data corresponding to the random measurement matrix to obtain predicted reproduced data. A Pearson correlation coefficient is calculated according to the actual reproduced data and the predicted reproduced data. When the Pearson correlation coefficient is not within a preset deviation range, the gene single-line sampling rate is adjusted, and the random measurement matrix is re-determined according to the gene overall sampling rate and the adjusted gene single-line sampling rate.
9. A multiple random coding based unknown cellular network expression observation system, characterized in that, The unknown cell network expression observation system comprises: A random matrix determination module is configured to determine a plurality of random measurement matrices based on a limited equidistant condition and a gene dimension of a measured unknown cell; each random measurement matrix comprises 0 values, positive values and negative values; A gene name determination module is configured to determine, for each random measurement matrix, a plurality of gene name groups according to the random measurement matrix and the gene dimension of the measured unknown cell; the number of gene name groups is the same as the number of rows of the random measurement matrix; A PCR reaction module is configured to perform multiplex PCR reactions on the measured unknown cell based on the plurality of gene name groups to determine a plurality of Ct data sets of the measured unknown cell; each Ct data set comprises Ct positive values and Ct negative values; An observation value determination module is configured to determine gene observation values by performing relative quantity calculation according to the Ct positive values and the Ct negative values; a plurality of gene observation values constitute a gene observation matrix; A dictionary training module is configured to perform dictionary training on a plurality of gene observation matrices corresponding to the plurality of random measurement matrices based on a blind compressed sensing algorithm to obtain an unknown cell gene dictionary; specifically comprising: An initial module dictionary matrix and an initial module activity matrix are randomly generated according to the plurality of gene observation matrices corresponding to the plurality of random measurement matrices; An orthogonal matching pursuit algorithm is used to iteratively calculate the initial module activity matrix corresponding to each gene observation matrix; When the number of iterations of the initial module activity matrix reaches a first preset number of iterations, the initial module activity matrices corresponding to the plurality of gene observation matrices after iteration are connected to obtain a secondary activity matrix; A proximal gradient algorithm is used to iteratively calculate the initial module dictionary matrix corresponding to each gene observation matrix; When the number of iterations of the initial module dictionary matrix reaches a second preset number of iterations, the initial module dictionary matrices corresponding to the gene observation matrices after iteration are connected to obtain a secondary dictionary matrix; The secondary activity matrix is iterated using the orthogonal matching pursuit algorithm, and the secondary dictionary matrix is iterated using the proximal gradient algorithm; When the number of iterations of the secondary activity matrix and the number of iterations of the secondary dictionary matrix both reach a third preset number of iterations, the secondary activity matrix after iteration and the secondary dictionary matrix after iteration are respectively output; the secondary activity matrix after iteration and the secondary dictionary matrix after iteration constitute an unknown cell gene dictionary; A gene expression quantity calculation module is configured to calculate gene expression quantities of the measured unknown cell according to the unknown cell gene dictionary and the gene observation matrices corresponding to the plurality of random measurement matrices.
Citation Information
Patent Citations
Gene expression filling data acquisition method and device and storage medium
CN113782093A
Cell communication network identification method and device, equipment and storage medium
CN114722988A