A GPU-based method for fast optimization of gene co-expression network calculations
By adopting a fast GPU-based optimization operation method in gene coexpression network analysis, the problem of low processing efficiency of multidimensional gene coexpression matrix in the prior art is solved, and more efficient calculations and more accurate results are achieved.
Patent Information
- Application Number
- CN202111404214.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-24
AI Technical Summary
When processing multidimensional gene coexpression matrix, the dimensionality reduction efficiency is low and difficult, resulting in large calculation and low efficiency.
Using the fast optimization operation gene co-expression network method based on GPU, a weighted correlation matrix coefficient iteration module is established by establishing a data reading module, a correlation matrix coefficient iteration module and a level numerical iteration module, and using the multi-core computing power provided by the GPU, the correlation matrix coefficients are processed and iteratively updated to build a weighted correlation network.
Through GPU acceleration, the period and waiting time of gene coexpression network analysis are significantly reduced, the computational efficiency is improved, and more biologically significant results can be obtained.
Smart Images

Figure CN113920003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technology, and in particular to a method for quickly optimizing and computing a gene co-expression network based on a GPU. Background Art
[0002] Since gene co-expression data usually measure the expression values of thousands of genes in dozens of samples, for such a large matrix, the amount of computation required to process gene data is extremely large and inefficient, so many bioinformatics tools have been developed for data analysis and annotation. Using network biology methods to analyze and mine high-throughput gene expression data has become an important research direction in bioinformatics. The gene co-expression matrix is a multidimensional matrix, and the computational optimization of multidimensional matrices and complex networks has always been a hot topic for researchers at home and abroad.
[0003] Rudolph et al. studied the directed WS small-world network model and gave a new analytical framework and algebraic expression for describing its adjacency matrix, which can accurately describe the asymmetry index and clustering coefficient of the WS small-world network. Galiceanu et al. studied polymers with scale-free network characteristics and analyzed the influence of stiffness parameters and network size on their dynamic characteristics. In terms of new network construction, Peng et al. proposed a multidimensionally growing deterministic small-world network model and gave an accurate analytical expression for calculating the characteristic path length of the network model. Emmerich et al. studied the structural and functional characteristics of embedded scale-free networks and compared the similarities and differences between embedded scale-free networks and non-embedded scale-free networks. In 2017, Sang Aijun et al. proposed a motion vector estimation method based on the multidimensional vector matrix transformation domain, introducing the multidimensional vector matrix theory, transformation theory, and the theoretical derivation of the energy concentration plane formed by the moving target in the transformation domain. The introduction of this method has played a certain role in optimizing data processing and multidimensional matrix operations. In the same year, Han Jianxun from China University of Petroleum used API, OpenMP, MPI and PPL, four multi-core parallel algorithms, to compare and analyze a fast calculation method suitable for multi-dimensional matrices. Zhao Lifeng and Huang Yiwen from the School of Science at Nanjing University of Posts and Telecommunications proposed a K-shortest path algorithm based on matrix operations. The shortest path problem is a classic problem in complex networks, and its solution algorithms are endless, each with its own advantages and disadvantages.
[0004] At present, most mainstream methods for processing multidimensional matrices are to reduce the dimensionality of multidimensional matrices, thereby using Dijkstra algorithm, Ford algorithm, Floyd algorithm, etc. to process matrices. However, these methods are not satisfactory due to the low efficiency and difficulty of dimensionality reduction. Summary of the invention
[0005] The object of the present invention is to solve the defects existing in the prior art, and a method for rapidly optimizing the operation of a gene co-expression network based on GPU is proposed.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for rapidly optimizing the operation of a gene co-expression network based on GPU, comprising the following steps:
[0008] Step 1: Establish a data reading module, a correlation matrix coefficient iteration module, and a rank value iteration module, and use GPU to provide computing power to construct a gene co-expression network;
[0009] Step 2: Use the data reading module to read the relevant data of the input file;
[0010] Step 3: Select a corresponding digital model through the correlation matrix coefficient iteration module to process and iteratively update the coefficients of the correlation matrix for the data read above;
[0011] Step 4: According to the user's selection, directly calculate, calculate after matrix iteration, calculate simultaneously by iteration, and perform conventional calculation on the data obtained in Step 2 and Step 3 through the rank value iteration module and the correlation matrix coefficient iteration module respectively;
[0012] Step 5: Obtain the arrays and rank values calculated in Step 4;
[0013] Step 6: Use the quicksort method to sort and output the results.
[0014] Further, for Step 2 above, the step process of the data reading module reading relevant data is successively: reading the total number of genes, reading the first row of the correlation matrix, dividing the names of different genes, reading the correlation coefficients, and saving them in the format of a two-dimensional array.
[0015] Further, for Step 3 above, the types of digital models are: three models, namely, the same reduction model, the segmentation change model, and the same increase model.
[0016] Further, for Step 3 above, the process of the correlation matrix coefficient iteration module for iterative update:
[0017] Read the two-dimensional array storing the coefficients;
[0018] Calculate the value of each element after one mathematical calculation;
[0019] Judge whether the difference between two iterations is greater than 0.001;
[0020] According to the judgment result, select to update the content of the two-dimensional array or recalculate.
[0021] Furthermore, the process for calculating the array in the above step 4 through the level value iteration module includes the following steps:
[0022] Initialize the transition probability matrix;
[0023] Assign values according to coefficients;
[0024] Initialize the structure array;
[0025] Calculate the LR value and determine whether the difference in the calculation result or the number of iterations meets the stopping condition;
[0026] If it does not meet the requirements, the LR value will be calculated again. If it meets the requirements, the data will be saved in the result data.
[0027] Furthermore, the correlation matrix coefficient iteration module processes the gene co-expression data through a GPU.
[0028] Furthermore, in the above step 4, direct calculation is performed: that is, the coefficients of the correlation matrix are not iterated, but the level values of the connection relationship are directly iterated;
[0029] Calculation after matrix iteration: that is, after iterating the coefficients of the correlation matrix, the level values are iterated according to the new connection relationship;
[0030] Simultaneous iterative calculation: that is, the rank value is calculated once for each iteration of the coefficient of the correlation matrix until the iteration of the correlation matrix is completed;
[0031] General calculation: used to calculate the network graph of one-way edges.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] GPU is used to process gene co-expression networks to obtain gene nodes with high connectivity, which is convenient for subsequent research on gene expression networks. The coefficients of the correlation matrix are processed through three mathematical models to achieve three functions. The running time of the three functions is gradually increasing, but the final results are becoming more and more accurate.
[0034] The iterative idea is introduced to use old values to recursively infer new values, so that it can adapt to gene co-expression data under different conditions, and then construct a weighted association network and obtain more biologically meaningful results;
[0035] Since GPU has multiple cores, using it in gene co-expression network analysis to process gene co-expression matrices can greatly reduce the analysis cycle and shorten the analysis waiting time. The required hardware and software environment is relatively simple and easy to install. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0037] Figure 1 A flowchart of the method for rapidly optimizing the computation of gene co-expression networks based on GPU proposed by the present invention;
[0038] Figure 2 This is a logic diagram of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0039] Figure 3 This is a flow logic diagram of a data reading module of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0040] Figure 4 It is a flow logic diagram of the correlation matrix coefficient iteration module of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0041] Figure 5 It is a flow logic diagram of the hierarchical numerical iteration module of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0042] Figure 6 It is a function curve after the iteration convergence of the same reduction model in the correlation matrix coefficient iteration module of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0043] Figure 7 It is a function curve after iterative convergence of the separation change model in the correlation matrix coefficient iteration module of the GPU-based fast optimization operation gene co-expression network method proposed in the present invention;
[0044] Figure 8 It is a function curve after the iteration convergence of the same increase model in the correlation matrix coefficient iteration module of the GPU-based fast optimization operation gene co-expression network method proposed by the present invention;
[0045] Fig. 9 This is a comparison chart of the processing speed of the GPU-based gene co-expression network fast optimization operation method proposed in the present invention under different iteration times of the same matrix;
[0046] Fig.10 This is a comparison chart of the processing speed of the GPU-based fast optimization operation gene co-expression network method proposed in the present invention under different scale matrix environments. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0048] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0049] Example
[0050] Reference Figure 1-2 , a method for fast optimization and operation of gene co-expression network based on GPU, comprising the following steps:
[0051] Step S101: establishing a data reading module, a correlation matrix coefficient iteration module and a level value iteration module, and using a GPU to provide computing power to construct a gene co-expression network;
[0052] Step S103: using a data reading module to read relevant data of the input file;
[0053] Step S105: the read data is processed and iteratively updated by selecting a corresponding digital model through a correlation matrix coefficient iteration module to process the coefficients of the correlation matrix;
[0054] Step S107: the data obtained in step S103 and step S105 are respectively directly calculated, calculated after matrix iteration, calculated simultaneously and calculated conventionally by the level value iteration module and the correlation matrix coefficient iteration module according to the user's choice;
[0055] Step S109: Obtain the array and level values calculated in step S107;
[0056] Step S111: Use quick sorting to sort and output the results.
[0057] Reference Figure 3 In a specific embodiment of the present application, in the above step S103, the steps of the data reading module reading the relevant data are as follows: reading the total number of genes, reading the first row of the correlation matrix, dividing the names of different genes, reading the correlation coefficients and saving them in the format of a two-dimensional array.
[0058] Reference Figure 6-8In the specific embodiment of the present application, in the above step S105, the types of digital models are: a same reduction model, a separation change model and a same enlargement model, wherein:
[0059] The iterative equation used in the reduced model is:
[0060] The iterative equation used in the partition change model is:
[0061] The iterative equation used in the same growth model is:
[0062] Reference Figure 4 , in a specific embodiment of the present application, in the above step S105, the correlation matrix coefficient iteration module performs iterative updating process:
[0063] Read the two-dimensional array containing the coefficients;
[0064] Calculate the value of each element after a mathematical calculation;
[0065] Determine whether the difference between two iterations is greater than 0.001;
[0066] Based on the judgment result, choose to update the contents of the two-dimensional array or recalculate.
[0067] Specifically, if the difference between two iterations is greater than 0.001, the value of each element after a mathematical calculation is recalculated, and if the difference between two iterations is less than 0.001, the content of the two-dimensional array is updated.
[0068] Reference Figure 5 In a specific embodiment of the present application, the process for calculating the array through the level value iteration module in the above step S107 includes the following steps:
[0069] Initialize the transition probability matrix;
[0070] Assign values according to coefficients;
[0071] Initialize the structure array;
[0072] Calculate the LR value and determine whether the difference in the calculation result or the number of iterations meets the stopping condition;
[0073] If it does not meet the requirements, the LR value will be calculated again. If it meets the requirements, the data will be saved in the result data.
[0074] In a specific embodiment of the present application, the correlation matrix coefficient iteration module processes the gene co-expression data through a GPU.
[0075] Specifically, since GPU has multiple cores, using it in gene co-expression network analysis and processing gene co-expression matrix can greatly reduce the analysis cycle and shorten the analysis waiting time.
[0076] In a specific embodiment of the present application, in the above step S107, direct calculation is performed: that is, the coefficients of the correlation matrix are not iterated, but the level values of the connection relationship are directly iterated;
[0077] Calculation after matrix iteration: that is, after iterating the coefficients of the correlation matrix, the level values are iterated according to the new connection relationship;
[0078] Simultaneous iterative calculation: that is, the rank value is calculated once for each iteration of the coefficient of the correlation matrix until the iteration of the correlation matrix is completed;
[0079] General calculation: used to calculate the network graph of one-way edges.
[0080] In order to better understand the technical solution of the present invention, the following is a further description in conjunction with experiments and drawings.
[0081] experiment
[0082] Object: The object of comparison is a common CPU-based gene co-expression network.
[0083] Content: Comparison of the processing speed of the gene co-expression network based on GPU optimization calculations.
[0084] The processing speed comparison of the object and the gene co-expression network based on GPU optimization calculation under different iteration times of the same matrix is shown in Fig. 9 .
[0085] For comparison of the processing speed of the object and the GPU-based fast optimization operation of the gene co-expression network in different matrix environments, see Fig.10 .
[0086] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A GPU-based method for fast optimization of gene co-expression network calculations. It is characterized in that The following steps are involved: Step 1: Establish a data reading module, a correlation matrix coefficient iteration module, and a level value iteration module, and use GPU to provide computing power to build a gene co-expression network; Step 2: Use the data reading module to read the relevant data of the input file; Step 3: The read data is processed and iteratively updated by selecting a corresponding digital model through a correlation matrix coefficient iteration module to process the coefficients of the correlation matrix; Step 4: The data obtained in step 2 and step 3 are respectively directly calculated, calculated after matrix iteration, calculated simultaneously and calculated conventionally through the level value iteration module and the correlation matrix coefficient iteration module according to the user's choice; Step 5: Get the array and level values calculated in step 4; Step 6: Use quick sorting to sort and output the results; In the above step 2, the steps of the data reading module reading the relevant data are as follows: reading the total number of genes, reading the first row of the correlation matrix, dividing the names of different genes, reading the correlation coefficients and saving them in the format of a two-dimensional array; In the above step 3, the types of digital models are: same reduction model, separation change model and same enlargement model; The process of iterative updating of the correlation matrix coefficient iteration module in step 3 above is as follows: Read the two-dimensional array containing the coefficients; Calculate the value of each element after a mathematical calculation; Determine whether the difference between two iterations is greater than 0.001; According to the judgment result, choose to update the content of the two-dimensional array or recalculate; The process for calculating the array in the above step 4 through the level value iteration module includes the following steps: Initialize the transition probability matrix; Assign values according to coefficients; Initialize the structure array; Calculate the LR value and determine whether the difference in the calculation result or the number of iterations meets the stopping condition; If it does not meet the requirements, the LR value will be calculated again. If it meets the requirements, the data will be saved in the result data. The correlation matrix coefficient iteration module processes the gene co-expression data through the GPU; Used in the above step 4, direct calculation: that is, directly iterating the level value of its connection relationship without iterating the coefficient of the correlation matrix; Calculation after matrix iteration: that is, after iterating the coefficients of the correlation matrix, the level values are iterated according to the new connection relationship; Simultaneous iterative calculation: that is, the rank value is calculated once for each iteration of the coefficient of the correlation matrix until the iteration of the correlation matrix is completed; Conventional computing: used to calculate network graphs with one-way edges; There are three types of digital models: same reduction model, separation change model and same increase model, among which: The iterative equation used in the reduced model is: The iterative equation used in the partition change model is: The iterative equation used in the same growth model is: If the difference between two iterations is greater than 0.001, the value of each element after a mathematical calculation is recalculated. If the difference between two iterations is less than 0.001, the contents of the two-dimensional array are updated.
Citation Information
Patent Citations
Self-adaptive gene regulation grid construction method and device
CN113160890A