Genome data analysis method

Through adaptive multidimensional spatial compression algorithm and gene variant-phenotype association analysis method, the problems of inaccurate processing and inefficient computing in traditional genomic data analysis are solved, and efficient data compression and disease marker screening are achieved.

CN120199337AActive Publication Date: 2025-06-24XIDIAN GRP HOSPITAL
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510677106.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Traditional genomic data storage and processing techniques have problems such as inaccurate processing, low computational efficiency and low analysis accuracy in the genomic data analysis process.

Method used

Adaptive multidimensional spatial compression algorithm is used to intelligently compress the genomic data, and correlation analysis is performed through the gene variant-phenotype association analysis method to screen out potential disease markers.

Benefits of technology

Through adaptive multidimensional spatial compression algorithm and variant differential coding mechanism, the storage requirements of genomic data are significantly reduced and the efficiency of data compression is improved. At the same time, the gene variant-phenotype association analysis method improves the accuracy of screening of disease markers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199337A_ABST
    Figure CN120199337A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a genome data analysis method. The method comprises the following steps: preprocessing genome data, and carrying out intelligent compression processing on the preprocessed genome data through a self-adaptive multi-dimensional space compression algorithm to obtain compressed genome data; performing gene variation detection and filtering on the compressed genome data to obtain varied genome data, and performing function annotation to obtain genome data with an annotation result; performing gene expression analysis on the genome data with the annotation result to obtain gene expression data; based on the genome data and the gene expression data with annotation results, correlation analysis is carried out through a gene variation-phenotype correlation analysis method, and potential disease markers are screened out. The technical problems that in the genome data analysis process of a traditional genome data storage and processing technology, genome data processing is not accurate, the calculation efficiency is low, and the analysis accuracy is low are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method for genomic data analysis. Background Art

[0002] With the rapid development of genomics research, the scale of genomic data has been continuously expanding. Especially in the fields of large-scale genomic sequencing and personalized medicine, the amount of generated data has been growing exponentially. Traditional data storage and processing technologies can no longer meet this growing data demand. Especially in the application of high-throughput sequencing technology, a large amount of raw genomic data needs to undergo strict variant detection, data compression, and storage, and these data must be efficiently queried and analyzed.

[0003] However, traditional genomic data storage and processing technologies have the following technical problems: during the genomic data analysis process, the genomic data is inaccurately processed, the computational efficiency is low, and the analysis accuracy rate is low. Summary of the Invention

[0004] The present invention provides a method for genomic data analysis to solve the technical problems that in the process of genomic data analysis by traditional genomic data storage and processing technologies, the genomic data is inaccurately processed, the computational efficiency is low, and the analysis accuracy rate is low.

[0005] A method for genomic data analysis according to the present invention specifically includes the following technical solutions: A method for genomic data analysis includes the following steps: S1. Preprocess the genomic data to obtain the preprocessed genomic data; perform intelligent compression processing on the preprocessed genomic data through an adaptive multi-dimensional space compression algorithm to obtain the compressed genomic data; perform gene variant detection and filtering on the compressed genomic data to obtain the variant genomic data, and perform functional annotation to obtain the genomic data with annotation results; S2. Perform gene expression analysis on the genomic data with annotation results to obtain gene expression data; based on the genomic data with annotation results and the gene expression data, perform correlation analysis through a gene variant-phenotype association analysis method to screen out potential disease markers.

[0006] Preferably, the S1 specifically includes: During the implementation of the adaptive multi-dimensional space compression algorithm, regard the preprocessed genomic data as a point cloud in a high-dimensional space, and apply high-dimensional space mapping to map the preprocessed genomic data to a low-dimensional space to obtain the mapped genomic data.

[0007] Preferably, the S1 specifically includes: Divide the mapped genomic data into data blocks, adopt an adaptive local compression strategy for the data blocks, dynamically select a compression method based on the local characteristics of the data blocks, and perform local compression.

[0008] Preferably, the S1 specifically includes: During the local compression process, define a local compression function, adjust the compression strategy through weighted summation and logarithmic transformation, and obtain the local compression cost of the data block; the specific formula of the local compression function is as follows: ; Where is the th data block the storage space required after local compression, representing the local compression cost of the data block; is the total number of feature dimensions considered during the compression process; is the weighting factor, indicating the weight of the th feature during the compression process; is the scaling factor, used to control the contribution degree of each variant feature to the compression effect; is the variant function of the mapped genomic data in the th data block on the th feature, calculated from specific features in the genomic sequence; Based on the local compression cost of the data block, perform compression processing to obtain the locally compressed data block.

[0009] Preferably, the S1 specifically includes: Based on the local compression cost of the data block, combined with the similarity metric between the locally compressed data blocks, perform global compression on the mapped genomic data to obtain the globally compressed genomic data, and the specific formula is: ; Where is the globally compressed genomic data; is the total number of mapped genomic data blocks; is the weight of the th locally compressed data block; is the th data block 's local compression cost; is the th locally compressed data block and the th locally compressed data block the similarity metric between them; is the adjustment parameter.

[0010] Preferably, the S1 specifically includes: Introduce a mutation difference coding mechanism, calculate the difference coding of mutation points based on the mutation points and adjacent mutation points in the globally compressed genomic data; introduce a global index for data position positioning; based on the local compression cost of data blocks, the globally compressed genomic data, the difference coding of mutation points, and the global index, obtain the compressed genomic data.

[0011] Preferably, the S2 specifically includes: Record the mutation data in the genomic data with annotation results as gene mutation data; the gene mutation-phenotype association analysis method, based on the gene mutation data and gene expression data, combined with disease phenotype data, dynamically adjusts the contributions of various types of data to the screening of disease markers through a non-linear weighted recursion method.

[0012] Preferably, the S2 specifically includes: During the implementation of the gene mutation-phenotype association analysis method, preprocess the gene mutation data, and perform standardization processing on the gene expression data and disease phenotype data; calculate the association degrees between the gene mutation data, gene expression data, and disease phenotype data to obtain the preliminary weighting factors of the gene mutation data and the preliminary weighting factors of the gene expression data, and dynamically update the weighting factors through a recursive method.

[0013] Preferably, the S2 specifically includes: Based on the weighting factors, model the relationship between the gene mutation data and the disease phenotype data through a non-linear regression model, and introduce a regularization term to optimize the non-linear regression model to obtain an optimized non-linear regression model; based on the optimized non-linear regression model, calculate the mutation-phenotype association degree of genes and sort them to select the genes most relevant to the disease phenotype as potential disease markers.

[0014] The beneficial effects of the technical solution of the present invention are: 1. Through the adaptive multi-dimensional space compression algorithm and the mutation difference coding mechanism, the genomic data is compressed into an extremely small storage space, which not only reduces the storage requirements but also effectively retains key information such as gene mutations and gene expression changes. By adopting an adaptive compression strategy, the compression method is dynamically adjusted according to the local characteristics of data blocks and the global redundancy relationship, making the data compression more efficient. In particular, it can effectively compress the mutation information and the redundant parts of gene sequences, improving the overall compression ratio.

[0015] 2. Conduct correlation analysis through the gene mutation-phenotype association analysis method. Based on gene mutation data, gene expression data, and disease phenotype data, dynamically adjust the contributions of various types of data to the screening of final disease markers in a non-linear weighted recursive manner to improve the accuracy of correlation analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a genomic data analysis method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0019] The following specifically describes the specific solution of a genomic data analysis method provided by the present invention with reference to the accompanying drawings.

[0020] Refer to the appendix Figure 1 , which shows a flowchart of a genomic data analysis method provided by an embodiment of the present invention. The method includes the following steps: S1. Preprocess the genomic data to obtain preprocessed genomic data; perform intelligent compression processing on the preprocessed genomic data through an adaptive multi-dimensional space compression algorithm to obtain compressed genomic data; perform gene mutation detection and filtering on the compressed genomic data to obtain mutated genomic data, and perform functional annotation to obtain genomic data with annotation results; Preprocess the genomic data (such as removing adapter sequences, trimming low-quality sequences, and removing contaminated sequences) to obtain preprocessed genomic data; the technical means used in the preprocessing process are all well-known to those skilled in the art and will not be elaborated here.

[0021] The preprocessed genomic data is intelligently compressed through an adaptive multi-dimensional space compression algorithm to obtain compressed genomic data. The adaptive multi-dimensional space compression algorithm uses an adaptive compression strategy to maximize data compression. In addition, considering the particularity of genomic data (such as the existence of variant information), differential coding is performed on the variant part of the preprocessed genomic data to further compress the variant data. At the same time, the adaptive multi-dimensional space compression algorithm supports efficient random access to the compressed data. The specific implementation process of the adaptive multi-dimensional space compression algorithm is as follows: First, the preprocessed genomic data is regarded as a point cloud in a high-dimensional space, denoted as , where represents the point cloud; is the point cloud in the th data point, represented as a -dimensional vector, containing metadata such as the actual information of the gene sequence, variant position, variant type, sequence repeatability, etc.; is the total number of data points in the point cloud. To obtain a compact compressed representation, high-dimensional space mapping is applied, and the preprocessed genomic data is mapped to a new low-dimensional space through existing space embedding methods (such as PCA or t-SNE). This mapping should not only consider the local structure (such as the similarity between similar sequences) but also maintain the global structure (such as the distribution of variant points in the genome). The mapping of each data point is represented as: ; where is the representation of the th data point in the low-dimensional space, that is, the mapped genomic data; is the dimension of the low-dimensional space; is the space embedding function obtained through space embedding methods such as PCA or t-SNE; is the value of the th data point in the th dimension of the low-dimensional space; Furthermore, after high-dimensional space mapping, there is a large amount of duplicate information and structural redundancy in the local area. Therefore, local compression needs to be performed on the mapped genomic data. The mapped genomic data is divided into data blocks using expert experience combined with specific scenarios. An adaptive local compression strategy is adopted for each data block, and the compression method is dynamically selected according to the local characteristics of the data block (such as repeatability, variant density, etc.) to perform local compression, obtaining the locally compressed data blocks; The goal of local compression is to make the representation of genomic data more compact through quantization and encoding. Therefore, by combining the various feature dimensions of the data blocks, a local compression function is defined, and the compression strategy is adjusted through weighted summation and logarithmic transformation to obtain the local compression cost of the data blocks. The specific formula of the local compression function is as follows: ; where, is the storage space required for the th data block after local compression, that is, the local compression cost of the data block, which is an index to measure the compression efficiency. The smaller the value, the better the compression effect; is the total number of feature dimensions considered in the compression process; is the weighting factor, indicating the weight of the th feature in the compression process, used to adjust the influence of each feature on the compression effect, and is determined according to the expert experience method; is the scaling factor, used to control the contribution degree of each variant feature to the compression effect, and determines the influence size of the variant feature in the compression process; is the mutation function of the mapped genomic data in the th data block on the th feature, which is calculated from specific features in the genomic sequence (such as gene mutations, insertions / deletions (InDels) mutations, changes in gene expression, or other numerical values related to biological features), such as , where, represents certain statistics related to the variant feature in the th data block , is the number of types of statistics related to the variant feature, is the weight coefficient related to the th type of statistics, is the feature function related to the th data block on the th type of statistics. For example, if it is the mutation frequency, the frequency of each mutation can be used as ; ; represents the change in gene expression in the th data block , is the coefficient related to the change in gene expression, is the gene expression difference degree, indicating the difference in the expression levels of genes under different conditions (such as comparing the expression differences under two conditions). For example, if the expression values of a gene under different conditions are respectively and , then . The purpose of the above local compression function is to adjust the compression intensity according to the density and importance of each feature. By introducing logarithmic transformation, the feature values with large numerical variations can be effectively compressed while retaining important structural information; Based on the local compression cost of data blocks, existing compression techniques such as Huffman coding, LZ77, arithmetic coding, etc. are selected according to the expert experience method for compression processing to obtain the locally compressed data blocks; Global compression is performed based on the locally compressed data blocks; the goal of global compression is to further reduce the storage amount of genomic data by analyzing the global redundancy relationship between the mapped genomic data. Considering the global repeatability and variant data existing in the mapped genomic data, combined with the similarity measurement between the locally compressed data blocks, the local compression costs of all data blocks are weighted and summed to obtain the globally compressed genomic data. The specific formula is: ; where, is the globally compressed genomic data; is the total number of mapped genomic data blocks; is the weight of the th locally compressed data block, indicating the importance of this locally compressed data block in global compression, , is the variability of the th locally compressed data block, reflecting the number of variations; is the size of the th locally compressed data block; , , are weighting coefficients used to control the influence degree of different factors on weight calculation and are determined according to the expert experience method; is the entropy value of the th locally compressed data block, indicating the information amount of the locally compressed data block; is the th locally compressed data block and the th locally compressed data block The similarity measurement between them indicates the similarity in structure or content between these two data blocks; is a regulation parameter used to control the influence of similarity measurement on the compression result.

[0022] For variations in genomic data such as single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels), a variation difference encoding mechanism is introduced. Based on the globally compressed genomic data, the data is further compressed. Especially at consecutive variation positions, redundant information can be effectively reduced. The specific content of the variation difference encoding mechanism is as follows: For the variation points in the globally compressed genomic data and adjacent variation points , the difference encoding of the variation point is calculated through the difference encoding formula. The difference encoding formula is as follows: ; where, is the difference encoding of the th variation point in the globally compressed genomic data; and are the positions of the current and the previous variation points in the globally compressed genomic data respectively; is the type encoding of the th variation point in the globally compressed genomic data (e.g., 1 for SNP, 2 for InDel, etc.); Through difference encoding, the variation data is compressed into a more concise representation, thereby further improving the compression efficiency. Especially for those variations with small changes, difference encoding can significantly reduce the storage requirements; Furthermore, to support efficient random access, an existing global index structure is adopted, which can quickly locate any position in the data, support efficient query and decompression operations, and denote the global index as . Through the global index, users can directly jump to a specific position in the compressed data for random access. The decompression process will only be performed on the required data blocks, rather than decompressing the entire data; Based on the local compression cost of data blocks, the globally compressed genomic data, the difference encoding of variation points, and the global index, the compressed genomic data is obtained, ; Perform gene variant detection and filtering on the compressed genomic data. Use efficient variant detection tools, such as GATK or SAMtools, to identify SNPs (single nucleotide polymorphisms) and Indels (insertion-deletion variations) in the compressed genomic data, and screen out potential false positives and low-quality variations through filtering criteria (such as depth, quality value, etc.) preset according to expert experience method, so as to obtain the variant genomic data. Further, perform functional annotation on the variant genomic data, that is, by docking with known gene libraries (such as RefSeq, Ensembl, etc.), understand the biological significance of each variant. Functional annotation not only helps to determine the potential impact of the variant on gene function, but also helps to mine disease-related variants, and finally obtain the genomic data with annotation results.

[0023] S2. Perform gene expression analysis on the genomic data with annotation results to obtain gene expression data; based on the genomic data with annotation results and gene expression data, perform correlation analysis through the gene variant-phenotype association analysis method to screen out potential disease markers.

[0024] Perform gene expression analysis on the genomic data with annotation results. The genomic data with annotation results contains variant information (such as SNPs, Indels, etc.), and has already marked the gene position, function information, etc.; use differential expression analysis tools (such as DESeq2, EdgeR, etc.) to analyze the genomic data with annotation results, calculate the expression differences of each gene under different conditions, and evaluate the statistical significance of its expression differences to obtain gene expression data, including the expression profiles and differentially expressed genes of the genomic data under different conditions.

[0025] Further, perform correlation analysis through the gene variant-phenotype association analysis method. The gene variant-phenotype association analysis method is based on the variant data (gene variant data) in the genomic data with annotation results, gene expression data, and disease phenotype data from existing databases, and dynamically adjusts the contributions of various types of data to the final disease marker screening in a non-linear weighted recursive manner. The specific implementation process of the gene variant-phenotype association analysis method is as follows: First, preprocess the gene variant data by using technical means well-known to those skilled in the art to remove low-frequency variants and variants with poor quality, and use the gene variant data for subsequent analysis, represented as the genotype of each variant site; use the gene expression data Standardize; Standardize the existing disease phenotype data to ensure that all disease phenotype data are under a unified framework, facilitating subsequent association analysis. For disease phenotype data, a binary classification (0 or 1) is used to represent the disease state. For example, "1" represents the disease group and "0" represents the control group; Furthermore, introduce a weighting factor to dynamically weight the contributions of gene mutation data and gene expression data to disease phenotype data, calculate the association degree between gene mutation data and disease phenotype data, and obtain the preliminary weighting factor of gene mutation data. The calculation formula for the preliminary weighting factor of gene mutation data is as follows: ; Among them, is the preliminary weighting factor of gene mutation data, reflecting the correlation between gene mutation data and disease phenotype data; is the total number of gene mutation data; is the mutation amount of the th gene in the gene mutation data; is the disease state of the th gene mutation data in the disease phenotype data; Similarly, calculate the preliminary weighting factor of gene expression data ; In order to make the weighting factor gradually tend to be stable, it is necessary to dynamically update the weighting factor recursively. The update formula for the weighting factor is as follows: ; ; Among them, and are the weighting factors for the current iteration (i.e., the th iteration); and are smoothing factors used to control the influence of the previous iteration result, obtained through experiments; and are the weighting factors for the previous iteration (i.e., the th iteration). Through the recursive update process, the weighting factor can be gradually optimized to make the weighted contributions of gene mutation data and gene expression data more accurate; Subsequently, use a non-linear regression model to model the relationship between gene mutation data and disease phenotype data to capture the possible non-linear association between the two, and obtain the mutation-phenotype association degree of the gene. The specific expression of the non-linear regression model is: ; Among them, is the mutation-phenotype association degree of the th gene; is the regression coefficient, determined through experiments; is the th gene variation in the gene expression data. By adding sine, logarithm, and square root functions, the non - linear regression model can more flexibly fit complex non - linear relationships, thereby more accurately capturing the implicit connection between gene variation data and disease phenotype data, making the non - linear regression model more robust and adaptable when processing data; To ensure the stable optimization of the non - linear regression model in multiple rounds of iteration, a regularization term is further introduced to prevent overfitting. For example, when minimizing the objective function, a regularization term of L2 norm is added; Finally, based on the optimized non - linear regression model above, recalculate the gene mutation - phenotype correlation degree, and sort it according to its value to obtain the sorted most relevant marker gene set, which is used as a candidate for disease marker screening, and select the gene most relevant to the disease phenotype as a potential disease marker for subsequent verification.

[0026] In summary, a genomic data analysis method is completed.

[0027] The order of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0028] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0029] The above - mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for genome data analysis, characterized in that, It includes the following steps: S1. Preprocess the genomic data to obtain the preprocessed genomic data; perform intelligent compression processing on the preprocessed genomic data through an adaptive multi-dimensional space compression algorithm to obtain the compressed genomic data; perform gene mutation detection and filtering on the compressed genomic data to obtain the mutated genomic data, and perform functional annotation to obtain the genomic data with annotation results; S2. Perform gene expression analysis on the genomic data with annotation results to obtain gene expression data; based on the genomic data with annotation results and the gene expression data, perform correlation analysis through a gene mutation-phenotype association analysis method to screen out potential disease markers.

2. The genomic data analysis method according to claim 1, wherein The S1 specifically includes: In the implementation process of the adaptive multi-dimensional space compression algorithm, regard the preprocessed genomic data as a point cloud in a high-dimensional space, and apply high-dimensional space mapping to map the preprocessed genomic data to a low-dimensional space to obtain the mapped genomic data.

3. A genomic data analysis method according to claim 2, characterized in that, The S1 specifically includes: Divide the mapped genomic data into data blocks, adopt an adaptive local compression strategy for the data blocks, dynamically select a compression method based on the local characteristics of the data blocks, and perform local compression.

4. A genomic data analysis method according to claim 3, wherein The S1 specifically includes: In the process of local compression, define a local compression function, adjust the compression strategy through weighted summation and logarithmic transformation to obtain the local compression cost of the data block; the specific formula of the local compression function is as follows: , Among them, is the th data block The storage space required after local compression, representing the local compression cost of the data block; is the total number of feature dimensions considered during the compression process; is the weighting factor, indicating the weight of the th feature during the compression process; is the scaling factor, used to control the contribution degree of each mutant feature to the compression effect; is the mutant function of the mapped genomic data in the th data block on the th feature, calculated from specific features in the genomic sequence; Based on the local compression cost of the data block, perform compression processing to obtain the locally compressed data block.

5. A genomic data analysis method according to claim 4, characterized in that The S1 specifically includes: Based on the local compression cost of the data block, combined with the similarity measure between the locally compressed data blocks, perform global compression on the mapped genomic data to obtain the globally compressed genomic data, and the specific formula is: , Among them, is the globally compressed genomic data; is the total number of mapped genomic data blocks; is the weight of the th locally compressed data block; is the local compression cost of the th locally compressed data block; and the th locally compressed data block; is the adjustment parameter.

6. The genomic data analysis method according to claim 5, characterized in that The S1 specifically includes: Introduce a mutation difference coding mechanism, calculate the difference coding of the mutation points based on the mutation points and adjacent mutation points in the globally compressed genomic data; introduce a global index for data position positioning; based on the local compression cost of the data block, the globally compressed genomic data, the difference coding of the mutation points, and the global index, obtain the compressed genomic data.

7. A genomic data analysis method according to claim 1, characterized in that The S2 specifically includes: Record the mutation data in the genomic data with annotation results as gene mutation data; the gene mutation-phenotype association analysis method, based on the gene mutation data and the gene expression data, combined with the disease phenotype data, dynamically adjusts the contribution of various types of data to the screening of disease markers in a non-linear weighted recursive manner.

8. A genomic data analysis method according to claim 7, wherein The S2 specifically includes: In the implementation process of the gene mutation-phenotype association analysis method, preprocess the gene mutation data, and standardize the gene expression data and the disease phenotype data; calculate the correlation degree between the gene mutation data, the gene expression data, and the disease phenotype data to obtain the preliminary weighted factors of the gene mutation data and the preliminary weighted factors of the gene expression data, and dynamically update the weighted factors through a recursive manner.

9. A method for genomic data analysis according to claim 8, characterized in that The S2 specifically includes: Based on the weighting factor, the relationship between gene mutation data and disease phenotype data is modeled through a non-linear regression model, and a regularization term is introduced to optimize the non-linear regression model, obtaining an optimized non-linear regression model; based on the optimized non-linear regression model, the mutation-phenotype association degree of genes is calculated and sorted, and the genes most relevant to the disease phenotype are selected as potential disease markers.

Citation Information

Patent Citations

  • Gene variation detection method based on pattern growth algorithm

    CN111243663A

  • Solid tumor mutant gene detection and analysis method and system

    CN115132276A

  • Gene sequencing data processing method and system

    CN119601083A

  • Parallel computing and data visualization analysis system for whole genome prediction

    CN119964655A

  • Methods and systems for storing genomic data in a file structure comprising an information metadata structure

    EP4226382A1