Method and apparatus for compressing genomic data using a gene mutation dictionary

By creating a gene mutation dictionary, genomic data is divided into biologically significant unit partitions, and mutant types are counted and numbered, which solves the problem of insufficient storage space for genomic data, and realizes efficient data compression and low-cost storage, supporting individualized medical care.

CN114930724BActive Publication Date: 2025-07-04MGI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980102589.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-31
Publication Date
2025-07-04
Estimated Expiration
2039-12-31

AI Technical Summary

Technical Problem

The existing genomic data storage methods cannot effectively compress the storage space, resulting in high storage costs and complex decompression processes, which cannot meet the needs of individualized medical care.

Method used

By creating a gene mutation dictionary, genomic data are divided into biologically significant unit partitions, mutant types of each unit partition are counted and numbered, and the encoding of the mutant type is used instead of the original data storage to generate a gene mutation dictionary.

Benefits of technology

It achieves about 25,000 times compression of genomic data, significantly reduces storage costs, and can quickly retrieve and restore gene information, supporting individualized medical care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930724B_ABST
    Figure CN114930724B_ABST
Patent Text Reader

Abstract

A method and device for compressing genomic data using a gene mutation dictionary. Specifically, it involves a method for creating a gene mutation dictionary, including: obtaining genomic sequence data of multiple individuals of a species and the reference genomic data of the species; aligning the genomic sequence data of multiple individuals to the reference genomic data respectively to obtain the mutation results of the genomic sequence data of each individual relative to the reference genomic data; dividing the genome of the species into several unit partitions with biological significance; according to the mutation results, statistically analyzing the mutant situations of each unit partition respectively to generate all mutant types of each unit partition in multiple individuals, and numbering the mutant types to obtain a gene mutation dictionary. The present invention solves the problem of genomic data compression, significantly reducing its storage capacity and greatly reducing the cost of storing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of genome data storage, and in particular to a method and device for compressing genome data using a gene mutation dictionary. Background Art

[0002] The human genome has been mapped, but the genomes of most individuals have not yet been mapped. Once the individual genomes are mapped, there will be a problem: there is simply not enough storage space in the world's computer systems to store this data. The volume of genomic data alone is one of the key obstacles to the true global popularization of genomics. The raw data size of the sequenced human genome may reach 200-300 gigabytes, while the analyzed genome data may occupy a full terabyte. If you want to build a whole genome library, the data volume alone will be a very thorny problem.

[0003] Storing human genome data is not just about storage. The key is not only to understand how genes usually interact with each other, but also to be able to apply genome data to individuals to achieve personalized medicine. For example, after drawing a genetic map that can be used by individuals, personalized medicine may usher in a prosperous era. If the existing genome data is stored in a coded way, you only need to look at your own genetic code to understand your genetic information, which will benefit the general public.

[0004] At present, the method to reduce data storage space is still through data compression, such as gzip and other data compression tools. However, the compression rate of data compression tools for general programs or texts is only about 60%. The compressed data still takes up a lot of storage space. In the face of the rapid growth of sequencing data, it is still a drop in the bucket.

[0005] Using data compression to save storage space on genomic data can reduce storage and hide data in a lossless and transparent way. However, if you want to view your own genetic information, you still need to decompress the data. It will be very troublesome to understand your personal genome information through repeated compression and decompression. Moreover, decompression still requires large storage to support it, and the cost has not been significantly reduced. Summary of the invention

[0006] The purpose of the present invention is to provide a method for creating a gene mutation dictionary and a method and device for compressing genome data using the gene mutation dictionary, which solves the problem of genome data compression, significantly reduces its storage volume, and greatly reduces the cost of storing data.

[0007] According to a first aspect of the present invention, the present invention provides a method for creating a gene mutation dictionary, comprising:

[0008] Obtain genomic sequence data of multiple individuals of a species and reference genomic data of the species;

[0009] Align the genomic sequence data of the multiple individuals to the reference genomic data respectively to obtain the mutation results of the genomic sequence data of each individual relative to the reference genomic data;

[0010] Divide the genome of the species into several biologically meaningful unit partitions;

[0011] According to the above mutation results, statistically analyze the mutant situations of each unit partition respectively, generate all mutant types of each unit partition among the multiple individuals, and number the mutant types to obtain the gene mutation dictionary, which includes multiple mutant types corresponding to each unit partition and their numbers.

[0012] In a preferred embodiment, the species is a human; the multiple individuals are more than 1000 human bodies.

[0013] In a preferred embodiment, the unit partitions include coding regions, non-coding regions, and genes.

[0014] In a preferred embodiment, the number of the biologically meaningful unit partitions is several thousand to several tens of thousands.

[0015] In a preferred embodiment, the number of the biologically meaningful unit partitions is 60,000, with an allowable error range of plus or minus 10%.

[0016] In a preferred embodiment, the unit partitions include 30,000 gene coding regions and 30,000 non-coding regions, and their numbers respectively allow an error range of plus or minus 10%.

[0017] In a preferred embodiment, the step of statistically analyzing the mutant situations of each unit partition respectively, generating mutant types and numbering the mutant types includes:

[0018] For each of the above unit partitions, sequentially take the mutation results of the multiple individuals as mutant types for numbering and counting in the order of individuals, where the counted number is the number of individuals supporting the mutant type, and if the mutation result of a later individual is the same as that of any previous individual, then adopt the mutant type and its number of the previous individual and add 1 to the count of that mutant type; if the mutation result of a later individual is different from the mutation results of all previous individuals, then a new mutant type is added to the dictionary, that is, regarded as a new mutant type and given a number and count, and finally all mutant types of each of the above unit partitions and the number and count of each mutant type are obtained.

[0019] In a preferred embodiment, the above method further includes:

[0020] Sorting each mutant type in descending order according to its count number, and renumbering the above mutant types in sequence.

[0021] According to a second aspect of the present invention, the present invention provides a device for creating a gene mutation dictionary, including:

[0022] A data acquisition unit, configured to acquire genomic sequence data of multiple individuals of a species and reference genomic data of the species;

[0023] A data comparison unit, configured to respectively compare the genomic sequence data of the above multiple individuals to the above reference genomic data to obtain a mutation result of the genomic sequence data of each individual relative to the reference genomic data;

[0024] A partition division unit, configured to divide the genome of the above species into several unit partitions with biological significance;

[0025] A dictionary generation unit, configured to respectively count the mutant situations of each unit partition according to the above mutation results, generate all mutant types of each unit partition in the above multiple individuals, and number the above mutant types to obtain the above gene mutation dictionary, where the gene mutation dictionary includes multiple mutant types corresponding to each unit partition and their numbers.

[0026] According to a third aspect of the present invention, the present invention provides a computer-readable storage medium, which includes a program, and the above program can be executed by a processor to implement the method as in the first aspect.

[0027] According to a fourth aspect of the present invention, the present invention provides a method for compressing genomic data using a gene mutation dictionary, including:

[0028] Acquiring genomic sequencing data of an individual, where the genomic sequencing data includes several unit partitions with biological significance;

[0029] Comparing the above genomic sequencing data to the gene mutation dictionary generated by the method in the first aspect to obtain mutant types and their numbers that are consistent with the mutant situations of each unit partition of the above individual;

[0030] Storing the numbers of the mutant types of each unit partition of the above individual instead of the actually measured mutations.

[0031] According to a fifth aspect of the present invention, the present invention provides a device for compressing genomic data using a gene mutation dictionary, including:

[0032] A data acquisition unit for acquiring genomic sequencing data of an individual, where the genomic sequencing data includes multiple biologically significant unit partitions;

[0033] A data comparison unit for comparing the genomic sequencing data to a gene mutation dictionary generated by the method of the first aspect to obtain the mutant types and their numbers that are consistent with the mutant conditions of each unit partition of the individual;

[0034] A data storage unit for storing the numbers of the mutant types of each unit partition of the individual instead of the actually measured mutations.

[0035] According to the sixth aspect of the present invention, the present invention provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the method of the fourth aspect.

[0036] According to the seventh aspect of the present invention, the present invention provides a method for restoring and utilizing genomic data compressed by a gene mutation dictionary, including:

[0037] Obtaining compressed genomic data, which is data compressed by using the gene mutation dictionary generated by the method of the first aspect and includes the numbers of mutant types of multiple unit partitions in the gene mutation dictionary;

[0038] Finding the numbers of mutant types of each unit partition in the gene mutation dictionary from the compressed genomic data;

[0039] Extracting the corresponding mutant types and their mutation results at each base site in the gene mutation dictionary according to the numbers of the mutant types.

[0040] According to the eighth aspect of the present invention, the present invention provides a device for restoring and utilizing genomic data compressed by a gene mutation dictionary, including:

[0041] A data acquisition unit for acquiring compressed genomic data, which is data compressed by using the gene mutation dictionary generated by the method of the first aspect and includes the numbers of mutant types of multiple unit partitions in the gene mutation dictionary;

[0042] A number acquisition unit for finding the numbers of mutant types of each unit partition in the gene mutation dictionary from the compressed genomic data;

[0043] A mutant extraction unit for extracting the corresponding mutant types and their mutation results at each base site in the gene mutation dictionary according to the numbers of the mutant types.

[0044] According to the ninth aspect of the present invention, the present invention provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the method as in the seventh aspect.

[0045] The method of the present invention solves the problem of compressing genomic data, and can achieve a compression ratio of about 25,000 times, significantly reducing its storage volume and greatly reducing the cost of storing data. With the rapid development of sequencing data, the storage volume of genomic data has almost reached an incredible level. By effectively encoding genomic data using the method of the present invention, the entire genome can be compressed to a size of about 120 kb. This is following the significant reduction in the cost of genome sequencing, and the storage cost of genomic data has also been greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of the method for creating a gene mutation dictionary in an embodiment of the present invention.

[0047] Figure 2 It is a statistical chart of the number of biotypes of all human genes in an embodiment of the present invention.

[0048] Figure 3 It is a block diagram of the structure of the device for creating a gene mutation dictionary in an embodiment of the present invention.

[0049] Figure 4 It is a flowchart of the method for compressing genomic data using a gene mutation dictionary in an embodiment of the present invention.

[0050] Figure 5 It is a block diagram of the structure of the device for compressing genomic data using a gene mutation dictionary in an embodiment of the present invention.

[0051] Figure 6 It is a flowchart of the method for restoring genomic data compressed using a gene mutation dictionary in an embodiment of the present invention.

[0052] Figure 7 It is a block diagram of the structure of the device for restoring genomic data compressed using a gene mutation dictionary in an embodiment of the present invention.

[0053] Figure 8 It is a list of mutant types of the Chr21_KRTAP12-4 gene in the 1000 Genomes Project in an embodiment of the present invention.

[0054] Figure 9 It is a schematic diagram of a typical gene mutation dictionary in an embodiment of the present invention.

[0055] Figure 10 It is a schematic diagram of the storage process of a human gene mutation dictionary in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific embodiments. In the following embodiments, many details are described to enable a better understanding of the present invention. However, those skilled in the art can easily recognize that some of these features can be omitted in different situations, or can be replaced by other materials or methods.

[0057] In addition, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in a manner obvious to those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment, and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0058] In the era of the continuous explosion of individual genomic data, many companies and institutions first think of solving the data compression problem through computers and high-performance storage. Since the data itself has not changed, only the data format has changed, the reduction in storage capacity is not obvious.

[0059] The method of the present invention is different from the usual form of data compression to reduce the storage space. Instead, it uses bioinformatics to compress the sequence information of nearly 3,000,000,000 bases of the human genome (about 3G) to only about 120KB, achieving a data compression of nearly 25,000 times without loss of genetic information.

[0060] The method of the present invention mainly includes three technical solutions for realizing the compression of genomic data:

[0061] (1) Create multiple types (Patterns) with biological significance, encode them with numbers, and generate a gene mutation dictionary. In some embodiments, there are approximately 30,000 genes in the human whole genome. After the sequencing of the 1000 Genomes Project, the 10,000 Genomes Project, or even the 1,000,000 Genomes Project is completed, all the mutants after determination are used to create a mutant dictionary for each gene as a unit. In some embodiments, for non-coding regions such as introns, every 100kb sequence can be used as a unit (for a 3G human genome, there will be 30,000 100kb units), and all the mutants after sequencing are also used to create a mutant dictionary for each gene as a unit. In this way, the human whole genome can be divided into 60,000 units, including 30,000 coding gene regions and 30,000 non-coding regions. In other embodiments, the number of units can also be appropriately increased or decreased according to the mutation types (patterns) of the genome. According to the types of mutants in each unit, 1 byte (2 8 = 256 types), or 2 bytes (2 16= 65,536 (65536 species), or 3 bytes (2 24 = 16,777,216 (16,777,216 species) to create a gene mutation dictionary, and then store this mutant dictionary in a large database.

[0062] (2) After the individual genome is sequenced, only the type (pattern) code of each unit is stored according to the gene mutation dictionary, instead of storing the actually measured mutations. If each type (pattern) is represented by 2 bytes, the whole genome of each person only requires a storage capacity of 60,000 * 2 = 120 kb. Thus, a data compression of nearly 25,000 times is achieved.

[0063] (3) Obtain gene mutation information. Using the stored gene type (pattern) code and the gene mutation dictionary, the results of gene sequencing can also be fully restored.

[0064] Such as Figure 1 As shown, the embodiment of the present invention provides a method for creating a gene mutation dictionary, including the following steps:

[0065] S101: Data acquisition

[0066] Obtain the genomic sequence data of multiple individuals of a species and the reference genomic data of this species.

[0067] The method for creating a gene mutation dictionary of the present invention is applicable to any species, especially species with the need to store individual genomic sequencing data, mainly the storage of human genomic data. The method of the present invention requires the genomic sequence data of multiple individuals, such as the genomic sequence data of more than 1000 people. Generally speaking, the more individual samples there are, the more comprehensive the mutant types in the created gene mutation dictionary will be. Generally speaking, in some embodiments, at least the genomic data information of 1000 people should be included, and it is best if there is genomic data information of more than one million people. Correspondingly, the reference genomic data is human reference genomic data.

[0068] S102: Data comparison

[0069] Align the genomic sequence data of the above multiple individuals to the above reference genomic data respectively to obtain the mutation results of each individual's genomic sequence data relative to the reference genomic data.

[0070] In the embodiment of the present invention, there are various methods to achieve genomic sequence alignment, such as sequence alignment software such as BWA or SOAP2. Through sequence alignment, the mutation results of each individual's genomic sequence data relative to the reference genomic data are obtained, and the mutation results are reflected in the differences between the bases of the individual's genome and the bases of the reference genome at each base position.

[0071] S103: Partition Division

[0072] Divide the genomes of the above-mentioned species into several unit partitions with biological significance.

[0073] In the embodiments of the present invention, the unit partitions with biological significance may be coding regions, non-coding regions, genes, etc. on the genome. In other embodiments, the unit partitions with biological significance may also be genes of all biotypes of humans, etc.

[0074] In the embodiments of the present invention, the number of unit partitions with biological significance is from thousands to tens of thousands. For example, in some embodiments, the number of unit partitions with biological significance is 60,000, and a tolerance range of plus or minus 10% is allowed. For example, there are 54,000 to 66,000 unit partitions with biological significance. In a preferred embodiment of the present invention, the unit partitions with biological significance include 30,000 gene coding regions and 30,000 non-coding regions, and their numbers are allowed a tolerance range of plus or minus 10% respectively.

[0075] For example, in one embodiment, according to the unit partitions with biological significance such as coding regions, non-coding regions, genes, etc., the entire genomic data of humans (about 3G) is divided into about 60,000 unit partitions, among which 30,000 unit partitions are gene coding regions and the other 30,000 unit partitions are non-coding regions, and each non-coding region is about 100kb in size.

[0076] S104: Lexicon Generation

[0077] According to the above mutation results, count the mutant situations of each unit partition respectively, generate all mutant types of each unit partition in the above-mentioned multiple individuals, and number the above mutant types to obtain the gene mutation lexicon, which includes multiple mutant types corresponding to each unit partition and their numbers.

[0078] In the embodiments of the present invention, a mutant type refers to a specific combination formed by all mutation sites on a specific unit partition with biological significance, including the positions of the mutation sites on the unit partition and base changes. There is at least one difference in mutation sites between any two mutant types. That is to say, as long as there is at least one difference in mutation sites, the two mutant types are different mutant types.

[0079] In the embodiments of the present invention, numbering the mutant types is to code the mutant types in sequence. A typical but non-limiting coding method is to combine the number of the unit partition with biological significance and the number of its corresponding mutant type as the code of a specific mutant type.

[0080] In one embodiment of the present invention, the step of separately counting the mutant situations of each unit partition, generating mutant types and numbering the mutant types includes the following steps: for each of the above unit partitions, sequentially numbering and counting the mutation results of the above-mentioned multiple individuals as mutant types in the order of individuals, where the counted number is the number of individuals supporting the mutant type, and if the mutation result of a later individual is the same as that of any previous individual, then the mutant type and its number of the previous individual are adopted and the count of this mutant type is incremented by 1; if the mutation result of a later individual is different from the mutation results of all previous individuals, then it is regarded as a new mutant type and a number and count are given, and finally all mutant types of each of the above unit partitions and the number and count quantity of each mutant type are obtained. Further, the above method further includes: sorting each mutant type from largest to smallest according to its count quantity, and renumbering the above mutant types in order.

[0081] For example, in one embodiment of the present invention, according to the mutation results analyzed in the previous step, the mutant situations of approximately 60,000 unit partitions of the human genome are separately counted. The statistical method for each unit partition (gene or intron) is as follows: in individual -1 (person -1), the aligned mutation result (the mutation combination of all mutation sites) is denoted as type -1 (pattern -1), and the count is denoted as 1; for the next individual -2 (person -2), if the alignment to this unit partition is the same as the mutation result aligned in individual -1 (person -1), then it is the same type, that is, still type -1 (pattern -1), and the count is 1 + 1 = 2, and so on; if individual -2 (person -2) aligns to this unit partition and its mutation result is different from that of individual -1 (person -1), then it is denoted as a new type (pattern), that is, type -2 (pattern -2), and the count is denoted as 1, until all individual samples in the database are counted. For the purpose of accelerating the subsequent retrieval speed of the gene mutation dictionary, in the preferred method of the present invention, each type (pattern) is sorted from largest to smallest according to its count, and the one with the highest count is redefined as type -1 (pattern -1), the one with the second highest count is defined as type -2 (pattern -2), and so on.

[0082] In the embodiments of the present invention, the division of unit partitions with biological significance can also be based on other biological criteria, such as counting the number of biotypes of all human genes. For example, Figure 2Shows the statistical count of the number of biotypes of all human genes.

[0083] Corresponding to the method for creating a gene mutation dictionary of the present invention, the present invention also provides a device for creating a gene mutation dictionary, as Figure 3 shown, including: a data acquisition unit 301, configured to acquire genomic sequence data of multiple individuals of a species and the reference genomic data of the species; a data alignment unit 302, configured to align the genomic sequence data of the multiple individuals to the reference genomic data respectively, to obtain the mutation results of the genomic sequence data of each individual relative to the reference genomic data; a partition division unit 303, configured to divide the genome of the species into several unit partitions with biological significance; a dictionary generation unit 304, configured to respectively count the mutant situations of each unit partition according to the mutation results, generate all mutant types of each unit partition in the multiple individuals, and number the mutant types, to obtain the gene mutation dictionary, where the gene mutation dictionary includes multiple mutant types corresponding to each unit partition and their numbers.

[0084] Those skilled in the art can understand that all or part of the functions of the above-mentioned implementation manners can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above-mentioned implementation manners are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium may include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc., and the above functions are implemented by a computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, the above-mentioned all or part of the functions can be implemented. In addition, when all or part of the functions in the above-mentioned implementation manners are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, and is saved to the memory of the local device by downloading or copying, or the system of the local device is updated in version. When the processor executes the program in the memory, the above-mentioned all or part of the functions in the implementation manners can be implemented.

[0085] Therefore, in an embodiment of the present invention, a computer-readable storage medium is provided, including a program that can be executed by a processor to implement the method for creating a gene mutation dictionary of the present invention.

[0086] Based on the gene mutation dictionary of the present invention, the present invention also provides a method for compressing genomic data using the gene mutation dictionary, as Figure 4 shown, including the following steps:

[0087] S401: Data acquisition

[0088] Obtain the genomic sequencing data of an individual, where the genomic sequencing data includes multiple biologically significant unit partitions.

[0089] S402: Data alignment

[0090] Align the above genomic sequencing data to the gene mutation dictionary generated by the method of creating a gene mutation dictionary of the present invention to obtain the mutant types and their numbers that are consistent with the mutant conditions of each unit partition of the above individual.

[0091] S403: Data storage

[0092] Store the numbers of the mutant types of each unit partition of the above individual instead of the actually measured mutations.

[0093] The present invention also provides a device for compressing genomic data using a gene mutation dictionary, as Figure 5 shown, including: a data acquisition unit 501 for obtaining the genomic sequencing data of an individual, where the genomic sequencing data includes multiple biologically significant unit partitions; a data alignment unit 502 for aligning the above genomic sequencing data to the gene mutation dictionary generated by the method of creating a gene mutation dictionary of the present invention to obtain the mutant types and their numbers that are consistent with the mutant conditions of each unit partition of the above individual; a data storage unit 503 for storing the numbers of the mutant types of each unit partition of the above individual instead of the actually measured mutations.

[0094] The present invention also provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the method for compressing genomic data using a gene mutation dictionary of the present invention.

[0095] The present invention also provides a method for restoring genomic data compressed using a gene mutation dictionary, as Figure 6 shown, including the following steps:

[0096] S601: Data acquisition

[0097] Obtain the compressed genomic data, which is the data compressed using the gene mutation dictionary generated by the method of creating a gene mutation dictionary of the present invention, and it includes the numbers of the mutant types of multiple unit partitions in the above gene mutation dictionary.

[0098] S602: Number acquisition

[0099] Find the numbers of the mutant types of each unit partition in the above gene mutation dictionary from the above compressed genomic data.

[0100] S603: Mutation extraction

[0101] Extract the corresponding mutant types in the gene mutation dictionary and their mutation results at each base site according to the numbers of the above mutant types.

[0102] The present invention also provides a device for restoring and utilizing genomic data compressed by a gene mutation dictionary, as Figure 7 shown, including: a data acquisition unit 701, configured to acquire the compressed genomic data, which is the data compressed by using the gene mutation dictionary generated by the method for creating a gene mutation dictionary of the present invention, and includes the numbers of mutant types in the gene mutation dictionary of multiple unit partitions; a number acquisition unit 702, configured to find the numbers of mutant types in the gene mutation dictionary of each unit partition from the above compressed genomic data; and a mutation extraction unit 703, configured to extract the corresponding mutant types in the gene mutation dictionary and their mutation results at each base site according to the numbers of the above mutant types.

[0103] According to the ninth aspect of the present invention, the present invention provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the method for restoring and utilizing genomic data compressed by a gene mutation dictionary of the present invention.

[0104] The technical solutions and effects of the present invention are described in detail below through embodiments. It should be understood that the embodiments are only exemplary and should not be construed as limiting the present invention.

[0105] Example 1: Creation of a gene mutation dictionary

[0106] (1) Collect a gene database. In the embodiment of the present invention, a thousand-person gene database is collected and gene data is collected.

[0107] (2) Align the collected gene data to the human reference genome.

[0108] (3) Divide the collected whole-genome data into approximately 60,000 units according to biologically significant unit partitions such as coding regions, non-coding regions, and genes, among which 30,000 units are gene coding regions and the other 30,000 units are non-coding regions, and each non-coding region is about 100 kb in size.

[0109] (4) According to the mutation results analyzed in the previous step, the mutant situations of approximately 60,000 unit partitions of the human genome are respectively counted. The statistical method for each unit partition (gene or intron) is as follows: In individual - 1 (person - 1), the aligned mutation results (the mutation combinations of all mutation sites) are recorded as type - 1 (pattern - 1), and the number (count) is recorded as 1; for the next individual - 2 (person - 2), when aligning to this unit partition, if it is consistent with the mutation results aligned in individual - 1 (person - 1), it is the same type, that is, still type - 1 (pattern - 1), and the number (count) is 1 + 1 = 2, and so on; if individual - 2 (person - 2) aligns to this unit partition and its mutation results are inconsistent with those of individual - 1 (person - 1), then it is recorded as a new type (pattern), that is, type - 2 (pattern - 2), and the number (count) is recorded as 1, until all individual samples in the database are counted. For the purpose of accelerating the subsequent retrieval speed of the gene mutation dictionary, in the preferred method of the present invention, each type (pattern) is sorted in descending order according to its number (count), and the one with the highest number (count) is re - defined as type - 1 (pattern - 1), the one with the second - highest number (count) is defined as type - 2 (pattern - 2), and so on.

[0110] In this embodiment, the variations of all coding regions (approximately 30,000 genes) of the whole genome information of 2504 people (data sourced from the 1000 Genomes Project) are analyzed. The variation analysis is performed on each gene of these 2504 people. According to the coding method in step (4) of this embodiment, each gene can finally obtain a maximum of 2504 mutant types (assuming that the variations of each person are different).

[0111] Taking the KRTAP12 - 4 gene as an example, the construction and statistics of the mutant dictionary are carried out on the VCF files of 2504 samples of the 1000 Genomes Project. The statistical results show that among the 2504 samples, 831 samples with the highest mutation frequency in the KRTAP12 - 4 gene have the same mutation type, that is, a homozygous mutation where the A base at position 460742022 on chromosome 21 mutates to the G base. The second - highest - frequency mutation has 318 (12.7%) samples. 16 samples have unique types. It is worth pointing out that there are only 33 mutant types (pattern) for the 2504 samples, and its data volume is further compressed by 75 times compared with the already greatly compressed VCF file. Figure 8 The schematic diagram of the results of the mutant type list (pattern list) of the Chr21_KRTAP12 - 4 gene in the 1000 Genomes Project in the embodiment of the present invention is shown.

[0112] Table 1 shows the statistical results of mutant types of 30 genes randomly selected from nearly 30,000 genes of 2,504 human samples in the Thousand-person Gene Database. Obviously, the number of mutant types of mutants is significantly lower than the number of samples. This indicates that for a certain mutant type, the mutant results of multiple individuals are exactly the same. If we store the mutant types instead of storing the mutants of each sample, the storage space required for this storage will be greatly reduced.

[0113] Statistical Results of Mutant Types of 30 Genes Randomly Selected in the Thousand-person Gene Database

[0114]

[0115]

[0116] In this embodiment, the genomic data of all people in the database are compared, and the mutant types are classified separately according to each unit partition. Each gene of the entire population is encoded in the above form and aggregated together to obtain a complete gene mutation dictionary. A typical gene mutation dictionary is similar to Figure 9 in form.

[0117] In this embodiment, taking the storage of 2 bytes (byte) for each unit partition as an example, 2 bytes (byte) can store 65,536 mutant types (pattern), which can store all biologically meaningful mutants. If the number of mutant types (pattern) in a certain unit partition is less than 65,536, the remaining mutant types (pattern) can be used for newly discovered types in the future. If the number of mutant types (pattern) in a certain unit exceeds 65,536, the mutant types (pattern) of this unit can be stored with 3 bytes, and the number of mutant types (pattern) at this time is 16,777,216.

[0118] Embodiment 2: Storage of Individual Human Genomic Data

[0119] After the individual human genome is sequenced, it is compared with the gene mutation dictionary according to each unit partition, as Figure 10 shown. According to the gene mutation dictionary, only the mutant type (pattern) encoding of each unit is stored, rather than storing the actually measured mutants.

[0120] Figure 10Shows a storage process of a human gene mutation dictionary. Taking the mutant types (patterns) of the first 100bp of genes OR4F29, SAMD11, LEMD2, GNG3, and KRTAP29-1 as examples, an individual is analyzed for these five genes. In gene 1, the alignment to pattern-7 is exactly the same; in gene 2, the alignment to pattern-2 is exactly the same; in gene 3, the alignment to pattern-5 is exactly the same; in gene 4, the alignment to pattern-2 is exactly the same; in gene 5, the alignment to pattern-4 is exactly the same. Then the mutation information of the 5 genes of this individual can be stored as the encoding of 5 mutant types. Storing each gene mutation information of a person as the encoding of a corresponding mutant type will save a large amount of storage.

[0121] In this embodiment, a typical storage form of a human mutant sequence is as follows: 1.7 2.1 3.5 4.2 5.4

[0127] … 60000.35

[0129] Among them, the number before the decimal point is the index number (index) of the unit partition, and the number after the decimal point is the index number (index) of the corresponding mutant type.

[0130] If the mutant type of each gene is represented by 2 bytes, and the whole genome of each person is divided into 60,000 genes, only 60000 * 2 = 120kb of storage is required, achieving nearly 25000-fold data compression.

[0131] Example 3: Restoration of gene mutant information

[0132] In this embodiment, using the encoding of the stored gene mutant type and the gene mutation dictionary, the result of gene sequencing can also be completely restored.

[0133] Since the region of each unit partition (such as a gene) in the whole genome corresponding to the gene mutation dictionary is fixed. The mutant corresponding to the mutant type of each unit partition is also fixed. Therefore, based on the stored gene unit partition index number and the corresponding mutant type index number, the real mutation can be very directly deduced.

[0134] Example 4: Application in Precision Medicine

[0135] Whole-genome sequencing of humans will become one of the physical examination items for humans. The storage of mutants of humans based on the gene mutation dictionary will be similar to Figure 9 the form of data. When a person is just born, a high-depth perfect genome sequencing is performed to establish his / her gene sequence reference. Then, a whole-genome sequencing is performed during the physical examination every year or at an appropriate time. Most of the time, there are no gene mutations. When there are mutations, the mutant types are significantly different from those at the time of his / her birth, so it is very easy to identify gene mutations.

[0136] As shown in Table 2, a storage example of the whole-genome sequencing physical examination items of an individual human is shown. Among them, the first column is the index number of the gene unit, and the second column and above are the number of index numbers of the mutant types corresponding to the index number of a certain gene unit.

[0137] Table 2 Storage Example of Mutants in Whole-Genome Sequencing of Humans in Precision Medicine

[0138]

[0139]

[0140] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the technical field to which the present invention belongs, according to the idea of the present invention, several simple deductions, deformations or substitutions can also be made.

Claims

1. A method for creating a gene mutation dictionary, characterized in that, The method includes: Obtaining genomic sequence data of multiple individuals of a species and reference genomic data of the species; Aligning the genomic sequence data of the multiple individuals to the reference genomic data respectively to obtain the mutation results of the genomic sequence data of each individual relative to the reference genomic data; Dividing the genome of the species into several biologically meaningful unit partitions; According to the mutation results, statistically analyzing the mutant situations of each unit partition respectively, generating all mutant types of each unit partition among the multiple individuals, and numbering the mutant types to obtain the gene mutation dictionary, which includes multiple mutant types corresponding to each unit partition and their numbers; Wherein, the unit partition includes at least one of a coding region, a non-coding region, and a gene.

2. The method according to claim 1, wherein The species is a human; the multiple individuals are more than 1000 human bodies.

3. The method according to claim 1, characterized in that The number of the biologically meaningful unit partitions is thousands to tens of thousands.

4. The method according to claim 1, wherein The number of the biologically meaningful unit partitions is 60,000, allowing an error range of plus or minus 10%.

5. The method according to claim 4, wherein The unit partition includes 30,000 gene coding regions and 30,000 non-coding regions, and their numbers allow an error range of plus or minus 10% respectively.

6. The method according to claim 1, characterized in that, The step of statistically analyzing the mutant situations of each unit partition respectively, generating mutant types and numbering the mutant types includes: For each of the unit partitions, sequentially taking the mutation results of the multiple individuals as mutant types for numbering and counting in the order of individuals, where the counted number is the number of individuals supporting the mutant type, and if the mutation result of a later individual is the same as that of any previous individual, then adopting the mutant type and its number of the previous individual and adding 1 to the count of the mutant type; if the mutation result of a later individual is different from the mutation results of all previous individuals, then adding a new mutant type to the dictionary, that is, taking it as a new mutant type and giving a number and a count, and finally obtaining all mutant types of each of the unit partitions and the number and count of each mutant type.

7. The method according to claim 6, characterized in that, The method further includes: Sorting each mutant type from largest to smallest according to its count and renumbering the mutant types in order.

8. An apparatus for creating a gene mutation dictionary, characterized in that, The device includes: A data acquisition unit for obtaining genomic sequence data of multiple individuals of a species and reference genomic data of the species; A data alignment unit for aligning the genomic sequence data of the multiple individuals to the reference genomic data respectively to obtain the mutation results of the genomic sequence data of each individual relative to the reference genomic data; A partition division unit for dividing the genome of the species into several biologically meaningful unit partitions, wherein the unit partition includes at least one of a coding region, a non-coding region, and a gene; A dictionary generation unit, configured to respectively count the mutant situations of each unit partition according to the mutation result, generate all mutant types of each unit partition in the multiple individuals, number the mutant types, and obtain the gene mutation dictionary, where the gene mutation dictionary includes multiple mutant types corresponding to each unit partition and their numbers.

9. A computer-readable storage medium, comprising a program that can be executed by a processor to implement the method according to any one of claims 1 to 7.

10. A method for compressing genomic data using a gene mutation dictionary, characterized in that, The method includes: Obtaining genomic sequencing data of an individual, where the genomic sequencing data includes multiple biologically significant unit partitions; Aligning the genomic sequencing data to the gene mutation dictionary generated by the method according to any one of claims 1 to 7, to obtain the mutant types and their numbers in the gene mutation dictionary that are consistent with the mutant situations of each unit partition of the individual; Storing the numbers of the mutant types of each unit partition of the individual instead of the actually measured mutations.

11. An apparatus for compressing genomic data using a gene mutation dictionary, characterized in that, The device includes: A data acquisition unit, configured to obtain genomic sequencing data of an individual, where the genomic sequencing data includes multiple biologically significant unit partitions; A data alignment unit, configured to align the genomic sequencing data to the gene mutation dictionary generated by the method according to any one of claims 1 to 7, to obtain the mutant types and their numbers that are consistent with the mutant situations of each unit partition of the individual; A data storage unit, configured to store the numbers of the mutant types of each unit partition of the individual instead of the actually measured mutations.

12. A computer-readable storage medium, comprising a program that can be executed by a processor to implement the method according to claim 10.

13. A method for restoring and utilizing genomic data compressed by a gene mutation dictionary, characterized in that, The method includes: Obtaining compressed genomic data, which is data compressed by using the gene mutation dictionary generated by the method according to any one of claims 1 to 7, and includes the numbers of the mutant types of multiple unit partitions in the gene mutation dictionary; Finding the numbers of the mutant types of each unit partition in the gene mutation dictionary from the compressed genomic data; Extracting the corresponding mutant types in the gene mutation dictionary and their mutation results at each base site according to the numbers of the mutant types.

14. An apparatus for restoring and utilizing genomic data compressed by a gene mutation dictionary, characterized in that, The device includes: A data acquisition unit, configured to obtain compressed genomic data, which is data compressed by using the gene mutation dictionary generated by the method according to any one of claims 1 to 7, and includes the numbers of the mutant types of multiple unit partitions in the gene mutation dictionary; A number acquisition unit, configured to find the numbers of the mutant types of each unit partition in the gene mutation dictionary from the compressed genomic data; A mutant extraction unit, configured to extract the corresponding mutant types in the gene mutation dictionary and their mutation results at each base site according to the numbers of the mutant types.

15. A computer-readable storage medium, comprising a program that can be executed by a processor to implement the method according to claim 13.

Citation Information

Patent Citations

  • Image conveying device and its method

    CN101335895A

  • Method for coding DNA sequence and device and computer readability medium

    CN1536068A

  • Sequence Data Analyzer, DNA Analysis System and Sequence Data Analysis Method

    US20170017717A1