Unified coding method, storage method and calculation method for large-scale genetic variation data

By designing the SWI encoding of the genotype, floating length integer encoding, BEG series storage encoding, dynamic translation framework and computing encoding, the problem of inefficient genotype data processing in the prior art is solved, and efficient storage and fast computing performance is achieved.

CN120220802APending Publication Date: 2025-06-27SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510191410.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology cannot simultaneously realize efficient coding, storage, access and calculation of genotype data, resulting in inefficient data analysis and cannot meet the processing needs of genomic data of super-large populations.

Method used

A unified encoding storage calculation method for large-scale genetic variant data is designed, including SWI encoding of genotypes, floating length integer encoding, BEG series storage encoding, dynamic translation framework and computing encoding, realizing efficient storage, fast access and high-performance computing.

Benefits of technology

It realizes efficient coding and storage of genotype data, significantly accelerates data access and computing performance, and can save storage space and computing costs in ultra-large-scale genomic data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220802A_ABST
    Figure CN120220802A_ABST
Patent Text Reader

Abstract

The invention discloses a unified coding, storing and calculating method for large-scale genetic variation data. The method comprises the following steps: a genotype coding and memory representation method: coding a genotype from a | b into a unique integer value; the floating length integer coding method is a method for coding an integer from a traditional fixed length byte to a floating byte and is used for recording an index of a sample; storage coding of a genotype sequence: converting the genotype from an integer value sequence into a byte coding sequence which can be directly stored by a computer; a dynamic translation framework of the genotype sequence: quickly loading the byte coding sequence stored in the hard disk in the step (3) into a memory in a low-load manner; and calculating and coding a genotype sequence: converting the genotype from an integer value sequence into 64-bit space coding accelerated calculation according to allele index (x)-ploidy index (y)-sample group (z). The method is of great significance in improving the data processing efficiency and promoting the development of large-queue genetics research and precise medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the interdisciplinary field of computer and genetics, and discloses a unified coding storage and calculation method for large-scale genetic variation data. Background Art

[0002] Genetic DNA sequence variations of an individual sample's genome are usually referred to as genotypes. Genotype coding technology is a method for numbering and representing alleles at gene loci in large-scale sample analysis. As a diploid species, humans have one allele on each of the two homologous chromosomes at each gene locus. These alleles correspond to different DNA sequences and have multiple variable forms. In analysis, these variable forms are numbered as REF (reference allele) and ALT (variant allele). An individual's genotype is represented as a / b, where a and b are numerical indices referring to the specific forms of the alleles. For example, in the allele list of A, T, G, C, 0 / 2 represents the genotype A / G. Although this coding method is suitable for human reading, it is not conducive to computer storage and calculation.

[0003] In the genetic pedigree analysis of disease phenotypes, the collection and analysis of large-scale samples are crucial. The progress and cost reduction of gene sequencing technology have spawned large-scale genome sequencing projects worldwide. These projects have not only broadened the boundaries of genetic research but also provided a solid data foundation for personalized medicine. For example, the UK Biobank (UKBB) project in the UK has covered 500,000 samples through whole-genome sequencing, involving more than 1 billion variant sites. Its genotype data is stored in a compressed variant call format (VCF), occupying a storage space of up to 1 PB. The China Kadoorie Biobank (CKB) project in China has also covered more than 510,000 samples, and the expected genotype data volume will also reach the PB level.

[0004] In the prior art, genotype data is obtained in the form of a VCF file after population mutation detection from the raw data (FASTQ) and saved in text form. Each genotype occupies at least 3 bytes of information (such as "0 / 0"). Such a text form is suitable for human reading but not suitable for computer storage and calculation. When genotypes are loaded into memory for downstream calculations (including statistical allele or genotype frequencies, linkage disequilibrium calculations for paired loci, genotype-phenotype association analysis), it takes a lot of time to translate genotypes from "text" into an integer or bit-coded form that can be calculated by a computer.

[0005] Internationally, the coding of genotypes is mainly divided into three major technologies:

[0006] ● Mutation count-based method: Convert "a / b" to "how many mutations are there". For example, encode the genotype "0 / 1" as 1, indicating that this genotype has 1 mutation; encode the genotype "1 / 1" as 2, indicating that this genotype has 2 mutations. This is the genotype encoding performed by PLINK and most software in genome-wide association studies.

[0007] ● Method of encoding genotypes as bitmaps: Encode the genotype "0 / 0" as "00", "0 / 1" as "01", "1 / 1" as "11", and ". / ." as "10". Genotype compression algorithms such as the GQT algorithm, BGT, PBWT, GTC, GTShark, PLINK-BED, etc. are all based on such techniques or their derivative methods. ● Method of encoding genotypes as integers: The binary format BCF of VCF splits the genotype "a / b" into two numbers, and then encodes them using 1 byte, 2 bytes, 3 bytes, etc. according to the value range.

[0008] In China, there is little research on genotype encoding. Currently, only the KGGSeq's bit-block genotype format (KED encoding) and Byte-Encoded of Genotype (BEG encoding) published by Li Miaoxin, Zhang Liubin, etc. in the Nucleic Acids Research journal and the Genome Biology journal have solved the problems of genotype phasing and scalability.

[0009] None of the above encodings can simultaneously achieve the following characteristics: high efficiency in computational encoding (having bitwise operation methods to accelerate downstream calculations), high efficiency in storage (being highly compressible), high efficiency in access (supporting fast random access in large encoding arrays), and lossless (supporting phasing and multi-allelic genotypes).

[0010] The objective of the present invention is to design a set of genotype encoding systems that can simultaneously achieve the above characteristics to meet the high-performance storage, access, management, and computational requirements of ultra-large-scale population genomic data. Summary of the Invention

[0011] Although these projects have achieved remarkable achievements in data collection, they still face challenges in the storage, access, calculation, and management of genotype data. Currently, there is a lack of lossless, unified, and efficient encoding technologies, which limits the efficiency and feasibility of data analysis. Therefore, the development of new genotype encoding technologies is of great significance for improving data processing efficiency and promoting the development of genetics research and precision medicine.

[0012] The technical solution of the present invention is as follows.

[0013] A unified coding storage and computing method for large-scale genetic variation data, comprising the following steps

[0014] (1) Genotype coding and in-memory representation method: Encode the genotype from a∣b into a unique integer value.

[0015] (2) Floating-length integer coding method: A method of encoding an integer from a traditional fixed-length byte (e.g., int is a fixed 4-byte representation, long is a fixed 8-byte representation) into floating bytes, used to record the index of the sample.

[0016] (3) Storage coding of genotype sequences (BEG series coding): Convert the sequence of integer values obtained from step (1) or step (2) of the genotype into a byte coding sequence that can be directly stored by a computer; according to the maximum number of alleles, types, frequencies, and distributions of the genotype, the present invention includes genotype sequence storage coding methods for different data characteristics; this coding saves memory overhead and storage overhead while maintaining the indexable property.

[0017] (4) Dynamic translation framework for genotype sequences: The byte coding sequence obtained in step (3) can be directly written to the hard disk. This framework proposes how to dynamically and on-demand translate the BEG series coding and the genotype binary coding method established by other programs into genotypes at specific positions in a memory-efficient (low-load) form without performing traversal and full data restoration.

[0018] (5) Computational coding of genotype sequences: Convert the sequence of integer values obtained from step (1) or step (2) of the genotype into a 64-bit spatial coding according to allele index (x)-haplotype index (y)-sample group (z). At this time, traversing genotypes, counting allele and genotype frequencies, performing genotype association calculations for paired loci, etc. can all be completed using bit operations. This coding is a lossless bit coding, which uses less memory space, has a fast calculation speed (the number of loop calculations is reduced by 64 times), and can be quickly converted with mainstream coding strategies.

[0019] In the above method, step (1) is specifically as follows

[0020] (I) Preheating of the character coding module

[0021] S1 First, according to the maximum number of allelic genotypes V set by the system max Pre-generate (V max +1) 2 possible diploid genotype objects (. / ., 0 / 0, 0 / 1, 1 / 1, etc.); the genotype object contains a coding value, a left genotype, and a right genotype value; V maxThe recommended value is 255, the maximum value is 4095, and the included allele indices are -1, 0, 1, …, V max -1, where -1 represents the missing genotype; the encoded value is an integer and uses the SWI encoding;

[0022] S2 stores the (V max +1) 2 genotype objects obtained by S1 into the G1 and G2 arrays. Among them, G1 establishes the mapping from the genotype encoded value (integer, wrapped sequential integer encoding, SWI, see Embodiment 1 for details) to the genotype object; G2 establishes the mapping from the left and right values of the genotype to the genotype object;

[0023] (For example, for the genotype 0 / 2, through G2, G2[1][3] is obtained, which is the object of 0 / 2 in memory. This object can obtain the encoded value 10, and the above object can also be obtained through G1

[10] )

[0024] (II) Text Processing and Genotype Encoding

[0025] S1 reads the text data (or encoded data in other formats) containing genotypes input by the user, usually a large matrix; the rows of the matrix represent genetic loci, and the columns represent samples;

[0026] For the genotype of each sample at each genetic locus, S2 parses to obtain the left genotype value i and the right genotype value j;

[0027] S3 converts the left and right genotype encoded values obtained in S2 into genotype objects through the G2 matrix, and obtains the SWI encoded value through this object (this step is the index access of the array, with an o(1) time overhead);

[0028] After all samples at a single genetic locus are parsed, a genotype memory object for this locus is generated. According to the number of variable genotype types, its storage container is byte[], short[] or int[].

[0029] In the above method, in step (2), the encoding steps are as follows:

[0030] S1: Check the length of the value. Determine the number of additional bytes n required to encode the value according to the input integer value i i , and allocate a byte array E with a length of n i +1, and write the length marker bit information; i , and write the length marker bit information;

[0031] S2: If i < 0, then make the integer i be i ← -(i + 1), and fill 1 in the sign bit;

[0032] S3: Convert the integer i into a binary sequence, fill it into the data bits, and pad the last byte to 8 bits with 0s.

[0033] In the above method, in step 2), during decoding, conversely, first obtain how many consecutive 1s there are at the head of the first byte, that is, the index where the first bit 0 appears, and then read the remaining bytes and convert them into an integer.

[0034] In the above method, in step (3), CBEG, EBEG, MBEG, HBEG, BEG, DBEG, and TBEG encodings are used. Such encodings can be directly written into a file, or written into a file after being connected to other compression methods.

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] The related technology of the present invention is prototyped using the Java language and implemented in the KGGA (https: / / pmglab.top / kgga) platform and GBCv3.0 (https: / / pmglab.top / gbc, the following performance tests are completed using the GBCv3.0 platform).

[0037] Compared with the existing genotype coding schemes, the SWI coding designs a coding and representation method for complex genotypes in large-scale genomes in a fast, simple, highly scalable, and backward-compatible form.

[0038] The BEG series of storage encodings can effectively save the storage space of genotypes, and the encoding is rapid and parallelizable, and can well reduce the storage space for genomic data of any scale ( Figure 4 , Figure 4 the metrics in Figure 5 , Figure 5 include compression ratio, compression speed, and decompression speed), and the effect of accelerating access ( Figure 5 , Figure 5 taking the access to sample genotypes as an example, based on the genotype coding and access methods implemented in the present invention, it is more than 6 times faster than peer tools). PLINK is currently the gold standard tool for genome-wide association analysis, and it is also the only tool in our Figure 3 tests that can process genotypes extremely quickly like the GBC platform. The storage space required by GBC is only 1 / 3 of that of PLINK.

[0039] The dynamic translation framework and genotype coding method for genotypes proposed by the present invention achieve the effect of being faster from VCF to GTB and then to PLINK-PGEN (the method of this patent) than from VCF to PLINK-PGEN (the method provided by PLINK) ( Figure 6 ), and in the processing of ultra-large-scale genomic data, it can save 3 times the computational cost compared to PLINK.

[0040] In addition, the genotype calculation and encoding method proposed by the present invention is comparable in terms of computing performance to the currently fastest tool PLINK in the industry ( Figure 7 ). The encoding of GBC also has the potential to access the CPU for vectorized computing and the GPU for matrix computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is the genotype encoding framework and system.

[0042] Figure 2 It is an example of VarInt encoding (taking 20250101 as an example).

[0043] Figure 3 It is a UML diagram of the genotype sequence dynamic translation framework.

[0044] Figure 4 It is a performance test chart for extracting sample subsets of each gradient on the genotype dataset of chromosome 21 in the UKBB whole genome and performing all genome compression tools (including cold compression and hot compression).

[0045] Figure 5 It is an effect diagram of accelerated access.

[0046] Figure 6 It is a processing speed test chart of GBC and PLINK on the genotype dataset of chromosome 1 in the UKBB whole genome.

[0047] Figure 7 It is a calculation encoding performance chart of GBC taking LD calculation as an example. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0049] (1) Encoding and in-memory representation method of genotype (SWI encoding)

[0050] For a genotype a∣b, the present invention proposes a wrapped sequential integer value (Wrapped Sequential Integer, SWI, as shown above Figure 1 on the right) to encode and represent different genotypes. SWI can also be obtained according to the following calculation formula:

[0051]

[0052] Among them, the missing genotype '.' is regarded as '-1' for the above encoding (in the representation of genotype storage files such as VCF or PLINK, the unilateral values of genotypes are., 0, 1, 2, 3. For example, for genotypes. / ., 0 / 2, 4 / 5,. / . is regarded as -1 / -1, and then encoded using the formula.); the unphased genotype 'a / b' is converted to

[0053] 'min{a,b}∣max{a,b}' and then encoded according to the above formula (min{a,b} is equivalent to the smaller value in genotype a∣b, and max{a,b} is equivalent to the larger value in genotype a∣b. For example, the phased genotype 4|2, when converted to unphased, is 2 / 4.); the haplotype 'a' is converted to 'a∣a' and then encoded according to the above formula. The encoding result of this formula is a non - negative and continuous integer. The encoding table is as shown in Figure 1 the right sub - figure. It is an encoding structure from the inside out. When encountering more allele forms, only the encoding range needs to be extended outward, thus achieving strong scalability and downward compatibility.

[0054] However, when implemented on a computer, if the above formula is used to obtain genotypes for each genotype, the computational complexity is quite large. Therefore, in this embodiment, "warming up" is performed first during program operation. According to the maximum number of allele genotypes V max (the recommended value is 255, the maximum value is 4095, and the included allele indices are 0, 1, …, V max -1), (V max +1) 2 possible genotype objects (including the missing genotype '.') are pre - generated and stored in the global G1 and G2 object arrays (as shown in Figure 1 ). Thus, for a given genotype 'a∣b', the left - hand encoding value a and the right - hand encoding value b are obtained through character parsing, and then the genotype object is obtained according to G2[a + 1, b + 1]. Conversely, if there is an integer encoding value i of a genotype, the genotype object can be obtained according to G1[i], and then the left - hand encoding value a and the right - hand encoding value b can be obtained (in fact, the SWI encoding of the genotype is also its index in the G1 array). This part establishes a fast mapping method among the character representation, integer encoding, and memory object of genotypes.

[0055] Subsequently, when parsing a specific genotype file, according to the number of variable alleles at a locus (for example, the number of variable alleles can be obtained through the REF and ALT fields in VCF, BCF, and PLINK - PGEN files), the following three array - based variable genotype containers are allocated:

[0056] ●LiteGenotypes: A lightweight genotype container that uses byte[] as the underlying container to store the encoded values of genotypes. When the number of variable genotypes is in the range of [0, 15], the integer encoded values of genotypes are in the range of [0, 255], and can be represented by 1-byte integers.

[0057] ●Genotypes: A regular genotype container that uses short[] as the underlying container to store the encoded values of genotypes. When the number of variable genotypes is in the range of [0, 255], the integer encoded values of genotypes are in the range of [0, 65535], and can be represented by 2-byte integers.

[0058] ●LargeGenotypes: A heavyweight genotype container that uses int[] as the underlying container to store the encoded values of genotypes; when the number of variable genotypes ≥ 256, 4-byte integers can be used for representation.

[0059] They all implement the IGenotypes interface, exposing the "get genotype by index" and "set genotype by index" APIs to users. Internally, the conversion between genotypes and integers is implemented through a G1 array. When bridging with other programs or programming languages, only by agreeing on the G1 array encoded in SWI within these programs or languages can the fast transfer and parsing of genotypes be achieved.

[0060] (2) Floating-length integer encoding method

[0061] In large sample sequences, the proportion of rare mutation sites (i.e., mutation sites with a mutation frequency < 1%) is very high. Storing such genotype sequences by encoding the indices of samples with mutations is a more efficient approach. The index range of samples depends on the number of samples. In a genotype sequence containing dozens of samples, its sample index does not exceed 100 and can be represented by 1 byte; in a genotype sequence containing hundreds of thousands of samples, its maximum sample index requires 3 bytes for representation. Although a fixed 4-byte length can accurately represent the integer value of the index, it introduces a large amount of redundant data. For this reason, a floating-length integer encoding method is designed in this embodiment.

[0062] This encoding method encodes an integer value into byte sequences of different lengths. The first byte of the integer reserves n (n ≤ 8) bits of 1 and the following 0s (if 8 additional bytes are required, there is no need for the following 0s), indicating that an additional n bytes are needed to store the data. Then the integer is saved with an offset of the reserved bits.

[0063] Table 1: Floating-length integer encoding method

[0064]

[0065]

[0066] The specific encoding steps are as follows:

[0067] Step 1: Check the length of the value. Determine the number of additional bytes n required to encode the value based on the input integer value i. i , the allocation length is n i +1 Byte array E i , and write the length mark information;

[0068] Step 2: If i<0, let the integer i be i←-(i+1) and fill the sign bit with 1;

[0069] Step 3: Convert the integer i into a binary sequence and fill it into the data bits, using 0 to fill the last byte to make it 8 bits.

[0070] When decoding, we first obtain the number of consecutive 1 bits in the header of the first byte (that is, the index where the first bit 0 appears), then read the remaining bytes and convert them into integers.

[0071] Through the above steps, an integer is effectively encoded into byte codes of different lengths. The length of the integer value is determined in the first byte. This encoding method only needs to be read twice when decoding (get the first byte, parse the total length; get the remaining bytes and bit information). The traditional method (Varint encoding in Google Protobuf) must be read and requested byte by byte to translate into the original integer value.

[0072] (III) Storage coding of genotype sequence (BEG series coding)

[0073] When the genotype is stored on disk, we do not want to perform a complicated text parsing process the next time it is used, so the genotype data needs to be stored in binary format. At the same time, there are many redundant structures in the genotype sequences of a large number of samples, and the distribution of genotype sequences and genotype frequencies need to be considered to streamline the space design. To this end, this patent also includes the following 7 built-in storage encoding schemes (all integer values ​​below, unless otherwise specified, are encoded using VarInt in the second part):

[0074] CBEG( C Constant Byte-Encoded Genotype: Constant genotype byte encoding. When all genotypes are the same, store the SWI encoded value of the genotype. If there is one exception genotype,

[0075] The SWI encoding value of the exceptional genotype and the index of the exceptional genotype are further stored.

[0076] ·EBEG( E numerated Byte-Encoded Genotype): Enumerated genotype encoding. When more than 95% of the genotypes in the genotype sequence are of the same type, store the SWI code value of the major allele genotype, the SWI code value of the minor allele genotype 1, the sample index 1 of the minor allele genotype 1, the sample index 2 of the minor allele genotype 1, …, -1 (end marker), the SWI code value of the minor allele genotype 2, the sample index 1 of the minor allele genotype 2, the sample index 2 of the minor allele genotype 2, …,

[0077] -1, etc.

[0078] ●MBEG( M aximized Byte-Encoded Genotype): Maximized genotype byte encoding. When there are only two genotypes in the genotype sequence (e.g., only 0 / 0 and 0 / 1), set the high-frequency genotype to bit value 0 and the low-frequency genotype to bit value 1. Store the SWI code value of the high-frequency genotype, the SWI code value of the low-frequency genotype, and the bit-encoded genotype byte sequence (bitmap container).

[0079] ●HBEG( H alf Byte-Encoded Genotype): Half-byte genotype encoding. When the number of variable alleles is less than 4, the maximum number of possible genotype types for each sample is at most 16 (including the missing genotype). In this case, store the genotype sequence as 4 bits (0.5 bytes) / genotype.

[0080] ●BEG( B yte-Encoded Genotype): Single-byte genotype encoding. When the number of variable alleles is less than 15, the maximum number of possible genotype types for each sample is at most 256 (including the missing genotype). In this case, store the genotype sequence as 1 byte / genotype.

[0081] ●DBEG( D ouble Byte-Encoded Genotype): Double-byte genotype encoding. When the number of variable alleles is less than 255, the maximum number of possible genotype types for each sample is at most 65536 (including the missing genotype). In this case, store the genotype sequence as 2 bytes / genotype.

[0082] ●TBEG( T(Triple Byte-Encoded Genotype): Triple-byte genotype encoding. When the number of variable alleles is less than 4095, there are at most 16,777,216 possible genotype types for each sample (including missing genotypes). In this case, the genotype sequences are stored at 3 bytes per genotype.

[0083] MBEG, HBGE, BEG, DBEG, and TBEG are all fixed-length genotype encoding methods. Their encoding methods determine that for the extraction of specific genotypes, the offset can be calculated based on their lengths for genotype parsing and restoration, without having to convert the genotype from the encoded sequence back to the SWI array (or the three variable genotype containers mentioned in the first part) for extraction. Although CBEG and EBEG do not support direct address translation, in large samples, they can usually reduce the data that originally required more than 100,000 bytes to only a few hundred bytes, which is quite efficient in terms of space. When using them, only the encoded sequence needs to be restored to a more memory-efficient structure, and low-load and fast access can also be achieved (for example, using a dictionary table Map for storage).

[0084] (IV) Dynamic Translation Framework for Genotype Sequences

[0085] When performing subsequent analysis on genomic files of a large number of samples, almost rarely all sample genotype data is used. In the past, on the one hand, it was due to technical and hardware resource limitations caused by the lack of effective encoding, and on the other hand, it was because when analyzing case-control samples, it was usually necessary to accurately screen samples with specific phenotypes. However, in the past, when implementing sample subset screening, all genotype encoding sequences had to be read, restored to genotype objects, and accessed through indexing, which undoubtedly caused huge time overhead.

[0086] This embodiment includes a dynamic translation framework for genotype sequences, and the UML diagram is as Figure 2 shown. Specifically, genotype sequences are usually encoded as long, memory-continuous binary data. When reading a locus, record the physical address of the encoded sequence corresponding to the locus (bound to the file). When accessing a specific genotype through the get(int index) API, according to different genotype types, trigger the corresponding address translation (map to the offset byte and perform disk I / O).

[0087] When implementing in a specific programming language, the disk I / O during address translation should include cache reading (for example, using BufferReader in Java). And the genotypes that have been read once should be saved by a cache structure (such as an array).

[0088] In addition, this dynamic translation framework can also be applied to the translation of other genotypes. Taking PLINK-PGEN as an example, the genotype sequences read from this file are presented as P = [a0, b0, a1, b1, …, a n-1 , b n-1 , where a i , b i is the genotype a i ∣b i of sample i. When bridging the genotype in this format to the BEG coding system, only the API method of get(int index) needs to be designed: when obtaining the genotype of the specified sample i, it will be translated into G2[P[i << 1], P[(i << 1) + 1]]. This design enables the bridging between different genotypes to have a unified specification without having to perform the conversion of the entire genotype sequence.

[0089] (V) Computational Encoding of Genotype Sequences

[0090] When genotypes participate in downstream calculations, if only one traversal and scan are required, it can be performed at the SWI coding array level (using the variable genotype container designed in the first part, or the dynamic translation genotype method designed in the third and fourth parts). However, when a large number of pairwise calculations are involved and each locus participates in more than one calculation (that is, it needs to be scanned many times), more efficient computational encoding becomes particularly important.

[0091] Modern CPUs support bitwise operations (such as AND, OR, XOR) at the hardware level, and their execution efficiency is much higher than array operations (especially floating-point calculations). In addition, for the genotype sequence after bit encoding, its bitwise operations can operate on multiple genotypes at once (in a 64-bit computer device, 64 sample pairs can be calculated at once), greatly improving performance. This embodiment proposes a computational encoding of a new type of genotype sequence, which is described as follows:

[0092] For polyploid species (humans are diploid species), the haplotypes it contains are Y = {0, 1, …}. When analyzing N samples, they are divided into coding groups of 64 samples each, resulting in coding groups in total, and the index set of the coding groups is recorded as

[0093] For the specified variant locus i, the variable allele index list X i generates an indicator sequence of valid allele genotypes (i.e., non-missing genotypes) and a genotype sequence where indicates that the genotype of the y-th haplotype of sample 64z + j is non-missing, The genotype of the y-th haplotype of sample 64z + j is x. Without considering semi-calls (i.e., there is no half-missing genotype such as. / 0), can be reduced to The above bit is a bit position, representing the j-th bit of the long integer value That is, Performing a bitwise AND operation on multiple numbers is similar to the meaning of the summation symbol ∑ in mathematics, and the same applies to performing a bitwise OR operation on multiple numbers;

[0094] This computational encoding retains all genotype signals while having a high-performance bitwise operation method and can quickly convert to the bit encoding of other programs. Based on this encoding, the common computational methods of genotypes can be optimized as follows:

[0095] 1) Swap the genotypes of allele index j and allele index k (set the high-frequency allele

[0096] to REF and the low-frequency allele to ALT):

[0097]

[0098] 2) Set allele index j to allele index k (k ∈ {0, 1}), and set the non-allele index

[0099] j to 1 - k (for binary allele conversion during GWAS analysis):

[0100]

[0101] 3) Retain the specified set of alleles (The order in X i is the same as that in X i ), and the rest are set to missing genotypes (binary allele conversion of PLINK - BED, retaining the top K high-frequency

[0102] alleles): where & represents the bitwise AND operation in the computer;

[0103]

[0104] 4) Count the number of the specified allele a (calculate allele frequency): bitcount calculates the number of bit positions with 1 in the binary sequence of the given integer value; for example, the binary of 255 is 0000000011111111, and bitcount(255) = 8;

[0105]

[0106] 5) Count the number of specified genotypes a|b|… (calculate the homozygosity rate, heterozygosity rate, genotype frequency):

[0107]

[0108] 6) Count the number of effective genotypes (non-empty genotypes, used to calculate the number of effective alleles, allele

[0109] frequency):

[0110]

[0111] 7) Weight according to the number of specified alleles a carried (taking diploid as an example, w k represents the

[0112] weight given when k mutations are carried at this

[0113]

[0114] Based on the above coding method, the following shows 3 common optimized expressions for paired calculation of population genotypes commonly used in human genetics research. Since this coding reduces the number of traversals, the speed can usually be increased by more than 64 times.

[0115] 1) Calculate the linkage disequilibrium coefficient using the D’ method

[0116] Suppose Aa, Bb are alleles at two loci. If the two loci are independent and not related, then the frequency (combination) of each genotype can be given by the allele frequencies:

[0117] P AB = P A P B , P aB = P a P B P Ab = P A P b , P ab = P a P b

[0118] The deviation between the actual value and the theoretical value is called "linkage disequilibrium" and is measured by D’:

[0119] D = P ab - P a × P b

[0120]

[0121] Or the correlation coefficient:

[0122]

[0123] The corresponding bitwise operation expression is:

[0124]

[0125] 2) Using the MI method to measure the conditional correlation coefficient between variant sites. Mutual Information (MI) is a measure of the mutual dependence between two random variables. Use MI to measure the LD coefficient between two sites:

[0126]

[0127] Normalized NMI coefficient:

[0128]

[0129] Conditional Mutual Information (CMI) is an extension of mutual information and is used to measure the correlation between two other random variables given a third random variable:

[0130]

[0131] Mutual information can be used to calculate the linkage disequilibrium coefficient, can also be extended to the calculation between multi-allelic genotype sites, can also be extended to the joint calculation between multiple sites, and can also calculate the LD effect under a specified phenotype. Past coding methods could not perform such calculation acceleration, and the present invention realizes this for the first time:

[0132] The corresponding bitwise operation expression:

[0133]

[0134] P a· = P ab + P aB , P A· = P Ab + P AB , P ·b = P ab + P Ab , P ·B = P aB + P AB

[0135]

[0136] H a· = P a· logPa· , H A· = P A. logP A. , H ·b = P ·b logP .b , H ·B = P ·B logP ·B ,

[0137] MI = H aa + H aB + H Ab + H AB ,

[0138] 3) Calculate the linkage disequilibrium coefficient using the Pearson correlation coefficient method

[0139] The calculation of the linkage disequilibrium coefficient mentioned in the above 1) and 2) is sensitive to phased (phase information). When calculating LD for unphased genotypes, genotype correlation is usually used for calculation:

[0140] Given x si represents the number of mutant genotypes of sample s at locus i (for example, 1 when 0 / 1, 2 when 1 / 1). The genotype correlation calculation between locus i and locus j:

[0141]

[0142] The corresponding bitwise operation expression:

[0143]

[0144]

[0145] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for unified coding, storage and calculation of large-scale genetic variation data, characterized in that: The following steps are included (1) Genotype encoding and memory representation method: Encode the genotype from a|b into a unique integer value; (2) Floating-length integer encoding method: a method of encoding an integer from traditional fixed-length bytes to floating bytes, which is used to record the index of the sample; (3) Storage encoding of genotype sequence: converting the genotype from the integer value sequence obtained in step (1) or step (2) into a byte code sequence that can be directly stored by a computer; (4) Dynamic translation framework of genotype sequence: the byte code sequence stored in the hard disk obtained in step (3) is loaded into the memory with low load and fast speed; (5) Calculation and encoding of genotype sequence: The integer value sequence of genotype obtained from step (1) or step (2) is converted into a 64-bit spatial encoding according to allele index (x)-ploidy index (y)-sample group (z).

2. A large-scale genetic variation data unified coding storage calculation method as claimed in claim 1, characterized in that: Step (1) is specifically as follows: (I) Character encoding module preheating S1 first calculates the maximum number of alleles V set by the system max Pre-generated (V max +1) 2 possible diploid genotype objects; genotype objects contain SWI encoding value, left genotype, right genotype value; V max The recommended value is 255, the maximum value is 4095, and the allele indexes included are -1, 0, 1, ..., V max -1, where -1 indicates a missing genotype; the encoding value is an integer, using SWI encoding; S2 converts the (V max +1) 2 The genotype objects are stored in the G1 and G2 arrays; G1 establishes the mapping from the genotype encoding value (integer, SWI encoding) to the genotype object; G2 establishes the mapping from the left and right values ​​of the genotype to the genotype object; 2. Text processing and genotype encoding S1 reads the text data containing genotypes input by the user, usually a large matrix; the rows of the matrix represent genetic loci and the columns represent samples; S2 analyzes the genotype of each sample at each genetic locus to obtain the genotype value i on the left and the genotype value j on the right; S3 converts the left and right genotype coding values ​​obtained in S2 into genotype objects through the G2 matrix, and obtains the SWI coding values ​​through this object; After all samples of a single genetic locus are parsed, a genotype memory object for the locus is generated. Depending on the number of variable genotypes, its storage container is byte[], short[] or int[].

3. A large-scale genetic variation data unified coding storage calculation method as claimed in claim 1, characterized in that: In step (2), the encoding step is: S1: Check the length of the value; determine the number of additional bytes n required to encode the value based on the input integer value i i , the allocation length is n i +1 Byte array E i , and write the length mark information; S2: If i<0, let the integer i be i←-(i+1) and fill the sign bit with 1; S3: Convert the integer i into a binary sequence and fill it into the data bits, using 0 to fill the last byte to make it 8 bits.

4. A large-scale genetic variation data unified coding storage calculation method as claimed in claim 1, characterized in that: In step 2), when decoding, first obtain the number of consecutive 1 bits in the header of the first byte, that is, the index where the first bit 0 appears, and then read the remaining bytes and convert them into integers.

5. A large-scale genetic variation data unified coding storage calculation method as claimed in claim 1, characterized in that: In step (3), CBEG, EBEG, MBEG, HBEG, BEG, DBEG, and TBEG encoding are used.

Citation Information

Cited By

  • River crab parent breeding data management system and method

    CN121393578A