A multi-base encoding and readout method for composite base DNA storage
Through multi-digit encoding and readout methods, combined with packet error correction code and information entropy method, the problems of information capacity and sequencing recovery cost in DNA storage are solved, and data recovery under high density storage and low coverage are realized.
Patent Information
- Application Number
- CN202411726207.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the existing DNA storage technology, the information capacity based on standard bases is limited, which is difficult to meet the high-density storage needs, and the sequencing recovery of complex bases requires high sequencing coverage, which increases the cost.
Multi-digit encoding and reading methods are used to map data into DNA sequences containing standard bases and complex bases using grouping error correction codes, and character classification and error correction decoding are performed through information entropy and log-likelihood methods to achieve high-reliability data recovery.
The information capacity per synthesis cycle is improved to 3 bits, reducing the sequencing coverage requirement, and achieving efficient data recovery and low-cost storage.
Smart Images

Figure CN119649874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA data storage, and in particular to a multi-base encoding and reading method for composite base DNA storage. Background Art
[0002] With the continuous growth of global data volumes, traditional magnetic, optical, and electrical storage media are unable to meet the demand for long-term data storage. Therefore, the exploration of new storage media has become a hot research area. In 2012, researchers proposed storing digital information in synthetic DNA using high-throughput DNA synthesis and sequencing technologies, a concept known as DNA data storage. The main advantages of DNA storage are high storage density and long media lifespan, making it considered a data storage technology with great potential. Using a large number of short oligonucleotides to form an oligonucleotide pool for data storage has become the predominant model for DNA data storage. DNA data storage involves two basic steps: writing data into synthetic DNA and reading data from the synthesized oligonucleotides. The writing step primarily involves adding coding redundancy for error correction, converting the information sequence into a 4-base sequence, and synthesizing specific DNA chains. The reading step involves using sequencing technology to obtain sequence reads from a DNA sample, reconstructing the original base sequence from multiple sequenced copies using a merging algorithm, and then decoding and recovering the information.
[0003] The core of DNA data storage technology is the use of a DNA synthesizer to synthesize artificially designed DNA sequences for data storage. This technology is limited by the cost of DNA synthesis. However, there are only four natural bases. In the process of converting binary data into DNA, each base has a theoretical information capacity of 2 bits (e.g., A=00, T=01, G=10, C=11). Information capacity here refers to the amount of digital information that can be stored in a single synthesis cycle. The type of base and encoding method are key factors affecting DNA storage density. Research by Erlich et al. on DNA storage fountain coding based on a quaternary encoding system achieved the highest information capacity of 1.57 bits for standard bases to date and demonstrated that DNA storage density can reach 215PB / g.
[0004] DNA synthesis technology based on phosphoramidite chemistry synthesizes numerous identical molecules simultaneously during the synthesis process, creating multiple copies of each designed DNA sequence. These copies represent identical information, leading to information redundancy. However, by mixing nucleotides at specific ratios during synthesis to create a "composite DNA alphabet," this method can effectively store more information in each synthesis cycle, reducing the number of synthesis cycles. Writing a composite base at a fixed position in a DNA sequence is equivalent to synthesizing multiple oligonucleotide copies of the sequence. The generation of composite bases offers unique advantages in high-throughput synthesis. First, it combines the four standard DNA nucleotides (adenine A, thymine T, guanine G, and cytosine C) and precisely controls their ratios, creating mixed sites with varying base ratios within the synthesized DNA chain. This means that at each base position, a specific ratio of mixed bases can replace the original single base. This method allows for the simultaneous participation of multiple bases, resulting in higher information density. The randomness and diversity of these composite bases make them ideal for applications in high-throughput sequencing and complex information encoding. They can simultaneously represent multiple possibilities at a fixed position, significantly increasing the redundancy of information representation and sequence diversity.
[0005] In 2019, researchers proposed using composite bases generated by mixing in equal proportions to increase information capacity. They introduced 11 composite bases mixed in equal proportions into the DNA code, effectively increasing information capacity to 3.37 bits per symbol (Scientific Reports, 2019). After synthesis, composite bases exhibit base diversity at the same position in different sequences. During data readout, if appropriate composite base reconstruction algorithms are not employed, distinguishing composite bases from standard bases based solely on the observed ratios of the four bases at the same position requires high sequencing coverage. In their composite base recovery experiments, they did not employ a specialized algorithm, but instead determined the original base alphabet based solely on the most frequently observed base type. This resulted in error-free recovery requiring a sequencing depth of at least 250×, which increases sequencing costs. A recent study proposed increasing logical density by using a composite DNA alphabet mixed in arbitrary proportions. This core approach consists of a set of four standard DNA nucleotides combined in predetermined, free-range proportions (Nature Biotechnology, 2019). Theoretically, the composite alphabet can be expanded indefinitely. However, due to the limitations of synthesis technology and precision, it is impractical to synthesize in exact proportions. Their work achieved a logical density of 1.96 bits / nt using six composite letters and 4.29 bits / nt using 20 composite letters. However, the scale of the experiment was small, encoding only a 42nt composite letter sequence. In their work, they designed a relative entropy-based method to distinguish between different composite letters, which reduced sequencing coverage to some extent, achieving 100% error-free recovery of data at 48× coverage. Summary of the Invention
[0006] The present invention provides a multi-base encoding and readout method for composite base DNA storage, achieving matching of the multi-base encoding model with the composite base, breaking through the existing storage limit and increasing the theoretical information capacity limit of 2 bits per synthesis cycle to 3 bits. Detailed description is given below:
[0007] A multi-base encoding and reading method for composite base DNA storage, the method comprising the following steps:
[0008] (1) At the composite base data writing end, a group error correction code is used to encode the user data as a whole to obtain a codeword. The codeword is converted into a multi-base symbol sequence according to a mapping method in which each m bits corresponds to a multi-base symbol d. The codeword is then divided into short segments. Standard base letters and composite base letters corresponding to the multi-base symbols are selected according to a user-defined method to obtain a DNA sequence containing standard bases and composite bases. A highly reliable address sequence and primer sequence are added at both ends to obtain a composite base oligonucleotide pool for storing data.
[0009] (2) At the composite base data reading end, with the help of second-generation high-throughput sequencing, the composite bases of each synthesis site are converted into sequencing reads composed of standard bases. According to the primer sequence and address sequence, the unique position of the sequencing read is determined and clustered. Then, the base observation frequency in each address cluster is counted and converted into information entropy. Based on the separable property of entropy, the character set of each synthesis site is divided into groups. Furthermore, in each character set, the maximum likelihood is calculated between the observation frequency and the candidate character set to detect standard bases and composite bases. Finally, the inferred characters are sent to the error correction decoder to correct the remaining substitution errors and erasure errors and restore the original data.
[0010] The step (1) is specifically as follows:
[0011] (1.1) Using a binary or multi-bit block error correction code long code C1(n1,k1), the digital file of length k1 bits is encoded as a codeword of length n1 bits. Both the binary and multi-bit encoding processes are calculated on a bit basis, where k1 and n1 are positive integers representing the length of information bits and the length of code bits, respectively.
[0012] (1.2) Divide each m bits of the codeword into a group and map it into a multi-ary symbol d according to a predetermined mapping rule, where the multi-ary symbols have a total of 2 m The multi-ary symbol sequence of length n1 / m is obtained, and then divided into S symbol sequences of length L, corresponding to S payload DNA sequences, where m is a positive integer not less than 3, d∈[0,2 m -1], S*L=n1 / m;
[0013] (1.3) According to the mapping relationship between the predetermined symbols and DNA bases, the standard base set {A, T, G, C}, the composite base {Y, M, K, R, S, W, R, B, V, H, D, N} and the composite bases {Y i ,M i ,K i ,R i ,S i ,W i ,R i ,B i ,V i ,H i ,D i ,N i}, i represents a variety of composite bases mixed with different proportions of standard bases, select 2 m The letters are used to represent the encoded multi-base symbol sequence to obtain the payload DNA sequence, and the payload DNA sequence contains standard bases and composite bases. The vector representing the mixing ratio of all four standard DNA nucleotides in the composite base during the synthesis process can be expressed as:
[0014] σ=(σ A ,σ T ,σ G ,σ C ),
[0015] Among them, the vector σ satisfies σ A +σ T +σ G +σ C =1;
[0016] (1.4) Adding address sequences and primer sequences to both ends of S data-carrying sequences of length L to generate an oligonucleotide pool containing standard bases and composite bases, wherein the address sequences and primer sequences are generated from the four standard bases {A, T, G, C};
[0017] (1.5) Based on the ratio of the standard bases corresponding to the composite bases or a single standard base, each site of each information sequence is synthesized in sequence by DNA synthesis, and all the synthesized sequences are mixed into a pool to obtain a composite oligonucleotide pool.
[0018] Wherein, the step (1.5) is specifically as follows:
[0019] According to the vector σ=(σ A ,σ T ,σ G ,σ C ), synthesize each site in turn, for the standard base, A=(1,0,0,0), T=(0,1,0,0), G=(0,0,1,0), C=(0,0,0,1), composite base Indicates that the composite base Y is synthesized by mixing the standard base T and the standard base C in the same proportion. It is synthesized by mixing the standard base T and the standard base C in a ratio of 2:1, and the same applies to other cases;
[0020] The step (2) is specifically as follows:
[0021] (2.1) Amplifying the synthesized composite oligonucleotide pool, preparing the library, and performing second-generation high-throughput sequencing to obtain sequencing reads composed of standard bases;
[0022] (2.2) Identify the primer sequence and address sequence of the sequencing read, determine the unique address number of the sequencing read, obtain all DNA sequences corresponding to the same primer and address sequence, and aggregate them into clusters;
[0023] (2.3) Align the r data-bearing region sequences within each cluster and count the frequencies of the four standard bases {A, T, G, C} at each different synthesis site i corresponding to the r sequences, denoted as: r i,A ,r i,T ,r i,G ,r i,C , where the observation frequency satisfies
[0024] (2.4) Based on the base frequencies observed at each synthesis site, calculate the frequency probability distribution p of the four standard bases i =(p i,A ,p i,T ,p i,G ,p i,C ), and calculate the information entropy, and then divide different types of characters into different sets based on the separable characteristics of entropy;
[0025] (2.5) Within each set, the normalized maximum log-likelihood method is used to compare the observed frequencies with the candidate character set to accurately detect standard bases and composite bases. Each inferred character is then demapped into m bits in turn and concatenated into a block code of length n1 bits. A highly reliable error correction decoding method is then used to correct any remaining substitution and erasure errors, restoring the original digital file.
[0026] The step (1.4) is specifically as follows:
[0027] (3.1) Select the address information of length k2 in the range [1, S] to represent the order information of S multi-ary sequences, where Then, the short block error correction code C2(n2,k2) is used to encode the address bits of length k2 to generate a short block error correction code codeword of length n2 bits, and a check bit that can represent the parity check is added to form an encoded address bit sequence of length L1;
[0028] (3.2) Group each two bits and map the encoded address bit sequence of length L1 into an address DNA sequence of length L2 according to the mapping rules: 00→A, 01→T, 10→G, 11→C, and add it to the 5' end of the data-bearing sequence;
[0029] (3.3) Generate a parity check bit based on the address bits of length k2, cyclically left-shift the original address bits of length k2 by one bit to obtain a left-shifted address bit sequence of length k2, and then follow the concatenation rule of "parity check bit - original address bit - parity check bit - left-shifted address bit" to obtain a coded address bit sequence of length L1.
[0030] (3.4) According to the same mapping rules as step (3.2), the encoded address bit sequence obtained in step (3.3) is converted into an address DNA sequence of length L2, which is added to the 3' end of the data-bearing sequence. Primer sequences for amplification are also added to both ends. The address sequence and primer sequences only use the standard four bases (A, T, G, C) to generate an oligonucleotide pool containing standard bases and composite bases.
[0031] The step (2.4) is specifically as follows:
[0032] (4.1) The base frequency distribution r observed based on the r data-carrying sequences in each cluster i , calculate the frequency probability distribution of the four standard bases, expressed as p i =(p i,A ,p i,T ,p i,G ,p i,C ), and then calculate the information entropy at each fixed point based on the observed frequency distribution The frequency probability distribution satisfies p i,α represents the probability of character α appearing at synthesis position i in the r data-carrying sequence;
[0033] (4.2) Based on the separable property of information entropy, the set segmentation method is adopted. According to the information entropy results at each site, the three character types of standard bases, composite bases, and composite bases with different proportions are classified. And when m=3, all characters are divided into three character sets according to the information entropy results, namely, the standard base character set S1={A,T,G,C}, the two-base hybrid character set S2={Y,M,K,R,S,W} and the three-base hybrid character set S3={H,B,V,D}.
[0034] The step (2.5) is specifically as follows:
[0035] (5.1) In each character set, according to the observed frequency distribution r i , and then the standard vectors σ=(σ A ,σ T ,σ G ,σ C ), calculate the log-likelihood function l(σ,r i ), which is used to infer the character in the set that is closest to the encoded character, and can be expressed as:
[0036]
[0037] (5.2) The log-likelihood function l(σ,r i) is normalized, and the character with the smallest result is regarded as the detected character. Each detected character is then mapped to m bits, and a highly reliable error correction decoding algorithm is executed to correct the remaining substitution errors and erasure errors and restore the original digital file.
[0038] The beneficial effects of the technical solution provided by the present invention are as follows: the present invention proposes a multi-base encoding method and information representation model that matches composite bases, which can achieve highly reliable data recovery; proposes a highly reliable address sequence design method based on short block error correction codes, which can resist synthesis and sequencing errors and achieve more accurate multi-copy read segment clustering; when reading data, it is proposed to use the observed frequency distribution to calculate information entropy for composite base character set segmentation to achieve reliable character classification of different types, and within each character set, based on the log-likelihood method, infer the original multi-base character, and combine linear block code decoding to perform residual error correction to achieve error-free reading of user data. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A block diagram of a multi-base encoding and readout method for complex base DNA storage;
[0040] Figure 2 Flowchart of the multi-bit group error correction coding method designed for the present invention;
[0041] Figure 3 Schematic diagram of selecting standard bases and composite bases as storage media for the present invention;
[0042] Figure 4 Schematic diagram of the address DNA sequence based on the short packet error correction code designed for the present invention;
[0043] Figure 5 Schematic diagram of the structure of the oligonucleotide molecule comprising standard bases and composite bases designed for the present invention;
[0044] Figure 6 Schematic diagram of the DNA synthesis process based on the introduction of bases in the present invention;
[0045] Figure 7 Flowchart for performing double-ended primer identification and high-reliability label identification in the present invention;
[0046] Figure 8 Flowchart of the composite base detection method based on set segmentation designed for the present invention;
[0047] Figure 9 A graph showing the ratio of effective reads for primer identification and address identification of sequencing reads provided by the present invention;
[0048] Figure 10 The length distribution diagram of the sequencing reads after primer truncation provided by the present invention;
[0049] Figure 11 A distribution diagram of the number of copies within a cluster after read segment clustering after label identification provided by the present invention;
[0050] Figure 12 A distribution diagram of information entropy provided by the present invention;
[0051] Figure 13 The normalized log-likelihood distribution graph provided by the present invention;
[0052] Figure 14 This is a bit error rate distribution diagram after reconstruction of the composite base character provided by the present invention. DETAILED DESCRIPTION
[0053] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.
[0054] This invention, leveraging the high number of synthesized molecules per synthesis site in high-throughput DNA synthesis technology, proposes a polynomial encoding and readout method for composite base DNA storage. This method overcomes existing storage limitations, increasing the theoretical information capacity limit of 2 bits per synthesis cycle to 3 bits. The invention uses polynomial encoding to map user data into a polynomial symbol sequence, which is then converted into an oligonucleotide pool containing standard bases and composite bases, which serves as the storage medium. During data readout, multiple sequencing copies are clustered based on a designed highly reliable address sequence. In combination with the multi-copy nature of the encoded polynomial character sequence and the frequency distribution of the four standard bases observed at each fixed site, a method is proposed to calculate information entropy using frequency probability distribution. This method is used to distinguish different types of composite bases using a set partitioning approach. Within each alphabet set, the original polynomial alphabet is inferred from a candidate alphabet based on the log-likelihood method. Compared to directly determining the original base alphabet based on the most frequently observed base type, this method has the advantage of a lower inference error rate. Combined with the designed highly reliable error correction decoding, data recovery is achieved even under low coverage conditions.
[0055] The present invention proposes a multi-base encoding and readout method for composite base DNA storage, designs a multi-base encoding method and information representation model, and can achieve matching with composite bases; proposes an address sequence design method with strong error correction capability based on short block error correction code, which can resist synthesis and sequencing errors and achieve more accurate multi-copy read segment clustering; when reading data, it is proposed to use the observed frequency distribution to calculate information entropy for composite base character set segmentation to achieve reliable character classification of different types, and infer the original multi-base character within each character set based on the log-likelihood method, and combine it with linear block code decoding to perform residual error correction to achieve error-free reading of user data.
[0056] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0057] This embodiment describes in detail a multi-base encoding and reading method for composite base DNA storage proposed by the present invention. Figure 1 , specifically including the following steps:
[0058] (1) At the composite base data writing end, a group error correction code is used to encode the user data as a whole to obtain a codeword. The codeword is converted into a multi-base symbol sequence according to a mapping method in which each m bits corresponds to a multi-base symbol d. The codeword is then divided into short segments. Standard base letters and composite base letters corresponding to the multi-base symbols are selected according to a user-defined method to obtain a DNA sequence containing standard bases and composite bases. A highly reliable address sequence and primer sequence are added at both ends to obtain a composite base oligonucleotide pool for storing data.
[0059] (2) At the composite base data reading end, with the help of second-generation high-throughput sequencing, the composite bases of each synthesis site are converted into sequencing reads composed of standard bases. According to the primer sequence and address sequence, the unique position of the sequencing read is determined and clustered. Then, the base observation frequency in each address cluster is counted and converted into information entropy. Based on the separable property of entropy, the character set of each synthesis site is divided into groups. Furthermore, in each character set, the maximum likelihood is calculated between the observation frequency and the candidate character set to detect standard bases and composite bases. Finally, the inferred characters are sent to the error correction decoder to correct the remaining substitution errors and erasure errors and restore the original data.
[0060] The step (1) is specifically as follows:
[0061] (1.1) Figure 2 As shown, a binary or multi-bit block error correction code long code C1(n1,k1) is used to encode the digital file of length k1 bits into a codeword of length n1 bits. The binary code and multi-bit code encoding process are both calculated based on bits, where k1 and n1 are both positive integers, representing the length of information bits and the length of coding bits, respectively.
[0062] (1.2) Divide each m bits of the code word into a group and map it into a multi-ary symbol d according to a predetermined mapping rule, where the multi-ary symbols have a total of 2 m The multi-ary symbol sequence of length n1 / m is obtained, and then divided into S symbol sequences of length L, corresponding to S payload DNA sequences, where m is a positive integer not less than 3, d∈[0,2 m -1], S*L=n1 / m;
[0063] (1.3) Figure 3As shown, according to the mapping relationship between the predetermined symbols and DNA bases, from the standard base set {A, T, G, C}, the composite base {Y, M, K, R, S, W, R, B, V, H, D, N} and a variety of composite bases with different proportions {Y i ,M i ,K i ,R i ,S i ,W i ,R i ,B i ,V i ,H i ,D i ,N i}, i represents a variety of composite bases mixed with different proportions of standard bases, select 2 m The letters are used to represent the encoded multi-base symbol sequence to obtain the payload DNA sequence, and the payload DNA sequence contains standard bases and composite bases. The vector representing the mixing ratio of all four standard DNA nucleotides in the composite base during the synthesis process can be expressed as:
[0064] σ=(σ A ,σ T ,σ G ,σ C ),
[0065] The vector σ satisfies σ A +σ T +σ G +σ C =1;
[0066] (1.4) Figure 4 As shown, address sequences and primer sequences are added to both ends of S data-carrying sequences of length L to generate an oligonucleotide pool containing standard bases and composite bases, where the address sequence and primer sequence are generated by the four standard bases {A, T, G, C}. The structure of the oligonucleotide molecule is shown in Figure 5 As shown;
[0067] (1.5) Figure 6 As shown, according to the ratio of the standard bases corresponding to the composite bases or a single standard base, each site of each information sequence is synthesized in sequence by DNA synthesis, and all the synthesized sequences are mixed into a pool to obtain a composite oligonucleotide pool.
[0068] Wherein, the step (1.5) is specifically as follows:
[0069] According to the vector σ=(σ A ,σ T ,σ G ,σ C), synthesize each site in turn, for the standard base, A=(1,0,0,0), T=(0,1,0,0), G=(0,0,1,0), C=(0,0,0,1), composite base Indicates that the composite base Y is synthesized by mixing the standard base T and the standard base C in the same proportion. It is synthesized by mixing the standard base T and the standard base C in a ratio of 2:1, and the same applies to other cases;
[0070] The step (2) is specifically as follows:
[0071] (2.1) Amplifying the synthesized composite oligonucleotide pool, preparing the library, and performing second-generation high-throughput sequencing to obtain sequencing reads composed of standard bases;
[0072] (2.2) Figure 7 As shown, the primer sequence and address sequence of the sequencing read are identified, the unique address number of the sequencing read is determined, and all DNA sequences corresponding to the same primer and address sequence are obtained and aggregated into clusters;
[0073] (2.3) Align the r data-bearing region sequences within each cluster and count the frequencies of the four standard bases {A, T, G, C} at each different synthesis site i corresponding to the r sequences, denoted as: r i,A ,r i,T ,r i,G ,r i,C , where the observation frequency satisfies
[0074] (2.4) Figure 8 As shown, based on the base frequencies observed at each synthetic site, the frequency probability distribution p of the four standard bases is calculated. i =(p i,A ,p i,T ,p i,G ,p i,C ), and calculate the information entropy, and then divide different types of characters into different sets based on the separable characteristics of entropy;
[0075] (2.5) Within each set, the normalized maximum log-likelihood method is used to compare the observed frequencies with the candidate character set to accurately detect standard bases and composite bases. Each inferred character is then demapped into m bits in turn and concatenated into a block code of length n1 bits. A highly reliable error correction decoding method is then used to correct any remaining substitution and erasure errors, restoring the original digital file.
[0076] The step (1.4) is specifically:
[0077] (3.1) Select the address information of length k2 in the range [1, S] to represent the order information of S multi-ary sequences, where Then, the short block error correction code C2(n2,k2) is used to encode the address bits of length k2 to generate a short block error correction code codeword of length n2 bits, and a check bit that can represent the parity check is added to form an encoded address bit sequence of length L1;
[0078] (3.2) Group each two bits and map the encoded address bit sequence of length L1 into an address DNA sequence of length L2 according to the mapping rules: 00→A, 01→T, 10→G, 11→C, and add it to the 5' end of the data-bearing sequence;
[0079] (3.3) Generate a parity check bit based on the address bits of length k2, cyclically left-shift the original address bits of length k2 by one bit to obtain a left-shifted address bit sequence of length k2, and then follow the concatenation rule of "parity check bit - original address bit - parity check bit - left-shifted address bit" to obtain a coded address bit sequence of length L1.
[0080] (3.4) According to the same mapping rules as step (3.2), the encoded address bit sequence obtained in step (3.3) is converted into an address DNA sequence of length L2, which is added to the 3' end of the data-bearing sequence. Primer sequences for amplification are also added to both ends. The address sequence and primer sequences only use the standard four bases (A, T, G, C) to generate an oligonucleotide pool containing standard bases and composite bases.
[0081] The step (2.4) is specifically as follows:
[0082] (4.1) The base frequency distribution r observed based on the r data-carrying sequences in each cluster i , calculate the frequency probability distribution of the four standard bases, expressed as p i =(p i,A ,p i,T ,p i,G ,p i,C ), then based on the observed frequency distribution, the information entropy at each fixed point is calculated, which can be expressed as:
[0083]
[0084] Among them, the frequency probability distribution satisfies p i,α represents the probability of character α appearing at synthesis position i in the r data-carrying sequence;
[0085] (4.2) Based on the separable property of information entropy, the set segmentation method is adopted. According to the information entropy results at each site, the three character types of standard bases, composite bases, and composite bases with different proportions are classified. And when m=3, all characters are divided into three character sets according to the information entropy results, namely, the standard base character set S1={A,T,G,C}, the two-base hybrid character set S2={Y,M,K,R,S,W} and the three-base hybrid character set S3={H,B,V,D}.
[0086] The step (2.5) is specifically as follows:
[0087] (5.1) In each character set, according to the observed frequency distribution r i , and then the standard vectors σ=(σ A ,σ T ,σ G ,σ C ), calculate the log-likelihood function l(σ,r i ), which is used to infer the character in the set that is closest to the encoded character, and can be expressed as:
[0088]
[0089] (5.2) The log-likelihood function obtained in step (5.1) is normalized, and the character with the smallest result is regarded as the detected character. Each detected character is then mapped to m bits, and a highly reliable error correction decoding algorithm is executed to correct the remaining substitution errors and erasure errors and restore the original digital file. Specific embodiments
[0091] Specific examples are given below to illustrate the feasibility of the multi-base encoding and readout method for composite base DNA storage provided by the present invention.
[0092] This embodiment utilizes a multi-ary LDPC code defined in the Galois field GF(64) with strong error correction capability, the number of symbols is 3780, the degree of the variable node is 2, the error correction performance is excellent, and an octal symbol mapping is designed. Every 3 bits are mapped to an octal symbol and then mapped to a standard base or a composite base. In the design of the present invention, the conversion from quaternary encoding to octal encoding does not bring about the loss of bit and symbol conversion, the coding efficiency is 1, and the theoretical information capacity of each synthesis cycle is increased from 2 bits per base to 3 bits per base. Specifically, the present invention designs a NB-LDPC (22680, 7560) codeword defined in the Galois field GF(64) with a code rate of 1 / 3, and a digital information of size 7560 bits is added with 15120 bits of redundant bits to generate a long codeword. Then, according to the rule of mapping every 3 bits to 1 character, 7560 coding characters containing standard bases and composite bases are obtained.
[0093] Furthermore, the present invention divides the designed 7560 coded characters into 126 sequence segments of 60 characters in length, uses 8 types of nucleotides including standard bases and composite bases to represent polynomial sequences, adds address sequences and primer sequences at both ends of the 126 polynomial sequences, and generates an oligonucleotide pool including standard bases and composite bases.
[0094] The specific steps of the address sequence designed in this embodiment are as follows:
[0095] (1.1) Select address information ranging from 1 to 126 with a length of 7 bits to represent the sequence information of the designed 126 multi-ary sequences, where 2 7 >126, then use the short block error correction code convolutional code BCH (15, 7) to encode the 7-bit address information to generate a 15-bit short block error correction code codeword, and further add a check bit that can represent the parity check to the codeword to form a 16-bit coded address bit sequence;
[0096] (1.2) Group each two bits and map the 16-bit coded address bit sequence into an 8-base address DNA sequence according to the mapping rules: 00→A, 01→T, 10→G, 11→C, and add it to the 5' end of the data-bearing sequence;
[0097] (1.3) Generate a parity check bit based on the original 7-bit address information. Then, cyclically shift the original 7-bit address bits left by one bit to obtain a 7-bit address information after the left shift. Then, according to the concatenation rule of "parity bit - original address bit - parity bit - left-shifted address bit", obtain a 15-bit encoded address bit sequence.
[0098] (1.4) According to the same mapping rules as step (1.2), the encoded address bit sequence obtained in step (1.3) is converted into an 8-base address DNA sequence, which is added to the 3' end of the data-bearing sequence. A 20-base primer sequence for amplification is added to both ends to generate an oligonucleotide pool containing standard bases and composite bases.
[0099] The present invention synthesizes an oligonucleotide pool comprising standard bases and composite bases using a high-throughput parallel synthesis method, and realizes the synthesis of composite bases by adjusting the reagent ratio of standard DNA and adopting fuzzy synthesis. In the specific scheme, the present invention designs two kinds of design schemes of composite bases according to different standard DNA nucleotide ratios, as shown in Table 1. Scheme 1 is a composite base constructed by two standard bases in a ratio of 1: 1, and Scheme 2 is a composite base constructed by three standard bases in a ratio of 1: 1: 1. Six kinds of composite bases (R / Y / M / K / S / W) are designed in Scheme 1, which constitute an octal model with the standard base (A / T); in Scheme 2, the composition of the composite base is adjusted, and the composite base (H / B / V / D) is constructed, which constitutes an octal model with the standard base (A / T / G / C).
[0100] The present invention uses a synthetic oligonucleotide pool as a storage medium, undergoes multiple rounds of amplification reactions, library construction, and second-generation double-end sequencing to generate sequencing reads of 150 nucleotides in length. The method of the present invention for recovering data from double-end sequencing reads is specifically as follows:
[0101] (2.1) Splice the double-ended sequencing reads and align them with the designed 20-nucleotide double-ended primers to select the sequencing reads with the correct primers. Then, based on the boundary positions of the double-ended primers, intercept the address sequence part and the data-carrying sequence part.
[0102] (2.2) Compare the double-ended address sequence with the designed 16-nucleotide address DNA sequence obtained by combining the two-end 8-nucleotide address sequences. Filter the sequencing reads with the correct address sequence. Based on the boundary position of the address sequence, intercept the data-carrying sequence portion and cluster the reads to obtain multiple copies of sequencing reads for each specific sequence. Figure 9 The ratio of primer identification to address identification is shown;
[0103] (2.3) Align the r data-carrying sequences in the cluster corresponding to each address sequence, and count the frequency distribution of the four standard bases at each fixed site r i (r i,A ,r i,T ,r i,G ,r i,C ),in, Figure 10The length distribution of sequencing reads after primer truncation is shown. Figure 11 The results of the distribution of the number of copies within the cluster after read segment clustering after label identification are shown.
[0104] Table 1 Types of designed composite bases and their corresponding relationships with standard bases
[0105]
[0106] Furthermore, based on the observed base frequency distribution, a set segmentation method is used to classify different types of characters. Then, a normalized log-likelihood method is used to distinguish different characters within each character set, reconstructing standard bases and composite bases. Furthermore, each reconstructed character is mapped to 3 bits, and highly reliable error correction decoding is performed to correct remaining substitution errors and erasure errors, restoring the original digital file. The specific steps are as follows:
[0107] (3.1) The base frequency distribution r observed based on the r data-carrying sequences in each cluster i , calculate the frequency probability distribution of the four standard bases, expressed as p i =(p i,A ,p i,T ,p i,G ,p i,C ),in p i,α Indicates the probability of character α appearing at the synthesis position i in the r data-carrying sequence; then calculates the information entropy at each fixed position The information entropy results are as follows Figure 12 As shown;
[0108] (3.2) Using the set segmentation method, according to the information entropy results at each site, the three character types are classified into standard bases, composite bases, and composite bases with different ratios. When m = 3, all characters are segmented into three character sets based on the information entropy results, namely, the standard base character set S1 = {A, T, G, C}, the two-base hybrid character set S2 = {Y, M, K, R, S, W}, and the three-base hybrid character set S3 = {H, B, V, D};
[0109] (3.3) In each character set, according to the observed frequency distribution r i , and then the standard vectors σ=(σ A ,σ T ,σ G ,σ C ) Calculate the log-likelihood function Used to infer the character in the set that is closest to the encoded character;
[0110] (3.4) Normalize the log-likelihood function obtained in step (3.3) and regard the character with the smallest result as the reconstructed character. Figure 13 The normalized log-likelihood results of the characters "K" and "H" are displayed. Each reconstructed character is then mapped to 3 bits, and a highly reliable error correction decoding algorithm is performed to correct the remaining substitution errors and erasure errors and restore the original digital file.
[0111] Figure 14 The figures show the number of missing bits and erroneous bits for two example storage schemes at different sequencing coverages. At a coverage of 10x, the missing error rate is as high as 10%, which is beyond the erasure correction capability of the LDPC code designed by the present invention. At this coverage, the character derivation error rate of the three-base mixed coding scheme is significantly higher than that of the composite base coding scheme composed of a mixture of two bases. As the sequencing coverage increases, the erasure error and substitution error gradually decrease, and at a sequencing coverage depth of 50x, the erasure error and substitution error are already as low as 1% or even lower.
[0112] Those skilled in the art will understand that the accompanying drawings are only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0113] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-base encoding and readout method for composite base DNA storage, characterized in that: The method comprises the following steps: (1) At the composite base data writing end, a block error correction code is used to encode the user data as a whole to obtain a codeword. The codeword is converted into a multi-base symbol sequence according to a mapping method in which each m bits corresponds to a multi-base symbol d. The codeword is then divided into short segments. Standard base letters and composite base letters corresponding to the multi-base symbols are selected to obtain a DNA sequence containing standard bases and composite bases. A highly reliable address sequence and primer sequence are added at both ends to obtain a composite base oligonucleotide pool for storing data. (2) At the composite base data readout end, the second-generation high-throughput sequencing converts the composite bases of each synthesis site into sequencing reads composed of standard bases. Based on the primer sequence and address sequence, the unique position of the sequencing read is determined and clustered. Then, the frequency of base observations in each address cluster is counted and converted into information entropy. Based on entropy separability, the characters of each synthesis site are divided into subsets. Within each character set, the maximum likelihood is calculated for the observation frequency and the candidate character set, and the standard bases and composite bases are detected. Finally, the inferred characters are sent to the error correction decoder to correct the remaining substitution errors and erasure errors and restore the original data.
2. The multi-base encoding and reading method for composite base DNA storage according to claim 1, characterized in that: The step (1) is specifically as follows: (1.1) Using a binary or multi-bit block error correction code long code C1(n1,k1), the digital file of length k1 bits is encoded as a codeword of length n1 bits. Both the binary and multi-bit encoding processes are calculated on a bit basis, where k1 and n1 are positive integers representing the length of information bits and the length of code bits, respectively. (1.2) Divide each m bits of the code word into a group and map it into a multi-ary symbol d according to a predetermined mapping rule, where the multi-ary symbols have a total of 2 m The multi-ary symbol sequence of length n1 / m is obtained, and then divided into S symbol sequences of length L, corresponding to S payload DNA sequences, where m is a positive integer not less than 3, d∈[0,2 m -1], S*L=n1 / m; (1.3) According to the mapping relationship between the predetermined symbols and DNA bases, the standard base set {A, T, G, C}, the composite base {Y, M, K, R, S, W, R, B, V, H, D, N} and the composite bases {Y i ,M i ,K i ,R i ,S i ,W i ,R i ,B i ,V i ,H i ,D i ,N i }, i represents a variety of composite bases mixed with different proportions of standard bases, select 2 m The letters are used to represent the encoded multi-base symbol sequence to obtain the payload DNA sequence, and the payload DNA sequence contains standard bases and composite bases. The vector representing the mixing ratio of all four standard DNA nucleotides in the composite base synthesis process can be expressed as σ=(σ A ,σ T ,σ G ,σ C ), the vector σ satisfies σ A +σ T +σ G +σ C =1; (1.4) Adding address sequences and primer sequences to both ends of S data-carrying sequences of length L to generate an oligonucleotide pool containing standard bases and composite bases, wherein the address sequences and primer sequences are generated from the four standard bases {A, T, G, C}; (1.5) synthesizing each site of each information sequence in sequence using DNA synthesis according to the ratio of the standard bases corresponding to the composite bases or a single standard base, and mixing all the synthesized sequences into a pool to obtain a composite oligonucleotide pool; Wherein, the step (1.5) is specifically as follows: According to the vector σ=(σ A ,σ T ,σ G ,σ C ), synthesize each site in turn, for the standard base, A=(1,0,0,0), T=(0,1,0,0), G=(0,0,1,0), C=(0,0,0,1), composite base Indicates that the composite base Y is synthesized by mixing the standard base T and the standard base C in the same proportion. It is synthesized by mixing the standard base T and the standard base C in a ratio of 2:
1.
3. The multi-base encoding and reading method for composite base DNA storage according to claim 1, characterized in that: The step (2) is specifically as follows: (2.1) Amplifying the synthesized composite oligonucleotide pool, preparing the library, and performing second-generation high-throughput sequencing to obtain sequencing reads composed of standard bases; (2.2) Identify the primer sequence and address sequence of the sequencing read, determine the unique address number of the sequencing read, obtain all DNA sequences corresponding to the same primer and address sequence, and aggregate them into a cluster; (2.3) Align the r data-bearing region sequences within each cluster and count the frequencies of the four standard bases {A, T, G, C} at each different synthesis site i corresponding to the r sequences, denoted as: r i,A ,r i,T ,r i,G ,r i,C , where the observation frequency satisfies (2.4) Based on the base frequencies observed at each synthesis site, calculate the frequency probability distribution p of the four standard bases i =(p i,A ,p i,T ,p i,G ,p i,C ), and calculate the information entropy, and then divide different types of characters into different sets based on the separable characteristics of the information entropy; (2.5) Within each set, the normalized maximum log-likelihood method is used to compare the observed frequencies with the candidate character set to accurately detect standard bases and composite bases. Each inferred character is then demapped into m bits in turn and concatenated into a block code of length n1 bits. A highly reliable error correction decoding method is then used to correct any remaining substitution and erasure errors, restoring the original digital file.
4. The multi-base encoding and reading method for composite base DNA storage according to claim 2, characterized in that: The step (1.4) is specifically as follows: (3.1) Select the address information of length k2 in the range [1, S] to represent the order information of S multi-ary sequences, where Then, the short block error correction code C2(n2,k2) is used to encode the address bits of length k2 to generate a short block error correction code codeword of length n2 bits, and a check bit representing the parity check is added to form an encoded address bit sequence of length L1; (3.2) Group each two bits and map the encoded address bit sequence of length L1 into an address DNA sequence of length L2 according to the mapping rule MR = 00 → A, 01 → T, 10 → G, 11 → C, and add it to the 5' end of the data-bearing sequence; (3.3) Generate a parity check bit based on the address bits of length k2. Circularly left-shift the original address bits of length k2 by one bit to obtain a left-shifted address bit sequence of length k2. Then, follow the concatenation rule of "parity check bit - original address bit - parity check bit - left-shifted address bit" to obtain a coded address bit sequence of length L1. (3.4) According to the mapping rule MR, the encoded address bit sequence obtained in step (3.3) is converted into an address DNA sequence of length L2, which is added to the 3' end of the data-bearing sequence. Primer sequences for amplification are also added at both ends. The address sequence and primer sequences use the standard four bases (A, T, G, C) to generate an oligonucleotide sequence containing standard bases and composite bases.
5. The multi-base encoding and reading method for composite base DNA storage according to claim 3, characterized in that: The step (2.4) is specifically: (4.1) The base frequency distribution r observed based on the r data-carrying sequences in each cluster i , calculate the frequency distribution of the four standard bases, expressed as p i =(p i,A ,p i,T ,p i,G ,p i,C ), then based on the observed frequency distribution, the information entropy at each fixed point is calculated, expressed as Frequency distribution satisfies p i,α represents the probability of character α appearing at synthesis position i in the r data-carrying sequence; (4.2) Based on the information entropy, the set segmentation method is used to classify the three character types of standard bases, composite bases, and composite bases with different proportions according to the information entropy results at each site; when m = 3, all characters are divided into three character sets according to the information entropy results, namely the standard base character set S1 = {A, T, G, C}, the two-base hybrid character set S2 = {Y, M, K, R, S, W} and the three-base hybrid character set S3 = {H, B, V, D}.
6. The multi-base encoding and reading method for composite base DNA storage according to claim 3, characterized in that: The step (2.5) is specifically as follows: (5.1) In each character set, according to the observed frequency distribution r i , and then the standard vectors σ=(σ A ,σ T ,σ G ,σ C ) calculates the log-likelihood function l(σ,r i ) Estimate the character in the set that is closest to the encoded character. The log-likelihood ratio function is expressed as (5.2) The log-likelihood function l(σ,r i ) is normalized, and the character with the smallest result is regarded as the detected character. Each detected character is then mapped to m bits, and a highly reliable error correction decoding algorithm is executed to correct the remaining substitution errors and erasure errors and restore the original digital file.
Citation Information
Patent Citations
DNA storage coding method and device based on decimal system, and readable storage medium
CN114974429A
Space layering DNA storage method of large-scale oligonucleotide pool
CN118841094A