Gc content locally balanced dna storage encoding method

By combining quaternary Hamming codes and quaternary VT codes, a locally balanced DNA encoding method is constructed, which solves the problem of unstable GC content in DNA storage and achieves efficient error correction and low redundancy DNA storage encoding.

CN116405693BActive Publication Date: 2025-11-25SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310399619.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-11-25
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

Existing DNA storage encoding methods cannot achieve local balance of GC content under the premise of error correction, which makes DNA sequences prone to errors during storage.

Method used

An encoding method combining quaternary Hamming codes and quaternary VT codes is adopted. By constructing a set of codewords with globally stable GC content, and by using splicing technology to ensure the local stability of GC content of DNA fragments under each sliding window, Hamming codes and VT codes are combined to correct base substitution, insertion and deletion errors.

Benefits of technology

It achieves local balance of GC content in DNA sequences during storage, reduces the error rate, and can correct single base substitution, insertion, and deletion errors with low redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405693B_ABST
    Figure CN116405693B_ABST
Patent Text Reader

Abstract

The application discloses a DNA storage coding method with local GC content balance, which adopts a four-element sequence coding method for original data, and adopts a splicing method to splice a GC content globally balanced code word set for several times, so that a new four-element code satisfies the GC content local balance, so that the four-element DNA sequence composed of A, T, C and G can be more conveniently corresponded, and the base substitution error and the base insertion and deletion error can be respectively corrected by combining a Hamming code and a VT code. Furthermore, the GC content of the DNA coding can be finally ensured to be globally stable at 40% to 60% by selecting the code word satisfying the GC content in the Hamming code, and the DNA segment in each fixed length sliding window satisfies the GC content stable at 30% to 70%.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information coding technology of distributed storage, in particular to a DNA storage coding method for local balance of guanine-cytosine (GC) content in DNA sequence. BACKGROUND

[0002] DNA storage technology refers to converting data information into a DNA sequence composed of four bases of adenine (A), guanine (G), cytosine (C) and thymine (T). In order to ensure the reliability of data storage, an encoding method suitable for DNA storage system needs to be designed. On the one hand, the error rate of the encoded DNA sequence is required to be low in the process of DNA synthesis, storage and sequencing. On the other hand, the encoded DNA sequence is required to be able to correct the substitution, insertion and deletion errors of bases after error. The error rate of the DNA sequence is related to the content of guanine (G) and cytosine (C) in the sequence. The GC content of the whole DNA sequence is stably around 50% (referred to as global balance of GC content), which is beneficial to reduce the error rate. Further, for each fixed length sliding window, the GC content of the DNA sub-fragment intercepted by the sliding window is stably within the interval of around 50% (referred to as local balance of GC content), and the DNA sequence will be more stable and less prone to error.

[0003] The current DNA storage coding method is divided into three categories. The first category directly uses classical error correction codes (RS codes, etc.) or insertion and deletion codes (Levenshtein codes, etc.) to correct base errors, but does not consider the constraint of GC content. The second category of coding gives a DNA code that satisfies the GC balance, but does not have error correction capability. The third category of code combines error correction properties and GC content constraints, and simultaneously achieves GC balance and error correction. However, these methods only satisfy the global balance of GC content, and do not satisfy the local balance of GC content. Therefore, under the premise of error correction, how to construct a DNA coding method with local balance of GC content is still a problem to be solved. SUMMARY

[0004] The present application proposes a DNA storage coding method with local balance of GC content, which can solve the limitation that the existing coding method cannot achieve local balance of GC content under the premise of error correction. The encoded DNA sequence satisfies that the GC content of each DNA sequence is globally stable at 40% to 60%, and for each fixed length sliding window, the corresponding DNA fragment maintains the GC content stable at 30% to 70%, which can greatly reduce the error rate of DNA sequence in storage, and also has the error correction capability of resisting substitution, insertion and deletion errors of single base.

[0005] The present application is realized by the following technical scheme:

[0006] This invention relates to a method for storing and encoding DNA with locally balanced GC content, comprising the following steps:

[0007] Step 1: Obtain raw data information: Convert the binary information sequence of the video, image, and other data to be stored into a quaternion sequence for encoding, with each group consisting of two bits.

[0008] Step 2: Quaternary Hamming code encoding: Constructing a finite field containing four elements Where: finite field The generator α satisfies: α 2 +α+1=0; Encode the quaternion sequence to be encoded in step 1 using quaternion Hamming codes, resulting in a codeword set Ham(r, 4). Then... Wherein: the code length of the Hamming code Dimension Minimal Hamming distance d H =3, r is a positive integer representing the codimensional of the quaternary Hamming code.

[0009] The Hamming code described above can correct a one-bit substitution error.

[0010] Step 3: Establish a finite field The mapping relationship to the DNA base set {A, T, C, G} is: 0→C, 1→G, α→A, α 2 →T, such that the finite field The GC content of the upper quaternion sequence is equivalent to the content of 0 and 1 in it.

[0011] Step 4: Obtain codewords in the codeword set Ham(r, 4) whose GC content is globally stable between 40% and 60%, and form a set. Specifically, it includes:

[0012] 4.1) Select one of Ham(r, 4) A parity check matrix H of order r is given, such that the first r columns are linearly independent.

[0013] 4.2) Extract all codewords c = (c1, ..., c4) from Ham(r, 4). l ) is divided into: c = (c I c A c B ), where: c I = (c1, ... c) r ), c A =(c r+1 , ...c r+m ), c B =(c r+m+1 c l ), m is a positive integer between 1 and lr; according to H·cT =0, when c A c B When determined, c I That is the only certainty.

[0014] 4.3) Select codewords from Ham(r, 4) to form a set. Specifically: c A component c r+1 c r+m The value can be 0 or 1, c B component c r+m+1 c l The value can be α or α 2 c I The component belongs to gather The number of 0s and 1s contained therein ranges from m to m+r.

[0015] 4.4) Calculate the set Number of code words To ensure that the GC content of the codewords in the set selected in step 4.3) is globally stable between 40% and 60%, where the value of m satisfies 0.4l≤m≤0.6lr.

[0016] Step 5: For the set Further encoding yields a quadruple code with locally balanced GC content. Specifically: Select a positive integer t≥8, and for... Perform t concatenations and augmentations to obtain the quaternion code. in: The code length n = tl, the number of codewords

[0017] Step 6: Select Quaternion Code The local parameter l′ = 8(l-1), where l′ represents the sliding window length, and under a sliding window of arbitrary length l′, the GC content of the DNA fragment remains between 30% and 70%, ensuring... All codewords satisfy GC content local balance, specifically including:

[0018] 6.1) Translate the quadruple code Any codeword c is divided into t blocks: c = (c (1) c (2) c (t) ), where: each block c (i) =(c i,1 c i,l )belong

[0019] 6.2) Consider a sliding window of arbitrary length l', denote the local segment of the codeword c under this window as Since l' > 2l, contains at least two consecutive blocks; when contains s consecutive blocks: c (i+1) ,..., c (i+s) , and the last a components of block c (i) and the first b components of block c (i+s+1) , then where: s≥2, 0≤a, b≤l-1.

[0020] 6.3) Since the GC content of each block c (i) is 40% to 60%, the number of 0s and 1s in the local segment ranges from 0.4sl to 0.6sl+a+b.

[0021] 6.4) Calculate the GC content of the local segment to where: That is, the GC content of is 30% to 70%, so all codewords of the quaternary code satisfy that the GC content of the DNA segment under the sliding window of arbitrary length l' is stable at 30% to 70%, which is a locally balanced quaternary code with GC content, and can greatly reduce the error rate in storage.

[0022] Step 7: Establish a mapping relationship from the finite field to the quaternary set {0, 1, 2, 3}: 0→0, 1→1, α→2, α 2 →3, which isomorphically converts the quaternary code to the quaternary code over {0, 1, 2, 3}

[0023] Step 8: Encode the quaternary code using the quaternary VT code: Select integers a and b as the check parameters of the quaternary VT code, where: 0≤a≤n-1, 0≤b≤3, encode using the (a, b) type quaternary VT code, and obtain the quaternary DNA code

[0024] Step 9: According to the mapping relationship: 0→C, 1→G, 2→A, 3→T, convert the quaternary sequence in the quaternary DNA code to a DNA sequence composed of A, T, C, G for storage, and the sequence length is n.

[0025] Step 10: Calculate the number of codewords of the quaternary DNA code and the redundancy.

[0026] Preferably, the global stability of GC content in step 4 is 40% to 60%, which means that for a quaternary vector x of length l belonging to The number of 0s and 1s in x is 0.4l to 0.6l.

[0027] Preferably, after the block of codewords in step 4.2), c I corresponds to the first r columns of the check matrix H, c A corresponds to the r+1th column to the r+mth column of H, c B corresponds to the r+m+1th column to the lth column of H. According to the selection of H, the first r columns are linearly independent, so c l is uniquely determined. A , c B is uniquely determined.

[0028] Preferably, the selected codewords in step 4(3) have components with values of 0 or 1 in c A and c I , where: the components of c A all have values of 0 or 1, and the number of components of c I with values of 0 or 1 is at least 0 and at most r. Therefore, the number of components with values of 0 and 1 in these codewords is m to m+r.

[0029] Preferably, the calculation of the number of codewords of in step 4(4) is achieved by the following method:

[0030] Take m components from the last l-r components of the codeword vector to form c A , and the components of c A belong to , the number of which is . The remaining l-r-m components form c B , and the components of c B belong to , the number of which is l-r-, . According to step 4.2), when c A and c B are determined, the entire codeword vector is uniquely determined. According to step 4(3), the number of 0s and 1s in the codeword c is N, which satisfies m≤N≤m+r, in order to ensure that the GC content of c is globally stable at 40% to 60%, then 0.4l≤m≤N≤m+r≤0.6l, that is, 0.4l≤m≤0.6l-r. The size of the set is ​

[0031] Preferably, each codeword of the quaternary code in step 5 contains t blocks, and each block belongs to Then each codeword of the quaternary code satisfies that the GC content is globally stable in the range of 40% to 60%.

[0032] Preferably, the local parameter l' of the quaternary code in step 6 represents a sliding window length satisfying the GC content locally balanced, i.e. for a quaternary vector x of length n, the consecutive sub-segments of x intercepted by any sliding window of length l' belong to Then the number of 0s and 1s in is in the range of 0.3l' to 0.7l'.

[0033] Preferably, the selection rule of the check parameters a and b of the quaternary VT code in step 8 includes:

[0034] ① For a quaternary sequence x = (x1,..., xn) ∈ {0, 1, 2, 3} n , construct its binary companion sequence as n wherein: a1 takes the value of 0, for 2≤i≤n, when x i ≥ x i-1 , a i takes the value of 1; when x i < x i-1 , a i takes the value of 0.

[0035] ② For According to the pigeonhole principle, there exist integers a and b such that: By traversing all integers a between 0 and n-1 and integers b between 0 and 3, a set of integers a and b satisfying can be selected.

[0036] Preferably, the DNA code obtained in step 8 satisfies the GC content locally balanced. All codewords of the DNA code are codewords in the quaternary code , thus satisfying the GC content locally balanced.

[0037] Preferably, the number of codewords of the DNA code obtained in step 10 is and and wherein: the DNA code Redundancy Where n = tl is the code length of the DNA code. The redundancy is finally calculated as R = log2 n + O(1) bits.

[0038] This invention relates to a DNA decoding method based on the above-mentioned encoding and local equilibrium of GC content, comprising:

[0039] Step a: Decoding process: The DNA sequence after DNA sequencing is processed according to the sequence C→0, G→1, A→2, T→3.

[0040] Convert to quaternary sequence y = (y1, ..., y2) n′ )∈{0,1,2,3} n′ The sequence length is n′.

[0041] Step b: When n′=n, a one-bit substitution error has occurred: Rewrite y in the form of t blocks: y=(y (1) y (2) , ..., y (t) If at most one block in y has a one-bit substitution error, then for each block in y... (1) , ..., y (t) The Hamming code decoder is used to decode the code, and the original Hamming codeword quaternion sequence is obtained, which is then restored to binary data for reading.

[0042] Step c: When n′ = n+1 or n-1, a bit insertion or deletion error is determined: using the decoding algorithm of the quaternary VT code, the quaternary sequence y is decoded to obtain the original VT code codeword c of length n. The t blocks of c are all error-free Hamming codewords, which are further restored to binary data for reading.

[0043] Preferably, in step b, the DNA code It can correct single base substitution errors. It uses a Hamming code decoding algorithm for decoding, with a decoding complexity of O(nlogn), where n is the code length.

[0044] Preferably, in step c, the DNA code It can correct single base insertion or deletion errors. It uses a quaternary VT code decoding algorithm for decoding, with a decoding complexity of O(n), where n is the code length.

[0045] Technical effect

[0046] Compared with the prior art using a single error correction code (Reed-Solomon code) or insertion and deletion code (Levenshtein code) correction, the present application uses quaternary sequence encoding of the original data, which can more conveniently correspond to the quaternary DNA sequence composed of A, T, C and G, and corrects base substitution errors and base insertion and deletion errors by combining Hamming code and VT code. Further, by selecting the code word in the Hamming code that satisfies the GC content, the GC content of the DNA code can be finally ensured to be globally stable at 40% to 60%, laying a foundation for constructing a quaternary DNA code with locally balanced GC content. Secondly, the present application uses a splicing method to splice the code word set with globally balanced GC content to obtain a new quaternary code that satisfies the locally balanced GC content, and the DNA fragment under each fixed length sliding window satisfies the GC content stable at 30% to 70%. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 The flowchart of the present application. DETAILED DESCRIPTION

[0048] As Figure 1 shown, the present embodiment relates to a DNA storage encoding method satisfying locally balanced GC content, including the following steps:

[0049] Step 1: Obtain original data information. The binary information sequence of the data to be stored, such as video and image, is converted into a quaternary sequence to be encoded, one group for every two bits.

[0050] Step 2: Quaternary Hamming code encoding. A finite field containing four elements is constructed, wherein: α satisfies α 2 +α+1=0 is the generator of the finite field . The redundancy number r of the quaternary Hamming code is selected to be 3. The quaternary sequence to be encoded in step 1 is encoded using the quaternary Hamming code, and the code word set after encoding is Ham(3, 4). Then wherein: the code length l of the Hamming code is 21, the dimension k is 18, and the minimum Hamming distance d H =3 to correct one-bit substitution error.

[0051] Step 3: Establish the mapping relationship of the finite field to the DNA base set {A, T, C, G}: 0→C, 1→G, α→A, and α 2 →T. Then the GC content of the above quaternary sequence is equivalent to the content of 0 and 1.

[0052] Step 4: Obtain the codewords in Ham(3,4) whose GC content is globally stable between 40% and 60%, and form a set Specifically comprising:

[0053] (1) Select a check matrix H of Ham(3,4), H contains 3 rows and 21 columns, so that the first 3 columns are linearly independent.

[0054] (2) Divide all the codewords c=(c1,...,c l ) in Ham(3,4) into three blocks as follows: c=(c I , c A , c B ), where: c I =(c1, c2, c3), c A =(c4,...c m ), c B =(c m+1 ,...,c 21 ), where: m is a positive integer taken from 4 to 20. According to H·c T =0, when c A , c B are determined, c I is uniquely determined.

[0055] (3) Select the codewords from Ham(3,4) to form The components c4,...,c m of c A take values 0 or 1, the components c B ,...,c m+1 of c 21 take values α or α 2 , and the components of c I belong to The components of each codeword that take values 0 or 1 are in c A and c I , so the number of 0 and 1 contains m to m+3.

[0056] (4) Ensure that the GC content of the codewords selected in (3) is globally stable between 40% and 60%. Here, as long as the number of 0 and 1 N in the codeword satisfies: 8.4≤m≤N≤m+3≤12.6, i.e. m=9. Calculate the number of codewords

[0057] Step 5: Further encode to obtain a quaternary code with locally balanced GC content Select a positive integer t=8, and perform 8 times of splicing expansion on to obtain a quaternary code​ in: The code length n = 21t = 168, the number of codewords Each codeword consists of 8 The codewords in the code are composed of GC content, which is therefore globally stable at 40% to 60%.

[0058] Step 6: Select Quaternion Code The local parameter l′ = 160, where l′ represents the sliding window length. Ensure... All codewords satisfy local equilibrium of GC content: within a sliding window of arbitrary length 160, the GC content of the DNA fragment remains between 30% and 70%. Specifically, this includes:

[0059] (1) Any codeword c consists of 8 blocks: c = (c (1) c (2) c (8) ), where: each block c (i) =(c i,1 c i,l )belong

[0060] (2) Select any local segment of codeword c with a length of l′ = 160. Since l′>2l=42, It contains at least two consecutive blocks. Where: s≥2, 0≤a, b≤20.

[0061] (3) Due to each block c (i) If the GC content is 40% to 60%, then the local fragment The number of 0s and 1s is between 0.4sl and 0.6sl+a+b, where l = 21.

[0062] (4) Calculate local segments GC content. The GC content is to in: Right now The GC content ranges from 30% to 70%. Therefore... All codewords are available. DNA fragments under any sliding window of length l′ = 160 satisfy the GC content being stable between 30% and 70%, which is a quadruple code with locally balanced GC content, which can greatly reduce the error rate in storage.

[0063] Step 7: Establish a finite field The mapping relationship to the four-element set {0, 1, 2, 3} is: 0→0, 1→1, a→2, a 2 →3. Then the four-element code is isomorphically converted into the four-element code on {0, 1, 2, 3}

[0064] Step 8: encoding using the four-element VT code. The four-element DNA code is obtained by encoding using the (0, 0) type four-element VT code, that is, wherein the four-element DNA code is a subcode of , the GC content of each code word is globally stabilized at 40% to 60%, and the GC content of the DNA fragment under any l' = 160 length sliding window is locally stabilized at 30% to 70%, which is a four-element code with locally balanced GC content.

[0065] Step 9: according to the mapping relationship: 0→C, 1→G, 2→A, 3→T, the quaternary sequence in the four-element DNA code is converted into a DNA sequence composed of A, T, C, G for storage, and the sequence length is n = 168.

[0066] Step 10: calculate the code word number and redundancy of . The code word number of is wherein n = 168, t = 8. The redundancy of is bits.

[0067] Step 11: decoding process. According to the relationship C→0, G→1, A→2, T→3, the DNA sequence after DNA sequencing is converted into a quaternary sequence y = (y1,..., y n′ ) ∈ {0, 1, 2, 3} n′ , and the sequence length is n', and the size relationship between n' and 168 is compared.

[0068] Step 12: if n' = 168, it is determined that a one-bit substitution error occurs. y is written in the form of 8 blocks: y = (y (1) , y (2) ,..., y (8) ), and at most one block in y has a one-bit substitution error. y (1) ,..., y (8) are decoded using the Hamming code decoder to obtain the four-element sequence of the original Hamming code word, and then the binary data is read.

[0069] Step 13: If n' = 169 or 167, it is determined that one bit insertion or deletion error occurs. The quaternary VT code decoding algorithm is used to decode the quaternary sequence y to obtain the original VT code word c with a length of 168. The 8 sub-blocks of c are all error-free Hamming code words, which are further restored into binary data for reading.

[0070] The global GC content of the encoded DNA sequence of the embodiment is stably maintained at 40% to 60%, and the GC content of the DNA segment under each window with a fixed length of 160 is locally stably maintained at 30% to 70%, which can greatly reduce the error rate of the DNA sequence in synthesis and sequencing. According to the decoding algorithm of the Hamming code and the VT code, the DNA encoding method can resist one bit substitution or insertion or deletion error. Through simulation experiments on the MATLAB 2021 platform, for the case of r = 2 and t = 8 in the embodiment, about bytes of data are stored, and about bytes of redundant data are required, and the redundancy ratio is 22.6%. It is stated again that the embodiment is only used as an example of the general method, and in fact, the Hamming code redundancy dimension r is a randomly generated positive integer, which is greater than or equal to 8, and the splicing parameter t is any positive integer greater than or equal to 8.

[0071] The application sequentially uses the quaternary Hamming code, splicing encoding and quaternary VT code to encode the quaternary data sequence of the original information to be stored, and constructs a DNA encoding scheme for correcting single base insertion, deletion and substitution errors. The scheme ensures that the global GC content of the DNA sequence is stably maintained at 40% to 60%, and the GC content of the DNA segment under each window with a fixed length is stably maintained at 30% to 70%, which meets the local balance of GC content and can greatly reduce the error rate of the DNA sequence in synthesis and sequencing. The application processes the quaternary Hamming code word by block processing to find the code word space with a global GC content of 40% to 60%. Subsequently, the code word space satisfying the global balance of GC content is spliced several times by using the splicing technology, and then the quaternary code with a local balance of GC content is obtained, so that the GC content of the DNA segment under each fixed length sliding window can be stably maintained at 30% to 70%. At the same time, the minimum Hamming distance of the quaternary code is greater than or equal to 3, which ensures that one bit substitution error can be corrected. Finally, the quaternary code is encoded by using the quaternary VT code to obtain the quaternary DNA code, so that the code can also correct one bit insertion and deletion error.

[0072] ​​Compared with the previous coding scheme, the DNA coding method has the characteristics of local balance of GC content, and the GC content of the coded DNA sequence is globally stable at 40% to 60%, and for a fixed length sliding window, the GC content of the DNA fragment under each window is stable at 30% to 70%, which can greatly reduce the error rate of the DNA sequence in the storage process, and can correct one base substitution, insertion, deletion error. In addition, the coding scheme has low redundancy, which can save coding cost. The present application makes up for the limitation of the previous DNA coding method that cannot achieve error correction and local balance of GC content at the same time, and is a new DNA coding scheme.

[0073] The above specific embodiments can be adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present application, the protection scope of the present application is subject to the claims and is not limited by the above specific embodiments, and each implementation method within the scope is subject to the constraints of the present application.

Claims

1. A GC-content locally balanced DNA storage encoding method, characterized in that, It comprises the following steps: Step 1: obtaining original data information: converting binary information sequence of video and image data to be stored into quaternary sequence to be encoded, every two bits as a group; Step 2: Quaternary Hamming code encoding: constructing a finite field containing four elements Wherein: the generator element a of the finite field satisfies: a 2 + a + 1 = 0; encode the quaternary sequence to be encoded in step 1 using a quaternary Hamming code, and the code word set after encoding is Ham(r, 4); then Wherein: the code length of the Hamming code dimension The minimum Hamming distance d H = 3, r is a positive integer, representing the co-dimension of the quaternary Hamming code; Step 3: Establishing a finite field Mapping to the set of DNA bases {A, T, C, G}: 0→C, 1→G, a→A, a 2 →T, such that the finite field The GC content of the upper quartet sequence is equivalent to the content of 0 and 1 in it; Step 4: Obtain the codewords in the codeword set Ham(r, 4) whose GC content is globally stable at 40% to 60% to form the set Step 5: Set collection Further encoding to get GC content locally balanced quaternary code Specifically: select a positive integer t ≥ 8, and Do t times of splicing expansion to get quaternary code Wherein: The code length n = tl, the code word number Step 6: Selecting the quaternary code wherein: l' denotes the length of the sliding window and the GC content of the DNA fragments remains between 30% and 70% in the sliding window of arbitrary length l' to ensure that all codewords of the quaternary code satisfy the local balance of GC content. Step 7: Establishing Finite Fields Mapping to the four-element set {0, 1, 2, 3}: 0→0, 1→1, a→2, a 2 →3, converts the four-element code isomorphically to the four-element code Step 8: encoding the quaternary code using the quaternary VT code Encoding: select integers a and b as the check parameters of the quaternary VT code, where: 0≤a≤n-1, 0≤b≤3, and encode the quaternary DNA code using the (a, b) type quaternary VT code Encoding, to obtain a quaternary DNA code Step 9: Convert the quaternary sequence in the quaternary DNA code into a DNA sequence composed of A, T, C, G according to the mapping relationship: 0→C, 1→G, 2→A, 3→T, and store the sequence with a length of n; Step 10: Calculate the number of codewords and redundancy of the quaternary DNA code of the quaternary DNA code.

2. The GC content locally balanced DNA storage encoding method of claim 1, wherein, The step 4 specifically comprises: 4.1) Select one of Ham(r, 4) whose parity check matrix H has the first r columns linearly independent. 4.2) Select one of Ham(r, 4) whose parity check matrix H has the first r columns linearly dependent. 4.2) Extract all codewords c = (c1, ..., c4) from Ham(r, 4). l ) is divided into: c = (c I c A c B ), where: c I = (c1, ... c) r ), c A =(c r+1 , ...c r+m ), c B =(c r+m+1 c l ), m is a positive integer between 1 and lr; according to H·c T =0, when c A c B When determined, c I That is the only certainty; 4.3) selecting the codewords from Ham(r,4) to form the set Specifically: c A The components c r+1 ,..., c r+m Take 0 or 1, c B The components c r+m+1 ,..., c l Take α or α 2 , c I The components belong to The number of 0 and 1 contained in the set m to m+r; 4.4) Calculate the set of codewords to ensure that the GC content of the codewords of the set selected in step 4.3) is globally stable between 40% and 60%, wherein: m takes a value such that 0.4l < m < 0.6l - r.

3. The GC content locally balanced DNA storage encoding method of claim 2, wherein, The number of code words of c is realized by taking m components from the last l-r components of the code word vector to form c A , and the components of c A belong to with the number of kinds; the remaining l-r-m components form c B , and the components of c B belong to with the number of l-r-m kinds; according to step 4.2), when c A and c B are determined, the entire code word vector is uniquely determined; according to step 4.3), the number N of 0 and 1 contained in the code word c satisfies m≤N≤m+r, in order to ensure that the GC content of c is globally stable at 40% to 60%, then 0.4l≤m≤N≤m+r≤0.6l, that is, 0.4l≤m≤0.6l-r; the size of the set 4. The GC content locally balanced DNA storage encoding method of claim 1, wherein, The step 6 specifically comprises: 6.1) divide any code word c of the quaternary code into t sub-blocks: c = (c (1) , c (2) ,..., c (t) ), where: each sub-block c (i) = (c i,1 ,..., c i,l ) belongs to 6.2) Consider a sliding window of arbitrary length l', let the local segment of the codeword c under this window be Since l' > 2l, contains at least two consecutive segments; when contains s consecutive segments: c (i+1) ,..., c (i+s) and the last a components of segment c (i) and the first b components of segment c (i+s+1) , then where: s > 2, 0 < a, b < l - 1; 6.3) Since the GC content of each chunk c (i) is 40% to 60%, the number of 0s and 1s in the local segment ranges from 0.4sl to 0.6sl + a + b; 6.4) Calculate local segment GC content to wherein: that is GC content of 30% to 70%, so all codewords of the quaternary code satisfy the GC content stability of 30% to 70% under the sliding window of any l' length, are the quaternary codes with local balance of GC content, and can greatly reduce the error rate in storage.

5. The GC content locally balanced DNA storage encoding method of claim 4, wherein, The local parameter l' of the step 6 quaternary code satisfies the GC content local balance, i.e. for a quaternary vector x of length n belonging to the set of quaternary vectors of length n satisfying the GC content local balance, the length of the sliding window l' is such that any consecutive sub-segment of x of length l' belongs to the set of quaternary vectors of length l' satisfying the GC content local balance then the number of 0 and 1 in x is comprised between 0.3l' and 0.7l'.

6. The GC content locally balanced DNA storage encoding method of claim 1, wherein, The selection rule of the check parameters a and b of the quaternary VT code in the step 8 comprises: n )∈{0,1,2,3} n , its binary companion sequence is constructed as wherein a1=0, and for 2≤i≤n, ai=1 if xi≥xi-1, and ai=0 if xi i i-1 i i i-1 i ​​​​​​​ ii. For According to the pigeonhole principle, there exist integers a and b such that: A set of integers a and b that satisfy can be chosen by iterating over all integers a between 0 and n-1 and integers b between 0 and 3.

7. The GC content locally balanced DNA storage encoding method of claim 1, wherein, The step 10 is based on the DNA code obtained in step 8 the number of code words and and wherein: the DNA code is calculated the redundancy wherein: n = tl is the code length of the DNA code; the redundancy R = log2 n + O(l) bits is finally calculated.

8. A GC content locally balanced DNA coding method based on the coding method of any one of claims 1-7, characterized in that, It comprises: Step a: decoding process: The DNA sequence after DNA sequencing is converted into a quaternary sequence y = (y1,..., yn) e {0, 1, 2, 3} of length n according to the rules C→0, G→1, A→2, T→3. n′ n , sequence length n'.​ Step b: When n' = n, then a single bit substitution error is determined to have occurred: write y in the form of t blocks: y = (y (1) , y (2) , ..., y (t) ), then at most one block of y has a single bit substitution error, decode y (1) , ..., y (t) using a Hamming code decoder to obtain the quaternary sequence of the original Hamming code word, and then restore it to binary data reading; Step c: when n'=n+1 or n-1, it is judged that one-bit insertion or deletion error occurs: the quaternary sequence y is decoded by using the decoding algorithm of the quaternary VT code to obtain the original VT code word c with a length of n; all the t blocks of c are error-free hamming code words, which are further restored into binary data for reading.

Citation Information

Patent Citations

  • Image feature quantification method, system and device and storage medium

    CN114332495A

  • DNA storage coding method based on Hamming-VT

    CN115242255A