A DNA coding method capable of correcting multiple insertion and deletion errors

By combining the DNA encoding method constructed by generalized Helberg code, the problem that existing DNA storage methods cannot effectively correct the multi-bit insertion and deletion errors is solved, and efficient correction of multi-bit errors is achieved, and data reliability and storage efficiency are improved.

CN116994659BActive Publication Date: 2025-06-13YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310972865.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-06-13
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

The existing DNA storage methods have limitations in correcting insertion and deletion errors, and cannot effectively correct multi-bit errors, resulting in limited reliability and efficiency of data storage.

Method used

A DNA encoding method combined with a generalized Helberg code is adopted. By defining a 4-element finite domain and a specific weight sequence, the code that can correct the error of multiple insertion and deduce the specific expression of ωn+1 in the generalized Helberg code to calculate the upper bound of the redundant value.

Benefits of technology

It realizes efficient correction of multiple insertion and deletion errors, improves the data reliability and storage efficiency of DNA storage, reduces redundancy, and thus improves the performance of DNA encoding methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994659B_ABST
    Figure CN116994659B_ABST
Patent Text Reader

Abstract

The invention discloses a DNA coding method capable of correcting multiple insertion and deletion errors, which relates to the DNA storage method in the field of data storage. Aiming at the problem of synchronization errors caused by insertion and deletion errors in DNA storage; by utilizing the properties of the generalized Helberg code, a DNA coding method capable of correcting multiple insertion and deletion errors is constructed; at the same time, the expression of ω n+1 in the generalized Helberg code is also deduced to facilitate the calculation of the upper bound of the redundancy value of this coding method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a DNA storage method in the field of data storage, and particularly to a novel DNA encoding method capable of correcting multiple insertion and deletion errors. Background Art

[0002] DNA-based data storage is a new method for long-term preservation of digital data. DNA storage means encoding an information sequence into a DNA sequence composed of four bases A, T, C, and G using DNA as a medium and endowing it with error correction capabilities. Due to significant advancements in biotechnology such as DNA synthesis and sequencing, DNA storage has attracted extensive attention in recent years for its advantages of high storage density, long preservation time, low maintenance cost, and easy access.

[0003] However, during the DNA storage process, base insertions, deletions, and substitutions (collectively referred to as insertion and deletion errors) often occur. Insertion and deletion errors are a type of synchronization error caused by the loss of position information in the message. The main purpose of DNA encoding is to introduce a certain amount of redundancy through some encoding methods so that insertion and deletion errors can be corrected. The present invention proposes a DNA encoding method capable of correcting multiple insertion and deletion errors. This encoding algorithm combines the construction method of generalized Helberg codes, where generalized Helberg codes can correct multiple insertion and deletion errors.

[0004] The performance of an encoding method is mainly determined by the code rate and error correction ability of the code. Among them, the code rate can be compared by calculating the redundancy. The lower the redundancy, the higher the code rate, and the better the performance of the code. A q-ary code with a code length of n and the number of codewords of the code The redundancy calculation formula is: In order to calculate the redundancy expression of this encoding method, the present invention also derived the specific expression of ω n+1 in the generalized Helberg code.

[0005] Currently, most DNA encoding methods can only correct one or two insertion and deletion errors, and there is still a lack of DNA encoding methods that can correct multiple insertion and deletion errors. Summary of the Invention

[0006] In view of the problem of synchronization errors caused by insertion and deletion errors in DNA storage, the present invention proposes a DNA encoding method capable of correcting multiple insertion and deletion errors.

[0007] The technical solution of the present invention is a DNA encoding method capable of correcting multiple insertion and deletion errors, and this method includes:

[0008] Step 1: Define a 4-ary (n, M 1 , d1 ) Code Specifically include:

[0009] Step 1.1: Define a 4 - element finite field Among them, α satisfies α 2 +α + 1 = 0, define The quaternary information sequence of length n on is: Construct A bijective mapping τ to the DNA base set ∑ DNA ={A, T, C, G}: τ(0)=C, τ(1)=G, τ(α)=A, τ(1 + α)=T;

[0010] Step 1.2: Define a 4 - element code The code length is n, the number of codewords is M 1 , and the minimum Hamming distance is d 1 , that is, (n, M 1 , d 1 ) code;

[0011] Step 2: Construct a 4 - element (n, M 2 , d 2 ) code Capable of correcting multiple insertion - deletion errors, specifically including:

[0012] Step 2.1: Select a positive integer s, 1 ≤ s < n, and define the weight sequence as W = {ω 1 , ω 2 ...}, where the definition of ω i is: when i ≤ 0, ω i = 0; when i ≥ 1,

[0013] Step 2.2: Select a positive integer m ≥ ω n+1 , 0 ≤ r < m, and construct the code where M(c)≡r mod m means that the remainder of M(c) divided by m is r; The code length of is n, the number of codewords M 2 ≤M 1 , and the minimum Hamming distance d 2 ≥d 1 ;

[0014] Step 2.3: Satisfy the construction of the generalized Helberg code. According to the properties of the Helberg code, Correct s insertion - deletion errors, so its Hamming distance d 2 ≥s + 1;

[0015] Step 2.4: Calculate The redundant expression, according to the pigeonhole principle: Redundancy is a determined code, and the number of codewords M 1 is given. So, the smaller m is, the smaller the redundancy R is; therefore, m = ω is selected n+1 , but ω is not given in the construction of the generalized Helberg code n+1 of the specific expression. When the code length n is relatively large, calculating ω n+1 will become very difficult. Therefore, the expression of ω n+1 is deduced as follows:

[0016] According to the definition of ω i , then there is

[0017]

[0018] Let to get:

[0019]

[0020] It is transformed into matrix form:

[0021]

[0022] Let the matrix

[0023] Diagonalize the matrix A: A = PΛP -1 , where P is the matrix composed of the eigenvectors of A, and Λ = diag(λ 1 , λ 2 ,..., λ s ), λ i is the eigenvalue of A, 1 ≤ i ≤ s, and all eigenvalues of A satisfy the equation:

[0024] λ s - 3λ s-1 - 3λ s-2 - … - 3 = 0

[0025] Therefore, it is obtained:

[0026]

[0027] So, according to the above formula, by diagonalizing the matrix A and calculating the first s values of ω i , the specific expression of ω n+1 can be obtained; thus, the redundant expression can be simplified: The upper bound of the redundancy value of this coding method under specific parameters can be calculated using the redundant expression.

[0028] The present invention utilizes the properties of the generalized Helberg code to construct a DNA coding method that can correct multiple insertion and deletion errors; at the same time, the expression of ω in the generalized Helberg code is also deduced to facilitate calculating the upper bound of the redundancy value of this coding method. n+1 BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0030] Embodiment 1:

[0031] Step 1: Set the code length n = 1023, and let be a quaternary (n, M 1 , d 1 ) BCH code, specifically including:

[0032] 1.1 Let α be a primitive element in the finite field , and m i (x) represent the minimal polynomial of α i . Then deg(m i (x)) ≤ 5. According to the definition of the BCH code, its generating polynomial g(x) = (x - 1)m 1 (x)m 2 (x)m 3 (x), so deg(g(x)) ≤ 16, and the dimension k of

[0033] ≥ n - deg(g(x)) = 1007. Therefore 2 , d 2 ) code that can correct multiple insertion and deletion errors, specifically including:

[0034] 2.1 Select a positive integer s = 2, and the weight sequence is W = {ω 1 , ω 2 ...}, where ω i is: when i ≤ 0, ω i = 0; when i ≥ 1, select a positive integer m = ω n+1 , 0 ≤ r < m, and construct with a code length n = 1023.

[0035] 2.2 Satisfies the construction of the generalized Helberg code. Therefore, according to the properties of the Helberg code, it can correct s insertion and deletion errors. Therefore, its Hamming distance d 2 ≥ s + 1.

[0036] 2.3 Calculate the value of ω n+1 below. According to the definition of ω i , ω n+1 = 1 + 3(ω n + ω n-1 ). Then we have

[0037]

[0038] Let Therefore, the above equation can be expressed as:

[0039] t n+1 = 3(t n + t n-1 )

[0040] Convert it into matrix form:

[0041]

[0042] Let the matrix

[0043] Diagonalize the matrix A: A = PΛP -1 , where P is the matrix composed of the eigenvectors of A, Λ = diag(λ 1 , λ 2 ), here

[0044] Therefore, we get:

[0045]

[0046] Among them

[0047] Thus, we get ω n+1 = -0.427×(-0.791288) n+1 + 0.296×(3.79129) n+1 - 0.2. Therefore, the upper bound of redundancy can be obtained:

[0048] Implementation Case Two:

[0049] Step 1: Set the code length n = 4095. Let be a 4-ary (n 1 , M 1 , d 1 ) BCH code, specifically including:

[0050] 1.1 Let α be a primitive element in a finite field and m i (x) represent the minimal polynomial of α i , then deg(m i (x)) ≤ 6. According to the definition of BCH code, its generator polynomial g(x) = (x - 1)m 1 (x)m 2 (x)m 3 (x)m 5 (x), then deg(g(x)) ≤ 25, and its dimension k ≥ n - deg(g(x)) = 4070. Therefore

[0051] Step 2, construct a 4 - ary (n, M 2 , d 2 ) code that can correct multiple insertion - deletion errors, specifically including:

[0052] 2.1 Select a positive integer s = 3, and the weight sequence is W = {ω 1 , ω 2 ...}, where ω i is: when i ≤ 0, ω i = 0; when i ≥ 1, select a positive integer m = ω n+1 , 0 ≤ r < m, and construct a code with length n = 4095.

[0053] 2.2 It satisfies the construction of the generalized Helberg code. Therefore, according to the properties of the Helberg code, it can correct s insertion - deletion errors, so its Hamming distance d 2 ≥ s + 1.

[0054] 2.3 Next, calculate the value of ω n+1 . According to the definition of ω i , ω n+1 = 1 + 3(ω n + ω n-1 + ω n-2 ). Then we have

[0055]

[0056] Let So the above formula can be expressed as:

[0057] t n+1= 3(t n + t n-1 + t n-2 )

[0058] Convert it into matrix form:

[0059]

[0060] Let the matrix

[0061] Diagonalize the matrix A: A = PΛP -1 , where P is the matrix composed of the eigenvectors of A, Λ = diag(λ 1 , λ 2 , λ 3 ), here

[0062] Therefore, we get:

[0063]

[0064] Where

[0065] Thus, we get:

[0066] ω n+1 = 0.2682×(3.9514) n+1 + (-0.0689 - 0.0145i)×(-0.4757 + 0.73i) n+1 + (-0.0689 + 0.0145i)×(-0.4757 - 0.73i) n+1 - 0.125

[0067] Therefore, the upper bound of redundancy can be obtained:

Claims

1. A DNA encoding method capable of correcting multiple insertion and deletion errors, the method comprises: Step 1: Define a 4-tuple (n, M 1 , d 1 ) code Specifically including: Step 1.1: Define the quaternary finite field where α satisfies α 2 +α + 1 = 0, and define the quaternary information sequence of length n over Construct a bijection τ from DNA to the DNA base set Σ DNA ={A, T, C, G}: τ(0) = C, τ(1) = G, τ(α) = A, τ(1 + α) = T; Step 1.2: Define a quaternary code The code length is n, and the number of codewords is M 1 , and the minimum Hamming distance is d 1 , that is, (n, M 1 , d 1 ) code; Step 2: Construct a 4-ary (n, M 2 , d 2 ) code that can correct multiple insertion and deletion errors, specifically including: Step 2.1: Select a positive integer s, where 1 ≤ s < n, and define the weight sequence as W = {ω 1 , ω 2 …}, where the definition of ω i is as follows: when i ≤ 0, ω i = 0; when i ≥ 1, Step 2.2: Select a positive integer m ≥ ω n+1 , 0 ≤ r < m, and construct a code where M(c) ≡ r mod m means that the remainder of M(c) divided by m is r; The code has a length of n, the number of codewords M 2 ≤ M 1 , the minimum Hamming distance d 2 ≥ d 1 ; Step 2.3: Satisfy the construction of the generalized Helberg code. According to the properties of the Helberg code, Correct s insertion and deletion errors. Therefore, its Hamming distance d 2 ≥ s + 1; Step 2.4: Calculate the redundant expression of, according to the pigeonhole principle: Redundancy is a determined code, and the number of its codewords M 1 has been given. Therefore, the smaller m is, the smaller the redundancy R is. Thus, select m = ω n+1 , and deduce the expression of ω n+1 : According to the definition of ω i , we have Let Obtain: converting into matrix form: Let the matrix Diagonalize matrix A: A = PΛP -1 , where P is the matrix composed of the eigenvectors of A, Λ = diag(λ 1 , λ 2 , …, λ s ), λ i are the eigenvalues of A, 1 ≤ i ≤ s, and all the eigenvalues of A satisfy the equation: λ s -3λ s-1 -3λ s-2 -…-3=0 Therefore, it is obtained that: Therefore, according to the above formula, by diagonalizing matrix A and calculating the first s ω i values, the specific expression of ω n+1 is obtained; simplify the redundant expression: Use the redundant expression to calculate the upper bound of the redundancy value of this coding method under specific parameters.

Citation Information

Patent Citations

  • Novel DNA storage coding method

    CN115642923A