High code rate dna encoding method and system satisfying gc local balance and run constraint
Patent Information
- Application Number
- CN202410699705.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-31
AI Technical Summary
但现在大多数DNA编码的方法都不具备生物约束性质,或者具备少量约束性质的同时牺牲了大量码率,对于能够具备多个生物约束性质的高码率DNA编码方法仍然缺乏
[0055]第一、本发明提出了一种同时满足GC局部平衡和游程约束的高码率DNA编码方法,该方法利用了数学归纳法的思想,对满足约束的码字进行了明确计算,并通过排序实现了对码字的编解码,这种编码方法具有极高的码率,可以接近码率上限。
Smart Images

Figure CN118675625B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data storage, and in particular relates to a high-code-rate DNA encoding method and system that satisfies GC local balance and run-length constraints. Background Technology
[0002] In recent years, DNA storage technology has attracted widespread attention due to its advantages such as high storage density, long storage time, low maintenance cost, and easy access. DNA-based data storage is a new method for long-term preservation of digital data. DNA storage refers to encoding information sequences into DNA sequences composed of four bases: adenine (A), thymine (T), guanine (G), and cytosine (C) using a pre-designed encoding scheme. These sequences are then synthesized into oligonucleotides or long DNA fragments to allow for long-term storage. To retrieve data, DNA sequencing is used to obtain the original ATCG sequences from the synthesized DNA.
[0003] Errors frequently occur in the base sequence during DNA synthesis and sequencing, and the probability of errors during DNA storage is closely related to the DNA sequence structure. A series of identical bases is called a base run. Studies have shown that a base run longer than six bases significantly increases the incidence of substitution and deletion errors; therefore, long runs should be avoided in DNA strands. Furthermore, GC content refers to the proportion of G and C bases in a DNA strand. Studies have also demonstrated that excessively high or low GC content leads to more errors during synthesis and sequencing. Therefore, most studies control the GC content in DNA strands to around 50%, a process known as GC equilibrium. GC equilibrium can be further divided into global GC equilibrium and local GC equilibrium. Global GC equilibrium refers to the GC content being balanced throughout the entire DNA strand, while local GC equilibrium refers to the GC content being balanced at a given integer... At any length of The GC content in the subsequence also reached equilibrium. Obviously, GC local equilibrium is also GC global equilibrium. GC local equilibrium is often used in PCR amplification.
[0004] Therefore, when encoding DNA, additional biological constraints need to be added to better suit practical storage requirements. However, most current DNA encoding methods lack biological constraints, or they have a few constraints but sacrifice a significant amount of code rate. High-code-rate DNA encoding methods that possess multiple biological constraints are still lacking. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a high-code-rate DNA encoding method and system that satisfies GC local balance and run-length constraints.
[0006] This invention is implemented as follows: a high-rate DNA encoding method that satisfies GC local balance and run-length constraints, the method comprising:
[0007] Step 1: Define a value that satisfies GC ∈ - global balance and satisfies -4-yuan code with travel restrictions And calculate using the ideas of subset partitioning and mathematical induction. The specific number of codewords;
[0008] Step 2: For Sort the characters by size and establish a relationship with integers. The mapping between them, sorted by size, achieves the mapping between them. Encoding and decoding;
[0009] Step 3: [Regarding...] The codewords are recombined to construct a system that satisfies GC(m, δ)-local equilibrium and satisfies -4-yuan code with travel restrictions And can be obtained through a sequence of integers. Encode and decode.
[0010] Furthermore, step 1 specifically includes:
[0011] Step 1.1: Construction To the DNA base set ∑ DNA = {A, T, C, G} bijective τ: τ(0) = C, τ(1) = G, τ(2) = A, τ(3) = T; let [n] denote the set of integers {1, 2, ..., n};
[0012] Step 1.2: Let The above 4-tuple of length n is: c i Let represent the coordinate of the i-th position in sequence c; define the GC weight of C as wt. GC (c) = |{i∈[n]:c i =0}|+|{j∈[n]:c j =1}|, if wt GC If (c)∈[(0.5-∈)n, (0.5+∈)n], where 0<∈<1, then the codeword c is said to satisfy GC-∈ global balance; if for any subsequence b of length m in c, wt GC (b) ∈ [(0.5-∈)m, (0.5+∈)m], then the codeword c is said to satisfy GC(m,∈)-local balance; if the codeword c can be divided into t characters of length t... A continuous subsequence, i.e., c = (c1, c2, ..., c3). t ), and each subsequence c i satisfy Then the codeword c is said to satisfy - Segmented balancing;
[0013] Step 1.3: Consecutive identical elements are called a run. Any sequence can consist of several runs. If the length of any run in a sequence c does not exceed a certain value, then the sequence is considered complete. Then the sequence c is said to satisfy... -Travel restrictions; here it is used To represent simultaneously satisfying GC∈-global balance and - A set of 4-ary codewords with run-length restrictions;
[0014] Step 1.4: Use S to represent All lengths not exceeding Let Γ(s) represent the elements that make up the run s in the set of runs s. Let A(n, s, g) represent the set of all codewords with length n, last run s, and GC weight g, and then let Φ(n, s, g) represent the number of codewords in it, i.e. Φ(n, s, g) = |A(n, s, g)|;
[0015] Step 1.5: Theorem 1 is given: It can be decomposed into
[0016]
[0017] The calculation method for Φ(n, s, g) is as follows:
[0018]
[0019] Ω(·) is 1 when the condition in the parentheses is true, and 0 otherwise.
[0020] Furthermore, the proof of Theorem 1 is as follows:
[0021] first, It is obvious that A(n, s, g) can be divided into several non-overlapping sets, where for any s ∈ S, g ∈ [(0.5-∈)n, (0.5+∈)n]. When n ≤ |s|, the value of Φ(n, s, g) can be easily obtained through the definition. When n > |s|, a recursive relation is used for calculation. Let the codeword c = c′||s, where || represents the concatenation of two sequences, i.e., the last run of c is s, and the sequence after removing the run is c′. Let s′ represent the last run of c′, satisfying Γ(s′) ≠ Γ(s). Then c ∈ A(n, s, g) if and only if c′ ∈ A(n-|s|, s', g-wt). GC (s)), therefore the last equation can also be obtained; it can be calculated using Theorem 1. Specific number of code words
[0022] Furthermore, step 2 specifically includes:
[0023] Step 2.1: For s∈S, define Θ(s)=Γ(s)+4(|s|-1); for a pair of tuples Define (s′, g′) < (s, g) if g′ < g or g′ = g and Θ(s′) < Θ(s); for any two distinct quadruple sequences The last runs are s1 and s2, and c1 < c2 is defined if one of the following two conditions is met:
[0024] (1)(s1, wt) GC (c1))<(s2,wt GC (c2))
[0025] (2)
[0026] Among them, Λ t (c) represents the first t bits of codeword c;
[0027] Step 2.2: According to the definition in Step 2.1, any two different quadruple sequences There are only two cases: c1 < c2 or c1 > c2. Therefore, we can... If all codewords in the given text are sorted by size, then there must exist a bijective λ:
[0028]
[0029] λ represents the integer k and The mapping of the k-th smallest codeword c in the code; the following describes how to calculate the mapping result;
[0030] Step 2.3: Let Then for any integer There exists a function You can return The k-th smallest codeword in the array; Lemma 2 is given: for any integer There must exist a unique one. satisfy and So there are
[0031]
[0032] in g′=g-wt GC (s); the proof is as follows:
[0033] Let φ0 = 0, let
[0034]
[0035] Where i∈[1,m]; it is easy to see that φ i As i monotonically increases, therefore for any integer... There must exist an i∈[1, m] such that So We can obtain A(n, s) i g i ) The smallest sequence; and because A(n, s) i g i The sequences in the sequence have the same last run s i Further, we can obtain
[0036]
[0037] Proof complete;
[0038] Step 2.4: Similarly, there exists a function It is possible to return c in The order in the sequence; similar to step 2.3, Lemma 3 is given: for any codeword Let s represent the last run of c, and g = wt GC (c), then we have
[0039]
[0040] in g′=g-wt GC (s); The proof here is similar to step 2.3, and will not be repeated here;
[0041] Step 2.5: Based on Lemma 2 and Lemma 3, it can be achieved and Encoding and decoding between the k-th smallest codewords; for example: let... The code length is 4, and it satisfies the global balance sum of ∈ = 0.1. Travel restrictions, if you want to get The 60th smallest codeword in the middle; combine the calculated runs to finally obtain The 60th smallest codeword in the code is c = 3112; similarly, to calculate c = 3112 in... The order in which the calculations are performed Accumulate to get c in The order is k = 1 + 1 + 18 + 40 = 60.
[0042] Furthermore, step 3 specifically includes:
[0043] Step 3.1: Given Lemma 4: If the codeword satisfy - If the codeword c is segmented and balanced, then it also satisfies GC(m, δ) - local balance, where and
[0044] Step 3.2: Definition According to the definition in step 1.2, Satisfies GC(n, ∈) - piecewise balance; according to Lemma 4, Simultaneously satisfying GC(m, δ)-local equilibrium, where 2n < m ≤ nt, in addition, satisfy -Travel restrictions;
[0045] Step 3.3: Similar to step 2, Encoding and decoding can be performed using a sequence of integers; for any codeword c = (c1, c2, ..., c) t ),in Let λ be the mapping defined in step 2, λ -1 If λ is the corresponding inverse mapping, then λ -1 (c) = λ(c1, c2, ..., c t )=(k1,k2,...,k t ),in Therefore, using the encoding / decoding algorithm in step 2, it is also possible to achieve (k1, k2, ..., k t ) and (c1, c2, ..., c t Encoding between )
[0046] Step 3.4: It has an extremely high bitrate, approaching the upper limit of 2 when ∈ = 0.1. The bitrates for different code lengths are as follows: (Bitrate calculation method is log2|C) L | / nt).
[0047] Another object of the present invention is to provide a high-code-rate DNA coding system that satisfies GC local balance and run-length constraints for implementing the aforementioned high-code-rate DNA coding method that satisfies GC local balance and run-length constraints, the system comprising:
[0048] Define a computation module to define a system that satisfies GC ∈ - global balance and satisfies -4-yuan code with travel restrictions And calculate using the ideas of subset partitioning and mathematical induction. The specific number of codewords;
[0049] The sorting encoding and decoding module, connected to the definition calculation module, is used for... Sort the characters by size and establish a relationship with integers. The mapping between them, sorted by size, achieves the mapping between them. Encoding and decoding;
[0050] The combined decoding module, connected to the sorting encoding and decoding module, is used for... The codewords are recombined to construct a system that satisfies GC(m, δ)-local equilibrium and satisfies -4-yuan code with travel restrictions And can be obtained through a sequence of integers. Encode and decode.
[0051] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the high code rate DNA encoding method that satisfies GC local balance and run-length constraints.
[0052] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the high code rate DNA encoding method that satisfies GC local balance and run-length constraints.
[0053] Another objective of this invention is to provide an information data processing terminal for implementing the high code rate DNA coding system that satisfies GC local balance and run-length constraints.
[0054] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0055] First, this invention proposes a high-rate DNA encoding method that simultaneously satisfies GC local balance and run-length constraints. This method utilizes the idea of mathematical induction to explicitly calculate the codewords that satisfy the constraints, and realizes the encoding and decoding of the codewords through sorting. This encoding method has an extremely high code rate, which can approach the upper limit of the code rate.
[0056] Secondly, most current DNA encoding methods lack biological constraints or sacrifice code rate while possessing a few constraints. High-code-rate DNA encoding methods that possess multiple biological constraints are still lacking. This invention effectively solves the current problem by employing the idea of mathematical induction to explicitly calculate codewords that satisfy the constraints and to achieve encoding and decoding of codewords through sorting. This encoding method has an extremely high code rate, which can approach the upper limit of the code rate.
[0057] Third, as supporting evidence of the inventiveness of this invention, it is also reflected in the following important aspects:
[0058] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:
[0059] This invention constructs a DNA encoding method that meets two stringent biological constraints, has an extremely high code rate, and provides an effective encoding and decoding algorithm. It fits the needs of practical applications and can be applied to certain DNA storage systems with high requirements for biological constraints and code rate. It can effectively reduce the error rate of DNA sequences during the synthesis and sequencing process and has high availability.
[0060] (2) The technical solution of this invention fills a technical gap in the industry both domestically and internationally:
[0061] Most current DNA encoding methods lack biological constraints or sacrifice a significant amount of code rate while possessing only a few constraints. This invention constructs a novel encoding method that not only satisfies GC local balance and run-length constraints but also boasts an extremely high code rate. Furthermore, it provides an effective encoding and decoding method, filling a technological gap in the field of DNA encoding.
[0062] (3) The technical solution of the present invention solves a technical problem that people have long wanted to solve but have never been able to solve successfully:
[0063] In DNA coding, DNA sequences need to meet certain biological constraints to reduce the probability of errors. However, there is a trade-off between biological constraints and code rate; the more biological constraints are met and the more stringent the constraints, the lower the sequence code rate, which is a pain point in DNA coding. The coding method constructed in this scheme simultaneously satisfies two stringent constraints: GC local balance and run-length constraints, and has a code rate close to the upper bound, while ensuring the availability and reliability of the storage system, thus effectively overcoming this pain point.
[0064] Fourth, in the fields of bioinformatics and genetic engineering, efficient and accurate encoding and decoding of DNA sequences has always been a technical challenge. Traditional DNA encoding methods often struggle to simultaneously satisfy the requirements of GC local equilibrium (i.e., the ratio of guanine G to cytosine C in the sequence remains balanced within a local region) and run constraints (i.e., the maximum length of consecutive identical bases in the sequence is limited). This limitation leads to low encoding efficiency, high decoding error rates, and potential interference with biological functions, severely restricting the development of DNA data storage and gene editing applications.
[0065] This invention proposes a high-rate DNA encoding method that satisfies GC local balance and run-length constraints. Through ingenious algorithm design and the application of mathematical induction, it effectively solves the aforementioned problems. The method first defines a globally balanced 4-ary code and calculates the specific number of codewords by partitioning subsets and using mathematical induction. Next, by sorting the codewords by size and establishing a mapping relationship with integers, an efficient and accurate encoding and decoding process is achieved. Finally, by recombinating the codewords, a 4-ary code that satisfies local balance and run-length constraints is constructed, further improving the encoding efficiency and accuracy.
[0066] The significant technological advancements of this invention are mainly reflected in the following aspects: First, by introducing the concepts of GC local equilibrium and run-length constraints, this invention improves the accuracy and reliability of DNA encoding and reduces the risk of interference with biological functions. Second, by employing mathematical induction and integer mapping methods, this invention achieves an efficient and accurate encoding and decoding process, significantly improving encoding efficiency. Finally, the encoding method of this invention has an extremely high code rate, approaching the theoretical upper limit, providing strong technical support for applications such as DNA data storage and gene editing.
[0067] This invention has broad application prospects, playing a vital role not only in bioinformatics and genetic engineering but also extending to fields such as data storage and information security. By employing the DNA encoding method of this invention, we can store and transmit biological information data more efficiently and accurately, providing strong support for life science research and applications. Furthermore, this invention brings new technical ideas and solutions to fields such as DNA data storage and gene editing, possessing significant theoretical and practical value. Attached Figure Description
[0068] Figure 1 This is a flowchart of a high-rate DNA encoding method that satisfies GC local balance and run-length constraints, provided by an embodiment of the present invention.
[0069] Figure 2 This is a structural diagram of a high-rate DNA coding system that satisfies GC local balance and run-length constraints, provided in an embodiment of the present invention.
[0070] Figure 3 C G Partial code results;
[0071] Figure 4 C G The distribution of GC content and maximum run length of the middle codewords;
[0072] Figure 5 C L Partial code results;
[0073] Figure 6 C L The distribution of GC content and maximum run length of the middle codewords. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0075] Example 1: DNA Data Storage
[0076] DNA, as a high-density storage medium, has a storage capacity far exceeding that of traditional electronic storage devices. However, in order to achieve efficient DNA data storage, the encoding method must meet both biological constraints to ensure the accuracy of synthesis and retrieval, and a high code rate to maximize storage capacity.
[0077] 1) Data encoding:
[0078] The method of this invention encodes digital information to generate DNA sequences that conform to GC local equilibrium and run constraints. Encoding and decoding of codewords through sorting ensures that the generated DNA sequences are both efficient and stable.
[0079] 2) Data storage and retrieval:
[0080] The encoded DNA sequence is synthesized and stored in a suitable medium. High-throughput sequencing technology is used to read the stored DNA sequence, and the original data is decoded and reconstructed using sorting methods.
[0081] 3) Advantages:
[0082] The high-code-rate DNA encoding method of this invention approaches the upper limit of code rate, significantly improving the storage capacity per unit length of DNA sequence. It meets biological constraints, ensuring the stability and accuracy of the DNA sequence during storage and retrieval, and reducing the risk of data loss and errors.
[0083] Example 2: Gene Synthesis and Gene Editing
[0084] Gene synthesis and gene editing technologies have wide applications in the biomedical and bioengineering fields. However, to ensure the efficiency and accuracy of these technologies, the synthesized DNA sequences need to meet specific biological constraints (such as GC content balance and avoidance of specific sequence patterns). Traditional DNA coding methods often sacrifice code rate to meet these biological constraints, thus limiting coding efficiency and information capacity.
[0085] 1) Design and synthesize genes:
[0086] The high-code-rate DNA encoding method of this invention is used to generate DNA sequences that conform to GC local equilibrium and run constraints. High-code-rate codewords satisfying these constraints are explicitly calculated using mathematical induction, ensuring the stability and efficiency of the encoded DNA sequences during biosynthesis.
[0087] 2) Gene editing vectors:
[0088] The guide RNA (gRNA) sequences required for designing gene editing tools such as CRISPR-Cas9 are generated using the method of this invention, improving editing efficiency and specificity while avoiding non-specific cleavage and off-target effects.
[0089] 3) Advantages:
[0090] High-code-rate encoding methods approach the code-rate limit, allowing more information to be encoded in the same length of DNA sequence, thus improving the efficiency of gene synthesis and editing; they also meet biological constraints, ensuring that the synthesized DNA sequence has high stability and low toxicity in biological systems.
[0091] As can be seen from the above embodiments, the DNA encoding method proposed in this invention can be effectively applied to bioinformatics fields such as gene editing and sequence analysis, and has broad application prospects.
[0092] The technical solution of this invention is a novel high-code-rate DNA encoding method that satisfies GC local balance and run-length constraints. The method includes:
[0093] Step 1: Define a globally balanced system that satisfies GC ∈ -1 and satisfies -4-yuan code with travel restrictions And calculate using the ideas of subset partitioning and mathematical induction. The specific number of codewords.
[0094] Step 1.1: Construction To the DNA base set ∑ DNA Let τ be a bijection of {A, T, C, G}: τ(0) = C, τ(1) = G, τ(2) = A, τ(3) = T. Let [n] denote the set of integers {1, 2, ..., n}.
[0095] Step 1.2: Let The above 4-tuple of length n is: c i Let represent the coordinate of the i-th position in sequence c. The GC weight of c is defined as wt. GC (c) = |{i∈[n]:c i =0}|+|{j∈[n]:c j =1}|, if wt GCIf (c)∈[(0.5-∈)n, (0.5+∈)n], where 0<∈<1, then the codeword c is said to satisfy GC-∈ global balance. If for any subsequence b of length m in c, wt GC (b)∈[(0.5-∈)m, (0.5+∈)m], then the codeword c is said to satisfy GC(m,∈)-local balance. If the codeword c can be divided into t characters of length t... A continuous subsequence, i.e., c = (c1, c2, ..., c3). t ), and each subsequence c i satisfy The codeword c is said to satisfy GC. - Segmented balancing.
[0096] Step 1.3: Consecutive identical elements are called a run. Any sequence can consist of several runs. If the length of any run in a sequence c does not exceed a certain value, then the sequence is considered complete. Then the sequence c is said to satisfy... -Travel restrictions. (This is used here.) To represent simultaneously satisfying GC∈-global balance and - A set of 4-ary codewords with run-length restrictions.
[0097] Step 1.4: Use S to represent All lengths not exceeding Let Γ(s) represent the elements that make up the run s in the set of runs s. Let A(n, s, g) represent the set of all codewords with length n, last run s, and GC weight g, and then let Φ(n, s, g) represent the number of codewords, i.e., Φ(n, s, g) = |A(n, s, g)|.
[0098] Step 1.5: Theorem 1 is given: It can be decomposed into
[0099]
[0100] The calculation method for Φ(n, s, g) is as follows:
[0101]
[0102] Where Ω(·) is 1 when the condition in the parentheses is true, and 0 otherwise. The proof of Theorem 1 is as follows:
[0103] first, It is obvious that A(n, s, g) can be divided into several non-overlapping sets, where for any s ∈ S, g ∈ [(0.5-∈)n, (0.5+∈)n]. When n ≤ |s|, the value of Φ(n, s, g) can be easily obtained through the definition. When n > |s|, a recursive relation is used for calculation. Let the codeword c = c′||s, where || represents the concatenation of two sequences, i.e., the last run of c is s, and the sequence after removing the run is c′. Let s′ represent the last run of c′, satisfying Γ(s′) ≠ Γ(s). Then c ∈ A(n, s, g) if and only if c′ ∈ A(n-|s|, s′, g-wt). GC Therefore, the last equation can also be obtained.
[0104] It can be calculated using Theorem 1. Specific number of code words
[0105] Step 2: For Sort the characters by size and establish a relationship with integers. The mapping between them, sorted by size, achieves the mapping between them. Encoding and decoding.
[0106] Step 2.1: For s∈S, define Θ(s)=Γ(s)+4(|s|-1). For a pair of tuples… Define (s′, g′) < (s, g) if g′ < g or g′ = g and Θ(s′) < Θ(s). For any two distinct quadruple sequences... The last runs are s1 and s2, and c1 < c2 is defined if one of the following two conditions is met:
[0107] (1)(s1, wt) GC (c1))<(s2,wt GC (c2))
[0108] (2)
[0109] Among them, Λ t (c) represents the first t bits of codeword c.
[0110] Step 2.2: According to the definition in Step 2.1, any two different quadruple sequences There are only two cases: c1 < c2 or c1 > c2. Therefore, we can... Sort all codewords in the given set by size. Then there must exist a bijective λ:
[0111]
[0112] λ represents the integer k and The mapping of the k-th smallest codeword c in the dataset. The following describes how to calculate the mapping result.
[0113] Step 2.3: Let Then for any integer There exists a function You can return The k-th smallest codeword. Lemma 2 is given: For any integer... There must exist a unique one. satisfy and So there are
[0114]
[0115] in g′=g-wt GC (s). The proof is as follows:
[0116] Let φ0 = 0, let
[0117]
[0118] Where i∈[1, m]. It is easy to see φ i As i monotonically increases, therefore for any integer... There must exist an i∈[1, m] such that So We can obtain A(n, s) i g i ) The smallest sequence. And because A(n, s) i g i The sequences in the sequence have the same last run s i Further, we can obtain
[0119]
[0120] The proof is complete.
[0121] Step 2.4: Similarly, there exists a function It is possible to return c in The order of the codewords. Similar to step 2.3, Lemma 3 is given: For any codeword Let s represent the last run of c, and g = wt GC (c), then we have
[0122]
[0123] in g′=g-wt GC(s). The proof here is similar to step 2.3 and will not be repeated.
[0124] Step 2.5: Based on Lemma 2 and Lemma 3, it can be achieved and Encoding and decoding between the k-th smallest codewords. For example: Let... The code length is 4, and it satisfies the global balance sum of ∈ = 0.1. Travel restrictions, if you want to get The 60th smallest codeword is calculated as follows:
[0125] Table 1 Calculation The process of the 60th smallest codeword
[0126]
[0127] By combining the calculated runs, we finally obtain... The 60th smallest codeword is c = 3112. Similarly, to calculate c = 3112 in... The order of operations and calculations is as follows:
[0128] Table 2 calculates c = 3112 in... The order in
[0129]
[0130] The calculated Accumulate to get c in The order is k = 1 + 1 + 18 + 40 = 60.
[0131] Step 3: [Regarding...] The codewords are recombined to construct a system that satisfies GC(m, δ)-local equilibrium and satisfies -4-yuan code with travel restrictions And can be obtained through a sequence of integers. Encode and decode.
[0132] Step 3.1: Given Lemma 4: If the codeword satisfy - If the codeword c is segmented and balanced, then it also satisfies GC(m, δ) - local balance, where and
[0133] Step 3.2: Definition According to the definition in step 1.2, It satisfies GC(n, ∈) - piecewise balance. According to Lemma 4, Simultaneously satisfying GC(m, δ)-local equilibrium, where 2n < m ≤ nt, in addition, satisfy - Travel restrictions.
[0134] Step 3.3: Similar to step 2, Encoding and decoding can be performed using a sequence of integers. For any codeword... c = (c1, c2, ..., c) t ),in Let λ be the mapping defined in step 2, λ -1 If λ is the corresponding inverse mapping, then λ -1 (c) = λ(c1, c2, ..., c t )=(k1,k2,...,k t ),in Therefore, using the encoding / decoding algorithm in step 2, it is also possible to achieve (k1, k2, ..., k t ) and (c1, c2, ..., c t The encoding between )
[0135] Step 3.4: It has an extremely high bitrate, approaching the upper limit of 2 when ∈ = 0.1. When t=2, the bitrates for different code lengths are as follows: (Bitrate calculation method is log2|C L | / nt)
[0136] Table 3∈=0.1, Code rate at t=2 with different code lengths
[0137] 200 1.995424 400 1.995775 600 1.995768
[0138] II. Application Examples. To demonstrate the inventiveness and technical value of the technical solution of this invention, this section provides application examples of the technical solution of the claims on specific products or related technologies.
[0139] Based on the DNA encoding method mentioned in this invention, we can consider the following specific application example and its implementation scheme:
[0140] Application Example 1: DNA Storage
[0141] 1) Environment setup: Store text data in DNA base sequences.
[0142] 2) Preparation of the codeword set: First, set parameters m and δ, and select parameters n, ∈, δ ... t, calculate the sum satisfying GC∈-global equilibrium according to the method of Theorem 1. -Run constraints All codewords are processed and sorted by size. Then, Lemma 4 is used to... The codewords are combined to obtain a code that satisfies GC(m, δ)-local balance.
[0143] 3) Raw data processing: Establishing text data and... A one-to-one mapping that transforms data into The code words in the text.
[0144] 4) Data mapping and storage: The encoded 4-gram sequence is converted into a DNA sequence and then synthesized and stored in DNA.
[0145] 5) Integration and Testing: The synthesized DNA is subjected to multiple PCR amplifications and sequencing to obtain the contained DNA sequences. The DNA sequences are converted back into quaternary sequences, compared with the original data, and the error rate is calculated. Properties such as local equilibrium and run restriction are also determined.
[0146] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0147] The embodiments of the present invention have achieved some positive results during the research and development or use process, and have indeed great advantages compared with the prior art. The following content describes them in conjunction with the data, charts and other information of the experimental process.
[0148] Simulations were conducted based on the designed experiment, with parameters n = 10, ε = 0.1, l = 3, and t = 3. First, the set of quadruple codewords C satisfying GC 0.1-global balance and 3-run length constraint was obtained. G The number of codewords meeting the criteria is 653,896, with a code rate of 1.931870. Some codeword results are shown below. Figure 3 As shown, C G Partial code results. Figure 4 C GThe distribution of GC content and maximum run length of the codewords in the middle part; then, the distribution of C... G All codewords in C are encoded and decoded using an integer k, for example, 1→[3,3,1,3,1,0,2,1,2,2], 2→[1,2,3,3,1,0,3,0,0,1], 3→[1,3,1,2,0,2,3,3,2,0]. G All codewords in the set are recombined to obtain a set C of codewords that satisfy GC(m,δ)-local balance. L Where 20 < m ≤ 30, The number of codewords that meet the criteria is 279592837827867136, and the code rate is also 1.931870. Some codeword results are shown below. Figure 3 As shown, C L Partial code results; Figure 6 C L The distribution of GC content and maximum run length of the middle codewords.
[0149] The simulation results show that C L The global GC content of all codewords in the dataset is concentrated between 40% and 60%, and the maximum run length does not exceed 6, thus satisfying the 6-run limit. Then, for C... L All codewords in the code are encoded and decoded using an integer sequence (k1, k2, k3), for example:
[0150] (1,2,3)→[3,3,1,3,1,0,2,1,2,2,1,2,3,3,1,0,3,0,0,1,1,3,1,2,0,2,3,3,2,0].
[0151] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A high-rate DNA encoding method that satisfies GC local equilibrium and run-length constraints, characterized in that, The method includes: Step 1: Define a 4-ary code that satisfies GC global balance and run-length constraint, and calculate the number of specific codewords by partitioning subsets and mathematical induction; Step 1.1: Construction To DNA base set double shot : ;use express The set of integers; Step 1.2: Let The upper length is The 4-tuple sequence is: , Represents a sequence The Position coordinates; definition The GC weight is ,if ,in This is called code. Satisfy GC- Global balance; if for Any one of the lengths is subsequence of Both are available. This is called code. Satisfy GC Local balance; if the code is in place It can be divided into A length of A continuous subsequence, i.e. And each subsequence satisfy This is called code. Satisfy GC - Segmented balancing; Step 1.3: Consecutive identical elements are called a run. Any sequence can consist of several runs. If a sequence... The length of any of the journeys does not exceed Then it is called a sequence. satisfy Trip restrictions; here used To represent simultaneously satisfying GC Global balance and A set of 4-yuan codewords with run-length restrictions; Step 1.4: Use represent All lengths not exceeding The collection of tour itineraries, if ,use To represent the composition The elements of this tour are... ;use The representative length is The last leg of the journey is And the weight of GC is The set of all codewords, then using This represents the number of its codewords, i.e. ; Step 1.5: Theorem 1 is given: It can be decomposed into ; in, The calculation method is as follows: ; in The value is 1 when the condition in parentheses is true, and 0 otherwise; the proof of Theorem 1 is as follows: first, It can obviously be divided into several non-overlapping groups. For any , ;when hour, The value can be easily obtained through definition; while when At that time, a recursive relationship was used for calculation, and the codeword was set. , This represents the concatenation of two sequences, i.e. The last leg of the journey is The sequence after removing runs is ,use represent The last leg of the trip, satisfying ,So If and only if Therefore, the last equation can also be obtained; It can be calculated using Theorem 1. Specific number of code words ; Step 2: Sort the codewords by size and establish a mapping with integers to achieve encoding and decoding; Step 3: Recombine the codewords to construct a 4-ary code that satisfies GC local balance and run-length constraint, and can be encoded and decoded using integer sequences; Step 4: Using the obtained codeword set that satisfies GC local equilibrium and run-length constraint, the DNA sequences in the study area are encoded and decoded to support applications such as gene editing and sequence analysis in bioinformatics.
2. The method as described in claim 1, characterized in that, Step 2 specifically includes: Step 2.1: Define the size relationship between codewords; Step 2.2: Sort the codewords according to their size and establish a mapping between them and integers; Step 2.3: Find the corresponding codeword using an integer index to perform encoding; Step 2.4: Find the corresponding integer index through the codeword to achieve decoding.
3. The method as described in claim 1, characterized in that, Step 3 specifically includes: Step 3.1: Prove that codewords that satisfy GC segment balance also satisfy GC local balance; Step 3.2: Define the set of codewords that satisfy GC segment balance and run-length constraints; Step 3.3: Encode and decode codewords that satisfy GC local balance and run-length constraints using the integer sequence and the mapping relationship in Step 2; Step 3.4: Calculate and prove that the bit rate achieved by this method is close to or reaches the theoretical upper limit.
4. A high-rate DNA coding system that implements the method as described in any one of claims 1-3 and satisfies GC local equilibrium and run-length constraints, characterized in that, The system includes: Define a computation module to define a GC-compliant module. Global balance and satisfy 4-yuan code with travel restrictions And calculate using the ideas of subset partitioning and mathematical induction. The specific number of codewords; The sorting encoding and decoding module, connected to the definition calculation module, is used for... Sort the characters by size and establish a relationship with integers. The mapping between them, sorted by size, achieves the mapping between them. Encoding and decoding; The combined decoding module, connected to the sorting encoding and decoding module, is used for... The codewords are recombined to construct a GC-compliant structure. Local equilibrium and satisfy 4-yuan code with travel restrictions And can be obtained through a sequence of integers. Encode and decode.
5. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the high code rate DNA encoding method that satisfies GC local balance and run-length constraints as described in any one of claims 1-3.
6. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the high-code-rate DNA encoding method satisfying GC local balance and run-length constraints as described in any one of claims 1-3.
7. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the high code rate DNA coding system that satisfies GC local balance and run-length constraints as described in claim 4.