Encoding method and system with global gc balance and run constraint and capable of correcting burst indel errors

CN122844860APending Publication Date: 2026-09-29UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610678512.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

上述工作分别提供了DNA约束预编码和元突发删除纠错的基础工具,但尚未形成一种先将原始四元信息序列编码为满足GC全局平衡与游程长度限制的预编码码字,再以该预编码码字作为外层突发插入删除纠错结构输入的统一编码方法

Benefits of technology

[0051]本方案将生化适配与物理级纠错统一于编码设计,使DNA存储系统在较高编码效率下依然能够稳健应对生化约束和突发插入删除双重挑战,从源头提升数据写入和读出的成功率,对推进DNA数据存储的实用化和大规模部署具有实质性的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122844860A_ABST
    Figure CN122844860A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of information storage technology and discloses an encoding method and system that achieves GC global balance and run-length constraints while correcting burst insertion / deletion errors. The invention converts the outer auxiliary information into an interleaved auxiliary tail segment that satisfies GC balance and run-length constraints. Further, it combines the quaternary pre-coded codeword with the interleaved auxiliary tail segment through connection symbols and buffers to obtain the final codeword, which simultaneously satisfies η-GC global balance and -run-length constraints. The decoder determines the type of burst insertion / deletion error based on the change in the length of the received sequence and recovers the quaternary pre-coded codeword through finite candidate sequence enumeration and outer auxiliary information verification, thereby recovering the original quaternary information sequence. This invention is applicable to encoding scenarios in DNA storage that simultaneously require biochemical constraint satisfaction and burst insertion / deletion error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to, but is not limited to, the field of information storage technology, and particularly relates to an encoding method and system that achieves global GC balance and run-length constraints and can correct sudden insertion and deletion errors. Background Technology

[0002] DNA storage is a technology that uses DNA molecules (composed of the bases A, T, C, and G) as a storage medium. Due to its ultra-high density and long storage time, it is considered a revolutionary direction for next-generation storage technology. DNA storage achieves long-term data preservation by encoding digital information into DNA sequences and synthesizing them into DNA molecules. However, in practical applications of DNA storage, DNA sequences are susceptible to various types of errors during synthesis, storage, and sequencing, the most common of which include base insertion, deletion, and substitution errors. These errors can lead to deviations in the original data during retrieval, thereby reducing the reliability and accuracy of the stored data.

[0003] Further research revealed that the biochemical properties of DNA sequences have a significant impact on the error correction capabilities of storage systems. Excessively high or low GC content can lead to abnormal thermal stability of the DNA strand, while excessively long homopolymer run lengths (typically required to be less than 6 in DNA storage channels) increase the probability of secondary structure formation, thereby increasing the error rate during storage.

[0004] To address these errors, DNA storage coding technology needs to incorporate redundancy mechanisms, using error-correcting codes to resist errors and ensure the integrity of stored data. When designing error-correcting codes, in addition to considering error-correction capability, specific constraints of the DNA sequence must be met, such as GC-content balance and run-length constraint. GC balance requires the total ratio of G and C bases in the DNA sequence to be maintained at approximately 50% (allowing for a maximum tolerance for deviation). To avoid instability in the physicochemical properties of DNA molecules due to too many or too few GC base pairs, run length limitation restricts the length of consecutive identical bases to reduce signal interference or errors that may occur during sequencing. Therefore, designing DNA storage coding systems that simultaneously satisfy global GC balance, run length limitation, and error correction capabilities has become an important research direction for improving the reliability of storage technologies.

[0005] Although existing research has proposed various error-correcting code design methods for DNA storage, most methods still have limitations. For example, some methods focus on satisfying GC balance and run-length constraints, but lack efficient error correction capabilities for insertion and deletion errors; others, while able to correct insertion and deletion errors, fail to simultaneously satisfy GC balance and run-length constraints. Furthermore, existing error-correcting codes still have significant room for optimization in terms of redundancy and encoding / decoding complexity, making it difficult to meet the requirements of low redundancy and efficient encoding / decoding in practical applications.

[0006] Error correction is a particularly significant challenge in DNA storage. Insertion and deletion errors typically occur burstily, for example, during DNA synthesis or sequencing, due to the instability of polymerase chain reaction (PCR) amplification or equipment noise interference, potentially leading to the insertion or deletion of multiple consecutive bases. Traditional error-correcting code design methods, such as Hamming codes or Reed-Solomon codes, are primarily designed for substitution errors and perform poorly in handling insertion and deletion errors. Therefore, designing an error-correcting code that simultaneously satisfies GC balance, run-length constraints, and the ability to efficiently correct bursty insertion and deletion errors remains a core challenge in the field of DNA storage.

[0007] In existing research, Nguyen et al.[1] proposed a class of capacity approximation constraint coding methods for DNA storage. This method uses techniques such as differential transformation, continuous zero substring substitution, prefix flipping, and balanced indexing to make the coding sequence meet the homopolymer run length constraint and GC content constraint, and further supports editing error correction. Song, Cai, and Quek[2] proposed The corrected length in the meta-sequence does not exceed The encoding construction for sudden deletion errors employs VT-type constraint functions for locating the deletion position, shift VT-type constraint functions, and auxiliary functions for recovering the deleted substring. The aforementioned work provides DNA-constrained precoding and... While basic tools for burst deletion correction exist, a unified encoding method has yet to be developed that first encodes the original quaternary information sequence into a pre-encoded codeword that satisfies global GC balance and run-length constraints, and then uses this pre-encoded codeword as input to the outer burst insertion / deletion correction structure. Therefore, it is still necessary to further construct an encoding method that can maintain homopolymer run-length constraints and global GC content constraints while correcting burst insertion / deletion errors, building upon the aforementioned basic tools.

[0008] References: [1] TT Nguyen, K. Cai, KA Schouhamer Immink, and HM Kiah, “Capacity-Approaching Constrained Codes With Error Correction for DNA-BasedData Storage,” IEEE Transactions on Information Theory, vol. 67, no. 8, pp.5602–5613, 2021. [2] W. Song, K. Cai, and TQS Quek, “New Construction of q-aryCodes Correcting a Burst of at Most Deletions,” in Proceedings of the 2024IEEE International Symposium on Information Theory (ISIT), 2024, pp. 1101–1106. Summary of the Invention To address the problems existing in the prior art, this invention provides an encoding method and system that satisfies GC global balance and run length limitations and can correct burst insertion and deletion errors.

[0009] This invention is implemented as follows: an encoding method that satisfies GC global balance and run-length limitations while correcting burst insertion and deletion errors. The method specifically includes: Step 1: Receive the quaternion information sequence to be stored. Its symbol comes from the quadruple alphabet. .

[0010] Step 2, for the four-element information sequence Perform a difference transform to obtain the difference sequence. ; in the difference sequence Add an auxiliary symbol 0 to the end to obtain the extended difference sequence; detect and delete the first occurrence of a sequence with a length of 0 in the extended difference sequence. consecutive zero-signed substrings Then, add a position identifier indicating the start position of the substring to the end of the sequence; repeat the above detection, deletion, and addition steps until the resulting extended difference sequence does not contain any substrings. Subsequently, an inverse difference transform is performed on the resulting extended difference sequence to obtain the sequence that satisfies... - Middle codewords for run length limitations .

[0011] Step 3, define a flip function for converting symbols between the non-GC symbol set and the GC symbol set; for the intermediate codeword Search for a balanced index i in the index set determined by zero, the intermediate codeword length, and a preset step size, such that for the intermediate codeword... The sequence obtained after applying the flip function to the first i symbols satisfies the preset GC global balance deviation condition; the balance index i is represented as a quadruple index sequence, and the flip function is applied to the quadruple index sequence symbol by symbol to obtain a flipped index sequence; the quadruple index sequence and the flipped index sequence are interleaved bit by bit to generate an interleaved index segment, which itself satisfies strict GC balance and the run length does not exceed the limit. .

[0012] Step 4: Based on the value of the balanced index i, insert a disruptive symbol during the combination of the intermediate codeword after prefix flipping and the interleaved index segment to prevent lengths exceeding the specified values ​​from occurring at the flip boundary, end access position, or index segment access position. The same-sign consecutive runs; and two balance supplementary signs are appended to the end of the sequence to form a sequence that simultaneously satisfies the GC global balance deviation condition and -Complex constraint codewords with run-length limits .

[0013] Step 5, based on the composite constraint codeword Calculate the outer auxiliary information used for decoding and verification. The outer auxiliary information Encode the quaternary auxiliary sequence; perform the flip function on the quaternary auxiliary sequence to obtain a flipped auxiliary sequence, and interleave the quaternary auxiliary sequence and the flipped auxiliary sequence bit by bit to generate an interleaved auxiliary tail segment. The interlacing auxiliary tail section It satisfies strict GC balance and the run length does not exceed .

[0014] Step 6, select ordered symbol pairs ,in ,and and These belong to different sets within the GC symbol set and the non-GC symbol set, respectively; the ordered symbol pairs are repeated continuously and alternately. t / 2 Next, forming a buffer zone The buffer zone This is used to assist in determining the affected region of sudden insertion / deletion errors, and to ensure that the buffers themselves satisfy strict GC balance. -Run length limit.

[0015] Step 7, select a pair of connector symbols, the pair of connector symbols being derived from the quadratic alphabet. Remove the ordered symbol pairs from The last two symbols constitute the composite constraint codeword. The last symbol determines the order of the concatenation symbol pair, making the first concatenation symbol different from the composite constraint codeword. The last symbol, the second concatenation symbol is different from the first concatenation symbol, and the second concatenation symbol is different from the buffer space. The first symbol.

[0016] Step 8, convert the composite constraint codeword The connection symbol pair, the buffer space and the interlaced auxiliary tail section By sequentially piecing them together, the final code can be obtained. The final codeword Simultaneously satisfy -GC global balancing and - Run length limit, and able to correct for a length not exceeding Sudden insertion error or sudden deletion error.

[0017] Step 9: Map the final codeword according to the preset base mapping rules. Each quaternion in the sequence is mapped to a DNA base to obtain a DNA base sequence; in one system implementation, the DNA base sequence is provided to a DNA synthesizer to perform synthesis.

[0018] Furthermore, the decoder determines the type of burst insertion / deletion error based on the difference between the received sequence length and the final codeword length; and recovers the outer auxiliary information based on the alignment status between buffers and the alignment status of the interleaving auxiliary tail segment. Composite constraint codewords are recovered by enumerating finite candidate sequences and verifying with outer auxiliary information. Subsequently, the inverse operations of GC global balancing precoding and run-length limiting precoding are performed sequentially to restore the original quaternion information sequence. .

[0019] Furthermore, representing the balanced index as a quadruple index sequence and generating interleaved index segments specifically includes: The index representation length is determined based on the allowed GC global balance deviation. Map the balanced index to a length equal to The quadruple sequence is obtained; the flip function is applied to the quadruple sequence symbol by symbol to obtain the flipped sequence; the quadruple sequence and the flipped sequence are arranged alternately by symbol, so that the symbol of the quadruple sequence is located in the odd position and the symbol of the flipped sequence is located in the even position, forming the interleaved index segment.

[0020] Furthermore, the outer auxiliary information includes positive integers. Sum of remainders ,in To make the composite constraint codeword Input the preset integer value function modulo The remainder obtained; the positive integer. The preset integer value function is set as follows: for values ​​of the same length not exceeding The different candidate sequences corresponding to the burst deletion output, and their corresponding modulo The remainders are different.

[0021] Furthermore, the insertion of the disruption symbol and the addition of the balance supplement symbol include: When the balanced index is greater than zero and less than the length of the intermediate codeword, a first disruptive symbol is inserted between the prefix flipped portion and the unflipped suffix. This first disruptive symbol is different from the last symbol of the prefix and the first symbol of the suffix. A second disruptive symbol is inserted between the resulting sequence and the interleaving index segment. This second disruptive symbol is different from its immediate preceding symbol and the first symbol of the interleaving index segment. A first balanced supplementary symbol and a second balanced supplementary symbol are appended to the end of the sequence. The GC assignment of the first balanced supplementary symbol is opposite to that of the first disruptive symbol and different from that of its immediate preceding symbol. The GC assignment of the second balanced supplementary symbol is opposite to that of the second disruptive symbol and different from that of the first balanced supplementary symbol.

[0022] When the balance index is zero, a first break symbol is inserted before the intermediate codeword. This first break symbol is different from the first and last symbols of the intermediate codeword. The subsequent operations are the same as when the balance index is greater than zero and less than the length of the intermediate codeword.

[0023] When the balance index is equal to the length of the intermediate codeword, a first destruction symbol and a second destruction symbol are inserted sequentially between the fully flipped codeword and the interleaving index segment. The first destruction symbol is different from the last symbol of the flipped codeword, and the second destruction symbol is different from the first destruction symbol and the first and second symbols of the interleaving index segment. The subsequent operations are the same as when the balance index is greater than zero and less than the length of the intermediate codeword.

[0024] Another object of the present invention is to provide a decoding method for recovering the original quaternary information sequence from a DNA sequencing signal, wherein the DNA sequencing signal originates from a received sequence obtained by DNA synthesis, storage, and sequencing according to the generated final codeword sequence, characterized in that it includes: The received sequence is obtained, and the difference between its length and the preset length of the final codeword sequence is calculated. A positive difference is used to determine burst insertion, a negative difference is used to determine burst deletion, and a zero difference is used to determine no burst error.

[0025] The positions between the buffers are identified in the received sequence, the alignment status between the buffers is determined by the alternating repetition pattern of the marker symbol pairs, and the alignment of the interleaving auxiliary tail window at the end of the sequence is determined by detecting whether the odd-numbered symbols are equal to the adjacent even-numbered symbols after the flipping function.

[0026] When the interleaving auxiliary tail segments are aligned, the odd-numbered symbols of the tail segment window are extracted to form a quaternion auxiliary sequence, and the outer auxiliary information is decoded and recovered.

[0027] If the error type is burst insertion and the interleaving auxiliary tail segment is aligned, the insertion length t* is determined by the length difference, and a portion of the received sequence from the beginning to the (n+t*)th symbol is extracted, where n is the preset length of the composite constraint codeword. A candidate sequence is generated by deleting all possible consecutive t* symbols from this portion, and the first n symbols of the received sequence are used as supplementary candidates. The candidate sequence is verified using the outer auxiliary information, and the only candidate that satisfies the verification is used as the composite constraint codeword.

[0028] If the error type is burst deletion and the interleaving auxiliary tail segment is aligned, the deletion length t* is determined by the absolute value of the length difference. The first n symbols of the received sequence are first checked as candidates. If they fail, the first nt* symbols of the received sequence are truncated. Candidate sequences are generated by inserting all possible quadruple substrings of length t* at all possible positions. The outer auxiliary information is used for verification to uniquely determine the composite constraint codeword.

[0029] If the interleaving auxiliary tail segment is not aligned, and the buffers are aligned within the expected range and the first n symbols of the received sequence satisfy the GC global balance deviation condition and When run length is limited, the first n symbols of the received sequence are directly extracted as the composite constraint codeword.

[0030] Remove the last two balanced supplementary symbols from the determined composite constraint codeword, extract the odd-numbered bits of the interleaved index segment to recover the quad-index sequence, and parse to obtain the balanced index; identify and delete the inserted disruptive symbols according to the balanced index, perform the inverse operation of the flip function on the corresponding prefix, and recover the intermediate codeword.

[0031] Perform a differential transform on the intermediate codeword, and sequentially read position identifiers from the end of the resulting sequence. The insertion length is determined based on the position identifiers. The continuous zero-symbol substring is processed until a predetermined tail identifier is detected, and then the inverse differential transformation is performed to obtain the original quaternion information sequence.

[0032] Furthermore, the candidate sequence generation step in the case of sudden insertion further includes: Let the insertion length t* be equal to the length difference, and take the first n+t* symbols of the received sequence to form the source sequence.

[0033] For i from 0 to n, candidate sequences are obtained by deleting t* consecutive symbols starting from position i+1 in the source sequence, and the original first n symbols of the received sequence are taken as one of the candidates.

[0034] Furthermore, the candidate sequence generation step under the sudden deletion scenario further includes: Let the deletion length t* be equal to the absolute value of the length difference. If the first n symbols of the received sequence fail the verification, then take the first nt* symbols of the received sequence as the base sequence.

[0035] For i from 0 to nt*, in all sequences of length t* of the four-letter alphabet, each sequence is inserted after the i-th symbol of the base sequence to form a candidate sequence, and the outer auxiliary information is used for verification.

[0036] Another object of the present invention is to provide a DNA data storage system, comprising: The encoder is configured to receive a quaternary information sequence to be stored, execute an encoding method, and output a DNA base sequence.

[0037] A DNA synthesizer, connected to the encoder, is used to receive the DNA base sequence and synthesize the corresponding DNA molecule.

[0038] A DNA sequencer is used to sequence the stored DNA molecules and output the received sequence.

[0039] A decoder, connected to the DNA sequencer, is configured to receive the received sequence and execute the decoding method to recover the original quaternary information sequence.

[0040] Furthermore, the encoder includes: The run-length precoding unit is used to perform differential transformation, consecutive zero substitution and position identifier appending, and inverse differential transformation to output the intermediate codeword.

[0041] The GC balancing coding unit is used to search for the balancing index, perform prefix flipping, generate the interleaved index segment, insert the breaking symbol and the balancing supplement symbol, and output the composite constraint codeword.

[0042] An auxiliary tail segment generation unit is used to calculate the outer auxiliary information and generate the interleaving auxiliary tail segment.

[0043] The codeword concatenation unit is used to construct the buffer space, select the connection symbol pairs, and concatenate them to obtain the final codeword sequence.

[0044] Furthermore, the decoder includes: The error detection unit is used to determine the error type based on the difference between the received sequence length and the preset length.

[0045] The window detection unit is used to detect the alignment between buffers and the alignment of the interleaving auxiliary tail segment, and to recover the outer auxiliary information.

[0046] The candidate enumeration verification unit is used to construct a set of candidate sequences under sudden errors and to perform verification using outer auxiliary information to determine the composite constraint codeword.

[0047] The inverse precoding unit is used to recover the balance index from the composite constraint codeword, delete the disruptive symbol and the balance supplement symbol, perform inverse flipping, and recover the original quaternary information sequence through inverse run-length precoding.

[0048] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows: The coding method proposed in this invention addresses the data read / write failure problem caused by biochemical constraints and burst errors in the entire DNA data storage process. Through an integrated design at the coding level, a single codeword sequence can simultaneously satisfy two major biochemical constraints: global GC balance and run length limitation. It also has the ability to correct a burst insertion or deletion error whose length does not exceed a preset threshold, thereby directly improving the integrity and reliability of data read / write in the physical synthesis, preservation, and sequencing stages.

[0049] In terms of biochemical constraint assurance, this method uniformly maps quaternary information sequences to sequences with GC content strictly controlled within a preset deviation range. This eliminates uneven DNA strand hybridization efficiency and sequencing coverage depth shifts caused by local GC content distortion, ensuring the uniformity of synthesis and sequencing processing. Simultaneously, through continuous zero substitution and position identifier recovery mechanisms, the length of any consecutive runs with the same symbol in the sequence is limited to within a preset threshold, fundamentally preventing the formation of homopolymer regions. Homopolymers are a major biochemical cause of sequencing signal slippage, phase misalignment, and insertion / deletion errors. Run limitation eliminates them at the sequence level, thus significantly reducing the coding failure rate caused by unmet biochemical constraints after sequencing, and reducing the time and reagent costs associated with data rewriting and physical resynthesis.

[0050] In terms of burst error correction, this method embeds outer auxiliary information into the codeword tail segment in a self-balancing manner and constructs a specific alternation pattern between buffers. This allows the decoder to quickly identify the error type solely through changes in the length of the received sequence and the window alignment state, without relying on external indexes, and uniquely recover the original codeword by verifying a limited number of candidate sequences. This mechanism can correct any single burst insertion or deletion error introduced during synthesis, storage, and sequencing, with a length not exceeding a preset threshold. Because the correction targets consecutive symbols that are lost or redundant in segments, rather than just single symbol errors, it has the direct ability to correct physical errors commonly found in DNA storage, such as long fragment copy slip and missing nanopore sequencing signals, effectively suppressing error propagation and improving the integrity of information recovery.

[0051] This solution integrates biochemical adaptation and physical-level error correction into the coding design, enabling DNA storage systems to robustly cope with the dual challenges of biochemical constraints and sudden insertions and deletions while maintaining high coding efficiency. It improves the success rate of data writing and reading from the source, and has substantial technical effects on promoting the practical application and large-scale deployment of DNA data storage.

[0052] First, this invention establishes an encoding process that first performs biochemical constraint precoding, and then connects to an outer burst insertion / deletion error correction structure. Specifically, this invention first encodes the original quaternion information sequence into intermediate codewords that satisfy run-length constraints, and then further obtains codewords that simultaneously satisfy... -GC global balancing and - A quaternary pre-coded codeword with run-length constraints; subsequently, this quaternary pre-coded codeword is used as input to the outer error correction structure to calculate auxiliary information. Thus, the outer error correction structure does not directly act on the unconstrained original information sequence, but rather on the pre-coded codeword that has already met the biological constraints, thereby avoiding the problem of destroying the pre-constrained coding logic after redundant error correction access.

[0053] Second, this invention, through the coordination of buffer spaces, connection symbols, and interleaving auxiliary tail segments, ensures that the final codeword maintains GC global balance and run length limitations even after incorporating outer auxiliary information. The buffer spaces employ an alternating symbol structure, which assists the decoder in determining the affected region of sudden insertion / deletion errors. Connection symbols are adaptively selected based on the last symbol of the quaternary precoded codeword and the symbols in the buffer spaces to avoid excessively long consecutive runs of the same symbol at connection positions. The interleaving auxiliary tail segment is obtained symbol-by-symbol interleaving of the outer auxiliary sequence and its fully inverted sequence, inherently satisfying strict GC balance, and... Timely satisfaction - Run length limitation. Therefore, the final codeword can simultaneously take into account the recoverability of outer layer error correction auxiliary information and the basic biochemical constraints in DNA storage.

[0054] Third, this invention can correct a length not exceeding This invention addresses sudden insertion or deletion errors. For sudden insertion errors, the decoder can uniquely determine the original quaternary precoded codeword in the candidate sequence based on changes in the received sequence length, buffer positions, and outer auxiliary information recovered from the interleaving auxiliary tail segment. For sudden deletion errors, the decoder can determine whether the quaternary precoded codeword is affected based on the alignment relationship between buffers and the interleaving auxiliary tail segment, and if affected, recover the deleted substring using the outer auxiliary information. Therefore, this invention, while satisfying GC global balance and run-length limitations, further provides recovery capabilities against sudden insertion and deletion errors.

[0055] Fourth, the decoding and encoding processes of this invention have a clear inverse correspondence. The encoding end sequentially executes run-length-limited precoding, GC global balancing precoding, outer auxiliary information generation, and final codeword construction; the decoding end first recovers the quaternary precoded codeword based on the final codeword structure, then performs the inverse operation of GC global balancing precoding to recover intermediate codewords, and finally performs the inverse operation of run-length-limited precoding to recover the original quaternary information sequence. This structure helps reduce ambiguity in the decoding process and improves the feasibility and clarity of the coding system's engineering deployment.

[0056] Fifth, the final codewords generated by this invention simultaneously satisfy GC global balance and run length constraints, which helps to reduce the adverse effects of unsuitable GC content and long consecutive runs with the same symbol on DNA synthesis, amplification, and sequencing processes. Compared with encoding methods that only consider error correction capability without simultaneously constraining GC content and run length, this invention is more suitable for application scenarios in DNA storage that require both biochemical constraints and burst insertion / deletion error correction.

[0057] Furthermore, as supporting evidence of the inventiveness of this invention, it is also reflected in the following important aspects: (1) The technical solution of this invention has good application and transformation value. DNA storage systems need to simultaneously consider data recovery reliability, sequence biochemical adaptability, and encoding / decoding feasibility. This invention addresses these requirements by providing an encoding scheme that balances global GC balance, run-length limitations, and burst insertion / deletion error correction capabilities. This scheme can be used in scenarios such as long-term archive data storage, biomedical data storage, and high-density molecular storage. This scheme helps improve the reliability and practicality of DNA storage encoding systems and provides a reusable encoding structure for related system designs.

[0058] (2) This invention addresses the problem that constraint satisfaction and burst insertion / deletion error correction are difficult to simultaneously achieve in existing DNA storage coding schemes, and proposes a unified construction method. In existing schemes, one type of method mainly focuses on biochemical constraints such as GC balance and run length limits, while another type of method mainly focuses on insertion / deletion error correction capabilities. If error correction redundancy is directly appended to the constrained codeword, it is easy to disrupt GC balance or run length limits; if error correction coding is performed first and then constraint processing is performed, it may affect the decodeability of the synchronous error correction structure. This invention constructs constrained quadratic pre-coding codewords first, and then appends buffer intervals, connection symbols, and interleaving auxiliary tail segments to them, so that constraint preservation and burst insertion / deletion error correction are in the same coding framework.

[0059] (3) This invention resolves the structural conflict between "biochemical constraint preservation" and "synchronization error recovery" in DNA storage coding. Sudden insertion and deletion errors can disrupt the symbol alignment of sequences, making it difficult for the decoder to directly determine the affected position; at the same time, DNA storage codewords need to meet GC balance and run length constraints, and cannot arbitrarily add redundancy. This invention uses buffer-assisted localization of the error-affected region, uses interleaved auxiliary tail segments to preserve outer auxiliary information, and uses connection symbols to avoid excessively long runs at splicing positions, thereby coordinating constraint preservation, auxiliary information readability, and sudden insertion and deletion recovery in a unified structure.

[0060] (4) This invention does not simply use existing constraint coding and burst deletion error correction methods side by side, but rather redesigns the interface between the two. Specifically, this invention limits the outer auxiliary information to those that already satisfy... -GC global balancing and The run-length-constrained quadratic precoding codeword is used as input for computation, and this auxiliary information is constrained and incorporated into the final codeword through interleaving auxiliary tail segments and buffers. This structure avoids GC bias and run-length corruption that may be caused by direct appending of outer redundancy, enabling the final codeword to maintain biological constraints while possessing burst insertion and deletion recovery capabilities. Attached Figure Description

[0061] Figure 1 This is a flowchart of an encoding method for GC global balancing and run-length constraint that can correct sudden insertion and deletion errors, provided in an embodiment of the present invention.

[0062] Figure 2 The construction provided by the embodiments of the present invention is from To the set of composite constraints Flowchart of the GC global balanced precoding mapping method.

[0063] Figure 3 This is a block diagram of a coding system module that provides GC global balance and run-length constraints and can correct burst insertion and deletion errors, as provided in the embodiments of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0065] like Figure 1 As shown, the encoding method for GC global balancing and run-length constraint that can correct burst insertion and deletion errors provided in this embodiment of the invention specifically includes: S1, constructing from the original set of four-element information sequences To the run length limit set The precoding mapping is obtained to satisfy - Middle codewords for run length limitations .

[0066] S2, based on the run-length limitation precoding, further construct from To the set of composite constraints The GC globally balanced precoding map obtains the result that simultaneously satisfies -GC global balancing and - Quadruple precoding codewords with run length limitations .

[0067] S3, using the quadratic precoded codeword As input to the outer error correction structure, it is used to compute the quadratic precoding codeword used to uniquely determine the quadratic precoding codeword in burst insertion / deletion decoding. outer auxiliary information The outer auxiliary information is then converted into an interleaved auxiliary tail segment that satisfies GC balance and run-length constraints. .

[0068] S4, the quadratic precoded codeword Connecting symbols, buffer zones, and interleaving auxiliary tail segments Combine them in sequence to construct the final codeword. , making the final code Simultaneously satisfy -GC global balancing - Run length limit, and able to correct for a length not exceeding Sudden insertion error or sudden deletion error.

[0069] S5 is designed for the final codeword set. The decoding method involves the decoder determining the error type based on changes in the length of the received sequence, recovering the outer auxiliary information based on the interleaved auxiliary tail segment, and uniquely recovering the quaternary precoded codeword through finite candidate sequence enumeration and verification using the outer auxiliary information. Then, the inverse operations of GC global balancing precoding and run-length limiting precoding are performed sequentially to restore the original input sequence. .

[0070] The embodiment of the present invention provides S1 specifically including: S101, Define the quad alphabet ,in and use Indicates by The length of the elements in the middle is The set of all quaternions. Define the set of DNA nucleotides. And establish a one-to-one mapping from the quaternary alphabet to the DNA nucleotide set. In one specific implementation, the following is adopted: The mapping method. Among them, the differential transformation, continuous zero substring replacement and position identifier recovery mechanism can be implemented using the DNA constraint coding idea disclosed in reference [1]. In this embodiment of the invention, the precoding process is used as a run-length constraint precoding interface to convert the original quaternary information sequence into a sequence that satisfies the constraint coding method. - The intermediate codeword with run length limitation.

[0071] S102, Let the input quadruple sequence be... ,in It is an odd number. (The last part is a repetition of the previous one and can be omitted.) Perform a difference transform to obtain the difference sequence. ,in = ;when hour, The corresponding inverse difference transform is denoted as... ,in , .

[0072] S103, in the difference sequence Adding an auxiliary symbol 0 to the end yields the extended difference sequence. ,at this time .remember ,but It is an even number.

[0073] S103.1, if There is no continuous occurrence in it A subsequence of zeros, i.e. Not included ,in If the condition is not met, the replacement operation will not be performed, and the output will be directly changed. At this point, due to the difference sequence There is no continuity in The output sequence contains zeros, therefore... There is no length exceeding consecutive runs with the same sign, thus Belongs to the run length limit code set .

[0074] S103.2, if There exists a subsequence If so, a replacement operation is performed on the current extended difference sequence. Specifically, let the current extended difference sequence be represented as... ,in For the first occurrence of a consecutive A substring of zeros, starting at position 0 ,and Delete this consecutive string. A substring of zeros is generated, and a position identifier is added to the end of the sequence. ,in , ,and Starting position The quaternion representation. Repeat the above replacement operation until the resulting difference sequence no longer contains. Final output and define Therefore, the output sequence is... .

[0075] S104, to make the position identifier Capable of covering all potential replacement locations, in In the case of sequence length With run length limit parameter Satisfy As a specific implementation method, when At that time, the above relationship gives The theoretical upper limit is 3077. Furthermore, when... and At that time, for The codewords can be replaced using the above method to achieve run-length limit precoding.

[0076] like Figure 2 As shown, S2 in this embodiment of the invention specifically includes: S201, in the mapping method, quaternion symbols 0 and 1 correspond to DNA bases A and T, respectively, and quaternion symbols 2 and 3 correspond to DNA bases C and G, respectively.

[0077] Therefore, the set of non-GC symbols is defined as follows: Define the GC symbol set as The prefix flipping, balanced indexing, index representation and its fully flipped interleaving structure involved in this step and subsequent S202 to S206 can be implemented using the GC content balancing processing idea disclosed in reference [1]; the embodiment of the present invention further unifies the rule of breaking the symbol insertion and the access method of the outer burst insertion and deletion structure into the same encoding process.

[0078] S202, Define the flip function: This is used to perform conversions between non-GC symbols and GC symbols. Specifically, let... .thus, Will The symbols in the map are mapped to In, and will The symbols in the map are mapped to In the middle. For any sequence and index ,in ,definition To The former Symbol application The resulting sequence; when hour, ;when hour, Indicates to Flip all the symbols.

[0079] S203, defines GC weights; For any sequence Define its GC weight as ,Right now China belongs to The number of symbols. Let. ,like satisfy Then it is called satisfy -GC global balancing condition.

[0080] S204, Let the balanced index set be: Let the balanced index set be Of which, only a maximum of The elements obtained from S1. In the set Middle search balanced index ,make satisfy -GC global balancing condition. If multiple indexes meet the condition, the first index is selected as the GC global balancing condition. of - Balanced index. (Note) .

[0081] S205, for balanced index Perform quaternion representation: make Balance the index set Each index in the array uniquely maps to a string of length . A four-element sequence. Let the index be... The corresponding quad index sequence is ,in , Perform a complete reversal on each symbol of the quad index sequence to obtain... Subsequently, the quad-index sequence and its fully inverted sequence are interleaved bit by bit according to their signs to obtain the interleaved index segment. Because of each pair of symbols and There just happens to be one that belongs to The other belongs to ,therefore The length is GC weight is Therefore It inherently satisfies strict GC balancing; at the same time, due to ,and Internally, no consecutive runs of the same sign exceeding 2 will be generated. hour, satisfy -Run length limit.

[0082] S206, Perform sequence encoding: Based on the balanced index The value of will With interleaved index segments Combine the elements and insert break symbols at positions where the run-length limit might be violated to obtain the desired result. -GC global balancing and - Quadruple precoding codewords with run length limitations Specifically, let the first destruction symbol be... The second destruction symbol is ; in forming the second intermediate sequence Afterwards, Add two balance supplement symbols at the end and . and Select according to the following rules: If Then select ;like Then select and demand Not equal to The last sign; if Then select ;like Then select and demand .because and Both contain two symbols, the above and Both can be selected.

[0083] S206.1, when and At that time, I recorded for The length is prefix, for The length is The suffix. Because and All internal aspects are maintained - Travel length limit, but the connection point between the two may generate more than consecutive runs of the same sign, therefore in and Insert the first destruction symbol between .set up for The last sign, for The first symbol, select The first intermediate sequence is obtained. Furthermore, in With interleaved index segments Insert a second destruction symbol between .set up for The last sign, for The first symbol, select The second intermediate sequence is obtained. According to the rules described in S206, in Add two balance supplement symbols at the end and Output The sequence simultaneously satisfies -GC global balancing and -Run length limit.

[0084] S206.2, when hour, There is no internal join boundary between the flipped prefix and the non-flipped suffix. To maintain a consistent codeword structure and avoid violating run-length constraints at subsequent joins, in Insert the first destructive symbol at the front end .set up for The first symbol, select The first intermediate sequence is obtained. Furthermore, in With interleaved index segments Insert a second destruction symbol between .set up for The last sign, for The first symbol, select The second intermediate sequence is obtained. According to the rules described in S206, in Add two balance supplement symbols at the end and Output The sequence simultaneously satisfies -GC global balancing and -Run length limit.

[0085] S206.3, when hour, There is no unreversed suffix. To avoid With interleaved index segments The connection point violates the travel length limit, at which and Insert the first destruction symbol between them. Second destruction symbol .set up for The last sign, for The first symbol, select , and select The second intermediate sequence is obtained. According to the rules described in S206, in Add two balance supplement symbols at the end and Output The sequence simultaneously satisfies -GC global balancing and -Run length limit.

[0086] S206.4, let Define the set of quaternary composite constraint codes. For all the quaternion precoded codewords output by the above encoding process The set, i.e. Therefore, the quadratic precoded codeword output by S2 belong .

[0087] The embodiment of the present invention provides S3 specifically including: S301, let the quaternion composite constraint codeword output by S2 be... ,in , Indicates length is Simultaneously satisfy -GC global balancing and - A set of quaternary composite constraint codes with run length restrictions.

[0088] S302, addressing the upper bound of burst length Construct outer auxiliary information to distinguish candidate sequences Specifically, constructing and quadratic precoding codewords. Related verification functions and positive integers and order Represent the outer auxiliary information as a pair of tuples. The decoding stage utilizes... and The original quadratic precoding codeword is uniquely determined by verifying the finite candidate sequences. .

[0089] S303, the outer auxiliary information is used during burst deletion decoding to supplement the received sequence with a length not exceeding [a certain value]. The candidate set formed by consecutive quaternion substrings is uniquely determined Simultaneously, the outer auxiliary information is also used for burst insertion decoding. For burst insertion errors, the decoder removes a sequence of lengths not exceeding [a certain value] from the received sequence. Candidate sequences are formed from consecutive substrings, and then... The candidate sequences are verified. Therefore, burst insertion recovery can be transformed into a deletion-based verification process with a finite number of candidate sequences.

[0090] S304, outer auxiliary information Convert to a fixed-length quaternion auxiliary sequence ,in A positive integer determined according to a preset quaternion representation rule. The quaternion auxiliary sequence... For storage and This enables the decoding end to... The outer auxiliary information is obtained by unique parsing in the middle. At the upper bound of the fixed burst length In one implementation, Satisfy .

[0091] S305, for Perform a full flip to obtain ,in The flip function defined in S2 satisfies Subsequently, and Interweave symbol by symbol to obtain interleaved auxiliary tail segments. .

[0092] S306, due to each pair of symbols and There is exactly one in the GC symbol set. Another one belongs to the non-GC symbol set. Therefore, the interlacing auxiliary tail section The length is GC weight is Therefore It inherently satisfies strict GC balancing. Meanwhile, because... Intertwined auxiliary tail section Internally, no consecutive runs of the same sign exceeding 2 will be generated; when hour, satisfy -Run length limit.

[0093] The embodiment of the present invention provides S4, which specifically includes: S401, construct buffer zone Select the marker pair and order ,in Indicates the symbol pair Continuous repetition Next. Because each symbol pair There is exactly one symbol in the middle that belongs to Another symbol belongs to Therefore, between buffers It inherently satisfies strict GC balance; at the same time, Therefore No internal length exceeding consecutive runs with the same symbol.

[0094] S402, Construct a pair of connection symbols and .set up ,Right now and Taken from a pair of symbols that do not belong to the symbol pair The other two quaternions. Based on the quaternion composite constraint codeword The last sign Sure and The order: if There is one and only one symbol that is not equal to Then the symbol is denoted as Another symbol is denoted as ;like Neither of the two symbols is equal to Then, one of them will be selected according to the preset order. Another as This guarantees ,and .

[0095] S403, converts the quadruple composite constraint codeword output by S2. , Connecting symbols , Connecting symbols Buffer Zone and the interleaved auxiliary redundancy segment obtained by S3 Combine them in order to obtain the final codeword. ,in .in, ensure and The connection position will not be extended. The final journey; ensure and The connection position does not form a length exceeding The same symbolic run length; and because ,and The first symbol is Therefore ,thereby and The connection point will not extend the run. Therefore, , , and The connection points will not produce lengths exceeding [a certain value]. consecutive runs with the same symbol.

[0096] S404, due to satisfy -GC global balancing, link symbol pairs There is exactly one symbol in the middle that belongs to , buffer Satisfying strict GC balancing, interleaved auxiliary redundant segments It also satisfies strict GC balance, therefore the final code... The GC bias only comes from .make Then the final code satisfy -GC global balancing condition.

[0097] S405, due to satisfy -Travel length limit, and , , , buffer Alternating symbol sequences, interleaved auxiliary redundant segments Each pair of symbols in the symbol set consists of two different symbols that are flipped together. , , , and Neither the internal structure nor the connection points will produce lengths exceeding [a certain value]. consecutive runs with the same sign. Specifically, and Even if the same symbol appears at the connection point, it will at most form a continuous run of the same symbol with a length of 2; due to the embodiments of the present invention The connection still meets the requirements. - Run length limit. Therefore, the final code... satisfy -Run length limit.

[0098] S406, Define the final global code set For all the final codewords obtained through the above method The set, i.e. Therefore, the final codeword output by S4 .

[0099] The embodiment of the present invention provides S5, which specifically includes: S501, Let the received sequence be... The decoder first calculates the length difference. ,in For the final typing The length. If If, under the sudden insertion / deletion error model considered in this embodiment of the invention, it is determined that no sudden insertion / deletion error has occurred; if If this occurs, a sudden insertion error is determined, and the insertion length is [value missing]. ; like If this occurs, a sudden deletion error is determined, and the deletion length is [value missing]. Since this embodiment of the invention considers a length not exceeding... The sudden insertion error or sudden deletion error, therefore satisfying .

[0100] S502, order Indicates the buffer space The length of the buffer. The decoder defines a marker window and a tail window. The marker window is used to detect buffer gaps. The alignment status, the tail window is used to detect the interlacing auxiliary tail segment. The alignment status. Since sudden insertion / deletion errors can cause window position shifts, the decoder can check the buffer within a preset offset range. The location of its appearance; for the tail window, the decoder prioritizes reading the end of the received sequence with a length of [missing information]. window . satisfy , This is called the tail window. Alignment; when aligning the tail window, recover the quaternion auxiliary information sequence from the odd-numbered position symbols. ,Right now , and by Parsing yields outer auxiliary information .

[0101] S503, when At that time, directly from the received sequence The former Extracting quaternary composite constraint codewords from symbols ,Right now Then, it enters S508 to perform the precoding inverse operation.

[0102] S504, when At that time, burst insertion decoding is performed. If the tail window If aligned, the outer auxiliary information is restored from the tail window. Then, a burst insertion candidate set is constructed. Specifically, let... , For each ,from Delete location to A continuous substring of length is obtained. candidate sequences At the same time, the received sequence will be... a symbol The error occurred in the quaternary compound constraint codeword. The subsequent candidate sequences. For all the above candidate sequences, perform outer auxiliary information verification; if a certain candidate sequence... satisfy Then the candidate sequence is used as the recovered quaternary composite constraint codeword. Based on the uniqueness of the outer auxiliary information, the candidate sequence that satisfies the verification condition is unique.

[0103] S505, when And the rear window When misalignment occurs, a sudden insertion error is considered that may affect the interleaving auxiliary tail. In this case, if the buffer... Consistent within the preset position range, and before receiving the sequence. The symbols satisfy If the GC global balance and run-length limit are met, then the quaternion composite constraint codeword is determined. Unaffected, output .

[0104] S506, when At that time, perform a sudden deletion decoding. If the tail window If aligned, the outer auxiliary information is restored from the tail window. Then, a burst deletion candidate set is constructed. Specifically, the previous sequence is first... a symbol The error occurred in the quaternary compound constraint codeword. The candidate sequence is then used, and its performance is checked. If satisfied, output... If not satisfied, then let For each And each ,Will Insert into The After [number] symbols, we get a length of [length]. candidate sequences If a candidate sequence satisfies Then the candidate sequence is used as the recovered quaternary composite constraint codeword. Based on the uniqueness of the outer auxiliary information, the candidate sequence that satisfies the verification condition is unique.

[0105] S507, when And the rear window When misalignment occurs, a sudden deletion error is considered that may affect the interleaving auxiliary tail segment. In this case, if the buffer... Maintain alignment within the preset position range, and before receiving the sequence. The symbols satisfy If the GC global balance and run-length limit are met, then the quaternion composite constraint codeword is determined. Unaffected, output .

[0106] S508, after completing the above steps, the decoding end obtains the recovered quaternion composite constraint codeword. Subsequently, regarding Perform the inverse operation of S2. Specifically, first delete the last two balancing supplementary symbols. and Then, based on the preset index segment length Extracting interleaved index segments .Depend on Odd-position sign recovery index quadruple sequence ,Depend on Even position sign recovery And it is verified by checking whether the sign at the even position is equal to the flip value of the sign at the corresponding odd position. Then by... Restore Balanced Index .

[0107] S509, based on the balanced index Delete the corrupted symbol inserted in S2 and perform a flip function on the corresponding prefix. The inverse operation. Because For any Established, therefore applied again. This is the reverse flip operation. If and Then delete the first destructive symbol between the reversed prefix and the unreversed suffix. and for the former The symbol is executed again. ;like Then delete the first corrupted symbol inserted at the front. If prefix reversal is not performed; Then delete the location located at With interleaved index segments The first destructive symbol between Second destruction symbol and for all previous The symbol is executed again. This restores the run-length limit intermediate codeword output by S1. .

[0108] S510, for the recovered intermediate codeword Perform the inverse operation of S1. Specifically, perform a difference transformation on y to obtain the expanded difference sequence after replacement; then, read the position identifiers sequentially from the end of the expanded difference sequence, and restore the deleted continuous sequence according to the positions indicated by the position identifiers. Substring. Repeat this process until the terminal auxiliary symbol is 0; after deleting the auxiliary symbol, perform an inverse difference transform on the resulting difference sequence to recover the original quaternary information sequence. .

[0109] like Figure 3 As shown, the coding system provided in this embodiment of the invention, which features GC global balance and run-length constraints and can correct burst insertion and deletion errors, specifically includes: The run-length-limited precoding module is used to process the original quaternion information sequence. Perform differential transformation, consecutive zero substring replacement, and position identifier appending operations, and obtain the desired result through inverse differential transformation. - Middle codewords for run length limitations .

[0110] The GC globally balanced precoding module is used in the intermediate codeword Based on this, perform prefix reversal, balanced index encoding, complete index sequence reversal and interleaving, disruptive symbol insertion, and balanced supplementary symbol addition operations to obtain a result that simultaneously satisfies... -GC global balancing and - Quadruple precoding codewords with run length limitations .

[0111] The outer auxiliary information generation module is used to generate the four-element precoded codeword. As input, calculate outer auxiliary information used to locate and recover sudden insertion / deletion errors. and will Convert to quaternary auxiliary sequence .

[0112] The interleaving-assisted tail generation module is used to generate the quaternary auxiliary sequence. Perform a complete flip and The completely reversed sequence is interleaved sign-by-sign to obtain a sequence that satisfies strict GC balance and satisfies - Interlacing auxiliary tail section with travel length limitation .

[0113] The final codeword construction module is used to construct the buffer space. and connection symbol pair , and the quadratic precoded codeword , Connecting symbols , Connecting symbols Buffer Zone and interwoven auxiliary tail section Combine them in order to obtain the final codeword. ,in Simultaneously satisfy -GC global balancing and - Run length limit, and able to correct for a length not exceeding Sudden insertion error or sudden deletion error.

[0114] The decoding module is used to decode the received sequence. Length variation, buffer space Position and interlacing auxiliary tail section Outer auxiliary information recovered from Recover quaternary precoded codewords Then, the inverse operations of GC global balancing precoding and run-length limiting precoding are performed sequentially to recover the original quaternion information sequence. .

[0115] This invention provides an encoding method that incorporates global GC balance and run-length constraints, and can correct burst insertion and deletion errors. The method takes a 195-byte original four-ary information sequence as input and sequentially performs run-length-constrained precoding, global GC balance precoding, outer auxiliary information generation, final codeword construction, and decoding recovery steps, specifically including: Step 1: Construct a structure that satisfies the maximum run length. Quadruple run length limit code Specifically, it includes: 1.1, Define the quad alphabet ,in and use Indicates by The set of all quaternions of length 195, consisting of elements in the set. Define the set of DNA nucleotides. And establish a one-to-one mapping from the quaternary alphabet to the DNA nucleotide set. In this embodiment, the following is adopted: The mapping method is as follows. Therefore, quaternions 2 and 3 correspond to DNA bases C and G respectively, and belong to GC symbols; quaternions 0 and 1 correspond to DNA bases A and T respectively, and belong to non-GC symbols.

[0116] 1.2, Let the original quaternion information sequence be... .right Perform a difference transform to obtain the difference sequence. ,in ;when hour, The corresponding inverse difference transform is denoted as... ,in , .

[0117] 1.3, in difference sequences Adding an auxiliary symbol 0 to the end yields the extended difference sequence. ,at this time .remember ,but It is an even number.

[0118] 1.3.1, if There is no subsequence containing four consecutive zeros, i.e. Not included If the replacement operation is not performed, the intermediate codeword will be output directly. At this point, due to the extended difference sequence The sequence does not contain four consecutive zeros, therefore the output sequence is... There are no consecutive runs of the same sign with a length exceeding 4, therefore Belongs to the run length limit code set .

[0119] 1.3.2, if There exists a subsequence If so, a replacement operation is performed on the current extended difference sequence. Specifically, let the current extended difference sequence be represented as... ,in The first occurrence of a consecutive string of four zeros, starting at position 1. ,and Delete the four consecutive zero substrings and add a position identifier to the end of the sequence. ,in , ,and Starting position The quaternion representation. Repeat the above replacement operation until the resulting extended difference sequence no longer contains. The final output is the intermediate codeword. and define Therefore, the output sequence is... .

[0120] 1.4, to make the location identifier Capable of covering all potential replacement locations, in and In the case of sequence length Satisfy Therefore, this embodiment takes... At that time, location identifier This can cover all available starting positions in the above replacement operation. Because... The resulting intermediate codewords meet the run length constraints commonly used in DNA storage.

[0121] Step 2, construct a structure that simultaneously satisfies Run-length limit and 0.05-GC global balance quaternary composite constraint code Specifically, it includes: 2.1, Define the set of non-GC symbols Define the GC symbol set Define the flip function. This is used to perform conversions between non-GC symbols and GC symbols. Specifically, let... .thus, Will The symbols in the map are mapped to In, and will The symbols in the map are mapped to middle.

[0122] 2.2, for any sequence and index ,in ,definition To The former The symbol applies a flip function. The resulting sequence. When hour, ;when hour, Indicates to Flip all the symbols.

[0123] 2.3, for any sequence Define its GC weight as ,Right now China belongs to The number of symbols. When At that time, if satisfy Then it is called The 0.05-GC global balance condition is met.

[0124] 2.4, Let the balanced index set be... Only elements with a maximum number of 196 are retained. Because... Therefore For the result obtained in step 1 In the set Middle search balanced index ,make The 0.05-GC global balancing condition is met. If multiple indexes meet the condition, the first one is selected as the index. The 0.05-balanced index, and recorded .

[0125] 2.5, Balanced Index Perform quaternion representation. Let Balance the set of indices. Each index in the array uniquely maps to a quaternion sequence of length 2. Let the index be... The corresponding quad index sequence is ,in Perform a complete reversal on each symbol of the quad index sequence to obtain... Subsequently, the quad-index sequence and its fully inverted sequence are interleaved bit by bit according to their signs to obtain the interleaved index segment. Because of each pair of symbols and There just happens to be one that belongs to The other belongs to ,therefore The length is 4, and the GC weight is 2, therefore It inherently satisfies strict GC balancing; at the same time, due to ,in ,and Internally, no consecutive runs of the same sign exceeding 2 will be generated. hour, satisfy -Run length limit.

[0126] 2.6, Perform sequence encoding. Based on the balanced index. The value of will With interleaved index segments By combining the results and inserting violation symbols at locations where run-length limits might be violated, a global balance satisfying 0.05-GC is obtained. Quadruple precoding codewords with run length constraints Specifically, let the first destruction symbol be... The second destruction symbol is ; in forming the second intermediate sequence Afterwards, Add two balance supplement symbols at the end and . and Select according to the following rules: If Then select ;like Then select and demand Not equal to The last sign; if Then select ;like Then select and demand .because and Both contain two symbols, the above and Both can be selected.

[0127] 2.6.1, when and At that time, I recorded for The length is prefix, for The length of ′ is The suffix. Because and All internal aspects are maintained There is a run length limit, but the connection between the two may result in more than 4 consecutive runs of the same sign, therefore in and Insert the first destruction symbol between .set up for The last sign, for The first symbol, select The first intermediate sequence is obtained. Furthermore, in With interleaved index segments Insert a second destruction symbol between .set up for The last sign, for The first symbol, select The second intermediate sequence is obtained. According to the rules described in section 2.6, in Add a balance supplement symbol at the end and Output This sequence simultaneously satisfies 0.05-GC global balance and Travel length limit.

[0128] 2.6.2, when hour, There is no internal join boundary between the flipped prefix and the non-flipped suffix. To maintain a consistent codeword structure and avoid violating run-length constraints at subsequent joins, in Insert the first destructive symbol at the front end .set up for The first symbol, select The first intermediate sequence is obtained. Furthermore, in With interleaved index segments Insert a second destruction symbol between .set up for The last sign, for The first symbol, select The second intermediate sequence is obtained. According to the rules described in section 2.6, in Add a balance supplement symbol at the end and Output This sequence simultaneously satisfies 0.05-GC global balance and Travel length limit.

[0129] 2.6.3, when hour, There is no unreversed suffix. To avoid With interleaved index segments The connection point violates the travel length limit, at which 'and Insert the first destruction symbol between them. Second destruction symbol .set up for The last sign, for The first symbol, select , and select The second intermediate sequence is obtained. According to the rules described in section 2.6, in Add a balance supplement symbol at the end and Output This sequence simultaneously satisfies 0.05-GC global balance and Travel length limit.

[0130] 2.6.4, Define the set of quaternary composite constraint codes For all the quaternion precoded codewords output by the above encoding process The set, i.e. Therefore, the quadruple precoded codewords output in step 2... belong .

[0131] Step 3, using the quadratic precoding codewords output in Step 2 As input to the outer error correction structure, the outer auxiliary information is calculated and constrained, specifically including: 3.1 Let the quadratic precoding codeword output in step 2 be... The outer error correction structure is based on Instead of the original four-element information sequence, the input is... For input.

[0132] 3.2 Calculate the outer auxiliary information used to recover from sudden deletion errors. The outer auxiliary information The method disclosed in reference [2] can be used. The idea of ​​meta-burst deletion error correction is implemented. Specifically, there is a mapping. and with Corresponding positive integer ,make For any and Distinct sequences with a common length not exceeding 10 as the output of burst deletion All have Different from others .therefore, It can be represented as a pair Used to uniquely determine from candidate sequences affected by burst deletion during the decoding phase. .

[0133] 3.3, outer auxiliary information Convert to a fixed-length quaternion auxiliary sequence ,in The aforementioned Stored according to the preset quaternion representation rules and This enables the decoding end to... The unique parsing yielded and .

[0134] 3.4, regarding Perform a full flip to obtain ,in This refers to the flip function defined in step 2. Subsequently, and Interweave symbol by symbol to obtain interleaved auxiliary tail segments. .

[0135] 3.5, due to each pair of symbols and There just happens to be one that belongs to The other belongs to Therefore, the interlacing auxiliary tail section The length is GC weight is ,in Therefore It inherently satisfies strict GC balancing. Meanwhile, because... Intertwined auxiliary tail section Internally, no consecutive runs of the same sign exceeding 2 will be generated; when hour, satisfy -Run length limit.

[0136] Step 4: Introduce the linker symbol and buffer to construct a structure that simultaneously satisfies 0.05-GC global balance. Run length is limited, and the final codeword can correct a sudden insertion error or sudden deletion error of length not exceeding 10, specifically including: 4.1, Let the upper bound of the burst length be... Select the marker pair and make the buffer zone In this embodiment, it is possible to take... ,but Since each symbol pair has exactly one symbol belonging to 03. Another symbol belongs to Therefore, between buffers It inherently satisfies strict GC balancing; furthermore, 0 and 3 alternate, therefore... No more than consecutive runs with the same symbol.

[0137] 4.2 Constructing Connector Pairs and .set up .when Sometimes, Based on the quaternary pre-encoded codeword The last sign Sure and The order: if There is one and only one symbol that is not equal to Then the symbol is denoted as Another symbol is denoted as ;like Neither of the two symbols is equal to Then, one of them will be selected according to the preset order. Another as This guarantees ,and .

[0138] 4.3, convert the quadratic pre-encoded codewords , Connecting symbols , Connecting symbols Buffer Zone and interwoven auxiliary tail section Combine them in order to obtain the final codeword. .because The length is 204. The length is 10, and the connecting symbol is... and The total length is 2, with interlaced auxiliary tail section The length is ,in Therefore, the final codeword length is .

[0139] 4.4, due to Satisfying 0.05-GC global balancing, linking symbol pairs There is exactly one symbol in the middle that belongs to , buffer Satisfying strict GC balance, interleaved auxiliary tail segment It also satisfies strict GC balance, therefore the final code... The GC bias only comes from .make Then the final code satisfy -GC global balancing condition.

[0140] 4.5, due to satisfy Travel length limit, and , , , buffer It is an alternating symbol sequence with interleaved auxiliary tail segments. Each pair of symbols in the symbol set consists of two different symbols that are flipped together. and Neither the internal structure nor the connection points will produce lengths exceeding [a certain value]. consecutive runs with the same sign. Specifically, and Even if the same sign appears at the connection point, it will at most form a consecutive run of the same sign with a length of 2, which still satisfies the condition. Travel length limit.

[0141] 4.6, Define the final global code set For all the final codewords obtained through the above method The set, i.e. ,in Therefore, the final word count... .

[0142] Step 5, design for the final codeword The decoding methods specifically include: 5.1, Let the received sequence be... The decoder first calculates the length difference. ,in For the final typing The length. Under the single burst insertion / deletion error model considered in this embodiment, if If so, it is determined that no sudden insertion / deletion error has occurred; if If so, a sudden insertion error is determined to have occurred; if If so, a sudden deletion error is determined to have occurred. Let , but .

[0143] 5.2, the length of the end of the received sequence detected at the decoding end is... tail window Does it satisfy the complete reversal relationship of the interlacing auxiliary tail segment? satisfy , This is called the tail window. Alignment. When aligning the tail window, the quaternion auxiliary information sequence is recovered from the odd-numbered position symbols. , and by Parsing yields outer auxiliary information .

[0144] 5.3, when At that time, directly from the received sequence Extracting the first 204 symbols into a quadruple pre-coded codeword ,Right now Then proceed to step 5.7 to perform the precoding inverse operation.

[0145] 5.4, ​​when And the rear window During alignment, burst insertion decoding is performed. Let... For each ,from Delete location to The continuous substrings are used to obtain a candidate sequence of length 204. Simultaneously, the first 204 symbols of the received sequence will be... The error occurred in the quadruple precoding codeword The subsequent candidate sequences. External auxiliary information verification is performed on all candidate sequences; if a certain candidate sequence... satisfy Then the candidate sequence is used as the recovered quadratic precoding codeword. Based on the uniqueness of the outer auxiliary information, the candidate sequence that satisfies the verification condition is unique.

[0146] 5.5, when And the rear window During alignment, burst deletion decoding is performed. First, the first 204 symbols of the received sequence are processed. The error occurred in the quadruple precoding codeword The candidate sequence is then used, and its performance is checked. If satisfied, output... If not satisfied, then let , For each And each ,Will Insert into The After several symbols, a candidate sequence of length 204 is obtained. If a candidate sequence satisfies Then the candidate sequence is used as the recovered quadratic precoding codeword. Based on the uniqueness of the outer auxiliary information, the candidate sequence that satisfies the verification condition is unique.

[0147] 5.6, when the tail window When misaligned, the decoder detects the gap between buffers. Whether it remains consistent within the preset position range. If the buffers remain aligned and the first 204 symbols of the received sequence satisfy... If GC global balancing and run-length limits are met, then it is determined that sudden insertion / deletion errors have not affected the quaternary precoding codewords. Output Otherwise, the received sequence is determined not to fall under the correctable single burst insertion / deletion error scenario as defined in this embodiment.

[0148] 5.7 After completing the above steps, the decoder obtains the recovered quaternion precoded codeword. Subsequently, regarding Perform the reverse operation of step 2. Specifically, first delete the last two balancing supplementary symbols. and Then, based on the preset index segment length Extracting interleaved index segments .Depend on Odd-position sign recovery index quadruple sequence ,Depend on Even position sign recovery And it is verified by checking whether the sign at the even position is equal to the flip value of the sign at the corresponding odd position. Then by... Restore Balanced Index .

[0149] 5.8, Based on the balanced index Delete the corrupted symbol inserted in step 2, and perform a flip function on the corresponding prefix. The inverse operation. Because For any Established, therefore applied again. This is the reverse flip operation. If and Then delete the first destructive symbol between the reversed prefix and the unreversed suffix. and for the former The symbol is executed again. ;like Then delete the first corrupted symbol inserted at the front. If prefix reversal is not performed; Then delete the location located at With interleaved index segments The first destructive symbol between Second destruction symbol And execute again on all the first 196 symbols. This restores the run-length limit intermediate codeword output in step 1. .

[0150] 5.9, regarding the recovered intermediate codewords Perform the reverse operation of step 1. Specifically, for Perform a difference transformation to obtain the replaced extended difference sequence. ; then from The position identifiers are read sequentially from the end. And restore the deleted contiguous blocks according to the positions indicated by the position identifiers. Substring. Repeat this process until the terminal auxiliary symbol is 0; after deleting the auxiliary symbol, perform an inverse difference transform on the resulting difference sequence to recover the original quaternary information sequence. .

[0151] 5.10, the recovered original quaternion information sequence Through mapping function It is converted into a DNA nucleotide sequence, and the decoding is completed.

[0152] The coding system provided in this invention, which features GC global balance and run-length constraints and can correct burst insertion and deletion errors, specifically includes: The run-length-limited precoding module is used to process the original quaternion information sequence. Perform differential transformation, consecutive zero substring replacement, and position identifier appending operations, and obtain the desired result through inverse differential transformation. intermediate codewords with run length limit .

[0153] The GC globally balanced precoding module is used in the intermediate codeword Based on this, prefix reversal, balanced index encoding, complete index sequence reversal and interleaving, disruptive symbol insertion, and balanced supplementary symbol addition operations are performed to obtain a result that simultaneously satisfies 0.05-GC global balance and Quadruple precoding codewords with run length constraints .

[0154] The outer auxiliary information generation module is used to generate the four-element precoded codeword. As input, calculate outer auxiliary information used to locate and recover sudden insertion / deletion errors. and will Convert to quaternary auxiliary sequence .

[0155] The interleaving-assisted tail generation module is used to generate the quaternary auxiliary sequence. Perform a complete flip and The completely reversed sequence is interleaved sign-by-sign to obtain a sequence that satisfies strict GC balance and satisfies Interlacing auxiliary tail section with travel length limitation .

[0156] The final codeword construction module is used to construct the buffer space. and connection symbol pair , and the quadratic precoded codeword , Connecting symbols , Connecting symbols Buffer Zone and interwoven auxiliary tail section Combine them in order to obtain the final codeword. ,in and Simultaneously satisfy -GC global balancing and It has a run length limit and can correct a sudden insertion error or sudden deletion error with a length not exceeding 10.

[0157] The decoding module is used to decode the received sequence. Length variation, buffer space Position and interlacing auxiliary tail section Outer auxiliary information recovered from Recover quaternary precoded codewords Then, the inverse operations of GC global balancing precoding and run-length limiting precoding are performed sequentially to recover the original quaternion information sequence. .

[0158] I. Specific application areas or related products of this invention.

[0159] This invention can be applied to the encoding and decoding ends of DNA data storage systems, and is particularly suitable for long-term archive data storage, biomedical data storage, high-density cold data storage, and molecular storage scenarios requiring resistance to sudden insertion and deletion errors. The encoding end can convert the digital information to be stored into a DNA sequence that meets GC global balance and run-length constraints; the decoding end can recover the original quaternary information sequence based on the received sequence obtained from sequencing, thereby improving the data recovery reliability of the DNA storage system under insertion and deletion synchronization errors.

[0160] II. Evidence related to the technical effects obtained by the embodiments of the present invention.

[0161] This invention proposes a quaternion encoding method for DNA storage. This method uses a raw quaternion information sequence of length 195. For input, construct intermediate codewords that satisfy the run-length limit in sequence. Quadruple precoding codewords that simultaneously satisfy GC global balance and run-length constraints and further As input to the outer burst insertion / deletion error correction structure, the final codeword is constructed. Therefore, the embodiments of the present invention enable the final codeword to simultaneously satisfy 0.05-GC global balance, It has a run length limit and can correct a sudden insertion error or sudden deletion error with a length not exceeding 10.

[0162] Specifically, the technical effects of the embodiments of the present invention can be supported by the following steps and structural relationships.

[0163] Step 1: Construct a structure that satisfies intermediate codewords with run length limit .

[0164] In this embodiment, a quadratic alphabet is first defined. And establish a quaternion alphabet to DNA nucleotide set. mapping ,in Therefore, symbols 2 and 3 correspond to GC symbols, and symbols 0 and 1 correspond to non-GC symbols.

[0165] For the original four-element information sequence First, perform a difference transformation to obtain the difference sequence. And then Adding an auxiliary symbol 0 at the end yields the extended difference sequence. .like If there are no four consecutive zeros, the intermediate codeword can be obtained directly through inverse differential transform. ;like If there are four consecutive zeros, then delete the first occurrence. Substring and add a position identifier at the end Replace in this way, repeating the process until... No longer contains Since there are no four consecutive zeros in the difference sequence, the corresponding inverse difference output... There will be no consecutive runs of the same sign with a length exceeding 4, therefore .

[0166] The technical effect of this step is that, through differential transformation and replacement of consecutive zero substrings, the original quaternary information sequence is transformed into one that satisfies... The run-length-limited intermediate codewords provide input that satisfies the run-length limit for subsequent GC-balanced precoding.

[0167] Step 2: Construct a system that simultaneously satisfies 0.05-GC global balance and Quadruple precoding codewords with run length constraints .

[0168] In this embodiment, a GC symbol set is defined. Non-GC symbol set and define the flip function. ,in This flip function is used to switch between GC symbols and non-GC symbols.

[0169] For the result obtained in step 1 In a balanced set of indexes Middle search balanced index The sequence after prefix reversal The 0.05-GC global balancing condition is met. Then, the balancing index is... Encoded as a length of Quad index sequence The index sequence is then interleaved symbol by symbol with its completely reversed sequence to obtain the interleaved index segment. Since each pair of symbols in the interleaved index segment consists of one GC symbol and one non-GC symbol, the index segment itself satisfies strict GC balance, and... This will not violate the run length limit.

[0170] In this embodiment, destructive symbols are set at the prefix inversion boundary and the index segment access location, respectively. and Used to avoid connection locations exceeding Consecutive runs with the same sign; a balance supplement sign is added at the end. and This is used to compensate for GC class bias introduced by corrupted symbols. After the above processing, the output quadratic precoded codeword is generated. , and have .

[0171] The technical advantage of this step is that, without violating the run-length limit, the intermediate codewords are... Convert to simultaneously satisfy 0.05-GC global balancing and Quadruple precoding codewords with run length constraints .Should It serves as the input to the subsequent outer burst insertion / deletion error correction structure, rather than directly connecting the original information sequence to the outer error correction structure.

[0172] Step 3: Use quaternion pre-encoded codewords Generate outer auxiliary information for input And construct an interlaced auxiliary tail section .

[0173] In this embodiment, outer auxiliary information Using quadruple pre-encoded codewords The information is calculated based on the input. This auxiliary information can be represented as verification information for recovering from sudden deletion errors, and further converted into a fixed-length quaternion auxiliary sequence. ,in .

[0174] Subsequently, Perform a full flip to obtain and will and Symbol-by-symbol interweaving yields the interweaving auxiliary tail segment. Because each pair and There is exactly one symbol in the middle that belongs to Another symbol belongs to ,therefore The GC weight is exactly half its length, therefore It satisfies strict GC balancing. Meanwhile, because... , Internally, no consecutive runs of the same sign exceeding 2 will be formed, therefore... The travel length limit must be met.

[0175] The technical advantage of this step is that the outer auxiliary information is not directly appended as a regular redundant string, but rather formed into an interleaved auxiliary tail segment that satisfies GC balance and run-length limitations after complete flipping and symbol-by-symbol interleaving. This avoids redundant information in the outer layer from disrupting the biochemical constraints of the final codeword.

[0176] Step 4: Construct the final codeword by connecting the symbols and the buffer. .

[0177] In this embodiment, the upper bound of the burst length is set as follows: Select buffer space ,in The length is 10. This buffer consists of alternating 0s and 3s, where 0 belongs to the... ,3 belongs to ,therefore It inherently satisfies strict GC balance and will not produce long contiguous runs of the same sign.

[0178] To avoid quaternary pre-encoding codewords Between the buffer zone In this embodiment, an excessively long travel distance is formed at the connection point. Selecting connection symbols and and according to The last sign Sure and The order makes and The final codewords are constructed in sequence as follows: .in, The length is 204. and The total length is 2. The length is 10. The length is Therefore, the final codeword length is recorded as... ,but .

[0179] Since x satisfies the 0.05-GC global balance, There is exactly one symbol in the middle that belongs to , buffer and interwoven auxiliary tail section All satisfy strict GC balance, therefore the final codewords The GC bias only comes from .make Then the final code satisfy -GC global balancing conditions. Also, due to... satisfy Travel length limit, , , Not equal to buffer space The first symbol, and and Internally, no more than four consecutive runs of the same sign will be formed, therefore the final codeword Also satisfies Travel length limit.

[0180] The technical effect of this step is that, by combining the connection symbols, buffers, and interleaving auxiliary tail segments, the final codeword can still maintain GC global balance and run length limit after receiving the outer layer error correction auxiliary information, and provides a structural basis for the decoder to judge the region affected by sudden insertion and deletion errors.

[0181] Step 5: Design for the final code The decoding method recovers the original quaternary information sequence. .

[0182] In this embodiment, the decoding end first determines the sequence based on the received sequence. Length and final codeword length The error type is determined by the difference in length. If the length increases, a sudden insertion error is determined to have occurred; if the length decreases, a sudden deletion error is determined to have occurred; if the length remains unchanged, no sudden insertion or deletion error is determined to have occurred under the sudden insertion / deletion error model considered in this embodiment.

[0183] Subsequently, the decoding end checks the marker window in the received sequence. Is it equal to the buffer space? Determine whether the error affects the quaternary precoding codeword The area in question; and also by inspecting the tail window. To determine if the complete flipping interlacing relationship between adjacent odd and even positions is satisfied, and whether the interlacing auxiliary tail segments remain aligned. When the region is unaffected, the decoder can directly extract the first 204 symbols of the received sequence. ;when When the area is affected, the decoder uses the interleaving auxiliary tail segment to recover the outer auxiliary information. And unique recovery through candidate sequence verification. .

[0184] In recovering the quaternary precoded codeword Then, the decoder first performs the reverse operation of step 2: deleting the balanced supplementary symbol. and Extract interleaved index segments Restore balanced index and according to After deleting the corresponding destructive symbols, perform a reverse flip to restore the intermediate codewords. Subsequently, the decoding end performs the reverse operation of step 1: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Perform differential transformation and read the end position identifier. And gradually restore the replaced The substring is processed until the terminal auxiliary symbol is 0; after deleting the auxiliary symbol, an inverse difference transform is performed to recover the original quaternary information sequence. .

[0185] The technical advantage of this step is that the decoding process and the encoding process form a clear inverse correspondence, enabling the recovery of the quaternion pre-encoded codewords first. Then restore the intermediate code. Finally, the original quaternary information sequence is recovered. This ensures a closed-loop encoding and decoding process.

[0186] In summary, this embodiment constructs a method that simultaneously satisfies run-length limitation precoding, GC global balanced precoding, outer auxiliary information interleaving processing, buffer settings, and linker symbol selection. -GC global balancing and Final codeword for run length limit It can correct burst insertion or deletion errors of no more than 10 units in length. This scheme avoids the problem that directly attaching outer redundancy may disrupt GC balance and run-length constraints, while ensuring that the burst insertion / deletion error correction structure is consistent with the biochemical constraints in DNA storage.

[0187] Technological advancements: This invention establishes a constraint-preserving error correction structure for burst insertion / deletion errors in DNA storage. Instead of directly adding error-correcting redundancy to the original quaternary information sequence, this structure first obtains a quaternary pre-encoded codeword that simultaneously satisfies global GC balance and run-length constraints, and then uses this pre-encoded codeword as input to the outer error correction structure. In this way, the access of outer auxiliary information does not bypass the pre-constraint encoding process, which helps maintain the consistency between the encoding and decoding processes.

[0188] This invention, through the collaborative design of buffer spaces, connection symbols, and interleaving auxiliary tail segments, enables the final codeword to meet the biochemical constraints of DNA storage while possessing burst insertion / deletion error correction capabilities. Specifically, the buffer spaces can employ an alternating symbol structure, for example... or Used to assist the decoder in determining the affected area of ​​sudden insertion / deletion errors; connection symbols and The remaining quaternions are selected from those not belonging to the inter-buffered symbol pairs, and their order is determined according to the last symbol of the quaternary precoded codeword to avoid consecutive runs of the same symbol exceeding the run length limit at the connection point; the interleaving auxiliary tail segment is obtained symbol-by-symbol interleaving of the outer auxiliary sequence and its fully reversed sequence, thus satisfying strict GC balance, and in The travel length limit must be met.

[0189] The embodiments of the present invention can correct a length not exceeding The invention addresses sudden insertion or deletion errors. For sudden insertion errors, the decoder can uniquely recover the original quaternary precoded codeword from the candidate sequence based on changes in the received sequence length, buffer positions, and outer auxiliary information recovered from the interleaving auxiliary tail segment. For sudden deletion errors, the decoder can determine whether the quaternary precoded codeword is affected based on the alignment relationship between buffers and the interleaving auxiliary tail segment, and if affected, recover the deleted substring using the outer auxiliary information. Thus, this embodiment of the invention achieves reliable recovery for a single sudden insertion or deletion error while maintaining global GC balance and run-length limitations.

[0190] This invention also constrains the outer auxiliary information. The outer auxiliary information is first represented as a quaternion auxiliary sequence, and then interleaved symbol by symbol with its completely reversed sequence to form an auxiliary tail segment. This processing method ensures that the auxiliary tail segment itself has strict GC balance properties and avoids the formation of long consecutive runs of the same sign, thereby solving the problem that directly attaching error correction redundancy may disrupt GC balance or run length limitations.

[0191] The encoding and decoding processes in this invention have a clear inverse correspondence. The encoding end sequentially executes run-length-constrained precoding, GC globally balanced precoding, outer auxiliary information generation, and final codeword construction. The decoding end first recovers the quaternary precoded codeword based on the final codeword structure, then performs the inverse operation of GC globally balanced precoding and the inverse operation of run-length-constrained precoding, ultimately recovering the original quaternary information sequence. This process structure is clear and easy to implement in DNA storage systems.

[0192] Upper bound of fixed burst length In such cases, the outer-layer decoding process of this invention can be completed through finite candidate sequence enumeration and auxiliary information verification, achieving achievable computational complexity. Compared with encoding methods that only focus on error correction capabilities without simultaneously maintaining GC global balance and run-length constraints, this invention is more suitable for DNA storage applications that simultaneously require biochemical constraint satisfaction and burst insertion / deletion error correction capabilities.

[0193] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0194] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A coding method for achieving global GC balance and run-length constraint while correcting burst insertion / deletion errors, characterized in that: include: Step 1: Receive the quaternion information sequence to be stored. Its symbol comes from the quadratic alphabet. ; Step 2, for the four-element information sequence Perform a difference transform to obtain the difference sequence. ; In the difference sequence Adding an auxiliary symbol 0 to the end of the sequence yields the extended difference sequence. Detect and delete the first occurrence of a sequence with a length of [length missing]. consecutive zero-signed substrings Then, add a position identifier indicating the start position of the substring to the end of the sequence; repeat the above detection, deletion, and addition steps until the resulting extended difference sequence does not contain any substrings. Subsequently, an inverse difference transform is performed on the resulting extended difference sequence to obtain the sequence that satisfies... - Middle codewords for run length limitations ; Step 3, define a flip function for converting symbols between the non-GC symbol set and the GC symbol set; for the intermediate codeword Search for a balanced index i in the index set determined by zero, the intermediate codeword length, and a preset step size, such that for the intermediate codeword... The sequence obtained after applying the flip function to the first i symbols satisfies the preset GC global balance deviation condition; the balance index i is represented as a quadruple index sequence, and the flip function is applied to the quadruple index sequence symbol by symbol to obtain a flipped index sequence; the quadruple index sequence and the flipped index sequence are interleaved bit by bit to generate an interleaved index segment, which itself satisfies strict GC balance and the run length does not exceed the limit. ; Step 4: Based on the value of the balanced index i, insert a disruptive symbol during the combination of the intermediate codeword after prefix flipping and the interleaved index segment to prevent lengths exceeding the specified values ​​from occurring at the flip boundary, end access position, or index segment access position. The same-sign consecutive runs; and two balance supplementary signs are appended to the end of the sequence to form a sequence that simultaneously satisfies the GC global balance deviation condition and - Composite constraint codewords with run-length limits ; Step 5, based on the composite constraint codeword Calculate the outer auxiliary information used for decoding and verification. The outer auxiliary information Encode the quaternary auxiliary sequence; perform the flip function on the quaternary auxiliary sequence to obtain a flipped auxiliary sequence, and interleave the quaternary auxiliary sequence and the flipped auxiliary sequence bit by bit to generate an interleaved auxiliary tail segment. The interlacing auxiliary tail section It satisfies strict GC balance and the run length does not exceed ; Step 6, Select ordered symbol pairs ,in ,and and These belong to different sets within the GC symbol set and the non-GC symbol set, respectively; the ordered symbol pairs are repeated continuously and alternately. t / 2 Next, forming a buffer zone The buffer zone This is used to assist in determining the affected region of sudden insertion and deletion errors, and to ensure that the buffers themselves satisfy strict GC balance. - Travel length limit; Step 7, select a pair of connector symbols, the pair of connector symbols being derived from the quadratic alphabet. Remove the ordered symbol pairs from The last two symbols constitute the composite constraint codeword. The last symbol determines the order of the concatenation symbol pair, making the first concatenation symbol different from the composite constraint codeword. The last sign of the second concatenation symbol is different from the first concatenation symbol, and the second concatenation symbol is different from the buffer space. The first symbol; Step 8, convert the composite constraint codeword The connection symbol pair, the buffer space and the interlaced auxiliary tail section By sequentially piecing them together, the final code can be obtained. The final codeword Simultaneously satisfy -GC global balancing and - Run length limit, and able to correct for a length not exceeding Sudden insertion error or sudden deletion error; Step 9: Map the final codeword according to the preset base mapping rules. Each quaternion in the sequence is mapped to a DNA base to obtain a DNA base sequence; in one system implementation, the DNA base sequence is provided to a DNA synthesizer to perform synthesis.

2. The encoding method for GC global balancing and run-length constraint that can correct burst insertion and deletion errors according to claim 1, characterized in that, Representing the balanced index as a quad index sequence and generating interleaved index segments specifically includes: The index representation length k is determined based on the allowed GC global balance deviation; Map the balanced index to a quadruple sequence of length k; Perform the flipping operation symbol by symbol on the quaternion sequence to obtain the flipped sequence; The quaternion sequence and the reverse sequence are arranged alternately by sign, such that the sign of the quaternion sequence is in an odd position and the sign of the reverse sequence is in an even position, forming the interleaved index segment.

3. The encoding method for GC global balancing and run-length constraint that can correct burst insertion and deletion errors according to claim 1, characterized in that, The outer auxiliary information includes a positive integer γ and a remainder ρ, where ρ is the remainder obtained by modulo γ after inputting the composite constraint codeword into a preset integer value function; the positive integer γ and the preset integer value function are set such that for different candidate sequences corresponding to the burst deletion output of the same length not exceeding t, their remainders modulo γ are different.

4. The encoding method for GC global balancing and run-length constraint, and capable of correcting burst insertion and deletion errors, as described in claim 1, is characterized in that... The insertion of the disruption symbol and the addition of the balance supplement symbol include: When the balanced index is greater than zero and less than the length of the intermediate codeword, a first disruptive symbol is inserted between the prefix flipped portion and the unflipped suffix. This first disruptive symbol is different from the last symbol of the prefix and the first symbol of the suffix. A second disruptive symbol is inserted between the resulting sequence and the interleaving index segment. This second disruptive symbol is different from its immediate preceding symbol and the first symbol of the interleaving index segment. A first balanced supplementary symbol and a second balanced supplementary symbol are appended to the end of the sequence. The GC assignment of the first balanced supplementary symbol is opposite to that of the first disruptive symbol and different from that of its immediate preceding symbol. The GC assignment of the second balanced supplementary symbol is opposite to that of the second disruptive symbol and different from that of the first balanced supplementary symbol. When the balance index is zero, a first destruction symbol is inserted before the intermediate codeword. This first destruction symbol is different from the first and last symbols of the intermediate codeword. The subsequent operations are the same as when the balance index is greater than zero. When the balance index is equal to the length of the intermediate codeword, a first destruction symbol and a second destruction symbol are inserted sequentially between the fully flipped codeword and the interleaving index segment. The first destruction symbol is different from the last symbol of the flipped codeword, and the second destruction symbol is different from the first destruction symbol and the first and second symbols of the interleaving index segment. The subsequent operations are the same as when the balance index is greater than zero.

5. A decoding method for recovering the original quaternary information sequence from a DNA sequencing signal, wherein the DNA sequencing signal originates from a received sequence obtained by DNA synthesis, storage, and sequencing of the final codeword sequence generated according to any one of claims 1 to 4, characterized in that, include: The received sequence is obtained, and the difference between its length and the preset length of the final codeword sequence is calculated. A positive difference is used to determine burst insertion, a negative difference is used to determine burst deletion, and a zero difference is used to determine no burst error. The positions between the buffers are identified in the received sequence, the alignment status between the buffers is determined by the alternating repetition pattern of the marker symbol pairs, and the alignment of the interleaving auxiliary tail window at the end of the sequence is determined by detecting whether the odd-numbered symbols are equal to the adjacent even-numbered symbols after the flipping function. When the interleaving auxiliary tail segments are aligned, the odd-numbered symbols of the tail segment window are extracted to form a quaternion auxiliary sequence, and the outer auxiliary information is decoded and recovered. If the error type is burst insertion and the interleaving auxiliary tail segment is aligned, the insertion length t* is determined by the length difference. A portion of the received sequence from the beginning to the (n+t*)th symbol is extracted, where n is the preset length of the composite constraint codeword. A candidate sequence is generated by deleting all possible consecutive t* symbols from this portion, and the first n symbols of the received sequence are used as supplementary candidates. The candidate sequence is verified using the outer auxiliary information, and the only candidate that satisfies the verification is taken as the composite constraint codeword. If the error type is burst deletion and the interleaving auxiliary tail segment is aligned, the deletion length t* is determined by the absolute value of the length difference. The first n symbols of the received sequence are first checked as candidates. If they fail, the first nt* symbols of the received sequence are truncated. Candidate sequences are generated by inserting all possible quadruple substrings of length t* at all possible positions. The outer auxiliary information is used to uniquely determine the composite constraint codeword. If the interleaving auxiliary tail segment is not aligned, and the buffers are aligned within the expected range and the first n symbols of the received sequence satisfy the GC global balance deviation condition and When run length is limited, the first n symbols of the received sequence are directly extracted as the composite constraint codeword; Remove the last two balanced supplementary symbols from the determined composite constraint codeword, extract the odd-numbered bits of the interleaved index segment to recover the quad-index sequence, and parse to obtain the balanced index; identify and delete the inserted disruptive symbols according to the balanced index, perform the inverse operation of the flip function on the corresponding prefix, and recover the intermediate codeword; Perform a differential transform on the intermediate codeword, and sequentially read position identifiers from the end of the resulting sequence. The insertion length is determined based on the position identifiers. The continuous zero-symbol substring is processed until a predetermined tail identifier is detected, and then the inverse differential transformation is performed to obtain the original quaternion information sequence.

6. The decoding method for recovering the original quaternary information sequence from a DNA sequencing signal according to claim 5, characterized in that, The candidate sequence generation step under the sudden insertion scenario further includes: Let the insertion length t* be equal to the length difference, and take the first n+t* symbols of the received sequence to form the source sequence; For i from 0 to n, candidate sequences are obtained by deleting t* consecutive symbols starting from position i+1 in the source sequence, and the original first n symbols of the received sequence are taken as one of the candidates.

7. The decoding method for recovering the original quaternary information sequence from a DNA sequencing signal according to claim 5, characterized in that, The candidate sequence generation step under the sudden deletion scenario further includes: Let the deletion length t* be equal to the absolute value of the length difference. If the first n symbols of the received sequence fail the verification, then take the first nt* symbols of the received sequence as the base sequence. For i from 0 to nt*, in all sequences of length t* of the four-letter alphabet, each sequence is inserted after the i-th symbol of the base sequence to form a candidate sequence, and the outer auxiliary information is used for verification.

8. A DNA data storage system, characterized in that, include: An encoder configured to receive a quaternary information sequence to be stored, execute the encoding method as described in any one of claims 1 to 4, and output a DNA base sequence; A DNA synthesizer, connected to the encoder, is used to receive the DNA base sequence and synthesize the corresponding DNA molecule; A DNA sequencer is used to sequence the stored DNA molecules and output the received sequence. A decoder, connected to the DNA sequencer, is configured to receive the received sequence and perform the decoding method as described in any one of claims 5 to 7 to recover the original quaternary information sequence.

9. The DNA data storage system according to claim 8, characterized in that, The encoder includes: The run-length precoding unit is used to perform differential transformation, consecutive zero substitution and position identifier appending, and inverse differential transformation to output the intermediate codeword; The GC balancing coding unit is used to search for the balancing index, perform prefix flipping, generate the interleaving index segment, insert the breaking symbol and the balancing supplement symbol, and output the composite constraint codeword. An auxiliary tail segment generation unit is used to calculate the outer auxiliary information and generate the interleaving auxiliary tail segment; The codeword concatenation unit is used to construct the buffer space, select the connection symbol pairs, and concatenate them to obtain the final codeword sequence.

10. The DNA data storage system according to claim 8, characterized in that, The decoder includes: An error detection unit is used to determine the error type based on the difference between the length of the received sequence and the preset length. A window detection unit is used to detect alignment between buffers and alignment of interleaving auxiliary tail segments, and to recover outer auxiliary information. A candidate enumeration verification unit is used to construct a candidate sequence set under sudden errors and use outer auxiliary information to perform verification to determine the composite constraint codeword; The inverse precoding unit is used to recover the balance index from the composite constraint codeword, delete the disruptive symbol and the balance supplement symbol, perform inverse flipping, and recover the original quaternary information sequence through inverse run-length precoding.