Image DNA storage method for improving information density and biological stability

By employing wavelet domain visual optimization compression, dual-mode DNA encoding, and multi-level error control in the image inpainting algorithm, the problems of limited coding density and insufficient biological stability in DNA storage are solved, achieving efficient image data storage and high-quality reconstruction.

CN121837397APending Publication Date: 2026-04-10ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing DNA storage methods suffer from limitations in encoding density, insufficient biological stability, and reduced reconstruction quality under high error rates when used for image data storage.

Method used

An image inpainting algorithm employing wavelet domain visual optimization compression, dual-mode DNA coding, and multi-level error control, combined with discrete wavelet transform, adaptive quantization, dual-mode base mapping, and cross-Transformer network, achieves reduced data redundancy, balanced GC content, and image inpainting under high error rates.

Benefits of technology

It achieves efficient compression and encoding while ensuring image quality, improves information density and biological stability, maintains reconstruction quality under high error rate conditions, and is suitable for DNA storage in fields such as image archiving, cultural heritage protection, and medical imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837397A_ABST
    Figure CN121837397A_ABST
Patent Text Reader

Abstract

The invention provides an image DNA (deoxyribonucleic acid) storage method for improving information density and biological stability, which comprises the following steps of: partitioning an input image, performing two-dimensional discrete wavelet transform, performing quantization processing on a wavelet coefficient by adopting a self-adaptive quantization strategy, and converting the wavelet coefficient into a binary sequence to be coded; the method comprises the following steps: dividing a binary sequence to be coded into a byte according to every 8 bits, and carrying out dual-mode DNA coding by adopting an odd-even alternating mode and a segmentation structure of'feature segment-coding segment 'to generate a DNA information segment sequence; performing error correction on the coded DNA sequence based on preliminary correction of Hamming distance and RS code secondary error correction; marking the DNA sequence which still does not pass the CRC verification after multiple rounds of error correction as an unrepairable sequence, wherein the image block corresponding to the unrepairable sequence is an error block; and in combination with context information, the neighborhood features of the error blocks are learned by using a cross Transform context repair network, and the error blocks are repaired. According to the method, better balance is achieved in the aspects of image reconstruction quality, coding density and biological compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of image DNA storage, and more particularly to an image DNA storage method that improves information density and biological stability. Background Technology

[0002] The exponential growth of global data volume presents new challenges to high-density, low-energy-consumption, and highly reliable data storage media. Traditional magnetic and optical storage technologies are gradually approaching their physical limits in terms of capacity, lifespan, and energy efficiency. Against this backdrop, DNA molecules, with their extremely high theoretical storage density (approximately 455 EB / g), excellent chemical stability, and long-term preservation capabilities, are considered the ideal medium for achieving ultra-high-density data storage.

[0003] The concept of DNA storage can be traced back to the 1960s, proposed by Norbert Wiener and Mikhail Neiman. In 1988, Joe Davis, through the "Microvenus" project, first successfully wrote information, validating the feasibility of DNA as an information carrier. With advancements in DNA synthesis, sequencing, and information encoding technologies, DNA storage research has gradually moved from the proof-of-concept stage to practical application exploration. To ensure the feasibility of biological experiments and sequence stability, DNA encoding typically needs to meet several constraints, such as maintaining a GC content between 40% and 60%, and limiting the consecutive length of the same base to no more than 3–4.

[0004] Early studies often employed direct mapping between binary and base pairs, such as [00, 01, 10, 11] → [A, C, G, T]. While simple to implement, this method is prone to uneven GC content and long homopolymers, increasing sequencing error rates. In 2012, Church et al. proposed a binary base coding method, reducing such errors through binary grouping. In 2015, Grass et al. introduced triplet coding to optimize sequence characteristics. In 2017, Erlich et al. proposed the "fountain code," significantly improving coding robustness and storage efficiency. In 2022, the BGI research team proposed a "yin-yang coding / decoding" strategy, improving synthesis and sequencing compatibility through alternating coding rules. Furthermore, some studies have attempted to introduce degenerate bases to further increase coding density, but this also presents technical challenges to synthesis accuracy and sequencing depth.

[0005] Image storage, as an important research direction in DNA storage, possesses high fault tolerance and compressibility. The human visual system (HVS) is insensitive to small noise, enabling effective reconstruction of image data under high fault tolerance conditions. In recent years, image DNA storage technology has developed rapidly. For example, Li et al. proposed the IMG-DNA scheme in 2021, which suppresses the propagation of insertion and deletion errors through an error barrier mechanism; the HL-DNA scheme was proposed in 2022, which improves sequence stability through rotational coding; and in 2024, Wang et al. proposed the base128 coding scheme, which adaptively generates a coding table through statistical features to control GC content and homopolymer length. Although these methods have made some progress in terms of information density and biocompatibility, problems still exist, such as limited coding efficiency, unstable homopolymer control, and decreased image reconstruction quality under high error rate environments.

[0006] Therefore, DNA-based molecular storage technology is considered an important direction for future massive information storage due to its ultra-high storage density, long-term stability, and low energy consumption. However, existing DNA storage solutions still suffer from limitations in image data storage, including limited coding density, insufficient biological stability, and decreased reconstruction quality under high error rates. Summary of the Invention

[0007] To address the technical problems of limited coding density, insufficient biological stability, and degraded reconstruction quality under high error rates in existing DNA storage methods for image data, this invention proposes an image DNA storage method that improves information density and biological stability. The method employs wavelet domain visual optimization compression at the data processing layer to reduce data redundancy, dual-mode mapping at the coding layer to stabilize GC content, and multi-level error control and deep learning-based image inpainting algorithms at the error correction layer, ensuring high reconstruction quality and reliability even under high error rates.

[0008] To achieve the above objectives, the technical solution of the present invention is as follows: a method for storing image DNA with improved information density and biological stability, comprising the following steps:

[0009] Step 1: Data compression processing, based on discrete wavelet transform and adaptive quantization strategy to reduce data redundancy and convert it into a binary sequence to be encoded.

[0010] Step 2: Based on the GC content constraint and the maximum homopolymer length constraint, construct a dual-mode DNA encoding strategy based on byte index parity, namely a dual-mode base mapping strategy: divide the binary sequence to be encoded into 8 bits per byte, and encode it using an alternating parity mode and a segmented structure of "feature segment - coding segment" to generate a DNA information segment sequence that conforms to biological constraints;

[0011] Step 3: Input preprocessing and sequence structure design for dual-mode DNA coding

[0012] Step 4: Based on the preliminary correction of Hamming distance and the second-order error correction of RS code, the encoded DNA sequence is corrected.

[0013] Step 5: Sequences that still fail the CRC check after multiple rounds of error correction are marked as unrepairable sequences, and the image blocks corresponding to unrepairable sequences are erroneous blocks; combined with context information, the neighborhood features of the damaged image blocks are learned through the Cross Transformer Context Repair Network (CTNet) to repair the image blocks that still have distortion.

[0014] The method to compress data into a binary sequence is to compress the input image using discrete wavelet transform: the input image is divided into blocks of 16×16 pixels, and a two-dimensional discrete wavelet transform (DWT) is performed on each image block to convert the spatial domain pixel information into a wavelet coefficient domain containing low-frequency coefficients (LL) and high-frequency coefficients (LH, HL, HH).

[0015] Based on the perceptual characteristics of the human visual system (HVS), an adaptive quantization strategy is adopted to perform low-distortion uniform quantization on low-frequency coefficients that carry the core structure of the image, and to perform non-uniform quantization on high-frequency coefficients that correspond to edge texture details.

[0016] The quantized coefficients are directly converted into binary sequences, and after adding a header containing image block indexes and size information, they are concatenated into a fixed 28-byte compressed byte stream.

[0017] The binary sequence to be encoded obtained in step one is mapped according to the rule of "8-bit byte → 5 base".

[0018] Step two involves constructing a dual-mode base mapping strategy based on byte index parity to encode and map binary sequences, yielding the corresponding base sequences.

[0019] The binary sequence is grouped into 8-bit units. Odd-indexed bytes use mode A with 40% GC content, and even-indexed bytes use mode B with 60% GC content, ensuring that the global GC content is stable at 50% ± 2%.

[0020] During the mapping process, the high 3 bits of each byte are extracted as feature segments. The GC and AT positions in the 5 base positions are allocated according to preset rules. At the same time, the homopolymer length is controlled to be ≤3nt to avoid the formation of error-prone adjacent base combinations, thus generating a DNA information segment sequence that meets the constraints of biosynthesis and sequencing.

[0021] The dual-mode DNA encoding includes Mode A and Mode B, and the GC base position distribution rules are as follows:

[0022] ;

[0023] The high 3 bits of each byte are extracted as a feature segment, and the allocation scheme of GC bits and AT bits in the 5 base positions is determined according to the GC base position distribution rules of pattern A and pattern B.

[0024] Extract the lower 5 bits of the byte as the encoding segment, and generate specific bases according to the following rules: If the base position is a GC bit: the corresponding bit of the encoding segment is 0 when it is mapped to G, and 1 when it is mapped to C; if the base position is an AT bit: the corresponding bit of the encoding segment is 0 when it is mapped to A, and 1 when it is mapped to T.

[0025] In step three, to meet the constraints and information encoding requirements of synthetic biology, each DNA sequence consists of the following regions: information segment (stores encoded data), primer segment (20 nt on each side, used for PCR amplification), index segment (used for DNA fragment sorting), and check segment (stores CRC code to detect errors).

[0026] The method for correcting DNA errors based on Hamming distance and RS code in step four is as follows:

[0027] The read DNA sequences (length L) are grouped into 5-base groups, and the Hamming distance between each group and all codewords in the encoding table is calculated. A threshold of 0 is set: if the Hamming distance between a group of sequences and a codeword in the encoding table is 0, the group is considered correct; if the minimum Hamming distance is greater than 0, it indicates that there may be an error.

[0028] First, the error type is initially determined by comparing the actual sequencing sequence length with the standard DNA sequence length. Sequences shorter than the standard length are identified as deletion errors; sequences longer than the standard length are identified as insertion errors.

[0029] Secondly, locate the first base group with a non-zero Hamming distance (i.e., the first error base group): For missing errors, try inserting all possible bases (A, C, G, T) at the beginning of the error group, and select the scheme that minimizes the sum of Hamming distances of subsequent groups for correction; for insertion errors, try deleting bases at each position in the group, and select the scheme that minimizes the sum of Hamming distances of subsequent groups; if the sequence lengths are the same but the Hamming distance of a certain group is not 0, and the minimum Hamming distance is ≤1, then replace the group with the codeword with the closest distance in the encoding table;

[0030] The entire sequence is then divided into 5-nt units again, and the Hamming distance between each group and the codeword in the encoding table is recalculated. The above steps are repeated until the Hamming distance of all segments is 0.

[0031] If the Hamming distance of some codewords is less than 2, insertion / deletion errors may still cause the sequence to partially match the encoding table. Therefore, RS codes are needed for global error correction to address cross-byte errors and errors that have not been corrected by local error correction.

[0032] The implementation method of RS code level 2 error correction is as follows: the DNA sequence is grouped into units of 5 bases and mapped to information symbols on a finite field; a generator polynomial-based encoder is used to generate check symbols and construct system codewords; at the decoding end, the synodus of the received codeword is calculated to detect errors; if an error is detected, the error position is located by solving the error position polynomial and the error value is calculated, thereby completing the error correction; for error blocks that still cannot be repaired, the error blocks are repaired by combining context information and using a cross-Transformer context repair network to learn the neighborhood features of the error blocks.

[0033] The image restoration mechanism in step five is as follows:

[0034] During decoding, a dual detection is first performed by combining a local error correction mechanism and a CRC check: if the sequence passes the CRC check, or is restored to consistency after local error correction, it is determined to be a valid sequence and enters the image reconstruction stage;

[0035] If a sequence fails the CRC check after multiple rounds of correction, it is marked as an unrepairable sequence, and its corresponding image block will be further processed at the subsequent image level.

[0036] The image regions corresponding to irreparable sequences are defined as error blocks, and the set is represented as follows: .in This represents the k-th image patch. This serves as a marker of validity.

[0037] To address these erroneous blocks, the CTNet network is introduced. Through a combination of serial blocks, parallel blocks, and residual blocks, it learns the contextual dependencies between erroneous blocks and their neighborhoods, achieving multi-scale feature aggregation and texture detail restoration. The specific computation process is as follows: .in, Indicates an error block Spatial neighborhood characteristics, This is the result of the reconstruction after repair.

[0038] Compared with existing technologies, the beneficial effects of this invention are as follows: It proposes a highly efficient DNA encoding scheme for image data—DNA-GCoder. In the preprocessing stage, a compression strategy combining block discrete wavelet transform (DWT) with visual perception optimization is adopted to achieve a compression ratio of approximately 2.2:1 while ensuring image quality (PSNR>38dB), thereby effectively reducing the cost of DNA synthesis and sequencing. In the encoding stage, a dual-mode collaborative encoding strategy is designed. Through the mapping from 8 to 5 bases and the alternating allocation mechanism of GC content from 40% to 60%, the GC content of each 10nt fragment is maintained at 50%, and the global GC content is maintained at 50%±2%, balancing information density and biological stability. Simultaneously, a unified sequence structure including information segments, index segments, check segments, and primer segments is constructed to achieve ordered data organization and reliable error correction. To mitigate the impact of DNA degradation and sequencing errors on image reconstruction, this invention introduces an image denoising method using a cross-transformer network, leveraging contextual features to enhance the repair capability of damaged image patches. Experimental results show that the DNA-GCoder of this invention achieves a coding density of 3.01 bits / nt, and even at a 5% error rate, the peak signal-to-noise ratio of the reconstructed image remains at 29 dB. Compared with existing methods, this invention achieves a better balance in image reconstruction quality, coding density, and biocompatibility, providing a feasible approach for the application of DNA storage in image archiving, cultural heritage preservation, and medical imaging.

[0039] The preprocessing stage of this invention employs a block-based discrete wavelet transform combined with a visual perception-optimized quantization strategy to reduce data redundancy; the encoding stage uses dual-mode DNA encoding to map the compressed byte stream into a DNA sequence that conforms to biological constraints, achieving local independent mapping and supporting error correction design; error correction analysis completes local error detection and preliminary correction through a sliding window mechanism under encoding constraints, and combines Reed-Solomon (RS) codes to achieve cross-byte-level global error correction; at the same time, a context-based inpainting network based on cross-Transformer further restores damaged image patches, improving the overall reconstruction quality and system robustness. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a flowchart of the present invention.

[0042] Figure 2This is a schematic diagram of the DNA sequence design for this invention.

[0043] Figure 3 The diagram shows the DNA sequence characteristics of this invention, where (a) is a comparison of local GC content distribution and (b) is a comparison of homopolymer length distribution.

[0044] Figure 4 The robustness analysis diagram of the encoding of this invention is shown in the figure, where (a) is the binary data recovery rate under replacement error, (b) is the binary data recovery rate under insertion and deletion errors, and (c) is a comparison diagram of the error correction performance of the dual-mode encoding scheme at an error rate of 1-10%.

[0045] Figure 5 This is a comparison chart of image restoration quality at a 0.5% error rate according to the present invention.

[0046] Figure 6 The image restoration metrics of the different schemes under different error rates of the present invention are compared, where (a) is SSIM and (b) is PSNR.

[0047] Figure 7 This is a schematic diagram of image reconstruction analysis according to the present invention, wherein (a) is a reconstructed image under different error rates, and (b) is a comparison diagram of effectiveness on partially damaged image blocks. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] like Figure 1As shown, an image DNA storage method that improves information density and biological stability consists of three core steps: data compression, dual-mode DNA encoding, and error correction analysis under encoding constraints. First, the data compression module divides the input image into blocks and performs a two-dimensional discrete wavelet transform (DWT) on each block, converting spatial domain pixel information into wavelet coefficient domain. After generating a wavelet coefficient sequence, a header (containing image block index and size information) is added. A visual perception-optimized quantization strategy is used to reduce data redundancy, providing efficient input for subsequent DNA encoding. Subsequently, dual-mode DNA encoding maps the compressed byte stream to a biologically constrained DNA sequence, achieving local independent mapping and supporting error correction design. Finally, error correction analysis, under encoding constraints, completes local error detection and preliminary correction through a sliding window mechanism, and combines Reed-Solomon (RS) codes to achieve cross-byte-level global error correction. Simultaneously, a context-based inpainting network based on cross-Transformers further restores damaged image blocks, improving overall reconstruction quality and system robustness. Specifically, this invention includes the following steps:

[0050] Step 1: Data Compression: Divide the input image into blocks and perform a two-dimensional discrete wavelet transform. Use an adaptive quantization strategy to process the wavelet coefficients and convert the quantized wavelet coefficients into a binary sequence to be encoded.

[0051] To improve the storage efficiency of image data in DNA storage while ensuring information integrity, efficient compression of the input image is necessary. Specifically, the input image is first divided into 16×16 pixel sub-blocks; then, a two-dimensional discrete wavelet transform (DWT) is performed on each sub-block, converting the spatial domain pixels into a wavelet coefficient domain representation. Wavelet coefficients can be divided into low-frequency coefficients (reflecting the main structural features of the image and sensitive to visual quality) and high-frequency coefficients (describing edge and texture information and having lower visual sensitivity). Based on the perceptual differences of the human visual system (HVS) for different frequency information, this invention employs adaptive quantization strategies for low-frequency and high-frequency coefficients respectively, as specifically implemented below:

[0052] Low-frequency coefficient quantization: Low-frequency coefficients reflect the core structure of the image and are sensitive to visual quality. Low-distortion uniform quantization is adopted. By minimizing the trade-off function between quantization distortion and visual perception loss, it is ensured that the quantized low-frequency coefficients can accurately preserve the main outline and key details of the image.

[0053] High-frequency coefficient quantization: High-frequency coefficients correspond to detailed information such as image edges and textures. HVS has a high tolerance for distortion in these coefficients and employs non-uniform quantization. A quantization step size is designed based on the visual threshold matrix, and high-frequency coefficients are graded according to their energy: high-energy high-frequency coefficients (contributing significantly to texture) use a smaller quantization step size; low-energy high-frequency coefficients (low visual sensitivity) use a larger quantization step size. Simultaneously, a quantization threshold screening is introduced, directly setting high-frequency coefficients with amplitudes below the HVS perception threshold to zero, further reducing data redundancy. Adaptive adjustment of quantization parameters: The coefficients of each sub-band after DWT decomposition are divided into blocks; the just-perceptible distortion threshold JND for each block is calculated; for low-frequency blocks: quantization is performed according to the process of "background brightness → step size selection → offset correction" to ensure distortion ≤ JND; for high-frequency blocks: quantization is performed according to the process of "sub-band level → direction recognition → block mean → step size adaptation," allowing for moderate redundancy removal; post-quantization screening: high-frequency coefficients with absolute values ​​less than the just-perceptible distortion threshold JND are removed, while visually important coefficients are retained to achieve visual perception optimization compression. The quantized coefficients are directly converted into binary sequences for encoding.

[0054] To efficiently store and reliably reconstruct compressed image data in DNA media, this invention serializes the binary sequence obtained from quantization coefficient conversion. Specifically, the binary sequence from quantization coefficient conversion is first segmented, that is, the image block data is divided into fixed-length binary data segments, and fields containing header and check information are appended. To ensure that the length of the encoded DNA sequence does not exceed 200 nt, meeting the length limitations of current DNA synthesis technology, the length of each binary sequence segment is set to 32 bytes. This 32-byte structured design includes the following components: 2 bytes of header information for random oligonucleotide sequencing; 2 bytes of cyclic redundancy check (CRC) code for detecting random errors that may occur during encoding and decoding; and 28 bytes of core data for encoding binary information. The final compressed byte stream follows the structure of "2 bytes of header information + 28 bytes of core data information + 2 bytes of CRC check information".

[0055] Step 2: Dual-mode DNA encoding: The binary sequence to be encoded is divided into 8-bit segments into one byte, and encoded using an alternating odd-even mode and a segmented structure of "feature segment – ​​coding segment" to generate a DNA information segment sequence that conforms to biological constraints.

[0056] Before mapping the compressed binary sequence to a DNA sequence, the data is grouped into bytes, and each byte is independently mapped to a fixed-length DNA sequence, with each byte corresponding to 5 bases. This fixed-length mapping ensures a clear interface structure and facilitates subsequent error location and correction.

[0057] This invention proposes a dual-mode DNA encoding scheme, employing an alternating odd-even pattern: odd-indexed bytes use mode A (GC content of 40%), while even-indexed bytes use mode B (GC content of 60%). By alternating between the two modes, the synthesized DNA fragment from any two consecutive bytes maintains a macroscopic GC content of 50%, effectively avoiding sequence instability caused by excessively high or low GC content.

[0058] Furthermore, certain base combinations (such as "GCG", "CGC", "ATA", "TAT", etc.) are prone to generating high error rates during sequencing. To address this issue, this invention reduces sequencing errors at the coding level by decreasing the frequency of high-risk base combinations, thereby improving the reliability and stability of DNA storage.

[0059] At the single-byte level, DNA encoding adopts a segmented structure of "feature segment – ​​coding segment", as shown in Table 1.

[0060] Table 1. GC base position distribution rules for Pattern A and Pattern B

[0061]

[0062] The feature segment (high 3 bits) is used for indexing, determining the allocation sites of G / C bases in the 5 base positions corresponding to that byte. Two G / C positions are allocated in mode A, and three G / C positions are allocated in mode B. The encoding mapping table in Table 2 is shared between the encoder and decoder, and modes prone to forming long homopolymers or high-risk k-mers are avoided during the design phase.

[0063] Table 2 Dual-mode DNA coding mapping table

[0064]

[0065] The lower 5 bits of the encoding segment correspond to 5 base positions, determining the specific base type. When a base position is indexed as a G / C position, the binary bits are 0→G and 1→C; when it is an A / T position, the bits are 0→A and 1→T.

[0066] This segmented design controls GC content while suppressing the formation of long homopolymers, thereby improving the biological stability and sequencing accuracy of the DNA sequence. Taking the 16-bit data "0000000011111111" as an example: The first byte (00000000) is located in an odd position, using mode A. The high 3 bits = 000 determine the G / C position distribution {1,3}; the low 5 bits 00000 encode the sequence GAGAA. The second byte (11111111) is located in an even position, using mode B. The high 3 bits = 111 determine the G / C position distribution {2,4,5}; the low 5 bits 11111 encode the sequence TCTCC. The synthesized two-byte sequence is GAGAATCTCC, where G / C and A / T each occupy 5 positions, the GC content is 50%, and no homopolymers with more than 4 positions appear, meeting the requirements of biosynthesis and sequencing design.

[0067] This dual-mode DNA encoding scheme ensures high-density information storage while taking into account biosafety and operability, laying a reliable foundation for subsequent decoding, error correction, and information recovery processes.

[0068] The specific implementation method of the dual-mode DNA encoding using the byte-level independent mapping strategy is as follows:

[0069] 1. Byte splitting and mode determination: The compressed binary sequence to be encoded is split into 8-bit units. The encoding mode is determined according to the index position of the byte in the binary sequence (counting from 1): odd-indexed bytes adopt mode A (GC content 40%), and even-indexed bytes adopt mode B (GC content 60%).

[0070] 2. Feature Segment Parsing and GC Bit Allocation: Extract the high 3 bits of each byte as a feature segment. Based on Table 1 (GC base position distribution rules for Mode A and Mode B), determine the allocation scheme of GC bits and AT bits in the 5 base positions. For example, when the feature segment is 000 in Mode A, the GC bits are allocated to the 1st and 3rd positions, and the rest are AT bits; when the feature segment is 111 in Mode B, the GC bits are allocated to the 2nd, 4th, and 5th positions, and the rest are AT bits.

[0071] 3. Encoding Segment Mapping and Base Generation: Extract the lower 5 bits of the byte as the encoding segment, and generate specific bases according to the following rules: If the base position is a GC bit: the corresponding bit in the encoding segment is 0 and mapped to G, and 1 and mapped to C; if the base position is an AT bit: the corresponding bit in the encoding segment is 0 and mapped to A, and 1 and mapped to T.

[0072] 4. Mapping Verification and Correction: After generating the base sequence, verify whether its GC content meets the pattern requirements (pattern A 38%-42%, pattern B 58%-62%), and check for the presence of more than three consecutive identical bases. If it does not meet the constraints, remap by adjusting the GC bit allocation scheme corresponding to the feature segment (selecting the suboptimal solution within the allowable range in Table 1) to ensure that the 5-base sequence generated by each independent mapping meets the biological constraints.

[0073] To meet the constraints and information encoding requirements of synthetic biology, each encoded DNA sequence consists of an information segment, an index segment, a check segment, and primers at both ends, with a fixed total length of 200 nt to meet the experimental constraints of PCR amplification and high-throughput sequencing. Figure 2 As shown.

[0074] Information segment: 140 nt in length, used to store 28 bytes of core data information. During encoding, the 28 bytes of compressed data are mapped into a 140 nt DNA information segment sequence (28 bytes × 5 nt / byte = 140 nt) according to the "8-bit → 5-base" rule using the dual-mode DNA encoding method described in step three.

[0075] Index segment: 10 nt in length, corresponding to 2 bytes of header information, used to identify the position of each DNA sequence in the original image. It consists of two parts: the "image patch index" and the "information index". The image patch index (5 nt, corresponding to 1 byte) divides the original image data into 16×16 pixel units, assigning a unique number to each image patch to locate the sequence globally within the image. The information index (5 nt, corresponding to 1 byte) is used to mark the order of multiple sequences within the same image patch, ensuring accurate reconstruction of the data in the original order during the decoding stage.

[0076] Checksum segment: 10 nt in length, corresponding to 2 bytes of CRC checksum data, used for data integrity verification and error detection. The CRC-16 algorithm is used to uniformly verify both the information segment and the index segment. During decoding, if the checksum fails, the system will automatically trigger a local error correction procedure, thus implementing a complete data protection process.

[0077] Primer segments: 40 nt in length. To ensure the efficiency and stability of PCR amplification and sequencing, specific primers of 20 nt are added to both ends of each sequence. Based on the method of synthesizing DNA sequences according to the corresponding base sequences, primers are added to both ends of the base sequence, an index is added between the base sequence ends and the primers, and a check site is added between the base sequence ends and the primers to form the DNA sequence.

[0078] Overall, the above design effectively integrates the three major objectives of data localization, local error correction, and biocompatibility under a fixed length constraint. This processing method has the following advantages: (1) it significantly reduces data redundancy while preserving the main visual features of the image, thereby improving overall storage efficiency; (2) each byte is independently mapped to a fixed-length DNA sequence, facilitating local error detection and correction; (3) it provides structured data input for subsequent encoding and decoding stages, thereby enhancing the reliability and scalability of the entire DNA information storage scheme.

[0079] Step 3: Correct the encoded DNA sequence using preliminary correction based on Hamming distance and RS code secondary error correction.

[0080] To improve the reliability of DNA sequences during storage and sequencing, the dual-mode DNA encoding scheme incorporates multi-layer constraints and local error correction design, including byte-level constraints, cross-byte homopolymer control, and sliding window handling strategies for insertion / deletion errors, thereby constructing a sequence-level local error correction framework.

[0081] Byte-level mapping and encoding constraints: Each 8-bit data byte is independently mapped to a 5-base DNA sequence (odd-indexed bytes use pattern A, and even-indexed bytes use pattern B), forming a pattern A / B encoding mapping table (as shown in Table 2). Byte-independent mapping ensures that a single base substitution error will not propagate to other bytes, thus facilitating local error localization.

[0082] The high 3 bits of the feature segment determine the GC bit distribution position during the encoding stage. Even if the original binary information cannot be accessed during the decoding stage, the Hamming distance can be used to make a preliminary judgment on abnormal sequences. The high 3 bits of the feature segment are directly bound to the GC bit distribution (for example, the GC bits of mode A are fixed at the 1st and 3rd bits). If the GC bit position of the received sequence does not match the feature segment, it can quickly locate whether it is a "feature segment mapping error" or a "coding segment mapping error", reducing the error correction cost of traversing the entire sequence, avoiding invalid subsequent decoding calculations, and reducing computing power consumption.

[0083] Intra-byte Hamming distance constraint: The high 3 bits of the feature segment further restrict the legal codeword combinations for a single byte. When a base substitution error occurs, the Hamming distance can be used to quickly locate and correct it. The specific steps are as follows:

[0084] 1. Grouping and Hamming distance calculation: The read DNA sequences (length L) are grouped into 5-base groups, and the Hamming distance between each group of sequences and all codewords in the coding mapping table is calculated;

[0085] 2. Error type determination: Calculate the Hamming distance between the current 5-base group and each codeword in the encoding mapping table shown in Table 2, and record the minimum Hamming distance and the corresponding candidate codeword; if the minimum Hamming distance is 0, it is determined that there is no substitution error and the group is directly retained; if the minimum Hamming distance is 1, it is determined that there is a single base substitution error; if the minimum Hamming distance is ≥2, it is temporarily marked as a complex error and handled by the subsequent error correction mechanism.

[0086] 3. Precise error location: Compare the candidate codeword with the base sequence of the current group to find the base sites that are different at the corresponding positions. Combine the GC / AT position distribution defined by the feature segment to verify the error type: If the error site is at the GC position, it can only be a substitution between G and C; if it is at the AT position, it can only be a substitution between A and T, excluding the possibility of cross-type substitution.

[0087] 4. Constraint Compliance Correction: The current group is replaced with the candidate codeword corresponding to the minimum Hamming distance. If there is only one candidate codeword corresponding to the minimum Hamming distance, the 5-base group is directly replaced with the candidate codeword. If there are multiple candidate codewords (with the same Hamming distance and all being the minimum value), the coding consistency of adjacent groups is considered for screening: Candidate codewords with no homopolymer exceeding the limit (≤3nt) in base connection with the preceding group and the following group, and no error-prone adjacent base combinations (such as "GCG" and "ATA") are given priority.

[0088] Cross-byte homopolymer control: Through alternating odd and even modes (Mode A: 40% GC content, Mode B: 60% GC content) and optimized GC bit mapping rules, the overall GC content of DNA sequences synthesized from any two consecutive bytes is 50%, while avoiding more than four consecutive identical bases. This mechanism suppresses the propagation of sequencing errors caused by homopolymers, prevents local substitution errors from accumulating along the sequence, and enhances overall fault tolerance.

[0089] To address potential base insertion / deletion errors during sequencing or synthesis, single-byte Hamming distance cannot directly correct them. Therefore, a sliding window method is introduced to correct these errors. Specifically, a sliding window is applied across the received DNA sequence in 5-base units, progressively comparing the sequence against the dual-mode DNA coding map shown in Table 2 and calculating the Hamming distance. When the sequence length within the window does not match the standard codeword length, an attempt is made to insert or delete a base to generate a candidate sequence. Each candidate sequence must satisfy the GC / AT distribution constraints of its corresponding mode. The candidate sequence with the smallest Hamming distance and that satisfies the coding constraints is selected as the corrected result; otherwise, the window continues to slide and the operation is repeated. This method fully utilizes the coding map constraints shown in Table 2 and the alternation of odd and even modes to effectively limit the spread of insertion / deletion errors between adjacent bytes.

[0090] Local error detection: Substitution errors are identified using Hamming distance. The sequenced DNA segments are grouped by nt (unit). Based on the parity of the original byte index corresponding to each base segment, codewords in the corresponding mode of the dual-mode coding table are matched. A Hamming distance threshold of 1 is set. If the minimum Hamming distance of a group is 0, no substitution error is considered, and the group is retained. If the minimum Hamming distance is 1, a single-base substitution error is considered. If the minimum Hamming distance is ≥2, it is not initially considered a simple substitution error and is handled by subsequent error correction mechanisms to avoid false corrections. Insertion / deletion errors are detected using a sliding window analysis of abnormal length or positional offsets.

[0091] Potential correction strategies: Replacement errors are directly corrected based on the principle of minimum Hamming distance. The Hamming distance of the base sequence after grouping into 5nt blocks is calculated and compared with the codewords in the coding table. When a replacement error is determined, if there is only one candidate codeword with the minimum Hamming distance, the 5nt block is directly replaced with that candidate codeword. If there are multiple candidate codewords (with the same minimum Hamming distance), the coding consistency of adjacent blocks is considered for screening: priority is given to candidate codewords with no homopolymer excess (≤3nt) in base connection with the preceding and following blocks, and no error-prone adjacent base combinations (such as "GCG" and "ATA"). Insertion / deletion errors are corrected by adjusting the sequence through a sliding window to make it conform to the coding constraints again. It should be noted that the Hamming distance of some codewords is less than 2, and insertion / deletion errors may still cause the sequence to locally match the coding table. Therefore, a global error correction mechanism (such as RS code) is needed to further ensure data reliability. The specific steps are as follows:

[0092] Reed-Solomon (RS) error correction code is a commonly used forward error correction code. RS coding is based on polynomial operations over finite fields, encoding information symbols by mapping them to polynomial coefficients. The process involves first dividing the data to be transmitted into blocks of information symbols, then calculating redundant symbols using a generator polynomial, which, together with the original information symbols, constitute the coded data. At the receiving end, polynomial interpolation and error correction localization are used to detect the location and number of errors, thereby correcting the erroneous symbols.

[0093] This local error correction mechanism, together with the image restoration module, constitutes a sequence-image joint restoration framework, the effectiveness of which will be verified in simulation experiments.

[0094] Step 4: DNA sequences that still fail CRC check after multiple rounds of error correction are marked as unrepairable sequences, and the image blocks corresponding to unrepairable sequences are called error blocks. For these error blocks, the Cross Transformer Context Repair Network (CTNet) is used to learn the neighborhood features of the error blocks in combination with context information, and then the residual distortion error blocks are repaired.

[0095] After multiple rounds of error correction and repair, the image blocks are reassembled according to the original index, and then the image is restored by inverse transformation (such as inverse DWT) to be consistent with (or nearly consistent with) the original input, and finally a high-fidelity complete image is reconstructed.

[0096] To achieve integrity verification and high-quality recovery of image data during the decoding stage, this invention designs a two-level recovery mechanism: error detection and local error correction are implemented at the sequence level, and compensatory repair is performed at the image level through a deep learning model, thereby achieving collaborative protection of data error correction and visual recovery.

[0097] In the sequence structure design, each DNA sequence is appended with a 2-byte Cyclic Redundancy Check (CRC) code for verifying data integrity during decoding. During decoding, a dual detection process is performed, combining local error correction and CRC verification. If the DNA sequence passes the CRC check or is restored to consistency after local error correction, it is considered a valid sequence and proceeds to the recombination and image reconstruction stage. If the DNA sequence fails the CRC check after multiple rounds of error correction attempts, it is marked as an unrepairable sequence, and its corresponding image block undergoes further processing at the image level. CRC verification not only provides redundancy verification at the sequence level but also provides a reliable index for error localization and subsequent image restoration.

[0098] Regions determined to be irreparable at the sequence level are denoted as image blocks of error blocks, and the set is defined as follows:

[0099] (1)

[0100] in, This represents the k-th erroneous block. For error block The validity indicator. Represents the set of error blocks The number of elements, i.e., the total number of irreparable image patches. The range of values ​​for is: Total number of image blocks (Total number of image blocks = Original image width × Original image height / (16 × 16)), Represents the set of error blocks The k-th image block in the image, each image block corresponds to a 16×16 pixel sub-block in the original image, and its data comes from 28 bytes of compressed data of the information segment in the DNA sequence.

[0101] This invention introduces the CTNet network to extract the neighborhood spatial context features of erroneous blocks and fuse sequential blocks ( ), parallel blocks ( ) and residual blocks ( This structure enables multi-scale spatial feature aggregation and pixel detail reconstruction.

[0102] (2)

[0103] Wherein, S(b) k ) indicates error block b k Spatial neighborhood information, i.e., selecting the wrong block b k The normal image patch within a 3×3 range is used as input features, which includes the spatial structure, texture details and pixel association information around the error patch; These are the operation functions for serial blocks. These are the operation functions for parallel blocks. These are the operation functions for the residual block; This is the result of the reconstruction after restoration.

[0104] Serial block ( Based on the "deep search" design concept, this architecture is constructed by fusing an enhanced residual architecture with a Transformer mechanism. Its basic architecture employs an enhanced residual structure, consisting of alternating stacks of linear transformation layers and nonlinear activation layers. The linear transformation layers are responsible for adjusting feature dimensions and initially extracting structural information, while the nonlinear activation layer (GELU) introduces nonlinear feature representation capabilities. Simultaneously, a multi-head self-attention Transformer mechanism is embedded to strengthen cross-location feature associations by calculating global dependencies between pixels. The overall process adopts a serial flow of "feature extraction - Transformer enhancement - residual fusion" to ensure sufficient capture and deepening of core structural features surrounding erroneous blocks; parallel blocks ( Based on the "breadth-first search" design concept, this architecture is constructed using three heterogeneous parallel networks. Three structurally distinct base networks (lightweight convolutional networks, deep convolutional networks, and dilated convolutional networks) are selected as parallel branches to form heterogeneous feature extraction channels. Each branch extracts features from three dimensions: local details, mid-level structure, and global context. Multi-level feature fusion is achieved through feature concatenation and interactive computation. Furthermore, Transformer interaction units are embedded within the branches to strengthen pixel-level correlation information between features from different branches, preventing the loss of key details and improving the network's adaptability to complex image scenes. Residual blocks ( Based on the concept of pre-activation residual learning, it is the core output unit of image reconstruction. It adopts a "BatchNorm-activation function-convolution" pre-activation structure to reduce the gradient vanishing problem and improve feature transfer efficiency. The core contains two 3×3 convolutional layers. The first layer is responsible for processing parallel blocks. The output fused features undergo noise suppression and feature optimization. The second layer is responsible for mapping the optimized features to reconstructed features in the image pixel space, while introducing direct residual connections to convert serial blocks... The output deep structural features are element-wise added to the current layer's optimized features, preserving the original effective information, and finally outputting clear reconstructed image patch features.

[0105] The CTNet network can effectively learn the relationship between local texture and global structure, and can still generate images with continuous structure and natural appearance even with a small proportion of data loss or sequencing errors.

[0106] CRC checksums serve as part of data redundancy during the encoding phase and are also used in the image restoration phase to mark error locations and trigger the CTNet network for targeted repair. Sequence-level error correction ensures data-level reliability, while image-level restoration enhances visual integrity. This dual-layer design effectively strengthens the robustness and recoverability of the DNA image storage system.

[0107] To verify the feasibility and effectiveness of the proposed DNA-GCoder scheme in image storage and decoding, a multi-level simulation experiment was designed. The analysis comprehensively considered aspects such as biological constraint satisfaction, coding robustness, error correction performance, and image reconstruction quality to evaluate its stability, reliability, and storage density in the actual DNA information writing and reading process.

[0108] The experimental data came from the image datasets DIV2K, BSD, and Set12. Representative image samples were selected from these datasets, covering various structural features such as complex textures, smooth regions, and grayscale gradients, to comprehensively reflect the adaptability of the algorithm under different image types. The experimental procedure consisted of four parts: (1) DNA sequence feature analysis to verify whether the GC content and homopolymer length of the coding sequence met biological constraints; (2) Robustness test of the coding algorithm to analyze the decoding recovery capability under substitution, insertion, and deletion errors; (3) System-level multilayer error correction performance evaluation to examine the recovery effect under different error rates; (4) Image reconstruction and visual enhancement performance analysis to evaluate the final image quality and biological compatibility of combining error correction and CTNet networks.

[0109] All experiments were conducted in the Python 3.8 environment, and data were constructed and encoded according to the designed DNA sequence structure (information segment, index segment, check segment, and primer segment) to ensure that the experimental results were consistent with the aforementioned system model.

[0110] Based on the proposed DNA sequence structure and dual-mode DNA encoding strategy, the biocompatibility of the DNA-GCoder of this invention was evaluated from two aspects: local GC content and homopolymer length. Using a 95.2kB Mona Lisa image as the test object, it was divided into 3,482 segments of 28 bytes each, which were then encoded to generate 3,484 DNA sequences of 140 nt each. Comparisons were made between two representative schemes: Blawat (from the literature [Blawat M, Gaedke K, Huetter I, et al. (2016). Forwarderror correction for dna data storage. Procedia Computer Science, 80, 1011-1022.]) and DNA Fountain (from the literature [Erlich Y, & Zielinski D. (2017). Dnafountain enables a robust and efficient storage architecture. Science, 355(6328), págs. 950-954.]), and the results are as follows: Figure 3 As shown.

[0111] Local GC content equilibration can reduce the probability of DNA secondary structure formation, thereby improving PCR amplification and sequencing efficiency. Statistical analysis was performed on the first 1500 nt of each protocol using a 15nt sliding window; the results are as follows: Figure 3 As shown in (a), the results show that the local GC content of the DNA-GCoder is stable at 45%–55%, with an average of 50.02% and a fluctuation range of ≤5%, with no abnormal windows exceeding 40%–60%. This is due to the dual-mode alternating odd-even encoding mechanism, which maintains global base balance during continuous byte splicing. In contrast, the local GC content of Brawat fluctuates between 20% and 73%, with 32% of the windows exceeding the biological constraints, demonstrating the significant advantage of the DNA-GCoder of this invention in terms of GC balance.

[0112] Homopolymer length affects the accumulation of sequencing errors, as shown in the following results. Figure 3As shown in (b), experimental results show that the maximum homopolymer length of DNAFountain is approximately 3–4 nt; the maximum homopolymer length of Blawat and DNA-GCoder is ≤3 nt, but Blawat homopolymers are more concentrated, with 3 nt homopolymers occurring more frequently. The DNA-GCoder of this invention, through the dispersion constraint of a preset GC position table, suppresses cross-byte homopolymer aggregation. Even if consecutive bytes are all "0"s or all "1", it can still maintain GC / AT alternation and balance, significantly reducing the risk of read errors caused by homopolymers.

[0113] In summary, the DNA-GCoder of this invention exhibits excellent performance in local GC balance and homopolymer length control, providing a stable foundation for subsequent coding robustness and error correction performance.

[0114] The robustness of DNA storage systems depends not only on error correction mechanisms but also on the fault tolerance of the coding structure itself. During sequencing, substitution, insertion, and deletion errors can easily lead to sequence shifts or synchronization failures; therefore, coding-level fault-tolerant design is crucial. This invention uses MonaLisa images as test subjects, encoding them into 3,484 160nt sequences. Under a simulated environment, 0.1%–2% substitution, insertion, and deletion errors were randomly injected to evaluate the recovery capability of the DNA-GCoder under different error types and intensities, and compared with the DNA Fountain scheme.

[0115] like Figure 4 As shown in (a), under a single substitution error mode, the DNA-GCoder of this invention consistently achieves a data recovery rate >95%, an improvement of approximately 7% compared to the control scheme. This advantage stems from the dual-mode parity alternation and local Hamming distance constraint mechanism, which enables the coding structure to identify and locally correct substitution errors.

[0116] Insertion / deletion errors can easily lead to synchronization failures. Because insertion / deletion errors alter the length of the entire sequence and propagate the error to subsequent base sequences, they significantly reduce data recovery rates. The binary data recovery rates under insertion and deletion errors are as follows: Figure 4 As shown in (b), experimental results demonstrate that even when faced with such problems, the DNA-GCoder scheme of this invention still exhibits superior robust data recovery capabilities compared to the DNA Fountain scheme. This is due to the sliding window byte alignment mechanism, which automatically adjusts the alignment offset and prevents error propagation.

[0117] Randomly injecting a mixture of substitution, insertion, and deletion errors (ratio 27:1:1) within the error rate range of 0.1%–10% to simulate actual sequencing errors on the Illumina MiSeq platform, such as... Figure 4As shown in (c). The results show that the overall recovery rate is >90% at 1%–2% error rate; about 72% at 5%; and drops to 40% at 10%, indicating that the coding self-recovery mechanism is approaching its limit in high-noise environments.

[0118] By employing byte-level independent mapping, local Hamming distance constraints, and sliding window alignment, the DNA-GCoder of this invention constructs a naturally fault-tolerant coding structure, ensuring high data recovery accuracy at low to medium error rates and providing a reliable foundation for multi-layer error correction.

[0119] To further improve image recovery capabilities in high error rate scenarios, the DNA-GCoder of this invention introduces a multi-layer error correction system, including CRC check, Hamming correction, sliding window positioning, and RS code cross-byte error correction, realizing a fault-tolerant mechanism from sequence-level "detection-positioning-repair".

[0120] Taking three 256×256 grayscale images—Lena, Fly_grey, and Baboon—as examples, substitution, insertion, and deletion errors were simulated under a total error rate of 0.5%–5% (ratio 8:1:1). This was compared with Grass (from the literature [Grass RN, Heckel R, Puddu M, et al. (2015). Robust chemical preservation of digital information on DNA in silica with error-correcting codes. Angew Chem Int Ed Engl, 54(8),2552-2555.]), Blawat, and HL-DNA (from the literature [Li Y, Du DH, Ou L, et al. (2022, October). HL-DNA: a hybrid lossy / lossless encoding scheme to enhance DNA storage density and robustness for images. In 2022 IEEE 40th International Conference on Computer Design (ICCD) (pp. 434-442)]. A systematic comparison was conducted between DNA-GCoder and existing encoding schemes such as [IEEE.] and DNA-QLC---from the literature [Zhen Y, Cao B, Zhang X, et al. (2024). DNA-QLC: an efficient and reliable image encoding scheme for DNA storage. BMC genomics, 25(1), 266.] to evaluate the image restoration quality of DNA-GCoder under different error rates.

[0121] At an error rate of 0.5%, CRC checksum, combined with a sliding window and RS code mechanism, eliminates approximately 95% of base errors. For example... Figure 5As shown, the decoded image achieved a PSNR of 34.2 dB and an SSIM of 0.92. In contrast, the images restored by the Grass and Blawat schemes had PSNR values ​​below 29 dB and SSIM values ​​below 0.6, exhibiting significant detail blurring and noise interference. HL-DNA still suffers irreversible detail loss under rotational coding and lossy compression, limiting its restoration performance. DNA-QLC's error correction mechanism is prone to error propagation when faced with insertion / deletion errors, thus affecting the integrity of image restoration. Clearly, under the same error rate conditions, DNA-GCoder's image restoration performance is significantly superior to existing comparative schemes.

[0122] Within a high error rate range of 1%–5%, using Lena images as the primary test subject, the stability of the scheme was verified through 100 independent repeated experiments. Figure 6 As shown, the PSNR and SSIM scores of the recovered images obtained by the DNA-GCoder of this invention remain consistently stable and are both higher than those of other comparative coding schemes, demonstrating excellent robustness in high-error-rate image recovery. The error correction performance of the DNA-GCoder is mainly attributed to three aspects: byte-level independent mapping and coding table constraints limit the propagation of substitution errors; sliding window processing of insertion / deletion errors ensures that cross-byte errors can be locally corrected; and RS code global error correction provides a final guarantee for the repair of complex errors. Multi-layer error correction significantly improves data consistency under high error rates, providing reliable input for visual-level enhancement.

[0123] While sequence-level multi-level error correction can repair most base errors, local high-density errors or consecutive insertions / deletions can still produce residual noise and blurred details. To further improve image reconstruction quality, this invention introduces a CTNet-based deep convolutional visual restoration algorithm to predict and compensate for residual error blocks after multi-level error correction.

[0124] The CTNet network utilizes inter-pixel spatial correlation and contextual information to recover damaged image structure, enhance texture details, and improve overall visual quality. For a given test image, Figure 7 The image reconstruction results are shown under different error rates. Experimental results show that after introducing the CTNet network, the PSNR of the reconstructed image is improved by 0.05–4.2 dB (orange curve), and the SSIM is improved by 0.03–0.12 (blue curve). The optimization effect of image reconstruction under high error rate conditions is more significant. Taking the Lena image at a 2% error rate as an example, the CTNet network can effectively reduce artifacts and block noise in the reconstructed image, significantly enhance image clarity and visual readability, and make the reconstructed image closer to the original image.

[0125] By combining with CRC checksum, the DNA-GCoder of this invention achieves end-to-end recovery from sequence-level fault tolerance to vision-level optimization: CRC accurately locates residual errors, and CTNet performs semantic compensation and detail optimization, forming a three-layer optimization architecture. This framework ensures sequence data integrity and global image consistency, while improving visual quality and structural fidelity, providing a reliable cross-layer image reconstruction solution for DNA image storage in high-error-rate environments.

[0126] In summary, visual-level reconstruction achieves comprehensive optimization from sequence fault tolerance to high-quality visual-level reconstruction in terms of compensating for residual errors, restoring local structures, optimizing texture details, improving readability, enhancing generalization ability, and maintaining structural continuity. This provides a solid technical guarantee for high-fidelity image reconstruction in DNA storage scenarios.

[0127] Table 3 compares the performance of the DNA-GCoder of this invention with other schemes in key performance indicators such as biological constraints, error correction strategies, and net coding density (number of information bits / nucleotides). Church is referenced from [Church GM, Gao Y, & Kosuri S. (2012). Next-generation digital information storage in DNA. Science, 337(6102), 1628-1628.], and Goldman is referenced from [Goldman N, Bertone P, Chen S, et al. (2013). Towards practical, high-capacity, low-maintenance information storage in synthesized DNA. Nature, 494(7435), 77-80.]. Compared to earlier DNA coding schemes, DNA-GCoder achieves precise control of GC content (stabilized at 50%) through a dual-mode DNA coding rule, thereby significantly optimizing DNA coding constraints, reducing the complexity of subsequent synthesis and sequencing, and to some extent reducing potential errors. Although the DNA-QLC scheme of this invention has already exceeded the net information density threshold of 2.00 bits / nt, DNA-GCoder further improves upon this, achieving even higher information storage efficiency. Unlike schemes that rely on external error correction mechanisms (such as forward error correction, RS codes, and barrier error correction), DNA-GCoder utilizes the inherent Hamming distance characteristic of the encoding table to achieve basic error correction, eliminating the need to introduce additional redundant error correction sequences beyond the mapping relationship between binary data and DNA sequences, thus improving data reading efficiency and increasing encoding density. In summary, the DNA-GCoder scheme satisfies GC balance and homopolymer constraints while balancing high encoding efficiency and biological stability. This demonstrates that this invention achieves synergistic optimization of encoding efficiency and biological compatibility, making it particularly suitable for image storage scenarios.

[0128] Table 3 Comparison of the present invention with representative storage schemes

[0129]

[0130] (Note: "*" indicates that this scheme uses a compression algorithm.)

[0131] This invention addresses the core issues of limited information density and insufficient error correction capabilities in image-based DNA storage by proposing a novel encoding scheme—DNA-GCoder. First, wavelet transform is used to compress the image, effectively reducing data redundancy and improving encoding efficiency. Simultaneously, a dual-mode DNA encoding scheme is designed, using alternating odd and even sequences and GC distribution constraints to avoid potentially biologically unstable sequences, ensuring the compatibility and stability of the generated DNA sequence during synthesis and sequencing. Experimental results show that the DNA-GCoder of this invention can stably control the local GC content within the range of 40%–60%, with an overall GC content approaching 50%, and strictly control the maximum homopolymer length, thereby significantly improving sequence balance and diversity, providing a solid foundation for the reliable operation of DNA storage systems under biological constraints.

[0132] Regarding fault tolerance and recovery mechanisms, this invention combines CRC checksum and RS code to achieve cross-byte global error correction, and introduces image restoration based on the deep learning-based CTNet network to further improve reconstruction quality. Experiments show that this mechanism can effectively resist substitution, insertion, and missing errors, significantly reducing the risk of complete data loss. However, for applications with extremely high requirements for detail restoration, such as medical imaging or deep neural network parameter storage, the predictive restoration of the CTNet network still has certain uncertainties, posing a challenge to high-precision image storage.

[0133] Furthermore, the DNA-GCoder of this invention employs a lossy compression strategy to enhance information density. In summary, the DNA-GCoder of this invention demonstrates good potential in improving coding efficiency, enhancing biocompatibility, and improving error correction and recovery performance, thereby maximizing the information density of DNA storage while ensuring image quality, providing a feasible solution for large-scale high-fidelity image storage.

[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for storing image DNA that improves information density and biological stability, characterized in that, The steps are as follows: Step 1: Divide the input image into blocks and perform two-dimensional discrete wavelet transform. Use an adaptive quantization strategy to quantize the wavelet coefficients and convert the quantized wavelet coefficients into a binary sequence to be encoded. Step 2: Divide the binary sequence to be encoded into 8 bits per byte, and use a dual-mode DNA encoding with alternating odd and even patterns and a segmentation structure of "feature segment-coding segment" to generate a DNA information segment sequence that conforms to biological constraints. Step 3: Correct the encoded DNA sequence using preliminary correction based on Hamming distance and RS code secondary error correction; Step 4: DNA sequences that still fail the CRC check after multiple rounds of error correction are marked as unrepairable sequences, and the image blocks corresponding to unrepairable sequences are error blocks; By combining contextual information, the network learns the neighborhood features of erroneous blocks and repairs them using the cross-Transformer context repair network.

2. The image DNA storage method for improving information density and biological stability according to claim 1, characterized in that, The input image is divided into 16×16 pixel sub-blocks; a two-dimensional discrete wavelet transform is performed on each sub-block to convert the spatial domain pixels into a wavelet coefficient domain representation; the wavelet coefficients include low-frequency coefficients and high-frequency coefficients; The adaptive quantization strategy for wavelet coefficients is as follows: low-distortion uniform quantization is used for low-frequency coefficients to minimize the trade-off between quantization distortion and visual perception loss; non-uniform quantization is used for high-frequency coefficients. The quantization step size is designed based on the visual threshold matrix, and the high-frequency coefficients are classified according to their energy: high-energy high-frequency coefficients use a smaller quantization step size; low-energy high-frequency coefficients use a larger quantization step size. High-frequency coefficients with amplitudes below the perception threshold are directly set to zero. The coefficients of each sub-band after the two-dimensional discrete wavelet transform are divided into blocks. The just-perceptible distortion threshold JND of each block is calculated, and high-frequency coefficients with absolute values ​​less than the just-perceptible distortion threshold JND are removed.

3. The image DNA storage method for improving information density and biological stability according to claim 1 or 2, characterized in that, Before converting to the binary sequence to be encoded, the quantized wavelet coefficients are serialized. The implementation method is as follows: the odd-even alternation mode is implemented by using mode A with a GC content of 40% for odd index bytes and mode B with a GC content of 60% for even index bytes. The implementation method of the segmentation structure of "feature segment – ​​coding segment" is as follows: The high 3 bits of each byte are extracted as a feature segment, and the allocation scheme of G / C and A / T bits in the 5 base positions is determined according to the GC base position distribution rules of pattern A and pattern B. Extract the lower 5 bits of the byte as the encoding segment to determine the base type: when the base position is indexed as a G / C bit, binary bit 0 is mapped to base G and binary bit 1 is mapped to base C; when the base position is indexed as an A / T bit, binary bit 0 is mapped to base A and binary bit 1 is mapped to base T.

4. The image DNA storage method for improving information density and biological stability according to claim 3, characterized in that, In mode A, 2 G / C bits are allocated, and in mode B, 3 G / C bits are allocated. The GC base position distribution rules of Mode A and Mode B ; After generating the DNA information segment sequence, verify whether the GC content meets the pattern requirements, and check whether there are more than three consecutive identical bases. If it does not meet the constraints, adjust the GC position allocation scheme corresponding to the feature segment and remap it.

5. The image DNA storage method for improving information density and biological stability according to claim 3 or 4, characterized in that, The quantized wavelet coefficients are segmented into fixed-length binary data segments, and fields containing header and check information are added to obtain the binary sequence to be encoded. The length of the binary sequence to be encoded is 32 bytes, which includes 2 bytes of header information, 28 bytes of core data and 2 bytes of CRC check information. The dual-mode DNA encoding in step two processes the 28 bytes of core data. Each DNA sequence encoded by the dual-mode DNA consists of an information segment, an index segment, a check segment, and primers at both ends, with a total length of 200 nt; The information segment is 140 nt long and stores the 140 nt DNA information segment sequence encoded according to the "8-bit → 5-base" rule for 28 bytes of core data. The index segment is 10 nt long and corresponds to 2 bytes of header information. It consists of two parts: an image block index and an information index. The image block index is the number of the image block, and the information index is used to mark the order of multiple sequences within the same image block. The check segment is 10 nt long and corresponds to 2 bytes of CRC check data. The CRC-16 algorithm is used to perform unified verification on the information segment and the index segment. The primer segment is 40 nt long, with 20 nt primers added to both ends of each DNA sequence.

6. The image DNA storage method for improving information density and biological stability according to claim 5, characterized in that, The steps of the preliminary correction method for the Hamming distance are as follows: 1) Group the read DNA sequences into 5-base groups; 2) Calculate the Hamming distance between the current 5-base block and each codeword in the encoding mapping table, and record the minimum Hamming distance and the corresponding candidate codeword; if the minimum Hamming distance is 0, it is determined that there is no substitution error and the current 5-base block is retained; if the minimum Hamming distance is 1, it is determined that there is a single base substitution error; if the minimum Hamming distance is ≥2, it is marked as a complex error; the encoding mapping table is an 8-bit binary encoding of 5-base sequences in mode A and mode B; 3) Compare the candidate codeword with the base sequence of the current 5-base group to find the base sites that are different at the corresponding positions. Combine the GC / AT position distribution defined by the feature segment to verify the error type: if the error site is at the G / C position, it can only be a substitution between G and C; if it is at the A / T position, it can only be a substitution between A and T. 4) Replace the current 5-base group with the candidate codeword corresponding to the minimum Hamming distance. If there is only one candidate codeword corresponding to the minimum Hamming distance, directly replace the current 5-base group with the candidate codeword. If there are multiple candidate codewords with the same Hamming distance, prioritize the candidate codeword that has no homopolymer excess and no error-prone adjacent base combinations in the base connection with the preceding and following groups.

7. The image DNA storage method for improving information density and biological stability according to claim 6, characterized in that, To address base insertion / deletion errors during sequencing or synthesis, a sliding window method is introduced for error correction. The method involves sliding a window across the received DNA sequence in 5-base units, progressively comparing the coding map and calculating the Hamming distance. When the sequence length within the window does not match the standard codeword length, a base is inserted or deleted to generate a candidate sequence. Each candidate sequence must satisfy the GC / AT distribution constraints of its corresponding pattern. The candidate sequence with the smallest Hamming distance and that satisfies the coding constraints is selected as the correction result; otherwise, the window continues to slide and the operation is repeated. Substitution errors are identified using Hamming distance. The obtained DNA information segment sequences are grouped into units of 5 bases. Based on the parity of the original byte index corresponding to the base segment in the group, the codewords of the corresponding pattern in the encoding mapping table are matched. If the minimum Hamming distance of the group is 0, it is determined that there is no substitution error and the group is directly retained. If the minimum Hamming distance is 1, it is determined that there is a single base substitution error. If the minimum Hamming distance is ≥2, it is not determined to be a simple substitution error. Replacement errors are corrected based on the principle of minimum Hamming distance. The Hamming distance between the base sequence after grouping into 5 bases and the codeword in the coding mapping table is calculated. When a replacement error is determined, if there is only one candidate codeword corresponding to the minimum Hamming distance, the 5-base group is directly replaced with the candidate codeword. If there are multiple candidate codewords with the same Hamming distance, the candidate codeword with no homopolymer excess and no error-prone adjacent base combinations in the base connection with the preceding and following groups is given priority.

8. The image DNA storage method for improving information density and biological stability according to claim 6 or 7, characterized in that, The method for implementing RS code level 2 error correction is as follows: the DNA sequence is grouped into units of 5 bases and mapped to information symbols over a finite field; a generator polynomial-based encoder is used to generate check symbols and construct the system codeword; At the decoding end, the syndrome of the received codeword is calculated to detect errors; If an error is detected, the error location is located by solving the error location polynomial and the error value is calculated, thereby completing the error correction. For error blocks that still cannot be repaired; By combining contextual information, the network learns the neighborhood features of erroneous blocks and repairs them using the cross-Transformer context repair network.

9. The image DNA storage method for improving information density and biological stability according to claim 8, characterized in that, The cross-Transformer context repair network fuses serial blocks, parallel blocks, and residual blocks to achieve multi-scale spatial feature aggregation and pixel detail reconstruction, resulting in a restored reconstruction. ; Wherein, S(b) k ) indicates error block b k Spatial neighborhood information, error block b k A normal image patch within a 3×3 radius is used as the input feature. These are the operation functions for serial blocks. These are the operation functions for parallel blocks. These are the operation functions for the residual block.

10. The image DNA storage method for improving information density and biological stability according to claim 9, characterized in that, The set of error blocks is as follows: ;in, This represents the k-th erroneous block. For error block Validity markers Represents the set of error blocks The number of elements in each error block This corresponds to a 16×16 pixel sub-block in the original image; The serial block adopts an enhanced residual structure, which includes alternating stacks of linear transformation layers and nonlinear activation layers. The linear transformation layer is responsible for feature dimension adjustment and preliminary extraction of structural information, while the nonlinear activation layer introduces nonlinear feature expression capability and embeds a multi-head self-attention mechanism to calculate the global dependency relationship between pixels and strengthen cross-position feature association. The parallel block selects three structurally differentiated basic networks as parallel branches. The basic networks include lightweight convolutional networks, deep convolutional networks, and dilated convolutional networks. Each branch extracts features from three dimensions: local details, mid-level structure, and global context. Multi-level feature fusion is achieved through feature concatenation and interactive operations. Furthermore, Transformer interactive units are embedded within the branches to enhance pixel correlation information between features of different branches. The residual block adopts a pre-activation structure of BN-activation function-convolution, which contains two 3×3 convolutional layers. The first layer is responsible for noise suppression and feature optimization of the fused features output by the parallel block, and the second layer is responsible for mapping the optimized features to the reconstructed features in the image pixel space. Direct residual connections are introduced to add the deep structural features output by the serial block to the optimized features of the current layer element by element, and output clear reconstructed image block features.