Quaternary dna storage method and system based on entropy coding and rotation constraint coding, storage medium and terminal

By combining quaternary tANS entropy coding and rotation coding, the problems of low storage density and difficulty in simultaneously satisfying biochemical constraints in DNA storage methods are solved, achieving efficient data compression and biochemical constraint compliance, and improving coding stability and storage density.

CN122157810APending Publication Date: 2026-06-05TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-03-27
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing DNA storage methods struggle to balance high storage density, strict biochemical constraints, and low decoding complexity. They suffer from low storage density, low compression efficiency, and encoding schemes that cannot simultaneously meet biochemical constraints.

Method used

A quaternary DNA storage method based on entropy coding and rotation constraint coding is adopted. The input data is converted into a quaternary sequence through quaternary tANS entropy coding, and mapped to a DNA sequence that conforms to biochemical constraints, including CG content and homopolymer length constraints, through quaternary rotation coding rules, while avoiding specific harmful sequences.

Benefits of technology

It achieves high storage density and meets biochemical constraints while reducing decoding complexity, improving encoding compression efficiency and decoding stability, and achieving high storage density and biochemical constraint compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157810A_ABST
    Figure CN122157810A_ABST
Patent Text Reader

Abstract

The application discloses a quaternary DNA storage method and system based on entropy coding and rotation constraint coding, a storage medium and a terminal. First, compared with traditional binary coding, quaternary coding can improve storage density, thereby storing more data in smaller space. Second, dynamic rotation coding ensures that the DNA sequence meets strict biochemical constraints, including homopolymer length limitation, GC content balance and avoidance of specific harmful sequences, optimizes the synthesis and sequencing process of DNA, and reduces the error rate. In addition, the use of tANS coding provides higher data compression efficiency, further reducing storage costs. These technical innovations not only improve the performance of large-scale data storage, but also enhance the security and accuracy of data decoding, ensuring the integrity of data after long-term storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of DNA storage technology, specifically to a quaternary DNA storage method, system, storage medium, and terminal based on entropy encoding and rotation constraint encoding. Background Technology

[0002] DNA molecules, as a novel storage medium, have shown broad development prospects in addressing the current predicament of exponential growth in data volume. The principle of DNA storage is to encode binary information into DNA sequences, then synthesize DNA molecules and store them, and recover the original binary information through sequencing decoding when reading. As an information preservation encoding method, entropy encoding in the field of DNA storage has shortcomings in balancing high storage density, strict biochemical constraints, and low decoding complexity. Current research still faces the following challenges: (1) low storage density; (2) low compression efficiency; (3) most encoding schemes cannot simultaneously satisfy specific biochemical constraints.

[0003] In conclusion, it is necessary to improve existing DNA storage methods. Summary of the Invention

[0004] In view of the technical problem mentioned in the background art that existing methods of entropy encoding in the field of DNA storage are insufficient in balancing high storage density, strict biochemical constraints and low decoding complexity, the purpose of this invention is to provide a quaternary DNA storage method, system, storage medium and terminal based on entropy encoding and rotation constraint encoding.

[0005] To achieve the objectives of this invention, the technical solution provided by this invention is as follows: First aspect This invention provides a quaternary DNA storage method based on entropy encoding and rotation constraint encoding, comprising the following: (1) Perform quaternary tANS entropy encoding on the input data to convert the input data into a quaternary sequence; (2) The quaternary sequence is mapped through the quaternary rotation coding rule so that the generated DNA sequence meets the biochemical constraints, including: meeting the CG content constraints and homopolymer length constraints, and being able to avoid specific harmful sequences.

[0006] Furthermore, the step of performing quaternary tANS entropy encoding on the input data to convert the input data into a quaternary sequence specifically includes the following: (1.1) Statistical symbol frequency: The input data is a discrete source containing n independent symbols. Count and output source symbols Types of symbols and probability mass function Wherein, the definition is as follows: the data comes from a frequency count The distribution of the representation, where each Symbols Frequency of occurrence; (1.2) Initialization State and State Range Setting: The initialization state is set to 0. tANS restricts the allowed state values ​​to a fixed range, which is set as the state space. ,in is a positive integer, The number of symbol types is a parameter that determines the size of the state space; at the same time, it is based on the actual probability of the source-coded symbols. ,in For symbols The probability is used to assign the state range corresponding to each symbol. It is necessary to make As close as possible To achieve probability normalization approximation; (1.3) Design symbolic extension functions: based on the symbolic extension function... Frequency of occurrence and given interval , state space Weighted by sign probability Perform dynamic partitioning and design symbolic extension functions. Establish each state To source symbols The mapping of (1≤t≤n) is achieved by discretizing the symbols through probability redistribution, ensuring that the frequency of symbol occurrence is positively correlated with the state distribution density; (1.4) Construct the state transition table: based on Construct encoding and decoding functions and The encoding and decoding lookup tables are constructed by the encoding and decoding functions. During construction, if the current state is... Within a given symbol range Inside, directly through Mapping to the next state, if it's not within the given range, calculate the number of bits needed for a right shift. Then Move right Position, that is At the same time, will be removed The bits are stored in a bitstream variable, and then... Complete the encoding to obtain the next state; (1.5) Encoding: From the initial state Begin with the input symbol sequence Perform iterative encoding; each iteration passes through the given current state. and symbols to be encoded Look up the encoding table to get the next state. and the corresponding binary output segment The process continues until the last symbol is encoded, at which point the final output is the termination status. and synthesized binary sequences ,symbol Indicates will Assemble them sequentially; (1.6) Initial state Convert to Bit binary number Termination status Convert to Bit binary number Append to binary sequence At the end, an extended sequence is formed. , as a state range marker during decoding; calculate The length is such that if it is odd, zero bits are padded at the end to generate... Grouped by double bits Will Convert to quaternary sequence Q and add padding marker bits. The presence or absence of padding is indicated by 1 for padding and 0 for no padding, ultimately generating a quaternary encoded result. As input to the dynamic rotation encoding module, this conversion method ensures decoding integrity through state metadata encapsulation, achieves lossless conversion through parity alignment, and the final output quaternary sequence contains both the original data stream and encoding / decoding metadata. and .

[0007] Furthermore, in step (2), the quaternary rotation encoding rule is established based on the dynamic mapping relationship between the quaternary number system and DNA bases; after quaternary tANS encoding, a quaternary sequence is obtained, and the input sequence is defined as a set of quaternary numbers. The initial mapping rule is: Specifically: map 0 to A, 1 to C, 2 to T, and 3 to G; the dynamic rotation mapping rule is as follows: Specifically, for each subsequent character starting from the second character... The corresponding DNA bases Depends on the previously generated base and the currently entered number The design logic of this rule is to divide the leading base into two groups: the purine group (A, G) and the pyrimidine group (T, C). When the input is A or G, the outputs for 0, 1, 2, and 3 are C, T, CG, and TA respectively, indicating the current base. When the input is T or C, the outputs are A, G, AT, and GC respectively for inputs of 0, 1, 2, and 3.

[0008] Further, in step (2), meeting the CG content constraint means that the CG content is concentrated between 49% and 51%; the homopolymer length constraint means that the maximum homopolymer length is 1; and the specific harmful sequence means that the sequence set is .

[0009] Second aspect Corresponding to the above method, the present invention also provides a quaternary DNA storage system based on entropy coding and rotation constraint coding, comprising: a conversion unit and a mapping unit; The conversion unit is used to perform quaternary tANS entropy encoding on the input data and convert the input data into a quaternary sequence. The mapping unit is used to map the quaternary sequence through quaternary rotation coding rules so that the generated DNA sequence meets biochemical constraints, including: meeting CG content constraints and homopolymer length constraints, and being able to avoid specific harmful sequences.

[0010] Furthermore, the step of performing quaternary tANS entropy encoding on the input data to convert the input data into a quaternary sequence specifically includes the following: (1.1) Statistical symbol frequency: The input data is a discrete source containing n independent symbols. Count and output source symbols Types of symbols and probability mass function Wherein, the definition is as follows: the data comes from a frequency count The distribution of the representation, where each Symbols Frequency of occurrence; (1.2) Initialization State and State Range Setting: The initialization state is set to 0. tANS restricts the allowed state values ​​to a fixed range, which is set as the state space. ,in is a positive integer, The number of symbol types is a parameter that determines the size of the state space; at the same time, it is based on the actual probability of the source-coded symbols. ,in For symbols The probability is used to assign the state range corresponding to each symbol. It is necessary to make As close as possible To achieve probability normalization approximation; (1.3) Design symbolic extension functions: based on the symbolic extension function... Frequency of occurrence and given interval , state space Weighted by sign probability Perform dynamic partitioning and design symbolic extension functions. Establish each state To source symbols The mapping of (1≤t≤n) is achieved by discretizing the symbols through probability redistribution, ensuring that the frequency of symbol occurrence is positively correlated with the state distribution density; (1.4) Construct the state transition table: based on Construct encoding and decoding functions and The encoding and decoding lookup tables are constructed by the encoding and decoding functions. During construction, if the current state is... Within a given symbol range Inside, directly through Mapping to the next state, if it's not within the given range, calculate the number of bits needed for a right shift. Then Move right Position, that is At the same time, will be removed The bits are stored in a bitstream variable, and then... Complete the encoding to obtain the next state; (1.5) Encoding: From the initial state Begin with the input symbol sequence Perform iterative encoding; each iteration passes through the given current state. and symbols to be encoded Look up the encoding table to get the next state. and the corresponding binary output segment The process continues until the last symbol is encoded, at which point the final output is the termination status. and synthesized binary sequences ,symbol Indicates will Assemble them sequentially; (1.6) Initial state Convert to Bit binary number Termination status Convert to Bit binary number Append to binary sequence At the end, an extended sequence is formed. , as a state range marker during decoding; calculate The length is such that if it is odd, zero bits are padded at the end to generate... Grouped by double bits Will Convert to quaternary sequence Q and add padding marker bits. The presence or absence of padding is indicated by 1 for padding and 0 for no padding, ultimately generating a quaternary encoded result. As input to the dynamic rotation encoding module, this conversion method ensures decoding integrity through state metadata encapsulation, achieves lossless conversion through parity alignment, and the final output quaternary sequence contains both the original data stream and encoding / decoding metadata. and .

[0011] Furthermore, during the execution of the mapping unit, the quaternary rotation encoding rule is established based on the dynamic mapping relationship between the quaternary number system and DNA bases; after quaternary tANS encoding, a quaternary sequence is obtained, and the input sequence is defined as a set of quaternary numbers. The initial mapping rule is: Specifically: map 0 to A, 1 to C, 2 to T, and 3 to G; the dynamic rotation mapping rule is as follows: Specifically, for each subsequent character starting from the second character... The corresponding DNA bases Depends on the previously generated base and the currently entered number The design logic of this rule is to divide the leading base into two groups: the purine group (A, G) and the pyrimidine group (T, C). When the input is A or G, the outputs for 0, 1, 2, and 3 are C, T, CG, and TA respectively, indicating the current base. When the input is T or C, the outputs are A, G, AT, and GC respectively for inputs of 0, 1, 2, and 3.

[0012] Furthermore, when executing the mapping unit, compliance with the CG content constraint means that the CG content is concentrated between 49% and 51%; the homopolymer length constraint means that the maximum homopolymer length is 1; and the specific harmful sequence means that the sequence set is... .

[0013] Third aspect This invention provides a storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding.

[0014] Fourth aspect The present invention provides an electronic terminal, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: To achieve a balance between high storage density, compliance with biochemical constraints, and low decoding complexity, this invention proposes a quaternary biochemical constraint coding method based on entropy coding and rotation coding. Unlike Goldman's ternary tree structure, this scheme innovatively constructs a joint coding mechanism of quaternary entropy coding and dynamic rotation coding. In the entropy coding stage, quaternary tANS coding is introduced, utilizing a state transition mechanism to overcome the Shannon limit and achieve efficient arithmetic compression for long symbol sequences. The input data stream is mapped to a quaternary sequence of 0, 1, 2, 3 numbers, and tANS coding achieves higher compression density through continuous approximation of the probability distribution. Next, a mapping is established between the quaternary sequence and the biochemically compliant DNA sequence, ensuring that the generated DNA sequence meets CG content constraints and homopolymer length constraints, and can avoid specific sequences. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the entire encoding process provided in an embodiment of the present invention; Figure 2 This is an example of a quaternary tANS encoding and decoding process provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0018] like Figure 1As shown, the method provided by this invention firstly encodes the input data using quaternary tANS to obtain the corresponding quaternary sequence, thereby compressing and converting the original data. Entropy encoding, as an efficient data compression technique, can significantly reduce the redundancy of input data, thus achieving high storage density. Subsequently, the proposed quaternary rotation encoding algorithm is used to process the quaternary sequence, ensuring that the final generated DNA sequence not only meets the triple biochemical constraints (CG content, homopolymer length, and specific sequence avoidance) but also maintains the advantage of high storage density after data compression. Because directly using rotation encoding would sacrifice some compression ratio due to mapping rule limitations, entropy encoding first reduces data redundancy, and then rotation encoding introduces necessary redundancy to adapt to specific biochemical constraints, thus enabling the encoding method to achieve high storage density while meeting biochemical constraints. Specifically, as follows: This invention proposes a quaternary tANS encoding. First, tANS pre-generates a complete encoding and decoding lookup table, embedding the state transition process within the table. This significantly reduces runtime computational complexity and latency. Compared to other ANS variants that require multiple state expansions or bitstream processing, tANS achieves real-time data processing with higher efficiency. Second, tANS has a fixed, finite state range design and can output a bitstream. The encoding result is easily interfaced with subsequent quaternary conversion modules, maintaining the consistency of the overall data stream. Furthermore, tANS exhibits good adaptability to the probability distribution of input data. By rationally designing the encoding-decoding association table, it can accurately approximate the actual probability distribution of symbols, improving decoding stability and robustness while maintaining a high compression ratio.

[0019] Therefore, tANS's efficiency, predictability, and stability make it the ideal ANS encoding choice for high-density biochemical constraint coding. Since the tANS encoder outputs a binary sequence and the final integer state, while the subsequent dynamic rotation encoding module receives quaternary input, a binary-to-quaternary conversion mechanism is required. The encoding process is as follows:

[0020] The complete process is described below: 1. Statistical symbol frequency: The input data is a discrete source containing n independent symbols. Count and output source symbols Types of symbols and probability mass function The following definition is given: Assume the data comes from a frequency counting... The distribution of the representation, where each Symbols Frequency of occurrence; 2. Initialization State and State Range Setting: Typically, the initial state is set to 0. tANS restricts the allowed state values ​​to a fixed range, which is defined as the state space. ,in is a positive integer, The number of symbol types is a parameter that determines the size of the state space. Simultaneously, it is determined by the actual probability of the source-coded symbols. ,in For symbols The probability is used to assign the state range corresponding to each symbol. It is necessary to make As close as possible To achieve probability normalization approximation; 3. Design symbolic extension functions: based on the symbolic extension function... Frequency of occurrence and given interval , state space Weighted by sign probability Perform dynamic partitioning and design symbolic extension functions. Establish each state To source symbols The mapping of (1≤t≤n) is achieved by discretizing the symbols through probability redistribution, ensuring that the frequency of symbol occurrence is positively correlated with the state distribution density; 4. Construct the state transition table: based on Construct encoding and decoding functions and The encoding and decoding lookup tables are constructed by the encoding and decoding functions. During construction, if the current state is... Within a given symbol range Inside, directly through Mapping to the next state, if it's not within the given range, calculate the number of bits needed for a right shift. Then Move right bit (i.e.) At the same time, will be removed The bits are stored in the bitstream variable. Then, through... Complete the encoding to obtain the next state; 5. Encode: From the initial state Begin with the input symbol sequence Perform iterative encoding. Each iteration uses the current state as a given. and symbols to be encoded Look up the encoding table to get the next state. and the corresponding binary output segment The process continues until the last symbol is encoded, at which point the final output is the termination status. and synthesized binary sequences ,symbol" " indicates that will Assemble them sequentially; 6. Set the initial state Convert to Bit binary number Termination status Convert to Bit binary number Append to binary sequence At the end, an extended sequence is formed. , as a state range marker during decoding; calculate The length is such that if it is odd, zero bits are padded at the end to generate... Grouped by double bits Will Convert to quaternary sequence Q and add padding marker bits. The presence of padding (1 for padding, 0 for no padding) is used to generate the final quaternary encoded result. This serves as the input to the dynamic rotation encoding module. This conversion method ensures decoding integrity through state metadata encapsulation, achieves lossless conversion through parity alignment, and the final output quaternary sequence contains both the original data stream and encoding / decoding metadata. and .

[0021] Figure 2 This is an example of a quaternary tANS encoding and decoding process provided in an embodiment of the present invention.

[0022] The quaternary rotation encoding under biochemical constraints is as follows: The rotation encoding algorithm proposed in this invention is based on the dynamic mapping relationship between the quaternary number system and DNA bases. After quaternary tANS encoding, a quaternary sequence is obtained, and the input sequence is defined as a set of quaternary numbers. The initial mapping rule is The specific mapping rules are as follows: map 0 to A, 1 to C, 2 to T, and 3 to G; the dynamic rotation mapping rules are as follows: The specific process is as follows: For the subsequent characters starting from the second character... The corresponding DNA bases Depends on the previously generated base and the currently entered number The design logic of this rule is to divide the leading base into two groups: the purine group (A, G) and the pyrimidine group (T, C). When the input is A or G, the outputs for 0, 1, 2, and 3 are C, T, CG, and TA respectively, indicating the current base. When the input is T or C, inputs of 0, 1, 2, and 3 will output A, G, AT, and GC respectively. The output base sequence is... According to the core rules in Table 1, the specific process is explained as follows: (1) Initial mapping: first character Mapped to bases using an initial fixed mapping rule. ,Right now

[0023] (2) Encoding: Subsequent bases The mapping is based on the current quaternary character. Mapped bases with the previous one It was jointly decided to adopt the dynamic rotation mapping rule. (See Table 1 for example rules) Mapped to ,Right now:

[0024]

[0025] Dynamic rotation mapping rules The quaternary rotation encoding scheme in this invention meets three biochemical constraints: first, the maximum homopolymer length is 1; second, the CG content is concentrated between 49% and 51%; and third, it can avoid some unwanted sequences, the specific set of which is as follows: .

[0026] This invention is the first to introduce tANS encoding in the entropy coding stage, providing a higher compression ratio than Huffman coding. After encoding the source data using tANS, an integer and a bitstream are obtained, which are then converted together into a quaternary number sequence. At the biochemical constraint adaptation level, a quaternary dynamic rotation coding rule is designed to obtain a DNA sequence from the quaternary number sequence through rotation coding, ensuring that the homopolymer length is strictly limited to 1, the CG content is around 50%, and that specific sequences are avoided.

[0027] To verify the encoding results of this scheme, four text files and one audio file were selected for the encoding experiment. The text files were: Martin Luther King Jr.'s "I Have a Dream" English speech, the Chinese document of "GB / T19001-2016 Quality Management System Requirements", the English document of "The Little Prince", and the Chinese document of "Romance of the Three Kingdoms". The audio file was taken from the MUSDB18 music dataset, specifically the song "Heart Peripheral". The text files were directly encoded after being read. The audio file was first quantized, mapping the floating-point numbers [-1,1] to integers between 0 and 255 before encoding. Specific information is shown in the table below.

[0028]

[0029] Table 1 Information on files to be encoded In the quaternary tANS rotation encoding experiment, the parameters were adjusted according to the number of possible string types to achieve the optimal bitrate. The parameters selected for encoding "I Have a Dream" were... The parameters selected in GB / T 19001-2016 Quality Management System Requirements The parameters selected by The Little Prince The parameters selected in Romance of the Three Kingdoms The parameters selected in *Heart Peripheral* .

[0030] Table 2 shows the results of quaternary tANS rotational encoding. The results show that the bitrate of all four text files exceeded 2 bits / nt, the CG content of all files was around 50%, and the error was less than 0.1%. The differences in encoding results between different files are due to variations in their content. It can be seen that the highest bitrate is still for the relatively professional Chinese text "GB / T 19001-2016 Quality Management System Requirements," at 3.7 bits / nt, while the lowest is for audio files at 1.58 bits / nt. It is worth noting that the encoding time for Chinese text is longer than that for English text, especially for "Romance of the Three Kingdoms," which took 748.7839 seconds, significantly longer than other texts. This is because tANS encoding requires first constructing a finite state range based on the types of input characters, and then constructing an encoding / decoding table based on this finite state range. Chinese text involves far more types of characters than English text, and "Romance of the Three Kingdoms" involves many rare characters. Statistics show that "Romance of the Three Kingdoms" contains 3904 character types. The experimental parameter L = 655636, while "GB / T..." The "19001-2016 Quality Management System Requirements" specifies 829 character types. The experiment used parameter L = 1024. The other two English texts have fewer character types, with an initial parameter L = 256. This results in the encoding / decoding table for "Romance of the Three Kingdoms" being much larger than that for the other three texts, thus increasing the encoding time exponentially.

[0031]

[0032] Table 2. Quaternary tANS Rotation Encoding Results Corresponding to the above method, the present invention also provides a quaternary DNA storage system based on entropy coding and rotation constraint coding, comprising: a conversion unit and a mapping unit; The conversion unit is used to perform quaternary tANS entropy encoding on the input data and convert the input data into a quaternary sequence. The mapping unit is used to map the quaternary sequence through quaternary rotation coding rules so that the generated DNA sequence meets biochemical constraints, including: meeting CG content constraints and homopolymer length constraints, and being able to avoid specific harmful sequences.

[0033] In addition, the present invention provides a storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding.

[0034] In addition, the present invention provides an electronic terminal, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding.

[0035] Finally, it should be noted that the above embodiments are merely illustrative and explanatory of the present invention, and are not intended to limit the present invention to the scope of the described embodiments. Furthermore, those skilled in the art will understand that the present invention is not limited to the above embodiments, and many more variations and modifications can be made based on the teachings of the present invention, all of which fall within the scope of protection claimed by the present invention.

Claims

1. A quaternary DNA storage method based on entropy encoding and rotation constraint encoding, characterized in that, Including the following: (1) Perform quaternary tANS entropy encoding on the input data to convert the input data into a quaternary sequence; (2) The quaternary sequence is mapped through the quaternary rotation coding rule so that the generated DNA sequence meets the biochemical constraints, including: meeting the CG content constraints and homopolymer length constraints, and being able to avoid specific harmful sequences.

2. The quaternary DNA storage method based on entropy encoding and rotation constraint encoding according to claim 1, characterized in that, The step of performing quaternary tANS entropy encoding on the input data, converting the input data into a quaternary sequence, specifically includes the following: (1.1) Statistical symbol frequency: The input data is a discrete source containing n independent symbols. Count and output source symbols Types of symbols and probability mass function Wherein, the definition is as follows: the data comes from a frequency count The distribution of the representation, where each Symbols Frequency of occurrence; (1.2) Initialization State and State Range Setting: The initialization state is set to 0. tANS restricts the allowed state values ​​to a fixed range, which is set as the state space. ,in is a positive integer, The number of symbol types is a parameter that determines the size of the state space; at the same time, it is based on the actual probability of the source-coded symbols. ,in For symbols The probability is used to assign the state range corresponding to each symbol. It is necessary to make As close as possible To achieve probability normalization approximation; (1.3) Design symbolic extension functions: based on the symbolic extension function... Frequency of occurrence and given interval , state space Weighted by sign probability Perform dynamic partitioning and design symbolic extension functions. Establish each state To source symbols The mapping of (1≤t≤n) is achieved by discretizing the symbols through probability redistribution, ensuring that the frequency of symbol occurrence is positively correlated with the state distribution density; (1.4) Construct the state transition table: based on Construct encoding and decoding functions and The encoding and decoding lookup tables are constructed by the encoding and decoding functions. During construction, if the current state is... Within a given symbol range Inside, directly through Mapping to the next state, if it's not within the given range, calculate the number of bits needed for a right shift. Then Move right Position, that is At the same time, will be removed The bits are stored in a bitstream variable, and then... Complete the encoding to obtain the next state; (1.5) Encoding: From the initial state Begin with the input symbol sequence Perform iterative encoding; each iteration passes through the given current state. and symbols to be encoded Look up the encoding table to get the next state. and the corresponding binary output segment The process continues until the last symbol is encoded, at which point the final output is the termination status. and synthesized binary sequences ,symbol Indicates will Assemble them sequentially; (1.6) Initial state Convert to Bit binary number Termination status Convert to Bit binary number Append to binary sequence At the end, an extended sequence is formed. , as a state range marker during decoding; calculate The length is such that if it is odd, zero bits are padded at the end to generate... Grouped by double bits Will Convert to quaternary sequence Q and add padding marker bits. The presence or absence of padding is indicated by 1 for padding and 0 for no padding, ultimately generating a quaternary encoded result. As input to the dynamic rotation encoding module, this conversion method ensures decoding integrity through state metadata encapsulation, achieves lossless conversion through parity alignment, and the final output quaternary sequence contains both the original data stream and encoding / decoding metadata. and .

3. The quaternary DNA storage method based on entropy encoding and rotation constraint encoding according to claim 2, characterized in that, In step (2), the quaternary rotation encoding rule is established based on the dynamic mapping relationship between the quaternary number system and DNA bases; after quaternary tANS encoding, a quaternary sequence is obtained, and the input sequence is defined as a set of quaternary numbers. The initial mapping rule is: Specifically: map 0 to A, 1 to C, 2 to T, and 3 to G; the dynamic rotation mapping rule is as follows: Specifically, for each subsequent character starting from the second character... The corresponding DNA bases Depends on the previously generated base and the currently entered number The design logic of this rule is to divide the leading base into two groups: the purine group (A, G) and the pyrimidine group (T, C). When the input is A or G, the outputs for 0, 1, 2, and 3 are C, T, CG, and TA respectively, indicating the current base. When the input is T or C, the outputs are A, G, AT, and GC respectively for inputs of 0, 1, 2, and 3.

4. The quaternary DNA storage method based on entropy encoding and rotation constraint encoding according to claim 3, characterized in that, In step (2), meeting the CG content constraint means that the CG content is concentrated between 49% and 51%; the homopolymer length constraint means that the maximum homopolymer length is 1; and the specific harmful sequence means that the sequence set is .

5. A quaternary DNA storage system based on entropy encoding and rotational constraint encoding, characterized in that, It includes the following: a conversion unit and a mapping unit; The conversion unit is used to perform quaternary tANS entropy encoding on the input data and convert the input data into a quaternary sequence. The mapping unit is used to map the quaternary sequence through quaternary rotation coding rules so that the generated DNA sequence meets biochemical constraints, including: meeting CG content constraints and homopolymer length constraints, and being able to avoid specific harmful sequences.

6. A quaternary DNA storage system based on entropy encoding and rotation constraint encoding according to claim 5, characterized in that, The step of performing quaternary tANS entropy encoding on the input data, converting the input data into a quaternary sequence, specifically includes the following: (1.1) Statistical symbol frequency: The input data is a discrete source containing n independent symbols. Count and output source symbols Types of symbols and probability mass function Wherein, the definition is as follows: the data comes from a frequency count The distribution of the representation, where each Symbols Frequency of occurrence; (1.2) Initialization State and State Range Setting: The initialization state is set to 0. tANS restricts the allowed state values ​​to a fixed range, which is set as the state space. ,in is a positive integer, The number of symbol types is a parameter that determines the size of the state space; at the same time, it is based on the actual probability of the source-coded symbols. ,in For symbols The probability is used to assign the state range corresponding to each symbol. It is necessary to make As close as possible To achieve probability normalization approximation; (1.3) Design symbolic extension functions: based on the symbolic extension function... Frequency of occurrence and given interval , state space Weighted by sign probability Perform dynamic partitioning and design symbolic extension functions. Establish each state To source symbols The mapping of (1≤t≤n) is achieved by discretizing the symbols through probability redistribution, ensuring that the frequency of symbol occurrence is positively correlated with the state distribution density; (1.4) Construct the state transition table: based on Construct encoding and decoding functions and The encoding and decoding lookup tables are constructed by the encoding and decoding functions. During construction, if the current state is... Within a given symbol range Inside, directly through Mapping to the next state, if it's not within the given range, calculate the number of bits needed for a right shift. Then Move right Position, that is At the same time, will be removed The bits are stored in a bitstream variable, and then... Complete the encoding to obtain the next state; (1.5) Encoding: From the initial state Begin with the input symbol sequence Perform iterative encoding; each iteration passes through the given current state. and symbols to be encoded Look up the encoding table to get the next state. and the corresponding binary output segment The process continues until the last symbol is encoded, at which point the final output is the termination status. and synthesized binary sequences ,symbol Indicates will Assemble them sequentially; (1.6) Initial state Convert to Bit binary number Termination status Convert to Bit binary number Append to binary sequence At the end, an extended sequence is formed. , as a state range marker during decoding; calculate The length is such that if it is odd, zero bits are padded at the end to generate... Grouped by double bits Will Convert to quaternary sequence Q and add padding marker bits. The presence or absence of padding is indicated by 1 for padding and 0 for no padding, ultimately generating a quaternary encoded result. As input to the dynamic rotation encoding module, this conversion method ensures decoding integrity through state metadata encapsulation, achieves lossless conversion through parity alignment, and the final output quaternary sequence contains both the original data stream and encoding / decoding metadata. and .

7. A quaternary DNA storage system based on entropy encoding and rotation constraint encoding according to claim 6, characterized in that, When executing the mapping unit, the quaternary rotation encoding rule is established based on the dynamic mapping relationship between the quaternary number system and DNA bases; after quaternary tANS encoding, a quaternary sequence is obtained, and the input sequence is defined as a set of quaternary numbers. The initial mapping rule is: Specifically: map 0 to A, 1 to C, 2 to T, and 3 to G; the dynamic rotation mapping rule is as follows: Specifically, for each subsequent character starting from the second character... The corresponding DNA bases Depends on the previously generated base and the currently entered number The design logic of this rule is to divide the leading base into two groups: the purine group (A, G) and the pyrimidine group (T, C). When the input is A or G, the outputs for 0, 1, 2, and 3 are C, T, CG, and TA respectively, indicating the current base. When the input is T or C, the outputs are A, G, AT, and GC respectively for inputs of 0, 1, 2, and 3.

8. A quaternary DNA storage method based on entropy encoding and rotation constraint encoding according to claim 7, characterized in that, When executing the mapping unit, compliance with the CG content constraint means that the CG content is concentrated between 49% and 51%; the homopolymer length constraint means that the maximum homopolymer length is 1; and the specific harmful sequence means that the sequence set is .

9. A storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding as described in any one of claims 1-4.

10. An electronic terminal, characterized in that, The electronic terminal includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the quaternary DNA storage method based on entropy encoding and rotation constraint encoding as described in any one of claims 1-4.