DNA coding method and device with biological constraint

By employing a phased training strategy optimized with centralized encoding and biologically constrained loss function, the floating-point information sources of wireless sensor networks are directly mapped to biological base sequences. This solves the cumbersome DNA encoding problem in existing technologies, improves the efficiency of DNA data storage, and reduces the error rate.

CN120954503AActive Publication Date: 2025-11-14JIMEI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511484229.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-14
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing DNA encoding processes are cumbersome and have complex hardware implementations when processing floating-point information sources in wireless sensor networks, resulting in high error rates and low efficiency, and they cannot be naturally characterized as biological base sequences.

Method used

A progressive, phased training strategy is adopted. Through the collaborative design of centralized encoding, symmetric quantization threshold, and symmetric encoding/decoding network structure, floating-point number sequences are mapped to quaternary compressed codewords. A biological constraint loss function is introduced to optimize the GC content and distribution uniformity of the base sequence.

Benefits of technology

It simplifies the DNA encoding process, improves data storage efficiency, reduces the error rate, and ensures uniform base distribution with good distortion performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954503A_ABST
    Figure CN120954503A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a DNA coding method and device with biological constraint. In the method, a progressive staged training strategy is adopted, in the first stage, training is carried out with the purpose of minimizing distortion of an original information source and a reconstructed information source, and through collaborative design including centralized coding, a symmetric quantization threshold value and a symmetric network structure, a compressed sequence naturally meets the ideal GC content when represented as an ATCG base sequence. In the second stage, the loss function of the biological constraint is added, so that the base sequence meets the biological base arrangement characteristic while low distortion is kept. According to the method, the floating-point number sequence is naturally represented as the biological base sequence through the structural design, the base distribution is uniform, the distortion performance is good, the DNA data storage efficiency is improved, and the error rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of information processing, specifically to biological information storage and communication coding characterization technology, and more particularly to a biologically constrained DNA coding method and apparatus. Background Technology

[0002] The conventional DNA encoding process is as follows: First, the digital file undergoes binary conversion and error correction encoding; then, the binary data is converted to quaternary encoding (0→A, 1→T, 2→C, 3→G), segmented into short fragments, and then chemically synthesized and stored; during retrieval, the data is restored through PCR amplification and high-throughput sequencing. When this process is applied to process aggregated data from wireless sensor networks, the shortcomings of its "discrete" design become particularly prominent. Data collected by multiple sensor nodes in the network (such as temperature, humidity, and vibration signals) is aggregated at the central node, and its statistical characteristics can usually be accurately modeled as Gaussian mixture models or Gaussian model sources. These source data are essentially floating-point sequences.

[0003] However, existing processes force these floating-point data sources to undergo multiple independent and non-cooperative steps, such as quantization and binary encoding, before they can be finally mapped into base sequences. This process is cumbersome and requires complex hardware implementation. More seriously, the front-end digital encoding process is completely disconnected from the back-end biochemical process, and the generated sequences may not conform to the "biological constraints" necessary for DNA synthesis and sequencing, leading to increased error rates and reduced efficiency in the physical implementation stage.

[0004] Therefore, how to naturally characterize floating-point sequences into biological base sequences, ensuring uniform base distribution and good distortion performance, thereby improving the efficiency of DNA data storage and reducing the error rate, is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] This invention provides a biologically constrained DNA encoding and apparatus that can improve the efficiency of DNA data storage and reduce the error rate.

[0006] The first aspect of this invention provides a biologically constrained DNA encoding method, comprising: Acquire monitoring data collected by multiple sensor nodes in a wireless sensor network, and model the monitoring data as a sequence of raw floating-point numbers that follows a normal distribution; The original floating-point sequence is mapped to quaternary compressed codewords through a lossy source coding module. The lossy source coding module adopts a collaborative design of centralized coding, symmetric quantization threshold and symmetric encoding and decoding network structure, so that the quaternary compressed codewords have the basis to be naturally converted into ATCG base sequences. The quaternary compressed codeword is reconstructed into a reconstructed floating-point sequence through the lossy source decoding module; The first stage of training determines the distortion loss based on the mean square error between the minimized original floating-point sequence and the reconstructed floating-point sequence, and iteratively optimizes the lossy source coding module and the lossy source decoding module until the first preset training round is reached, so that the ATCG base sequence corresponding to the quaternary compressed codeword naturally meets the preset ratio of GC content requirements. The second stage of training introduces a biological constraint loss function. A composite loss function is obtained by minimizing the weighted sum of the distortion loss and the biological constraint loss function. The lossy source coding module and the lossy source decoding module are iteratively optimized until the second preset training round is reached. The biological characteristics of the quaternary compressed codeword are optimized through the biological constraint loss function to obtain the optimized quaternary compressed codeword. The optimized quaternary compressed codewords are converted into ATCG base sequences for DNA synthesis and storage.

[0007] Optionally, centralized encoding converts the quaternary codeword {0,1,2,3} into a set of values ​​{-1.5,-0.5,0.5,1.5} that are symmetric about the zero point, where the AT base set is mapped to the negative number field and the CG base set is mapped to the positive number field, so that the mean of the quaternary compressed codeword tends to 0, which naturally satisfies the requirement that the GC content is a preset ratio.

[0008] Optionally, the symmetric quantization threshold adopts a threshold configuration symmetric about zero. The original floating-point sequence is mapped to initial log-likelihood ratio information through the encoding network of the lossy source coding module. The initial log-likelihood ratio information is iteratively updated through the node connection relationship of the preset original model low-density parity check code in the lossy source coding module to obtain the updated log-likelihood ratio information. The value range of the updated log-likelihood ratio information is divided into four intervals symmetric about zero, each interval corresponding to a quaternary codeword, so that the probability of continuous values ​​falling into the positive interval and the negative interval is equal.

[0009] Optionally, the symmetric codec network structure adopts a zero-mean symmetric distribution weight initialization method, combined with a mean square error loss function, so that the output quaternary compressed codeword distribution naturally tends to zero mean.

[0010] Optionally, the biological constraint loss function includes: The homopolymer penalty value is calculated using the adjacent difference of the numerical difference between adjacent positions in the centered quaternary sequence; The biological constraint loss function is derived by averaging the homopolymer penalty values ​​at all adjacent locations across the batch and sequence dimensions.

[0011] Optionally, the optimized quaternary compressed codewords are converted into ATCG base sequences for DNA synthesis and storage, including: The optimized quaternary compressed codewords are converted into ATCG base sequences using a preset mapping rule, where 0 corresponds to A, 1 corresponds to T, 2 corresponds to C, and 3 corresponds to G.

[0012] A second aspect of the present invention provides a biologically constrained DNA coding device, comprising: The data acquisition module is used to acquire monitoring data collected by multiple sensor nodes in the wireless sensor network and model the monitoring data as a sequence of raw floating-point numbers that follows a normal distribution. The lossy source coding module is used to map the original floating-point number sequence into quaternary compressed codewords. The lossy source coding module adopts a collaborative design of centralized coding, symmetric quantization threshold and symmetric encoding and decoding network structure, so that the quaternary compressed codewords have the basis to be naturally converted into ATCG base sequences. The lossy source decoding module is used to reconstruct a sequence of reconstructed floating-point numbers from quaternary compressed codewords. The training module is used for the first stage of training. The reconstruction loss target is to minimize the mean square error between the original floating-point sequence and the reconstructed floating-point sequence. The lossy source coding module and the lossy source decoding module are iteratively optimized until the first preset training round is reached, so that the ATCG base sequence corresponding to the quaternary compressed codeword naturally meets the preset ratio of GC content requirements. The second stage of training introduces a biological constraint loss function, with the goal of minimizing the weighted sum of the reconstruction loss and the biological constraint loss function as the composite loss objective. The lossy source coding module and the lossy source decoding module are iteratively optimized until the second preset training round is reached. The biological characteristics of the quaternary compressed codeword are optimized through the biological constraint loss function to obtain the optimized quaternary compressed codeword. The conversion module is used to convert the optimized quaternary compressed codewords into ATCG base sequences for DNA synthesis and storage.

[0013] A third aspect of the present invention provides a biologically constrained DNA coding device, comprising: One or more processors; A memory on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the biologically constrained DNA encoding method as described in any of the preceding claims.

[0014] A fourth aspect of the present invention provides a computer storage medium for storing a program, which, when executed, is used to implement the method of biologically constrained DNA encoding as described in any of the preceding claims.

[0015] This invention discloses a biologically constrained DNA encoding method and apparatus. The method employs a progressive, phased training strategy. The first phase aims to minimize the distortion between the original and reconstructed information sources. Through a collaborative design incorporating centered encoding, symmetric quantization thresholds, and a symmetric network structure, the compressed sequence naturally satisfies the ideal GC content when represented as an ATCG base sequence. In the second phase, a biologically constrained loss function is added to maintain low distortion while ensuring the base sequence meets biological base arrangement characteristics. This invention, through structured design, naturally represents floating-point sequences as biological base sequences, exhibiting uniform base distribution and good distortion performance, thereby improving the efficiency of DNA data storage and reducing the error rate. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic flowchart of a biologically constrained DNA encoding method provided in an embodiment of the present invention; Figure 2 A schematic flowchart of another biologically constrained DNA encoding method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a biologically constrained DNA coding device provided in an embodiment of the present invention; Figure 4 A schematic diagram of the structure of a biologically constrained DNA coding device provided in an embodiment of the present invention; Figure 5 This is a system performance comparison chart provided for an embodiment of the present invention. Detailed Implementation

[0018] This invention provides a biologically constrained DNA encoding and device that can solve the technical problem of long-term and reliable archiving of massive aggregated monitoring data in wireless sensor networks. It can efficiently represent floating-point monitoring data aggregated by the network central node, which can be modeled as Gaussian distribution, as biological bases, and realize the characterization of floating-point DNA.

[0019] This invention proposes a biologically constrained DNA encoding method that directly maps floating-point data from sensor networks into quaternary base sequences that meet biological arrangement requirements, fundamentally simplifying and improving the DNA encoding process. Furthermore, this invention uses structured design to naturally characterize floating-point sequences into biological base sequences, ensuring uniform base distribution and good distortion performance, thus providing a more efficient encoding and storage solution for DNA archiving of large-scale IoT data.

[0020] See Figure 1 This figure is a schematic flowchart of a biologically constrained DNA encoding method provided by an embodiment of the present invention. The biologically constrained DNA encoding method provided by this embodiment of the present invention can be implemented, for example, through the following steps S101-106.

[0021] Step S101: Acquire monitoring data collected by multiple sensor nodes in the wireless sensor network, and model the monitoring data as a sequence of original floating-point numbers that follows a normal distribution.

[0022] In this embodiment of the invention, monitoring data collected by multiple sensor nodes in a wireless sensor network are aggregated, and the aggregated detection data is modeled as a sequence of original floating-point numbers that follows or approximately follows a standard normal distribution. Input lossy source coding module .

[0023] Step S102: Map the original floating-point sequence into quaternary compressed codewords using the lossy source coding module.

[0024] In this embodiment of the invention, the original floating-point sequence n is compressed and represented as a k-dimensional quaternary sequence u, and the source representation satisfies the set {0, 1, 2, 3}. A collaborative design employing centralized coding, symmetric quantization thresholding, and a symmetric encoding / decoding network structure enables the quaternary compressed codewords to naturally convert into ATCG base sequences.

[0025] In this embodiment of the invention, the centralized encoding includes converting the quaternary codeword {0,1,2,3} into a zero-symmetric set of values ​​{-1.5,-0.5,0.5,1.5}, where the AT base set is mapped to the negative domain and the CG base set is mapped to the positive domain, making the mean of the quaternary compressed codeword tend to 0, naturally satisfying the requirement that the GC content is a preset ratio. The symmetric quantization threshold adopts a zero-symmetric threshold configuration, mapping the original floating-point sequence to initial log-likelihood ratio information through a fully connected neural network in the lossy source coding module. This initial log-likelihood ratio information is iteratively updated through the node connection relationship of the preset original model low-density parity-check code in the lossy source coding module to obtain the updated log-likelihood ratio information. The value range of the updated log-likelihood ratio information is divided into four zero-symmetric intervals, each interval corresponding to a quaternary codeword, making the probability of continuous values ​​falling into the positive and negative intervals equal. The symmetric codec network structure uses a zero-mean symmetric weight initialization method, combined with a mean square error loss function, so that the output quaternary compressed codeword distribution naturally tends toward zero mean.

[0026] Specifically, through a collaborative design incorporating centralized encoding, symmetric quantization thresholding, and a symmetric encoding / decoding network structure, the GC content of the output sequence can be naturally stabilized within a preset ideal range without imposing additional biological constraints. Centralized encoding converts the quaternary codewords {0, 1, 2, 3} into a zero-symmetric set of values ​​{-1.5, -0.5, 0.5, 1.5}. Following the quaternary-to-base mapping rule {0→A, 1→T, 2→C, 3→G}, the AT base set {0, 1} is mapped to the negative domain, and the CG base set {2, 3} is mapped to the positive domain. The symmetric quantization threshold uses a zero-symmetric threshold configuration (e.g., T1=-1.0, T2=0.0, T3=1.0), ensuring that the probability of a continuous value falling into the positive interval is equal to its probability of falling into the corresponding negative interval. Symmetric encoder-decoder network structures, by employing zero-mean symmetric weight initialization methods (such as Xavier initialization, He initialization, or other initialization strategies suitable for the chosen network architecture) combined with the MSE loss function, naturally tend towards a zero-mean output distribution during training. The symmetric network structure ensures that the encoder tends to output mathematically zero-symmetric continuous values; symmetric quantization thresholds map these symmetric continuous values ​​to discrete quaternary codewords; and centralized encoding directly correlates this mathematical symmetry with the biological characteristics of the bases (GC / AT). This guarantees that, without any specific GC constraints, the GC content of the quaternary sequence u is within a preset ideal range (e.g., 45%-55%).

[0027] Step S103: Reconstruct the quaternary compressed codeword into a reconstructed floating-point sequence using the lossy source decoding module.

[0028] Step S104: Determine the distortion loss based on the mean square error of the minimized original floating-point sequence and the reconstructed floating-point sequence, and iteratively optimize the lossy source coding module and the lossy source decoding module.

[0029] In this embodiment of the invention, a progressive, phased training strategy is employed for end-to-end joint optimization. The training process is divided into two phases. The first phase involves building the basic capabilities using reconstruction distortion loss, with the goal of enabling the network to possess efficient compression and reconstruction capabilities, and obtaining an ideal GC content base through the collaborative mechanism in step S102.

[0030] In one implementation of this invention, the homopolymer penalty value is calculated using an exponential decay function based on the adjacent difference of the numerical differences between adjacent positions in the centered quaternary sequence; and the biological constraint loss function is obtained based on the average value of the homopolymer penalty values ​​of all adjacent positions in the batch and sequence dimensions.

[0031] Specifically, a centralized encoding is implemented for a k-dimensional quaternary sequence u, i.e. Based on centralized sequence Define the biological constraint loss used in the second training phase. For a centralized compressed sequence of dimension (B, L) Calculate the absolute difference between all adjacent positions, expressed as: ; Where B is the batch size. For sequence length, This represents the centered value at the j-th position in the i-th sample sequence. It is based on the adjacent difference. The homopolymer penalty is calculated using an exponential decay function, expressed as: ; The final calculation of the biological constraint loss function is the average of the penalty values ​​at all neighboring positions across the batch and sequence dimensions, expressed as: ; in, The attenuation coefficient is not specifically limited in this embodiment. For example, it can be 5. The specific value can be set according to actual needs.

[0032] Subsequently, the centralized compressed sequence is input into the lossy source decoding module. Obtain the reconstructed information source, and define system distortion as the reconstructed information source. Compared with the original source The mean squared loss is the sole objective function for the first stage of training, and the distortion loss is the result. Represented as: .

[0033] Step S105: Obtain the composite loss function by minimizing the weighted sum of the distortion loss and the biological constraint loss function, and iteratively optimize the lossy source coding module and the lossy source decoding module.

[0034] In this embodiment of the invention, the second stage of training is a bio-constrained fusion. Benefiting from the favorable initial state of the first stage, this stage can focus on homopolymer suppression and base distribution uniformity without the need for an additional GC content control term. A composite loss function is introduced for training, focusing on optimizing homopolymer suppression and base distribution uniformity while maintaining low distortion. The bio-constrained loss function includes calculating the homopolymer penalty value based on the neighbor difference of the numerical difference between adjacent positions in the centered quaternary sequence using an exponential decay function; the bio-constrained loss function is obtained by averaging the homopolymer penalty values ​​of all adjacent positions across the batch and sequence dimensions.

[0035] Specifically, in the second training phase, while maintaining a low distortion loss Simultaneously, utilizing biological constraint loss based on adjacent nucleotide difference penalty The resulting gradient-optimized deep encoder-decoder constructs a composite loss function defined as follows: = +

[0036] in, The weighting factor for biological constraint loss is not specifically limited in this embodiment. For example, it can be 5. The specific value can be set according to actual needs.

[0037] Step S106: Convert the optimized quaternary compressed codewords into ATCG base sequences for DNA synthesis and storage.

[0038] In this embodiment of the invention, the optimized quaternary compressed codeword is converted into an ATCG base sequence using a preset mapping rule, where 0 corresponds to A, 1 corresponds to T, 2 corresponds to C, and 3 corresponds to G.

[0039] Specifically, after completing two phases of progressive training, the lossy source coding module... It can acquire biological target sequences ubio with local diversity and uniform global base distribution. It minimizes compression distortion during the reconstruction of lossy source decoded segments, while satisfying the GC content within a preset range. It has the characteristics of low long homopolymer content and uniform distribution of the four bases.

[0040] To further verify the coding efficiency of this invention from a theoretical perspective, Figure 5 The system performance was further verified within the framework of RD theory. This theoretical limit represents the optimal performance boundary that any compression algorithm can achieve on this type of source. Figure 5 As shown, Figure 5 This invention provides a system performance comparison chart. To accurately simulate the intensity signal acquired by a LiDAR sensor and processed by normalization, the experiment used a signal with corresponding statistical characteristics (mean 0, variance 0). The source model of the proposed invention was described. Experimental conditions were set as follows: using the P-LDPC reference code AR3A, and setting the compression ratio R to three times the code rate RB. Results show that the quaternary system is closer to the theoretical limit, indicating that the quaternary architecture of this invention can more effectively approximate the theoretical optimal performance. This strongly suggests that the quaternary architecture of this invention has higher coding efficiency and can approach the theoretical optimal compression performance with less performance loss, thereby more fully exploring and utilizing the effective information in LiDAR signals.

[0041] Now combined Figure 2 To explain, Figure 2 This is a flowchart illustrating a biologically constrained DNA encoding method provided in an embodiment of the present invention. It clearly demonstrates the end-to-end process of mapping floating-point monitoring data aggregated from a wireless sensor network into DNA base sequences that meet biological characteristics. The process employs a progressive, staged training strategy, first building basic compression and reconstruction capabilities, then integrating biological constraints to optimize sequence characteristics, and finally outputting a target sequence that can be directly used for DNA synthesis. Figure 1 The process uses floating-point monitoring data (input source) aggregated by a wireless sensor network. Starting from ), quaternary compressed codewords are generated through lossy source coding. Then, it is reconstructed into a floating-point sequence through lossy source decoding. Subsequently, the parameters of the encoding / decoding modules are optimized through two-stage training, ultimately outputting the biologically constrained target sequence ubio (ATCG base sequence). The core of the process is the end-to-end collaborative optimization of "floating-point source → quaternary codeword → biologically constrained sequence", which solves the cumbersome problem of the traditional DNA encoding "separate design" (binary conversion → error correction → base mapping).

[0042] The process starts from the input information source. This refers to monitoring data (such as temperature, humidity, and vibration signals) collected by multiple sensor nodes in a wireless sensor network. After being aggregated at the central node, this data is modeled as a floating-point sequence that follows a standard normal distribution.

[0043] Input source Through the lossy source coding module Mapped to quaternary compressed codewords ,make The corresponding GC content naturally stabilizes within the ideal range: Quaternary codeword Through the lossy source decoding module Reconstructed into a sequence of floating-point numbers (Reconstructing the information source).

[0044] From the original source With reconstructing the source of information The mean square error is used as the distortion loss to measure the compression effect of encoding / decoding.

[0045] Determine if the required number of training rounds has been reached. If the preset number of training rounds for the first stage has not been reached, the information source will be reconstructed. Feedback to the lossy source coding module Continue to iterate and optimize the lossy source coding module. and lossy source decoding module The parameters are set; if the preset number of training rounds for the first stage is reached, the second stage of training will begin.

[0046] Keep the input source This remains unchanged to ensure training continuity.

[0047] Through the lossy source coding module Generate quaternary codewords Through the lossy source decoding module Refactoring (Consistent with the first phase).

[0048] Maintaining low distortion loss Simultaneously, utilizing biological constraint loss based on adjacent nucleotide difference penalty The resulting gradient optimizes the deep encoder-decoder, and constructs a composite loss function. Defined as: = + , This is the weighting factor for biological constraint loss.

[0049] Determine if the required number of training rounds has been reached. If not, continue iteratively optimizing the lossy source coding module. and lossy source decoding module The parameters are set; if the preset number of training rounds for the second stage is reached, the training ends.

[0050] After the second phase of training, the lossy source coding module Generated quaternary codewords Optimized into biological target sequence ubio (ATCG base sequence).

[0051] Based on the methods provided in the above embodiments, this invention also provides a biologically constrained DNA encoding device, see [link to previous embodiment]. Figure 3 The figure is a schematic diagram of a biologically constrained DNA coding device provided in an embodiment of the present invention.

[0052] The DNA encoding device 300 with biological constraints provided in this embodiment of the invention includes: a data acquisition module 301, a lossy source encoding module 302, a lossy source decoding module 303, a training module 304, and a conversion module 305.

[0053] The data acquisition module 301 is used to acquire monitoring data collected by multiple sensor nodes in a wireless sensor network and model the monitoring data as a sequence of raw floating-point numbers that follows a normal distribution. The lossy source coding module 302 is used to map the original floating-point number sequence into quaternary compressed codewords. The lossy source coding module adopts a collaborative design of centralized coding, symmetric quantization threshold and symmetric encoding and decoding network structure, so that the quaternary compressed codewords have the basis to be naturally converted into ATCG base sequences. Lossy source decoding module 303 is used to reconstruct a quaternary compressed codeword into a reconstructed floating-point sequence; Training module 304 is used for the first stage of training. The goal of minimizing the mean square error between the original floating-point sequence and the reconstructed floating-point sequence is to reconstruct the loss. It iteratively optimizes the lossy source coding module and the lossy source decoding module until the first preset training round is reached, so that the ATCG base sequence corresponding to the quaternary compressed codeword naturally meets the preset ratio of GC content requirements. The second stage of training introduces a biological constraint loss function, with the goal of minimizing the weighted sum of the reconstruction loss and the biological constraint loss function as the composite loss objective. The lossy source coding module and the lossy source decoding module are iteratively optimized until the second preset training round is reached. The biological characteristics of the quaternary compressed codeword are optimized through the biological constraint loss function to obtain the optimized quaternary compressed codeword. The conversion module 305 is used to convert the optimized quaternary compressed codewords into ATCG base sequences for DNA synthesis and storage.

[0054] In one possible implementation, the lossy source coding module 302 is specifically used for: Centralized encoding converts the quaternary codeword {0,1,2,3} into a set of values ​​symmetric about zero: {-1.5,-0.5,0.5,1.5}. The AT base set is mapped to the negative number field, and the CG base set is mapped to the positive number field, so that the mean of the quaternary compressed codeword tends to 0, which naturally satisfies the requirement that the GC content is at a preset ratio.

[0055] In one possible implementation, the lossy source coding module 302 is specifically used for: The symmetric quantization threshold adopts a threshold configuration symmetric about the zero point. The original floating-point sequence is mapped to the initial log-likelihood ratio information through the coding network of the lossy source coding module. The initial log-likelihood ratio information is iteratively updated through the node connection relationship of the low-density parity check code of the original model graph in the lossy source coding module to obtain the updated log-likelihood ratio information. The value range of the updated log-likelihood ratio information is divided into four intervals symmetric about the zero point. Each interval corresponds to a quaternary codeword, so that the probability of continuous values ​​falling into the positive interval and the negative interval is equal.

[0056] In one possible implementation, the lossy source coding module 302 is specifically used for: The symmetric codec network structure uses a zero-mean symmetric weight initialization method, combined with a mean square error loss function, so that the output quaternary compressed codeword distribution naturally tends toward zero mean.

[0057] In one possible implementation, training module 304 is specifically used for: The homopolymer penalty value is calculated using the adjacent difference of the numerical difference between adjacent positions in the centered quaternary sequence; The biological constraint loss function is derived by averaging the homopolymer penalty values ​​at all adjacent locations across the batch and sequence dimensions.

[0058] In one possible implementation, the conversion module 305 is specifically used for: The optimized quaternary compressed codewords are converted into ATCG base sequences using a preset mapping rule, where 0 corresponds to A, 1 corresponds to T, 2 corresponds to C, and 3 corresponds to G.

[0059] Since the biologically constrained DNA encoding device 300 is a device corresponding to the biologically constrained DNA encoding method provided in the above method embodiments, the specific implementation of each module of the biologically constrained DNA encoding device 300 is based on the same concept as in the above method embodiments. Therefore, for the specific implementation of each module of the biologically constrained DNA encoding device 300, please refer to the description of the biologically constrained DNA encoding method in the above method embodiments, and it will not be repeated here.

[0060] This invention also provides a biologically constrained DNA encoding device, the device comprising: a processor and a memory; The memory is used to store instructions; The processor is configured to execute the instructions in the memory to perform the biologically constrained DNA encoding method mentioned in the above embodiments.

[0061] It should be noted that the hardware structure of the biologically constrained DNA encoding device provided in the embodiments of the present invention can be as follows: Figure 4 The structure shown, Figure 4 This is a schematic diagram of the structure of a biologically constrained DNA coding device provided in an embodiment of the present invention.

[0062] Please see Figure 4 As shown, the biologically constrained DNA encoding device 400 includes a processor 410, a communication interface 420, and a memory 430. The number of processors 410 in the biologically constrained DNA encoding device 400 can be one or more. Figure 4 Taking a processor as an example, in this embodiment of the invention, the processor 410, communication interface 420, and memory 430 can be connected via a bus system or other means. Figure 4 Taking the connection between China and Israel via the bus system 440 as an example.

[0063] Processor 410 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. Processor 410 may further include hardware chips. These hardware chips may be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0064] The memory 430 may include volatile memory, such as random-access memory (RAM); the memory 430 may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 430 may also include a combination of the above types of memory.

[0065] Optionally, the memory 430 stores an operating system and programs, executable modules, or data structures, or subsets thereof, or extended sets thereof. The programs may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business processes and handling hardware-based tasks. The processor 410 can read the programs from the memory 430 to implement the biologically constrained DNA encoding method provided in this embodiment of the invention.

[0066] The bus system 440 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus system 440 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0067] This invention also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform the methods mentioned in the above embodiments.

[0068] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods mentioned in the above embodiments.

[0069] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A biologically constrained DNA encoding method, characterized in that, include: Acquire monitoring data collected by multiple sensor nodes in a wireless sensor network, and model the monitoring data as a sequence of raw floating-point numbers that follows a normal distribution; The original floating-point sequence is mapped to quaternary compressed codewords through a lossy source coding module. The lossy source coding module adopts a collaborative design of centralized coding, symmetric quantization threshold and symmetric encoding and decoding network structure, so that the quaternary compressed codewords have the basis for natural conversion into ATCG base sequences. The quaternary compressed codeword is reconstructed into a reconstructed floating-point sequence using a lossy source decoding module; The first stage of training determines the distortion loss based on the minimized mean square error between the original floating-point sequence and the reconstructed floating-point sequence, and iteratively optimizes the lossy source coding module and the lossy source decoding module until the first preset training round is reached, so that the ATGG base sequence corresponding to the quaternary compressed codeword naturally meets the preset ratio of GC content requirements. The second stage of training introduces a biological constraint loss function. A composite loss function is obtained by minimizing the weighted sum of the distortion loss and the biological constraint loss function. The lossy source coding module and the lossy source decoding module are iteratively optimized until the second preset training round is reached. The biological characteristics of the quaternary compressed codeword are optimized through the biological constraint loss function to obtain the optimized quaternary compressed codeword. The optimized quaternary compressed codewords are converted into ATCG base sequences for DNA synthesis and storage.

2. The biologically constrained DNA encoding method according to claim 1, characterized in that, The centralized encoding converts the quaternary codeword {0,1,2,3} into a set of values ​​symmetric about zero: {-1.5,-0.5,0.5,1.5}. The AT base set is mapped to the negative number field, and the CG base set is mapped to the positive number field, so that the mean of the quaternary compressed codeword tends to 0, which naturally satisfies the requirement that the GC content is a preset ratio.

3. The biologically constrained DNA encoding method according to claim 1, characterized in that, The symmetric quantization threshold adopts a threshold configuration symmetric about zero. The original floating-point sequence is mapped to initial log-likelihood ratio information through the encoding network of the lossy source coding module. The initial log-likelihood ratio information is iteratively updated through the node connection relationship of the preset original model low-density parity check code in the lossy source coding module to obtain the updated log-likelihood ratio information. The value range of the updated log-likelihood ratio information is divided into four intervals symmetric about zero, each interval corresponding to a quaternary codeword, so that the probability of continuous values ​​falling into the positive interval and the negative interval is equal.

4. The biologically constrained DNA encoding method according to claim 1, characterized in that, The symmetric encoding / decoding network structure adopts a zero-mean symmetric distribution weight initialization method, combined with a mean square error loss function, so that the output quaternary compressed codeword distribution naturally tends to zero mean.

5. The biologically constrained DNA encoding method according to claim 1, characterized in that, The biological constraint loss function includes: The homopolymer penalty value is calculated using the adjacent difference of the numerical difference between adjacent positions in the centered quaternary sequence; The biological constraint loss function is obtained by averaging the homopolymer penalty values ​​of all adjacent positions across the batch and sequence dimensions.

6. The biologically constrained DNA encoding method according to claim 1, characterized in that, The step of converting the optimized quaternary compressed codewords into ATCG base sequences for DNA synthesis and storage includes: The optimized quaternary compressed codeword is converted into an ATCG base sequence using a preset mapping rule, where 0 corresponds to A, 1 corresponds to T, 2 corresponds to C, and 3 corresponds to G.

7. A biologically constrained DNA coding device, characterized in that, include: The data acquisition module is used to acquire monitoring data collected by multiple sensor nodes in a wireless sensor network and model the monitoring data as a sequence of raw floating-point numbers that follows a normal distribution. The lossy source coding module is used to map the original floating-point number sequence into quaternary compressed codewords. The lossy source coding module adopts a collaborative design of centralized coding, symmetric quantization threshold and symmetric encoding and decoding network structure, so that the quaternary compressed codewords have the basis for natural conversion into ATCG base sequences. The lossy source decoding module is used to reconstruct the quaternary compressed codeword into a reconstructed floating-point number sequence; The training module is used for the first stage of training to minimize the mean square error between the original floating-point sequence and the reconstructed floating-point sequence as the reconstruction loss target, and to iteratively optimize the lossy source coding module and the lossy source decoding module until the first preset training round is reached, so that the ATGG base sequence corresponding to the quaternary compressed codeword naturally meets the preset ratio of GC content requirements. The second stage of training introduces a biological constraint loss function, with the goal of minimizing the weighted sum of the reconstruction loss and the biological constraint loss function as the composite loss objective. The lossy source coding module and the lossy source decoding module are iteratively optimized until the second preset training round is reached. The biological characteristics of the quaternary compressed codeword are optimized through the biological constraint loss function to obtain the optimized quaternary compressed codeword. The conversion module is used to convert the optimized quaternary compressed codewords into ATCG base sequences for DNA synthesis and storage.

8. The biologically constrained DNA coding device according to claim 7, characterized in that, The conversion module is specifically used for: The optimized quaternary compressed codeword is converted into an ATCG base sequence using a preset mapping rule, where 0 corresponds to A, 1 corresponds to T, 2 corresponds to C, and 3 corresponds to G.

9. A biologically constrained DNA coding device, characterized in that, The device includes: a processor and a memory; The memory is used to store instructions; The processor is configured to execute the instructions in the memory to perform the method according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, Including instructions that, when run on a computer, cause the computer to perform the method described in any one of claims 1-6 above.

Citation Information

Patent Citations

  • Method and device for storing and reading DNA of information

    CN117789788A

  • High-bit-rate DNA coding method and system meeting GC local balance and run constraint

    CN118675625A

  • DNA image transmission method and device based on deep joint source channel coding

    CN120529090A

  • Semantic intelligent enhanced DNA storage method for Internet of Things

    CN120673849A