Semantic intelligent enhanced DNA storage method for Internet of Things

By extracting semantic information from the original image and encoding it into DNA base sequences, and combining it with a multi-read screening mechanism, the problems of excessive redundant information and high bit error rate in existing DNA image storage methods are solved, and efficient and robust DNA image storage is achieved.

CN120673849APending Publication Date: 2025-09-19YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510786094.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing DNA image storage methods, the original image contains too much redundant information, resulting in low storage efficiency; the bit error rate during DNA synthesis and sequencing is high, affecting the quality of image restoration, and there is a lack of semantic intelligence and molecular channel robustness design.

Method used

Using a semantic intelligence-enhanced DNA storage method, semantic information is extracted from the original image to generate a semantically aware image, which is then encoded as a DNA base sequence. Combined with a multi-read screening mechanism, high-reliability sequences are selected from the sequenced copies for decoding and image reconstruction.

Benefits of technology

It effectively improves the compression efficiency and robustness of DNA image storage, enhances fault tolerance by eliminating redundant background information and selecting high-quality sequences, and is suitable for efficient semantic storage in IoT scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673849A_ABST
    Figure CN120673849A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic intelligent enhanced DNA storage method for the Internet of Things, and belongs to the crossing field of artificial intelligence and biological information storage, the compression efficiency and robustness of DNA image storage are effectively improved through introduced semantic extraction and a multi-read screening mechanism of design, key information areas in images are extracted through semantic extraction, and the compression efficiency and robustness of DNA image storage are improved. The method has the advantages that redundant backgrounds are eliminated, coding lengths are reduced, high-quality sequences are selected from sequencing copies by the aid of scoring and sequence analysis through a multi-read screening mechanism, fault tolerance is enhanced, the method is applicable to efficient semantic storage in scenes of the internet of things, the problem of low storage efficiency due to excessive redundant information in original images in existing DNA image storage methods is solved, and the method is applicable to high-efficiency semantic storage in scenes of the internet of things. And a high bit error rate exists in the DNA synthesis and sequencing process, so that the image recovery quality is influenced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection of artificial intelligence and biological information storage, and specifically relates to a semantic intelligence-enhanced DNA storage method for the Internet of Things. Background Art

[0002] With the explosive growth of digital image, video, and sensor data, traditional electronic storage media (such as hard drives and flash memory) are increasingly unable to meet the demands of long-term archiving and high-density storage in terms of capacity, lifespan, and stability. DNA, as a naturally high-density, long-lived, and biodegradable information carrier, is becoming a key research direction for next-generation non-volatile data storage media, due to its theoretical potential to achieve petabyte-level storage densities per gram and storage stability exceeding millennia.

[0003] In recent years, researchers have conducted extensive research in DNA image storage, primarily using methods such as convolutional coding, redundant error correction, and base mapping to convert images into DNA sequences for synthesis and sequencing. However, traditional DNA image storage methods commonly suffer from the following problems: First, at the image data level, the original image contains a large amount of redundant background information that is irrelevant to the task, which increases storage costs and hinders the extraction of core semantics in downstream tasks. Second, at the molecular channel level, DNA synthesis, PCR amplification, and sequencing processes are prone to introducing multiple types of errors such as insertions, deletions, and substitutions, and redundant design makes it difficult to balance efficiency and stability. Finally, existing methods often focus on data restoration accuracy and lack content selection and optimization from the perspective of "information value," making it difficult to meet the dual requirements of intelligent compression and robust perception in semantically driven IoT scenarios.

[0004] Furthermore, sequencing generates a large number of redundant copies (multi-reads) in the DNA channel output, leading to inconsistencies and noise contamination between different copies. Traditional methods often use random selection of single sequences or voting correction for recovery. These methods fail to fully utilize the potential redundancy of multi-read data and cannot effectively identify the most valuable fragments, resulting in unstable recovery accuracy and susceptibility to contaminated sequences.

[0005] In recent years, with the advancement of artificial intelligence (AI), particularly in semantic understanding capabilities of large models, image semantic extraction and attention mechanisms have been widely applied to content understanding and representation compression. Introducing semantically driven information representation into the DNA encoding process is expected to achieve "content-aware" selective compression, thereby reducing redundancy and improving DNA storage efficiency.

[0006] In summary, existing DNA image storage methods have yet to effectively integrate semantic intelligence with robust molecular channel design, lacking integrated semantic compression, intelligent filtering, and robust decoding mechanisms. Therefore, a new DNA image storage method combining semantic representation, large-scale model perception, multi-read screening, and deep learning end-to-end optimization is urgently needed to meet the practical needs of intelligent, efficient, and stable storage in the future Internet of Things environment. Summary of the Invention

[0007] In order to solve the problems raised in the above background technology, the present invention provides a semantic intelligent enhanced DNA storage method for the Internet of Things, so as to solve the problem of excessive redundant information in the original image in the existing DNA image storage method, resulting in low storage efficiency, and the problem of high bit error rate in the DNA synthesis and sequencing process, which affects the image recovery quality.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A semantic intelligence-enhanced DNA storage method for the Internet of Things, comprising the following steps:

[0010] S1: Extract semantic information from the original image to obtain a semantically perceived image;

[0011] S2: Encode semantically-aware images into DNA base sequences through feature extraction and base mapping;

[0012] S3: Channel simulation is performed on the DNA base sequence. The channel simulation includes DNA synthesis, PCR amplification, and DNA sequencing. During the channel simulation, random errors such as insertions and deletions are added to the DNA base sequence. Multiple copies of the noisy base sequence are generated and stored in parallel, completing the storage process of the original image.

[0013] S4: Set up a multi-read screening mechanism, and use scoring and sequence analysis to select multiple reliable base sequences from the stored base sequence copies to form a sequence set;

[0014] S5: Decode and map the sequence set into a reconstructed image to complete the reconstruction process of the original image.

[0015] Preferably, S1 is specifically:

[0016] The open source large model SAM, which includes an image encoder, a prompt encoder, and a mask decoder, is used to segment the original image and generate a set of semantic masks:

[0017] ;

[0018] Among them, M iis the mask of the i-th target, n is the number of targets in the image, x is the original image, and H×W represents the length and width of the mask;

[0019] Each mask M i Multiply the original image pixel by pixel to get the corresponding semantic segment s i :

[0020] ;

[0021] The attention fusion network ASI is introduced to identify and fuse the most informative semantic target region index set ℒ from the semantic fragments to generate the final semantic perception image x scm :

[0022] ;

[0023] Among them, s i Represents the i-th semantic segment.

[0024] Preferably, S2 is specifically: a DNA image coding network f using deep source-channel joint coding θ The semantically aware image x scm Mapped to a DNA base sequence of length k , that is: z=f θ( x scm) , where A is adenine, T is thymine, G is guanine, C is cytosine, z k is the kth base in the DNA base sequence.

[0025] Preferably, S4 is specifically:

[0026] From the channel simulation, we get v base sequences. Due to the random errors of base insertion and deletion, we set the maximum ratio of base insertion to be γ ins , the maximum ratio of base deletion is γ del , the length of the base sequence is arrive Fluctuations within a certain range are recorded as ≈k:

[0027] ;

[0028] in, represents the jth base sequence, is a combination of v base sequences;

[0029] Constructing a scoring matrix , and initialize all elements to -v;

[0030] First, perform the isotope base comparison and scoring: for the j-th base sequence and its base index i∈[1,k], count the number of bases in the column that match the j-th base sequence in the i-th column. Same number of bases, write the value into Q j,i , Q j,i is the value of the jth row and ith column of the scoring matrix. If the i-th position is missing in a reading, the initial -v is retained as a penalty score;

[0031] Calculate the average score q for each sequence j :

[0032] ;

[0033] q j Sort in descending order, select the first w sequence indexes with the highest scores to form set I, and the determined index set is recorded as I prev , for the sequences with parallel scores that cannot determine I, that is, the index set P performs K-mers comparison scoring: for each parallel sequence with an original length of approximately k , generate overlapping substrings of length K on it with a sliding step size of 1: ;

[0034] in ;

[0035] Construct the corresponding K-mers scoring matrix for each parallel sequence j∈P , excluding missing or overlong segments, substring , in the same column Count all parallel sequences with The number of identical substrings is recorded as If a sequence has a missing base at this position, Assign -v as a strong penalty and calculate the average K-mers score for each parallel sequence:

[0036] ;

[0037] The parallel sequence set P is Rearrange in descending order and take the first , and finally the index that passes the K-mers rescreening will be compared with the initial selected index I prev Merge to form a complete set of w sequences Z.

[0038] Preferably, S5 is specifically: using a DNA image decoding network g with deep source-channel joint decoding φ Map the sequence set Z matrix back to the reconstructed image ,Right now: .

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] This application effectively improves the compression efficiency and robustness of DNA image storage by introducing semantic extraction and designing a multi-read screening mechanism. It extracts key information areas in the image through semantic extraction, eliminates redundant background, and reduces the encoding length. The multi-read screening mechanism uses scoring and sequence analysis to select high-quality sequences from sequencing copies to enhance fault tolerance. This application is suitable for efficient semantic storage in the Internet of Things scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flowchart of the application;

[0042] Figure 2 Schematic diagram of DNA storage image restoration quality under different multi-read screening numbers including PSNR and SSIM indicators;

[0043] Figure 3 Schematic diagram of the performance comparison of the DNA image coding method with deep source-channel joint coding (DJSCC-DNA). DETAILED DESCRIPTION

[0044] To facilitate those skilled in the art to understand the technical content of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] Example 1

[0046] like Figure 1 As shown, a semantic intelligence enhanced DNA storage method for the Internet of Things includes the following steps:

[0047] S1: Extract semantic information from the original image to obtain a semantically perceived image. Specifically:

[0048] The open source large model SAM, which includes an image encoder, a prompt encoder, and a mask decoder, is used to segment the original image and generate a set of semantic masks:

[0049] ;

[0050] Among them, M i is the mask of the i-th target, n is the number of targets in the image, x is the original image, and H×W represents the length and width of the mask;

[0051] Each mask M i Multiply the original image pixel by pixel to get the corresponding semantic segment s i :

[0052] ;

[0053] The attention fusion network ASI is introduced to identify and fuse the most informative semantic target region index set ℒ from the semantic fragments to generate the final semantic perception image x scm :

[0054] ;

[0055] Among them, s i Represents the i-th semantic segment.

[0056] S2: Encode the semantically perceived image into a DNA base sequence through feature extraction and base mapping, specifically: a DNA image encoding network f using deep source-channel joint coding θ The semantically aware image x scm Mapped to a DNA base sequence of length k , that is: z=f θ( x scm) , where A is adenine, T is thymine, G is guanine, C is cytosine, z k is the kth base in the DNA base sequence.

[0057] S3: Channel simulation is performed on the DNA base sequence. The channel simulation includes DNA synthesis, PCR amplification, and DNA sequencing. During the channel simulation, random errors such as insertions, deletions, and substitutions are added to the DNA base sequence. Multiple copies of the noisy base sequence are generated and stored in parallel, completing the storage process of the original image.

[0058] S4: Set up a multi-read screening mechanism and use scoring and sequence analysis to select multiple reliable base sequences from the stored base sequence copies to form a sequence set. Specifically:

[0059] From channel simulation or actual sequencing, we obtain v base sequences of approximately k length. Due to the random errors of base insertion and deletion, we set the maximum base insertion ratio to be γ ins , the maximum ratio of base deletion is γ del , the length of the base sequence is arrive Fluctuations within a certain range are recorded as ≈k:

[0060] ;

[0061] in, represents the jth base sequence, is a combination of v base sequences;

[0062] Constructing a scoring matrix , and initialize all elements to -v;

[0063] First, perform the isotope base comparison and scoring: for the j-th base sequence and its base index i∈[1,k], count the number of bases in the column that match the j-th base sequence in the i-th column. Same number of bases, write the value into Q j,i , Q j,i is the value of the jth row and ith column of the scoring matrix. If the i-th position is missing in a reading, the initial -v is retained as a penalty score;

[0064] Calculate the average score q for each sequence j :

[0065] ;

[0066] q j Sort in descending order, select the first w sequence indexes with the highest scores to form set I, and the determined index set is recorded as I prev , for the sequences with parallel scores that cannot determine I, that is, the index set P performs K-mers comparison scoring: for each parallel sequence with an original length of approximately k , generate overlapping substrings of length K on it with a sliding step size of 1: ;

[0067] in ;

[0068] Construct the corresponding K-mers scoring matrix for each parallel sequence j∈P , excluding missing or overlong segments, substring , in the same column Count all parallel sequences with The number of identical substrings is recorded as If a sequence has a missing base at this position, Assign -v as a strong penalty and calculate the average K-mers score for each parallel sequence:

[0069] ;

[0070] The parallel sequence set P is Rearrange in descending order and take the first , and finally the index that passes the K-mers rescreening will be compared with the initial selected index I prev Merge to form a complete set of w sequences Z.

[0071] S5: Decode the sequence set and map it to the reconstructed image to complete the reconstruction process of the original image. Specifically, the DNA image decoding network g is used for deep source-channel joint decoding. φ Map the sequence set Z matrix back to the reconstructed image ,Right now: .

[0072] In this embodiment, the present application builds an encoding end including an encoder module and a semantic extraction module, and a decoding end including a multi-read screening module and a semantic decoding module to perform image encoding into base sequences and DNA base sequences decoding into images. After completing the construction of the encoding and decoding end modules, all trainable parameters are first optimized uniformly through end-to-end training. To ensure that the system has stable semantic recognition capabilities and good image restoration performance, the model training adopts a two-stage strategy, as follows:

[0073] First, a training dataset was constructed using the semantic extraction module model. This module consists of the open-source large-scale SAM model and the attention fusion network ASI. The SAM and ASI modules use publicly available pre-trained weights and require no further training. The trained semantic extraction module is then used to process the raw image data to batch generate a corresponding semantically perceived image dataset. This dataset serves as the training input for the semantic encoder and decoder, forming the foundation for training the backbone model.

[0074] Subsequently, the semantic encoding network and semantic decoding network are jointly trained. The semantic image output by the semantic extraction module is fed into the encoder, where features are extracted using a convolutional neural network and a DNA base sequence is generated. The base sequence undergoes DNA channel noise simulation to generate multiple sequencing copies with errors. The multi-read screening module selects highly reliable sequences from these copies and uses them as input to the decoder to recover the image.

[0075] During training, image reconstruction error is the primary optimization objective, and a biologically constrained loss function is introduced to constrain the GC content and homopolymer length of the generated base sequences, thereby improving the feasibility of synthesis and sequencing. Ultimately, all encoder and decoder parameters are jointly optimized end-to-end using backpropagation under a unified loss function. This training strategy effectively improves the system's overall performance in terms of image compression rate and image restoration quality.

[0076] like Figure 2 and Figure 3As shown, the DNA storage image restoration quality under different multi-read screening numbers including PSNR and SSIM indicators is respectively demonstrated, as well as the performance comparison of the DNA image coding method of deep source channel joint coding (DJSCC-DNA). It can be seen that the application effectively improves the compression efficiency and robustness of DNA image storage by introducing semantic extraction and designing a multi-read screening mechanism. The key information areas in the image are extracted through semantic extraction, redundant background is eliminated, and the coding length is reduced. The multi-read screening mechanism uses scoring and sequence analysis to select high-quality sequences from the sequencing copies to enhance fault tolerance. The application is suitable for efficient semantic storage in the Internet of Things scenario, and solves the problem of low storage efficiency caused by excessive redundant information in the original image in the existing DNA image storage method, as well as the problem of high bit error rate in the DNA synthesis and sequencing process, which affects the image restoration quality.

Claims

1. A semantic intelligence enhanced DNA storage method for the Internet of Things, characterized by: The following steps are involved: S1: Extract semantic information from the original image to obtain a semantically perceived image; S2: Encode semantically-aware images into DNA base sequences through feature extraction and base mapping; S3: Channel simulation is performed on the DNA base sequence. The channel simulation includes DNA synthesis, PCR amplification, and DNA sequencing. During the channel simulation, random errors such as insertions and deletions are added to the DNA base sequence. Multiple copies of the noisy base sequence are generated and stored in parallel, completing the storage process of the original image. S4: Set up a multi-read screening mechanism, and use scoring and sequence analysis to select multiple reliable base sequences from the stored base sequence copies to form a sequence set; S5: Decode and map the sequence set into a reconstructed image to complete the reconstruction process of the original image.

2. The semantic intelligence enhanced DNA storage method for the Internet of Things according to claim 1, characterized in that: S1 is specifically: The open source large model SAM, which includes an image encoder, a prompt encoder, and a mask decoder, is used to segment the original image and generate a set of semantic masks: ; Among them, M i is the mask of the i-th target, n is the number of targets in the image, x is the original image, and H×W represents the length and width of the mask; Each mask M i Multiply the original image pixel by pixel to get the corresponding semantic segment s i : ; The attention fusion network ASI is introduced to identify and fuse the most informative semantic target region index set ℒ from the semantic fragments to generate the final semantic perception image x scm : ; Among them, s i Represents the i-th semantic segment.

3. The semantic intelligence enhanced DNA storage method for the Internet of Things according to claim 2, characterized in that: S2 is specifically: DNA image coding network f using deep source-channel joint coding θ The semantically aware image x scm Mapped to a DNA base sequence of length k , that is: z=f θ( x scm) , where A is adenine, T is thymine, G is guanine, C is cytosine, z k is the kth base in the DNA base sequence.

4. The semantic intelligence enhanced DNA storage method for the Internet of Things according to claim 3 is characterized in that: S4 is specifically: From the channel simulation, we get v base sequences. Due to the random errors of base insertion and deletion, we set the maximum base insertion ratio to be γ ins , the maximum ratio of base deletion is γ del , the length of the base sequence is arrive Fluctuations within a certain range are recorded as ≈k: ; in, represents the jth base sequence, is a combination of v base sequences; Constructing a scoring matrix , and initialize all elements to -v; First, perform the isotope base comparison and scoring: for the j-th base sequence and its base index i∈[1,k], count the number of bases in the column that match the j-th base sequence in the i-th column. Same number of bases, write the value into Q j,i , Q j,i is the value of the jth row and ith column of the scoring matrix. If the i-th position is missing in a reading, the initial -v is retained as a penalty score; Calculate the average score q for each sequence j : ; q j Sort in descending order, select the first w sequence indexes with the highest scores to form set I, and the determined index set is recorded as I prev , for the sequences with parallel scores that cannot determine I, that is, the index set P performs K-mers comparison scoring: for each parallel sequence with an original length of approximately k , generate overlapping substrings of length K on it with a sliding step size of 1: ; in ; Construct the corresponding K-mers scoring matrix for each parallel sequence j∈P , excluding missing or overlong segments, substring , in the same column Count all parallel sequences with The number of identical substrings is recorded as If a sequence has a missing base at this position, Assign -v as a strong penalty and calculate the average K-mers score for each parallel sequence: ; The parallel sequence set P is Rearrange in descending order and take the first , and finally the index that passes the K-mers rescreening will be compared with the initial selected index I prev Merge to form a complete set of w sequences Z.

5. The semantic intelligence enhanced DNA storage method for the Internet of Things according to claim 4, characterized in that: S5 is specifically: DNA image decoding network g using deep source-channel joint decoding φ Map the sequence set Z matrix back to the reconstructed image ,Right now: .

Citation Information

Cited By

  • DNA coding method and device with biological constraint

    CN120954503A

  • High-voltage power supply DNA sequencing visual detection method and system

    CN121280430A