Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

53 results about "Dna storage" patented technology

DNA storage is the process of encoding and decoding binary data onto and from synthesized strands of DNA (deoxyribonucleic acid). In nature, DNA molecules contain genetic blueprints for living cells and organisms.

Preserving solution for stably preserving sample DNA (deoxyribonucleic acid) at normal temperature and preparation method of preserving solution

The invention discloses a preserving fluid for stably preserving sample DNA at normal temperature and a preparation method thereof, and belongs to the technical field of biological sample preservation. The preserving fluid comprises a lysis system, a nucleic acid protection system and a buffering and stabilizing system, the cracking system comprises a composite surfactant and an enzymolysis auxiliary agent; the composite surface active agent is prepared from polyether polyol fatty acid ester and cocamidopropyl hydroxy sulfobetaine; the enzymolysis auxiliary agent comprises lysozyme Lyso-V and protease K; the nucleic acid protection system comprises a nitrogen heterocyclic polyamine-carboxylic acid derivative and dextran sulfate; the nitrogen heterocyclic polyamine-carboxylic acid derivative comprises 1, 4, 7, 10-tetraazacyclododecane-N, N ', N' ', N ''tetraacetic acid, 1, 4, 7-triazacyclononane-N, N', N''-triacetic acid, disodium ethylene diamine tetraacetate-nitrogen heterocyclic derivative, and diethylenetriamine pentaacetic acid-piperazine derivative, and the nitrogen heterocyclic polyamine-carboxylic acid derivative comprises 1, 4, 7, 10-tetraazacyclododecane-N, N ', N '', N'' tetraacetic acid, 1, 4, 7-triazacyclononane-N, N', N ''-triacetic acid. The buffering and stabilizing system comprises an amphoteric buffering agent and a polymer stabilizer; the amphoteric buffering agent comprises 2-(N-morpholino) ethanesulfonic acid and N-tri (hydroxymethyl) methylglycine.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +1

DNA sequence reconstruction method and system based on multi-scale attention and contrast learning

The invention discloses a DNA sequence reconstruction method and system based on multi-scale attention and contrast learning, and relates to the technical field of DNA storage data reconstruction. Comprising the following steps: collecting a plurality of DNA sequence copies, screening out abnormal length sequences, and constructing a standardized clustering data set; performing one-hot coding and filling processing on the DNA sequence; extracting context dependent features and cross-sequence variation features; an Inter-Sequence multi-head attention mechanism is constructed to calculate the similarity between the sequences, and a weighted sequence tensor is generated; a global dependency relationship in the sequence is extracted through an Intra-Sequence multi-head attention mechanism; local offset features caused by insertion and deletion errors are extracted through a multi-size convolutional network; inputting a double-layer long-short-term memory network for sequence-level modeling, and outputting base reconstruction probability distribution; and constructing positive and negative sample pairs, calculating comparison loss, combining cross entropy loss to form a joint loss function, and outputting a high-precision DNA sequence reconstruction result. The method has high accuracy and robustness under the conditions of complex noise and multiple types of errors.
Owner:DALIAN UNIV

Invasive species identification method and device based on deep learning and DNA storage, and electronic equipment

The invention provides an invasive species identification method and device based on deep learning and DNA storage and electronic equipment, and relates to the technical field of invasive organism prevention and control, and the method comprises the steps: obtaining a to-be-identified invasive species image, and generating a first DNA sequence of a to-be-identified invasive species in the invasive species image through a deoxyribonucleic acid DNA encoder; obtaining a plurality of second DNA sequences to be hybridized from the invasive species database, and obtaining the hybridization yield of the first DNA sequence and each second DNA sequence by using a hybridization yield predictor; and determining the species category corresponding to the DNA sequence with the highest hybridization yield as the species category of the invasive species to be identified. According to the method provided by the invention, the strong feature extraction capability of deep learning is combined with the advantages of ultrahigh density and ultra-long stability of DNA as a data storage medium, so that end-to-end mapping and recognition from a species image to an exclusive DNA sequence thereof are realized.
Owner:BINZHOU MEDICAL COLLEGE

Multi-channel extensible automatic sample loading system oriented to DNA storage and calculation and control method

The invention discloses a multi-channel extensible automatic sample loading system for DNA storage and calculation and a control method. The system comprises a gas path control device, a multi-channel sample introduction device, a liquid flow control device, a multi-liquid-path liquid collection device, a micro-fluidic chip and an upper computer, operation control over the whole system is provided through the upper computer, a user can adjust the sample injection sequence and the reaction time of each channel, the liquid flow and the gas flow of each channel are monitored and adjusted in real time in an experiment, and multi-channel and high-precision automatic sample injection is achieved. According to the invention, accurate injection and automatic control of samples are realized, the efficiency and precision of sample treatment are greatly improved, and high automation, high-precision sample injection and multi-channel low cross contamination control capability are realized; the expandability is high, and modular expansion to hundreds of channels or even hundreds of channels is supported; meanwhile, the method is compatible with various application scenes, and is suitable for the fields of high-throughput DNA calculation, DNA information storage, biomedical detection and the like.
Owner:SHANGHAI JIAOTONG UNIV

Method for prolonging data storage time and application

PendingCN121343982ADNA preparationDNA stabilityEngineering
The invention relates to the technical field of DNA information storage, and discloses a method for prolonging data storage time and application, the method comprises the following steps: data is converted into a DNA sequence with a preset length, DNA is synthesized and stored, primary amino groups in basic groups of the DNA are independently modified by protective groups, the protective groups are selected from-COR, and the primary amino groups in the basic groups of the DNA are independently modified by the protective groups. R is selected from alkyl of C1-C6, substituted or unsubstituted naphthenic base of C3-C8 or substituted or unsubstituted phenyl. According to the method for prolonging the data storage time, the degradation rate of DNA can be remarkably inhibited, the stability of the DNA can be improved so as to realize long-time stable storage of the data, and the method is particularly suitable for forming a DNA storage scene with limited conditions, namely, the method is suitable for DNA synthesis in large companies with advanced technologies and equipment and also suitable for large companies with advanced technologies and equipment. The method is also suitable for DNA synthesis of small companies, and has application prospects and potential.
Owner:SHANGHAI DYNASTYGENE CO

A DNA storage method and data information storage entity capable of realizing data random lossless reading

The application provides a DNA storage method and a data information storage entity capable of realizing random lossless reading of data, and the DNA storage method comprises the following steps: S1, preparation of a DNA double-stranded structure; S2, preparation of a data information storage entity; S3, random lossless reading of DNA data; and S4, recycling of the data information storage entity. According to the DNA storage method provided by the application, the random access function is realized by using the DNA double-stranded structure, and the long-term storage function is realized by using the outer wrapping hydrogel material, random reading in a file system with a large number of files is realized, and the original data is ensured not to be lost, so that the application has a wide application prospect in the field of DNA storage.
Owner:XIANGFU LAB

DNA sequence assembly method and system based on dynamic variable-order unitg-level graph

The invention discloses a DNA sequence assembly method and system based on a dynamic variable order unitg-level graph, and relates to the technical field of bioinformatics and DNA storage. According to the method, a pseudo genome and a source perception k-mer index are constructed, a Mid-Max and Min-Mid two-stage variable order expansion strategy is adopted, a k value is dynamically adjusted to enhance the graph structure connectivity, and the problems that under the condition of low coverage rate or high error rate, an existing de Bruijn graph method is prone to breakage and path fuzziness is prone to being generated are effectively solved. According to the method, a hidden path is accurately repaired through node connection and splitting operation, and redundancy k-mer is filtered in combination with index continuity and prefix similarity, so that the assembly integrity and accuracy are improved, and meanwhile, the robustness of DNA data reconstruction is remarkably enhanced; the method is suitable for various high-noise and low-coverage-rate scenes such as genome assembly and DNA storage and reconstruction, and particularly shows excellent performance in practical application with high requirements on data integrity and reliability.
Owner:DALIAN UNIV

Design method of file system architecture oriented to DNA storage block equipment

The invention discloses a design method of a file system architecture oriented to DNA storage block equipment, and belongs to the technical field of DNA storage. The problem of providing more convenient, efficient and abundant DNA storage services for upper-layer users is solved. The method comprises the following steps: constructing a physical layer in a file system oriented to DNA storage block equipment, defining a primer block, and constructing a method for reading and writing DNA data input by DNA data input equipment oriented to the DNA storage block equipment on the basis of the primer block; constructing a file system layer in the file system oriented to the DNA storage block equipment, and establishing a method for converting the primer blocks into logic data blocks in the file system layer; and constructing the file system oriented to the DNA storage block device and a read-write delay optimization method of the DNA data input device, and completing the overall design of the file system oriented to the DNA storage block device. According to the invention, the data access efficiency of DNA storage is improved.
Owner:HARBIN INST OF TECH

DNA storage medium-oriented scalable vector graphics (SVG) image coding method and system

The invention provides a DNA storage medium-oriented scalable vector graphics (SVG) image coding method and system, and the method comprises the steps: reading an SVG file, analyzing the SVG image to obtain a document tree containing a plurality of nodes, and distributing structure information for representing the structure position of the nodes in the document tree for the nodes; coding the nodes to generate node fragments, wherein the node fragments comprise label codes generated according to labels of the nodes and attribute codes generated according to attributes of the nodes; aggregating the plurality of node fragments according to the label codes thereof to form at least one aggregation block, the aggregation block comprising header information and the plurality of node fragments, the header information being used for indexing the node fragments contained therein; the at least one aggregate block is converted into at least one DNA base sequence, and error correction information for error verification or repair is added to the at least one DNA base sequence. According to the method, the encoding length is reduced while the semantic integrity of the SVG is maintained, the error-resistant capability is improved, and the progressive decoding capability is provided.
Owner:SHANGHAI JIAOTONG UNIV

A biomimetic mineralized material with high stability for protecting DNA and a preparation method thereof

The application discloses a kind of biomimetic mineralization materials with high stable protection DNA and preparation method thereof.DNA molecule is self-assembled with divalent metal ion, and DNA / metal nanoparticle is obtained; polyelectrolyte layer is wrapped on the surface of DNA / metal nanoparticle by layer-by-layer assembly technology, and DNA / metal@LBL nanoparticle is obtained; polydopamine layer is deposited on the surface of DNA / metal@LBL nanoparticle, and DNA / Fe@LBL@PDA particle is obtained, which is biomimetic mineralization material with high stable protection DNA.The biomimetic mineralization material has the advantages of simple construction, high density and long-term stability, and can effectively protect the stability and integrity of DNA molecule in harsh external storage environment (such as ROS, high temperature, high salt, alkaline, humid, biological solvent and nuclease environment).The present application can provide technical support for DNA high-density storage and long-term stability storage, and has wide application value in the fields of biological medicine, genetic engineering, disease treatment and DNA storage.
Owner:FUZHOU UNIV

Robust multi-sequence reconstruction method based on maximum a posteriori probability in DNA storage

The application discloses a robust multi-sequence reconstruction method based on maximum posterior probability in DNA storage, and comprises the following steps: decoding each sequence in a cluster by using an improved BCJR decoder; converting each sequence in the cluster into corresponding time sequence representation; determining the weight of each sequence; and deducing a MAP algorithm formula of robust multi-sequence decoding, and integrating the sequence weight into joint posterior probability calculation to inhibit the influence of outlying sequences. By proposing a base IDS channel model, the BCJR decoding algorithm is improved, so that the defects that the existing statistical inference sequence reconstruction method is only suitable for binary error correction code are overcome. Therefore, the application can process the quaternary error correction of the base form in the DNA sequence.
Owner:TIANJIN UNIV

DNA (deoxyribonucleic acid) data storage molecular tag based on nanopore as well as preparation method and application of DNA data storage molecular tag

The invention relates to a DNA data storage molecular tag based on nanopores and a preparation method and application thereof, and belongs to the technical field of DNA storage. The molecular tag is prepared by modifying azido polysaccharide onto an alkyne-containing DNA chain through a click chemical reaction. Hybridizing and fixing the molecular tag and a DNA bracket chain containing a specific notch structure, and scanning by using a glass nanopore; when the molecular tag passes through the nanopore, a characteristic ionic current blocking signal is generated, and accordingly data coding is achieved. The molecular tag disclosed by the invention is low in cost and simple and convenient to operate; polysaccharide with wide sources and efficient click chemistry are utilized, so that the preparation cost and complexity are remarkably reduced; the physicochemical properties of polysaccharide molecules are stable, the consistency of read signal amplitudes is ensured, the decoding accuracy is high, the label structure stability is good, and a guarantee is provided for long-term storage of data; based on a reversible hybridization storage structure, data updating can be realized without synthesizing a new DNA chain.
Owner:CHONGQING UNIV OF POSTS & TELECOMM +1

DNA storage encoding method based on graph convolution network and self-attention mechanism

The application discloses a DNA storage coding method based on a graph convolution network and a self-attention mechanism, and belongs to the technical field of coding in DNA storage. Specifically, a DNA coding sequence meeting a combination constraint condition is predicted, first, existing DNA coding is screened and data is cleaned to construct a DNA storage coding training set; second, a prediction model based on a graph convolution neural network and a self-attention mechanism is trained, and the self-attention mechanism is used to capture the relationship of local DNA coding; then, the coding data processed into a graph is input into the prediction model to perform coding prediction meeting the combination constraint; finally, a DNA storage coding set meeting the condition is output. The application constructs a DNA storage coding training set, trains a graph convolution self-attention neural network, better captures the relationship between codings, and adopts a learning-based prediction model to perform DNA storage coding, so that the application has high coding efficiency when processing codings with complex constraints.
Owner:DALIAN UNIV OF TECH

DNA sequencing read clustering method and system based on representation learning

The invention belongs to the crossing field of DNA digital storage and bioinformatics, and discloses a DNA sequencing read clustering method and system based on characterization learning, and the method comprises the steps: carrying out the preprocessing of an original DNA sequencing read, and obtaining a sequencing read and a corresponding variant 1 and variant 2 set; carrying out characterization learning on the DNA sequencing read segments through a deep learning model based on the sequencing read segments and the corresponding variant 1 and variant 2 sets; and based on the DNA sequencing read after characterization learning, realizing clustering of the DNA sequencing read through a fine tuning model. According to the method, the problem of clustering difficulty caused by sequencing errors in DNA storage is solved.
Owner:GUANGZHOU UNIVERSITY

DNA storage method based on data superposition pseudorandom sequence index

The invention discloses a DNA (deoxyribonucleic acid) storage method based on data superposition pseudorandom sequence index, which comprises the following steps of: firstly, performing exclusive OR on a pseudorandom sequence and a sparse coding sequence to generate an addressable short fragment oligonucleotide sequence; then, by using the superposed pseudo-random sequence as an index, correlating the pseudo-random sequence hidden in the sequencing read with the known pseudo-random sequence at the boundary of each pseudo-random sequence by using a small sliding window to realize quick positioning of the read so as to determine the position of the read on the whole pseudo-random sequence; and finally, reconstructing a sparse coding sequence by adopting a majority voting algorithm, and realizing error-free recovery of data through sparse code decoding and iterative decoding. The method has the advantages that sequence positioning can be achieved only through correlation operation at the boundary of each pseudo-random sequence, local label sequences do not need to be added to oligonucleotide molecules, data recovery performance reduction caused by label sequence damage, sequence breakage and the like is avoided, and rapid data recovery can be achieved under low sequencing coverage.
Owner:TIANJIN UNIV SYNTHETIC BIOLOGY FRONTIER RES INST

Method and system for realizing DNA (Deoxyribonucleic Acid) storage by aiming at multi-rule rotation coding of Chinese text

The invention discloses a method and a system for realizing DNA storage by aiming at multi-rule rotation coding of Chinese texts, and relates to the technical field of DNA storage. Encoding the Chinese text by using a five-stroke font input method; mapping is carried out according to the positive and negative code tables, and an interval balance and dynamic detection strategy is introduced in the mapping process to control GC content distribution in intervals and the probability of occurrence of homopolymers between the intervals. Scattering and recombining the sequence by using block coding, and compressing by using RLE coding, wherein the RLE coding generates a sequence file and a run-length file; performing GC content constraint on the sequence file by using cross coding; and br compression is adopted for the run-length file to further improve the compression ratio. And generating a DNA sequence through rotary coding. According to the method, five-stroke coding, positive and negative code table mapping and sequence reconstruction strategies are introduced, so that the information storage density is higher, the local GC content of the DNA sequence is more stable, the homopolymer length is smaller, the unexpected motif proportion is lower, and the data has higher reliability and safety in the storage process.
Owner:DALIAN UNIV

Iterative decoding method and system of VT code without prior information in DNA storage and storage medium

The application discloses a kind of DNA storage in prior information VT code iterative decoding method, system and storage medium, it is related to DNA storage technical field.The method is first obtained systematic code word sequence by VT code, generates noisy reception sequence by ID model simulation channel transmission;Again, joint grid chart is constructed and initialized, and soft information and bit decision information are obtained by completing SISO decoding calculation, and channel parameters are updated based on MAP drift tracking;After convergence by iterative control, multiple independent estimated code words are obtained, and finally original information sequence is recovered by multi-sequence majority voting fusion.The application constructs decoding-estimation closed-loop iteration framework, does not need prior channel parameters, can adaptively learn channel characteristics, capture error hot spot, improve decoding robustness by combining multi-sequence fusion, realize original data high reliability recovery, and promote DNA storage practicality.
Owner:TIANJIN UNIV

A normal-temperature stable sample DNA storage solution and a preparation method thereof

This invention discloses a preservation solution for preserving sample DNA at room temperature and its preparation method, belonging to the field of biological sample preservation technology. The preservation solution comprises: a lysis system, a nucleic acid protection system, and a buffering and stabilizing system; the lysis system comprises a complex surfactant and an enzymatic hydrolysis aid; the complex surfactant comprises: polyether polyol fatty acid ester and cocamidopropyl hydroxysulfonate betaine; the enzymatic hydrolysis aid comprises lysozyme Lyso-V and proteinase K; the nucleic acid protection system comprises: nitrogen-heterocyclic polyamine-carboxylic acid derivatives and dextran sulfate; the nitrogen-heterocyclic polyamine-carboxylic acid derivatives comprise: 1,4,7,10-tetraazacyclododecane-N,N',N'',N'''tetraacetic acid, 1,4,7-triazacyclononane-N,N',N''-triacetic acid, disodium ethylenediaminetetraacetate-nitrocyclic derivative, and diethylenetriaminepentaacetic acid-piperazine derivative; the buffering and stabilizing system comprises: an amphoteric buffer and a polymeric stabilizer; the amphoteric buffer comprises: 2-(N-morpholino)ethanesulfonic acid and N-tris(hydroxymethyl)methylglycine.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL +1

End-to-end DNA data storage coding method and system based on multi-constraint loss

The invention discloses an end-to-end DNA data storage coding method and system based on multiple constraint loss, and relates to the technical field of DNA storage. A complementary matrix convolution detection mechanism is designed in a coding stage, a hairpin structure is found through the complementary matrix convolution detection mechanism, a micronizable loss function is constructed, and a micronizable hairpin structure constraint loss function is realized. An interpretable error predictor based on Transform is trained through real sequencing data, a penalty loss function is constructed by means of an attention mechanism error-prone motif mask matrix to achieve error-prone motif suppression, and then GC content loss and homopolymer loss are fused to form a multi-constraint loss function. According to the method, an end-to-end architecture with a Transform auto-encoder as a main body is constructed, mapping of image data and a DNA sequence is realized through differentiable arithmetic coding, data reconstruction is completed in combination with noise injection and reverse decoding, and in a training stage, a deep Q network is adopted to carry out adaptive joint adjustment on weights of a multi-constraint loss function and mean square error reconstruction loss, so that data reconstruction is completed. And end-to-end optimization of model parameters is realized.
Owner:DALIAN UNIV

DNA sequence direct encryption method and system based on automaton cryptography

The invention discloses a DNA sequence direct encryption method and system based on automaton cryptography, and relates to the technical field of DNA storage security. By designing a DS-mealy machine and a DS-cellular automaton, direct encryption of a DNA sequence is achieved, and the problems that in an existing DNA storage encryption method, compatibility with a coding model is poor, safety is insufficient, performance is low, and expandability is weak are solved. According to the method, sequence-level encryption is realized through base diffusion and rotation operation, so that base distribution is more uniform and random, and the security of DNA storage data is improved; meanwhile, the space resource consumption of encryption and decryption is reduced, the processing speed is increased, and the method is suitable for various DNA storage scenes and particularly suitable for storage of sensitive information such as medical data with high requirements for safety and efficiency.
Owner:DALIAN UNIV

A method and system for converting ancient text sequences for DNA storage data based on five-stroke coding

The application discloses a kind of ancient text sequence conversion method and system for DNA storage data based on five-stroke encoding, comprising the following steps: ancient text is converted into five-stroke encoding alphabet sequence according to word;Five-stroke encoding alphabet sequence is compressed by Huffman ternary, and ternary digital string is obtained;Ternary digital string is converted into DNA base sequence by dynamic mapping rule;DNA base sequence is segmented, and positioning code is added to each segment, and the final DNA sequence set suitable for DNA synthesis and high-throughput sequencing is output;The final DNA sequence set is synthesized, sequenced, and the ancient text is restored by reverse operation, and its reliability is verified.The method can efficiently and accurately convert ancient text to DNA storage sequence, ensure biological compatibility and data integrity restoration, and meet the long-term DNA storage needs of ancient text.
Owner:NANJING UNIV OF SCI & TECH

Image DNA storage method for improving information density and biological stability

The invention provides an image DNA (deoxyribonucleic acid) storage method for improving information density and biological stability, which comprises the following steps of: partitioning an input image, performing two-dimensional discrete wavelet transform, performing quantization processing on a wavelet coefficient by adopting a self-adaptive quantization strategy, and converting the wavelet coefficient into a binary sequence to be coded; the method comprises the following steps: dividing a binary sequence to be coded into a byte according to every 8 bits, and carrying out dual-mode DNA coding by adopting an odd-even alternating mode and a segmentation structure of'feature segment-coding segment 'to generate a DNA information segment sequence; performing error correction on the coded DNA sequence based on preliminary correction of Hamming distance and RS code secondary error correction; marking the DNA sequence which still does not pass the CRC verification after multiple rounds of error correction as an unrepairable sequence, wherein the image block corresponding to the unrepairable sequence is an error block; and in combination with context information, the neighborhood features of the error blocks are learned by using a cross Transform context repair network, and the error blocks are repaired. According to the method, better balance is achieved in the aspects of image reconstruction quality, coding density and biological compatibility.
Owner:ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY

Methods and apparatus for DNA storage coding and decoding and rules thereof

The invention discloses a method and a device for DNA storage coding and decoding and rules thereof. The method comprises the following steps: performing single-molecule sequencing on a reference sequence to obtain actual sequencing data of single-molecule sequencing; comparing the actual sequencing data with reference data of the reference sequence, counting the frequency of sequencing errors of each sequence fragment with the length of k in the actual sequencing data, and calculating the proportion of the sequencing errors of each sequence fragment with the length of k in the actual sequencing data, namely the error rate; and taking the sequence fragments of which the error rates exceed a threshold value as limiting conditions to be eliminated. According to the method provided by the invention, the DNA storage coding and decoding steps are simplified, and the complexity of data processing is reduced through the time sequence of the threshold elimination step.
Owner:SHENZHEN HUADA GENE INST

A DNA storage readout method based on multiple hidden reference sequences

ActiveCN120596016BA-DNATesting Methods
The application discloses a DNA storage reading method based on multiple hidden reference sequences, which constructs multiple hidden reference sequences to realize reliable recovery of large fragment DNA from a read containing insertion / deletion errors; the constructed hidden reference sequence changes the reading of large fragment DNA storage from scratch into a low complexity resequencing problem; wherein, the watermark reference sequence quickly identifies low error reads, the skeleton reference sequence identifies reads containing insertion / deletion errors, and the decoding feedback reference sequence identifies reads for filling low coverage areas; the forward-backward algorithm of each read corrects the insertion / deletion errors in stages to realize reliable data recovery. The application realizes reliable data reading from reads containing insertion / deletion errors, and has the advantages that the constructed multiple hidden reference sequences identify reads with complementary characteristics of error rates in stages, the forward-backward algorithm of each read corrects the insertion / deletion errors in stages, and gradual reading is realized.
Owner:TIANJIN UNIV

Coding method and decoding method applied to DNA storage and related equipment

The embodiment of the invention provides an encoding method, a decoding method and related equipment applied to DNA storage. The encoding method comprises the following steps: acquiring a to-be-encoded first base sequence and a first mapping relation; the first mapping relation is a mapping relation between the base fragments with various lengths and the coded values; matching the first base sequence with a base fragment contained in the first mapping relation to obtain a matching result; dividing the first base sequence into a plurality of segments to be coded according to a matching result; according to a first mapping relationship, converting each fragment to be coded in the first base sequence into a coded value to obtain a numerical sequence; and converting the numerical sequence into a second base sequence. According to the encoding method and the decoding method provided by the embodiment of the invention, DNA storage is carried out, and higher confidentiality and safety are achieved.
Owner:BEIJING BOE TECH DEV CO LTD +1

High-reliability DNA storage coding and decoding method and related device

The invention discloses a high-reliability DNA storage coding and decoding method and a related device, and the method comprises the steps: obtaining a to-be-coded text information sequence, and filling a two-dimensional information matrix with the to-be-coded text information sequence; reading a character sequence in the two-dimensional information matrix from different direction dimensions, and generating a first DNA sequence and a second DNA sequence according to a preset direct mapping relation by using a reading result; sequentially segmenting the first DNA sequence and the second DNA sequence to obtain a plurality of short segments, endowing each short segment with a dimension label according to the direction dimension corresponding to each short segment, and endowing each short segment with a unique index according to the segmentation sequence of each short segment; according to the method and the related device, DNA storage coding and decoding can be carried out on Chinese information, and the technical problems that in the prior art, information coding and decoding efficiency is low, an error correction mechanism is complex, and cascade data are prone to being lost are solved.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Spatially layered DNA storage method for large-scale oligo pools

The present disclosure discloses a spatially layered DNA storage method for large-scale oligonucleotide pools, employing a DNA spatially layered coding method to enable real-time data readout; the unordered DNA strands are spatially organized into an addressable base array, and the live data are encoded chronologically into sequential coding layers, wherein bases are mapped to crosscutting identical positions across all strands; for recovery, a live and accelerated approach to spatially form a coding layer is provided, and the error correction codes are utilized to fill the base gap, enabling continuous, real-time streaming; a layer-wise spatial-temporal recovery method is presented to facilitate an error-free data stream, spatially achieving instant consensus of multiple signals within a layer, and temporally updating flow signals via the previous successfully decoded layers; the error correction and readout methods provided by the present disclosure can match the sequencing process, achieving simultaneous sequencing and real-time decoding.
Owner:TIANJIN UNIV

DNA data fingerprint generation method, device, equipment, medium and product

The invention discloses a DNA data fingerprint generation method, device and equipment, a medium and a product, and relates to the technical field of data fingerprint and biological information crossing. The method comprises the following steps: preprocessing a target DNA sequence file to extract a plurality of K-mer subsequences, filtering an extraction result to obtain a plurality of residual K-mer subsequences of which the coverage is higher than or equal to a coverage threshold, and selecting N feature subsequences from a filtering result, and finally, standardizing and organizing the selected sub-sequence and the coverage / and the hash value of the selected sub-sequence into a data fingerprint of the target DNA sequence file, so that efficient, compact and powerful data fingerprint capturing and data copying tracking traceability capabilities can be provided for a DNA information storage system, that is, the problems of data rapid identification and content retrieval and verification in the field of DNA information storage are solved; key support is provided for practicability of the DNA storage technology, and practical application and popularization are facilitated.
Owner:TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI

Neural network enhanced base drift type DNA storage error correction coding method

The invention relates to the technical field of coding, and particularly discloses a neural network enhanced base drift type DNA storage error correction coding method. The method comprises the following steps: firstly, mapping binary data into a DNA (Deoxyribose Nucleic Acid) coding unit meeting multi-dimensional biological constraints such as GC (Gas Chromatography) content, homopolymer length, orthogonality and repeated substring, and establishing a coding dictionary (Codebook); insertion and deletion errors introduced in the DNA synthesis, storage and sequencing process are simulated, and detection is carried out through a classification model combining a one-dimensional convolutional neural network and Transform. And for the sequence which is judged to have the insertion / deletion error, comparing a sliding window with a coding dictionary, positioning the error based on an editing distance, and executing base drift type error correction. According to the method, insertion / deletion errors up to 1.1% can be effectively repaired under the condition that no extra error correction code exists, the average decoding time is shortened by about 9% under the 2% error rate, the error correction capability and the calculation efficiency are both considered, and the method is suitable for a large-scale DNA data storage system.
Owner:TIANJIN UNIV

A Method and System for Evaluating DNA Storage Sequencing Depth Based on Channel Simulation

This invention discloses a method and system for evaluating DNA storage sequencing depth based on channel simulation, addressing the problems of inaccurate predictions and limited guidance in existing uniform distribution models. The method includes: when sequencing data is available, fitting real data to obtain log-normal distribution parameters μ and σ; when sequencing data is unavailable, obtaining μ and σ through simulation modeling based on experimental parameters; and combining these parameters to calculate the decoding ratio of the coding strand in the noiseless channel and the sequencing depth boundary in the noisy channel. The system includes input, storage channel modeling, sequencing depth calculation, and output modules. The model of this invention closely reflects reality, provides accurate predictions, reduces DNA storage and retrieval costs, improves the success rate of single-sequencing decoding, and is easy to use, facilitating the practical application of the technology.
Owner:TIANJIN UNIV SYNTHETIC BIOLOGY FRONTIER RES INST