The invention relates to the technical field of coding, and particularly discloses a neural network enhanced base drift type
DNA storage
error correction coding method. The method comprises the following steps: firstly, mapping binary data into
a DNA (
Deoxyribose Nucleic Acid) coding unit meeting multi-dimensional biological constraints such as GC (
Gas Chromatography) content, homopolymer length, orthogonality and repeated
substring, and establishing a coding dictionary (
Codebook);
insertion and deletion errors introduced in the
DNA synthesis, storage and sequencing process are simulated, and detection is carried out through a classification model combining a one-dimensional
convolutional neural network and Transform. And for the sequence which is judged to have the
insertion / deletion error, comparing a sliding window with a coding dictionary, positioning the error based on an editing distance, and executing base drift
type error correction. According to the method,
insertion / deletion errors up to 1.1% can be effectively repaired under the condition that no extra error
correction code exists, the average decoding time is shortened by about 9% under the 2% error rate, the error correction capability and the calculation efficiency are both considered, and the method is suitable for a large-scale
DNA data storage system.