Information encoding and decoding method and apparatus for DNA storage of video data, device, and medium
By performing segmented parallel encoding of video data and using bit-base mapping rules, the problem of difficult to take into account both encoding speed and density in the prior art is solved, and efficient video data DNA storage is achieved.
Patent Information
- Application Number
- PCT/CN2023/142112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-12
AI Technical Summary
The existing DNA information storage technology is difficult to balance the encoding speed and encoding density, and it is impossible to implement a fast and high-density encoding method.
By segmenting the target video, using parallel encoding, combined with bit-base mapping rules, the video information is converted into efficient base sequences and stored through DNA synthesis.
It realizes fast encoding of video data while maintaining high encoding density, which not only improves encoding efficiency but also ensures encoding density.
Smart Images

Figure CN2023142112_12062025_PF_FP_ABST
Abstract
Description
Method, device, equipment and medium for encoding and decoding information stored in DNA of video data Technical Field
[0001] The present invention belongs to the technical field of DNA information storage, and in particular relates to an information encoding and decoding method, device, equipment and medium for DNA storage of video data. Background Art
[0002] DNA storage is a new field that integrates DNA synthesis and sequencing technology with computer storage. It stores digital information through the ordered combination of base pairs. The encoding algorithm converts the binary data 0 and 1 in the computer into a DNA sequence composed of four bases: A, T, C, and G. Then, by synthesizing DNA containing the specified base sequence, the data information can be stored. Compared with the storage media we commonly use, such as USB flash drives, CDs, hard drives, etc., DNA storage has the advantages of high storage density, long shelf life, low maintenance cost, and easy data backup.
[0003] The encoding methods involved in existing DNA information storage technologies are mainly divided into two categories: encoding and decoding methods based on predefined rule mapping tables and encoding and decoding methods that introduce screening mechanisms. Among them, although the former has a fast encoding speed, the encoding density is low, while the latter has a high encoding density but a more complex encoding method and is too slow. Technical issues
[0004] The present invention provides an information encoding and decoding method, device, electronic device and storage medium for video data DNA storage, aiming to propose a method with fast encoding speed and high encoding density to solve the problem that encoding speed and encoding density cannot be taken into account at the same time in existing methods. Technical Solutions
[0005] In order to solve the above technical problems, in the first aspect, the invention provides a method for encoding and decoding information stored in DNA of video data, the method comprising:
[0006] Splitting the target video to obtain multiple video segments of the target video;
[0007] Parsing each of the video clips to obtain video information of the video clip, and converting the video information of each of the video clips into a corresponding sequence to be encoded, wherein the sequence to be encoded is an N-ary sequence; the video information includes multimedia information and meta information;
[0008] Based on a bit-base mapping rule, encoding each of the sequences to be encoded into a plurality of base sequences corresponding to each of the video clips;
[0009] Corresponding preset primers are added to the plurality of base sequences, and DNA synthesis is performed on each of the base sequences after primer addition and stored to obtain DNA storage data of the target video; the preset primers are related to the positions of the video segments corresponding to the base sequences in the target video.
[0010] In a second aspect, the invention provides a DNA storage information encoding and decoding device, the device comprising
[0011] A splitting unit, configured to split a target video to obtain multiple video segments of the target video;
[0012] a parsing unit, configured to parse each of the video clips to obtain video information of the video clip, and convert the video information of each of the video clips into a corresponding sequence to be encoded, wherein the sequence to be encoded is an N-ary sequence; the video information includes multimedia information and meta information;
[0013] a mapping unit, configured to encode each of the sequences to be encoded into a plurality of base sequences corresponding to each of the video segments based on a bit-base mapping rule;
[0014] A synthesis unit is used to add corresponding preset primers to the multiple base sequences, perform DNA synthesis on each base sequence after adding the primers and store it to obtain the DNA storage data of the target video; the preset primers are related to the position of the video clip corresponding to the base sequence in the target video.
[0015] In a third aspect, the present invention further provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0016] In a fourth aspect, the present invention further provides a computer-readable storage medium having program data stored thereon, and the program data implements the above method when executed by a processor. Beneficial effects
[0017] Compared with the prior art, the present invention realizes fast encoding of the target video by segmenting the target video to be stored and using parallel encoding. On this basis, the adopted bit-base mapping rule can further improve the encoding efficiency of the file, and the encoding density under the mapping rule is also very high, thereby ensuring the encoding density and improving the encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] FIG1 is a schematic diagram of the main flow of an embodiment of a method for encoding and decoding information stored in DNA for video data provided by the present invention;
[0020] FIG2 is a schematic diagram of the main flow of another embodiment of a method for encoding and decoding information stored in DNA for video data provided by the present invention;
[0021] FIG3 is a schematic diagram of a sub-process of the embodiment shown in FIG2 ;
[0022] FIG4 is a schematic diagram of a sub-process of the embodiment shown in FIG1 ;
[0023] FIG5 is a schematic diagram of a sub-process of the embodiment shown in FIG4 ;
[0024] FIG6 is a schematic diagram of a sub-process of the embodiment shown in FIG5 ;
[0025] FIG7 is an example diagram of the first mapping table in the embodiment shown in FIG1 ;
[0026] FIG8 is an example diagram of a second mapping table in the embodiment shown in FIG2 ;
[0027] FIG9 is an example diagram of the third mapping table in the embodiment shown in FIG2 ;
[0028] FIG10 is a schematic diagram of the process of the bit-base mapping rule in the embodiment shown in FIG1 ;
[0029] FIG11 is a structural block diagram of a DNA storage encoding and decoding device for video data provided by the present invention;
[0030] FIG12 is a structural block diagram of the electronic device provided by the present invention. Best Mode for Carrying Out the Invention
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0032] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0033] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0034] 1 to 6 , in a first aspect, the present invention provides a DNA storage and encoding method for video data. As shown in FIG1 , an embodiment of a DNA storage and encoding method for video data is shown, and the method includes steps S100 to S400 .
[0035] S100: Split a target video to obtain multiple video segments of the target video.
[0036] In this embodiment, the target video is a file to be stored. The target video file is split into several video segments of varying or equal lengths. To efficiently divide the video into equal segments, the ffmpeg video processing tool is used to extract the segments. The segment parameter is used for segmentation while retaining timestamps. This facilitates integration of the video files during decoding. Segmenting the file facilitates subsequent segmented encoding, which improves encoding efficiency.
[0037] S200. Parse each of the video clips to obtain video information of the video clip, and convert the video information of each of the video clips into a corresponding sequence to be encoded, wherein the sequence to be encoded is an N-ary sequence; wherein the video information includes multimedia information and meta information.
[0038] In this embodiment, the original MP4 video is segmented into multiple video segments of equal or varying lengths. Multimedia information and meta-information are then extracted from these segmented video segments, specifically, the media information and specific video-related parameter information within the video file, such as encoding format, resolution, frame rate, and duration. Furthermore, the acquired multimedia information and meta-information for each video segment is converted into a binary sequence and stored. This binary sequence is then pre-encoded into a hexadecimal sequence to be encoded. In some embodiments, this multimedia information and meta-information can be directly converted into a hexadecimal sequence to be encoded, and this hexadecimal sequence can also be replaced with another hexadecimal sequence for subsequent encoding.
[0039] S300 : Based on a bit-base mapping rule, encode each of the sequences to be encoded into a plurality of base sequences corresponding to each of the video segments.
[0040] In this embodiment, the hexadecimal sequence to be encoded is converted into a base sequence based on the bit-base mapping rule. The base sequence is composed of base triplets, which are the basic unit of DNA information storage. The encoding rules based on the bit-base mapping rule have higher encoding efficiency.
[0041] S400. Add corresponding preset primers to the plurality of base sequences, perform DNA synthesis on each of the base sequences after adding the primers, and store the DNA to obtain DNA storage data of the target video; wherein the preset primers are related to the positions of the video segments corresponding to the base sequences in the target video.
[0042] In this embodiment, primers for identifying the positions of the video segments are added to the front and back ends of the base sequences of each acquired video segment. The base sequences after primer addition are then arranged in sequence to form a DNA sequence. In certain embodiments, such as this embodiment, the specific design principles of the primers are as follows: a length of 5 to 30 base pairs, a ratio of 40% to 60% for G and C bases, no more than three single-base repeats, and a melting temperature of 58-70 degrees Celsius, with a difference of no more than 4 degrees Celsius. Homology comparison is performed using blast software. The same pair of primers is used for base sequences within the same video segment, while different primer pairs are used for different video segments. This primer pair is used for retrieval during sequencing during the DNA sequence decoding process. The retrieval primers can directly read each video segment, enabling segmented reading of the video file, improving the efficiency and flexibility of reading. Modes for Carrying Out the Invention
[0043] In an exemplary embodiment, as shown in FIG2 , step S400 further includes step S500 : decoding the DNA storage data of the target video to obtain a target video segment;
[0044] Specifically, as shown in FIG3 , step S500 specifically includes the following steps S501 to S503:
[0045] S501: Sequence the DNA storage data of the target video to obtain the base sequence of the target video.
[0046] DNA sequencing is a technology that determines the sequence of DNA molecules and is used to obtain the base sequence of the target video.
[0047] S502: Acquire a base sequence of a target video segment under the target video according to the preset primers.
[0048] The base sequence of the target video segment under the target video can be obtained by searching the preset primers.
[0049] S503: Decode the base sequence of the target video segment based on the bit-base mapping rule to obtain each target video segment.
[0050] Based on the bit-base mapping rule, by decoding the base sequence of the target video clip, the video information of the target video clip can be obtained, by reverse parsing, the target video clip can be obtained, and by obtaining all video clips and combining them, the original target video can be further obtained.
[0051] It can be understood that steps S501-S503 are used for decoding the DNA storage data. The decoding process can be regarded as the inverse process of the above-mentioned encoding method. Based on the segmented storage of DNA storage data, the target video segment under the target video can be quickly read by retrieving the preset primers, thereby improving the reading efficiency and flexibility.
[0052] In an exemplary embodiment, as shown in FIG4 , step S300 specifically includes steps S310 - S320:
[0053] S310, segmenting each of the sequences to be encoded to obtain multiple subsequences to be encoded, and adding an index to each of the subsequences to be encoded;
[0054] In this embodiment, the sequence to be encoded is segmented. Optionally, the sequence to be encoded is divided into multiple subsequences to be encoded, each of which is 4 bits. For example, the hexadecimal sequence to be encoded is "eb3f27fa2676b0cc", which is segmented to obtain four subsequences to be encoded, "eb3f", "27fa", "2676", and "b0cc". An index is added before each subsequence to identify the position information of each subsequence to be encoded.
[0055] S320 : Perform parallel encoding on the subsequences to be encoded based on a bit-base mapping rule to obtain a plurality of base sequences corresponding to the video segments.
[0056] The coding subsequences obtained through the above steps are encoded in parallel to obtain multiple base sequences of the video segment, where the base sequences are composed of base triplets.
[0057] After the sequence to be encoded is segmented, parallel encoding of the sequence to be encoded can be achieved, further improving the encoding efficiency.
[0058] In an exemplary embodiment, as shown in FIG5 , step S320 specifically includes steps S321-S322
[0059] S321. Based on the index of each subsequence to be encoded, convert each N-ary subsequence to be encoded into three M-ary numbers by performing a remainder operation; and form an M-ary sequence from the multiple M-ary numbers.
[0060] After converting each N-ary subsequence to be encoded into a decimal number, the modulo calculation is performed on the value M three times in succession to obtain three M-ary numbers, where the value of M depends on the specific bit-base mapping rule.
[0061] In an optional embodiment, at least one RS error correction code is added to the M-ary sequence; specifically, several RS error correction codes are added at the end of the M-ary sequence. The RS error correction code is used to detect and ensure the integrity and accuracy of the M-ary sequence. The error correction capability of the RS error correction code depends on the number of RS error correction codes.
[0062] S322: Determine, based on a preset mapping table, a base triplet corresponding to each M-ary number in the M-ary sequence, so as to convert the M-ary sequence into the multiple base sequences.
[0063] The elements of the preset mapping table are different base triplets. Each element in the M-ary sequence is converted into a corresponding base triplet through the preset mapping table, and the M-ary sequence is finally converted into a base sequence composed of base triplets.
[0064] In an exemplary embodiment, as shown in FIG6 , step S322 specifically includes steps S3221 - S3223:
[0065] S3221. Calculate the median of the M-ary sequence and use the median as an offset value.
[0066] S3222: Selecting a corresponding base triplet in the preset mapping table according to the offset value and each M-ary number in the M-ary sequence;
[0067] S3223. Arrange the selected base triplets in sequence to form the base sequence.
[0068] In a specific implementation of this embodiment, the preset mapping table includes a first mapping table, a second mapping table, a third mapping table, and a fourth mapping table. The second mapping table and the third mapping table each include (M-1) / 2 base triplets, and the first mapping table includes 48-(M-1) base triplets. The first mapping table includes 4 odd-numbered triplets and 4 even-numbered triplets. The base triplets included in the first mapping table, the second mapping table, and the third mapping table are all different. The fourth mapping table is formed by alternating combinations of elements in the second mapping table and the third mapping table, plus a random base triplet.
[0069] The determining, according to the offset value and the element value of each element in the M-ary sequence, the base triplet in the preset mapping table includes:
[0070] randomly selecting a first base triplet, a second base triplet, a third base triplet, and a fourth base triplet from the first mapping table, wherein the first base triplet and the third base triplet are odd-numbered triplets, and the third base triplet and the fourth base triplet are even-numbered triplets;
[0071] When offset≠0, if n=0, the first base triplet and the third base triplet are selected alternately; if n=M-1, the second base triplet and the fourth base triplet are selected alternately;
[0072] When offset < (M-1) / 2, if offset > n ≥ 0, then select the n-th base triplet in the second mapping table; if M-2 > n ≥ offset + 20, then select the n-20th base triplet in the second mapping table; otherwise, select the n-offseth base triplet in the third mapping table;
[0073] When offset>(M-1) / 2, if M-2>n≥offset, then select the n-20th base triplet in the third mapping table; if offset-20>n≥0, then select the nth base triplet in the third mapping table; otherwise, extract the offset-n-1th base triplet in the second mapping table;
[0074] When offset=(M-1) / 2, if (M-1) / 2>n≥0, the n-th base triplet in the second mapping table is selected; otherwise, the n-20-th base triplet in the third mapping table is selected;
[0075] When offset=0, if n=0, the third base triplet and the fourth base triplet are selected alternately; if n<(M-1) / 2, the nth base triplet in the second mapping table is selected; otherwise, the n-20th base triplet in the third mapping table is selected;
[0076] Finally, the offset-th base triplet is selected from the fourth mapping table.
[0077] Wherein, n is the element value, M-1≥n≥0, n is an integer, offset is the offset value, and M is the base number of the above sequence, M=41.
[0078] Specifically, as shown in FIG7 , the first mapping table includes the following elements: odd-numbered bit triplets: ATA, TTA, AAT, TAT; and even-numbered bit triplets: CGC, GGC, CCG, GCG.
[0079] As shown in FIG8 , the second mapping table includes the following elements: ACA, TCA, AGA, TGA, CTA
[0080] , GTA, AAC, TAC, ATC, TTC, AAG, TAG, ATG, TTG, CAT, GAT, ACT, TCT, ACT, TGT.
[0081] As shown in Figure 9, the third mapping table includes the following elements: CCA, GCA, CGA, GGA, CAC, GAC, AGC, TGC, CTC, GTC, CAG, GAG, ACG, TCG, CTG, GTG, CCT, GCT, CGT, GGT.
[0082] The elements included in the fourth mapping table are as follows: ACA, CCA, TCA, GCA, AGA, CGA, TGA, GGA, CTA, CAC, GTA, GAC, AAC, AGC, TAC, TGC, ATC, CTC, TTC, GTC, AAG, CAG, TAG, GAG, ATG, ACG, TTG, TCG, CAT, CTG, GAT, GTG, ACT, CCT, TCT, GCT, AGT, CGT, TGT, GGT, ATA.
[0083] Understandably, in order to efficiently encode and decode, and to obtain base sequences that conform to DNA biochemical constraints, the above-mentioned mapping rules are designed from the perspective of single-base repeats and GC content, with four different mapping tables. The last two bases of each triplet in the second and third mapping tables are different, so that the maximum number of repeats of two triplets connected together does not exceed 3. There are a total of 48 base triplets that meet this requirement. The second and third mapping tables each have 20 base triplets, so the first mapping table includes the remaining 8 base triplets. The three mapping tables include 48 base triplets that meet the requirements. Each mapping table includes different base triplets. The base sequence obtained under this mapping rule can effectively avoid problems such as single-base repeats and GC content imbalance.
[0084] In order to better understand the encoding and decoding method of the present invention, FIG10 shows the complete mapping process under the bit-base mapping rule of an embodiment of the present invention. First, the hexadecimal sequence to be encoded "eb3f27fa2676b0cc" is obtained, and the hexadecimal sequence to be encoded is segmented to obtain four subsequences to be encoded "eb3f", "27fa", "2676", and "b0cc". After the subsequences to be encoded are converted into decimal numbers, 41 is modulo 41 for three consecutive times to obtain a forty-first order sequence with three remainders as a group, "35, 33, 35", "25 , 3, 6", "6, 35, 5", "37, 37, 26"; an RS error correction code is added to the obtained 4-undecimal sequence to correct the 4-undecimal sequence; it is calculated that the median of the 4-undecimal sequence is 30, that is, the offset value offset is set to 30, and based on the above-mentioned bit-base mapping rule, the 4-undecimal sequence is converted into the following base sequence according to the preset mapping table: "GTG, TCG, GTG, CTA, GGA, AGC, AGC, GTG, GAC, GCT, GCT, TGA, TAT, CCT, TCA, GAT".
[0085] The DNA storage encoding and decoding method of video data of the present invention realizes rapid encoding of the target video by segmenting the target video to be stored and using parallel encoding. On this basis, the bit-base mapping rule adopted can also further improve the encoding efficiency of the file, and the encoding density under the mapping rule is also very high, thereby ensuring the encoding density and improving the encoding efficiency.
[0086] As shown in FIG11 , an embodiment of the present invention further provides a DNA storage encoding and decoding device 100, which specifically includes a splitting unit 110, a parsing unit 120, a mapping unit 130, and a synthesis unit 140. The splitting unit is used to split a target video to obtain multiple video segments of the target video; the parsing unit is used to parse each of the video segments to obtain video information of the video segments, and convert the video information of each of the video segments into a corresponding sequence to be encoded, wherein the sequence to be encoded is an N-ary sequence; the video information includes multimedia information and meta-information; the mapping unit is used to encode each of the sequences to be encoded into multiple base sequences corresponding to each of the video segments based on a bit-base mapping rule; the synthesis unit is used to add corresponding preset primers to the multiple base sequences, perform DNA synthesis on each of the base sequences after the primers are added, and store them to obtain DNA storage data of the target video; the preset primers are related to the position of the video segment corresponding to the base sequence in the target video.
[0087] In an optional embodiment, the splitting unit 110 is further configured to combine the split video segments to obtain the original target video; the parsing unit 120 is further configured to parse the video information of the video segments to obtain the video segments; and the mapping unit 130 is further configured to decode the base sequence of the target video to obtain the target video's to-be-encoded sequence. Furthermore, the apparatus further includes a sequencing unit configured to sequence the DNA stored data of the target video to obtain the base sequence of the target video segment.
[0088] In an optional embodiment, the mapping unit 130 further includes a segmentation unit and an encoding unit, the segmentation unit being used to segment each of the sequences to be encoded to obtain a plurality of subsequences to be encoded, and adding an index to each of the subsequences to be encoded; the encoding unit being used to perform parallel encoding on each of the subsequences to be encoded based on a bit-base mapping rule to obtain a plurality of base sequences corresponding to each of the video clips.
[0089] In an optional embodiment, the encoding unit includes a conversion unit and a mapping unit, the conversion unit being used to convert each N-ary subsequence to be encoded into 3 M-ary numbers through a remainder operation based on the index of each subsequence to be encoded; an M-ary sequence is composed of multiple M-ary numbers; and the mapping unit is used to determine the base triplet corresponding to each M-ary number in the M-ary sequence based on a preset mapping table, so that the M-ary sequence is converted into the multiple base sequences.
[0090] In an optional embodiment, the encoding unit further includes an error correction unit, and the error correction unit is used to add at least one RS error correction code to the M-ary sequence.
[0091] In an optional embodiment, the error correction unit includes a calculation unit, a selection unit and a combination unit, the calculation unit is used to calculate the median of the M-ary sequence and use the median as an offset value; the selection unit is used to select the corresponding base triplets in the preset mapping table based on the offset value and each M-ary number in the M-ary sequence; the combination unit is used to arrange the selected base triplets in sequence to form the base sequence.
[0092] As shown in FIG12 , an embodiment of the present invention provides an electronic device 140, which includes a processor 141, a memory 142, a non-volatile memory 144, and a computer program 1441 stored in the non-volatile memory 144 and executable on the processor. When the processor 141 executes the computer program 1441, any embodiment of the current measurement method described above is implemented. Specifically, the electronic device 140 also includes an input / output interface 145 and an input / output device 146 connected thereto. The processor 141, the memory 142, the non-volatile memory 144, and the input / output interface 145 are connected via an internal bus 143.
[0093] An embodiment of the present invention provides a computer-readable storage medium. When instructions in the storage medium are executed by a processor of a terminal, the terminal is enabled to perform the above-mentioned encoding and decoding method.
[0094] It should be understood that in the embodiments of the present invention, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0095] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0096] It will be understood that the present invention is described by way of some embodiments, and it will be appreciated by those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present invention are intended to be protected by the present invention.
Claims
1. A DNA storage encoding and decoding method for video data, characterized in that, it includes: Splitting the target video to obtain multiple video segments of the target video; Parsing each of the video segments to obtain the video information of the video segment, and converting the video information of each video segment into a corresponding sequence to be encoded, where the sequence to be encoded is an N-ary sequence; the video information includes multimedia information and meta-information; Based on the bit-base mapping rule, encoding each of the sequences to be encoded into multiple base sequences respectively corresponding to each video segment; Adding corresponding preset primers to the multiple base sequences, performing DNA synthesis and storage on each of the base sequences after adding primers to obtain the DNA storage data of the target video; The preset primer is related to the position of the video segment corresponding to the base sequence in the target video.
2. The method according to claim 1, characterized in that, After obtaining the DNA storage data of the target video, the method further includes decoding the DNA storage data of the target video to obtain target video segments; The decoding the DNA storage data of the target video to obtain target video segments includes: Sequencing the DNA storage data of the target video to obtain the base sequence of the target video; According to the preset primer, obtaining the base sequence of the target video segment under the target video; Based on the bit-base mapping rule, decoding the base sequence of the target video segment to obtain each target video segment.
3. The method according to claim 1, characterized in that, The encoding each of the sequences to be encoded into multiple base sequences respectively corresponding to each video segment based on the bit-base mapping rule includes: Segmenting each of the sequences to be encoded respectively to obtain multiple subsequences to be encoded, and adding indexes to each of the subsequences to be encoded; Parallel encoding each of the subsequences to be encoded based on the bit-base mapping rule to obtain multiple base sequences respectively corresponding to each video segment.
4. The method according to claim 3, characterized in that, The parallel encoding each of the subsequences to be encoded based on the bit-base mapping rule to obtain multiple base sequences respectively corresponding to each video segment includes: Based on the indexes of each of the subsequences to be encoded, converting each N-ary subsequence to be encoded into 3 M-ary numbers through a modulo operation; forming an M-ary sequence from multiple M-ary numbers; Determining the base triplets corresponding to each M-ary number in the M-ary sequence based on a preset mapping table to convert the M-ary sequence into the multiple base sequences.
5. The method according to claim 4, characterized in that, After forming the M-ary sequence from multiple M-ary numbers, the method further includes: Adding at least one RS error correction code to the M-ary sequence.
6. The method according to claim 4, characterized in that, Determining the base triplets corresponding to each M-ary number in the M-ary sequence based on a preset mapping table, so as to convert the M-ary sequence into the plurality of base sequences, includes: Calculating the median of the M-ary sequence, and using the median as the offset value; According to the offset value and each M-ary number in the M-ary sequence, selecting the corresponding base triplets in the preset mapping table; Arranging the selected base triplets in sequence to form the base sequence.
7. The method according to claim 6, characterized in that the preset mapping table includes a first mapping table, a second mapping table, a third mapping table and a fourth mapping table. The second mapping table and the third mapping table each include (M - 1) / 2 base triplets, and the first mapping table includes 48 - (M - 1) base triplets. The first mapping table includes 4 odd-position triplets and 4 even-position triplets. The base triplets included in the first mapping table, the second mapping table and the third mapping table are all different. The fourth mapping table is formed by alternately combining the elements in the second mapping table and the third mapping table plus a random base triplet; Determining the base triplets in the preset mapping table according to the offset value and the element values of each element in the M-ary sequence, includes: Randomly selecting a first base triplet, a second base triplet, a third base triplet and a fourth base triplet from the first mapping table, wherein the first base triplet and the third base triplet belong to odd-position triplets, and the third base triplet and the fourth base triplet belong to even-position triplets; When offset≠0, if n = 0, then alternately select the first base triplet and the third base triplet; if n = M - 1, then alternately select the second base triplet and the fourth base triplet; When offset < (M - 1) / 2, if offset > n ≥ 0, then select the nth base triplet in the second mapping table; if M - 2 > n ≥ offset + 20, then select the (n - 20)th base triplet in the second mapping table; otherwise select the (n - offset)th base triplet in the third mapping table; When offset > (M - 1) / 2, if M - 2 > n ≥ offset, then select the (n - 20)th base triplet in the third mapping table; if offset - 20 > n ≥ 0, then select the nth base triplet in the third mapping table; otherwise extract the (offset - n - 1)th base triplet in the second mapping table; When offset = (M - 1) / 2, if (M - 1) / 2 > n ≥ 0, then select the nth base triplet in the second mapping table; otherwise select the (n - 20)th base triplet in the third mapping table; When offset = 0, if n = 0, then alternately select the third base triplet and the fourth base triplet; if n < (M - 1) / 2, then select the nth base triplet in the second mapping table; otherwise select the (n - 20)th base triplet in the third mapping table; Select the offsetth base triplet from the fourth mapping table; where n is the element value, 0 ≤ n ≤ M - 1, n is an integer, offset is the offset value, M is the base number of the above sequence, and M = 41.
8. A DNA storage encoding and decoding device for video data, characterized in that, it includes: a splitting unit for splitting a target video to obtain a plurality of video segments of the target video; an analysis unit for analyzing each of the video segments to obtain video information of the video segment, and converting the video information of each of the video segments into a corresponding sequence to be encoded, where the sequence to be encoded is an N - ary sequence; the video information includes multimedia information and meta - information; a mapping unit for encoding each of the sequences to be encoded into a plurality of base sequences respectively corresponding to each of the video segments based on a bit - base mapping rule; a synthesis unit for adding corresponding preset primers to the plurality of base sequences, performing DNA synthesis and storage on each of the base sequences after adding the primers to obtain DNA storage data of the target video; The preset primer is related to the position of the video segment corresponding to the base sequence in the target video.
9. An electronic device, characterized in that, it includes a memory and a processor, a computer program is stored on the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer - readable storage medium, characterized in that, program data is stored on the storage medium, and when the program data is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Encoding / decoding methods, apparatus and data processing apparatus
CN111095423B
DNA storage-oriented DNA sequence processing method and device and electronic equipment
CN113611364A
Encoding and decoding method and encoding and decoding device between binary information and base sequence for DNA data storage
WO2022109879A1
Apparatus and methods for embedding data in genetic material
WO2023129469A1