Data storage and retrieval method, device and medium using nucleic acid molecules
By generating alternative nucleic acid molecular address sequences that meet biological stability requirements and replacing the initial sequence, the problem of unstable nucleic acid molecular address sequences is solved, achieving more stable and efficient nucleic acid molecular data storage and retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-03-31
AI Technical Summary
Existing nucleic acid molecular address sequences are insufficient in terms of biological stability requirements, resulting in the inability to generate or the generation of unstable nucleic acid molecules, which affects the stability and effectiveness of data storage.
Multiple alternative nucleic acid molecule address sequences that meet biological stability requirements are generated, and the initial nucleic acid molecule address sequence is replaced by calculating similarity. Address molecule modules are pre-synthesized to improve generation efficiency and stability.
It improves the stability and effectiveness of nucleic acid molecular data storage, ensuring that the generated nucleic acid molecules are more stable under biological stability requirements, and supports efficient data storage and retrieval.
Smart Images

Figure CN117059178B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data storage, and in particular to data storage utilizing nucleic acid molecules. Background Technology
[0002] With the development of data storage technology, technologies that utilize nucleic acid molecules (including deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)) to store information have been proposed, with the aim of achieving higher storage density, longer storage life, and more reliable storage performance.
[0003] Nucleic acid molecules are long-chain polymers. DNA is composed of bases, deoxyribose, and phosphate. The bases include four bases: adenine (A), guanine (G), thymine (T), and cytosine (C), as well as other non-natural bases, which are arranged in a base sequence along the long DNA chain. RNA is composed of phosphate, ribose, and bases. The bases include four bases: adenine (A), guanine (G), cytosine (C), and uracil (U), as well as other non-natural bases, which are also arranged in a base sequence.
[0004] In nucleic acid molecule data storage technology, data can be encoded into base sequences, and nucleic acid molecules containing the corresponding base sequences can be obtained based on biotechnology, thereby realizing the storage of data in nucleic acid molecules. Summary of the Invention
[0005] Multiple nucleic acid molecules storing data constitute a nucleic acid molecule repository. To facilitate data retrieval within this repository, a nucleic acid molecule address sequence is typically assigned to each molecule for addressing. Taking DNA storage as an example, a DNA sequence can be designated as the DNA molecule address sequence, which is then used to synthesize a DNA molecule together with the corresponding content's DNA sequence.
[0006] In the data retrieval process, molecular address sequences are used as DNA molecular probes. DNA molecules that match the probes are searched in the nucleic acid molecular library through DNA molecular hybridization. The corresponding contents of the DNA molecules are then read, thereby achieving data retrieval.
[0007] Nucleic acid molecule address sequences can be set using various methods. However, in some cases, the set nucleic acid molecule address sequences may not meet the biological stability requirements of nucleic acid molecules, resulting in the inability to synthesize the corresponding nucleic acid molecules based on the nucleic acid molecule address sequences, or the synthesized nucleic acid molecules may be unstable.
[0008] This disclosure provides a method for generating nucleic acid molecule address sequences, which generates corresponding nucleic acid molecule address sequences under the condition of considering biological stability requirements, so that nucleic acid molecules generated based on the generated nucleic acid molecule address sequences are more stable, thereby improving the stability, effectiveness and practicality of nucleic acid molecule-based data storage.
[0009] Alternatively, for the generated nucleic acid molecule address sequence, a corresponding address molecule module can be pre-synthesized and stored, for example, in an address molecule module repository. Thus, in subsequent nucleic acid molecule generation processes, it is unnecessary to resynthesize the nucleic acid molecule address sequence; instead, the pre-synthesized address molecule module can be directly utilized, thereby improving the efficiency of nucleic acid molecule generation.
[0010] According to one aspect of this disclosure, a method for generating nucleic acid molecular address sequences is provided, comprising: generating a plurality of candidate nucleic acid molecular address sequences that meet the requirements of biological stability; and calculating the similarity between an initial nucleic acid molecular address sequence of input data and the plurality of candidate nucleic acid molecular address sequences, and replacing the initial nucleic acid molecular address sequence with the nucleic acid molecular address sequence among the plurality of candidate nucleic acid molecular address sequences that has the highest similarity to the initial nucleic acid molecular address sequence, as the nucleic acid molecular address sequence of the input data.
[0011] According to another aspect of this disclosure, a method for storing information in a nucleic acid molecule is provided, comprising: generating a nucleic acid molecule address sequence of input data to be stored in the nucleic acid molecule using a method for generating a nucleic acid molecule address sequence according to this disclosure; pre-synthesizing a plurality of address molecule modules for the plurality of candidate nucleic acid molecule address sequences; determining an address molecule module from the plurality of address molecule modules that corresponds to the nucleic acid molecule address sequence of the input data; and generating a nucleic acid molecule based on the address molecule module corresponding to the nucleic acid molecule address sequence of the input data and a content molecule module of the input data, so as to store the input data in the nucleic acid molecule.
[0012] According to another aspect of this disclosure, a method for retrieving information stored in nucleic acid molecules is provided, comprising: storing multiple input data into multiple nucleic acid molecules using the method for storing information in nucleic acid molecules according to this disclosure to generate a nucleic acid molecule repository including the multiple nucleic acid molecules; using the nucleic acid molecule address sequence of the input data in the multiple input data as a probe, retrieving one or more nucleic acid molecule sequences matching the probe in the nucleic acid molecule repository by molecular hybridization; and extracting the information stored in the retrieved one or more nucleic acid molecule sequences as the result of the retrieval.
[0013] According to another aspect of this disclosure, an apparatus for generating nucleic acid molecular address sequences is provided, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored in the memory to perform the following processes: generating a plurality of candidate nucleic acid molecular address sequences that meet biological stability requirements; and calculating the similarity between an initial nucleic acid molecular address sequence of input data and the plurality of candidate nucleic acid molecular address sequences, and replacing the initial nucleic acid molecular address sequence with the nucleic acid molecular address sequence among the plurality of candidate nucleic acid molecular address sequences that has the highest similarity to the initial nucleic acid molecular address sequence, as the nucleic acid molecular address sequence of the input data.
[0014] According to another aspect of this disclosure, an apparatus for storing information in a nucleic acid molecule is provided, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored in the memory to perform the following processes: generating a nucleic acid molecule address sequence of input data to be stored in the nucleic acid molecule using an apparatus for generating a nucleic acid molecule address sequence according to this disclosure; pre-synthesizing a plurality of address molecule modules for the plurality of candidate nucleic acid molecule address sequences; determining an address molecule module from the plurality of address molecule modules that corresponds to the nucleic acid molecule address sequence of the input data; and generating a nucleic acid molecule based on the address molecule module that corresponds to the nucleic acid molecule address sequence of the input data and a content molecule module of the input data to store the input data in the nucleic acid molecule.
[0015] According to another aspect of this disclosure, an apparatus for retrieving information stored in nucleic acid molecules is provided, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored in the memory to perform the following processes: storing a plurality of input data into a plurality of nucleic acid molecules using the apparatus for storing information in nucleic acid molecules according to this disclosure to generate a nucleic acid molecule repository including the plurality of nucleic acid molecules; using a nucleic acid molecule address sequence of the input data in the plurality of input data as a probe, retrieving one or more nucleic acid molecule sequences matching the probe in the nucleic acid molecule repository by molecular hybridization; and extracting information stored in the retrieved one or more nucleic acid molecule sequences as the result of the retrieval.
[0016] According to another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to perform a method for generating a nucleic acid molecule address sequence according to this disclosure. Attached Figure Description
[0017] The present disclosure will now be described in detail below with reference to the accompanying drawings, wherein the same reference numerals throughout the drawings denote the same or similar components. It should be understood that the drawings are not necessarily drawn to scale and are only used to illustrate exemplary embodiments of the present disclosure and should not be considered as limiting the scope of the disclosure. Wherein:
[0018] Figure 1 An exemplary configuration block diagram of an apparatus for generating nucleic acid molecule address sequences according to embodiments of the present disclosure is shown;
[0019] Figure 2 An exemplary flowchart of a method for generating nucleic acid molecule address sequences according to embodiments of the present disclosure is shown;
[0020] Figure 3 An exemplary flowchart of a method for generating a plurality of alternative nucleic acid molecular address sequences that meet biological stability requirements, according to embodiments of the present disclosure, is shown.
[0021] Figure 4 An exemplary flowchart of a method for generating nucleic acid molecule address sequences that retain semantic information of input data, according to embodiments of the present disclosure, is shown.
[0022] Figure 5 An exemplary flowchart of a method for generating an initial nucleic acid molecule sequence that retains semantic information of input data, according to embodiments of the present disclosure, is shown.
[0023] Figure 6 A schematic diagram of a method for storing information in a nucleic acid molecule according to an embodiment of the present disclosure is shown;
[0024] Figure 7 A schematic diagram of a method for retrieving information stored in nucleic acid molecules according to an embodiment of the present disclosure is shown;
[0025] Figure 8 An exemplary configuration of a computing device that can implement an embodiment of the present invention is shown. Detailed Implementation
[0026] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. However, it should be understood that the descriptions of the various exemplary embodiments are merely illustrative and are not intended to limit the technology of the present disclosure in any way. Unless otherwise specifically stated, the relative arrangement, expressions, and values of components and steps in the exemplary embodiments do not limit the scope of the present disclosure.
[0027] Figure 1 An exemplary configuration block diagram of an apparatus 1000 for generating nucleic acid molecular address sequences according to an embodiment of the present disclosure is shown.
[0028] like Figure 1 As shown, in some embodiments, device 1000 may include processor 1010. The processor 1010 of device 1000 provides various functions of device 1000. In some embodiments, the processor 1010 of device 1000 may be configured to perform the method 2000 of this disclosure for generating nucleic acid molecular address sequences (hereinafter referred to as...). Figure 2 (To be described). Specifically, such as Figure 1 As shown, the processor 1010 of the device 1000 may include a candidate sequence generation unit 1020 and a sequence replacement unit 1030, which are respectively configured to perform the following description. Figure 2 Steps S2010 and S2020 in method 2000 shown. Additionally, processor 1010 may also include other units (not shown) configured to perform the following. Figure 3-6 One or more of the methods shown.
[0029] It should be understood that Figure 1 The units of the illustrated device 1000 are merely logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. In actual implementation, the above modules can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.).
[0030] The processor 1010 of device 1000 can refer to various implementations of digital circuit systems, analog circuit systems, or mixed-signal (combination of analog and digital) circuit systems that perform functions in a computing system. The processing circuitry can include, for example, circuits such as integrated circuits (ICs), application-specific integrated circuits (ASICs), portions or circuits of a single processor core, the entire processor core, a single processor, programmable hardware devices such as field-programmable gate arrays (FPGAs), and / or systems that include multiple processors.
[0031] In some embodiments, device 1000 may further include a memory (not shown). The memory of device 1000 may store information generated by processor 1010, as well as programs and data for operation of processor 1010. The memory may be volatile memory and / or non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. In addition, device 1000 may be implemented at the chip level, or it may be implemented at the device level by including other external components.
[0032] Figure 2An exemplary flowchart of a method 2000 for generating nucleic acid molecule address sequences according to embodiments of the present disclosure is shown. This method can be used, for example, as... Figure 1 The device shown is 1000.
[0033] The nucleic acid molecules described in this disclosure include DNA molecules and RNA molecules. The base sequence of DNA molecules includes the four bases AGCT and other non-natural bases, while the base sequence of RNA molecules includes the four bases AGCU and other non-natural bases. In the following description, embodiments according to this disclosure are described using DNA molecules and their AGCT base sequences as examples. It should be understood that the embodiments of this disclosure can also be similarly applied to address sequence generation and storage technologies based on RNA molecules. Furthermore, the bases used are not limited to natural bases such as AGCT or AGCU; non-natural bases can also be used as needed. Additionally, in this disclosure, a nucleic acid molecule address sequence refers to a sequence of nucleic acid molecules used for addressing nucleic acid molecules; the length of the nucleic acid molecule address sequence can be preset as needed.
[0034] like Figure 2 As shown, in S2010, the candidate sequence generation unit 1020 generates multiple candidate nucleic acid molecule address sequences that meet the requirements of biological stability.
[0035] As mentioned above, in data storage using DNA molecules, the initial DNA molecule address sequence set for the input data may not meet the biological stability requirements of the DNA molecules, resulting in the inability to generate the corresponding DNA molecules or the generated DNA molecules being unstable. Therefore, in this disclosure, multiple alternative nucleic acid molecule address sequences that meet the biological stability requirements are pre-generated for replacement of the initial DNA molecule sequence in S2020 described later.
[0036] To ensure the biological stability of DNA sequences, the following constraints can be considered.
[0037] Continuity constraints
[0038] Continuity constraints refer to avoiding the repeated occurrence of the same base in the base sequence of a DNA molecule. In some embodiments, this can be achieved using the following continuity constraint evaluation function f. continuity (x) is calculated and constrained to make its value less than a pre-set threshold.
[0039]
[0040] Where x represents the DNA molecule's base sequence, and |x| represents the sequence length. T(a, T) value ) represents the threshold function, i.e., when a>T value Take 'a' if the condition is met, otherwise take 0.con The minimum counting threshold is 1 ≤ T con ≤n. cont b (x, i) represents the calculation for the longest consecutive base. When x i =x i-1 And x i ≠x i+1 When the time is reached, the length of the continuously counted bases is returned.
[0041] As an example, a single base can be defined as the smallest unit, and the same base unit can appear consecutively no more than 3 times. For example, in a base sequence like "......TTTTT......", the base T appears consecutively 5 times, thus failing the continuity constraint requirement.
[0042] As another example, we can define two bases as the smallest unit, and the same base unit can appear consecutively no more than three times. For example, in a base sequence like "...ACACACAC...", the base unit "AC" appears four times consecutively, thus violating the continuity constraint.
[0043] As another example, we can define three bases as the smallest unit, and the same base unit can appear consecutively no more than 3 times. For example, in a base sequence like "...ACTACT...", the base unit "ACT" appears twice consecutively, thus meeting the continuity constraint requirement.
[0044] Complexity constraints
[0045] Complexity constraints refer to the complexity resulting from the different base sequences within a DNA molecule. In some embodiments, this complexity constraint can be evaluated using the following complexity constraint evaluation function f. complexity (x) is used to calculate and constrain the value so that it is greater than a pre-set threshold.
[0046]
[0047] Where x represents the DNA molecule base sequence, |x| represents the sequence length, and A, T, C, and G represent the number of corresponding bases in the sequence.
[0048] GC content constraint
[0049] GC content constraint refers to the percentage of G and C bases in the total number of DNA bases in the DNA molecule's base sequence. In some embodiments, the GC content can be evaluated using the following GC content evaluation function f. GC (x) is used to perform calculations and constraints so that its value conforms to a given range.
[0050]
[0051] Where x represents the DNA molecule base sequence, |x| represents the sequence length, and G and C represent the number of corresponding bases in the sequence, respectively.
[0052] As an example, a global GC content ratio can be set, which is the ratio of the sum of the number of G and C cells in the DNA sequence to the total length. A given range can be set to [0.4, 0.6]. If the calculated global GC content ratio falls within this range, the GC content constraint is met; otherwise, it does not.
[0053] As another example, the proportion of local GC content can be set. For example, the constraints can be designed in stages based on experience: when the length of the sliding window is 10, the proportion of local GC satisfies [0.3, 0.7]; when the length of the sliding window is 8, the proportion of local GC satisfies [0.25, 0.75]; when the length of the sliding window is 6, the proportion of local GC satisfies [0.2, 0.8].
[0054] Secondary structural constraints
[0055] Secondary structure constraint (hairpin structure constraint) refers to the phenomenon that single-stranded DNA molecules may backfold to form a secondary structure resembling a hairpin, as shown below.
[0056]
[0057] Therefore, in some embodiments, the following secondary structure evaluation function f can be used. hairpin (x) is estimated and constrained to make its value less than a pre-set threshold.
[0058]
[0059] Where x represents the DNA molecule base sequence, s is the hairpin stem length, and S min R is the minimum stem length, r is the loop length of the hairpin, and R min The minimum ring length is T(a, T). value ) represents the threshold function, i.e., when a>T value Take 'a' if the x value is not 'a', otherwise take 0. cb(x) m x n Determine the two bases x m With base x n Whether they are complementary or not, if they are complementary, take 1; otherwise, take 0.
[0060] As an example, if bases A and T are complementary, and bases G and C are complementary, forming a series of consecutive anticomplementary sequences, such as the sequence "......ATGCA......TGCAT......" (where "ATGCA" and "TGCAT" are anticomplementary), then back folding may occur to form a secondary structure. Therefore, for example, it can be required that the longest consecutive anticomplementary sequence in the chain does not exceed 4 units in length, where the interval between consecutive anticomplementary sequences in a module is less than or equal to 2 units, which is acceptable. For example, in the sequence "......ATGCAACTGCAT......", the consecutive anticomplementary sequences "ATGCA" and "TGCAT" exceed 4 units in length, but their interval is 2 units (the interval is the base "AC"), thus meeting the secondary structure constraint requirement; in the sequence "......ATGCAATTCTGCAT......", the consecutive anticomplementary sequences "ATGCA" and "TGCAT" exceed 4 units in length, and their interval is 4 units (the interval is the base "ATTC"), thus not meeting the secondary structure constraint requirement.
[0061] Desolution temperature constraint
[0062] The melting temperature (Tm) refers to the temperature at which 50% of DNA molecules unwind and become single-stranded during the denaturation process. Factors affecting the melting temperature include the DNA base sequence, DNA concentration, and the pH of the solution. Empirical formulas exist for calculating the melting temperature (Tm), and in application, any method can be chosen to calculate Tm as needed. m (x) makes the melting temperature fall within a given range.
[0063] Empirical formulas based on sequence composition:
[0064] T m (x)=(A+T)×2℃+(C+G)×4℃ (5),
[0065] Empirical formula based on GC percentage content:
[0066]
[0067] Empirical formula based on nearest neighbor method:
[0068]
[0069] Where x represents the DNA molecule base sequence, and A, T, C, and G in equation (5) represent the number of bases in the sequence, respectively; |x| in equation (6) represents the sequence length, Na +The concentration of salt is generally taken as 50 mmol / L; ΔH°(x) in equation (7) represents the standard enthalpy change, and ΔS°(x) represents the standard entropy change. These two values can be found in the thermodynamic parameter table, for example. R represents the gas constant, with a value of 1.987 cal / kmol, and C T This represents the molar concentration of DNA molecules.
[0070] In some embodiments, in S2010, the candidate sequence generation unit 1020 can generate multiple candidate nucleic acid molecular address sequences based on the various biological stability constraints described above.
[0071] More specifically, the candidate sequence generation unit 1020 can generate multiple candidate nucleic acid molecule address sequences based on biological stability requirements defined by one or more of the following factors: GC content (corresponding to GC content constraints), indicating the percentage of C and G bases in the total number of bases in the nucleic acid molecule; single base continuity (corresponding to continuity constraints), indicating the number of consecutive occurrences of a single base in the nucleic acid molecule sequence; potential formation of secondary structures (corresponding to secondary structure constraints), indicating the possibility of backfolding the nucleic acid molecule sequence to form a secondary structure; nucleic acid molecule sequence complexity (corresponding to complexity constraints), indicating the complexity formed by different base sequence compositions; and melting temperature (corresponding to melting temperature constraints).
[0072] In addition, it should be understood that the requirements for biological stability are not limited to the constraints listed above. Those skilled in the art can design other constraints related to biological stability and generate multiple alternative nucleic acid molecule address sequences based on the corresponding constraints.
[0073] In this disclosure, the candidate nucleic acid molecular address sequences that meet the requirements of biological stability generated in S2010 are also referred to as "preferred sequences", and multiple candidate nucleic acid molecular address sequences are also referred to as a set of "preferred sequences".
[0074] Next, in S2020, the sequence replacement unit 1040 calculates the similarity between the initial nucleic acid molecule address sequence of the input data and the plurality of candidate nucleic acid molecule address sequences, and replaces the initial nucleic acid molecule address sequence with the nucleic acid molecule address sequence among the plurality of candidate nucleic acid molecule address sequences that has the highest similarity to the initial nucleic acid molecule address sequence, as the nucleic acid molecule address sequence of the input data.
[0075] In some embodiments, the input data is data to be stored in the nucleic acid molecule, which may include various types of data, such as text, images, videos, and audio. Furthermore, in this disclosure, the initial nucleic acid molecule address sequence can be obtained by any method, and this disclosure does not limit it. As mentioned above, such an initial nucleic acid molecule address sequence may not meet the biological constraints of nucleic acid molecules, resulting in the inability to generate the corresponding nucleic acid molecule, or the generated nucleic acid molecule may be unstable. Therefore, in S2020, the initial nucleic acid molecule address sequence is mapped to a "preferred sequence" that meets the biological stability requirements.
[0076] According to the method 2000 of this disclosure, the initial nucleic acid molecule address sequence is replaced by the "preferred sequence" in the "preferred sequence" set that has the highest similarity to the initial nucleic acid molecule address sequence, and used as the nucleic acid molecule address sequence of the input data. This makes the nucleic acid molecule address sequence of the replaced input data meet the requirements of biological stability, thereby making the nucleic acid molecules generated based on the nucleic acid molecule address sequence more stable. This can improve the stability, effectiveness and practicality of nucleic acid molecule-based data storage.
[0077] In some embodiments, for the set of "preferred sequences" generated in S2010, multiple address molecular modules (e.g., multiple DNA strands) corresponding to each of the multiple "preferred sequences" in the set of "preferred sequences" can be pre-synthesized and stored, for example, in an address molecular module repository. Thus, in the subsequent generation of nucleic acid molecules, it is not necessary to resynthesize the nucleic acid molecule address sequences; instead, the pre-synthesized address molecular modules can be directly used, thereby improving the efficiency of nucleic acid molecule generation.
[0078] In some embodiments, the "preferred sequences" generated in S2010 may be uniformly distributed, i.e., satisfying the uniformity requirement. The uniformity requirement means that the generated set of "preferred sequences" forms a uniform distribution within the possible combination space. More specifically, each "preferred sequence" (e.g., a DNA sequence) is a specific combination of bases, and all possible combinations constitute the design space. Uniformity refers to whether the distribution of sequences within this design space is broad and balanced.
[0079] From the perspective of the distance between any two DNA sequences, uniformity focuses on the distribution of relative distances between sequences throughout the entire sequence set. In a uniformly distributed DNA sequence set, the distances between all DNA sequence pairs should not be excessively concentrated or sparse. In a uniformly distributed sequence set, the distribution of these sequences in the design space will ensure that their distances span the entire possible range. This means that on a distance histogram, each distance interval, from minimum to maximum, will have approximately the same number of sequence pairs. Uniformity requires ensuring that the "preferred sequence" set covers the design space as comprehensively as possible, without neglecting or over-concentrating on any particular region within the design space.
[0080] By ensuring that the pre-made DNA sequences are nearly uniformly distributed in the design space, the resolution of addressing (e.g., the ability to distinguish content retrieval based on similarity) can be improved when they are used as the addresses of mapped DNA molecules.
[0081] Furthermore, it should be understood that the uniform distribution of DNA sequences described in this disclosure is not limited to a strictly uniform distribution, but is a flexible condition that can be limited by setting thresholds or related indicators in conjunction with actual inspection methods.
[0082] Next, refer to Figure 3 An exemplary flowchart of a method 3000 for generating a plurality of alternative nucleic acid molecule address sequences (“preferred sequence” set) that meet biological stability requirements, according to embodiments of the present disclosure, is described. This method 3000 can be used, for example, to implement... Figure 2 Step S2010 in Method 2000.
[0083] like Figure 3 As shown, in S302, parameter settings are performed. The goal is to generate m DNA sequences that meet biological stability requirements, with a sequence length of x. Additionally, constraints are set, such as no more than three consecutive single bases in the sequence and a GC ratio between 40% and 60%. Other necessary constraints can also be added. A set A of DNA sequences to be generated that meet the biological constraints is initialized.
[0084] In some embodiments, the sequence length x set in S302 can be the same as the length of the initial nucleic acid molecule address sequence of the input data, so as to facilitate the replacement of the initial nucleic acid molecule address sequence with the "preferred sequence". Alternatively, the sequence length x can be designed to be different from the length of the initial nucleic acid molecule address sequence, depending on actual needs.
[0085] Furthermore, the number of DNA sequences, m, to be generated can be specifically set as needed. A larger value for m results in more "preferred sequences," providing finer granularity for replacing the initial nucleic acid molecule address sequence. A smaller value for m results in fewer "preferred sequences," providing coarser granularity for replacing the initial nucleic acid molecule address sequence, but reducing the complexity of sequence generation and the potential complexity of generating the corresponding address molecule modules.
[0086] Next, in S304, n DNA sequences of length x are randomly generated. In S306-S310, based on the parameters set in S302, these n DNA sequences are filtered to meet biological stability requirements. For example, in S306, single-base continuation filtering is performed, retaining only DNA sequences with no more than 3 consecutive single bases; in S308, GC ratio filtering is performed, retaining only DNA sequences with a GC ratio between 40% and 60%; and in S310, other optional filtering conditions (such as secondary structure constraints, complexity constraints, and melting temperature constraints) are applied.
[0087] Next, in S312, deduplication is performed by checking whether the filtered sequences are duplicates of the sequences in the initialized set A, and removing the duplicate sequences. In S314, the remaining k DNA sequences that meet the requirements after deduplication in S312 are added to set A.
[0088] Next, in S316, it is determined whether the number of DNA sequences in set A is greater than the number of DNA sequences to be generated, m, set in S302. If the determination in S316 is "no", the process returns to S304 to continue sequence generation, filtering, and deduplication. Otherwise, if the determination in S316 is "yes", the process proceeds to S318.
[0089] In S318, the DNA sequence distribution is checked to ensure it meets the uniformity requirement, i.e., whether the generated DNA sequences are uniformly distributed. If the result of the check in S318 is "No", the process returns to S302, the parameter settings are adjusted, and subsequent steps are repeated. If the result of the check in S318 is "Yes", the process proceeds to S320, and all DNA sequences in set A are output as the "preferred sequences" to be generated.
[0090] Next, an example method for checking whether the DNA sequence distribution meets the uniformity requirement in S318 will be described in detail. This example method may include, for example, the following three steps:
[0091] a) Calculate pairwise distances: For a given set of DNA sequences, calculate the distance between each pair of sequences (any distance metric can be selected as needed, such as the various distance metrics described below), thus obtaining a set containing multiple distance values.
[0092] b) Constructing a distance histogram: Use the set of distance values obtained in the previous step to draw a distance histogram. Theoretically, if the DNA sequences are roughly uniformly distributed, the distance histogram will show approximately the same number of DNA sequence pairs within each distance interval. Therefore, the distance histogram can be used to roughly assess whether the generated set of DNA sequences conforms to a uniform distribution.
[0093] c) Statistical evaluation: Statistical tests are used to evaluate whether the distance distribution between DNA sequences in the observed distance histogram differs significantly from the expected uniform distribution (e.g., obtained statistically). For example, a predetermined threshold can be set, and if the difference between the distance distribution between DNA sequences in the observed distance histogram and the expected uniform distribution is below the predetermined threshold, the DNA sequence distribution is considered to meet the uniformity requirement.
[0094] It should be understood that the method for checking uniformity requirements given above is merely an example, and those skilled in the art can design other methods for checking the uniformity of DNA sequence distribution according to actual circumstances.
[0095] Furthermore, in method 3000, the process of checking whether the DNA sequences meet the uniformity requirements in step S318 is not necessary. In some embodiments, the process of S318 can also be omitted. If it is determined in S316 that the number of DNA sequences in set A is greater than m, the process directly proceeds to S320, and all DNA sequences in set A are output as the "preferred sequences" to be generated that meet the biological stability requirements.
[0096] Next, return to the reference. Figure 2 In some embodiments, for step S2020 of method 2000, there are several optional similarity calculation metrics, such as Hamming distance, edit distance, cosine similarity, Jaccard similarity, and sequence alignment. The similarity in S2020 can be calculated based on any of these similarity distances. Furthermore, these similarity calculation metrics can also be similarly used to calculate the distance between two DNA sequences when checking the uniformity of DNA sequence distribution in S318.
[0097] The following details the similarity calculation indicators mentioned above.
[0098] Hamming distance
[0099] The Hamming distance is a measure of the difference between two strings of equal length. Specifically, it represents the minimum number of substitutions required to make two strings equal. For DNA sequences, the Hamming distance can be understood as the number of bases that differ at corresponding positions in two DNA sequences of equal length. For example, for sequence A: ATCGA and sequence B: ATGGA, the Hamming distance is 1 because only the base at the third position is different.
[0100] Edit distance
[0101] Edit distance is a metric describing the difference between two strings. It represents the minimum number of edit operations required to make the two strings equal, including insertions, deletions, and substitutions. For DNA sequences, edit distance can be understood as the minimum number of base operations required to transform one DNA sequence into another. For example, for sequence A: ATCGA and sequence B: ATTCGA, the edit distance is 1 because it only requires inserting a 'T' into the second position of sequence A or deleting a 'T' from the second position of sequence B.
[0102] Cosine similarity
[0103] Cosine similarity is a method for calculating the cosine of the angle between two vectors in a multidimensional space. Suppose we have two vectors A and B, then the formula for calculating the cosine similarity between these two vectors is:
[0104] CosineSimilarity(A,B)=dot_product(A,B) / (norm(A)*norm(B)) (8),
[0105] Where dot_product(A,B) is the dot product of vectors A and B; norm(A) is the magnitude of vector A, i.e., the length of vector A; norm(B) is the magnitude of vector B, i.e., the length of vector B.
[0106] For example, consider two DNA sequences of length 4: "ATGC" and "ATGA". By mapping each base to a four-dimensional vector using one-hot encoding, for example: A->[1,0,0,0], T->[0,1,0,0], G->[0,0,1,0], C->[0,0,0,1]. Then these two sequences can be transformed into the following vector:
[0107] "ATGC"->[1,0,0,0,0,1,0,0,0,0,1,0,0,0,0,1]
[0108] "ATGA"->[1,0,0,0,0,1,0,0,0,0,1,0,1,0,0,0]
[0109] Furthermore, by using the above formula (8) to calculate the cosine similarity result, the cosine similarity between the two DNA sequences "ATGC" and "ATGA" can be obtained.
[0110] Jaccard similarity
[0111] Jaccard similarity is a metric used to measure the similarity between two sets. It is defined as the size of the intersection of the two sets divided by the size of their union. For DNA sequences, a DNA sequence can be viewed as a set of bases or a set of k-mers (k-mers represent the k consecutive bases in a sequence). For sequence A and sequence B, Jaccard similarity is defined as:
[0112]
[0113] Where |A∩B| is the number of k-mer sets shared by the two sequences, and |A∪B| is the total number of k-mer sets in the two sequences (excluding duplicates).
[0114] Global sequence alignment
[0115] Global sequence alignment can also be used to calculate DNA sequence similarity. Its core idea is to match as many parts of two sequences as possible, even if it requires inserting some gaps. Global sequence alignment attempts to find a way to align two DNA sequences to maximize their total match score. The alignment score can be calculated based on, for example, the number of matched base pairs, the number of mismatches, the number of insertions, and the number of deletions.
[0116] Global sequence alignment methods can be employed, for example, using the Needleman-Wunsch algorithm. This algorithm is a dynamic programming algorithm that creates a score matrix and then uses this matrix to find the optimal alignment.
[0117] Suppose there are two DNA sequences:
[0118] Sequence 1: ACTCGT
[0119] Sequence 2: CTCGTA
[0120] Global sequence alignment algorithms require defining scoring rules. For example, a match might score +1, a mismatch -1, and an insertion or deletion (i.e., creating a gap) -2. Using these scoring rules, the following optimal alignment results can be obtained, where "-" represents a gap:
[0121] Sequence 1: ACTCGT-
[0122] Sequence 2: -CTCGTA
[0123] Furthermore, by calculating the number of matches, gaps, and mismatches, and based on the scoring rules, the similarity result is obtained. In practical applications, the scoring rules and weights can be adjusted according to specific biological needs.
[0124] When selecting a similarity metric, it is necessary to accurately reflect the similarity between sequences, especially in regions where stable hybridization can occur. Different metrics may have different effects on molecular hybridization. For example, global sequence alignment may be more inclined to identify sequences that are highly similar overall, while methods such as local alignment or edit distance may have a higher preference for sequences in locally highly similar regions. Therefore, in this disclosure, one or more of the aforementioned similarity calculation metrics can be selected according to specific biological needs to calculate the similarity between the initial nucleic acid molecular address sequence and the "preferred sequence" in S2020.
[0125] The above describes the method for generating nucleic acid molecule address sequences disclosed herein, which ensures that the generated nucleic acid molecule address sequences meet biological stability requirements. These generated nucleic acid molecule address sequences can then be used for searching nucleic acid molecules in a nucleic acid molecule repository.
[0126] In some nucleic acid molecular data retrieval methods known to the inventor, molecular address sequences are used as DNA molecular probes. DNA molecules matching the probe are searched in a nucleic acid molecular library through DNA hybridization, and the corresponding content within those DNA molecules is read, thereby achieving precise data retrieval. The nucleic acid molecular sequences generated using the method of this disclosure can, for example, be used for such precise retrieval.
[0127] However, in some scenarios, there is a need for semantic-based fuzzy retrieval of data. For example, when a user wants to search for watchable movies, they don't necessarily enter a specific movie title for a precise search, but rather a fuzzy search keyword such as "comedy movie" to retrieve multiple movies that meet the criteria and then select the appropriate one to watch. In this scenario, precise retrieval based on molecular address sequences may not be applicable, and semantic-based fuzzy retrieval is required. The nucleic acid molecular sequences generated using the method of this disclosure can also be used for semantic-based fuzzy retrieval, for example.
[0128] Specifically, in some embodiments, the initial nucleic acid molecule sequence of the input data retains at least a portion of the semantic information of the input data. In such embodiments, since the substitution in S2020 preserves the similarity relationship between the initial nucleic acid molecule sequences, that is, after substitution in S2020, two similar initial nucleic acid molecule sequences will also have similar "preferred sequences". Therefore, the nucleic acid molecule address sequence of the substituted input data can also reflect the semantic information of the input data. Thus, the nucleic acid molecule address sequence generated in this way can both meet the requirements of biological stability and be used for semantic-based fuzzy retrieval.
[0129] Furthermore, when the set of "preferred sequences" conforms to a uniform distribution, using such a set of "preferred sequences" as the address to be mapped molecular addresses can result in higher addressing resolution (e.g., the ability to distinguish content retrieval based on similarity), making it more suitable for semantic-based fuzzy retrieval. That is, it is easier to retrieve multiple nucleic acid molecules that store similar semantic information in a single retrieval.
[0130] Next, we will refer to Figure 4 and Figure 5 The present disclosure provides an exemplary flowchart of a method for generating a nucleic acid molecule address sequence that retains semantic information of input data, according to embodiments of the present disclosure.
[0131] like Figure 4 As shown, method 4000 includes S4010. In S4010, an initial nucleic acid molecule sequence containing at least a portion of the semantic information of the input data is generated. Additionally, method 4000 includes S2010 and S2020, which are consistent with reference to... Figure 2 Steps S2010 and S2020 of method 2000 are the same and will not be repeated here. In method 4000, the initial nucleic acid molecule sequence generated in S4010 retains at least a portion of the semantic information of the input data, and the "preferred sequence" after replacement in S2020 retains the similarity relationship between the initial nucleic acid molecule sequences. Therefore, the nucleic acid molecule address sequence of the input data obtained in this way can also reflect the semantic information of the input data. Thus, the nucleic acid molecule address sequence generated in this way can not only meet the requirements of biological stability, but also be used for semantic-based fuzzy retrieval, enabling the retrieval of one or more fuzzy search results similar to the search information.
[0132] Figure 5 An exemplary flowchart of a method 5000 for generating an initial nucleic acid molecule address sequence that retains semantic information of input data, according to embodiments of the present disclosure, is shown. This method 5000 can, for example, be used to implement... Figure 4 Step S4010 in method 4000.
[0133] like Figure 5 As shown, in S5010, feature extraction is performed on the input data D to generate a first feature vector v1 of the input data D, which has a first dimension l1.
[0134] As described above, the input data is the data to be stored in the nucleic acid molecule, and can include various types of data, such as text, images, videos, and audio. Any feature extraction technique can be used to extract features from the input data, including but not limited to Scale Invariant Feature Transform (SIFT), Histogram of Oriented Gradients (HOG), and deep learning-based feature extraction. In this disclosure, since the first feature vector is generated based on feature extraction from the input data, the first feature vector retains at least a portion of the semantic information of the input data.
[0135] Next, in S5020, the first feature vector v1 of the input data D is transformed into the second feature vector v2 of the input data D, which has a second dimension l2.
[0136] In this disclosure, the transformation in S5020 ensures that similar data in the original data space retains their similarity after transformation. More specifically, in the transformation of this disclosure, for the first input data D1 and the second input data D2, the first feature vector v of D1... 1_D1 With the first eigenvector v of D2 1_D2 The higher the similarity, the better the transformation of the second feature vector v of D1. 2_D1 With the second eigenvector v of D2 2_D2 The higher the similarity, the better. Therefore, the transformed second feature vector not only retains at least some semantic information of the input data, but also reflects the similarity between the input data.
[0137] Furthermore, in this disclosure, the dimension l1 of the first feature vector and the dimension l2 of the second feature vector are not limited. l1 can be greater than, less than or equal to l2, and can be designed according to the actual situation to determine the size relationship between l1 and l2.
[0138] Next, in S5030, the initial nucleic acid molecule address sequence of the input data D is generated based on the second feature vector v2 of the input data D.
[0139] In some embodiments, the length L of the nucleic acid molecule address sequence can be preset, and the second dimension l2 of the converted second feature vector v2 can be set to L × 4, so that every 4 elements (e.g., 4 binary bits) of the second feature vector is mapped to one of the 4 base AGCTs, thereby generating the nucleic acid molecule address sequence of the input data. Furthermore, depending on the different mapping methods from the second feature vector to the base sequence, the relationship between the second dimension l2 of the second feature vector v2 and the length L of the nucleic acid molecule address sequence can also be different. For example, it can be set to l2 = L × 2, using two elements (e.g., 2 bits) in the second feature vector to indicate one of the base AGCTs. Alternatively, each element in the second feature vector can also have a one-to-one correspondence with a base AGCT, in which case the second feature vector can be directly used as the nucleic acid molecule address sequence. Furthermore, the elements of the second feature vector are not limited to discrete values; depending on the conversion method in S5020, the elements of the second feature vector can also be continuous values, and the mapping relationship from the second feature vector to the nucleic acid molecule address sequence can be designed accordingly.
[0140] According to method 5000, by extracting and transforming features from the input data, the semantics of the input data are embedded into an initial DNA molecule address sequence, which also preserves the similarity between the input data. Since the similarity of DNA molecule sequences and DNA hybridization are approximate problems—that is, a pair of DNA molecule sequences with higher similarity (e.g., closer Hamming distances) is more likely to hybridize—it can be approximated that the more similar a pair of DNA molecules is generated, the easier it is to achieve DNA strand matching during the retrieval process, thus retrieving more similar results. Therefore, the initial nucleic acid molecule address sequence generated according to the method of this disclosure retains the semantic information of the input data.
[0141] Furthermore, by replacing the initial nucleic acid molecule addresses that retain semantic information with the "preferred sequence" in S2020, the "preferred sequence" also retains the similarity relationship between the initial nucleic acid molecule addresses, thereby preserving the similarity between the input data. Therefore, the nucleic acid molecule address sequence of the input data after replacement in S2020 ("preferred sequence") satisfies both biological stability requirements and can be used for semantic-based fuzzy retrieval.
[0142] In some embodiments, as described above, the input data can be any of a variety of data types. Additionally, in some embodiments, the input data may include multimodal input data, such as input data composed of a combination of two or more of the data types described above.
[0143] In some embodiments, when the input data includes multimodal input data, in the feature extraction in S5010, the multimodal input data can be mapped to a unified embedding space according to a pre-trained Multimodal Large Language Model (MLLM), wherein the embedding space has a first dimension l1, and the vectors in the mapped embedding space are used as first feature vectors. This enables the alignment of the feature vector representations of the multimodal input data in the embedding space, facilitating subsequent processing. Furthermore, the embedding space mapped to by MLLM is typically a high-dimensional space. Therefore, in some embodiments, the dimension l1 of the embedding space can be set to be greater than the dimension l2 of the transformed second feature vector.
[0144] A suitable pre-trained MLLM can be selected for the data to be stored, and this disclosure does not particularly limit the type of MLLM. For example, if the data to be stored consists entirely of text and images, the CLIP model can be used. Alternatively, an MLLM can be trained manually to generate the embedding space based on the actual data type.
[0145] Next, the implementation of converting the first feature vector into the second feature vector in S5020 will be described in detail. In the following examples, the unit implementing the conversion in S5020 will sometimes be described as a semantic DNA encoder (e.g., which can be configured in...). Figure 1 In the apparatus 1000 shown, a first feature vector extracted from input data is converted into a second feature vector corresponding to a DNA address sequence. The second feature vector retains at least a portion of the semantics of the input data and preserves the similarity structure between the input data.
[0146] In some embodiments, the semantic DNA encoder can be implemented using a pre-trained autoencoder (AE) model. An AE model is an unsupervised learning model that uses the input data itself as supervision, based on backpropagation and optimization methods (such as gradient descent), to guide the neural network in attempting to learn a mapping relationship, thereby obtaining a reconstructed output. The AE model mainly consists of an encoder and a decoder. The encoder's role is to encode high-dimensional input into low-dimensional latent variables, enabling the neural network to learn the latent features of the input data. The decoder's role is to restore the latent variables of the hidden layers to their initial dimensions, thereby reconstructing the original input.
[0147] In some embodiments, the semantic DNA encoder of this application can be implemented using such a structure of the AE model. Specifically, since the training objective of the AE model is to minimize the reconstruction objective, the latent variables, as intermediate products of reconstruction, retain the features and similarities of the input vectors. Thus, the mapping relationship between the high-dimensional input and the low-dimensional latent variables of the AE model can be used to transform the first feature vector of the input data into the second feature vector.
[0148] Next, we will describe in detail the pre-training process for the AE model.
[0149] In some embodiments, a first feature vector of the training data can be input as an input vector X to the AE model. The training data can be pre-prepared data to be stored, and can be single-type data or multimodal data. The first feature vector of the training data can be generated in the same manner as step S5010 in the method S5000 described above. After the input vector X is input into the AE model, the AE model encodes the first feature vector into a latent variable h with a second dimension, and decodes the latent variable with the second dimension into a reconstructed vector X' with a first dimension.
[0150] Next, the AE model undergoes unsupervised learning, and the parameters of the AE model are modified to adjust the loss function. AE The value decreases. In some embodiments, the loss function can be defined based on the difference between the input vector X and the reconstructed vector X', i.e.
[0151] Loss AE =Loss recon (10),
[0152] Loss recon The reconstruction error represents the difference between the input vector X and the reconstructed vector X'. A common representation of the reconstruction error is the mean squared error. In this case, the loss function can be expressed as follows:
[0153] Loss AE =1 / nΣ||X-X'|| 2 (11),
[0154] Where n is the number of samples used as training data, ||·|| 2 This represents the Euclidean norm, which is the square root of the sum of the squares of the elements. Furthermore, depending on the nature of the input data, the reconstruction error can also be represented using other methods such as the negative log-likelihood.
[0155] After pre-training the AE model as described above, the first feature vector v1 of the input data D can be input into the pre-trained AE model, and the latent variables with the second dimension encoded by the pre-trained AE model can be used as the transformed second feature vector v2. Thus, the second feature vector v2 retains at least a portion of the semantic information of the input data D and preserves the similarity between the input data.
[0156] In some embodiments, the above-described AE model may include a Variational Autoencoder (VAE) model. Similar in structure to the AE model, the VAE model also includes an encoder and a decoder. The encoder's role is to encode the high-dimensional input into low-dimensional latent variables, thereby enabling the neural network to learn the most informative features. The decoder's role is to restore the latent variables of the hidden layers to their initial dimensions, thus reconstructing the original input. The pre-training process for the VAE model is similar to the pre-training process for the AE model described above, except that the loss function can also be defined based on the difference between the distribution of the latent variables and a predefined prior distribution.
[0157] Specifically, the loss function of a VAE consists of two parts: the reconstruction error and the KL (Kullback-Leibler) divergence, which can be represented as follows:
[0158] Loss VAE =Loss recon +Loss KL (12),
[0159] Among them, Loss recon The reconstruction error is represented by a function similar to that of the loss function in the Advanced Effects (AE) model, and will not be elaborated upon here; Loss KL KL divergence loss measures the difference between the distribution of a latent variable and a predefined prior distribution. The prior distribution can be, for example, a standard normal distribution or a uniform distribution. The distribution of the latent variable is output by the encoder of the VAE, and the goal of KL divergence is to make the distribution of the latent variable encoded by the encoder as close as possible to the prior distribution.
[0160] In some embodiments, in order to adapt the AE / VAE model to discrete data, or in order to make its latent space discrete, the Gumbel-Softmax method can be used to discretize the AE / VAE model to generate discrete latent variables.
[0161] In this case, the encoder portion of the AE / VAE can be modified to output parameters of the discrete latent variables. Specifically, the encoder output of the AE / VAE includes a set of logits, which are used to define a multinomial distribution. Then, noise from the Gumbel distribution can be sampled and added to the logits to obtain perturbed logits. Next, the Gumbel-Softmax method is used to process the perturbed logits through a softmax function, thus obtaining a continuous approximation of the discrete latent variables. The formula for generating the discrete latent variables is as follows:
[0162] Z=softmax((logits+gumbel_noise) / τ) (13),
[0163] Here, gumbel_noise represents the noise of the Gumbel distribution, and τ is a temperature parameter used to control the degree of discretization. When τ is close to 0, Gumbel-Softmax is closer to true discrete sampling; when τ is large, the distribution becomes more uniform, and the degree of discretization is lower.
[0164] In the above embodiment where the Gumbel-Softmax method is used to discretize the AE / VAE model to generate discrete latent variables, the loss function remains unchanged. It should be understood that for the AE model, the loss function is Equation (10) or (11), and for the VAE model, the loss function is Equation (12). By discretizing the AE / VAE model using the Gumbel-Softmax method, the DNA sequence can be mapped to a discrete latent space, thereby making each element of the second feature vector a discrete value, which facilitates the subsequent generation of nucleic acid molecular address sequences based on the second feature vector.
[0165] The above describes an embodiment of implementing a semantic DNA encoder using a pre-trained AE model. In some embodiments, the semantic DNA encoder can also be implemented based on the Local Sensitive Hashing (LSH) algorithm.
[0166] LSH is an approximate nearest neighbor search algorithm for large-scale datasets. It achieves approximate nearest neighbor search by projecting nearest neighbors into the same "bucket". LSH attempts to ensure that physically adjacent objects have the same "hash" value and are therefore placed in the same bucket.
[0167] In some embodiments, the first feature vector v1 of the input data D can be mapped to a hash value with a third dimension l3 based on the LSH algorithm, and this hash value can be mapped to a second feature vector v2 with a second dimension l2. Subsequently, an initial DNA address sequence is generated based on the second feature vector v2 obtained based on the LSH algorithm.
[0168] Specifically, as an example, the parameters of the LSH encoder can be defined to produce a sufficiently large range of hash values, which can then be converted into DNA sequences using a DNA mapper. For instance, the LSH encoder maps the first feature vector v1 of the input data D to a hash value, further dividing this hash value into L groups (L being the length of the DNA molecule address sequence), with each group taking values in an integer space, for example, between 0 and 3. This means the first feature vector v1 is mapped to a hash value in an L×4 space, and the DNA mapper then maps this hash value to a second feature vector v2. It should be understood that the range of hash values is not particularly limited and can be designed according to actual needs.
[0169] In the above mapping of the DNA mapper, for the first input data and the second input data, the higher the similarity between the hash values of the first input data and the second input data, the higher the similarity between the second feature vector of the mapped first input data and the second feature vector of the second input data. In some embodiments, the DNA mapper can use a simple mapping relationship (e.g., the mapping from hash values to the second feature vector v2 in the L×4 space described above) to maintain the original approximation relationship.
[0170] In the semantic DNA encoder based on the LSH algorithm, since the LSH encoder can encode physically adjacent objects into the same hash value, and the DNA mapper uses a strategy of mapping close hash values into similar vectors, the corresponding second feature vector v2 after mapping is also similar, and the resulting initial DNA address sequence is also similar.
[0171] Next, refer to Figure 6 The diagram illustrates a method 6000 for storing information in a nucleic acid molecule according to embodiments of the present disclosure.
[0172] like Figure 6 As shown, in method 6000, for the multiple candidate nucleic acid molecule address sequences generated in step S2010 of method 2000 or 3000, multiple address molecule modules are pre-synthesized and stored, for example, in [location missing]. Figure 6 The address shown is in the molecular module repository 6020.
[0173] Furthermore, for the input data D to be stored, a DNA address sequence 6010 of the input data D is generated using method 2000 or 3000 according to this disclosure. This DNA address sequence 6010 is the DNA address sequence after replacing the initial DNA address sequence of the input data D with a "preferred sequence".
[0174] Next, module determination 6030 is performed, which involves determining the address molecular module 6040 corresponding to the DNA address sequence 6010 of the input data from multiple address molecular modules (e.g., address molecular module repository 6020).
[0175] Furthermore, using DNA encoding method 6050, a corresponding DNA content molecular module 6060 is generated based on the input data D. Next, using biotechnology (e.g., molecular assembly technology), a DNA molecule 6070 is generated based on the DNA address molecular module 6040 and the DNA content molecular module 6060. Thus, the input data D is stored in the DNA molecule 6070. It should be understood that the aforementioned DNA encoding method 6050 and biotechnology (e.g., molecular assembly technology) can be chosen as appropriate, and this disclosure does not limit them.
[0176] Furthermore, the above method can also be applied to generate other DNA molecules from other input data to be stored. Additionally, such as... Figure 6 As shown, the multiple DNA molecules generated as described above can be stored in the DNA molecule repository 6080.
[0177] In this disclosure, since multiple address molecular modules (e.g., stored in address molecular module repository 6020) corresponding to multiple "preferred sequences" in the "preferred sequence" set are pre-synthesized, it is not necessary to resynthesize the DNA address sequence in the subsequent generation of DNA molecule 6070. Instead, the pre-synthesized address molecular module 6040 can be used directly, thereby improving the generation efficiency of nucleic acid molecules.
[0178] In addition, in some embodiments, the DNA content molecule module 6060 of the input data D can also be pre-synthesized using relevant pre-fabrication techniques. Therefore, in method 6000, nucleic acid molecular data can be generated directly using the pre-synthesized address molecule module 6040 and content molecule module 6060, thereby further improving the efficiency of the nucleic acid molecule generation and storage process.
[0179] Next, refer to Figure 7 This is a schematic diagram illustrating a method 7000 for retrieving information stored in nucleic acid molecules according to embodiments of the present disclosure.
[0180] like Figure 7As shown, in method 7000, a DNA molecule repository 6080 is pre-generated, which is, for example, using... Figure 6 The method shown generates 6000.
[0181] Next, using the DNA molecular address sequence from multiple input data sets as probe 7010, a search is performed in the DNA molecular repository 6080. One or more DNA molecular sequences matching probe 7010 are retrieved through molecular hybridization, and the information stored in these retrieved sequences is extracted as the search result 7020. Molecular data reading and decoding 7030 are then performed on the molecular search result 7020 to finally obtain the search result data D.
[0182] In some embodiments, in method 7000, the DNA molecular address sequence of the input data can be generated, for example, by method 2000 or 3000 of this disclosure. When the DNA molecular address sequence of the input data is generated according to method 3000, it retains at least a portion of the semantic information of the input data and preserves the similarity between the input data. Using the DNA molecular address sequence generated in this way for fuzzy retrieval, DNA molecules with molecular sequences similar to that DNA molecular address sequence can be retrieved.
[0183] In some embodiments, search information retrieved by a user can be obtained and used as input data to generate a DNA molecular address sequence of the search information using method 2000 or method 3000 according to this disclosure. Next, the generated DNA molecular address sequence of the search information can be used as a probe 7010 for searching in the DNA molecular repository 6080. Thus, a corresponding probe 7010 can be generated based on the user's search information, thereby performing the user's desired search. On the other hand, even if the user's search information is different from the address sequences of the generated DNA molecules in the DNA molecular repository 6080, a corresponding probe 7010 can be generated in real time.
[0184] Furthermore, the user's search information can be single-type search data, such as text keywords, images, or voice, or it can be search information composed of multimodal data. Additionally, even if the user's search information is single-type search data (e.g., text keywords), because the method for generating DNA molecular address sequences disclosed herein can support the generation of DNA molecular address sequences from cross-modal input data, it is possible to retrieve search results of other modalities similar to the user's input search data (e.g., images related to the keywords).
[0185] In some embodiments, multiple DNA molecular address sequences corresponding to multiple search information can be pre-generated using method 2000 or 3000 according to this disclosure. These multiple search information are search information that the user may be interested in. Additionally, the DNA molecular address sequence corresponding to the search information retrieved by the user can be determined from the pre-generated multiple nucleic acid molecular address sequences and used as a probe 7010 for retrieval in the DNA molecular repository 6080. In this embodiment, because multiple DNA molecular address sequences corresponding to multiple search information are pre-generated, the time required for real-time generation of DNA molecular address sequences based on the user's search information can be saved, thereby improving search efficiency.
[0186] Figure 8 An exemplary configuration is shown that enables a computing device 800 according to an embodiment of the present invention.
[0187] Computing device 800 is an example of a hardware device capable of applying the above aspects of the present invention. Computing device 800 can be any machine configured to perform processing and / or calculations. Computing device 800 can be, but is not limited to, a workstation, server, desktop computer, laptop computer, tablet computer, personal data assistant (PDA), smartphone, in-vehicle computer, or a combination thereof.
[0188] like Figure 8 As shown, computing device 800 may include one or more components that can be connected to or communicate with bus 802 via one or more interfaces. Bus 802 may include, but is not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. Computing device 800 may include, for example, one or more processors 804, one or more input devices 806, and one or more output devices 808. The one or more processors 804 may be any type of processor and may include, but is not limited to, one or more general-purpose processors or special-purpose processors (such as special-purpose processing chips). Processor 802 may, for example, correspond to... Figure 1The processor 1010 is configured to implement the functions of the various units of the present invention: the means for generating nucleic acid molecule address sequences, the means for storing information in nucleic acid molecules, and the means for retrieving information stored in nucleic acid molecules. The input device 806 can be any type of input device capable of inputting information to a computing device, and may include, but is not limited to, a mouse, keyboard, touchscreen, microphone, and / or remote controller. The output device 808 can be any type of device capable of presenting information, and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer.
[0189] The computing device 800 may also include or be connected to a non-transitory storage device 814, which may be any non-transitory storage device capable of storing data, and may include, but is not limited to, disk drives, optical storage devices, solid-state storage, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, compressed disks or any other optical media, cache memory and / or any other storage chip or module, and / or any other medium from which a computer may read data, instructions and / or code. The computing device 800 may also include random access memory (RAM) 810 and read-only memory (ROM) 812. ROM 812 may store executable programs, utilities, or processes in a non-volatile manner. RAM 810 provides volatile data storage and stores instructions related to the operation of the computing device 800. The computing device 800 may also include a network / bus interface 816 coupled to a data link 818. The network / bus interface 816 can be any kind of device or system capable of enabling communication with external devices and / or networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication devices and / or chipsets (such as Bluetooth). TM Equipment, IEEE 802.11 equipment, WiFi equipment, WiMax equipment, mobile cellular communication facilities, etc.
[0190] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to limit the scope of this disclosure. Unless the context explicitly indicates otherwise, the singular forms “a” and “the” as used herein are intended to include the plural forms as well. It should also be understood that the word “comprising”, as used herein, indicates the presence of the indicated feature, integral, step, operation, unit, and / or component, but does not preclude the presence or addition of one or more other features, integrals, steps, operations, units, and / or components, and / or combinations thereof. Furthermore, in the description of this disclosure, the terms “first,” “second,” etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or order. Additionally, in the description of this disclosure, unless otherwise stated, “a plurality of” means two or more.
[0191] In this specification, references to "embodiment" or similar expressions mean that a specific feature, structure, or characteristic described in connection with that embodiment is included in at least one specific embodiment of this disclosure. Therefore, the use of phrases such as "in an embodiment of this disclosure" and similar expressions in this specification does not necessarily refer to the same embodiment.
[0192] Those skilled in the art will understand that this disclosure can be implemented in various forms, such as a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microprogram code, etc.), or a software and hardware embodiment, hereinafter referred to as a "circuit," "module," "unit," or "system." Furthermore, this disclosure can also be implemented in any tangible media as a computer program product having computer-usable program code stored thereon.
[0193] The description herein is based on flowcharts and / or block diagrams of systems, apparatuses, methods, and computer program products according to specific embodiments of this disclosure. It will be understood that each block in each flowchart and / or block diagram, and any combination of blocks in the flowcharts and / or block diagrams, can be implemented using computer program instructions. These computer program instructions are executable by a machine comprising a processor of a general-purpose computer or a special-purpose computer, or other programmable data processing means, and are processed by the computer or other programmable data processing means to perform the functions or operations described in the flowcharts and / or block diagrams.
[0194] The accompanying drawings illustrate flowcharts and block diagrams showing the architecture, functionality, and operation of systems, apparatuses, methods, and computer program products achievable according to various embodiments of the present disclosure. Thus, each block in a flowchart or block diagram may represent a module, segment, or portion of program code, including one or more executable instructions to implement a specified logical function. It should also be noted that in some other embodiments, the functions described in a block may not be performed in the order shown in the figures. For example, two blocks illustrated as connected may actually be executed simultaneously, or in some cases, depending on the functions involved, they may be executed in the reverse order shown in the figures. Furthermore, it should be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware, or by a combination of dedicated hardware and computer instructions, to perform specific functions or operations.
[0195] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to market technology of the embodiments, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating a nucleic acid molecule address sequence, comprising: pre-generating a plurality of candidate nucleic acid molecule address sequences that meet a requirement of biological stability and are uniformly distributed; generating an initial nucleic acid molecule address sequence of input data that retains at least part of semantic information of the input data; and calculating a similarity of the initial nucleic acid molecule address sequence of the input data with the plurality of candidate nucleic acid molecule address sequences, replacing the initial nucleic acid molecule address sequence with a nucleic acid molecule address sequence of the plurality of candidate nucleic acid molecule address sequences that has the highest similarity with the initial nucleic acid molecule address sequence, as the nucleic acid molecule address sequence of the input data.
2. The method of claim 1, wherein, the similarity of the initial nucleic acid molecule address sequence with the plurality of candidate nucleic acid molecule address sequences is calculated based on any of the following similarity distances: Hamming distance, edit distance, cosine similarity, Jaccard similarity, and sequence alignment.
3. The method of claim 1, wherein, the requirement of biological stability is defined by one or more of the following factors: GC content indicating a percentage of the number of base C and base G over the total number of nucleic acid molecule bases, single base continuity indicating a number of times a single base appears consecutively in a nucleic acid molecule sequence, potential formation of secondary structure indicating a possibility of a nucleic acid molecule sequence folding back to form a secondary structure, nucleic acid molecule sequence complexity indicating a complexity of different base sequences formed, and melting temperature.
4. The method of claim 1, further comprising: generating the initial nucleic acid sequence of the input data that retains at least part of semantic information of the input data, wherein generating the initial nucleic acid sequence of the input data comprises: performing feature extraction on the input data to generate a first feature vector of the input data, the first feature vector having a first dimension; converting the first feature vector of the input data to a second feature vector of the input data, the second feature vector having a second dimension; and generating the initial nucleic acid address sequence of the input data from the second feature vector of the input data, wherein in the conversion, the higher the similarity of the first feature vector of the first input data with the first feature vector of the second input data, the higher the similarity of the converted second feature vector of the first input data with the second feature vector of the second input data.
5. The method of claim 4, wherein, the input data comprises multi-modal input data, and the feature extraction comprises: mapping the multi-modal input data to a unified embedding space according to a pre-trained multi-modal large language model, taking a vector in the mapped embedding space as the first feature vector, wherein the embedding space has the first dimension.
6. The method of claim 4, wherein, the conversion comprises: converting the first feature vector of the input data to the second feature vector of the input data using a pre-trained autoencoder model, wherein the autoencoder model is pre-trained by: inputting a first feature vector of the training data as an input vector to the autoencoder model to cause the autoencoder model to encode the first feature vector into a latent variable having a second dimension and to decode the latent variable having the second dimension into a reconstructed vector having the first dimension; and modifying parameters of the autoencoder model to cause a value of a loss function to decrease, wherein the loss function is defined based on a difference between the input vector and the reconstructed vector, wherein in the conversion, a first feature vector of the input data is inputted to the pre-trained autoencoder model, and a latent variable having a second dimension encoded by the pre-trained autoencoder model is taken as the converted second feature vector.
7. The method of claim 6, wherein, the autoencoder model comprises a variational autoencoder model, the loss function is further defined based on a difference between a distribution of the latent variable and a predefined prior distribution.
8. The method of claim 6, wherein, the latent variable is a discrete latent variable generated based on a Gumbel-Softmax method.
9. The method of claim 4, wherein, the conversion comprises: mapping the first feature vector of the input data into a hash value having a third dimension based on a locality sensitive hashing algorithm; and mapping the hash value into the second feature vector, wherein in the mapping of the hash value into the second feature vector, for a first input data and a second input data, the higher the similarity between the hash value of the first input data and the hash value of the second input data, the higher the similarity between the mapped second feature vector of the first input data and the second feature vector of the second input data.
10. A method for storing information in a nucleic acid molecule, comprising: generating a nucleic acid molecule address sequence of input data to be stored in the nucleic acid molecule using the method of any one of claims 1 to 9; pre-synthesizing a plurality of address molecule modules for the plurality of candidate nucleic acid molecule address sequences; determining an address molecule module corresponding to the nucleic acid molecule address sequence of the input data from the plurality of address molecule modules; and generating a nucleic acid molecule based on the address molecule module corresponding to the nucleic acid molecule address sequence of the input data and a content molecule module of the input data to store the input data into the nucleic acid molecule. the content molecule module of the input data is pre-synthesized.
11. The method of claim 10, wherein, 12. A method for retrieving information stored in a nucleic acid molecule, comprising: storing a plurality of input data into a plurality of nucleic acid molecules using the method of claim 10 or 11 to generate a nucleic acid molecule storage comprising the plurality of nucleic acid molecules; retrieving one or more nucleic acid molecule sequences matching a probe in the nucleic acid molecule storage by molecular hybridization using a nucleic acid molecule address sequence of input data in the plurality of input data as the probe; and extracting information stored in the retrieved one or more nucleic acid molecule sequences as a result of the retrieval.
13. The method of claim 12, further comprising: acquiring retrieval information for a user to perform the retrieval; using the method according to any one of claims 1 to 9 as a probe for searching in the nucleic acid molecule repository. and using the generated nucleic acid molecule address sequence of the search information as a probe for searching in the nucleic acid molecule repository.
14. The method according to claim 13, further comprising: pre-generating a plurality of nucleic acid molecule address sequences respectively corresponding to a plurality of search information using the method according to any one of claims 1 to 9; and determining, as a probe for searching in the nucleic acid molecule repository, a nucleic acid molecule address sequence corresponding to the search information searched by the user from the pre-generated plurality of nucleic acid molecule address sequences.
15. An apparatus for generating a nucleic acid molecule address sequence, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored on the memory to perform the following processing: pre-generating a plurality of candidate nucleic acid molecule address sequences that meet a requirement of biological stability and are uniformly distributed; generating an initial nucleic acid molecule address sequence of input data that retains at least part of semantic information of the input data; and calculating a similarity of the initial nucleic acid molecule address sequence of the input data with the plurality of candidate nucleic acid molecule address sequences, and replacing the initial nucleic acid molecule address sequence with a nucleic acid molecule address sequence of the plurality of candidate nucleic acid molecule address sequences that has the highest similarity with the initial nucleic acid molecule address sequence as the nucleic acid molecule address sequence of the input data.
16. The apparatus according to claim 15, wherein the similarity of the initial nucleic acid molecule address sequence with the plurality of candidate nucleic acid molecule address sequences is calculated based on any of the following similarity distances: Hamming distance, edit distance, cosine similarity, Jaccard similarity, and sequence alignment.
17. The apparatus according to claim 15, wherein the requirement of biological stability is defined by one or more of the following factors: GC content indicating a percentage of the number of bases C and bases G over the total number of bases of a nucleic acid molecule, mono-base continuity indicating a number of times a single base appears consecutively in a nucleic acid molecule sequence, potential formation of secondary structure indicating a possibility of a nucleic acid molecule sequence folding back to form a secondary structure, nucleic acid molecule sequence complexity indicating a complexity of different base sequences formed, and melting temperature.
18. The apparatus of claim 15, wherein, the processor is configured to execute the instructions stored on the memory to further perform: generating the initial nucleic acid molecule sequence of the input data that retains at least part of semantic information of the input data, wherein generating the initial nucleic acid molecule sequence of the input data comprises: performing feature extraction on the input data to generate a first feature vector of the input data, the first feature vector having a first dimension; converting the first feature vector of the input data to a second feature vector of the input data, the second feature vector having a second dimension; and generating the initial nucleic acid molecule address sequence of the input data according to the second feature vector of the input data, In the conversion, the higher the similarity between the first feature vector of the first input data and the first feature vector of the second input data, the higher the similarity between the second feature vector of the converted first input data and the second feature vector of the second input data.
19. The apparatus of claim 18, wherein, The input data includes multi-modal input data, and the processor is configured to execute instructions stored on the memory to further perform the following processing: According to the pre-trained multi-modal large language model, the multi-modal input data is mapped to a unified embedding space, and the vector in the mapped embedding space is taken as the first feature vector, wherein the embedding space has the first dimension.
20. The apparatus of claim 18, wherein, The conversion includes: using a pre-trained autoencoder model to convert the first feature vector of the input data into the second feature vector of the input data, wherein the autoencoder model is pre-trained by the following processing: inputting the first feature vector of the training data into the autoencoder model as an input vector, so that the autoencoder model encodes the first feature vector into a latent variable with a second dimension, and decodes the latent variable with the second dimension into a reconstructed vector with the first dimension; and modify the parameters of the autoencoder model to reduce the value of the loss function, wherein the loss function is defined based on the difference between the input vector and the reconstructed vector, wherein in the conversion, the first feature vector of the input data is input into the pre-trained autoencoder model, and the latent variable with the second dimension encoded by the pre-trained autoencoder model is taken as the second feature vector after conversion.
21. The apparatus of claim 20, wherein, the autoencoder model includes a variational autoencoder model, the loss function is further defined based on the difference between the distribution of the latent variable and a predefined prior distribution.
22. The apparatus of claim 20, wherein, the latent variable is a discrete latent variable generated based on a Gumbel-Softmax method.
23. The apparatus of claim 18, wherein, The conversion includes: mapping the first feature vector of the input data into a hash value with a third dimension based on a local sensitive hashing algorithm; and mapping the hash value into the second feature vector, wherein in the mapping of the hash value to the second feature vector, the higher the similarity between the hash value of the first input data and the hash value of the second input data, the higher the similarity between the second feature vector of the mapped first input data and the second feature vector of the second input data.
24. An apparatus for storing information in a nucleic acid molecule, comprising: a memory having instructions stored thereon; and a processor configured to execute instructions stored on the memory to perform the following processing: generating a nucleic acid molecule address sequence of input data to be stored in the nucleic acid molecule using the apparatus of any one of claims 15 to 23; pre-synthesizing a plurality of address molecule modules for the plurality of candidate nucleic acid molecule address sequences; determining, from the plurality of address molecule modules, an address molecule module corresponding to the nucleic acid molecule address sequence of the input data; and generating a nucleic acid molecule based on the address molecule module corresponding to the nucleic acid molecule address sequence of the input data and the content molecule module of the input data to store the input data into the nucleic acid molecule.
25. The apparatus of claim 24, wherein, The content molecule module of the input data is pre-synthesized.
26. An apparatus for retrieving information stored in a nucleic acid molecule, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored on the memory to perform the following processing: storing a plurality of input data into a plurality of nucleic acid molecules by using the apparatus according to claim 24 or 25 to generate a nucleic acid molecule repository comprising the plurality of nucleic acid molecules; retrieving one or more nucleic acid molecule sequences matching a probe in the nucleic acid molecule repository by molecular hybridization using a nucleic acid molecule address sequence of an input data in the plurality of input data as the probe; and extracting information stored in the one or more nucleic acid molecule sequences as a result of the retrieval.
27. The apparatus of claim 26, wherein, The processor is configured to execute the instructions stored on the memory to further perform the following processing: obtaining retrieval information for retrieval by a user; generating a nucleic acid molecule address sequence of the retrieval information by using the apparatus according to any one of claims 15 to 23 as input data; and using the generated nucleic acid molecule address sequence of the retrieval information as the probe for retrieval in the nucleic acid molecule repository. The processor is configured to execute the instructions stored on the memory to further perform the following processing:
28. The apparatus of claim 27, wherein, pre-generating a plurality of nucleic acid molecule address sequences respectively corresponding to a plurality of retrieval information by using the apparatus according to any one of claims 15 to 23; and determining, from the pre-generated plurality of nucleic acid molecule address sequences, a nucleic acid molecule address sequence corresponding to retrieval information for retrieval by the user as the probe for retrieval in the nucleic acid molecule repository.
29. A computer readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 9.