Information storage method and reading method based on polypeptide

By constructing a collection of polypeptide chains of positive and negative electrical guiding units, combining nanopore detection technology and polypeptide concentration encoding, the problems of high cost and low accuracy in the storage and reading of polypeptides are solved, and high-density and low-cost information storage and accurate reading are achieved.

CN120432020APending Publication Date: 2025-08-05NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604888.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing polypeptide information storage and reading technologies have problems such as high cost, complex operation and low reading accuracy. In particular, mass spectrometry technology is expensive and susceptible to amino acid modification and isomers. Nanopore technology has redundant information under the limitation of the polypeptide chain length.

Method used

Using a polypeptide-based information storage method, by constructing a collection of information storage polypeptide chains including positive and negative electrical guiding units, information storage and reading is performed using nanopore detection technology, and combining polypeptide concentration as a coding method, the information storage density is improved and the accuracy of reading is achieved.

Benefits of technology

It achieves higher information storage density and lower cost, and uses nanopore technology to accurately extract peptide information, avoiding interference from amino acid modification and isomers, and adapting to a higher storage density design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432020A_ABST
    Figure CN120432020A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of information storage, and relates to a polypeptide-based information storage method and a polypeptide-based information reading method.The polypeptide-based information storage method comprises the steps that 1, to-be-stored target information is obtained; 2) converting the target information into N-ary coding information, wherein N is greater than or equal to 2; 3) constructing a set of information storage polypeptide chains; and 4) assigning the N-ary coding information obtained in the step 2) into the information storage polypeptide chain set obtained in the step 3) in sequence, and completing information storage of the to-be-stored target information. The invention provides the polypeptide-based information storage method and reading method which are high in information storage density, more accurate in extraction and lower in cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information storage, and relates to an information storage method and a reading method based on biological molecules, and in particular to an information storage method and a reading method based on polypeptides. Background Art

[0002] With the rapid development of information technology, the amount of global data is exploding. Traditional silicon-based storage media, due to limited storage capacity and poor durability, are facing severe challenges. In contrast, biomolecule-based information storage technology, with its high storage density, long-term stability, and good environmental adaptability, is gradually becoming an emerging data storage solution.

[0003] In recent years, DNA storage technology has been widely developed and has become increasingly mature. However, compared to DNA, peptides have many advantages in information storage and have gradually attracted attention (Nature communications, 2021, 12(1): 4242). For example: (1) Higher storage density: Peptides can be composed of 20 natural amino acids and a variety of non-natural amino acids. Compared with DNA, which is composed of only four nucleotides, its richer monomer collection can provide a higher theoretical information storage density. (2) Stronger information carrying capacity: The monomer mass of peptides is lower than that of nucleotides, so more information can be stored under the same mass conditions. (3) Better stability. Studies have shown that peptides or proteins can still be detected and sequenced after millions of years, while DNA is often degraded. This high stability ensures the long-term preservation of information and reduces the cost of data transfer caused by the degradation of storage media.

[0004] However, the current information storage and reading based on peptides still mainly rely on mass spectrometry technology, but it has many limitations. First, high-resolution mass spectrometers are expensive and have high maintenance costs, which limits their application in large-scale data storage and reading. Second, mass spectrometry detection requires complex sample preparation processes, such as enzymatic hydrolysis, purification, and matrix-assisted methods, which increases time and operating costs. In addition, in complex peptide mixtures, mass spectrometry readings are susceptible to interference from mass shift, post-translational modifications (PTMs), and isomers, affecting reading accuracy (Nature Communications, 2021, 12(1): 5795). In contrast, nanopore technology exhibits excellent reading capabilities. This technology has ultra-high precision and has been proven to achieve accurate identification of 20 natural amino acids (Nature Methods, 2024, 21(4): 609-618). It can analyze peptide sequences in real time and efficiently, while having lower costs and more convenient operation methods. Therefore, nanopore technology is expected to become an ideal solution for peptide information storage and reading. However, due to the limitations of current peptide sequencing technology on amino acid chain length, a large amount of information usually requires multiple peptide chains to be stored simultaneously. In order to determine the reading order of these peptide chains, traditional sequence storage molecules often use additional monomers to encode address bits to index information (Nature Reviews Genetics, 2019, 20(8): 456-466). However, this method will lead to information redundancy. Therefore, exploring new reading technologies and storage methods to effectively solve this problem is still a difficult problem that needs to be overcome. Summary of the Invention

[0005] In order to solve the above technical problems existing in the background technology, the present invention provides a polypeptide-based information storage method and reading method with high information storage density, more accurate extraction and lower cost.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A polypeptide-based information storage method, comprising the following steps:

[0008] 1) Obtain target information to be stored;

[0009] 2) converting the target information into N-ary coded information, where N is greater than or equal to 2;

[0010] 3) Constructing a collection of information storage polypeptide chains;

[0011] 4) Sequentially assigning the N-ary coded information obtained in step 2) to the set of information storage polypeptide chains obtained in step 3) to complete the information storage of the target information to be stored.

[0012] Preferably, the collection of information storage polypeptide chains in step 3) includes multiple information storage polypeptide chains, each of which includes a positively charged guiding unit, an information storage unit and a negatively charged guiding unit; the positively charged guiding unit is connected to the negatively charged guiding unit through the information storage unit; the positively charged guiding unit is a compound formed by multiple positively charged amino acids through peptide bonds; the negatively charged guiding unit is a compound formed by multiple negatively charged amino acids through peptide bonds; the information storage unit is an amino acid sequence formed by basic constituent monomers through peptide bonds.

[0013] Preferably, the amino acids in the positively charged directing unit and the amino acids in the negatively charged directing unit are both natural or unnatural amino acids.

[0014] Preferably, the specific implementation of step 3) is:

[0015] 3.1) Selecting basic components for constructing the information storage unit;

[0016] 3.2) connecting the basic building blocks obtained in step 3.1) via peptide bonds to form an amino acid sequence;

[0017] 3.3) Connecting a positively charged guiding unit and a negatively charged guiding unit to the front and back ends of the amino acid sequence, respectively, to form an information storage polypeptide chain;

[0018] 3.4) Constructing a collection of information storage polypeptide chains based on multiple information storage polypeptide chains.

[0019] Preferably, the basic constituent monomers in step 3.1) are the same natural or non-natural amino acids or non-same natural or non-natural amino acids;

[0020] When the basic constituent monomers in step 3.1) are the same natural or non-natural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of different concentrations formed by the same basic constituent monomers, and integrating the multiple information storage polypeptide chains of different concentrations to form a collection of information storage polypeptide chains;

[0021] When the basic constituent monomers in step 3.1) are non-identical natural or non-natural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of the same concentration formed by non-identical basic constituent monomers, and integrating the multiple information storage polypeptide chains of the same concentration to form a collection of information storage polypeptide chains;

[0022] When the basic constituent monomers in step 3.1) are non-identical natural or non-natural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of non-identical concentrations formed by non-identical basic constituent monomers, and integrating the multiple information storage polypeptide chains of non-identical concentrations to form a collection of information storage polypeptide chains.

[0023] Preferably, the specific implementation of step 4) is:

[0024] 4.1) Obtaining the address bits of the N-ary coded information obtained in step 2);

[0025] 4.2) detecting all the information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 3) in the nanopore, respectively, to obtain blocking current signals of all the information storage polypeptide chains;

[0026] 4.3) Based on the blocking current signals of all the information storage polypeptide chains obtained in step 4.2), respectively obtaining the signal frequencies of different information storage polypeptide chains;

[0027] 4.4) Arranging the signal frequencies obtained in step 4.3) in order of highest to lowest frequencies to obtain an arrangement result;

[0028] 4.5) Mapping the high-low order results obtained in step 4.4) with the address bits obtained in step 4.1) in a front-to-back manner to complete the information storage of the target information to be stored.

[0029] Preferably, the different information storage polypeptide chains in step 4.2) are multiple information storage polypeptide chains of different concentrations formed from the same basic monomers, multiple information storage polypeptide chains of the same concentration formed from different basic monomers, or multiple information storage polypeptide chains of different concentrations formed from different basic monomers;

[0030] The blocking current signal includes a blocking current, a current standard deviation, a blocking time, and a signal frequency.

[0031] A polypeptide-based information reading method, comprising the following steps:

[0032] 1) obtaining a set of information storage polypeptide chains storing target information;

[0033] 2) detecting the set of information storage polypeptide chains storing target information obtained in step 1) based on nanopore detection technology to obtain the blocking current signal frequency of all information storage polypeptide chains in the set of information storage polypeptide chains; the blocking current signal frequency of all information storage polypeptide chains in the set of information storage polypeptide chains is the number of current signals of a single information storage polypeptide chain in the set of information storage polypeptide chains per unit time;

[0034] 3) Analyzing target information based on the blocking current signals of all information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 2).

[0035] Preferably, the collection of information storage polypeptide chains includes multiple information storage polypeptide chains, each of which includes a positively charged guiding unit, an information storage unit and a negatively charged guiding unit; the positively charged guiding unit is connected to the negatively charged guiding unit through the information storage unit; the positively charged guiding unit is a compound formed by multiple positively charged amino acids through peptide bonds; the negatively charged guiding unit is a compound formed by multiple negatively charged amino acids through peptide bonds; the information storage unit is an amino acid sequence formed by basic constituent monomers through peptide bonds.

[0036] Preferably, the step 3) specifically includes:

[0037] 3.1) parsing the signal frequencies of all the information storage polypeptide chains in the set of information storage polypeptide chains based on the blocking current signals of all the information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 2);

[0038] 3.2) determining the address bit of the target information based on the signal frequencies of all information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 3.1);

[0039] 3.3) Obtain target information based on the mapping relationship between the address bits and the information storage polypeptide chain.

[0040] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0041] The present invention provides a polypeptide-based information storage method and a reading method, wherein the polypeptide-based information storage method includes the following steps: 1) obtaining target information to be stored; 2) converting the target information into N-base coded information, where N≥2; 3) constructing a set of information storage polypeptide chains; 4) sequentially assigning the N-base coded information obtained in step 2) to the set of information storage polypeptide chains obtained in step 3), thereby completing the information storage of the target information to be stored. The present invention can not only utilize polypeptide sequences for information storage, but also use polypeptide concentration as a new encoding method to further improve the information storage density; it provides a new method for reading polypeptide information storage, and utilizes nanopore single-molecule detection technology to accurately extract information in a lower-cost and simpler manner, while not being limited by interference from amino acid modifications or isomers, and can theoretically adapt to the design of information polypeptide chains with higher storage density. The present invention provides a polypeptide-based information storage method and a reading method. Utilizing nanopore single-molecule detection technology, information can be accurately extracted in a lower-cost and simpler manner. It is not subject to interference from isomers and is not limited to natural amino acids. Natural amino acids and post-translationally modified amino acids (more than 400 types) can be used for storage, representing a wider range of storage unit types and theoretically adapting to higher storage density designs. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of the process of the polypeptide-based information storage method provided by the present invention;

[0043] Figure 2 This is a schematic diagram of the process of the polypeptide-based information reading method provided by the present invention;

[0044] Figure 3 is a schematic diagram of the information storage polypeptide chain used in the present invention;

[0045] Figure 4 This is a schematic diagram of the sequence coding of the information storage polypeptide chain used in the present invention;

[0046] Figure 5 Schematic diagram of the concentration coding of the information storage polypeptide chain used in the present invention;

[0047] Figure 6 Schematic diagram of the sequence design of the information storage polypeptide chain in Example 1 of the present invention and the information storage principle based on polypeptide sequence and concentration coding;

[0048] Figure 7 This is a single-molecule signal current distribution diagram of eight polypeptides with different sequence information under the same concentration conditions in Example 1 of the present invention;

[0049] Figure 8This is a comparison diagram of current signals caused by the relative movement of the information storage polypeptide chain GGG and the information storage polypeptide chain GGF in the nanopore in Example 1 of the present invention;

[0050] Figure 9 The single molecule signal current distribution diagram and single molecule signal frequency comparison diagram of the information storage polypeptide chain GGG and the information storage polypeptide chain GGF with equal concentrations a priori in Example 1 of the present invention;

[0051] Figure 10 This is the actual experimental data of information storage and reading in Example 1 of the present invention;

[0052] Figure 11 This is the single-molecule signal current distribution of the polypeptide composed of the extended encoded amino acids supplemented in Example 2 of the present invention;

[0053] Figure 12 3 is a graph showing the relationship between the concentration of the information storage polypeptide chain GGG and the signal frequency under different concentration conditions in Example 3 of the present invention. DETAILED DESCRIPTION

[0054] The following is a further description of the invention with reference to the accompanying drawings:

[0055] See also Figure 1 as well as Figure 2 The present invention provides a system for storing and reading polypeptide information, comprising: an information storage method and an information reading method. The information storage method involves storing data using a collection of polypeptide chains. The information reading method involves enabling the polypeptide chains to move relative to a nanopore, enabling nanopore detection technology to read and interpret the information stored in the polypeptides.

[0056] See also Figure 3 The information storage polypeptide chain provided by the present invention and which can be used for information storage comprises an information storage unit and a guide unit, and is composed of guide units with opposite electrical properties at both ends and an information storage unit in the middle. The information storage unit can be different types of amino acids with arbitrary electrical properties. Exemplarily, the guide units are located at both ends of the information storage unit (polypeptide chain) and have opposite electrical properties. Exemplarily, one side can contain positively charged lysine (K), arginine (R) and histidine (H), while the other side can contain negatively charged aspartic acid (D) and glutamic acid (E). The guide chain can use electrophoretic force to drive the information storage polypeptide chain to enter the nanopore in a directional manner, and balance the force acting on the information storage polypeptide chain in the pore, so that the information storage polypeptide chain can stay in the pore for a longer time, thereby improving the accuracy of nanopore reading. See. Figure 6Here, four arginines (R) and four aspartic acids (D) are used as guide strands. The information storage unit is located between the guide units and encodes information via sequences of at least two amino acids, which can include natural, unnatural, and modified amino acids. This is due to the nanopore's ability to accurately identify natural, unnatural, and modified amino acids.

[0057] See also Figure 4 The encoding method of the information storage unit is as follows: 1) N amino acids are selected as the basic building blocks of the information storage unit, and the N amino acids can be assigned values of 0, 1, 2, ..., N-1; 2) the content to be stored is converted into N-base coded information (for example, the stored content can be converted into N-base information according to ASCII or Unicode); 3) based on the coded information, the amino acid sequence of the information storage unit is generated according to the assigned amino acid values. Due to the large number of amino acids, including 20 standard natural amino acids and hundreds of unnatural amino acids and modified amino acids, a high information storage density can be guaranteed.

[0058] See also Figure 5 , where (a) is to use the frequency of single-molecule signals corresponding to different concentration conditions, from high to low, to assign values 0, 1, 2 to N-1 in sequence, to perform N-ary information encoding; (b) is to use the frequency of single-molecule signals corresponding to different sequence polypeptides under specific concentrations, from high to low, to assign values 0, 1, 2 to N-1 in sequence, to number the reading order of the polypeptide-carried information, that is, to perform address bit encoding. Exemplarily, the concentration of polypeptide chains is used to store information, and the information storage polypeptide chain set is as follows: Figure 5 (a) When it contains one polypeptide chain, its concentration gradient is greater than or equal to 2; Figure 5 (b) When two or more polypeptide chains are included, the concentrations of the different polypeptide chains may be the same or different. The stored information is assigned values from high to low based on the single-molecule signal frequencies obtained under the corresponding concentration conditions. Preferably, the polypeptide concentrations encode the order in which the polypeptides are read, i.e., the address bits.

[0059] Example 1

[0060] In order to realize information storage, the corresponding relationship between amino acids and sequences should be established first, that is, N kinds of amino acids are selected as the basic monomers of information storage units, and N kinds of amino acids can be assigned numbers such as 0, 1, 2..., N-1. Figure 6Here, two amino acids, glutamic acid (G) and phenylalanine (F), are selected as the basic monomers of the information storage unit, where glutamic acid (G) is assigned a value of "0" and phenylalanine (F) is assigned a value of "1". Three amino acids are used as a group to represent the storage unit of the polypeptide chain. Eight different combinations of polypeptide chains can be used to represent the binary code. The corresponding information coding sequence is shown in Table 1.

[0061] Table 1 Peptide sequence information coding table

[0062] name Binary sequence Information storage unit sequence Polypeptide chain 1 000 GGG Polypeptide chain 2 001 GGF Polypeptide chain 3 010 GFG Polypeptide chain 4 100 FGG Polypeptide chain 5 011 GFF Polypeptide chain 6 101 FGF Polypeptide chain 7 110 FFG Polypeptide chain 8 111 FFF

[0063] The nanopore detection signal corresponding to the peptide in this coding table is the basis for reading and decoding, so refer to Figure 7 The information-storing polypeptide chains obtained from Table 1 were added to the nanopore detection environment. Statistical analysis of the resulting current signals revealed the distribution of the encoded polypeptides in a scatter plot of the blocking current standard deviation (std) versus the blocking current degree (I / I0). Different colors represent the characteristic distribution of the polypeptides, and clearly distinguishable distributions were observed for the eight polypeptides. Subsequent observations of the same distribution at the same scatter plot position indicate the presence of the encoded polypeptide, allowing information to be read.

[0064] For example, in Figure 6 The content to be stored is the letter "H", which can be converted into binary code information according to the ASCII code as 001000. Since computers use the above ASCII code, which requires six-bit binary code, and since each information storage polypeptide chain represents three binary codes, two groups of information storage polypeptide chains are needed to jointly represent one character. Referring to Table 1, it can be seen that it is necessary to select information storage polypeptide chain 1 (hereinafter referred to as polypeptide chain 1) with an information storage unit sequence of GGG (binary sequence is 000) and an information storage polypeptide chain (hereinafter referred to as polypeptide chain 2) with an information storage unit sequence of GGF (binary sequence is 001) for information storage; and since 001 comes before 000 in the binary code, the address bit of the entire code "001" should be 0, and the address bit of the entire code "000" should be 1. The meaning of the address bits 0 and 1 here is different from the codes in the binary code table. In a specific embodiment, the assigned value of the polypeptide chain concentration is used as the address bit of the information to represent the meaning of 0 and 1. Since the polypeptide chain set contains two polypeptide chains, the concentration needs to be adjusted so that the signal frequency of polypeptide chain 2 (GGF) is higher than that of polypeptide chain 1 (GGG), so that the signal storage of 001000 can be achieved. Since the signal frequencies of polypeptide chains are not necessarily the same when the concentration is the same, in order to determine the concentration of the polypeptide chain that should be added during concentration encoding, the signal frequencies of polypeptide chain 1 and polypeptide chain 2 are first extracted. Figure 9 , place equal concentrations of polypeptide chain 1 and polypeptide chain 2 on Figure 2 In the information reading module based on nanopore detection, by applying voltage, the polypeptide chain enters the nanopore and moves relative to it, and the following can be observed: Figure 8 The extracted current signal contains a variety of characteristic parameters, such as blocking current, current standard deviation (std), blocking time, etc., which can be plotted as follows: Figure 9 The characteristic distribution of the polypeptide chains is determined by the scatter plot of the blocking current and the current standard deviation, that is, polypeptide chain 1 (GGG) is located at the center of the set with an I / I0 of 0.44, while polypeptide chain 2 (GGF) is located at the center of the set with an I / I0 of 0.41. In addition, the signal frequencies of different polypeptide chains can be calculated, that is, Figure 9 On the right, it can be seen that the signal frequency of polypeptide chain 1 is higher than that of polypeptide chain 2. According to the principle of polypeptide concentration encoding information (the encoding rule from high signal frequency to low signal frequency, in order to achieve the encoding of 001000, the signal frequency of polypeptide chain 2 (GGF) encoding 001 needs to be higher than that of polypeptide chain 1 (GGG) encoding 000. Since the frequency of polypeptide chain 1 is higher than that of polypeptide chain 2 at equal concentrations, the concentration of polypeptide chain 2 needs to be increased to make the frequency of polypeptide chain 2 higher than that of polypeptide chain 1), the encoding rule from high signal frequency to low signal frequency. It should be noted that in the actual storage process, it is necessary to try it out, and it is only necessary to ensure that the concentration of polypeptide chain 2 is increased so that the frequency of polypeptide chain 2 is higher than that of polypeptide chain 1. Exemplarily, the polypeptide set selected here for subsequent information storage is: the concentration ratio of polypeptide chain 1 to polypeptide chain 2 is polypeptide chain 1: polypeptide chain 2 = 1:3.

[0065] The above polypeptide collection, i.e. the concentration ratio of polypeptide chain 1: polypeptide chain 2 = 1:3, is placed Figure 2 In the information reading module based on nanopore detection, by applying voltage, multiple polypeptide chains are allowed to enter the nanopore and move relative to it, and the following can be observed: Figure 9 By extracting the current signal containing multiple characteristic parameters, such as blocking current, current standard deviation (std), etc., it is possible to draw the following Figure 10 Scatter plot of blocking current and current standard deviation, according to Figure 10 The characteristic distribution of polypeptides in the set is that polypeptide chain 1 (GGG) is located at the center point of the set with I / I0 of 0.44, while polypeptide chain 2 (GGF) is located at the center point of I / I0 of 0.41. It can be determined that the polypeptide set contains polypeptide chain 1 and polypeptide chain 2. At the same time, according to the encoding rules in Table 1, it can be known that the polypeptide set is 000 and 001. Then, through machine learning, the following information is obtained: Figure 10The number of signals from polypeptide chain 1 and polypeptide chain 2, which are characteristically distributed, indicates that the signal frequency of polypeptide chain 2 is higher than that of polypeptide chain 1. Therefore, the address of polypeptide chain 2 is assigned to 0, and the address of polypeptide chain 1 is assigned to 1. According to the sequence information encoding rules, the stored data of this polypeptide set is 001000, and the ASCII code shows that the stored information is the letter "H", which can be read.

[0066] Example 2

[0067] See also Figure 4 , since the encoding method of the information storage unit is: 1) select N kinds of amino acids as the basic components of the information storage unit, N kinds of amino acids can be assigned with numbers such as 0, 1, 2..., N-1. As a supplementary explanation, on the basis of the two polypeptides in the above specific embodiment 1, more kinds of amino acids are added to the information storage, see Figure 11 , the results of the information storage polypeptide chain RXD with the sequence N'-RRRRGGXGGDDDD-C' are given as an example, where X is glycine (G), lysine (K), phenylalanine (F), aspartic acid (E), and trimethylated lysine (K me3 The experimental results are derived from the results of equal concentrations of the aforementioned amino acid sequence passing through the nanopore. The characteristic parameters of the current signal, such as the blocking current degree (I / I0) and the blocking time (duration), can be extracted to plot the scatter plots. Figure 11 , determining the characteristic distribution of polypeptide chains. Based on the different color distributions, it can be seen that the nanopores are able to distinguish polypeptides with the aforementioned extended amino acid types.

[0068] Example 3

[0069] The following describes a method for storing information using a collection of information storage chains at different concentrations. If polypeptide chain 1 (GGG) is used as the information storage chain for data storage, the process should be as follows: first determine the signal frequency of different concentrations of polypeptide chains during nanopore detection. Figure 12As shown, each polypeptide concentration corresponds to a specific signal frequency. As the concentration increases, the signal frequency gradually increases. For example, 4 different polypeptide concentration values are selected for quaternary information encoding and storage, and are encoded as 0, 1, 2, and 3 in descending order according to the frequency corresponding to the concentration. For example, when the sample concentration is 25 μM, it is assigned to 0, 20 μM, 1, 10 μM, and 5 μM. The sample concentrations of different polypeptide samples are physically separated by a 96-well plate, and the samples are sampled and tested according to the physical spatial position, and the concentration of the sample corresponding to each position is read out in sequence to obtain its assignment information. During the information storage process, multiple peptide sets with different concentration values can be selected as units of information storage, such as 5μM (set 1), 10μM (set 2), 20μM (set 3), and 25μM (set 4). According to the signal frequencies generated, there is a relationship of set 4>set 3>set 2>set 1. Therefore, according to the frequency assignment from high to low, set 4 can be assigned 0, set 3 can be assigned 1, set 2 can be assigned 2, and set 1 can be assigned 3. Use physical partitions such as 96-well plates to save different sets in an independent and sequential space. During decoding, the samples can be tested in sequence according to the order of the space, and the concentrations in the independent space can be read out sequentially to obtain their assignment information.

Claims

1. A polypeptide-based information storage method, characterized in that: The polypeptide-based information storage method comprises the following steps: 1) Obtain the target information to be stored; 2) Converting the target information into N-ary coded information, where N is greater than or equal to 2; 3) Constructing a collection of information-storing polypeptide chains; 4) Sequentially assigning the N-ary coded information obtained in step 2) to the set of information storage polypeptide chains obtained in step 3) to complete the information storage of the target information to be stored.

2. The polypeptide-based information storage method according to claim 1, characterized in that: The collection of information storage polypeptide chains in step 3) includes a plurality of information storage polypeptide chains, each of which includes a positively charged guiding unit, an information storage unit, and a negatively charged guiding unit; the positively charged guiding unit is connected to the negatively charged guiding unit via the information storage unit; the positively charged guiding unit is a compound formed by a plurality of positively charged amino acids via peptide bonds; The negatively charged guiding unit is a compound formed by a plurality of negatively charged amino acids through peptide bonds; the information storage unit is an amino acid sequence formed by basic constituent monomers through peptide bonds.

3. The polypeptide-based information storage method according to claim 2, characterized in that: The amino acids in the positively charged guiding unit and the amino acids in the negatively charged guiding unit are both natural or unnatural amino acids.

4. The polypeptide-based information storage method according to claim 3, characterized in that: The specific implementation of step 3) is: 3.1) Selecting basic building blocks for constructing information storage units; 3.2) connecting the basic building blocks obtained in step 3.1) via peptide bonds to form an amino acid sequence; 3.3) Connecting a positively charged guiding unit and a negatively charged guiding unit to the front and back ends of the amino acid sequence, respectively, to form an information storage polypeptide chain; 3.4) Constructing a collection of information storage polypeptide chains based on multiple information storage polypeptide chains.

5. The polypeptide-based information storage method according to claim 4, characterized in that: The basic constituent monomers in step 3.1) are the same natural or non-natural amino acids or non-same natural or non-natural amino acids; When the basic constituent monomers in step 3.1) are the same natural or unnatural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of different concentrations formed from the same basic constituent monomers, and integrating the multiple information storage polypeptide chains of different concentrations to form a collection of information storage polypeptide chains; When the basic constituent monomers in step 3.1) are non-identical natural or non-natural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of the same concentration formed by non-identical basic constituent monomers, and integrating the multiple information storage polypeptide chains of the same concentration to form a collection of information storage polypeptide chains; When the basic constituent monomers in step 3.1) are non-identical natural or non-natural amino acids, the specific implementation method of step 3.4) is: obtaining multiple information storage polypeptide chains of non-identical concentrations formed by non-identical basic constituent monomers, and integrating the multiple information storage polypeptide chains of non-identical concentrations to form a collection of information storage polypeptide chains.

6. The polypeptide-based information storage method according to claim 5, characterized in that: The specific implementation of step 4) is: 4.1) Obtain the address bits of the N-ary coded information obtained in step 2); 4.2) All information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 3) are detected in the nanopore to obtain blocking current signals of all information storage polypeptide chains; 4.3) Based on the blocking current signals of all information storage polypeptide chains obtained in step 4.2), the signal frequencies of different information storage polypeptide chains are obtained respectively; 4.4) Arrange the signal frequencies obtained in step 4.3) in order of high to low to obtain a high to low arrangement result; 4.5) Mapping the high-low order results obtained in step 4.4) with the address bits obtained in step 4.1) from front to back to complete the storage of the target information to be stored.

7. The polypeptide-based information storage method according to claim 6, characterized in that: The different information storage polypeptide chains in step 4.2) are multiple information storage polypeptide chains of different concentrations formed from the same basic monomers, multiple information storage polypeptide chains of the same concentration formed from different basic monomers, or multiple information storage polypeptide chains of different concentrations formed from different basic monomers; The blocking current signal includes a blocking current, a current standard deviation, a blocking time, and a signal frequency.

8. A polypeptide-based information reading method, characterized in that: The polypeptide-based information reading method comprises the following steps: 1) Obtaining a collection of information storage polypeptide chains storing target information; 2) detecting the set of information storage polypeptide chains storing target information obtained in step 1) based on nanopore detection technology to obtain the blocking current signal frequency of all information storage polypeptide chains in the set of information storage polypeptide chains; the blocking current signal frequency of all information storage polypeptide chains in the set of information storage polypeptide chains is the number of current signals of a single information storage polypeptide chain in the set of information storage polypeptide chains per unit time; 3) Analyzing target information based on the blocking current signals of all information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 2).

9. The polypeptide-based information reading method according to claim 8, characterized in that: The collection of information storage polypeptide chains includes multiple information storage polypeptide chains, each of which includes a positively charged guiding unit, an information storage unit, and a negatively charged guiding unit; the positively charged guiding unit is connected to the negatively charged guiding unit via the information storage unit; the positively charged guiding unit is a compound formed by multiple positively charged amino acids through peptide bonds; The negatively charged guiding unit is a compound formed by a plurality of negatively charged amino acids through peptide bonds; the information storage unit is an amino acid sequence formed by basic constituent monomers through peptide bonds.

10. The polypeptide-based information reading method according to claim 8 or 9, characterized in that: The step 3) is specifically: 3.1) Based on the blocking current signals of all information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 2), the signal frequencies of all information storage polypeptide chains in the set of information storage polypeptide chains are parsed; 3.2) determining the address bit of the target information based on the signal frequencies of all information storage polypeptide chains in the set of information storage polypeptide chains obtained in step 3.1); 3.3) Obtain target information based on the mapping relationship between the address bits and the information storage polypeptide chain.