Data storage using proteins
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- THE HONG KONG POLYTECHNIC UNIV
- Filing Date
- 2023-07-20
- Publication Date
- 2026-05-27
AI Technical Summary
Existing peptide-based data storage methods face challenges such as high production costs, labor-intensive chemical synthesis, and limited storage density due to the need for address codes and error-correction overhead, which restricts the number of amino acid residues available for data storage.
A method for storing digital data in proteins by assigning specific binary sequences to amino acid residues, translating digital codes into amino acid sequences, and synthesizing proteins with peptide segments of up to 30 amino acid residues, using cell-based or cell-free expression systems to produce proteins with 100 to 2,000 amino acid residues.
This approach reduces production costs, increases data storage density, and enhances the stability and durability of stored data by utilizing proteins, which are more stable than DNA and can be expressed in various biological systems.
Smart Images

Figure CN2023108347_23012025_PF_FP_ABST
Abstract
Description
DATA STORAGE USING PROTEINS
[0001] REFERENCE TO SEQUENCE DISCLOSURE
[0002] Asequence listing file named “P25120PCT00_sequence_listing. xml” in ST. 26 format with a file size of 40kb created on July 19th, 2023 is incorporated herein by reference.TECHNICAL FIELD
[0003] The present invention relates to methods for digital data storage in proteins and data retrieval therefrom.BACKGROUND
[0004] Digital data storage in biomolecules such as DNA and peptides have been reported, including the present inventors’ U.S. patents under the patent numbers 11, 302, 421 (hereinafter as ‘421) and 11, 315, 203 (hereinafter as ‘203) . However, DNA normally has only 4 natural nucleotides, whereas unnatural DNA may contain 6 or even more nucleotides, and would require the enzymes for DNA sequencing. In addition, DNA is prone to degradation which makes it challenging for long-term data storage.
[0005] Compared with DNA, peptides can be selected from 20 canonical amino acids together with other non-canonical amino acids as monomers for data storage since peptide sequencing can be performed without enzyme recognition, thus in principle having a much higher density for data storage than DNA. Moreover, peptides in theory are less prone to degradation and more stable against temperature fluctuation than DNA. The previous patents ‘421 and ‘203 by the present inventors had shown that using chemically synthesized peptides could store and retrieve data including text files (848 bits) and a music file (13752 bits) with 40 and 511 18-mer peptides, respectively.
[0006] However, the peptide-based data storage and retrieval methods still face some problems such as high expected cost in mass production, labor-intensive chemical synthesis of peptides, and limited storage density by the number of amino acids in each peptide. In particular, the methods disclosed in ‘421 and ‘203 required incorporation of address codes and error-correction overhead in each data storage peptide, limiting the number of amino acid residues available for data storage. In ‘421 and ‘203, the N-and C-terminal amino acid selections of each data-bearing peptide sequence were subject to the proposed fragmentation and sequencing strategies, thereby hindering the data storage availability and design freedom of peptide sequence.
[0007] Aneed therefore exists for an improved method sharing the advantages of peptide-based data storage and retrieval methods while eliminating or at least diminishing the disadvantages and problems described above.SUMMARY OF INVENTION
[0008] Accordingly, a first aspect of the present invention provides a method for storing digital data in one or more proteins comprising:
[0009] providing the digital code that has been encoded from the digital data;
[0010] assigning each of a plurality of amino acid residues with a specific binary sequence composed of at least three bits each bit selected from “0” or “1” , wherein the plurality of amino acid residues is selected from at least eight different amino acid residues as data-bearing amino acid residues;
[0011] translating the digital code into an amino acid sequence according to the specific binary sequence assigned to each of the amino acid residues; and
[0012] synthesizing the one or more proteins each comprising a plurality of peptide segments, wherein each of the plurality of peptide segments comprises no more than 30 amino acid residues and the one or more proteins each contains 100 to 2,000 amino acid residues.
[0013] In certain embodiments, said synthesizing the one or more proteins is performed in a cell-based expression system comprising a biological cell selected from the group consisting of a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, and a mammalian cell.
[0014] In certain embodiments, said synthesizing the one or more proteins is performed in a cell-free expression system comprising a cell lysate.
[0015] In certain embodiments, the cell lysate is derived from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.
[0016] In certain embodiments, the specific binary sequence assigned to each of the data-bearing amino acid residues is “000” , “001” , “010” , “011” , “100” , “101” , “110” , or “111” .
[0017] In certain embodiments, the digital code further comprises at least one error-correction code prior to said translating the digital code into the amino acid sequence.
[0018] In certain embodiments, the at least one error-correction code comprises one or more of a repetition code, a convolutional code, a turbo code, a fountain code, a low-density parity-check (LDPC) code, a Reed–Solomon (RS) code, a Hadamard code, and a Hamming code, or any error-correction code disclosed in previous patents ‘421 and ‘203.
[0019] In certain embodiments, said synthesizing the one or more proteins comprises generating an expression vector comprising a nucleotide sequence encoding the amino acid sequence.
[0020] In certain embodiments, the expression vector comprises a prokaryotic expression vector and a eukaryotic expression vector.
[0021] In certain embodiments, each of the peptide segments comprises at least 6 amino acid residues.
[0022] In certain embodiments, each of the peptide segments comprises 16 to 21 amino acid residues, 18 to 21 amino acid residues, or any length of amino acid residues as long as structural requirements for a designated type of data storage protein are satisfied.
[0023] In certain embodiments, the one or more proteins have an average molecular weight of about 10 to 200 kDa.
[0024] In certain embodiments, the at least eight amino acid residues being the data-bearing amino acid residues are selected from the group consisting of leucine (L) , valine (V) , isoleucine (I) , glutamine (Q) , asparagine (N) , glutamic acid (E) , threonine (T) , serine (S) , alanine (A) , proline (P) , pyrrolysine (O) , histidine (H) , aspartic acid (D) , phenylalanine (F) , tyrosine (Y) , and glycine (G) .
[0025] In certain embodiments, phenylalanine (F) and arginine (R) are designated as N-and C-terminal amino acids of each of the peptide segments.
[0026] In certain embodiments, arginine (R) is substituted with lysine (K) at the C-terminus of each of the peptide segments.
[0027] In certain embodiments, at least one amino acid residue selected from the at least eight amino acid residues is designated in each of the peptide segments to indicate an address or order of the peptide segments in the protein.
[0028] In other embodiments, any two amino acid residues other than the at least eight amino acid residues are fixed at two positions in each of the peptide segments to indicate an address or order of the peptide segments in the protein.
[0029] In certain embodiments, each of the peptide segments comprises a collagen-like motif or a coiled-coil motif.
[0030] In certain embodiments, the plurality of peptide segments may not include transmembrane region or domain of the data storage protein.
[0031] In certain embodiments, each of the peptide segments comprises a peptide sequence of (Gly-Xaa-Yaa) n, wherein n is at least 4, Xaa and Yaa are independently or jointly selected from any one of the at least eight amino acid residues; or the plurality of peptide segments comprises a plurality of 7-amino acid peptides each comprising a peptide sequence of [ (R1) (Xaa) 2 (R1) (Xaa) 3] q, wherein R1 is independently I or L and q is 2 to 3.
[0032] A second aspect of the present invention provides a method for retrieving digital data from a data storage protein, where the method comprises:
[0033] providing the data storage protein or expressing the data storage protein from a storage medium, wherein the data storage protein comprises a plurality of peptide segments each independently comprising no more than 30 amino acid residues comprising at least eight different amino acid residues as data-bearing amino acid residues, wherein the data storage protein contains 100 to 2,000 amino acid residues;
[0034] sequencing the data storage protein to obtain amino acid sequences of the plurality of peptide segments;
[0035] converting the data-bearing amino acid residues into a plurality of binary sequences according to each specific binary sequence assigned to each of the data-bearing amino acid residues to obtain a digital code; and
[0036] decoding the digital code into digital data.
[0037] In certain embodiments, said expressing the data storage proteins from the storage medium comprises expressing the data storage protein from a cell-based expression system or a cell-free expression system.
[0038] In certain embodiments, the cell-based expression system comprises a biological cell selected from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.
[0039] In certain embodiments, the cell-free expression system is an mRNA-based translation system, a DNA-based transcription system, or a DNA-based translation system.
[0040] In certain embodiments, the cell-free expression system comprises a cell lysate.
[0041] In certain embodiments, the cell lysate is derived from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.
[0042] In certain embodiments, said sequencing comprises single-molecule protein sequencing, Edman degradation, or a mass spectrometry method.
[0043] In certain embodiments, the single-molecule protein sequencing comprises nanopore-based sequencing.
[0044] In certain embodiments, the mass spectrometry method comprises liquid chromatography with tandem mass spectrometry (LC-MS / MS) and matrix-assisted laser desorption / ionization time-of-flight / time-of-flight tandem mass spectrometry (MALDI-TOF / TOF MS) .
[0045] In certain embodiments, the data-storage protein has a molecular weight of about 10 to 200 kDa.
[0046] In certain embodiments, the present method further comprises checking the presence and order of one or more non-data bearing amino acid residues in each protein segment prior to said converting the data-bearing amino acid residues into the plurality of binary sequences according to each specific binary sequence assigned to each of the determined data-bearing amino acid residues to obtain a digital code.
[0047] In certain embodiments, the one or more non-data bearing amino acid residues in each protein segment comprises one or more amino acid residues corresponding to an error-correction code comprising one or more of repetition code, convolutional code, turbo code, fountain code, LDPC code, RS code, Hadamard code, and Hamming code, or any error-correction method disclosed in previous patents ‘421 and ‘203.
[0048] In certain embodiments, the specific binary sequence assigned to each of the data-bearing amino acid residues is composed of at least three bits each bit selected from “0” or “1” .
[0049] In certain embodiments, the specific binary sequence assigned to each of the data-bearing amino acid residues is selected from “000” , “001” , “010” , “011” , “100” , “101” , “110” , or “111” .
[0050] In certain embodiments, the at least eight amino acid residues being the data-bearing amino acid residues are selected from the group consisting of leucine (L) , valine (V) , isoleucine (I) , glutamine (Q) , asparagine (N) , glutamic acid (E) , threonine (T) , serine (S) , alanine (A) , proline (P) , pyrrolysine (O) , histidine (H) , aspartic acid (D) , phenylalanine (F) , tyrosine (Y) , and glycine (G) .
[0051] In certain embodiments, phenylalanine (F) and arginine (R) are designated as N-and C-terminal amino acids of each of the peptide segments.
[0052] In certain embodiments, arginine (R) is substituted with lysine (K) at the C-terminus of each of the peptide segments.
[0053] In certain embodiments, at least one amino acid residue selected from the at least eight amino acid residues is designated in each of the peptide segments to indicate an address or order of the peptide segments in the data storage protein.
[0054] In other embodiments, any two amino acid residues other than the at least eight amino acid residues are fixed at two positions in each of the peptide segments to indicate an address or order of the peptide segments in the protein.
[0055] In certain embodiments, each of the plurality of peptide segments comprises a collagen-like motif or a coiled-coil motif.
[0056] In certain embodiments, the plurality of peptide segments may not include transmembrane region or domain of the data storage protein.
[0057] In certain embodiments, each of the plurality of peptide segments comprises a peptide sequence of (Gly-Xaa-Yaa) n, wherein n is at least 4, Xaa and Yaa are independently or jointly selected from any one of the at least eight amino acid residues; or the plurality of peptide segments comprises a plurality of 7-amino acid peptides each comprising a peptide sequence of [ (R1) (Xaa) 2 (R1) (Xaa) 3] q, wherein R1 is independently I or L and q is 2 to 3.
[0058] Other aspects of the present invention include a system for designing and synthesizing a data storage protein described herein, and retrieving a corresponding digital data from the data storage protein, a storage medium for said data storage protein, a kit for expressing data storage protein, an expression vector and / or template comprising an encoding sequence thereof.
[0059] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Other aspects of the present invention are disclosed as illustrated by the embodiments hereinafter.BRIEF DESCRIPTION OF DRAWINGS
[0060] The appended drawings, where like reference numerals refer to identical or functionally similar elements, contain figures of certain embodiments to further illustrate and clarify the above and other aspects, advantages and features of the present invention. It will be appreciated that these drawings depict embodiments of the invention and are not intended to limit its scope. The invention will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0061] FIG. 1 schematically depicts methods for storing digital data in and retrieving from the data storage protein according to an exemplary embodiment of the present invention;
[0062] FIG. 2A shows the textual data in both Chinese and English to be encoded into digital code and stored in the data storage protein according to certain embodiments of the present invention;
[0063] FIG. 2B shows an example of protein sequence (Template I) designed according to the bit sequence map of Table 1 for storing the textual data as shown in FIG. 2A;
[0064] FIG. 2C shows where certain amino acids (dark gray boxes) other than eight selected amino acid residues (light gray boxes) are inserted into each segment of the protein sequence of FIG. 2B as address code for subsequent analysis by liquid chromatography with tandem mass spectrometry (LC-MS / MS) ;
[0065] FIG. 2D shows a liquid chromatography-mass spectrometry (LC-MS) chromatogram of tryptic peptides from the protein expressed by a cell-based expression system according to the protein sequence as shown in FIG. 2B; the maximum error for each peak was set to 25 ppm;
[0066] FIG. 2E shows an MS / MS spectrum of one of the tryptic peptides from the protein of FIG. 2B for determining the sequence with the home-made software; the maximum error for each peak was set to 25 ppm; masses were corrected to at least five decimal places;
[0067] FIG. 3A shows an image of SDS-PAGE gel electrophoresis for samples prepared according to previous patents (P) and proteins with the protein sequence as shown in FIG. 2B (I) which were expressed by a bacterial cell expression system according to Example 1; M denotes protein ladders (in KDa) ;
[0068] FIG. 3B shows an image of SDS-PAGE gel electrophoresis for proteins with the protein sequence of Templates II-a and II-b designed according to Example 2 and as shown in FIG. 3D, which were expressed by a bacterial cell expression system;
[0069] FIG. 3C shows an image of SDS-PAGE gel electrophoresis for proteins with the protein sequence of Template III designed according to Example 3 and as shown in FIG. 3D, which were expressed by a bacterial cell expression system;
[0070] FIG. 3D shows a comparison table of protein sequence between Template II-a, Template II-b, and Template III designed according to Examples 2 and 3, respectively;
[0071] FIG. 4 schematically depicts one of the conventional error-correction methods that can be applied in the present invention to ensure data integrity;
[0072] FIG. 5 is a flow chart illustrating the other conventional error-correction method that can be applied in the present invention to ensure data integrity.
[0073] Skilled artisans will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been depicted to scale.DETAILED DESCRIPTION OF THE INVENTION
[0074] It will be apparent to those skilled in the art that modifications, including additions and / or substitutions, may be made without departing from the scope and spirit of the invention. Specific details may be omitted so as not to obscure the invention; however, the disclosure is written to enable one skilled in the art to practice the teachings herein without undue experimentation.
[0075] Some portions of the description which follows are explicitly or implicitly presented in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithms and functional or symbolic representations are the means used by those skilled in the data processing arts to convey most effectively the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities, such as electrical, magnetic or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.
[0076] Unless specifically stated otherwise, and as apparent from the following, it will be appreciated that throughout the present specification, discussions utilizing terms such as “storing” , “retrieving” , “encoding” , “decoding” , “translating” , “converting” , “mapping” , “adding” , “appending” , “including” , “generating” , “comparing” , “determining” , “indicating” , “detecting” , “communicating” , or the like, refer to the action and processes of a computer system, or similar electronic device, that manipulates and transforms data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission or display devices.
[0077] The present specification also discloses apparatus for performing the operations of the methods. Such apparatus may be specially constructed for the required purposes, or may include a computer or other computing device selectively activated or reconfigured by a computer program stored therein. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various machines may be used with programs in accordance with the teachings herein. Alternatively, the construction of more specialized apparatus to perform the required method steps may be appropriate. The structure of a computer will appear from the description below.
[0078] In addition, the present specification also implicitly discloses a computer program, in that it would be apparent to the person skilled in the art that the individual steps of the method described herein may be put into effect by computer code. The computer program is not intended to be limited to any particular programming language and implementation thereof. It will be appreciated that a variety of programming languages and coding thereof may be used to implement the teachings of the disclosure contained herein. Moreover, the computer program is not intended to be limited to any particular control flow. There are many other variants of the computer program, which can use different control flows without departing from the spirit or scope of the disclosure.
[0079] Furthermore, one or more of the steps of the computer program may be performed in parallel rather than sequentially. Such a computer program may be stored on any computer readable medium. The computer readable medium may include storage devices such as magnetic or optical disks, memory chips, or other storage devices suitable for interfacing with a computer. The computer readable medium may also include a hard-wired medium such as exemplified in the Internet system, or wireless medium such as exemplified in the GSM mobile telephone system. The computer program when loaded and executed on a computer effectively results in an apparatus that implements the steps of the preferred method.
[0080] In embodiments of the present disclosure, use of the term ‘server’ may mean a single computing device or at least a computer network of interconnected computing devices which operate together to perform a particular function. In other words, the server may be contained within a single hardware unit or be distributed among several or many different hardware units.
[0081] Term Definitions
[0082] As used herein, the terms “peptide” and “peptide sequence” refer to a string of the amino acid residues, wherein the order of the assembly of the amino acid is the peptide sequence in the peptide. As used herein, the term “amino acid” refers to an organic compound that comprises an amine group (-NH2) and a carboxyl group (-COOH) . Amino acids can be natural, unusual, unnatural or synthetic (e.g., an amino acid analog) , wherein examples include, but are not limited to, D or L optical isomers, and amino acid analogs and peptidomimetics. Further examples of naturally occurring amino acids can be, but are not limited to, alanine, arginine, asparagine, aspartic acid, asparagine (or aspartic acid) , cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine or valine. In another example, groups of unnatural amino acids can be, but are not limited to, beta-homo-amino acids, and N-alkyl amino acids, such as alpha-methyl and alpha-dimethyl amino acids. Further examples of unnatural, unusual occurring or synthetic amino acids include, but are not limited to, citrulline, hydroxyproline, norleucine, 3-nitrotyrosine, nitroarginine, ornithine, naphtylalanine, methionine sulfoxide, methionine sulfone, cyclohexylalanine, ring-substituted phenylalanine, tyrosine or tryptophan derivatives (including but not limited to, for example, cyano-phenylalanine, o-tyrosine, m-tyrosine, hydroxy-tryptophan, methoxy-tryptophan) , or halogen-labelled amino acids derivatives (including but not limited to, for example, fluoro / chloro / bromo / iodo-phenylalanine, fluoro / chloro / bromo / iodo-tryptophan, fluoro / chloro / bromo / iodo-alanine) .
[0083] Additional examples of unnatural amino acids include 3- (2-chlorophenyl) -alanine, 3-chloro-phenylalanine, 4-chloro-phenylalanine, 2-fluoro-phenylalanine, 3-fluoro-phenylalanine, 4-fluoro-phenylalanine, 2-bromo-phenylalanine, 3-bromo-phenylalanine, 4-bromo-phenylalanine, homophenylalanine, 2-methyl-phenylalanine, 3-methyl-phenylalanine, 4-methyl-phenylalanine, 2, 4-dimethyl-phenylalanine, 2-nitro-phenylalanine, 3-nitro-phenylalanine, 4-nitro-phenylalanine, 2, 4-dinitro-phenylalanine, 1, 2, 3, 4-tetrahydroisoquinoline-3-carboxylic acid, 1, 2, 3, 4-tetrahydronorharman-3-carboxylic acid, 1-naphthylalanine, 2-naphthylalanine, pentafluorophenylalanine, 2, 4-dichloro-phenylalanine, 3, 4-dichloro-phenylalanine, 3, 4-difluoro-phenylalanine, 3, 5-difluoro-phenylalanine, 2, 4, 5-trifluoro-phenylalanine, 2-trifluoromethyl-phenylalanine, 3-trifluoromethyl-phenylalanine, 4-trifluoromethyl-phenylalanine, 2-cyano-phenylalanine, 3-cyano-phenylalanine, 4-cyano-phenylalanine, 2-iodo-phenylalanine, 3-iodo-phenylalanine, 4-iodo-phenylalanine, 4-methoxyphenylalanine, 2-aminomethyl-phenylalanine, 3-aminomethyl-phenylalanine, 4-aminomethyl-phenylalanine, 2-carbamoyl-phenylalanine, 3-carbamoyl-phenylalanine, 4-carbamoyl-phenylalanine, m-tyrosine, 4-amino-phenylalanine, styrylalanine, 2-amino-5-phenyl-pentanoic acid, 9-anthrylalanine, 4-t-butyl-phenylalanine, 3, 3-diphenylalanine, 4, 4'-diphenylalanine, benzoylphenylalanine, α-methyl-phenylalanine, α-methyl-4-fluoro-phenylalanine, 4-thiazolylalanine, 3-benzothienylalanine, 2-thienylalanine, 2- (5-bromothienyl) -alanine, 3-thienylalanine, 2-furylalanine, 2-pyridylalanine, 3-pyridylalanine, 4-pyridylalanine, 2, 3-diaminopropionic acid, 2, 4-diaminobutyric acid, allylglycine, 2-amino-4-bromo-4-pentenoic acid, propargylglycine, 4-aminocyclopent-2-enecarboxylic acid, 3-aminocyclopentanecarboxylic acid, 7-amino-heptanoic acid, dipropylglycine, pipecolic acid, azetidine-3-carboxylic acid, cyclopropylglycine, cyclopropylalanine, 2-methoxy-phenylglycine, 2-thienylglycine, 3-thienylglycine, α-benzyl-proline, α- (2-fluoro-benzyl) -proline, α- (3-fluoro-benzyl) -proline, α- (4-fluoro-benzyl) -proline, α- (2-chloro-benzyl) -proline, α- (3-chloro-benzyl) -proline, α- (4-chloro-benzyl) -proline, α- (2-bromo-benzyl) -proline, α- (3-bromo-benzyl) -proline, α- (4-bromo-benzyl) -proline, α-phenethyl-proline, α- (2-methyl-benzyl) -proline, α- (3-methyl-benzyl) -proline, α- (4-methyl-benzyl) -proline, α- (2-nitro-benzyl) -proline, α- (3-nitro-benzyl) -proline, α- (4-nitro-benzyl) -proline, α- (1-naphthalenylmethyl) -proline, α- (2-naphthalenylmethyl) -proline, α- (2, 4-dichloro-benzyl) -proline, α- (3, 4-dichloro-benzyl) -proline, α- (3, 4-difluoro-benzyl-proline, α- (2-trifluoromethyl-benzyl) -proline, α- (3-trifluoromethyl-benzyl) -proline, α- (4-trifluoromethyl-benzyl) -proline, α- (2-cyano-benzyl) -proline, α- (3-cyano-benzyl) -proline, α- (4-cyano-benzyl) -proline, α- (2-iodo-benzyl) -proline, α- (3-iodo-benzyl) -proline, α- (4-iodo-benzyl) -proline, α- (3-phenyl-allyl) -proline, α- (3-phenyl-propyl) -proline, α- (4-t-butyl) -benzyl) -proline, α-benzhydryl-proline, α- (4-biphenylmethyl) -proline, α- (4-thiazolylmethyl) -proline, α- (2-thiophenylmethyl) -proline, α- (5-bromo-2-thiophenylmethyl) -proline, α- (3-thiophenylmethyl) -proline, α- (2-furanylmethyl) -proline, α- (2-pyridinylmethyl) -proline, α- (3-pyridinylmethyl) -proline, α- (4-pyridinylmethyl) -proline, α-allyl-proline, α-propynyl-proline, γ-benzyl-proline, γ- (2-fluoro-benzyl) -proline, γ- (3-fluoro-benzyl) -proline, γ- (4-fluoro-benzyl) -proline, γ- (2-chloro-benzyl) -proline γ- (3-chloro-benzyl) -proline, γ- (4-chloro-benzyl) -proline, γ- (2-bromo-benzyl) -proline, γ- (3-bromo-benzyl) -proline, γ- (4-bromo-benzyl) -proline, γ- (2-methyl-benzyl) -proline, γ- (3-methyl-benzyl) -proline, γ- (4-methyl-benzyl) -proline, γ- (2-nitro-benzyl) -proline, γ- (3-nitro-benzyl) -proline, γ- (4-nitro-benzyl) -proline, γ- (1-naphthalenylmethyl) -proline, γ- (2-naphthalenylmethyl) -proline, γ- (2, 4-dichloro-benzyl) -proline, γ- (3, 4-dichloro-benzyl) -proline, γ- (3, 4-difluoro-benzyl) -proline, γ- (2-trifluoromethyl-benzyl) -proline, γ- (3-trifluoromethyl-benzyl) -proline, γ- (4-trifluoromethyl-benzyl) -proline, γ- (2-cyano-benzyl) -proline, γ- (3-cyano-benzyl) -proline, γ- (4-cyano-benzyl) -proline, γ- (2-iodo-benzyl) -proline, γ- (3-iodo-benzyl) -proline, γ- (4-iodo-benzyl) -proline, γ- (3-phenyl-allyl-benzyl) -proline, γ- (3-phenyl-propyl-benzyl) -proline, γ- (4-t-butyl-benzyl) -proline, γ-benzhydryl-proline, γ- (4-biphenylmethyl) -proline, γ- (4-thiazolylmethyl) -proline, γ- (3-benzothienylmethyl) -proline, γ- (2-thienylmethyl) -proline, γ- (3-thienylmethyl) -proline, γ- (2-furanylmethyl) -proline, γ- (2-pyridinylmethyl) -proline, γ- (3-pyridinylmethyl) -proline, γ- (4-pyridinylmethyl) -proline, γ-allyl-proline, γ-propynyl-proline, trans-4-phenyl-pyrrolidine-3-Carboxylic acid, trans-4- (2-fluoro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-fluoro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-fluoro-Phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-chloro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-chloro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-Chloro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-bromo-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-bromo-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-bromo-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-methyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-methyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-methyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-nitro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-nitro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-nitro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (1-naphthyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-naphthyl) -pyrrolidine-3-carboxylic acid, trans-4- (2, 5-dichloro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2, 3-dichloro-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-trifluoromethyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-trifluoromethyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-trifluoro methyl-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-cyano-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-cyano-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-cyano-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-methoxy-phenyl) - pyrrolidine-3-carboxylic acid, trans-4- (4-methoxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-hydroxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-hydroxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-hydroxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2, 3-dimethoxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3, 4-dimethoxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (3, 5-dimethoxy-phenyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-pyridinyl) -pyrrolidine-3-carboxylic acid , trans-4- (3-pyridinyl) -pyrrolidine-3-carboxylic acid, trans-4- (6-methoxy-3-pyridinyl) -pyrrolidine-3-carboxylic acid, trans-4- (4-pyridinyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-thienyl) -pyrrolidine-3-carboxylic acid, trans-4- (3-thienyl) -pyrrolidine-3-carboxylic acid, trans-4- (2-furanyl) -pyrrolidine-3-carboxylic acid, trans-4-isopropyl-pyrrolidine-3-carboxylic acid, and 4-phosphonomethyl-phenylalanine.
[0084] As used herein the term “amino acid analogs” refers to synthetic or semi-synthetic analogs of naturally occurring and unnatural amino acids, which include one or more chemical modifications, including, but not limited to hydrolysis, hydroxylation, alkoxylation, alkylation, homo-analogs, nor-analogs, cyclic analogs, arylation, nitration, halogenation, N-, O-, or S-alkylation, N-, O-, or S-acylation, dehydrogenation, oxidation, reduction, and decaroboxylation.
[0085] As used herein, the term “hydrophilic amino acids” refers to naturally occurring or unnatural amino acids having a hydrophilic side chain; such side chain may be uncharged, positively (cationic) or negatively charged (anionic) under normal physiological conditions, in particular at about pH 7.4. Uncharged hydrophilic amino acids can include, but are not limited to, serine, threonine, asparagine and glutamine; positively charged hydrophilic amino acids can be, but are not limited to, arginine, histidine and lysine, and the unnaturally occurring amino acid ornithine; negatively charged amino acids include aspartic acid and glutamic acid.
[0086] “Hydrophobic amino acids” refers to naturally occurring or unnatural amino acids having a hydrophobic and / or aromatic side chain, such as, but not limited to, alanine, valine, leucine, isoleucine, and the aromatic amino acids phenylalanine, tyrosine and tryptophan. Further included are unnaturally occurring amino acids such as, but not limited to, cyclohexylalanine.
[0087] As used herein, the term “peptide” refers to a polymer of 2 or more amino acids, amino acid analogs, peptidomimetics, or any combinations thereof. The subunits can be linked by peptide bonds. In another example, the subunits may be linked by other bonds such as, but not limiting to ester or ether bonds. Peptides also have a structure, wherein the structure can be generally understood to be linear or cyclic. In one example, a peptide can be, but is not limited to, dipeptides, tripeptides, oligopeptides or polypeptides. In another example, the peptide can be about 2 to 100 amino acids in length. In yet another example, the peptide can be about 2 to about 10 amino acids in length, about 5 to about 15 amino acids in length, about 10 to about 20 amino acids in length, about 15 to about 25 amino acids in length, about 20 to about 30 amino acids in length, about 25 to about 35 amino acids in length, about 30 to about 40 amino acids in length, about 35 to about 45 amino acids in length, about 40 to about 50 amino acids in length, about 45 to about 55 amino acids in length, about 50 to about 60 amino acids in length, about 55 to about 65 amino acids in length, about 60 to about 70 amino acids in length, about 65 to about 75 amino acids in length, about 70 to about 80 amino acids in length, about 75 to about 85 amino acids in length, about 80 to about 90 amino acids in length, about 85 to about 95 amino acids in length, about 90 to about 100 amino acids in length. In yet another example, the peptide can be 18 amino acids in length. Shorter peptides are cheaper to synthesize, and easier to be sequenced with reduced missed fragmentation, while longer peptides can store more data per peptide, reduce the number of peptides required for analysis, and reduce the address and error correction overhead.
[0088] As used herein, one or more amino acids in a peptide sequence represents a pattern of bits / symbols. A sequence of amino acids in a peptide sequence therefore represents a string of bits / symbols, which in turn represents part of the digital data. Therefore, the representation of digital data by peptide sequences is referred to herein as the digital data being stored in peptide sequences for simplicity sake. Since the digital data is expressed in a different form or system, the representation of digital data by peptide sequences can also be described as the digital data being encoded in peptide sequences. Therefore, the storing of digital data in peptide sequences is the same as the encoding of digital data in peptide sequences.
[0089] As used herein, the term “protein” refers to a molecule composed of amino acids and optionally other substance (s) such as carbohydrates, carbon, hydrogen, oxygen, nitrogen and sometimes sulfur. In some instances, a protein is a complex of multiple peptides with identical or different amino acid sequences each being optionally associated with one or more of the other substances than peptides. Protein structures can be defined by the sequence of amino acids (primary structure) , dihedral angles of peptide bonds among different amino acids (secondary structure) , folding of protein chains in space (tertiary structure) and association of different polypeptide molecules to form functional proteins (quaternary structure) . For simplicity, the “proteins” used for data storage or retrieval of data therefrom are based on their primary structure, i.e., the amino acid sequence of the proteins, rather than other protein structures as described herein. However, in various embodiments, some amino acid sequences representing regions such as transmembrane structures should be avoided.
[0090] As used herein, and analogous to the one or more amino acids in the peptide sequence as described herein, one or more amino acids in a protein sequence represents a pattern of bits / symbols. A sequence of amino acids in a protein sequence therefore represents a string of bits / symbols, which in turn represents part of the digital data. Therefore, the representation of digital data by protein sequences is referred to herein as the digital data being stored in protein sequences for simplicity sake. Since the digital data is expressed in a different form or system, the representation of digital data by protein sequences can also be described as the digital data being encoded in protein sequences. Therefore, the storing of digital data in protein sequences is the same as the encoding of digital data in protein sequences.
[0091] The following examples provide some proof-of-concept experiments, in which each data-bearing protein consists of peptide sequences with a specific design that could facilitate proteolytic cleavage and MS / MS sequencing. In the examples, data-bearing peptides are concatenated together to form proteins. The address codes are employed to ensure data retrieval. However, address codes are not required for data storage proteins which are sequenced by single-molecule sequencing method such as nanopore-based sequencing method. Variable length peptide segments are applied for future reduction of the address and error-correction overheads. To facilitate the expression and purification of certain data-storing proteins (e.g., protein templates I, II-a and II-b described herein) , transmembrane regions may be avoided by predicting if the structure of the proteins would contain such regions using web-based transmembrane prediction software. If transmembrane regions can be predicted during de novo design stage, then the protein would be re-encoded. FIG. 1 schematically depicts how digital data is basically stored in and retrieved from data storage proteins according to certain embodiments of the present invention.
[0092] EXAMPLES
[0093] Example 1 –Protein Template I
[0094] 1.1. Design, Synthesis and Data Storage of Protein Template I:
[0095] In accordance with the proposed data storage and retrieval schemes depicted in FIG. 1, raw data is first converted to binary digits. By assigning amino acids as specific sequences of digital bits, the binary digits are translated into amino acid sequences, which are incorporated into protein sequences with pre-designed templates. The data-bearing proteins are then expressed through cell-based systems or cell-free systems. To retrieve data, protein sequencing techniques are used for accurate sequence readout. Finally, the protein sequences obtained are converted back to binary digits and decoded back to the raw data.
[0096] In this example, proteins with pre-designed template were expressed in E. coli, and after expression the proteins were purified using immobilized metal affinity chromatography (IMAC) . The Chinese and English texts as shown in FIG. 2A were translated into a protein sequence with a total bit size of 1088 bits according to the specific binary sequence assigned to each of eight selected amino acid residues: S, T, E, Y, D, G, N and Q. Table 1 below provides each of the specific binary sequence (in 3 bits each) assigned to these eight amino acid residues.
[0097] Table 1 –Binary Sequences Assigned to Eight Selected Amino Acid Residues in Template I: *the IUPAC-IUB Joint Commission on Biochemical Nomenclature. Nomenclature and Symbolism for Amino Acids and Peptides.
[0098] Different from previous peptide-based storage method disclosed in ‘421 and ‘203, eight data-bearing amino acids were selected based on their hydrophobicity when designing the protein sequence de novo in this example. Table 2 shows the difference in selection of amino acid residues based on the consideration in this example compared with the previous peptide-based data storage method.
[0099] Table 2
[0100] As shown in FIG. 2B, the protein as designed consisted of 40 different peptide segments each with 16 to 20 amino acids (the amino acid sequence of the protein is represented by SEQ ID No: 41) . Since MS / MS was used in this example for sequencing, each peptide segment contained an address code and also the same N-and C-terminal amino acids (F and R, respectively) as the other segment. Details of where to insert the address code in each of those peptide segments is depicted in the table of FIG. 2C.
[0101] It is generally considered that proteins with highly hydrophobic regions / domains could reduce protein yield or production level. Therefore, a rational design of protein based on an optimal hydrophobicity of the protein backbone is one of the feasible ways to increase protein yield or production level. In certain embodiments, the data storage protein is generally hydrophilic and relatively more hydrophilic amino acids instead of relatively more hydrophobic amino acids are selected as the data-bearing amino acid residues. For example, binary sequence “011” in protein template I substitutes valine (V) used in the prior patents ‘421 and ‘203 into asparagine (N) because N is relatively more hydrophilic whereas V is very hydrophobic among 20 canonical amino acids. The backbones could be made through using algebraic equations to specify the geometry parametrically and building mathematical model. Preferably, the backbones construction model should allow tight packing between the amino acid side chains, satisfy the hydrogen-bonding potential of the protein backbone and have little strain in the backbone torsion angles. By understanding the rules of de novo protein design, it is possible to design templates to form well-folded, thermodynamically stable proteins or even proteins that perform various functions. In addition, the advanced approaches of protein structure prediction, e.g., AlphaFold or DeepTMHMM, could be used to assist protein design by confirming the protein structure and avoiding the undesired regions.
[0102] Corresponding encoding sequence of the protein template was then generated and codon-optimized before being inserted into an expression plasmid / vector. An error-correction code, e.g., the error-correction code disclosed in prior U.S. patents ‘421 and ‘203, where the disclosures of which are incorporated herein by reference, i.e., any one or combination of repetition code, convolutional code, turbo code, fountain code, low-density parity-check (LDPC) code, Reed-Solomon (RS) code, Hadamard code, and Hamming code, or those generated based on the digital code, or both the digital code and order-checking bits added into the digital code (if any) , could be incorporated into the protein sequence before corresponding gene insert including the encoding sequence of the protein template was generated. Suitable error-correction methods provide a capability of recovering the original digital data when errors occur in the data storage and retrieval processes. One of the error-correction methods proposed in ‘421 and ‘203 is schematically depicted in FIG. 4. The term “starting digital code” used therein referred to the digital code generated from the original digital data. The starting digital code proposed in ‘421 and ‘203 was a digital code without any error-correction codes. The proposed error-correction method added two redundant bits to each bit in the starting digital code, i.e., the error-correction method added redundancy to the starting digital code. For example, bit 0 in the starting digital code was added with additional 0 bits, while bit 1 in the starting digital code was added with addition 1 bits. The digital code with the error-correction code in that method was transformed from (0, 1) into (000, 111) . The term “redundancy” used therein referred to extra data that was generated for repetition of information or inclusion of additional information during the data storage / retrieval / transfer. By adding redundancy, error correction and detection could be advantageously achieved. That proposed error-correction method by ‘421 and ‘203 allows an error in any of the triplet of bits to be corrected by “majority vote” . That error-correction method also allows up to 2 bits of triplet omitted.
[0103] Another error-correction method proposed by ‘421 and ‘203 was based on order-checking bits, and one or more LDPC codes or one or more RS codes which were designed to correct errors during the synthesis, detection and sequencing of the peptides. For a general principle of that another error-correction method based on order-checking bits and one or more LDPC codes, a flow chart in FIG. 5 illustrates how it could be achieved. In one arrangement of that error-correction method, the order-checking bits were added into the starting digital code as redundant bits, which could contain information of the correct order of certain bits / symbols in the starting digital code. The value of the order-checking bits could be determined according to user-defined rules, for example, the order-checking bit of “1” was added to a symbol if the subsequent symbol had a lesser value. For example, a starting digital code had symbols of 32. An order-checking bit was then added according to the above example user-defined rules. As the first symbol had a greater value (i.e., “3” representing 3 bits “011” ) than the second symbol (i.e., “2” representing 3 bits “010” ) , a redundant bit “1” was added as an order-checking bit. The same process could be repeated for the second symbol to add an order-checking bit. Therefore, there could be one or more order-checking bits generated from the starting digital code. In one example, the generated order-checking bits could be added to the starting digital code and became part of the digital code before generating the one or more LDPC codes based on the starting digital code and the order-checking bits. LDPC code is a linear error-correction code constructed using a sparse bipartite graph. The encoding of LDPC codes into the digital code comprised: (1) constructing a sparse parity-check matrix, and (2) generating codewords with the matrix. A codeword contained information bits and parity bits, in which the parity bits were the redundant bits appended to the information bits. A parity bit was set to either 0 or 1, depending on the total number of the 1-bits in some of the information bits either even or odd, which could advantageously be used to detect and / or correct errors in the information bits. Alternatively, error-correction codes could include any one or combination of repetition code, convolutional code, turbo code, fountain code, LDPC code, RS code, Hadamard code, and Hamming code. Although the digital data was encoded in a digital code having a block of bits or symbols, the encoded digital code could be in any form determined by the user.
[0104] As an illustrative example of that error-correction method, digital data was encoded in a starting digital code with 850 information bits, b = {b1, b2, …, b850} . The starting digital code was to be processed using error-correction methods and then stored into 40 16-mer peptide sequences. In other words, a starting digital code of 850 information bits together with its error-correction codes was to be translated into 40 peptide sequences, each sequence having 16 amino acids, and each amino acid represented a 3-bit pattern or a symbol (i.e., S1, S2, …, S16) . Alternatively, the starting digital code may be reformatted into a matrix of sequences before adding the error-correction codes, in which each sequence (i.e., Seq #1, Seq #2, …, Seq #40) includes 16 symbols where 2 of the 16 symbols were used as an address pair (i.e., {Ai, 1, Ai, 2} , i = 1, 2, …, 40) , where i was the sequence number. The address pair may occupy the first 2 symbols of each peptide sequence, or any 2 symbols of each peptide sequence. The address pair of a peptide sequence (e.g., Seq #1) was used to indicate the position of the peptide sequence (e.g., Seq #1) in the order of the peptide sequences (e.g., Seq #1, Seq #2, etc. ) . For example, a peptide sequence Seq #j may have an address pair value of A1, 1 of 000 and A1, 2 of 001, which was combined to provide an address value of 000001. Further, the peptide sequence Seq #k may have an address pair value of A2, 1 of 000 and A2, 2 of 000, which was combined to provide an address value of 000000. Therefore, when retrieving the data from the synthesized peptide by sequencing, the address pairs of the peptide sequences indicated the correct order of the peptide sequences, and in turn allowed the digital code (inclusive of any error-correction codes) to be reconstructed based on order indicated by the address pairs. The addressing of the peptide sequences may also use more or less than 2 symbols of a peptide sequence.
[0105] The gene insert after customization including the encoding sequence of the protein template I and other address, error-correction, and order-checking codes / bits was cloned into the pET28a (+) vector via NdeI / XhoI site to generate pET28a (+) -PolyU construct. The recombinant plasmid was purchased from Sangon Biotech Corporation. Genetic construct, pET28a (+) -PolyU, was transformed into E. coli BL21 (DE3) strain for storage, transportation, or subsequent protein expression and other analysis.
[0106] 1.2. Expression of Protein Template I and Sequencing by LC-MS / MS:
[0107] Recombinant protein production was performed by inoculating a single colony into 10 mL LB media containing kanamycin and cultured at 37 ℃, 250 rpm overnight. The overnight culture was used to inoculate 200 mL culture that was incubated 37 ℃, 250 rpm until the cells were grown to an OD600 nm of approximately 0.6. Protein expression was induced with the addition of 0.5 mM isopropyl β-D-1-thiogalactopyranoside (IPTG) (BBI) , and the cells were grown at 37 ℃, 250 rpm for a further 6h. Cells were harvested by centrifugation (4000 rpm, 4 ℃, 20 min) . Cell pellets were resuspended in ice cold PBS with 1mM PMSF. Samples were sonicated using Branson SFX550 sonifier. The lysate was centrifuged (12,000 rpm, 4 ℃, 20 min) , and the supernatant was collected. The protein was purified using HisTrap HP (Cytiva) on an Pure FPLC (Cytiva) and dried using refrigerated CentriVap vacuum concentrator (Labconco) at 4 ℃ overnight.
[0108] The frozen protein pellet was dissolved with 50 mM ammonium bicarbonate and 10 mM of dithiothreitol (DTT) and incubated for 30 minutes at 37 ℃. The sample was incubated in dark with 20mM iodoacetamide (IAA) for 15 minutes. Trypsin (Promega) was added to the protein sample (protein to trypsin ratio = 50: 1) , and the solution was incubated overnight at 37 ℃. Trypsin digestion was stopped by adding 0.1%formic acid. The digest was desalted using C18 spin column (ThermoFisher Scientific) and vacuum-dried to obtain a digestion mixture.
[0109] The digestion mixture was then analyzed by liquid chromatography with tandem mass spectrometry (LC-MS / MS) , where the acquired MS / MS spectra were processed by a self-developed software for assignment of amino acid sequences (The result of TIC is shown in FIG. 2D, while an example of sequence assignment from an MS / MS spectrum is shown in FIG. 2E) . The LC-MS / MS analysis was performed according to the following protocol. Initially, the tryptic peptides were dissolved in 50%acetonitrile and 0.1%formic acid. The peptide mixtures were then separated using a Waters Acquity UPLC system with a C18 column (Agilent AdvanceBio Peptide Map, 2.1 × 150 mm, 2.7 μm particle size, pore size) . Mobile phase A was 0.2%formic acid in water and B was 0.2%formic acid in acetonitrile. The flow rate was 0.3 mL / min and the temperature was 55 ℃. The gradient changed from 10%B to 18%B at 0 to 2 min, from 18%B to 22%B at 2 to 8 min, from 22%B to 34%B at 8 to 48 min, from 34%B to 40%B at 48 to 64 min, from 40%B to 55%B at 64 to 75 min, from 55%B to 80%B at 75 to 78 min, and remained at 80%B from 78 to 83 min.
[0110] MS / MS analysis was performed using an Orbitrap Fusion Lumos mass spectrometer (ThermoFisher Scientific) operated in positive ion mode. The spray voltage for electrospray ionization was +3600 V, and both ion transfer tube temperature and vaporizer temperature were 280 ℃. In each cycle, a MS1 scan with m / z from 900 to 1400 Da was performed with a resolution of 30 K. Ions were selected for MS / MS with quadruple, using advanced peak determination (APD) with default charge of +2, top-speed mode with 3 s cycles, mass tolerance of 25 ppm, dynamic exclusion window of 4 s, and isolation window width of 1.6 or 0.7 Da. High-energy collision dissociation (HCD) at 28%of normalized collision energy with stepped collision energy of 5%was used for the fragmentation. MS / MS spectra were obtained with m / z from 240 to 2450 Da and a resolution of 15 K.
[0111] Table 3 below summarizes the sequencing results by LC-MS / MS for 40 different data-bearing peptides from the expressed protein by E. coli, each in a range of 14-18 amino acid residues (excluding the N-and C-terminal amino acids and address code, if any) .
[0112] Table 3
[0113] From the results in Table 3, about 91.9%of the amino acids were correctly recovered for their sequences, and by decoding with the proposed error-correction scheme (with a tolerance of 10%error employed to protect data integrity) , the original data could be fully retrieved, i.e., achieving almost 100%correctness. These preliminary results demonstrate the feasibility of storing data in and retrieving data from the proposed data storage proteins.
[0114] Example 2 –Design, Synthesis and Expression of Protein Templates IIa and IIb
[0115] Different from template I in Example 1, templates II-a and II-b in this example were designed to enhance yield or production level and protein stability, since the corresponding yield or production level of template I by E. coli is not high nor yield is low after purification (FIG. 3A) . When designing sequence of template II-a or II-b de novo, the sequence pattern was commensurate with that of collagen, one of the most stable proteins, i.e., the typical collagen-like (Gly-Xaa-Yaa) n sequences. A combination selected from twelve data-bearing amino acids (L, Q, N, D, E, T, V, S, Y, H, A and P) was selected according to the composition of collagen. Optionally, P can be substituted with O; L can be substitute with I. Therefore, the sequence pattern of template II-a or II-b can be represented by (Gly-Xaa-Yaa) n, where n is at least 4; Xaa and Yaa are independently or jointly selected from any one of the eight data-bearing amino acids according to certain embodiments. An amino acid sequence of the template II-ais represented by SEQ ID No: 42. An amino acid sequence of template II-b is represented by SEQ ID No: 43. Optionally, the sequence pattern may be commensurate with other collagen isoforms which are stable, and many of them may have hydroproline (Hyp) for Yaa in the foregoing formula for the typical collagen-like sequences. In that case, the sequence pattern may be (Gly-Xaa-Hyp) n, where Xaa is any one of the eight data-bearing amino acids. Corresponding binary sequences were assigned to different selected data-bearing amino acids, and then the digital data of interest were encoded into a protein where each peptide segment has a length of 3n amino acids. Since LC-MS / MS was used for sequencing, N-and C-termini of peptide segments were fixed as F and R, respectively; G was fixed at certain positions to improve protein yield and production; 2 amino acids selected from the 12 data-bearing amino acids, e.g., V and D, were used to indicate the address or order of peptide segments in this example. In the case where the protein is sequenced by any method involving trypsin digestion into peptide segments, R or K could be selected at the C-terminal amino acid. The incorporation of certain amino acids at different positions in each peptide segment in this example indicates the address (or order) of the plurality of peptide segments in the data storage protein. For example, V and D were used in protein template II-a (SEQ ID No: 42) to indicate the order of the peptide segments by putting these two amino acid residues in random positions of each peptide segment. More specifically, V and D were inserted at 4th and 6th positions in the first peptide segment (FGNVGDQGSAGEPGEQGPSGR) , at 4th and 7th positions in the second peptide segment (FGSVGEDGEPGQLGTTGEAGR) , 4th and 9th positions in the third peptide segment (FGTVGEEGDSGNNGAAGLQGR) , and at 6th and 7th positions in the tenth peptide segment (FGNLGVDGNPGLPGSNGEQGR) , respectively, of the protein template II-a. In another example, S was used in protein template II-b (SEQ ID No: 43) to indicate the address of each peptide segment in the protein. In certain embodiments, none of order-recognition code and address code is required when other sequencing method is employed such as single-molecule protein sequencing. In other words, amino acid residues required for trypsin digestion prior to sequencing such as R or K, those for encoding the order-recognition code, or those designated or fixed at certain positions for indicating address / order of the peptide segments in the data storage protein, can all be used for data-bearing amino acid residues, so long as the protein yield or production is not affected.
[0116] The encoding sequence of protein template II-a or II-b was inserted into the same expression plasmid and transfected into the same bacterial host cell as in Example 1 for protein expression. However, it should be understood that any competent expression vector or plasmid and host cell can be used for protein expression of this template without limiting to the same plasmid / host cell as in the preceding example.
[0117] Example 3 –Design, Synthesis and Expression of Protein Template III
[0118] Similar to protein template II-a or II-b in Example 2, protein template III was designed de novo according to one of the known protein structures, coiled-coil components, or coiled-coil domain, which is a structural motif in proteins by coiling 2-7 alpha-helices together to form a coiled-coil structure. They usually contain a repeated pattern, hxxhcxc, where h denotes hydrophobic amino acid residue and c denotes charged amino acid residue. The positions in each heptad repeat are usually represented by abcdefg, where a and d denote isoleucine (I) , leucine (L) , or valine (V) . An amino acid sequence of template III is represented by SEQ ID No: 44. Proteins containing this type of structures usually involve in certain biological functions such as gene expression, viral infection and oligomerization. For example, in a landmark study of archetypal coiled coil, GCN4, it has a repeated I and L at the a and d positions of each heptad motif (equivalent to 4 heptad repeats in each of the two 31-amino-acid long alpha helices) . Therefore, another combination of twelve data-bearing amino acids (H, T, S, N, A, D, Q, G, P, Y and E) was selected for this kind of protein structure. Corresponding binary sequences were assigned to different selected data-bearing amino acids, and then the digital data of interest were encoded into a protein where each peptide segment has a length of 7n amino acids. Since LC-MS / MS was used for sequencing, N-and C-termini of peptide segments were fixed as F and R, respectively; I or L was fixed on certain positions to mimic the coiled-coil heptad repeat sequence. In the case where the protein is subjected to trypsin digestion prior to sequence, R could be substituted with K at the C-terminus. The coiled-coil-like heptad repeat in this example may be represented by an amino acid sequence of [ (R1) (Xaa) 2 (R1) (Xaa) 3] q, wherein R1 is independently I or L and q is 2 to 3. The incorporation of certain amino acids at different positions in this example also serves similar purposes to those mentioned in Example 2, i.e., to indicate address or order of each of the peptide segments in the corresponding data storage protein. In certain embodiments, none of order-recognition code and address code is required when other sequencing method is employed such as single-molecule protein sequencing according to certain embodiments. FIG. 3D shows the difference between the protein template III in this example and the protein templates II-a and IIb in Example 2 in terms of their protein sequence.
[0119] The coding sequence of protein template III was inserted into the same expression plasmid and transfected into the same bacterial host cell as in the preceding examples for protein expression, but it should not be considered limiting to those as described herein. It is also understood that cell-free expression system can also be used for expression of any protein templates described herein, as long as it is competent for enabling data density and integrity of the data storage protein of the present invention.
[0120] The four protein templates (I, II-a, II-b and III) in Examples 1, 2 and 3, respectively, were expressed in E. coli and analyzed by SDS-PAGE gel electrophoresis. Images of SDS-PAGE gel for each of the protein templates are shown in FIGs. 3A-3C, respectively, in which “M” denotes protein ladders; “P” denotes the sample prepared according to previous patents ‘421 and ‘203; “I” denotes the sample prepared by expression of protein template I; II-a denotes the sample by expression of protein template II-a; II-b denotes the sample by expression of protein template II-b; and III denotes the sample by expression of protein template III. As can be seen from FIG. 3A, the sample prepared according to the previous patents ‘421 and ‘203 and that by expression of protein template I had no band and a significantly low yield of protein, respectively. Comparatively, samples by expression of protein templates II-a, II-b and protein template III, respectively (FIGs. 3B and 3C) , had higher yield or production level than that by expression of protein template I. Also, the recovery rate of the amino acid sequence of protein templates II-a and II-b in Example 2 and that of protein template III in this example were 98%, 100%, and 100%, respectively, which were significantly improved compared to protein template I in Example 1 (only 91.9%) .
[0121] Although the invention has been described in terms of certain embodiments, other embodiments apparent to those of ordinary skill in the art are also within the scope of this invention. Accordingly, the scope of the invention is intended to be defined only by the claims which follow.INDUSTRIAL APPLICABILITY
[0122] The present invention has numerous advantages over the conventional data storage medium / source. Firstly, by expressing the protein instead of purely chemical synthesis, the cost of storing and duplicating data would be reduced. For example, it has been reported that 4256 E. coli proteins could be expressed and purified at once, and even de novo designed proteins could be expressed in this way. Secondly, the larger size of protein would reduce the address and error-correction overhead in the encoding step, achieving higher data density. Thirdly, the de novo designed proteins would allow stable protein structure design which could improve the data durability. Fourthly, living organism could be used as storage devices, which would open many new possibilities. For example, given data-bearing proteins could be expressed under specific inducing conditions and be transported to certain location through genetic engineering. Fifthly, various techniques developed for protein labeling, modification, etc. can be utilized to perform specific functions, such as selective retrieval of information, cryptography or steganography. In addition, with developments in proteomics, protein sequencing has become routine. Single-molecule protein sequencing simplifies the protein sequence design and avoids the pre-processing steps such as typical trypsin digestion into peptides before tandem mass spectrometry (MS / MS) , allowing more flexible protein sequence design, higher throughput and higher data density in proteins than peptide-based data storage method. Finally, it is a low-energy consumption method for information storage, duplication and transmission.
Claims
1.A method for storing digital data in one or more proteins comprising:providing the digital code that has been encoded from the digital data;assigning each of a plurality of amino acid residues with a specific binary sequence composed of at least three bits each bit selected from “0” or “1” , wherein the plurality of amino acid residues is selected from at least eight different amino acid residues as data-bearing amino acid residues;translating the digital code into an amino acid sequence according to the specific binary sequence assigned to each of the amino acid residues; andsynthesizing the one or more proteins each comprising a plurality of peptide segments, wherein each of the plurality of peptide segments comprises no more than 30 amino acid residues and the one or more proteins each contains 100 to 2,000 amino acid residues.2.The method of claim 1, wherein said synthesizing the one or more proteins is performed in a cell-based expression system comprising a biological cell selected from the group consisting of a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, and a mammalian cell.3.The method of claim 1, wherein said synthesizing the one or more proteins is performed in a cell-free expression system comprising a cell lysate.4.The method of claim 3, wherein the cell lysate is derived from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.5.The method of claim 1, wherein the specific binary sequence assigned to each of the data-bearing amino acid residues is “000” , “001” , “010” , “011” , “100” , “101” , “110” , or “111” .6.The method of claim 1, wherein said synthesizing the one or more proteins comprises generating an expression vector comprising a nucleotide sequence encoding the amino acid sequence.7.The method of claim 6, wherein the expression vector comprises a prokaryotic expression vector and a eukaryotic expression vector.8.The method of claim 1, wherein the at least eight amino acid residues being the data-bearing amino acid residues are selected from the group consisting of leucine (L) , valine (V) , isoleucine (I) , glutamine (Q) , asparagine (N) , glutamic acid (E) , threonine (T) , serine (S) , alanine (A) , proline (P) , pyrrolysine (O) , histidine (H) , aspartic acid (D) , phenylalanine (F) , tyrosine (Y) , and glycine (G) .9.The method of claim 8, wherein at least one amino acid residue selected from the at least eight amino acid residues is designated in each of the plurality of peptide segments to indicate an address or order of the corresponding peptide segment in each of the one or more proteins.10.The method of claim 8, wherein any two amino acid residues other than the at least eight amino acid residues are fixed at two positions in each of the peptide segments to indicate an address or order of the peptide segments in the protein.11.The method of claim 8, wherein each of the plurality of peptide segments comprises a collagen-like motif or a coiled-coil motif.12.The method of claim 11, wherein each of the plurality of peptide segments comprises a peptide sequence of (Gly-Xaa-Yaa) n, wherein n is at least 4, Xaa and Yaa are independently or jointly selected from any one of the at least eight amino acid residues; or the plurality of peptide segments comprises a plurality of 7-amino acid peptides each comprising a peptide sequence of [ (R1) (Xaa) 2 (R1) (Xaa) 3] q, wherein R1 is independently I or L and q is 2 to 3.13.A method for retrieving digital data from a data storage protein comprising:providing the data storage protein or expressing the data storage protein from a storage medium, wherein the data storage protein comprises a plurality of peptide segments each independently comprising no more than 30 amino acid residues comprising at least eight different amino acid residues as data-bearing amino acid residues, wherein the data storage protein contains 100 to 2,000 amino acid residues;sequencing the data storage protein to obtain amino acid sequences of the plurality of peptide segments;converting the data-bearing amino acid residues into a plurality of binary sequences according to each specific binary sequence assigned to each of the data-bearing amino acid residues to obtain a digital code; anddecoding the digital code into digital data.14.The method of claim 13, wherein said expressing the data storage proteins from the storage medium comprises expressing the data storage protein from a cell-based expression system or a cell-free expression system.15.The method of claim 14, wherein the cell-based expression system comprises a biological cell selected from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.16.The method of claim 14, wherein the cell-free expression system is an mRNA-based translation system, a DNA-based transcription system, or a DNA-based translation system.17.The method of claim 16, wherein the cell-free expression system comprises a cell lysate.18.The method of claim 17, wherein the cell lysate is derived from a bacterial cell, an archaeal cell, a plant cell, a yeast cell, an insect cell, or a mammalian cell.19.The method of claim 13, wherein said sequencing comprises single-molecule protein sequencing, Edman degradation, or a mass spectrometry method.20.The method of claim 19, wherein the single-molecule protein sequencing comprises nanopore-based sequencing.21.The method of claim 19, wherein the mass spectrometry method comprises liquid chromatography with tandem mass spectrometry (LC-MS / MS) and matrix-assisted laser desorption / ionization time-of-flight / time-of-flight tandem mass spectrometry (MALDI-TOF / TOF MS) .22.The method of claim 13, wherein the specific binary sequence assigned to each of the data-bearing amino acid residues is composed of at least three bits each bit selected from “0” or “1” .23.The method of claim 13, wherein the specific binary sequence assigned to each of the data-bearing amino acid residues is selected from “000” , “001” , “010” , “011” , “100” , “101” , “110” , or “111” .24.The method of claim 13, wherein the at least eight amino acid residues being the data-bearing amino acid residues are selected from the group consisting of leucine (L) , valine (V) , isoleucine (I) , glutamine (Q) , asparagine (N) , glutamic acid (E) , threonine (T) , serine (S) , alanine (A) , proline (P) , pyrrolysine (O) , histidine (H) , aspartic acid (D) , phenylalanine (F) , tyrosine (Y) , and glycine (G) .25.The method of claim 24, wherein at least one amino acid residue selected from the at least eight amino acid residues is designated in each of the plurality of peptide segments to indicate an address or order of the corresponding peptide segment in the data storage protein.26.The method of claim 24, wherein any two amino acid residues other than the at least eight amino acid residues are fixed at two positions in each of the peptide segments to indicate an address or order of the peptide segments in the protein.27.The method of claim 24, wherein each of the plurality of peptide segments comprises a collagen-like motif or a coiled-coil motif.28.The method of claim 27, wherein each of the plurality of peptide segments comprises a peptide sequence of (Gly-Xaa-Yaa) n, wherein n is at least 4, Xaa and Yaa are independently or jointly selected from any one of the at least eight amino acid residues; or the plurality of peptide segments comprises a plurality of 7-amino acid peptides each comprising a peptide sequence of [ (R1) (Xaa) 2 (R1) (Xaa) 3] q, wherein R1 is independently I or L and q is 2 to 3.