OPTIMIZED PROTEIN THAT COMPRISES THE ESSENTIAL AMINO ACIDS FOR HUMAN NUTRITION.
Patent Information
- Application Number
- MX2021014268
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2041-11-19
AI Technical Summary
Current protein sources for human nutrition are diluted, lack essential amino acids, require excessive water for production, and have imbalanced amino acid compositions, leading to inefficiencies in muscle maintenance and increased environmental impact.
Development of an optimized protein comprising essential amino acids in balanced proportions using an algorithmic method, specifically designed to enhance muscle protein synthesis and reduce environmental footprint.
The optimized protein provides balanced essential amino acids for human nutrition, improving muscle maintenance and reducing resource consumption, while addressing the inefficiencies of traditional protein sources.
Abstract
Description
OPTIMIZED PROTEIN COMPRISING THE AMINO ACIDS ESSENTIAL FOR HUMAN NUTRITION FIELD OF INVENTION The present invention relates to the techniques and principles used in Biochemistry for the study and research of cells, as well as their chemical nature, in addition to the development of new chemical processes in biological systems of cells, tissues and organs, and more particularly, it relates to an optimized protein comprising the essential amino acids in the appropriate proportions for human nutrition. BACKGROUND OF THE INVENTION The decline in muscle function and muscle mass, also known as sarcopenia, is a natural phenomenon that occurs in humans starting around age 40. Every 10 years, a person loses an average of 8% of their muscle mass, so by age 65, they will have lost a fifth of their muscle mass, significantly impairing their ability to move independently. Sarcopenia not only limits independent movement but also promotes the development of metabolic syndrome, which in turn increases the risk of diabetes, cancer, and mortality, among other health problems. Sarcopenia can be slowed down by consuming large amounts of protein or essential amino acids. However, another natural phenomenon that occurs during aging is loss of appetite, which makes it difficult for these treatments to produce the desired results. There are several problems associated with natural sources of protein, namely: 1) are very diluted (for example, out of every 100 grams of beef, 30 grams are protein); 2) they do not possess the amounts of essential amino acids necessary for the human diet; 3) amino acids are not readily available, for example, less than 50% of the amino acids present in plant proteins are available; 4) They require large amounts of drinking water for their production, for example, to produce 1 kg of beef, 16,000 liters of water are required, 1 kg of chocolate requires 19,000 liters of water, among others. Amino acids are essential for life and their consumption is fundamental, not only as a source of energy, but also for the oxygenation of the body, the proper functioning of the immune system and the repair of tissues (National Research Council (2001) “Protein and Amino Acids”). Amino acids have been classified as essential and non-essential through animal growth monitoring and nitrogen balance assays. Essential amino acids are those whose carbon skeletons cannot be synthesized de novo by the body and therefore must be obtained from the diet (Wu et al., 2009; Wu et al., 2014). Of the twenty (20) essential amino acids (EAAs) required to build proteins, nine (9) are considered essential: histidine (His), isoleucine (Ie), leucine (Leu), lysine (Lys), methionine (Met), phenylalanine (Phe), threonine (Thr), tryptophan (Trp), and valine (Val). MA / a / ZUZl / U14Z0Ü On the other hand, non-essential amino acids (NEAAs) are also required for protein synthesis, but they are synthesized by tissues from other essential and non-essential amino acids. These include alanine (Ala), arginine (Arg), asparagine (Asn), aspartate (Asp), glutamate (Glu), glutamine (Gln), glycine (Gly), proline (Pro), and serine (Ser) for adult, non-carnivorous mammals (National Research Council (2001), “Protein and Amino Acids”). These NEAAs are classified as conditionally essential because their utilization rates exceed their synthesis rates under certain conditions, such as pregnancy, wounds, and infections. These are glutamate, glycine, glutamine, proline, and taurine in mammals (National Research Council (2001) “Protein and Amino Acids”). Even though elevated levels of glutamine, glutamate, and aspartate have been shown to have a neurotoxic effect (Wu et al., 2014).Cysteine and tyrosine, which are not synthesized by the body, are not considered essential because they can be formed from methionine and phenylalanine in the liver, respectively (Wu., 2013). If amino acids are not present in adequate proportions, protein synthesis is reduced and protein breakdown increases (Brunton et al., 1998; Zello et al., 1995; Duffy et al., 1981). Angiotensin-aldosterone systems (AAS) are primarily responsible for regulating muscle protein synthesis, with leucine playing the most important role in this process (Volpi et al., 2003; Garlick et al., 2005). Katsanos et al., 2006, reported that leucine has a unique role in stimulating protein synthesis in muscle in older adults. Protein intake is essential given its importance in transport, structural, regulatory, contractile, immunological, catalytic, and energetic functions. The nutritional value of proteins has been measured by parameters such as caloric content or weight, but this does not provide any information about the essential amino acid (EAA) content in the proteins consumed by humans (Tessari et al., 2016). The most important aspect of proteins, from a nutritional point of view, is their amino acid composition. The quality of a protein is primarily determined by its digestibility and bioavailability, as well as its essential amino acid (EAA) content (FAO / WHO, 1991). EAAs cannot be synthesized by the body and therefore must be obtained through diet to ensure the synthesis of proteins required by the human body. The World Health Organization (WHO) establishes the optimal proportions of EAAs required in the Recommended Daily Intake (RDA). The intake of high-quality protein containing essential amino acids (EAAs) that the body cannot synthesize is of paramount importance because these are key substrates for preserving or gaining muscle mass and for ensuring the synthesis of proteins required by the body (Millward et al., 2008). Identifying alternative protein sources that are affordable, have high nutritional value, and are sustainably produced has become a highly relevant issue due to the growing global demand for food (Hussein et al., 2017). New technologies are needed to find alternative protein supplements to replace traditional sources such as soybeans (Seo et al., 2008). Land availability for soybean cultivation is limited, and antinutritional factors such as trypsin inhibitors, lectins, and tannins, present in legumes like soybeans, have been reported to increase protein losses by inducing intestinal paralysis in animal models, which would result in less protein hydrolysis and thus reduced amino acid absorption (Salgado et al., 2002). In the 1990s, the concept of the ideal protein became popular among nutritionists at the University of Illinois (Stein et al., 1994). However, the idea of a perfect amino acid balance had been discussed since 1946 (Mitchell and Block, 1946), and to date, no protein containing essential amino acids (EAAs) in the proportion recommended for daily human intake has been reported. Current protein production systems rely on generating large quantities of protein, which requires extensive land use and results in high greenhouse gas emissions, rather than producing high-quality protein (Tessari et al., 2016). Tessari and colleagues re-evaluated the environmental footprint of protein production by considering a key factor in the quality of proteins consumed by humans: the essential amino acid (EAA) content. Land use was recalculated to produce 13 g of EAA or to meet the Recommended Dietary Allowance (RDA) for each EAA. This study concludes that producing high-quality plant protein in sufficient quantities to meet the RDA for EAA would require increased land use and higher greenhouse gas emissions, while land use for beef and soybeans would not change as much because they are closer to meeting the RDA requirements. FAO (Food and Agriculture Organization of the United Nations) estimates indicate that food production will need to increase by 70% by 2050, posing a significant challenge to the global capacity to provide sufficient food (Veldkamp et al., 2015). A 2018 study evaluated food production systems in terms of land use and greenhouse gas emissions if the Harvard Healthy Eating Program (HHEP) were adopted (Bahadur et al., 2018). The estimates show that current agricultural systems overproduce grains, fats, and sugars, but do not produce enough protein to meet the nutritional needs of the current population and the projected population growth of 7 to 9.8 billion people by 2050. Currently, there are no reports of any natural or synthetic protein with the recommended proportion of essential amino acids (EAAs). A protein with this composition would not only have a nutritional impact but could also result in more affordable costs and reduced land use compared to current production systems. In the prior art, several documents were found that are related to the subject matter of the present invention, such as International Patent Application No. PCT / US2013 / 071091 (International Publication No. WO 2014 / 081884), which proposes to genetically modify proteins to change their amino acid composition for health purposes; it also presents the expression of some selected proteins; in particular, in example 14, it refers to the production of the protein mainly secreted by Aspergillus niger (PDB: 3EQA, catalytic domain of glucoamylase) which was mutated in its loops with more than 4 amino acids in length to include essential amino acids (F, L, I, M, V, T, K, R, W); these mutations increased the content of essential amino acids by 41, 44 and 6% with respect to the original protein. An example is also given of a protein secreted by Bacillus that is enriched in essential amino acids (H, I, L, K and M).It acknowledges the low content of essential amino acids in natural sources and the various problems associated with the intake of natural proteins in the human diet. It suggests that it is possible to design polypeptides with the required essential amino acid composition, but recognizes the technical difficulty of achieving this.To reduce the risk of failure when building a protein with all the essential amino acids, they focus on designing proteins that supplement one or more amino acids that are necessary in the human diet to be fed using conservative substitutions (1: serine, threonine; 2: aspartic acid, glutamic acid; 3: asparagine, glutamine; 4: arginine, lysine; 5: isoleucine, leucine, methionine, alanine, valine; and, 6: phenylalanine, tyrosine, tryptophan); specifically they design proteins that have up to 5 times the proportion of essential amino acids (histidine, isoleucine, leucine, lysine, methionine, phenylalanine, threonine, tryptophan, and valine); or aromatic (phenylalanine, tryptophan, tyrosine, histidine and thyroxine); or branched (leucine, isoleucine and valine); or RQL (arginine, glutamine, and leucine) based on multiple alignments; substitutions were guided based on the frequency of occurrence of amino acids at each position.The selection of mutations was based on the estimation of six factors: amino acid probability (AALike), amino acid type probability (AATLike), position entropy (Spos), entropy per amino acid and position (SAATpos), relative free energy of folding (AAGfoId), and secondary structure identity (LoopID). They also describe polypeptides designed to lack certain amino acids (leucine, isoleucine, valine, arginine, histidine, lysine). These engineered polypeptides are 20 amino acids or longer and are designed to increase the amount of the selected amino acid(s). The engineered proteins resemble the initial protein by 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% in their amino acid sequences.These proteins are naturally secreted by the following microorganisms: Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Pichia pastoris, Corynebacterium, Synechocystis, and Synechococcus. However, the invention described in publication 081884 is based on a design that only allows for the addition or removal of essential amino acids, but does not permit modification to achieve a particular composition, as is permitted in the present invention. Furthermore, the invention in publication 081884 describes proteins selected from databases for its design, but none of these proteins are used in the present invention. On the other hand, International Patent Application No. PCT / US1998 / 006673 (International Publication No. WO 1998 / 045458) relates to the development of seeds and seed storage proteins that are improved in the quantity of amino acids essential for humans and animals. More specifically, this invention relates to the genetic engineering of the Brazil Nut 2S albumin seed storage protein to contain a higher percentage of essential amino acid residues. The expression of a gene encoding this modified seed storage protein in transgenic plants results in a greater accumulation of essential amino acids in the seeds of these plants. The production in plant seeds (soybeans) of proteins (albumin) with a high content of tryptophan, cysteine, and methionine is described.This is justified by the fact that these amino acids are essential for the human diet and plants have little of them. International patent application No. PCT / US1997 / 020441 (international publication No. WO 1998 / 020133) provides polypeptides comprising protease inhibitors with increased amounts of essential amino acids and nucleotides encoding these peptides. Transformed plants and seeds with enhanced nutritional value due to the expression of modified polypeptides are also provided. The production of proteins functioning as protease inhibitors with increased essential amino acid content (K, W, M, T) and their expression in plants to increase their nutritional value are described. The design was based on conservative substitutions. It is important to note that in the prior art several patent documents were found that refer to the use of essential amino acids, but none of them discuss or describe an invention that produces proteins with the balanced content of essential amino acids for the human diet, which is an important aspect of the invention described in the present patent application. On the other hand, genetic engineering has been used to modify plants (for example, corn) to increase the amount of the amino acid lysine they produce (https: / / www.ncbi.nlm.nih.aov / Dmc / articles / PMC2442549 / ). However, this approach does not solve the problems of the excessive water required for its production, provides only one of the 11 essential amino acids, and is produced in a plant whose proteins are highly diluted and poorly bioavailable to humans. Currently, the intake of free essential amino acids is used to address sarcopenia, particularly leucine, one of the twenty amino acids cells use to synthesize proteins (https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC3183816 / ). However, the long-term health consequences of free amino acid intake have not yet been evaluated. Animal studies show that free amino acid intake is not very efficient in animal nutrition (https: / / www.frontiersin.org / articles / 10.3389 / fped.2019.00563 / full) and may cause colitis and inflammation (https: / / Dubmed.ncbi.nlm.nih.aov / 29209321 / ). To elaborate further on the health risks to both humans and animals posed by the ingestion of free amino acids, a study conducted by the FDA in 1994 demonstrated that the lack of regulation for the sale of free amino acids led to the death of people due to the presence of contaminants in these preparations (https:www.ncbi.nlm.nih.gov / books / NBK209070 / ). Similarly, recent studies show that sustained intake of free amino acids over a long period negatively affects health and life expectancy in animals (https:www.sciencedirect.com / Science / Article / pii / S2468501119300082); (https: / / www.nature.com / articles / s42255-019-0059-2). In this regard, studies have already begun to evaluate whether the same occurs in humans. In particular, the invention described in international application PCT / US2013 / 071091, which describes the protein selected to improve the essential amino acid composition in A. Niger, glucoamylase (470 aa), which has the following proportion of essential amino acids: CS[C,M] Expected: 0.0105, CS[C,M] Observed:0.01950912344241975, Ratio:1.8580117564209284 CS[F,Y] Expected: 0.0175, CS[F,Y] Observed:0.12041872570381736, Ratio:6.881070040218134 ML / a / ZUZ 1 / U 14Z00 CS[H] Expected: 0.007, CS[H] Observed:0.010872084792543642, Ratio:1.5531549703633774 CS[I] Expected: 0.014, CS[I] Observed:0.04036903352408676, Ratio:2.8835023945776257 CS[K] Expected: 0.021, CS[K] Observed :0.020322422430736932, Ratio:0.9677344014636634 CS[L] Expected: 0.0273, CS[L] Observed :0.08298079113284501, Ratio:3.0395894187855315 CS[T] Expected: 0.0105, CS[T] Observed:0.07814890209095882, Ratio:7.442752580091315 CS[V] Expected: 0.0182, CS[V] 0bserved:0.060906466077756176, Ratio:3.3465091251514383 CS[W] Expected: 0.0028, CS[W] Observed :0.055358833891450694, Ratio:19.771012104089532. MA / a / ZUZl / U 14Z00 In other words, while some essential amino acids are present in the recommended ratio (K, H), others are present in excess (W, T, F / Y). Consequently, the invention focused solely on increasing the quantity of essential amino acids, rather than producing a protein with a balanced ratio for the human diet. Another issue is that the selected protein has enzymatic activity that may make it undesirable to ingest, so the authors propose some strategies to inactivate it. One problem that can arise from mutating amino acids important for enzyme activity is that the protein's stability depends on these amino acids. ML / a / ZUZ 1 / U1 As can be seen from the above discussion, the inventions described in applications Nos. PCT / US1998 / 006673 and PCT / US1997 / 020441 do not solve the problems of the excessive use of water required for their production and are produced in vegetables whose proteins are highly diluted and are poorly available to humans. BRIEF DESCRIPTION OF THE INVENTION The present invention relates to an optimized protein comprising the essential amino acids in the appropriate proportions for human nutrition, wherein said optimized protein was obtained using an algorithmic method, where from a selection in public databases of protein sequences, the sequence that most closely approximated containing the appropriate proportion of essential amino acids for human nutrition was selected, also considering prior evidence of its expression in heterologous systems and its nature to be secreted, the algorithmic method was developed that allowed making changes to said selected sequence so that it meets the greatest possible amount of essential amino acids in the appropriate proportion for human nutrition. To select the protein sequence that most closely approximated the appropriate proportion of essential amino acids, a search was conducted in public protein sequence databases. This search consisted of the following steps: 1) defining the desired proportion of essential amino acids per gram of total protein (PAAE), where this proportion is specified within a range of values (RVPAAE); 2) defining how many essential amino acids must meet the RVPAAE (EEA: Expected Essential Amino Acids), which can range from 1 to 9; 3) searching protein sequence databases for sequences that satisfy steps 1) and 2); 4) repeating steps 1), 2), and 3) for different possible EEA values; 5) selecting the proteins with the highest EEA values, preferably those proteins that satisfy a greater number of essential amino acids in the appropriate proportion for human nutrition. This procedure can be performed to supplement the amino acids present in a desired food (e.g., milk, egg, etc.) or to find a protein that by itself can satisfy the requirements of essential amino acids for human nutrition. Starting with the sequence selected in the previous procedure, changes were made to that sequence to make it contain the essential amino acids in the desired proportions, for which the optimization algorithm known as “simulated annealing” was used, where the steps of the optimization algorithm are as follows: (a) initialize the energy of the system; (b) calculate the value of the energy function of the sequence; (c) print the solution, as long as the value of the energy function of the sequence is equal to 0; otherwise continue with the following steps; (d) randomly select the positions in the sequence to be mutated; the number of positions is the method parameter specified in the next step e); (e) verify that the selected positions are among the positions susceptible to being changed specified by the corresponding method parameter in step b);(f) randomly select the essential amino acids to which the amino acids present in the selected positions will be replaced; (g) make the substitutions specified in step e) in the positions selected in step d); if the new sequence reduces the energy value, maintain it; otherwise, evaluate whether that sequence is maintained by applying the following formula:; g(previous energy - new energy) / temperature if the result is a number greater than a number chosen at random between 0 and 1, then the sequence is maintained for the next cycle; and, (h) repeat the above steps until the number of printed sequences is equal to the method parameter specified in step g), or a value in the system energy less than 1 has been reached. For obtaining the optimized protein of the present invention, from the universe of amino acid sequences found in the search, the following optimized amino acid sequences, established in the attached sequence list, were selected preferentially, but not limitingly, for said present invention: SEQ ID NO. 1; SEQ ID NO. 2; SEQ ID NO. 3; SEQ ID NO. 4; SEQ ID NO. 5; SEQ ID NO. 6; SEQ ID NO. 7; SEQ ID NO. 8; SEQ ID NO. 9; SEQ ID NO. 10; SEQ ID NO. 11; SEQ ID NO. 12; SEQ ID NO. 13; SEQ ID NO. 14; SEQ ID NO. 15; SEQ ID NO. 16; SEQ ID NO. 17; SEQ ID NO. 18; SEQ ID NO. 19; SEQ ID NO. 20, which comprise the proportions indicated for each essential amino acid (Observed and Required; Ratio indicates the relationship between Observed and Required); the identity in observed sequence with the template sequence is also indicated. In a further aspect of the present invention, from among said preferred optimized amino acid sequences, the amino acid sequences with which the experimental tests were carried out, SEQ ID NO. 12 and SEQ ID NO. 16, were most preferably selected. In another aspect of the present invention, the amino acid sequence named SEQ ID NO. 16 was selected even more preferentially to be expressed in the yeast Pichia pastoris under 2 different promoters: pGAP and pAOX1. The pGAPZ alpha plasmid was used to include the gene encoding for the SEQ ID NO protein. 16, using the codons of preferential use in the yeast Pichia pastorís. The corresponding nucleotide sequence for protein SEQ ID NO. 16 is as follows: 5'CAACCACCCAGGTCCACCAGGTCCACCAGGTCCACCAGTTTCTGCTATGTTGCCAGGT CCATTCGGTTTGCCAGGTTTCCCAGGTACTCCAGGTATGAAGGGTATTCAAGGTGAGA GAGGTTTGCCAGGTGAGAAGGGTGAGGTTGTTGCCAGGCCAGGCCAG GTGAGTCTAGATTGGGTCCACCAGGTTCTACTGGTTCTAGAGGTGTTCCAGGTCCACC AGGTAGACCAGGTGACTCTGGTATTAAG-3' Strains of the yeast Pichia pastoris X33 that overexpress the plasmid carrying the gene encoding the SEQ ID NO. 16 protein were isolated using different concentrations of the antibiotic zeocin (400 pg / mL, 800 pg / mL, and 1200 pg / mL). Two strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pAOX1 promoter, and three strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pGAP promoter. As already discussed in the background of the invention, the nutritional quality of a protein is mainly determined by its digestibility and essential amino acid (EAA) content, which cannot be synthesized by the body and, therefore, must be consumed in the diet. WlAia / ZVZl iv i ^zoo diet to ensure the synthesis of proteins required by the human body. The QMS establishes the optimal proportions of EAAs required in the human daily intake (RDA). OBJECTS OF THE INVENTION Taking into account the problems encountered in the prior art, it is an object of the present invention to provide an optimized protein that includes the essential amino acids in the proportions suitable for human nutrition. It is a further object of the present invention to provide the optimized protein that has the nutritional composition of essential amino acids (EAA) required in the human diet by bioinformatic analysis of existing databases and that has been previously reported by the WHO. Another object of the present invention is to provide the optimized protein that complements the EAA content of natural protein sources such as egg, milk, beef, chicken, pork, and fish. It is further an object of the present invention to generate variants of the optimized protein that include the EAAs required in human nutrition. It remains a further object of the present invention to clone the gene of the protein closest to the RDA into an expression vector. The foregoing and other objects, as well as the features and advantages of the present invention, will become more obvious when the embodiments of the present invention are described in greater detail and with reference to the accompanying drawings, which, besides forming part of the present invention, provide further understanding of said embodiments but do not constitute a limitation of the present invention. In the drawings, the same numerical references generally represent identical or similar parts or steps. BRIEF DESCRIPTION OF THE FIGURES The novel aspects that are considered characteristic of the present invention will be set forth in detail in the appended claims. However, the invention itself, both in terms of its organization and its method of operation, together with its other objects and advantages, will be better understood from the following detailed description of the embodiments of the present invention, when read in conjunction with the accompanying drawings, in which: Figure 1 is a graph showing the number of proteins in which the proportion of each of the 9 essential amino acids (EAA) is satisfied, where the 5 proteins analyzed were the following: synaphin, uncharacterized protein, vasotocin, type XII collagen and RECX_PSEMY regulatory protein. Figure 2 is a graph showing the proportions of each of the 9 EAAs in the 5 proteins found, where the proportion (1x-13x) of each EAA (T, M, F, K, H, V, I, L, W) in the uncharacterized protein of Macaca fascicularis (pink bars), the synaphin of Doryteuthis pealeii (blue bars), the collagen protein of Bos taurus (yellow bars), the regulatory protein of Pseudomonas mendocina (red bars) and the Vasotocin-neurophysin of Gallus gallus (green bars). Figure 3 is a graph showing the proportion of each amino acid (AA) in the Uniprot (SwissProt and TrEMBL) and PDB databases. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 4 shows the distribution of protein lengths annotated in SwissProt and PDB. The number of proteins in PDB (red dots) and SwissProt (blue dots) that fall within certain length ranges, increasing in increments of 50 amino acids, is shown. Figure 5 shows the distribution of protein lengths annotated in TrEMBL. The number of proteins (purple dots) in TrEMBL falls within certain length ranges, which increase in increments of 50 amino acids. Figure 6 shows the proportions of amino acids (AAs) in bacterial proteomes. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database multiplied by the average mass of the AA and the average mass of the annotated proteins. All 20 AAs were considered. Figure 7 shows the proportions of amino acids (AAs) in fungal proteomes. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 8 shows the proportions of amino acids (AAs) in plant proteomes. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 9 shows the proportions of amino acids (AAs) in animal proteomes. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 10 shows the proportions of amino acids (AAs) in viral proteomes. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 11 shows the proportions of amino acids (AAs) in the proteomes of animals, plants, fungi, viruses, and bacteria. The proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database * average mass of the AA / average mass of the annotated proteins. Figure 12 shows a photo of a Coomassie blue-stained acrylamide gel staining proteins obtained from the supernatant of Pichia pastoris cells expressing the SEQ ID NO. 16 sequence under the pGAP promoter. Figure 13 shows a photo of a Coomassie blue-stained acrylamide gel staining proteins obtained from the supernatant of Pichia pastoris cells expressing the SEQ ID NO. 16 sequence under the pAOX1 promoter. DETAILED DESCRIPTION OF THE MODALITIES OF THE INVENTION To this day, neither protein design nor engineering has been used to optimize the composition of essential amino acids for the human diet, but has simply been used to improve protein stability, change protein activity for use in industrial processes, and specifically recognize molecules, for example, in antibodies. The lack of nutritionally high-quality protein to feed the population, the environmental impact of current plant and animal protein production systems, and the low content of essential amino acids (EAAs) in plant proteins, which implies consuming a wide variety of plant protein sources to meet human requirements, were the reasons for which the present invention was carried out. In light of the foregoing, the inventors of the present invention carried out two primary activities, namely: i) a search for proteins of high nutritional value that do not depend on plant or animal systems; and ii) a proposal for a sustainable expression system for a protein for human consumption. Based on these two activities, in accordance with a particularly preferred embodiment of the present invention, an optimized protein comprising the essential amino acids in the proportions suitable for human nutrition was obtained, wherein said optimized protein was obtained using an algorithmic method developed by the inventors of the present invention. ML / a / ZUZ 1 / U1 Based on a selection from public databases of protein sequences, the sequence that most closely resembled containing the appropriate proportion of essential amino acids for human nutrition was selected. Considering also previous evidence of its expression in heterologous systems and its nature to be secreted, the algorithmic method was developed that allowed changes to be made to said selected sequence so that it fulfills the greatest possible amount of essential amino acids in the appropriate proportion for human nutrition. The optimization algorithm used is known as "simulated annealing." The idea behind this algorithm is to simulate what happens in metallurgical practice, where metals are first heated and then cooled, thereby removing impurities. The algorithm starts at a high temperature and ends when the function to be optimized has been satisfied, or when a minimum temperature has been reached. The steps of the optimization algorithm are as follows: a) start the system's energy; b) calculate the value of the energy function of the sequence; c) Print the solution, provided that the energy function value of the sequence is equal to 0; otherwise, continue with the following steps; d) randomly select the positions in the sequence to be mutated; the number of positions is the method parameter specified in the next step e); e) verify that the selected positions are among the positions that can be changed specified by the corresponding method parameter in step b); f) randomly select the essential amino acids that will replace the amino acids present in the selected positions; g) perform the substitutions specified in step e) in the positions selected in step d); if the new sequence reduces the energy value, maintain it; otherwise, evaluate whether that sequence is maintained by applying the following formula: (previous energy - new energy) / temperature if the result is a number greater than a number chosen at random between 0 and 1, then the sequence is maintained for the next cycle; and, h) repeat the previous steps until the number of printed sequences equals the method parameter specified in step g), or until the system energy value is less than 1; and i) Synthesize or produce the protein with the desired amino acid sequence or sequences. In accordance with the foregoing, to obtain the optimized protein of the particularly preferred modality of the present invention, a selection was made of sequences optimized in their essential amino acid composition. It is important to note that, using the algorithmic method described above, it was found that the solutions matched only within a range of values for the ratio of 1.9 to 3.1. As already mentioned, for obtaining the optimized protein of the present invention, from the universe of amino acid sequences found in the search, the following optimized amino acid sequences were preferably selected, but not limitingly, from the present invention, as established in the attached sequence list: SEQ ID NO. 1; SEQ ID NO. 2; SEQ ID NO. 3; SEQ ID NO. 4; SEQ ID NO. 5; SEQ ID NO. 6; SEQ ID NO. 7; SEQ ID NO. 8; SEQ ID NO. 9; SEQ ID NO. 10; SEQ ID NO. 11; SEQ ID NO. 12; SEQ ID NO. 13; SEQ ID NO. 14; SEQ ID NO. 15; SEQ ID NO. 16; SEQ ID NO. 17; SEQ ID NO. 18; SEQ ID NO. 19; SEQ ID NO. 20, which comprise the proportions indicated for each essential amino acid (Observed and Required; Ratio indicates the relationship between Observed and Required); the identity in observed sequence with the template sequence is also indicated. The letters in the sequences correspond to the 20 amino acids present in nature: A, Alanine; C, Cysteine; D, Aspartic acid; E, Glutamic acid; F, Phenylalanine; G, Glycine; H, Histidine; I, Isoleucine; K, Lysine; L, Leucine; M, Methionine; N, Asparagine; P, Proline; O, Glutamine; R, Arginine; S, Serine; T, Threonine; V, Valine; W, Tryptophan; Y, Tyrosine. Any of these preferred optimized amino acid sequences theoretically fulfills the condition of containing amino acids in the appropriate proportion for human nutrition. However, in a further aspect of the present invention, the amino acid sequences with which the experimental tests were carried out were selected from among these preferred optimized amino acid sequences, taking into account the following considerations: i) that the dispersion between the proportions obtained for each amino acid was not very large, therefore the most preferred sequences of the present invention are the amino acid sequences named SEQ ID NO. 12 and SEO ID NO. 16, which presented the least dispersion; where both sequences maintain the three-dimensional structure of the template protein according to the Zhang group's predictor called i-TASSER (https: / / zhanggroup.org / l-TASSER / ). This prediction anticipates that these variants will not have folding problems when expressed in a cell. In another aspect of the present invention, the amino acid sequence named SEQ ID NO. 16 was selected even more preferentially to be expressed in the yeast Pichia pastoris under 2 different promoters: pGAP and pAOX1; wherein pGAP is a system in which gene and protein expression is induced in the presence of glucose and pAOX1 in the presence of methanol. The methodological strategy followed to construct the DNA plasmids including the still more preferred amino acid sequence of the present invention SEQ ID NO. 16 under the pGAP and pAOX1 promoters is described below. Plasmids and strains used to validate optimized protein design: The pGAPZ alpha plasmid was used to include the gene encoding the SEQ ID NO. 16 protein, using the preferential codons in the yeast Pichia pastoris. The corresponding nucleotide sequence for the SEQ ID NO. 16 protein is as follows: 5'CAACCACCCAGGTCCACCAGGTCCACCAGGTCCACCAGTTTCTGCTATGTTGCCAGGT CCATTCGGTTTGCCAGGTTTCCCAGGTACTCCAGGTATGAAGGGTATTCAAGGTGAGA GAGGTTTGCCAGGTGAGAAGGGTGAGGTTGGTTTGCCAGGTCCACCAGGTCCACAAG GTGAGTCTAGATTGGGTCCACCAGGTTCTACTGGTTCTAGAGGTGTTCCAGGTCCACC AGGTAGACCAGGTGACTCTGGTATTAAG-3' The letters correspond to the 4 nucleotide bases present in DNA: A, Adenine; G, Guanine; T, Thymine; C, Cytosine. Strains of the yeast Pichia pastoris X33 that overexpress the plasmid carrying the gene encoding the SEQ ID NO. 16 protein were isolated using different concentrations of the antibiotic zeocin (400 pg / mL, 800 pg / mL, and 1200 pg / mL). Two strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pAOX1 promoter, and three strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pGAP promoter. As discussed in the background section of the invention, the nutritional quality of a protein is primarily determined by its digestibility and essential amino acid (EAA) content. EAAs cannot be synthesized by the body and therefore must be obtained through diet to ensure the synthesis of proteins required by the human body. The QMS establishes the optimal proportions of EAAs required in the Recommended Daily Intake (RDA). Currently, there are no reports of any natural or synthetic protein with the recommended proportion of essential amino acids (EAAs). A protein with this composition would not only have a nutritional impact but could also result in more affordable costs and reduced land use compared to current production systems. The present invention will be better understood from the following example, which is presented for illustrative purposes only, not as a limitation, to allow for a complete understanding of the embodiments of the present invention. This does not imply that other unillustrated embodiments do not exist and that they can be implemented based on the detailed description provided above. It is important to note that the data and experimental results obtained in the example described below are intended solely to provide the necessary elements for carrying out the invention and should not be considered as limiting its scope. EXAMPLE I. EXAMPLE / Searching for a protein in databases with the composition of EAAs required by humans. The Uniprot (96,757,994 sequences) and PDB (144,871 sequences) databases were consulted. From Uniprot, the 454,976 protein sequences comprising the 16,233 reviewed proteomes (SwissProt) were downloaded in FASTA format, as well as the 96,303,018 protein sequences comprising the 166,576 unreviewed proteomes (TrEMBL). The file containing the unreviewed protein sequences (96,303,018), which consisted of 54 GB, was fragmented into 223 files of 260 MB each. Each file was then analyzed using a Java-based code, which identified proteins that met the following criteria: • the proportion of EAAs closest to the requirement in the human daily intake of EAAs according to the QMS; • A protein length of up to 260 amino acids (AA), where this length was chosen considering that approximately 50% of the proteins listed in the databases fall within this length range, and that at lengths greater than 260 AA, proteins closer to the requirements were not found; instead, the number of substitutions required increased significantly, meaning that these proteins would have to be largely modified. Furthermore, the generation of protein mixtures showed that concatenating three sequences (each up to 260 AA) no longer yielded proteins close to the requirements because the EAA content was diluted at longer lengths. For these reasons, the search was limited to sequences of 1–260 AA. Proteins with the fewest AA substitutions were selected, taking into account a total protein change percentage of up to 20% (calculated as: the sum of the substitutions to be made / total length of AA of the protein). The Java code consisted of 3 methods, namely: 1. getSeqs: allowed reading the protein sequences in fasta format; 2. getMass: entered the average atomic mass of each AA according to those reported in: http: / / education.expasy.org / student projects / isotopident / htdocs / aa-list.html. 3. seqContainsEnoughAAE: This function allowed us to obtain the sequences that meet the daily requirement for each amino acid. For this purpose, the values previously established by the QMS for each amino acid were entered. For demonstration purposes, the calculations used in the code are shown for the case of the (K) proteins. However, each formula was applied with each of the 9 amino acids to find those proteins that satisfied the recommended daily intake. The first formula used to determine the Usinas content in the protein (given in atomic mass) was the following: Score K = (k*128.1741 Da) / SeqMass) where: k represents the number of Usinas found in the analyzed protein; 128.1741 Da is the average mass of the power plant; and, SeqMass is the sum of the average atomic masses of the amino acids that make up the protein. Subsequently, the required proportions of each essential amino acid (EAA) were obtained by dividing the RDA values of the QMS by 100 g of protein. Dividing by 100 g normalized the result to the grams of protein consumed. Taking the RDA value for lysine, which is 2100 mg, as an example: K ratio = 2.1 g / 100 g of protein Thus, the required proportion of lysine was 0.021. After obtaining the required proportions of the other EAAs, another formula was implemented in the code to determine the substitutions that had to be made in each protein to adjust its AA composition to the RDA: Ak = (EAA_K*SeqMass / 128.1741 Da)-k where: Ak represents the number of lysine substitutions; EAAK is the required proportion of lysine (0.021); SeqMass represents the average mass of the protein; and, 128.1741 Da is the average mass of the Power Plant The number of 39 Usinas that the protein initially contained was subtracted from the previous term, resulting in the number of K substitutions required to meet the RDA. Finally, the average mass of an AA for the protein in question was subtracted from each AA (this was obtained by dividing SeqMass by the number of AAs in that protein), thus giving the mass variation that should be used to adjust the mass of the new protein once the changes by Usinas were made: Ak2 = (128.1741 Da - massMProt)*Ak where: Ak2 represents the number of substitutions to be made in lysine taking into account the average mass of the protein; 128.1741 Da is the average mass of lysine; masaMProt is the average mass of the protein divided by the number of AA in that protein; and, Ak represents the number of lysine substitutions The code also received certain arguments, which could be modified at the time of execution: • The first of these was a ratio range that could go from 0.9-1.2 (1x), 1.9-3.1 (2-3x), or 2.9-4.1 (3-4x). A 1x ratio meant that the code would search for proteins with the required proportion of each EAA; while a 2-3x ratio meant searching for proteins with two or three times the required proportion of each EAA. Using a ratio range ensured that the proteins found did not deviate too much from their established RDA value; that is, using a minimum of 0.9 meant that each of the 9 EAAs could be 10% below its required 1x ratio, such that applying a maximum range of 1.2 meant that each EAA could be 20% above its required 1x ratio. • Secondly, there was the number of AEAs that were to be within the required proportion, which was abbreviated as the NAPR value (number of AEAs in the required proportion), where this value could range from 1 to 9, which corresponds to the number of AEAs that were to fall within the proportion ranges. For the execution of the code with the PDB and Uniprot databases, NAPR values ranging from 1 to 4 were used because using higher values did not produce positive results in the range of 0.9 to 1.2. The code was also executed using ranges of 1.9 to 3.1 and 2.9 to 4.1 with NAPR values of 8. Search for protein blends with the appropriate EAA composition. In addition to searching for a protein that met the EAA requirements, protein mixtures with the recommended EAA ratios were also sought. To do this, the previously developed code was used and modified to concatenate different protein sequences, generating 1:1 protein mixtures. The code was run against the Uniprot database, receiving as an argument whether the protein mixtures should consist of 2, 3, or more sequences. That is, if this parameter was set to 2, the code concatenated the sequences of protein 1 with the other sequences in the database, and then the sequence of protein 2 with all the other sequences, analyzing whether the generated mixtures met the EAA requirements. A minimum range of 0.9, a maximum range of 1.2, a NAPR value ranging from 4 to 9, and a maximum length of 260 AAs were used for each protein in the mixture. / Search for proteins that complement natural protein sources: egg, milk, meat, fish, pork and chicken. The Uniprot and PDB databases were consulted to find proteins that would complement the essential amino acid (EAA) content of natural protein sources such as egg, pork, milk, beef, fish, and chicken, using code developed in Java. First, the proportions of each EAA in these foods were calculated. The proportion of each EAA was obtained by dividing the grams of each EAA by 100 grams of the protein source. For example, 100 grams of beef contains 2,002 mg of lysine (WHO). Therefore, the proportion of potassium (K) in beef was 0.02002 (2.002 g / 100 g). The code searched for proteins that complemented the proportions of the nine essential amino acids (EAAs) for each of the five foods. Following the example of beef, it was 98 mg short of meeting its requirement. The code searched for proteins that contained these missing 98 mg and complemented its proportion (0.020022) to meet the required proportion of 0.02100. The code was run using a range of 0.9–1.2, a NAPR value of 1–8, and a maximum protein length of 260 and 900 amino acids. V Proportions in which each AA is found in the proteomes and in the Uniprot and PDB databases. After finding that none of the proteins in Uniprot and PDB had all 9 EAAs in the proportions recommended by the WHO, the proportions in which the AAs were found in the noted proteins were analyzed to see if these proportions were higher or lower than those required in the human daily intake. The proportions of the 20 AAs were calculated as follows: Proportion of each AA = number of times the AA was found in the proteins annotated in the database * average mass of the AA / total mass of the proteins annotated. The code also allowed the determination of the number of proteins falling within a certain amino acid (AA) length range. For this, an argument of 50 was used, which meant that the number of proteins within that length range, increasing by 50 AA increments, would be displayed. Using this code, the proportions of amino acids in the proteomes of the following organisms were analyzed: Acynonyx jubatus, Alligator mississippiensis, Bos taurus, dromedary camels, Gadus morhua, Galla gallas, Sas scrofa, Glycine max, Oryza sativa, Paenibacilla polymyxa, Rhodobacter spheroides, Schizophyllum commane, Taber melanosporum, Ustilago maydis, Lactobacillus casei, and Saccharomyces cerevisiae. WlAia / ZVZl IV l ^ZOO For the above, the proteomes were downloaded from the website of the National Center for Biotechnology Information (https: / / www.ncbi.nlm.nih.gov / aenome / ). V Synthesis of the gene that codes for the protein with the EAA composition required in the human daily intake. Bioinformatic analysis of the databases allowed us to identify the protein that most closely matched the EAA requirement. This was the type XII collagen protein from Bos taurus (Uniprot ID:P25508). The gene encoding this protein was optimized using codons from Pichia pastoris and was synthesized de novo by Gene Universal (229A Lake Dr. Newark, DE 19702) in the vectors pGAPZa A (constitutive) and pAOX1 (inducible). V Generation of variants of the Bos taurus collagen protein so that it has the EAA composition in the appropriate proportion. The algorithmic method described above allows for modifications to selected sequences to ensure they contain the highest possible number of essential amino acids in the correct proportions for human nutrition. This generates variants of the Bos taurus collagen protein that incorporate the necessary amino acid substitutions to achieve the recommended proportions of all nine essential amino acids. The algorithm accepts as arguments the protein positions where mutations are possible and the amino acids that can be substituted for those positions, excluding proline and glycine. These amino acids are not mutated to avoid folding problems, as they contribute to the helical structure of the collagen protein. The program also accepts as an argument the number of simultaneous substitutions to be performed to evolve the protein in each mutation cycle, allowing for the mutation of 1 to 20 amino acid residues at the same time.The algorithmic method yielded variant results by mutating the 20 AA and using a range of 1.9 - 3.1. This sequence was included in the pGAPZ alpha plasmid under the pGAP or pAOX1 promoter and was expressed in the Pichia pastoris strain X33. II. RESULTS The following tables show the proteins and protein complexes that most closely matched human amino acid requirements at 1x, 2x, 3x, and 4x ratios, as well as the proteins that best complemented the amino acid content of natural protein sources. The proportions of each amino acid found in the analyzed databases and proteomes are shown in graphs, and finally, a section is included with the experimental results of cloning the gene encoding the protein selected as closest to human requirements. / Proteins closest to the EAA requirement in SwissProt Running the Java code against the SwissProt database resulted in 4 proteins with a maximum of 260 amino acids that had up to 4 essential amino acids (EAAs) in the required 1x ratio of each EAA according to the QMS (range 0.9 - 1.2; NAPR = 3 and 4) (refer to Table 1). Four protein sequences were identified that best approximated the EAA requirement in the human diet when running the code with NAPR values of 4 (Macaca fascicularis, Pseudomonas mendocina, and Gallus gallus proteins) and 3 (Bos taurus protein). Table 1. Protein sequences identified in the database of SwissProt (454,976 sequences). Uniprot IP Length (AA) No. of Total Protein Change Organism protein substitutions (# of IL? of AA substitutions / length) Vasotocin neurophysin P24787 Gallus gallus 161 27 16.77% Regulatory protein Pseudomonas RECXPSEMY A4XWQ3 mendocins 150 25 16.66% Uncharacterized protein C16orf86 homolog (Fragment) Alpha-1 chain Q9GKT8 Macaca fascicularis 77 12 15.58% (XII) of collagen P25508 Bos taurus 86 7 8.13% As can be seen in Table 1, the type XII collagen protein from Bos taurus noted in the SwissProt database was the closest to the RDA, requiring 7 AA substitutions in an 86 AA protein, with a total change percentage of 8.13%, followed by the Macaca fascicularis protein with 15.58%, Pseudomonas mendocina with 16.66%, and finally the Gallus gallus protein with 16.77%. Proteins closest to the EAA requirement in PDB and TrEMBL The search for proteins in PDB yielded only one result, a protein with a maximum length of 260 amino acids and up to four essential amino acids (EAAs) in a 1x ratio (range 0.9–1.2) (see Table 2). In contrast, the unreviewed sequences from Uniprot (TrEMBL) yielded 911 proteins that most closely approximated the human daily requirement for EAAs. Running the code with NAPR values of 5–9 produced no results, which is interesting because this value indicates that the proteins in the analyzed databases (Uniprot and PDB) meet a maximum of four EAAs in the 1x ratio recommended for human daily intake. The other five EAAs fall outside this ratio, and therefore, certain substitutions must be made in the found proteins to meet the EAA content, since no protein with the required EAA ratio exists in nature. To see if the length range used of a maximum of 260 AA was interfering with the generation of results with NAPR values greater than 4, the code was run again using as an argument the search for proteins up to 1000 AA, obtaining that there were also no proteins in this length range that had NAPR values of 5-9. Table 2. Protein sequence identified in the PDB database (144,871 sequences). No. of Total Change Protein Organism Length AA substitutions of the protein AA Sinafina A Doryteuthis 79 11 13.92% pealeii The sequence in Table 2 was obtained using the code developed in Java using a NAPR value of 4 and a range of 0.9 - 1.2. On the other hand, as can be seen in Table 2, a sequence was identified in the PDB database that most closely approximated the EAA requirement in the human diet. This sequence identified in PDB corresponded to synaffin A from the squid Doryteuthis pealeii, with a total protein change percentage of 13.92%. / Selection of the protein with the EAA content in the ratio 1x closest (range=0.9-1.2) to that recommended for human consumption from the SwissProt, PDB and TrEMBL databases The proteins found in TrEMBL were discarded because most of these proteins are not characterized, their structure is unknown, they have not been expressed, and there is no experimental evidence that they are truly proteins or are ORFs. The sequences from the PDB and SwissProt databases that best approximated the RDA with a maximum total protein change percentage of 20% were evaluated, because this implied that fewer substitutions would have to be made to the protein and that it would be as close as possible to the RDA of EAA. Five proteins were found in which the number of amino acid substitutions required was minimal relative to the total protein length. The identified proteins correspond to the following organisms: Macaca fascicularis (12 substitutions, length = 77 aa), Doryteuthis pealeii (11 substitutions, length = 79 aa), Bos taurus (7 substitutions, length = 86 aa), Pseudomonas mendocina (25 substitutions, length = 150 aa), and Gallus gallus (27 substitutions, length = 161 aa). Figure 1 of the accompanying drawings illustrates a graph showing the number of proteins in which the proportion of each of the 9 EAAs is satisfied, where the 5 proteins analyzed were the following: synaphin, uncharacterized protein, vasotocin, type XII collagen and RECX_PSEMY regulatory protein. ML / a / ZUZ 1 / Ul 4Z00 The EAAT, M, I and F were found in the required 1x ratio in the uncharacterized protein of Macaca fascicularis, T, V, L and F in the synaffin of Doryteuthis pealeii, I, L and F in the collagen protein of Bos taurus, K, H, T and V in the regulatory protein of Pseudomonas mendocina, and K, H, V and I in the Vasotocin-neurophysin of Gallus gallus. Figure 2 of the accompanying drawings shows the proportions of each of the nine essential amino acids (EAAs) in the five proteins found. As can be seen in Figure 3, three of the five proteins met the amino acid requirements: threonine, phenylalanine, valine, and isoleucine. The collagen protein from Bos taurus was the only one that met the Recommended Dietary Allowance (RDA) for leucine, while the uncharacterized protein from Macaca fascicularis met the RDA for methionine. Vasotocin and the regulatory protein from Pseudomonas mendocina met the RDA for lysine and histidine, but none of the five proteins contained tryptophan in the proportion recommended by the WHO. All of them lacked whey, with the exception of the regulatory protein from Pseudomonas mendocina, which had seven times the RDA for whey. V Protein blends that satisfy the composition of 4 EAAs in a 1x ratio (range 0.9-1.2) Since no protein was found that met the RDA for all 9 EAAs, we proceeded to analyze whether there were protein blends that contained each of the 9 EAAs in the required 1x ratio. The blends contained up to 4 EAAs in the recommended ratio (see Table 3). ML / a / ZUZ 1 / U14Z00 wiA / aizvzi / ui has a length of 86 AA and a molecular weight of 8 kDa. The Uniprot database shows that there is experimental evidence at the protein level and it is known to have a stable structure. Furthermore, human collagen fragments of types 1, 2 and 3 ranging from 8.7 to 43 kDa have already been previously expressed in Pichia pastoris at high levels and secreted into the medium as single chains by the signal sequence of the yeast mating factor a (Nokelainen et al., 2001; Williams et al., 2008; He et al., 2015; Wang et al., 2014; Bin et al., 2011; Pakkanen et al., 2006) with yields of up to 14.8 g / l (Werten et al., 1999). The collagen protein would require 7 AA substitutions to have all the EAAs in the correct proportion, so we proceeded to generate in silico variants of the sequence of this protein. / Generation of variants of the collagen protein of Bos taurus The Java code developed allowed the identification of collagen protein variants by mutating amino acid residues in groups of 20, with the exception of proline and glycine. Interestingly, no variant was found with a 1x ratio of the desired proportion; instead, only 20 variants were found with 2 to 3 times the proportion required in the human daily intake, 9 of them with a 13.9% difference from the original collagen protein (see Table 4). Two of these sequences could be selected to evaluate the expression of the modified versions in Pichia pastoris. MA / a / 2U21 / U 14200 Table 4. Variants of the Bos taurus type XII collagen protein sequence. Of the 9 sequences found, the AA residues where the code introduced substitutions with respect to the WT sequence are shown in red, and the protein positions that the code could not mutate are marked in blue. WT sequence of type XII collagen protein from Bos taurus NQPGPPGPPGPPGSAGEPGPGGRPGFPGTPGMQGPQGERGLPGEXGERGLPG PPGPQGESRTGPPGSTGSRGPPGPPGRPGDSGIR Variants of collagen protein sequences of Bos taurus NKPGPPGPPGPPGSAFEPGPGGRPGFPGVPGMQGPQGERKLPGLMGERGLPGPPGPQ GESRTIPPGSTGHRVPPGPPGRPGKLVIR NQPGPPGPPGPPGSKGEPGPGGVPGFPGTPGMQMPQGERGLPGEXVKLHLPGPPGPQ GESRFGPPGSTGSRIPPGPPGRPGDKVIL NQPGPPPGPPGPPGKAGVPGPGVRPGFPGKPGMQMIQGERGLPGLKGLVGLPGPPGPQ GESRTGHPGSTGSRGPPGPPGRPGDFGIR NKVGPPGPPGPPGSFGEPGPGGRPGFPGTPGMQGPQGVRGLPGMXGERLLPGPPGPQ GESRTIKPGSKHSRGPPGPGRPGVSLIR MFPGPPGPPGPPGSKGHPGPGLRPGFPGTPGMQGVQKEVGLPGEXGVRGLPGPPGPQ GESRTGPPGSKGSRGIPGPPGRPGDLGIR HQPGPPGPPGPPGLAGKPGPVGKPGFPGTPGMQGPQVERMLPGEVGLRGLPGPPGPQ GESKFGPPGSTISRGPPGPGRPGDSGIR NQPGPPGPPGPPGSLVLPGPGGRPGFPGTPGMQGPQKEVGLPGEKGEKVLPGPPGPQ GESRIGMPGSTGSHGFPGPPGRPGDSGIR NQPGPPPGPPGPPGKAGEPGPGGRPGFPGTPGMQGPKFERGLPGEXKERVLPGPPGPQ GESVTGIPGSLGSRLPPGPPGRPGHVGIM NQPLPPGPPGPPGSKVEPPGGGRPGFPGVPGMQGPKGERGLPGHXVELGLPGPPGPQ GESRTGFPGSTKSRIPPGPPGRPGDMGIR No tryptophan (W) was introduced into any of the collagen protein variants because the required proportion of W was the lowest (0.0028) since the collagen protein length was only 86 AA. This meant that the required proportion of W was less than 1 tryptophan in the protein, making it impossible to find variants with 1x the required proportion of this EAA. - / Proteins that have 2 to 3 times the desired composition of EAAs (range of 1.9 - 3.1) in Uniprot and PDB When generating collagen variants, it was observed that it was not possible to generate a sequence that had exactly 1x the required composition of each EAA. Therefore, the databases were searched again to see if there were proteins that had 2 to 3 times the required proportion of each EAA, and not just once (1x) as is the case with the Bos taurus collagen protein (range of 0.9–1.2). It was found that there are no sequences that satisfy the 9 EAA proportions in the range of 1.9 to 3.1; however, there are 2 proteins in the SwissProt database that satisfy the requirement of up to 8 EAAs (excluding tryptophan) in a 2–3x proportion: • RS15 NATPD 30S ribosomal protein from Natronomonas pharaonis (Uniprot ID: Q3IQA3). W was found in a ratio of 3.7x: • Pre-mRNA CWC21 splicing factor from Ashbya gossypii (Uniprot ID:Q751G9). W was present in a proportion of Ox. ML / a / ZUZ 1 / U 14Z00 In TrEMBL, 141 sequences were found that had 8 EAAs in a range of 1.9 MA / a / ZUZl / U I4Z00 3.1. All these proteins had an EAA outside the desired proportion, some of them are like those shown in Table 5 below: Table 5. TrEMBL annotated proteins that satisfy 2 to 3 times the human required intake of EAAs. A NAPR value of 8, a range of 1.9-3.1, and a maximum length of 260 AA were used. ID Uniprot Protein Organism Protein length (AA) EAA out of proportion (1.9-3.1) A0A091CL33 Non-Fukomys protein 230 W=7.6x A0A1U8DU65 characterized Damarensis-associated protein Alligator 245 M=3.6x A0A1A8NEZ1 ribonucleoprotein SH3 domain binding to Nothobranchius sinensis protein 207 F=3.7x A0A0S7ENX5 rich in glutamic acid RNF37 pienaari Poeciliopsis 210 W=11.4x M4AJT2 (Fragment) Prolific transport Xiphophorus 210 W=8.2x K7G7M7 Intrafagellar 43 Syntaxin 7 maculatus Pelodiscus 258 l=7x A0A2K5NP26 Protein no Cercocebus sinensis 223 F=10.3x A0A2K5ZMJ3 J3MH86 characterized Uncharacterized protein Non-atys protein Mandrillus leucophaeus Oryza 223 212 F=10.4x H=0x A0A2H5N6J0 characterized Non-brachyantha protein Citrus unshiu 72 W=16.4x A0A287MZH9 characterized Uncharacterized protein Hordeum vulgare 240 T=6x V Proteins that have 3 to 4 times the desired composition of EAAs (range of 2.9-4.1) In SwissProt, 8 proteins were found to be missing only one EAA in the 3-4x ratio, as can be seen in Table 6 below: Table 6. Proteins annotated in SwissProt that satisfy a 3-4x ratio. An NAPR value of 8 and a range of 2.9 - 4.1 were used for Java code execution. ID Uniprot Protein Organism Protein Length (AA) EAA out of ratio (2.9-4.1) Q9VR89 RNA-binding protein pno1 Protein Drosophila melanogaster Lactobacillus reuteri 240 W=2.4x A5VIU4 ribosomal 30S S4 201 F=5.4x Q0C187 Methyltransferase E Hyphomonas neptunium 234 W=7.7x Q7VS88 Deformylase 2 Bordetella pertussis 170 V=5x Q87TT1 delta subunit of ATP synthase Pseudomonas syringae pv. 178 L=4.2x tomato Actin-related protein Q9D898 Mus musculus 153 M=2.7x Q089E2 LexA repressor Shewanella frigidimarine 204 1=5.3x P30334 Ribosomal hibernation factor Bradyrhizobium diazoeficiens 203 F=4.5x V Proteins that complement natural protein sources: egg, milk, beef, pork, fish and chicken. Proteins were sought that could serve as a dietary supplement for athletes, older adults, and people with innate protein metabolism disorders. To this end, proteins were sought that complemented the essential amino acid (EAA) content of natural sources in a 1:1 ratio. However, none of the proteins found were a good candidate because they only had 3 of the 9 EAAs in the recommended ratio, as shown in Table 7 below: Table 7. Proteins annotated in PDB that complement egg, milk and fish in a 1x ratio. NAPR values (2-3) and a range of 0.9 - 1.2 for Java code execution. Protein Length (AA) ID PBD Protein source that complements EAA within the required 1x ratio Vasopressin receptor V1a 84 1 ytv_M Egg 2 Ribosomal protein L37 85 5it7J¡ Egg 2 Pollen allergen Art v 1 108 2kpy_A Fish 2 Phospholipase A2 119 1 mh2_B Milk 3 Filamin-A 95 2mtp A Milk 3 Interleukin 11 177 4mhl_A Milk 3 Anticoagulant protein C2 85 1cou_A Milk 2 Xylanase D 87 1e5b_A Milk 2 The proteins found in PDB (Table 7) complemented egg and fish with up to 2 EAAs in the 1x ratio required in the human daily intake; whereas the proteins that complement milk do so with up to 3 EAAs in this required ratio. On the other hand, in PDB proteins were found that met the requirements of a greater number of EAAs as the required proportion increased (2-4x), as shown in Tables 8 and 9 below: Table 8. Proteins annotated in PDB that complement egg, milk, chicken, fish, beef, and pork in a 2-3x ratio. NAPR values (2-6) and a range of 1.9-3.1 for Java code execution. EAA Source within Protein Length (AA) ID Protein PDB to which the ratio 2- complements 2x required Ribosomal protein 94 4v8p_AA Beef 3 Helical repeat protein 239 5cwm_A Beef 3 Ethanol region transcription factor 65 1f4s_P Chicken 2 Vasopressin receptor pathway 84 1 ytv_M Chicken 2 Muscarinic toxin-like protein homolog 65 3hh7_A Chicken 2 Neurotrophin-4 130 1b8m_B Egg 4 Pleiotropin 136 2n6f_A Egg 4 Endoxylanase 151 1°8p_A Egg 4 L37 60S ribosomal protein 97 AugO_Lj Fish 3 Prokaryotic protein similar to Ubiquitin 68 3m9d_G Milk 6 MAB_3112 protein similar to ESAT-6 94 4¡0x_A Milk 6 Rab-3A interaction protein 78 4lhx_C Milk 6 Streptavidin 159 5f2b_A Milk 6 Uncharacterized protein 139 4lmi_A Pig 3 Table 9. Proteins annotated in PDB that complement beef, chicken, egg, fish, milk, and pork in a 3-4 x ratio. NAPR values (2-3) and a range of 2.94.1 for Java code execution. ML / a / ZUZ 1 / U 14Z00 EAA Source within Protein Length (AA) ID Protein PDB to which the 3-complements 4x ratio required Phospholipase A2 122 1faz_A Beef 3 Matrix metalloproteinase 2 72 1j7m_A Beef 3 Small nuclear ribonucleoprotein C 61 4pjo_L Beef 3 Clotting factor-binding protein B chain 123 1j34_B Chicken 2 Ribosomal protein L24E 66 11m1k_V Chicken 2 Tachylectin 2 236 1t12_A Egg 5 Restriction endonuclease K 166 3z¡5_A Egg 5 Eukaryotic translation initiation factor 3 128 5h7u_A Egg 5 Endoglucanase 181 1wc2_A Fish 3 Small nuclear ribonucleoprotein C 61 4pjo_l Fish 3 5(3) deoxyribonyletidase 197 1q91_A Milk 7 prgH protein 227 3gr1_G Milk 7 Na(+) H(+) exchange regulatory cofactor NHE-RF1 210 2krg_A Milk 7 Hsp-associated protein 90 110 21su_A Milk 7 Isolectin I agglutinin 89 1en2_A Pig 3 Alpha-1 (XX) collagen chain 104 2dkm_A Pig 3 eL29 245 51zw_b Pig 3 After analyzing proteins that complemented natural sources in 1x, 2-3x, and 3-4x ratios, milk was found to be the best complement, containing up to 7 essential amino acids (EAAs) in the 3-4x ratio. Complementary proteins contained the EAAs phenylalanine and tryptophan outside this ratio, with up to 8 or 9 times the required proportion of phenylalanine. The NHE-RF1 cofactor contained the EAAs lysine (4.4x) and methionine (4.3x) outside the 3-4x ratio. No proteins were found in the TrEMBL database that complemented beef, pork, fish, or chicken with NAPR values of 2-9. In SwissProt, it was observed that the proteins that complemented beef satisfied the requirement of 3 EAAs in a 1x ratio. In the case of milk, no protein complemented its EAA content; for pork and chicken, only 1 EAA was satisfied, while for egg and fish, 2 EAAs were satisfied, as can be seen in Table 10 below: Table 10. Proteins noted in SwissProt that complement beef, chicken, egg, and fish in a 1x ratio. NAPR values (1-3) and a range of 0.9-1.2 for Java code execution. EAA within Protein Organism Length (AA) ID Protein source of the required 1x ratio Abundant protein Arabidopsis Beef 29 of late embryogenesis Bacillus thaliana 225 195 Q9LW12 P39801 3 Beef 2 spore subtilis Homo sapiens tumor suppressor protein 212 Q2TAM9 Beef 2 tumor suppressor protein Bos taurus 242 P06836 Chicken 1 Protein 12-1 Mus musculus 130 Q9Z287 Chicken keratin-associated protein 1 Phospholemman Oryctolagus cuniculus 92 G1TZA0 Egg 2 Nuclear transition protein 2 Sus scrofa 137 P29258 Egg 2 Myelin-associated neurite growth inhibitor Bos taurus 195 1e5b_A Egg 2 Uncharacterized protein R346 Acanthamoeba polyphaga mimivirus 195 Q5UQT4 Fish 2 Since no proteins with more than 3 EAAs in a 1x ratio were found, proteins were sought in SwissProt that complemented natural sources with a ratio of 3 times greater than the recommended amount. Table 11. Proteins annotated in SwissProt that complement natural protein sources in a 2-3x ratio. NAPR values (3-5) and a range of 1.9-3.1 for Java code execution. Protein Organism Length ID Source EAA within the protein ratio 2-3x required Abundant protein 7 from Arabidopsis Q9627 Meat from late embryogenesis thaliana 10 Interaction protein Mus Q9D01 Meat from beef with PLK1 specific to musculus 178 1 M phase 3 Oryctolagus related protein 199 G1TZA Meat from MARCKS cuniculus 0 0 Nuclear transition protein 2 Sus scrofa 137 P29258 Chicken 3 Oryza sativa A3CG8 Cell wall glycine-rich structural protein subspecies loo 3 japonica 15 Glycine-rich protein from the secretory gland of ribaliUS 246 P08568 Egg 5 submandibular norvegicus Protein L37 60S ribosomal Bos taurus 97 P79244 Fish 3 Protein L37 60S Homo 97 P61928 Fish 3 ribosomal sapiens Mus Q9D01 interaction protein with PLK1 specific to musculus 178 1 Fish 3 M phase Secretin Homo 121 P09683 Pig 4 20 sapiens As shown in Table 11 above, SwissProt revealed proteins that complemented the essential amino acid (EAA) content of eggs, providing up to 5 EAAs at a 2-3x ratio. One of these, the glycine-rich structural protein of the Oryza sativa cell wall, could be a candidate protein for use as a dietary supplement. For pork, secretin from Homo sapiens was found to contain up to 4 EAAs at a 2-3x ratio. Beef, chicken, and fish showed a NAPR value of 3, while no complementary proteins were found for milk. Some of the proteins that complemented foods at a 2-3x ratio had also been observed to complement them at a 1x ratio.One example is the nuclear transition protein 2 of Sus scrofa, which was found to complement egg with a NAPR value of 2 in a 1x ratio, and chicken with a NAPR value of 3 in a 2-3x ratio. The proteins found in both the 1x and 2-3x ratios correspond to the same organisms: Bos taurus, Flattus norvegicus, Sus scrofa, Homo sapiens, Mus musculus, Oryctolagus cuniculus, and Arabidopsis thaliana. Furthermore, two abundant late embryogenesis proteins of Arabidopsis thaliana, 29 and 7, complemented the same food, beef, in 1x and 2-3x ratios, respectively. Proteins from Bos taurus, Homo sapiens, and Streptococcus pneumoniae complemented the egg with up to 5 essential amino acids (EAAs) at a ratio of 3–4x (see Table 12). The EAA content of pork was complemented by a protein with up to 4 EAAs at the recommended ratio. In the case of fish, proteins from various organisms were found, ranging from archaea, viruses, and bacteria to proteins from Bos taurus, Gallus gallus, Danio rerio, Oryctolagus cuniculus, Flattus norvegicus, and Xenopus tropicalis, with a value NAPR of 3. Table 12. Proteins annotated in SwissProt that complement natural protein sources at a 3-4x ratio. NAPR values (3-5) and a range of 2.9 - 4.1 for Java code execution. 5 Length (AA) 169 ID Q5FLW8 Protein source Beef EAA within the required 2-3x ratio 3 Protein Organism Lactobacillus acidophilus Ribosomal protein L35 50S Hsp9 Heat Shock protein S. pombe-associated protein 15-1 Mus musculus 68 150 P50519 Q9QZU5 Beef Beef 3 3 Keratin Chromosome 6 protein Emencella niduians 106 Q5BP05 Chicken 3 Uncharacterized protein S cerevisiae 128 P38216 Chicken Q YBR16W O Serine / arginine-rich splicing factor 7 Homo sapiens 238 Q16629 Egg 5 Delta subunit RNA Streptococcus 200 P66718 Egg 5 polymerase pneumoniae 10 Nuclear protein 2 Bos taurus 98 Q32PB4 Egg 5 Tegument protein Epstein Barr 217 P0C724 Fish 3 BKRF4 virus PPR111 protein ligase Xenopus 119 Q6GLB0 Fish 3 ubiquitin E3 tropicaiis Regulatory subunit 1A Oryctolagus 166 P01099 Fish 3 phosphatase 1 cuniculus Proline-rich cell wall protein 2 Giycine max Staphylococcus 230 P13993 Fish 3 T-transglycosylase SceD 1 s 238 Q4A0X5Fish 3 saprophyticus Small nuclear ribonucleoprotein C Danio rerio 159 Q8JGS0 Fish 3 High mobility cluster B3 protein Gallus gpllMS 202 P40618 Fish 3 Bos taurus complex E subunit 244 Q29RS4 Fish 3 1 c INO80 15 Glycine-rich secretory gland CA protein Rattus 246 P08568 Fish 3 submandibular norvegicus ATP synthase subunit B Stapbyiotermu 195 A3DNQ9 Fish 3 marine Rattus modulator 3 158 Q6MG88 Pig A protein signaling G norvegicus 4 / Proportions in which each AA is found in the proteins of the PDB and Uniprot databases (SwissProt and TrEMBL) The proportions in which each AA was found in the proteins annotated in PDB, SwissProt and TrEMBL were sought to see if this proportion was higher or lower than that required in the daily human intake, and thus, in this way, be able to analyze why the proteins found did not satisfy the requirements of the 9 EAAs. The proportions of each of the 20 amino acids (AAs) were very similar in the three databases analyzed (see Figure 3 in the accompanying diagrams). The AAs found in the lowest proportions were tryptophan, cysteine, methionine, and histidine. Leucine, glutamic acid, and arginine were found in the highest proportions. To obtain the graph shown in Figure 3, all 20 AAs (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y) were considered. The proportions recommended by the QMS for each of the nine essential amino acid supplements (EAAs) are shown in black dots, along with PDB (red dots), TrEMBL (purple dots), and SwissProt (blue dots). The proportions of each amino acid (AA) found in SwissProt, PDB, and TrEMBL, as well as the proportions recommended by the QMS for each essential amino acid (EAA), are shown in Table 13. The proportion of each AA was equal to the number of times the AA was found in the annotated proteins multiplied by the average mass of the AA divided by the average mass of the annotated proteins. The proportions are normalized to 1 g of protein. The EAAs found in the lowest proportions are shown in red, and the other EAAs (F, T, I, V, K, and L) are shown in blue. ML / a / ZUZ 1 4Z00 Table 13. Proportions of each AA in 1 gram of protein. ML / a / ZUZ 1 / U 14Z00 Promedio de la AA WHO SwissProt PDB TrEMBL proportion of each AA in the 3 data bases C 0.0136 0.0219 0.0111 0.0155 W 0.0028 0.0181 0.0215 0.0219 0.0205 M 0.0105 0.0278 0.0268 0.0281 0.0275 H 0.0070 0.0282 0.0319 0.0272 0.0291 G 0.0357 0.0440 0.0377 0.0391 N 0.0415 0.0418 0.0399 0.0410 P 0.0430 0.0396 0.0427 0.0417 Q 0.0460 0.0424 0.0438 0.0440 Y 0.0422 0.0487 0.0430 0.0446 T 0.0.105 0.0485 0.0500 0.0508 0.0497 F 0.0175 0.0506 0.0499 0.0523 0.0509 S 0.0545 0.0479 0.0525 0.0516 A 0.0515 0.0564 0.0588 0.0555 D 0.0563 0.0562 0.0570 0.0565 1 0.0140 0.0580 0.554 0.0583 0.0572 V 0.0182 0.0604 0.0611 0.0619 0.0611 K 0.0210 0.0675 0.0673 0.0575 0.0641 E 0.0793 0.747 0.0721 0.0753 R 0.0783 0.0724 0.0812 0.0773 L 0.0273 0.980 0.0891 0.1013 0.0961 The mean amino acid (AA) ratio was 0.047, indicating a uniform distribution of AAs. While the ratios of tryptophan (0.0205), methionine (0.0275), and histidine (0.0291) were among the lowest in the databases, they were still higher than the recommended daily intake for humans (0.0028 for tryptophan, 0.0105 for methionine, and 0.0070 for histidine). Therefore, the requirement for a greater number of AAs at higher ratios (2-4x) was met. In addition to AA ratios, the code also yielded the number of proteins within a given length range, increasing in increments of 50 AAs, as shown in Figure 4 of the accompanying illustrations. As can be seen in Figure 4, the PDB contains mostly annotated sequences of 1–50 AA (50,830), 51–100 AA (45,928), 101–150 AA (71,031), 151–200 AA (54,033), and 201–250 AA (61,246). The number of proteins found decreased with increasing AA length ranges, with proteins of 4851–4900, 4301–4350, 3951–4000, 3801–3850, and 2951–3000 AA being less common, each with only one protein found within those ranges. PDB contained 50,830 peptides of 1–50 amino acids, while SwissProt contained only 2,988 in this range. Similar to the PDB database, SwissProt showed a clear downward trend in the number of proteins with longer amino acid lengths, as seen in the 4901–4950 (5), 4601–4650 (6), and 4851–4900 (6) ranges. In the TrEMBL database (see Figure 5 of the accompanying diagrams), it can be seen that there are more proteins in the 101–150 AA range (14,327,667), followed by proteins in the 151–200 AA range (14,108,345) and the 201–250 AA range (13,734,838). There is a greater number of proteins in the 951–1000 AA range (5,384,100), and in smaller proportions, proteins are found within the 5051–5100 AA range (911), the 5001–5050 AA range (1,262), the 4851–4900 AA range (1,281), and the 4901–4950 AA range (1,322). Of the three amino acid length distributions analyzed (PDB, SwissProt, and TrEMBL), sequences of 1–250 amino acids are abundant. Restricting the search to proteins up to 260 amino acids resulted in even lower required proportions of essential amino acids (EAAs) such as whey, methionine, and hydrogen. For example, tryptophan is required at a ratio of 0.0028. In proteins up to 260 amino acids, this ratio is not met because the requirement is less than 1 tryptophan per protein, while the proportion of whey in the proteins is higher than required. Tryptophan is the amino acid required in the smallest proportion (0.0028) but also has the largest mass (186.2132 Da). To determine the possible minimum length the protein should have to satisfy the requirement for each amino acid, the following probabilistic calculation was used, taking into account the average mass of the 20 amino acids: P / n*118.7360=V where: P is the average atomic mass of the EAA: n is the minimum protein length to satisfy the required ratio of EAA; V is the required proportion of EAA as reported by the WHO; and, 118.7360 is the average mass of the 20 AA. Applying the formula above to tryptophan (n=186.2132 / (118.7360*0.0028)), the minimum possible length to satisfy its required proportion (0.0028) was found to be 560 amino acids. Therefore, it was unlikely that the nine essential amino acids (EAAs) would be present in the recommended proportions using a sequence search range of 1–260 amino acids. For the other EAAs, the minimum protein lengths to satisfy their requirements are shown in Table 14 below: ML / a / ZUZl / Ul 4Z00 Table 14. Minimum protein length to meet the requirements of the 9 EAA (W,K,H,T,M, V,I,L,F). The possible minimum protein length to satisfy the requirement of each EAA was obtained through probabilistic calculation: P / n*118.7360=V. Minimum length of EAA protein to meet your requirement W560 K51 H165 T81 M105 V46 I68 L40 F71 As can be seen in Table 14, using the length range of 1–260 amino acids (AA), the requirement of 8 essential amino acids (EAAs) was met, with the exception of tryptophan. In proteins up to 1000 AA, the requirement of more than 4 EAAs in a 1x ratio was also not met. One possible explanation for this is that the probability of finding proteins with all EAAs in the ratio recommended by the QMS is very low. The length distributions showed that there were approximately 3 x 10⁻⁵ sequences in SwissProt and 4.5 x 10⁻⁵ sequences in PDB of 1–1000 AA. For a 560 AA sequence to have a tryptophan, all possible 560 AA sequences that did not have W (19L560 sequences) would have to be found and each one added a W. The probability of finding such sequences was 1 / 19L559 and in a sample of 3-4.5χ10L5 sequences this was very unlikely. / Proportions of each AA at the proteome level After reviewing the proportions of amino acids (AAs) in the databases, these proportions were analyzed at the proteomic level. Proteomes from fungi, plants, animals, bacteria, and viruses were selected to determine if the human diet might be inducing a trend in the composition of essential amino acids (EAAs) in these organisms, with proportions closer to those required in the proteomes of organisms that form the basis of our food production, such as cattle, pigs, and chickens. The proteomes analyzed were as follows: • Bacteria: Paenibacillus polymyxa (Gram-positive bacterium capable of fixing nitrogen, used in biocontrol and antibiotic production), Rhodobacter spheroides (purple photosynthetic bacterium with diverse metabolism) and Lactobacillus casei (probiotic anaerobic bacterium). • Fungi: Tuber melanosporum (known as truffle, it is highly valued in gastronomy and of great economic value), Ustilago maydis (known colloquially as huitlacoche, it is a traditional part of Mexican food), Schizophyllum commune (saprophytic fungus of deciduous tree trunks) and Saccharomyces cerevisiae (yeast used industrially in the manufacture of bread, beer and wine). • Plants: Glycine max (known as soybeans, is widely used in human food), Oryza sativa (provides 20% of the world's food energy supply (FAO)). • Animals: Acynonyx jubatus (cheetah), Alligator mississippiensis (Mississippi alligator), Bos taurus (cow), Camelus dromedarius (Arabian camel), Gadus morhua (Norwegian cod), Gallus gallus (chicken) and Sus scrofa (pig). • Viruses: influenza A virus, human immunodeficiency virus 1 and SARS-CoV-2. ML / a / ZUZ 1 / U14Z0Ü Figure 6 of the accompanying drawings shows the amino acid (AA) proportions in bacterial proteomes, where the proportion of each AA is equal to the number of times the AA was found in the proteins recorded in the database multiplied by the average mass of the AA divided by the average mass of the recorded proteins. All 20 AAs were considered. The proportions in the proteomes of Paenibacillus polymyxa (blue dots), Rhodobacter spheroides (purple dots), and Lactobacillus casei (black dots) are shown. Figure 7 of the accompanying drawings shows the amino acid (AA) proportions in fungal proteomes, where the proportion of each AA is equal to the number of times the AA was found in the proteins recorded in the database multiplied by the average mass of the AA divided by the average mass of the recorded proteins. All 20 AAs were considered. The proportions in the proteomes of Tuber melanosporum (red dots), Ustilago maydis (purple dots), Schizophyllum commune (blue dots), and Saccharomyces cerevisiae (pink dots) are shown. Figure 8 of the accompanying drawings shows the proportions of amino acids (AA) in plant proteomes, where the proportion of each AA is equal to the number of times the AA was found in the proteins recorded in the database multiplied by the average mass of the AA divided by the average mass of the recorded proteins. All 20 AAs were considered. The proportions in the proteomes of Glycine max (green dots) and Oryza sativa (brown dots) are shown. Figure 9 of the accompanying drawings shows the proportions of amino acids (AAs) in animal proteomes, where the proportion of each AA is equal to the number of times the AA was found in the proteins recorded in the database multiplied by the average mass of the AA divided by the average mass of the recorded proteins. All 20 AAs were considered. The proportions in the proteomes of Acynonyx jubatus (black dots), Alligator mississippiensis (green dots), Bos taurus (red dots), Camelus dromedarios (brown dots), Gadus morhua (blue dots), Gallus gallos (orange dots), and Sos scrofa (pink dots) are shown. Figure 10 of the accompanying diagrams shows the amino acid (AA) proportions in viral proteomes, where the proportion of each AA is equal to the number of times the AA was found in the proteins annotated in the database multiplied by the average mass of the AA divided by the average mass of the annotated proteins. All 20 AAs were considered. The proportions in the proteomes of influenza A virus (purple dots), human immunodeficiency virus 1 (pink dots), and SARS-CoV-2 (blue dots) are shown. Figure 11 of the accompanying diagrams shows the proportions of amino acids (AAs) in the proteomes of animals, plants, fungi, viruses, and bacteria. The proportion of each AA is equal to the number of times the AA was found in the proteins recorded in the database multiplied by the average mass of the AA divided by the average mass of the recorded proteins. All 20 AAs were considered. The proportions in the proteomes of animals (purple dots), plants (black dots), fungi (red dots), viruses (gray dots), and bacteria (blue dots) are shown. wiA / a / zvzi iv i ^zoo Table 15. Average proportion and standard deviation of each AA in the 19 proteomes analyzed. The average proportion and standard deviation were calculated for each of the 20 AA from the proportions of the proteomes of animals (7), plants (2), fungi (4), viruses (3) and bacteria (3). 5 AA Average Proportion Standard Deviation A 0.0491 0.0152 C 0.0184 0.0112 D 0.542 0.0072 E 0.0756 0.0100 F 0.0467 0.0080 G 0.0344 0.0042 H 0.0315 0.0087 I 0.0517 0.0119 K 0.0643 0.0125 L 0.0927 0.0175 M 0.0296 0.0090 N 0.0393 0.0077 P 0.0487 0.0088 10 Q 0.0508 0.0083 R 0.0816 0.0143 S 0.0607 0.0107 T 0.0516 0.0057 V 0.0549 0.0108 w 0.0232 0.0082 Y 0.0403 0.0080 Analyzing proteomes reveals that amino acid (AA) compositions are conserved across all organisms. Standard deviation values for each of the 20 AAs showed very low dispersion from the average proportion. The proportions of AAs were very similar across different kingdoms (plants, animals, prokaryotes, and fungi), as well as in viruses. On average, the proportion of each AA deviated from the mean by between 0.0057 and 0.0175. Tryptophan, methionine, histidine, and cysteine were present in lower proportions, while leucine, arginine, and glutamic acid were present in higher proportions. Strains of the yeast Pichia pastoris X33 that overexpress the plasmid carrying the gene encoding the SEQ ID NO. 16 protein were isolated using different concentrations of the antibiotic zeocin (400 pg / ml, 800 pg / ml, and 1200 pg / ml). Two strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pAOX1 promoter, and three strains that grew at the highest concentration of zeocin expressing the SEQ ID NO. 16 protein gene under the pGAP promoter. These five selected strains were grown in rich medium (YPD: yeast extract, peptone, and glucose) for 18 hours, and then transferred to a 250 ml flask containing 50 ml of PTM1 salts medium at pH 4.0 with 2% glucose for strains with the pGAP promoter and 1.5% methanol for strains with the pAOX1 promoter. They were grown for 96 hours at 25°C with shaking at 250 rpm; samples were taken from these cultures every 24 hours. To verify that these strains produced the protein, the samples were centrifuged at 5,000 rpm, and the cells in the resulting pellet were discarded. The supernatant from these centrifugations was used to load denaturing polyacrylamide gels and stained with Coomassie blue. Figure 12 of the accompanying diagrams shows the results of these gels for the three strains with the pGAP promoter, while Figure 13 shows the two strains with the pAOX1 promoter. In both Figures 12 and 13, bands of different colors are visible on the far right, corresponding to molecular weight markers (ThermoFisher PageRuler™ Catalog No. 26619). The expected molecular weight of the protein is 8 kDa. The two lower bands on the molecular weight marker correspond to proteins of 10 and 15 kDa. It is noted that the native strain of Pichia pastoris X33 that was not transfected with the plasmid encoding the SEQ ID NO. 16 protein does not express any protein of that size. III. DISCUSSION V. Search for protein in databases with the composition of EAAs required by humans: Analysis of proportions and length distribution None of the 56,827,426 sequences annotated in PDB and Uniprot in a length range of 1-260 AA met the requirements of each of the 9 EAAs, having a maximum of 4 EAAs in the 1x ratio recommended in the human diet. This was found to be due to the proportions of amino acids (AAs) in the annotated proteins and the low probability of finding proteins of a certain length that met the requirements for each of the essential amino acids (EAAs). The EAAs required in the lowest amounts in the human daily intake, according to the QMS, were tryptophan, histidine, and methionine. These EAAs were found in the lowest proportions in the proteins of the analyzed databases (PDB, SwissProt, and TrEMBL), although in proportions greater than the 1x required. Therefore, by increasing the required proportion of each EAA by 2–4 times, proteins were found that met the RDA of at least 8 EAAs. Length distributions showed that 45% of the sequences annotated in PDB and Uniprot have 1-250 amino acids. The search for proteins with the composition MA / a / 2U21 / U 14200 required EAA had been carried out using a maximum length of 260 AA because a longer protein would have a higher content of non-essential AAs and more AA substitutions would have to be made to the protein to meet the EAA content. The possible minimum length that the proteins should have to satisfy the requirement of each EAA, calculated using the formula: P / n*118.7360=V was 560 AA for tryptophan. Because of this, it was highly unlikely that the code would find proteins shorter than 260 AA that met the requirement of 9 EAAs, since the required ratio of W was never met in this length range. Using a length of 1000 AA, it was also found that there were no proteins with more than 4 EAAs in the recommended 1x ratio due to the low probability of this occurring. For example, a minimum sequence of 560 AA containing tryptophan would occur only once in 19L560 sequences (1 / 19L560), and in a sample of 3-4.5×10¹⁵ sequences of 1-1000 AA, this was highly improbable. However, finding proteins that met the requirement of up to 8 EAAs showed that proteins are not present randomly, even when their AA composition is very similar. V Proportions in which each AA is found at the proteome level. A recent study titled "The Distribution of Biomass on Earth" has become the first major estimate of the total biomass on our planet (Bar-On and Phillips, 2018). This research shows that dietary choices have a significant impact on the habitats of living beings, where 60% of mammals on the planet are livestock (cows, sheep, goats, and pigs) and 70% of birds are poultry. This led to the hypothesis that the human diet might be inducing a trend in the amino acid composition of organisms, with higher levels of essential amino acids (EAAs) in the proteins of the organisms from which food production is based, such as cattle, pigs, chickens, etc. Using a bioinformatics approach made it possible to quickly analyze the proportions in which each of the 20 AAs were found, not only in the entire database, as had been done previously, but also at the proteome level. An interesting finding is that the composition of amino acids (AAs) is conserved in plants, animals, bacteria, fungi, and viruses. The 19 proteomes analyzed had very similar proportions of AAs. On average, the proportion of each AA deviated from the mean by between 0.0057 and 0.0175. Influenza A viruses, HIV, and SARS-CoV-2 also followed the same pattern in AA composition (L, R, and E in greater proportion and W, M, and H in lesser proportion). It had previously been observed that the probability of finding proteins with the desired composition of certain EAAs was very low; an example was the case of tryptophan, where the probability of finding a 560 AA sequence that had a W was 1 / 19L559. However, finding results of proteins with 8 EAAs in the recommended proportion showed that proteins are not found randomly, even though their AA composition appears to be random (0.047 on average for each AA), since there are other influencing factors, such as the bioenergetic cost of the AAs. MA / a / ZUZl / U1 Akashi and Gojobori published an article in the scientific journal PNAS in 2001, showing that the amino acid composition in Bacillus subtilis and E. coli reflected natural selection for improved metabolic efficiency. The total metabolic cost of biosynthesis in E. coli was obtained by considering the number of phosphate bonds in ATP and GTP molecules, as well as the number of hydrogen atoms available in NADH, NADPH, and FADH2 molecules, assuming a 2 P per H ratio. Tryptophan (W) is the most expensive amino acid, requiring 74.3 ATP molecules, followed by phenylalanine (52), histidine (38.3), methionine (34.3), isoleucine (32.3), thymine (30.3), leucine (27.3), arginine (27.3), cysteine (24.7), valine (23.3), proline (20.3), and threonine (18.7). Amino acids with a lower biosynthesis cost include glutamine (Q), glutamic acid (E), asparagine (N), aspartate (D), alanine (A), glycine (G), and serine (S). As can be observed, the metabolic pathways of essential amino acids (EAAs) are bioenergetically more expensive than those of non-essential amino acids, which is consistent with the results obtained in this project, where the amino acids found in the lowest proportions were tryptophan, methionine, and histidine. Tryptophan, which has the highest bioenergetic cost of the 20 amino acids, was one of the least abundant. Leucine, arginine, and glutamic acid, which have a lower bioenergetic cost, were found in the highest proportions in the analyzed proteomes. / Search for protein blends with the appropriate EAA composition. When searching for protein mixtures that would meet human requirements for EAAs, it was found that the RDA of up to 4 EAAs was met. If the sequences of more than two proteins were concatenated, the code no longer yielded positive results because, with a mixture of proteins, the essential amino acid (EAA) content would be diluted among many other non-essential amino acids. An example of this would be mixing 100 g of beef with 100 g of egg. The egg contains 1.001 g of lysine (K), which satisfies 47.66% of the daily requirement for lysine (2100 mg); while the beef has 2.002 g of lysine, which satisfies 95.33% of the requirement for this EAA. If these amounts are added together, it can be seen that the total amount ingested exceeds the total required amount of lysine. However, the proportion of lysine in that mixture is diluted. The fact that humans need to consume a greater amount of protein results in an increased load of certain amino acids. It is well known that many amino acid residues in proteins are susceptible to oxidation by various reactive oxygen species (ROS) and that these oxidized proteins accumulate during aging. Methionine and cysteine residues in proteins are particularly sensitive to this oxidation, so consuming proteins with high levels of these amino acids could be harmful to health (Stadtman, et al., 2005). An ideal protein, with the precise balance of essential amino acids (EAAs), would be of great importance to the general public, as well as to population groups sensitive to protein quality, such as older adults and athletes. Athletes need higher protein intake to preserve muscle mass, with a requirement of up to 2.4 g / kg. An athlete weighing 80 kg would need to consume approximately 192 g of protein per day. Natural protein sources like beef have ML / a / ZUZl / Ul 4Z00 contains approximately 22 g of protein per 100 g of meat, so to meet this protein requirement, an athlete would need to consume around 872 g of protein sources daily. This would be detrimental to health due to the increased amino acid (AA) load, which has been shown to be a source of essential fatty acids (EFAs), compounds that promote pro-inflammatory and pro-oxidative nephrotoxicity (Goldberg et al., 2004; Uribarri et al., 2005). In people with hypertension, diabetes, or older adults with impaired kidney function, having to consume so much protein to meet the daily EAA requirement could also have deleterious effects, with an increase in markers of kidney damage and a higher risk of developing cardiovascular disease (Wrone et al., 2003; Hoogeveen et al., 1998). V Search for proteins that complement natural protein sources: egg, milk, meat, fish, pork and chicken. When searching for proteins that would complement the EAA content of natural protein sources such as beef, chicken, egg, milk, pork, and fish, it was found that no sequence annotated in PDB, SwissProt, or TrEMBL in a length range of 1-260 AA could complement the EAA content of these foods to meet the required 1x ratio of each EAA. Similar to what was observed in the search for proteins with the appropriate EAA composition, the requirement for a greater number of EAAs was met as the proportion increased. At proportions 2–4 times higher than those required for each of the 9 EAAs, the natural sources that best complemented each other were found to be eggs, with up to 5 EAAs at the 2–3x and 3–4x proportions, and milk, with up to 7 EAAs at the 3–4x proportion, 6 EAAs at the 2–3x proportion, and 3 EAAs at the 1x proportion. For the proteins that complemented milk with 6 or 7 EAAs and eggs with 5 EAAs, it was analyzed whether these could serve as a dietary supplement for people with phenylketonuria; however, all of them had an excess of the amino acid phenylalanine. When analyzing the EAA content in 100 g of these foods, it was observed that milk is the least rich in EAAs (1795 mg EAA / 100g), along with eggs (6597 mg EAA / 100g); while beef (10,448 mg EAA / 100g) and chicken (11,858 mg EAA / 100g) have the highest EAA content. The fact that milk and eggs showed better results in protein supplementation could be due to these foods having the lowest EAA content, making it easier to supplement a food that is almost devoid of EAAs in the required proportion, since the proportions in which the EAAs are found exceed the required 1x ratio. / Selection of the protein closest to the optimal composition of EAAs and generation of collagen protein variants. The search for proteins suitable for human consumption in the databases led to the selection of Bos taurus collagen protein. This choice was based on previous results, which showed that no protein or protein mixture satisfies the required proportions of each of the nine essential amino acids (EAAs) at 1x. Collagen protein was chosen due to its proximity to the daily EAA requirement and the minimal number of amino acid substitutions required. Collagen fragments have already been expressed in Pichia pastoris, and it is a protein that is secreted into the extracellular environment, which is very useful in terms of purification because P. pastoris is a yeast that secretes few proteins (Cereghino et al., 2000). A very relevant issue regarding collagen protein is its lack of allergenicity. Previous immunoblotting studies have shown that collagen is not allergenic because it has no binding sites against human IgE antibodies (Wijaya et al., 2020; Hansen et al., 2004). The collagen protein would require the addition of 2 threonines, 1 histidine, 2 valines and the removal of 2 threonines to meet the RDA of each EAA, which would imply carrying out 7 substitutions in a protein of 86 AA. In several proteins, it has been observed that mutating a single AA residue is counterproductive to the structure and function of the protein (Purton et al., 2001). However, in this case, since it is a protein focused on human consumption, it is not required that the protein have a function, although the structure could be relevant for the secretion of the collagen protein, given that it has been seen that collagen proteins in random alpha helix conformation can be secreted into the medium, while collagen proteins in triple helix have not been able to be secreted, but remain intracellular (Nokelainen et al., 2001; Williams et al., 2008; He et al., 2015; Wang et al., 2014; Bin et al., 2011; Pakkanen et al., 2006). In the collagen variants generated in silico, proline and glycine were excluded as positions in the protein where mutations could occur to avoid folding problems, since these residues are what give the protein its helical structure. Nine variants were found that had 2 to 3 times the desired proportion of EAAs, with a percentage difference from the original collagen protein of 13.9% / Cloning of the gene encoding for the collagen protein of Bos taurus in an inducible and a constitutive expression vector. The P. pastoris system is proposed for collagen protein expression, as it is a GRAS (Generally Recognized as Safe) organism that has been approved for food production by the Food and Drug Administration (FDA) (Ahmad et al., 2014). Yields in recombinant human collagen protein production have reached up to 14.8 g / L in P. pastoris (Werten et al., 1999), and this system is widely used at an industrial level, scaling up from a flask to high-density cell cultures (Cereghino and Cregg., 2000). Induction of recombinant protein expression with methanol is the most common method (Cereghino and Cregg., 2000). However, the present invention proposes a glucose-inducible promoter, given the project's sustainable focus, as this protein is intended for human consumption and methanol induction would be toxic (Prielhofer et al., 2018). On the other hand, genetic engineering has been used to modify plants, such as corn, which have an increased production of the amino acid lysine (https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC2442549 / ). However, this approach does not solve the problems of the excessive water use required for its production, since it only provides one of the eleven essential amino acids and is produced in a plant whose proteins are highly diluted and poorly available due to the environment. MA / a / ZUZl / U14ZOO human. Currently, the intake of essential free amino acids, particularly leucine, is used to address the problem of sarcopenia. (http: / / www.ncbi.nlm.n¡h.aov / Dmc / articles / PMC3183816 / ). Finally, it is important to note that no protein in nature provides humans with the minimum amount of essential amino acids required in their daily diet. For the purposes of this invention, "minimum amount" means that if a human requires 20 grams of essential amino acids, obtaining that amount from natural sources would require consuming five or more times that amount in grams; the minimum would be 20 grams. This explains the need to consume large quantities of food, which has a negative impact on the environment. On the other hand, there are proteins that come close to containing the proportion of essential amino acids that can be used to change their essential amino acid composition with a minimal number of changes and achieve a protein optimized in composition and ingestible quantity for human nutrition. Although exemplary embodiments of the present invention have been described with reference to the drawings, it should be understood that these exemplary embodiments are merely illustrative and are not intended to limit the scope of the present invention.It is likely that an expert in the field may make various changes and modifications to these methods, but without departing from the true scope and spirit of the present invention, where such changes and modifications must be intended to be included within the scope of the present invention as written in the attached claims, such as using the algorithmic method developed by the inventors of the present invention, which, in addition to allowing obtaining the optimized protein comprising all the essential amino acids in human nutrition, allows obtaining any other proteins with different proportions of amino acids and for use in any other living being. In the particularly preferred modality described herein, it should be understood that the optimized protein comprising all the essential amino acids in human nutrition can be implemented in other ways, so that such modality of the method is merely illustrative. A person skilled in the art may understand that, in addition to the mutual exclusion of features, any possible combination may be adopted to incorporate and combine all the features disclosed by the description (including the appended claims, abstract, and drawings) and all processes or units of any method disclosed as such. Unless expressly stated otherwise, each feature disclosed by this description (including the appended claims, abstract, and drawings) may be replaced by an alternative feature that provides the same, equivalent, or similar purpose. Furthermore, those skilled in the art may understand that, even though some of the embodiments described herein comprise some features included in other embodiments, rather than other features, combinations of features from different embodiments are considered to fall within the scope of the present invention and constitute different embodiments. For example, in the claims, any of the embodiments for which protection is sought may be used in various combinations. References in the description to a modality or modalities indicate that the described modality may include a particular aspect, feature, structure, or characteristic, but not all modalities necessarily include that aspect, feature, structure, or characteristic. Furthermore, such phrases may, but do not necessarily, refer to the same modality mentioned elsewhere in the specification. Moreover, when a particular aspect, feature, structure, or characteristic is described in relation to a modality, it is within the knowledge of a person skilled in the art to affect or connect that aspect, feature, structure, or characteristic with other modalities, whether explicitly described or not. In other words, any element or characteristic may be combined with any other element or characteristic in different modalities, unless there is an obvious or inherent incompatibility, or it is specifically excluded. As such, an invention has been disclosed in terms of preferred embodiments thereof that fulfill each and every object of the present invention, as set forth above, and provide an optimized protein comprising all the essential amino acids in proportions suitable for human nutrition. Of course, a person skilled in the art may contemplate various changes, modifications, and alterations to the teachings of the present invention, but without departing from its intended spirit and scope. It is intended that the present invention be limited only by the terms of the appended claims. The terminology used herein is solely for the purpose of describing particular or preferred embodiments and is not intended to limit the invention. As described herein, the singular forms a, an, and the are intended to include the plural forms as well, unless the context clearly indicates otherwise. It shall be further understood that the terms comprise and / or comprising, when described in this specification, specify the presence of stated features, whole numbers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more additional features, whole numbers, steps, operations, elements, components, and / or groups thereof. As described herein, the term and / or includes any and all combinations of one or more of the associated enumerated elements.Throughout the description, unless explicitly stated otherwise, the word understand and variations such as includes or that includes shall be understood to imply the inclusion of the stated elements, but not the exclusion of any other element. Claims may be drafted to exclude any optional elements. As such, this statement is intended to serve as a basis for precedent regarding the use of exclusive terminology, such as solely, only, and the like, in connection with the mention of claim elements or the use of a negative limitation. The terms preferably, preferred, prefer, optionally, may, and similar terms are used to indicate that a referenced element, condition, or step is an optional (not required) feature of the invention. In summary, although the preceding detailed description of the present invention has referred to certain modalities of the optimized protein comprising all the essential amino acids in the proportions suitable for human nutrition, it should be emphasized that numerous modifications to these modalities are possible, but without departing from the true scope of the present invention, such that the characteristics described in the aforementioned modalities, shown in the figures and claimed in the claims chapter, as well as the characteristics of different modalities not described herein, may be used individually or in any arbitrary combination for the realization of the present invention.Therefore, it should be understood that the embodiments of the present invention are merely illustrative and are not intended to limit the scope of the present invention except as provided in the prior art and the appended claims.
Claims
1. An optimized protein for human nutrition, characterized in that it comprises one or more amino acid sequences selected from the group of sequences: SEQ ID NO. 1; SEQ ID NO. 2; SEQ ID NO. 3; SEQ ID NO. 4; SEQ ID NO. 5; SEQ ID NO. 6; SEQ ID NO. 7; SEQ ID NO. 8; SEQ ID NO. 9; SEQ ID NO. 10; SEQ ID NO. 11; SEQ ID NO. 12; SEQ ID NO. 13; SEQ ID NO. 14; SEQ ID NO. 15; SEQ ID NO. 16; SEQ ID NO. 17; SEQ ID NO. 18; SEQ ID NO. 19; SEQ ID NO. 20, either individually or in combinations of two or more, wherein said optimized protein comprises all the essential amino acids in proportions suitable for human nutrition.
2. The protein optimized for human nutrition according to claim 1, further characterized in that it comprises the amino acid sequences as set out in SEQ ID NO. 12 and SEQ ID NO. 16, wherein said optimized protein comprises all the essential amino acids in proportions suitable for human nutrition.
3. The protein optimized for human nutrition according to any of claims 1 or 2, further characterized in that it comprises the amino acid sequence as set out in SEQ ID NO. 16, wherein said optimized protein comprises all the essential amino acids in proportions suitable for human nutrition.
4. The protein optimized for human nutrition according to claim 3, further characterized in that the corresponding nucleotide sequence for the protein SEQ ID NO. 16 is as follows: 5'CAACCACCCAGGTCCACCAGGTCCACCAGGTCCACCAGTTTCTGCTATGTTGCCAGGT CCATTCGGTTTGCCAGGTTTCCCAGGTACTCCAGGTATGAAGGGTATTCAAGGTGAGA GAGGTTTGCCAGGTGAGAAGGGTGAGGTTGGTTTGCCAG GTGAGTCTAGATTGGGTCCACCAGGTTCTACTGGTTCTAGAGGTGTTCCAGGTTCCACC AGGTAGACCAGGTGACTCTGGTATTAAG-3' 5 .- The protein optimized for human nutrition according to claim 4, further characterized because the amino sequence of SEQ ID NO. 16 is expressed in the yeast Pichia pastorís under 2 different promoters: pGAP and pAOX1; where pGAP is a system in which gene and protein expression is induced in the presence of glucose and pAOX1 in the presence of methanol.
6. A method for selecting a protein sequence for human nutrition that most closely approximates the proportion of essential amino acids, characterized in that it comprises performing a search in public databases of protein sequences comprising the following steps: 1) defining the proportion of essential amino acids per gram of total protein that is desired to be found in any protein (PAAE), where this proportion is specified in a range of values (RVPAAE); 2) defining how many essential amino acids must comply with the RVPAAE (AAEE: Expected Essential Amino Acids), which can be from 1 to 9; 3) searching in databases of protein sequences, those that satisfy steps 1) and 2); 4) repeating steps 1), 2) and 3) for different possible values of AAEE;5) Select proteins with higher ESA content, preferably those proteins that satisfy a greater number of essential amino acids in the appropriate proportion for human nutrition; where, starting from the sequence selected in the previous procedure, changes are made to that sequence to make it contain the essential amino acids in the desired proportions of essential amino acids.
7. A method for synthesizing a protein optimized for human nutrition, characterized in that it comprises: (a) performing a search in public databases of protein sequences; (b) initializing the system energy; (c) calculating the value of the sequence energy function; (d) printing the solution, provided that the value of the sequence energy function is equal to 0; otherwise, proceeding to the following steps; (e) randomly selecting the positions in the sequence to be mutated; the number of positions being the method parameter specified in the following step f); (f) verifying that the selected positions are among the positions susceptible to being changed specified by the corresponding method parameter in step c); (g) randomly selecting the essential amino acids to which the amino acids present in the selected positions are to be replaced;(h) perform the substitutions specified in step f) at the positions selected in step e); if the new sequence reduces the energy value, retain it; otherwise, evaluate whether to retain that sequence by applying the following formula: e(previous energy - new energy) / temperature; if the result is greater than a number chosen at random between 0 and 1, then retain the sequence for the next cycle; (i) repeat the previous steps until the number of printed sequences equals the method parameter specified in step h), or a system energy value less than 1 has been reached; and (j) produce or synthesize a protein with the desired sequence(s).
8. The method for synthesizing a protein optimized for human nutrition according to claim 7, characterized in that the production or synthesis of the protein comprises the steps of: expressing the protein in the yeast Pichia pastoris under 2 different promoters: pGAP and pAOX1; wherein pGAP is a system in which the expression of the gene and the protein is induced in the presence of glucose and pAOX1 in the presence of methanol and wherein the corresponding nucleotide sequence for the protein SEQ ID NO. 16 is the following: 5'CAACCACCCAGGTCCACCAGGTCCACCAGGTCCACCAGTTTCTGCTATGTTGCCAGGT CCATTCGGTTTGCCAGGTTTCCCAGGTACTCCAGGTATGAAGGGTATTCAAGGTGAGA GAGGTTTGCCAGGTGAGAAGGGTGAGGTTGGTTTGCCAGGTCCACCAGGTCCACAAG GTGAGTCTAGATTGGGTCCACCAGGTTCTACTGGTTCTAGAGGTGTTCCAGGTCCACC AGGTAGACCAGGTGACTCTGGTATTAAG-3' 9. The method according to claims 7 and 8, characterized in that the search in public databases of protein sequences comprises the following steps: 1) defining the proportion of essential amino acids per gram of total protein that is desired to be found in any protein (PAAE), wherein this proportion is specified in a range of values (RVPAAE); 2) defining how many essential amino acids must comply with the RVPAAE (AAEE: Expected Essential Amino Acids), which can be from 1 to 9; 3) searching in protein sequence databases for those that satisfy steps 1) and 2); 4) repeating steps 1), 2), and 3) for different possible values of AAEE; 5) selecting the proteins whose AAEE are higher, preferably those proteins that satisfy a greater number of essential amino acids in the appropriate proportion for human nutrition;where additionally, based on the sequence selected in the previous procedure, changes are made to the sequence to make it contain the essential amino acids in the desired proportions of essential amino acids.; 10. A food formulation or supplement characterized in that it comprises the protein optimized for human nutrition in accordance with claims 1, 2 or 3, in combination with acceptable vehicles and / or excipients.
11. The food formulation or supplement, according to claim 10, for use in the treatment of a nutritional deficiency