Systems and devices for biomining and methods thereof

Cell-free biomining systems using peptide-metal binding interactions address inefficiencies in traditional mining by enhancing metal recovery rates and adapting to diverse geological conditions, offering a sustainable and efficient extraction method for precious metals.

WO2025250829A1PCT designated stage Publication Date: 2025-12-04GIRAFFE BIO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031487
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-01
Filing Date
2025-05-29
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Traditional mining methods for precious metals face challenges such as low extraction rates and yields, high costs, environmental impact, and inefficiencies in processing lower-grade deposits, with chemical catalysts and microbial processes exhibiting inconsistent performance across varying geological conditions.

Method used

Utilizing cell-free biomining systems that leverage peptide-metal binding interactions to identify and extract target metals through patterned arrays of peptide molecules on substrates, enabling efficient and adaptable metal recovery across different mining sites.

Benefits of technology

Enhances metal recovery rates, reduces water and chemical input reliance, and adapts to varying geological conditions, providing a cost-effective and sustainable solution for extracting precious metals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025031487_04122025_PF_FP_ABST
    Figure US2025031487_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides systems for biomining a target metal from a medium obtained from a geological sample and related methods thereof. In some embodiments, the system can comprise a substrate comprising a patterned array of peptide molecules on a surface of the substrate. The patterned array of peptide molecules can comprise a plurality of candidate peptide molecules for binding and extracting the target metal from the medium. The substrate can be utilized to identify (or design) one or more optimal peptide molecules for binding to and biomining the target metal from geological samples.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND DEVICES FOR BIOMINING AND METHODS THEREOFCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Patent Application No. 63 / 654,948, filed June 1, 2024, which is entirely incorporated herein by reference.BACKGROUND

[0002] Recent decades have seen a growing demand for precious metals. During a mining process, raw ores obtained from a mining site are processed to extract such metals of interest.SUMMARY

[0003] The present disclosure provides methods and systems for (i) generating (e.g., printing) a chip having an array of polypeptides and / or (ii) utilizing such array of the polypeptides to identify and / or further design one or more optimal polypeptides capable of binding (e.g., capturing) one or more metals of interest. In some embodiments, the one or more optimal polypeptides can be used (e.g., via presentation on a chromatography column) to bind and separate the one or more metals of interest in a larger scale. In some embodiments, the methods and systems provided herein can be utilized to generate the optimal polypeptide(s) that are specific, unique, and / or tailored for each mining site.

[0004] An aspect of the present disclosure provides a system for biomining a target metal from a medium obtained from a geological sample, the system comprising: (1) a substrate comprising a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and (2) a sensor configured to measure a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or another geological sample.

[0005] Another aspect of the present disclosure provides a method for biomining a target metal from a medium obtained from a geological sample, the method comprising: (a) contacting the medium with a substrate, wherein the substrate comprises a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and (b) measuring, via a sensor, a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable toidentify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or an additional geological sample.

[0006] Another aspect of the present disclosure provides a method for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the method comprising: (a) providing at least one amino acid or at least one peptide; and (b) directing, via an optical source of a printer, a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of the at least one amino acid or the at least one peptide on or adjacent to the surface of the substrate, thereby printing the patterned array of peptide molecules on the surface, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

[0007] Another aspect of the present disclosure provides a system for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the system comprising: (1) a printer comprising: an optical source configured to direct a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of at least one amino acid or at least one peptide one or adjacent to the surface of the substrate; and (2) a controller configured to instruct the optical source to direct the light towards the surface to print, via the conjugation, an array of peptide molecules on the surface, wherein the array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

[0008] Another aspect of the present disclosure provides a system for detecting a presence of a sample on a microarray, comprising: (1) a 3D printer, comprising a light engine; and (2) an optical pick-up unit; (3) a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate; (4) a reader to capture kinetic data for each of the one or more immobilized molecules; and (5) an artificial intelligence (Al) algorithm to generate a model based on codebase.

[0009] Another aspect of the present disclosure provides a system for detecting a presence of a sample on a microarray, comprising: (1) a 3D printer, comprising: a light engine; an optical pick-up unit; and a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate; (2) a reader to capture endpoint measurements for each of the one or more immobilized molecules; and (3) an artificial intelligence (Al) algorithm to generate a model based on codebase.

[0010] Another aspect of the present disclosure provides a method for detecting a presence of a sample on a microarray, comprising (a) providing a substrate; (b) positioning one or more molecule variants on a surface of the substrate; (c) performing analyte perfusion; (d) capturing kinetic data for the one or more molecule variants; (e) analyzing the captured kinetic data; (f) employing an artificial intelligence (Al) algorithm on the captured kinetic data; (g) generating one or more patterns from a plurality of molecular interactions of the one or more molecule variants; and (h) generating a codebase based on one or more patterns.

[0011] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0012] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0013] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0014] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanyingdrawings (also “Figure” and “FIG.” herein), of which:

[0016] FIG. 1 shows an example of a system for printing molecules onto a photo-sensitive surface;

[0017] FIG. 2A and FIG. 2B show examples of a method for detecting the presence of a sample on a microarray;

[0018] FIG. 3 shows an example of a technology platform to implement a system for conducting molecular interaction studies, where immobilized molecules on a chip surface interact with perfused analytes to generate valuable experimental datasets;

[0019] FIG. 4 shows an example of a flow of the technology platform of FIG. 3;

[0020] FIG. 5 shows an example of a system for biomining a target metal from a medium obtained from a geological sample;

[0021] FIG. 6 shows an example of a method for biomining a target metal from a medium obtained from a geological sample;

[0022] FIG. 7 shows an example of a method for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample;

[0023] FIG. 8 shows an example of a system for printing a patterned array of peptide molecules;

[0024] FIG. 9 shows an example of translation of RNA molecules to peptide molecules on a patterned array;

[0025] FIG. 10A shows an example of a method for generating one or more biomolecules for biomining a target metal;

[0026] FIG. 10B shows an example of a method for biomining a target metal from a medium obtained from a geological sample;

[0027] FIG. 11 shows an example of a method for biomining a target metal during a metal recovery process;

[0028] FIG. 12A and FIG. 12B show examples of an adsorption column for biomining a target metal from a medium obtained from a geological sample;

[0029] FIG. 13 shows a computer system that is programmed or otherwise configured to implement methods provided herein; and

[0030] FIG. 14 shows an example of a chip architecture for biomining a target metal from a medium obtained from a geological sample.DETAILED DESCRIPTION

[0031] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of exampleonly. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0032] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0033] Whenever the term “at most,” “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at most,” “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0034] The term “biomining,” when used herein, includes a mining operation wherein one or more biological materials is used. Example mining operations include excavating a geological site, obtaining a geological sample from the geological site, obtaining a metal-comprising mixture or medium (e.g., metal in solution, metal in suspension, metal in colloid, etc.) from the geological sample, or purifying / isolating / extracting one or more minerals from the solution.Alternatively, or in addition to, the term “biomining” can refer to isolation or extraction of one or more minerals, via the biological material(s), from a source that may or may not be a natural geological sample. For example, such source may be waste (e.g., metal recycling). The one or more minerals can include a metal, a rock, and / or a crystal. The biological material(s) may or may not comprise one or more living organisms. The biological material(s) may or may not be produced by a living organism (e.g., produced synthetically outside of a living organism). Nonlimiting examples of the biological material(s) may comprise one or more tissues, one or more cells, one or more cellular components, one or more enzymes (e.g., functional or non-functional variant), one or more peptides, one or more nucleic acids, or one or more small molecules.

[0035] The term “metal,” “metal of interest,” or “target metal,” as used interchangeably herein, includes metals and substances having metallic properties which can be obtained via biomining. A target metal can include alkali metals, alkaline earth metals, transition metals, rare earth metals, noble metals, alloys, and metalloids. A target metal can comprise a metal ion. Non-limiting examples of a chemical state of the target metal (e.g., during biomining) can include (i) metal ions in a medium (e.g., Cu2+, Li+, Al3+, Fe2+, etc.) or as part of salts (e.g., malachite (Cu2(COs)(OH)2), azurite (Cu3(COs)2(OH)2), cuprite (CU2O), Bornite (CusFeS^,spodumene (LiAl(SiOs)2), lepidolite (K(Li,Al Si,Al)4Oio(F,OH)2), petalite (LiAlSi40w), etc.), (ii) free metal as bulk metal or nanoparticles, (iii) metal complexes comprising metal ions bound to ligands (e.g., organic or inorganic ligands), (iv) metal oxides (e.g., CuO, Li2O, ZnO, etc.), and (v) metal alloys (e.g., bronze, brass, aluminum-lithium alloys (Al-Li), white gold, etc.).

[0036] The term “amino acid,” as used herein, refers to a molecule containing both an amino group and a carboxyl group. Suitable amino acids include, without limitation, both the D- and L- isomers of the naturally-occurring amino acids, as well as non-naturally occurring amino acids prepared by organic synthesis or other metabolic routes. The term amino acid, as used herein, includes without limitation, a-amino acids, natural amino acids, non-natural amino acids, and amino acid analogs. The term “a-amino acid” refers to a molecule containing both an amino group and a carboxyl group bound to a carbon which is designated the a-carbon. The abbreviation “b-” prior to an amino acid represent a beta configuration for the amino acid.

[0037] The terms “natural amino acid” or “naturally occurring amino acid,” as used interchangeably herein, refer to any one of the twenty amino acids commonly found in peptides synthesized in nature: alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and / or valine. The terms “non-natural amino acid” or “synthetic amino acid,” as used interchangeably herein, refer to an amino acid which is not one of the twenty amino acids commonly found in peptides synthesized in nature. Non-limiting examples of non-natural amino acids include azidohomoalanine (AHA), biphenylalanine, L-3,4- dihydroxyphenylalanine (L-DOPA), canavanine, O-Methyl-L-tyrosine, p-Boronophenylalanine (BP A), homopropargylglycine (Hpg), norvaline (Nva), norleucine (Nle), cyclohexylalanine (Cha), 2-aminoisobutyric acid (Aib), N-methylated residues, beta-amino acids, ornithine (Om), and / or citrulline (Cit).

[0038] The term “peptide,” when used herein, includes any molecule comprising two or more amino acids. The two or more amino acids can comprise natural amino acids and / or non-natural amino acids. Peptides can include at least one peptide bond between an amine group of a first amino acid and a carboxyl group of a second amino acid. The term “peptide” can include any naturally occurring or engineering polypeptide, protein, or enzyme. A peptide may or may not be a polypeptide (e.g., a polypeptide can comprise at least 51 or more amino acid residues).

[0039] The term “binding,” when used herein, can include any interaction between two or more molecules. Binding can occur as a result of one or more ionic bonds, one or more covalent bonds, one or more non-covalent bonds, or one or more intermolecular forces.

[0040] The term “substrate,” when used herein, can include any material having a surfacewhere one or more biological materials can be positioned, deposited, coupled to, or printed on (e.g., directly or indirectly). The substrate can comprise a material having a rigid, semi-rigid, or flexible surface. The substrate can comprise one or more films or layers. Each film or layer can comprise a metal, a crystal, a polymer, or any combination thereof.

[0041] The term “patterned array,” or “array,” as used interchangeably herein, can include a microarray, a nanoarray, a sequencing array formed as a patterned flow cell, and so forth. Such devices comprise sites at which analytes (e.g., biological materials, peptides, target metals) may be located for processing and analysis. The sites can be spatially disposed in a repeating pattern, a non-repeating pattern, or in a random arrangement on one or more surfaces of a substrate.Overview

[0042] Global electrification and renewable energy growth hinge on a reliable supply of precious metals, such as copper and lithium. As more industries rely on products made from such precious metals, including wiring, electronics, and infrastructure, traditional mines struggle to keep up with demand. Environmental concerns, regulatory pressure, and the diminishing grade of ore further complicate extraction. New approaches are urgently needed to unlock additional supply of these precious metals without exacerbating environmental harm.

[0043] A key challenge in mining is that many precious metals are extracted at limited rates and / or yields. Another key challenge is that high grade deposits may be largely exhausted, leaving only lower-grade deposits, which demand more energy and chemical inputs, thereby increasing costs. When used on the lower-grade deposits, traditional heap leaching or flotation can have a low rate and / or yield of extraction. For example, high-grade copper oxide deposits are largely exhausted today, leaving only low-grade sulfides like chalcopyrite. Traditional heap leaching or flotation can leave up to half the copper behind in tailings.

[0044] Conventional methods for precious metal extraction face important limitations. For example, chemical catalysts are associated with high input costs and hazardous byproducts. As another example, microbes exhibit inconsistent performance in harsh mine environments, and different mining sites’ distinct geology limits their efficacy. Current mining processes are also associated with environmental concerns. For example, large-scale crushing, acid usage, and water-intensive processes can harm local ecosystems. Water scarcity, especially in arid regions, can impose additional hurdles and costs.

[0045] Thus, there is an unmet need for metal extraction methods that are cost-effective, efficient, sustainable, and consistent across mining sites and geological conditions. Cell-free biomining systems and methods, as described herein, can provide higher metal recovery rates,reduce reliance on water and toxic reagents, and adapt quickly to varying geological conditions.

[0046] Accordingly, the present disclosure provides methods and systems for biomining target metals using biomolecule-metal (e.g., peptide-metal) binding interactions. Such interactions may be performed in a cell-free environment.

[0047] FIG. 10B schematically illustrates an example method for biomining a mineral (or target metals) from a medium obtained from a geological sample in accordance with example embodiments described herein. The method can comprise a discovery step. During the discovery step, the method can comprise providing a plurality of biomolecules (e.g., peptide candidates) and a scaffold (e.g., polymer substrates, such as polymer beads). The method can comprise a formulation step. During the formulation step, the method can comprise assembling the biomolecules and the scaffold (e.g., by synthesizing a sequence of biomolecules and coupling the sequence to the scaffold) to generate a biosorbent. The method can comprise a flotation step. During the flotation step, the biosorbent can be contacted with a medium comprising the mineral. The target metal can adsorb to the biosorbent. The method can comprise an analytical step. The analytical step can comprise measuring an adsorption of the mineral to the biosorbent over a period of time. The analytical step can further comprise comparing the adsorption of the mineral to the biomolecule with an adsorption of the mineral to a chemical reagent. In some cases, extraction or recovery rate of the mineral (or target metals) by the biosorbent as provided herein can be comparable or greater than that by a chemical process (e.g., hydrometallurgy such as acid leaching for copper, cyanide leaching for gold, etc.).

[0048] As provided herein, one or more identified peptide candidates may be used in target molecule identification and purification systems, such as, for example, column systems, bead conjugates, or membrane-bound platforms for industrial metal recovery or environmental remediation.

[0049] FIG. 11 schematically illustrates an example method for biomining a target metal during a metal recovery process in accordance with example embodiments described herein. The method can comprise providing the geological sample as a feed. The method can comprise grinding the geological sample. The method can optionally comprise a step 1102 of biomining the target metal during the grinding. The method can comprise filtering an output of the grinding based on particle size. The method can optionally comprise a step 1104 of biomining the target metal during the filtering. The method can comprise scavenging the target metal from a first portion of the output of the filtering. The method can optionally comprise a step 1106 of biomining the target metal during the scavenging. The method can comprise regrinding a second portion of the output of the filtering. The method can comprise extracting the target metal fromthe output of the regrinding via a first cleaner. The method can optionally comprise a step 1108 of biomining the target metal during the extracting at the first cleaner. The method can comprise further extracting the target metal from the output of the first cleaner via a second cleaner. The method can optionally comprise a step 1110 of biomining the target metal during the further extracting at the second cleaner.Detection of peptide-metal binding

[0050] Certain aspects of the present disclosure provide systems and method for biomining a target metal from a medium obtained from a geological sample.

[0051] In an aspect, the present disclosure provides a system for biomining a target metal from a medium (e.g., solution) obtained from a geological sample. The system can comprise a substrate. The substrate can comprise a patterned array of peptide molecules on a surface of the substrate. The patterned array of peptide molecules can comprise a plurality of candidate peptide molecules and extracting the target metal from the medium. The system can further comprise a sensor. The sensor can be configured to measure a metal binding performance for each of the plurality of candidate peptide molecules. The metal binding performance can be usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or another geological sample.

[0052] In another aspect, the present disclosure provides a method of use of any one of the systems provided herein. The method can be for biomining a target metal from a medium (e.g., solution) obtained from a geological sample. The method can comprise contacting the medium with a substrate. The substrate can comprise a patterned array of peptide molecules on a surface of the substrate. The patterned array of peptide molecules can comprise a plurality of candidate peptide molecules for binding and extracting the target metal from the medium. The method can further comprise measuring, via a sensor, a metal binding performance for each of the plurality of candidate peptide molecules. The metal binding performance can be usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or an additional geological sample.

[0053] FIG. 6 shows an example method 600 for biomining a target metal from a medium obtained from a geological sample. The method 600 can optionally comprise a step 602 of obtaining a medium from a geological sample. The method 600 can comprise a step 604 of contacting the medium with a substrate comprising a patterned array of candidate peptide molecules for binding a target metal. The method 600 can comprise a step 606 of measuring a metal binding performance for each of the candidate peptide molecules. The method 600 canoptionally include a step 608 of identifying at least one optimal peptide molecule for biomining the target metal.

[0054] In some embodiments, the target metal may or may not be a part of an organometallic molecule. In some embodiments, the target metal may or may not be a part of a synthetic molecule (e.g., a synthetic organometallic molecule, a synthetic metal nanoparticle, etc.).

[0055] In some embodiments, the contacting (e.g., between the medium and the substrate) and / or the measuring (e.g., the metal binding performance of one or more candidate peptide molecules) are performed at or near a same mining site as the geological sample. The same mining site can comprise a same mine, a same region of the mine, a same depth of the mine, a same excavation site, a same drilling site, a same geographical region, or a same country. Performing the contacting and / or the measuring at or near the same mining site as the geographical sample can allow optimal peptide molecules to be identified with faster turn-around times, compared to sending the sample to an external site for testing, which can improve time and cost efficiency. Performing the contacting and / or the measuring at or near the same mining site can also allow the peptide molecules to continuously adapt to current geological conditions (e.g., temperature, humidity, pH) of the mining site in real time. For example, as a depth of a mining site changes during excavation, the geological condition can change correspondingly.Performing the contacting and / or the measuring at or near the mining site can allow optimal peptide molecules to be identified based on the most current geological conditions, which can further improve mining efficiency. Alternatively, or in addition to, the contacting and / or the measuring can be performed not at the same mining site as the geological sample. The contacting and / or the measuring can be performed at a different mining site, such as a different mine, a different region of the same mine, a different excavation site, a different drilling site, a different geographical region, or a different country. The contacting and / or the measuring can be performed at a non-mining site, such as a surveying area, a planning area, a research area, or a laboratory.

[0056] In some embodiments, the geological sample and the another geological sample are from a same geological resource. The same geological resource can comprise a same mine, a same mining site, a mine for a same mineral, a same type of mine, a same mineral composition, a same geological formation, or a mine in a same geographical region. The same type of mine can comprise an underground mine, a surface mine, a placer mine, or an in-situ mine. The same mineral composition can be a same alloy, a same mixture of minerals, or a same ore. The same geological formation can be a same rock formation, a same vein or ore, or a same mineral deposit. Alternatively, or in addition to, the geological sample and the another geological samplecan be not from the same geological resources. The different geological resource can comprise a different mine, a different mining site, a mine for a different mineral, a different type of mine, a different mineral composition, a different geological formation, or a mine in a different geographical region.

[0057] In some embodiments, (i) the medium (e.g., for binding and extracting the target metal from the geological sample) has a pH that is less than about 7, ranges between about 2 and about 6, or ranges between about 3 and about 5; (ii) the medium comprises one or more metal chelating agents; and / or (iii) the medium comprises a plurality of mineral particles comprising the target metal, wherein (1) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (2) the plurality of minerals has an average particle size of at least about 1 micrometer, at least about 5 micrometers, or at least about 10 micrometers.

[0058] In some embodiments, (i) the method comprises, prior to (a), preparing the medium comprising the target metal at a pH that is less than about 7, ranges between about 2 and about 6, or ranges between about 3 and about 5; (ii) the method comprises, prior to (a), adding one or more metal chelating agents; and / or (iii) the method comprises adding a plurality of mineral particles comprising the target metal to the medium, wherein (1) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (2) the plurality of minerals has an average particle size of at least about 1 micrometer, at least about 5 micrometers, or at least about 10 micrometers.

[0059] The medium can comprise a filtered acidic aqueous medium having a pH between about 3 and about 5. The one or more metal chelating agents can include ethylenediaminetetraacetic acid (EDTA), ethylenediamine, citrate. The one or more metal chelating agents can aid in mobilizing loosely bound metals. The one or more metal chelating agents can be selectively included in the medium when loosely bound metals are present in the medium and / or the geological sample. The medium can comprise a disaggregated mineral slurry. The plurality of mineral particles can be present in the medium at a concentration of at least or at most about 1%, at least or at most about 2%, at least or at most about 5%, at least or at most about 10%, at least or at most about 15%, or at least or at most about 20% weight by volume.The plurality of minerals can have an average particle size of at least or at most about 1 micrometer, at least or at most about 5 micrometers (pm), at least or at most about 10 pm, at least or at most about 20 pm, at least or at most about 50 pm, at least or at most about 100 pm, at least or at most about 200 pm, at least or at most about 500 pm, or at least or at most about 1000 pm.The plurality of minerals can be filtered to remove particles having a particle size of more than about 10 pm. The plurality of minerals can be filtered to remove particles having a particle size of more than about 1 pm, about 2 pm, about 5 pm, about 10 pm, about 20 pm, about 50 pm, about 100 pm, about 200 pm, about 500 pm, or about 1000 pm. Filtering the plurality of minerals can prevent a channel for conducting the medium from being clogged. Filtering the plurality of minerals can ensure smooth perfusion of the medium across the patterned array.

[0060] In some embodiments, the measurement of the metal binding performance is nonfluorescence-based. Alternatively, or in addition to, the measurement of the metal binding performance can be fluorescence-based. A fluorescent-based measurement can comprise contacting the medium and / or the substrate with one or more fluorescent probes. The one or more fluorescent probes can have fluorescent activity associated with the metal binding performance. For example, the one or more fluorescent probes can be fluorescent when a metal molecule is bound to a peptide molecule and non-fluorescent otherwise. As another example, the one or more fluorescent probes can be fluorescent at a first wavelength when the metal molecule is bound to the peptide molecule, and fluorescent at a second wavelength different from the first wavelength when the metal molecule is not bound to the peptide molecule.

[0061] In some embodiments, the sensor comprises a surface plasmon resonance (SPR) sensor. The SPR sensor can generate an SPR signal based on the detected at least a portion of the light that is reflected from the another surface. An attribute of the signal can be dependent on a refractive index of the another surface. The refractive index can be indicative of the metal binding performance. For example, the another surface can have a higher or a lower refractive index when a metal molecule is bound to a peptide molecule compared to when the metal molecule is not bound to the peptide molecule. A number of metal molecules bound to a peptide molecule can be determined, based on the SPR signal. Alternatively, or in addition to, the sensor can comprise an interferometry sensor (e.g., biolayer interferometry (BLI) sensor) and / or a calorimetry sensor (e.g., isothermal titration calorimetry (ITC) sensor). The interferometry sensor can detect optical paths associated with two or more lights. The two or more lights can have a same light source. At least one light of the two or more lights can be reflected from the another surface of the substrate, while one or more other lights are not reflected from the another surface of the substrate. The interferometry sensor can generate a signal indicative of lengths of the respective optical paths of each of the two or more lights. The metal binding performance can be determined based on a difference between the lengths of the respective optical paths of the at least one light reflected from the another surface of the substrate and the one or more other lights not reflected from the another surface of the substrate. For example, the at least one componentlight reflected from the another surface of the substrate can have a longer or shorter optical path compared to the remaining component lights only if a metal molecule is bound to a peptide molecule. The calorimetry sensor can detect a heat emitted by the substrate. The emitted heat can be emitted in response to a stimulus applied to the substrate. An attribute of the emitted heat can be indicative of the metal binding performance. For example, the emitted heat can be higher or lower if a metal molecule is bound to a peptide molecule, compared to if the metal molecule is not bound to the peptide molecule. Alternatively or in addition to, sensor readouts (or chip readouts) may be obtained via localized surface plasmon resonance (LSPR) and / or optical absorption sensing system.

[0062] In some embodiments, the sensor comprises an optical source configured to direct a light comprising a single wavelength or a range of wavelengths to another surface of the substrate that is opposite of the surface. Alternatively, or in addition to, the optical source can be configured to direct the light to the surface. The surface can be a surface of the substrate comprising a transparent layer. Alternatively, or in addition to, the surface can be a surface of the substrate comprising a metal layer. The surface and / or the another surface can be a top surface facing the optical source. For example, the surface can be a top surface and the another surface opposite of the surface can be a bottom surface. The single wavelength or the range of wavelengths can correspond to infrared light, visible light, and / or ultraviolet light. The ultraviolet light can have a wavelength of about 10 nm to about 100 nm, about 100 nm to about 280 nm, about 280 nm to about 315 nm, and / or about 315 nm to about 400 nm. The visible light can have a wavelength of about 380 nm to about 450 nm, about 450 nm to about 500 nm, about 500 nm to about 565 nm, about 565 to about 590 nm, about 590 to about 625 nm, and about 625 to about 780 nm. The infrared light can have a wavelength of about 750 nm to about 1.4 pm, about 1.4 pm to about 3 pm, and / or about 3 pm to about 1 mm. In some cases, the single wavelength or the range of wavelengths can between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm. The single wavelength or the range of wavelengths can be at least or at most about 300 nm, at least or at most about 350 nm, at least or at most about 400 nm, at least or at most about 450 nm, at least or at most about 500 nm, at least or at most about 550 nm, or at least or at most about 600 nm.

[0063] In some embodiments, the light can comprise a laser light, a light-emitting diode (LED) light, an incandescent light, and / or a fluorescent light. In some cases, the laser light can have a wavelength range between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm. The laser light can have a wavelength of at least or at most about 300 nm, at least or at most about 350 nm, at least or at most about 400 nm, at least or atmost about 450 nm, at least or at most about 500 nm, at least or at most about 550 nm, or at least or at most about 600 nm. The laser light can comprise a 405 nm and / or a 532 nm laser. The wavelength or wavelength range of the laser light can be determined based on a linker group coupled to one or more amino acids and / or peptide molecules. In some cases, a power output of a laser light (e.g., an amount of energy the laser light emits per unit of time) can be at least or at most about 1 milliwatts (mW), at least or at most about 2 mW, at least or at most about 5 mW, at least or at most about 10 mW, at least or at most about 20 mW, at least or at most about 50 mW, at least or at most about 100 mW, at least or at most about 200 mW, or at least or at most about 500 mW. In some cases, a dimension (e.g., radius, diameter, circumference, etc.) of a crosssection of the laser light (e.g., a cross-section that is substantially parallel to a beaming direction of the laser light) may be at least or at most about 1 micrometer, at least or at most about 2 micrometers, at least or at most about 5 micrometers, at least or at most about 10 micrometers, at least or at most about 20 micrometers, at least or at most about 50 micrometers, at least or at most about 100 micrometers, at least or at most about 200 micrometers, or at least or at most about 500 micrometers. In some cases, a precision of the laser light can be controlled. The precision can be spectral precision as abovementioned (e.g., how narrowly defined the wavelength (color) of the laser is), directional precision (e.g., how much the laser beam spreads out over distance), spatial precision (e.g., how small a spot the laser can be focused to), temporal precision (e.g., how constant the power output is over time), etc. The precision of the laser light can be controlled over one or more axes or one or more degrees of freedom (e.g., x, y, z, pitch, yaw, and / or roll). The precision of the laser light over an axis (e.g., z-axis, where a plane defined by the z-axis is substantially perpendicular to the substrate) can be at most about ± 0.1 micrometers, at most about ± 0.2 micrometers, at most about ± 0.5 micrometers, at most about ± 1 micrometer, at most about ± 2 micrometers, at most about ± 5 micrometers, at most about ± 10 micrometers, or at most about ± 20 micrometers. Such high resolution printing may enhance high-resolution printing of the peptide array. For example, a laser source may direct a laser light at about 405 nanometers in wavelength, at between about 10 and about 50 mW power output, at less than or equal to about 100 micrometers beam diameter, and with a laser light precision of about ± 10 micrometers in X / Y-direction and about ± 2 micrometers in Z-direction.

[0064] In some embodiments, the sensor can comprise a detector configured to detect at least a portion of the light that is reflected from the another surface, to measure the metal binding performance. The detector can be configured to detect infrared light, visible light, and / or ultraviolet light. The detector can be configured to detect an attribute associated with the at least a portion of the light that is reflected, such as a wavelength, a frequency, or an intensity. Theoptical source and the detection may or may not be on the same side of the substrate.

[0065] In some embodiments, the metal binding performance comprises association between the target metal and each of the plurality of candidate peptide molecules. The association can comprise a bond between a metal molecule and a peptide molecule to form a metal-peptide complex via one or more of: a magnetic attraction, a covalent bond, and an ionic bond. The metal binding performance can comprise data points corresponding to a concentration of the metal-peptide complex at a plurality of time points. The concentration of the metal-peptide complex can be measured via the sensor. The metal binding performance can comprise data points corresponding to a concentration of the metal molecule and / or a peptide molecule at a plurality of time points. The concentration of the metal molecule and the concentration of the peptide molecule can be measured via the sensor or calculated based on the concentration of the metal-peptide complex, a known initial metal molecule concentration, and a known initial peptide molecule concentration. The association can be represented as an association constant (Ka) indicative of a strength of the association. A value of the association constant can calculated by dividing the concentration of the metal-peptide complex by the concentration of the metal molecule and the concentration of the peptide molecule, at equilibrium. A value of the association constant can be equal to a rate of binding between the metal molecule and the peptide molecule divided by a rate of unbinding between the metal molecule and the peptide molecule.

[0066] In some embodiments, the metal binding performance comprises dissociation between the target metal and each of the plurality of candidate peptide molecules. The dissociation can be represented as a dissociation constant (Kd) indicative of a propensity of a metal-peptide complex to disassociate into a metal molecule and a peptide molecule. The dissociation constant can be a reciprocal of the association constant. A value of the dissociation constant can be calculated by multiplying a concentration of the metal molecule with a concentration of the peptide molecule and dividing by a concentration of the metal-peptide complex, at equilibrium. A value of the dissociation constant can be equal to a rate of unbinding between the metal molecule and the peptide molecule, divided by a rate of binding between the metal molecule and the peptide molecule.

[0067] In some embodiments, the metal binding performance comprises both the association and the dissociation information as provided herein. The association and the dissociation information may be utilized at equal weights or at different weights. The association constant Ka and the dissociation constant Kd can be simultaneously monitored.

[0068] In some embodiments, the metal binding performance comprises the association and the dissociation, wherein the association is assigned a higher weight value than the dissociation.In some cases, the weight value assigned to the association can be greater than, substantially equal to, or less than the weight value assigned to the dissociation. The weight value of the association and / or the dissociation can be pre-determined. Alternatively, or in addition to, the weight value of the association and / or the dissociation can be adjusted based on the mining site, the geological sample, the target metal, and / or the plurality of candidate peptide molecules. Assigning the higher weight value to the association can allow a strong initial binding between the target metal and a candidate peptide molecule to be prioritized when assessing metal binding performance. Assigning a moderate weight value to the dissociation can maintain a moderate rate of unbinding between the target metal and the candidate peptide molecule. A moderate rate of unbinding can improve a material regeneration efficiency.

[0069] In some embodiments, a ratio of the assigned weight values to the association and the dissociation is between about 60:40 and about 90: 10, between about 60:40 and about 80:20, or about 70:30. The ratio of the assigned weight values can be about 60:40, about 65:35, about 70:30, about 75:25, about 80:20, about 85: 15, or about 90: 10. Out of a total weight scale of 100 that is distributed between different elements comprising the association (e.g., Ka) and the dissociation (e.g., Kd), the weight assigned to the association at be at least or at most about 10, at least or at most about 20, at least or at most about 30, at least or at most about 40, at least or at most about 50, at least or at most about 55, at least or at most about 60, at least or at most about 65, at least or at most about 70, at least or at most about 75, at least or at most about 80, at least or at most about 85, at least or at most about 90, at least or at most about 95, or at least or at most about 99.

[0070] In some embodiments, the system comprises a processor configured to execute an algorithm to identify the at least one optimal peptide molecule from the plurality of candidate peptide molecules based at least in part on the metal binding performance. The at least one optimal peptide molecule can comprise at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 25, at least or at most about 50, or at least or at most about 100 peptide molecule(s). The algorithm can identify the at least one optimal peptide molecule by identifying the at least one peptide molecule having a highest or a lowest metal binding performance. The at least one optimal peptide molecule can comprise at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 25, at least or at most about 50, or at least or at most about 100 optimal peptide molecule(s) (e.g., that are differentfrom one another). In some embodiments, the algorithm can select the at least one optimal peptide molecules based on the association constant Ka and / or the dissociation constant Kd. The at least one optimal peptide molecules can be selected to have a high Ka value and / or a moderate Kd value. The algorithm can comprise instructions to determine one or more additional properties associated with each of the plurality of candidate peptide molecules. The one or more additional properties can comprise one or more of: a metal binding performance for a metal different than the target metal, a thermal stability, and a chemical stability. The algorithm can be configured to predict the one or more additional properties. The algorithm can identify the at least one optimal peptide molecule based at least in part on the one or more additional properties. A same weight can be given to the metal binding performance as the one or more additional properties when identifying the at least one optimal peptide molecule. A higher or a lower weight can be associated with the metal binding performance, compared to the one or more additional properties. The algorithm can comprise instructions to determine a weight associated with the metal binding performance and / or the one or more additional properties.

[0071] In some embodiments, the processor is configured to execute the algorithm or another algorithm to rank the plurality of candidate peptide molecules based at least in part on their respective metal binding performance. The plurality of candidate peptide molecules can be ranked based on each peptide molecule’s association and / or dissociation with the target metal. The algorithm can rank the plurality of candidate peptide molecules from highest to lowest or lowest to highest based on the metal binding performance. The plurality of candidate peptide molecules can be ranked based on the association constant and / or the dissociation constant. The plurality of candidate peptide molecules can be ranked based on the association constant having a high value and / or the dissociation constant having a moderate value. The algorithm or the another algorithm can be configured to rank the plurality of candidate peptide molecules based at least in part on one or more additional properties, such as a metal binding performance for the target metal, a metal binding performance for a metal different than the target metal, a thermal stability, and a chemical stability. The metal binding performance and the one or more additional properties can have a same weight when ranking the plurality of candidate peptide molecules. The metal binding performance can have a higher weight or a lower weight compared to the one or more additional properties. The algorithm can comprise instructions to determine a weight associated with the metal binding performance and / or the one or more additional properties. In some embodiments, the processor is configured execute the algorithm or another algorithm to select the at least one optimal peptide molecule from the top about 50%, about 40%, about 30%, about 25%, about 20%, about 15%, about 10%, about 5%, about 2%, or about 1% of the rankedcandidate peptide molecules.

[0072] In some embodiments, the processor is configured execute the algorithm or another algorithm to design in silico at least one peptide variant comprising at least one amino acid modification as compared to the amino acid sequence of the at least one optimal peptide molecule. The algorithm or the another algorithm can comprise determining one or more peptide properties and / or one or more sequence patterns associated with high metal binding performance. The high metal binding performance can be evaluated based on the association constant and / or the dissociation constant. The high metal binding performance can be determined based on a high value of the association constant and / or a moderate value of the dissociation constant. The one or more peptide properties can comprise: a length, a size, a subsequence, a secondary structure, and / or a charge. The one or more sequence patterns can comprise: a specific amino acid at one or more positions, one or more amino acid having a specific property at one or more positions, and / or an amino acid subsequence. The amino acids having the specific property can comprise polar amino acids, nonpolar amino acids, positively charged amino acids, negatively charged amino acids, aromatic amino acids, nonaromatic amino acids, amino acids having a large side chain, and / or amino acids having a small side chain. The one or more peptide properties and / or the one or more sequence patterns associated with high metal binding performance can be determined by analyzing peptide properties and / or sequence patterns of one or more of the plurality of candidate peptide molecules selected based at least in part on their respective metal binding performance. For example, the one or more of the plurality of candidate peptide molecules can comprise highly ranked candidate peptide molecules and / or candidate peptide molecules identified to be optimal. The peptide variant can be designed in silico to comprise some or all of the one or more peptide properties and / or the one or more sequence patterns associated with high metal binding performance.

[0073] The at least one peptide variant can comprise at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 25, at least or at most about 50, or at least or at most about 100 peptide variants. Each of the at least one peptide variant can comprise a mutation at an amino acid position in a sequential order. For example, the at least one peptide variant can comprise 25 peptide variants each having a length of 20 amino acids, wherein each of the 20 peptide variants comprises a mutation at a different amino acid position. The at least one amino acid modification can comprise an addition, a deletion, and / or a substitution. The at least one amino acid modification can comprise a modification of at least one naturally occurring amino acid and / or at least onesynthetic amino acid.

[0074] In some embodiments, two peptide variants, as provided herein, can comprise at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 6, at least or at most about 7, at least or at most about 8, at least or at most about 9, at least or at most about 10, at least or at most about 11, at least or at most about 12, at least or at most about 13, at least or at most about 14, or at least or at most about 15 different amino acid residues, when compared to one another. In some embodiments, two peptide variants, as provided herein, can comprise a stereoisomeric variant (e.g., L- and D- stereoisomers) of the same amino acid residue (e.g., at the same amino acid residue site), when compared to one another.

[0075] In some embodiments, the in silico design of the at least one peptide variant is based at least in part on a site-directed mutagenesis. The site-directed mutagenesis can comprise sequentially providing a substitution at a single amino acid position of a peptide molecule. The site-directed mutagenesis can comprise in silico alanine scanning, glycine scanning, and / or proline scanning. Alanine scanning, glycine scanning, and proline scanning can comprise substituting an amino acid of the peptide molecule with alanine, glycine, and proline respectively. Alanine scanning, glycine scanning, and proline scanning can provide information about a contribution of the substituted amino acid to a structure, stability, and / or function of the peptide molecule. The site-directed mutagenesis can comprise a charge alteration. The charge alteration can comprise substituting a charged amino acid with an amino acid with the opposite charge, substituting a charged amino acid with an uncharged amino acid, and / or substituting an uncharged amino acid with a charged amino acid. In some cases, the charge alteration can provide information about the importance of the charge of the substituted amino acid. Additional non-limiting examples of site-directed mutagenesis includes saturation mutagenesis, deep mutational scanning (DMS), rational backbone redesign (e.g., Rosetta), and AlphaF old-guided site selection.

[0076] In some embodiments, the at least one amino acid modification comprises at least 2, at least 3, at least 4, or at least 5 amino acid modifications. The at least one amino acid amino acid modification can comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 50, or at least about 100 amino acid modification(s). In some embodiments, the at least one amino acid modification comprises at most 5, at most 4, at most 3, at most 2, or at most 1 amino acid modification(s). The at least one amino acid amino acid modification can comprise at most about 100, at most about 50, at most about 25, at most about 20, at most about 15, at most about10, at most about 5, at most about 4, at most about 3, at most about 2, or at most about 1 amino acid modification(s).

[0077] In some embodiments, the processor is configured to execute the algorithm or another algorithm to (i) determine a rate and / or yield of extraction of the target metal of one or more candidate peptide molecules and / or (ii) identify the at least one optimal peptide molecule based on the determined rate and / or yield of extraction. The rate of extraction can be determined by measuring the metal binding performance at a plurality of time points for the one or more candidate peptide molecules, and processing the metal binding performance measurements to determine a rate of binding for each of the one or more candidate peptide molecules. The rate of extraction can be determined based at least in part on the association constant Ka and / or the dissociation constant Kd. The yield of extraction can be determined by processing the metal binding performance to determine a number of target metal molecules bound to each the one or more candidate peptide molecules. Determining the yield of extraction can further comprise comparing the number of bound target metal molecules to a total number of target metal molecules in the medium. The algorithm or the another algorithm can be configured to identify the at least one optimal peptide molecule by determining the at least one peptide molecule having a highest metal binding performance, a highest rate of extraction, a highest yield of extraction, or any combination thereof. The algorithm or the another algorithm can be configured to assign each of the one or more candidate peptide molecules a score based on the metal binding performance, the rate of extraction, and / or the yield of extraction. The score can be assigned by one or more machine learning algorithms. The metal binding performance, the rate of extraction, and the yield of extraction can have the same weight when determining the score. The metal binding performance, the rate of extraction, and the yield of extraction can have different weights, the different weights determined by the algorithm or the another algorithm. The algorithm or the another algorithm can identify the at least one optimal peptide molecule by determining the at least one peptide molecule having the highest score.

[0078] In some embodiments, the processor is configured to execute a machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to identify the at least one optimal peptide molecule. The machine learning algorithm can be pre-trained on a plurality of amino acid sequences and corresponding metal binding performance labels. The machine learning algorithm can be fine-tuned on the collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules. The machine learning algorithm can be configured to predict metal binding performances for one or more of the respective amino acidsequences. The machine learning algorithm can be configured to identify the at least one optimal peptide molecule, based at least in part on the at least one optimal peptide molecule having a high predicted metal binding performance. The high metal binding performance can be determined based on the association constant and / or the dissociation constant. The high metal binding performance can correspond to a high value of the association constant and / or a moderate value of the dissociation constant.

[0079] In some embodiments, the processor is configured to execute the machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to design in silico the at least one peptide variant. The machine learning algorithm can be pre-trained on a plurality of training sequences. The plurality of training sequences can comprise a plurality of amino sequences having a high and / or low collected and / or predicted metal binding performance. The high metal binding performance can be determined based on the association constant and / or the dissociation constant. The high metal binding performance can correspond to a high value of the association constant and / or a moderate value of the dissociation constant. The machine learning can be trained and / or finetuned on one or more of the respective amino acid sequences of the candidate peptide molecules. The machine learning algorithm can be configured to identify a first plurality of respective amino acid sequences having a high metal binding performance. Alternatively, or in addition to, machine learning algorithm can be configured to identify a second plurality of respective amino acid sequences having low metal binding performance. The machine learning algorithm can analyze the first plurality and / or the second plurality of respective amino acid sequences to identify one or more peptide properties and / or one or more sequence patterns associated with high metal binding performance. For example, the machine learning algorithm can compare the first plurality and the second plurality of respective amino acids to identify one or more peptide properties and / or one or more sequence patterns appearing in the first plurality of respective amino acids and / or not appearing in the second plurality of respective amino acids. The one or more peptide properties can comprise: a length, a size, a subsequence, a secondary structure, and / or a charge. The one or more sequence patterns can comprise: a specific amino acid at one or more positions, one or more amino acid having a specific property at one or more positions, and / or an amino acid subsequence. The amino acids having the specific property can comprise polar amino acids, nonpolar amino acids, positively charged amino acids, negatively charged amino acids, aromatic amino acids, nonaromatic amino acids, amino acids having a large side chain, amino acids having a small side chain, or any combination thereof. The machine learning model can design in silico the at least one peptide variant based on the at least one peptide varianthaving the one or more peptide properties and / or one or more sequence patterns associated with high performance. The machine learning algorithm can be configured to predict metal binding performances for one or more designed peptide variants. The predicted metal binding performances can correspond to a predicted association constant and / or a predicted dissociation constant. The machine learning algorithm can design in silico the at least one peptide variant based on the at least one peptide variant having a high predicted metal binding performance.

[0080] In some embodiments, the machine learning algorithm is trained to (1) generate a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule, (2) utilize graph neural network (GNN) to analyze the plurality of kinetic graph embeddings and identify one or more amino acid residues or positions thereof that are associated with the association and dissociation between the target metal and a candidate peptide molecule, and (3) select the identified one or more amino acid residues or positions thereof for mutation to generate the at least one peptide variant. Identifying the one or more amino acid residues or positions can comprise identifying one or more amino acid residues or positions having a highest contribution to a binding energy. The highest contribution to the binding energy can be assessed based on an association / dissociation curve sensitivity analysis and / or an annotation. The annotation can be provided by one or more databases and / or one or more protein modeling software modules. Nonlimiting examples of such software modules can include AlphaFold, RoseTTAFold, ESMFold, OmegaFold, RaptorX, ProSPr, OpenFold, FoldDock, etc. Identifying the one or more amino acid residues or positions can comprise avoiding one or more amino acid positions having a stabilizing effect (e.g., a cysteine residue in a loop of a peptide molecule; a proline residue that provides a conformational constraint; a glycine residue that allows a tight turn or compact folding; tyrosine, tryptophan, and / or phenylalanine that provides TI- TI stacking and / or hydrophobic packing, etc.). The at least one peptide variant can be generated based on a mutation library builder. Non-limiting examples of the mutation library builder can include Rosetta library builder, PyRosetta, RosettaScripts, Protein Message Passing Neutral Network (ProteinMPNN), DeepScan, DeepMutScan, FoldX, EvoDesign, HotSpot Wizard, FireProt, MutateX, dTERMen, etc.

[0081] In some embodiments, the machine learning algorithm is configured to analyze the association and / or dissociation between the target metal and each of the plurality of candidate peptide molecules, to identify a pattern of a molecular interaction between the target metal and the plurality of candidate peptide molecules. The machine learning algorithm can be configured to predict an association or dissociation each of the plurality of candidate peptide molecules. Themachine learning algorithm can be pre-trained on a training dataset comprising one or more peptide molecules and corresponding association and / or dissociation. The machine learning algorithm can be fine-tuned on the association and / or dissociation between the target metal and each of the plurality of candidate peptide molecules, to identify the pattern of a molecular interaction between the target metal and the plurality of candidate. The machine learning algorithm can be configured to determine one or more quantitative values associated with the association and / or the dissociation, including an association constant, a dissociation constant, a rate of binding, a rate of unbinding, and / or a yield of extraction. The machine learning can comprise a kinetic model configured to fit one or more curves to the metal binding performance data. The kinetic model can determine a rate of binding and / or a rate of unbinding based on a slope of the curve at an initial time point. The kinetic model can be configured to determine the association constant and / or the dissociation constant based on the rate of binding and the rate of unbinding. The machine learning algorithm can be configured to determine the concentrations of the metal-peptide complex, the metal molecule, and the peptide molecule at equilibrium. The machine learning algorithm can be configured to determine the association constant and / or the dissociation constant based on the determined concentrations of the metal-peptide complex, the metal molecule, and the peptide molecule. The machine learning algorithm can be configured to determine the association constant and the dissociation constant simultaneously.

[0082] In some embodiments, the machine learning algorithm is trained to identify the pattern via (1) generation of a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule and (2) comparison of the plurality of kinetic graph embeddings. The plurality of kinetic graph embeddings can comprise a lower-dimensional representation of an input data. The input data can comprise the amino acid sequences of the candidate peptide molecules and the respective SPR curves. Generating the kinetic graph embeddings can comprise generating a knowledge graph. The knowledge graph can comprise a plurality of nodes. The plurality of nodes can represent amino acids, amino acid sequences and / or target metals. The knowledge graph can comprise a plurality of edges, each edge comprising a connection and / or interaction between two nodes of the plurality of nodes. The plurality of kinetic graph embeddings can represent one or more node properties and / or one or more edges of the knowledge graph. The plurality of kinetic graph embeddings can comprise a vector representation of the knowledge graph. The plurality of kinetic graph embeddings can represent relationships between amino acid sequences and association and / or dissociation with the target metal. Comparing the plurality of kinetic graph embeddings can comprise comparing one or more nodes and / or one or more edgesof a knowledge graph. Comparing the plurality of kinetic graph embeddings can comprise comparing two or more knowledge graphs.

[0083] In some cases, the machine learning algorithm as provided herein can generate a graph embedding representing the metal-peptide molecule binding interaction (e.g., the SPR data obtained by the sensor). In the graph embedding, for example, a node can represent one or more amino acid residues, an edge can represent one or more sequential (e.g., the order or sequence of amino acid residues in a peptide molecule) and / or energetic links (e.g., non-sequential interactions within a peptide molecule or between two or more peptide molecules, such as hydrogen bonds, van der Waals forces, electrostatic interactions, etc.). The graph embedding can be processed via one or more GNNs to model association and / or dissociation.

[0084] In some embodiments, the machine learning algorithm is further trained to identify the pattern via (2) clustering of two or more kinetic graph embeddings from the plurality of kinetic graph embeddings based on similarities and / or differences between the plurality of kinetic graph embeddings and / or (3) identification of one or more kinetic graph embeddings exhibiting an anomaly. The clustering can be an unsupervised clustering. The clustering can comprise time-series clustering. The clustering can comprise assigning each of the two or more kinetic graph embeddings to one of two or more clusters. The clustering can comprise minimizing a difference between kinetic graph embeddings assigned to a same cluster. The clustering can comprise maximizing a difference between kinetic graph embeddings assigned to different clusters. The identification of the anomaly can comprise flagging one or more amino acid sequences having an unexpected affinity profile. The unexpected affinity profile can be determined by comparing a predicted affinity profile and an experimental affinity profile. The unexpected affinity profile can be determined based on the clustering. The unexpected affinity profile can be determined based on a kinetic graph embedding of the amino acid sequence having a large difference from each of the two or more clusters. The unexpected affinity profile can be determined based on the kinetic graph embedding of the amino acid sequence being assigned to a different cluster than a predicted cluster. In some cases, outliers can be identified through clustering in embedding space, deviation from predicted kinetic profiles, and use of unsupervised methods (e.g., Isolation Forests and / or autoencoders).

[0085] In some embodiments, the machine learning algorithm is trained based at least in part on supervised learning using a labeled dataset of known interactions between the target metal and the plurality of candidate peptide molecules or other control peptide molecules. The machine learning algorithm can be trained based on input data and associated labels. The input data can comprise amino acid sequences of the plurality of candidate peptide molecules or other controlpeptide molecules. The input data can further comprise peptide properties of the plurality of candidate peptide molecules or other control peptide molecules, such as a length, a size, a subsequence, a secondary structure, or a charge. The associated labels can be indicative of the interaction between the target metal and the input peptide sequence, such as metal binding performance, association, or dissociation. Training the machine learning algorithm can comprise partitioning the input data and associated labels into a training dataset and a validation dataset. One or more weights of the machine learning algorithm can be optimized based on the training dataset, and a performance of the machine learning algorithm can be evaluated based on the validation dataset. One or more parameters of the machine learning algorithm can be updated based on the performance of the machine learning algorithm on the validation dataset. The one or more parameters can comprise: a model type, a model size, a batch size, a cross validation size, an activation, a loss function, an optimizer, a regularization, a learning rate, a number of iterations, and a kernel. The machine learning algorithm can be configured to output a predicted label based on an input peptide sequence. The machine learning algorithm can comprise a classification algorithm configured to output a discrete label indicative of the interaction between the target metal and the input peptide sequence. The machine learning algorithm can comprise a regression algorithm configured to output a continuous label indicative of a strength and / or a likelihood of the interaction between the target metal and the input peptide sequence.

[0086] In some embodiments, the machine learning algorithm is trained based at least in part on unsupervised learning to detect a hidden or unknown trend from the interactions between the target metal and the plurality of candidate peptide molecules. The machine learning algorithm can be trained based on input data without associated labels. The machine learning algorithm can comprise a clustering algorithm configured to identify one or more groups of similar peptide molecules based on peptide sequences and / or additional peptide properties as provided herein (e.g., a length, a size, a subsequence, a secondary structure, or a charge). The machine learning algorithm can detect the hidden or unknown trend by evaluating members of each of the one or more groups and determining shared peptide properties. The machine learning algorithm can be configured to detect the hidden or unknown trend by detecting two or more correlated peptide properties. The machine learning algorithm can comprise a dimensionality reduction algorithm configured to transform the data into a lower-dimensional representation. The dimensionality reduction algorithm can detect the hidden or unknown trend by identifying one or more peptide properties having a largest importance on a variance of the input data.

[0087] In some embodiments, the machine learning algorithm comprises reinforcement learning to dynamically adjust one or more performance characteristics in the analysis of thecollection of the metal binding performance, the identification of the at least one optimal peptide molecule, and / or the in silico design of the at least one peptide variant. The reinforcement learning can comprise providing one or more rewards or one or more penalties to the machine learning algorithm, based on a model performance. The model performance can be determined based on experimental data. For example, the at least one optimal peptide molecule and / or the at least one peptide variant designed in silico can be produced and corresponding metal binding performance can be obtained. If the corresponding metal binding performance for the optimal peptide and / or the least one optimal peptide molecule and / or the at least one peptide variant designed in silico is high, a reward can be provided to the machine learning algorithm, otherwise, a penalty can be provided. The high metal binding performance can comprise a high association constant and / or a moderate dissociation constant. The model performance can be determined based on a human supervisor. For example, the human supervisor can analyze the collection of the metal binding performance, and the results of the human supervisor analysis can be compared to the machine learning algorithm analysis. If the results of the human supervisor analysis and the machine learning algorithm analysis are similar, a reward can be provided to the machine learning algorithm, otherwise, a penalty can be provided. The model performance can be determined based on the machine learning algorithm. For example, the at least one peptide variant designed in silico by a first module of the machine learning algorithm can be an input to a second module of the machine learning algorithm to predict a metal binding performance. If the metal binding performance is predicted to be high, a reward can be provided to the first module, otherwise, a penalty can be provided. The machine learning algorithm can be trained continuously via a feedback loop. The machine learning algorithm can update one or more parameters, determine an effect of the update based on the received reward or penalty and adjust the one or more parameters accordingly. For example, the reinforcement learning can comprise increasing a weight, predicting an optimal peptide molecule, receiving a reward or penalty based on the evaluated performance, and further increasing the weight if the reward is received, or decreasing the weight if the penalty is received.

[0088] In some embodiments, the machine learning algorithm is configured to compare (i) the measured metal binding performance by the sensor and (ii) an in silico predicted metal binding performance, to improve its capability in predicting a metal binding performance of a candidate peptide molecule sequence. The machine learning algorithm can be configured to output the predicted metal binding performance for one or more candidate peptide molecules, then access the measured metal binding performance. The machine learning algorithm can determine a similarity or a difference between the measured metal binding performance and thepredicted metal binding performance. If a difference is determined, the machine learning algorithm can adjust one or more weights and / or one or more parameters based on the measured metal binding performance. The machine learning algorithm can comprise two or more models each having different weights and / or parameters. A model of the two or more models can be selected based on the model having a more similar predicted metal binding performance to the measured metal binding performance. The machine learning algorithm can be configured to continuously learn from measured data.

[0089] In some embodiments, the machine learning algorithm is part of a self-learning codebase that is configured to improve its performance in (i) the identification of the at least one optimal peptide molecule and / or (ii) the design of the at least one peptide variant. The codebase can store one or more measure metal binding performance and one or more corresponding peptide sequences. The codebase can store the machine learning algorithm. The codebase can store weights and / or parameters of the machine learning algorithm. The codebase can store instructions for training the machine learning algorithm. The codebase can be stored locally, on one or more servers, or on one or more cloud platforms. The codebase can be configured to automatically store the measured metal binding performance and corresponding amino acid sequences for the candidate peptide molecules. The codebase can be configured to execute the machine learning algorithm to i) identify the at least one optimal peptide molecule and / or ii) design the at least one peptide variant based on the stored metal binding performances and corresponding peptide sequences. The codebase can be configured to receive and store the measured metal binding performance of the identified at least one optimal peptide molecule and / or the at least one designed peptide variant. The codebase can be configured to update the weights and / or the parameters of the machine learning algorithm based on the received measured metal binding performance, thereby allowing the machine learning algorithm to self-leam.

[0090] The machine learning algorithm as provided herein can be utilized for detection of peptide-metal binding, peptide array printing, or both. Any operation (e.g., analysis of data, generation of one or more new candidate peptide sequences) of the machine learning algorithm can be performed substantially in real-time as the detection of the peptide-metal binding via the sensor. Alternatively, or in addition to, such operation can be performed at a subsequent time, e.g., prior to or during the process of peptide array printing.Peptide array printing

[0091] Certain aspects of the present disclosure provide systems and methods for printing a patterned array of peptide molecules for biomining a target metal from a medium (e.g., solution)obtained from a geological sample.

[0092] In an aspect, the present disclosure provides a method for printing a patterned array of peptide molecules for biomining a target metal from a medium (e.g., solution) obtained from a geological sample. The method can comprise providing at least one amino acid or at least one peptide. The method can further comprise directing, via an optical source of a printer, a light towards a surface of a substrate. The light can comprise a single wavelength or a range of wavelengths sufficient to effect conjugation of the at least one amino acid or the at least one peptide on or adjacent to the surface of the substrate. The patterned array of peptide molecules can thereby be printed on the surface. The patterned array of peptide molecules can comprise a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

[0093] In another aspect, the present disclosure provides a system for implementing any one of the methods provided herein. The system can be for printing a patterned array of peptide molecules for biomining a target metal from a medium (e.g., solution) obtained from a geological sample. The system can comprise a printer. The printer can comprise an optical source configured to direct a light towards a surface of a substrate. The light can comprise a single wavelength or a range of wavelengths sufficient to effect conjugation of at least one amino acid or at least one peptide one or adjacent to the surface of the substrate. The system can further comprise a controller configured to instruct the optical source to direct the light towards the surface to print, via the conjugation, an array of peptide molecules on the surface. The array of peptide molecules can comprise a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

[0094] FIG. 7 shows an example method 700 for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample. The method 700 can optionally comprise a step 702 of obtaining a medium from a geological sample. The method 700 can comprise a step 704 of providing at least one amino acid or at least one peptide. The method 700 can comprise a step 706 of directing a light towards a surface of a substrate, wherein the light is sufficient to print a patterned array of peptide molecules on the surface. The method 700 can comprise an optional step 708 of binding and extracting a target metal via the patterned array of peptide molecules.

[0095] In some embodiments, the providing, the directing, and / or the printing is performed at or near a same mining site as the geological sample. The same mining site can comprise a same mine, a same region of the mine, a same depth of the mine, a same excavation site, a same drilling site, a same geographical region, or a same country. Performing the providing, thedirecting, and / or the printing at or near the same mining site as the geographical sample can allow patterned arrays of peptide molecules to be printed with faster turn-around times, compared to receiving the patterned arrays from an external site the sample to an external site for testing, which can improve time and cost efficiency. Performing the providing, the directing, and / or the printing at or near the same mining site can also allow the patterned arrays to be adapted to be continuously adapted to current geological conditions (e.g., temperature, humidity, pH) of the mining site in real time. For example, as a depth of a mining site changes during excavation, the geological condition can change correspondingly. Performing the providing, the directing, and / or the printing at or near the mining site can allow the patterned array of peptide molecules to be printed based on the most current geological conditions, which can further improve mining efficiency. Alternatively, or in addition to, the providing, the directing, and / or the printing can be performed not at the same mining site as the geological sample. The providing, the directing, and / or the printing can be performed at a different mining site, such as a different mine, a different region of the same mine, a different excavation site, a different drilling site, a different geographical region, or a different country. The providing, the directing, and / or the printing can be performed at a non-mining site, such as a surveying area, a planning area, a research area, or a laboratory.

[0096] In some embodiments, the printer and the sensor are in a same unit and / or housing. The patterned array of peptide molecules can positioned at a same location during the printing and the measuring of the metal binding performance. Alternatively, or in addition to, the printer and the sensor can be in different units and / or housings. The patterned array of peptide molecules can be positioned at a first location for the printing, and moved to a second location for the measuring of the metal binding performance. The moving of the patterned array of peptide molecules can be performed manually and / or by an automated process.

[0097] In some embodiments, the wavelength or the range of wavelengths of the light is sufficient to activate the surface or an amino acid bound to the surface for the conjugation. The wavelength or the range of wavelengths can correspond to infrared light, visible light, and / or ultraviolet light. The ultraviolet light can have a wavelength of about 10 nm to about 100 nm, about 100 nm to about 280 nm, about 280 nm to about 315 nm, and / or about 315 nm to about 400 nm. The visible light can have a wavelength of about 380 nm to about 450 nm, about 450 nm to about 500 nm, about 500 nm to about 565 nm, about 565 to about 590 nm, about 590 to about 625 nm, and about 625 to about 780 nm. The infrared light can have a wavelength of about 750 nm to about 1.4 pm, about 1.4 pm to about 3 pm, and / or about 3 pm to about 1 mm. In some cases, the wavelength or the range of wavelengths can be between about 350 nanometers (nm)and about 550 nm or between about 400 nm and about 550 nm. The wavelength or the range of wavelengths can be at least or at most about 300 nm, at least or at most about 350 nm, at least or at most about 400 nm, at least or at most about 450 nm, at least or at most about 500 nm, at least or at most about 550 nm, or at least or at most about 600 nm.

[0098] In some embodiments, the surface and / or the amino acid comprises a photo-labile protecting group (PPG), wherein the light cleaves the PPG to activate the surface and / or the amino acid. The surface can be pre-coated with a stable linker comprising the PPG. The PPG can allow site- selective conjugation. The PPG can be cleaved selectively at a wavelength or a range of wavelengths of light. The PPG can protect a carboxyl at the surface or the amino acid. The PPG can prevent a bond formation between the protected carboxyl group and one or more amine groups. When the PPG is cleaved, a peptide bond can form between the carboxyl group and the one or more amine groups. Alternatively, or in addition to, the PPG can protect an amine group at the surface or the amino acid and prevent a bond formation between the protected amine group and one or more carboxyl groups. When the PPG is cleaved, a peptide bond can form between the amine group and the one or more carboxyl groups. The PPG can allow for Merrifield synthesis of the candidate peptide molecules.

[0099] In some embodiments, the PPG can comprise nitroveratryloxy carbonyl. In some embodiments, the PPG can comprise an o-nitrobenzyl derivative (e.g., nitroveratryloxycarbonyl or “NVOC”, 2-(2-Nitrophenyl)-propyloxycarbonyl or “NPPOC”, 2, 3 -dimethyl -2, 3 -dinitrobutane or “DMNB”, etc.), p-hydroxyphenacyl (pHP), a coumarin-4-ylmethyl group, a carbonyl group, a benzyl group, an arylcarbonylmethyl group, a phenacyl group, an alkylphenacyl group, a hydroxyphenacyl group, a benzoin group, a nitroaryl group, a nitrophenyloxycarbonyl group, a nitroanilide group, an arylmethyl group, a hydroxyarylmethyl group, and / or a metal-containing group.

[0100] In some embodiments, the PPG can comprise a photolabile protecting group that is activatable by a green light (e.g., comprising a wavelength of about 510 nanometers). In some cases, the PPG can comprise a boron-dipyrromethene (BODIPY)-based PPG. The light absorption properties (e.g., at or around 510 nanometers in wavelength) of the BODIPY derivatives can be finely tuned through structural modifications. The BODIPY derivatives can used for controlled release of one or more functional groups, such as, for example, alcohols and carboxylic acids. The high molar absorptivity and photostability of the BODIPY derivatives may provide precise spatial and / or temporal control over the surface for generating a patterned array of peptide molecules.

[0101] In some embodiments, the PPG can comprise a xanthene-derived PPG, such as thosederived from fluorescein. The xanthene-derived PPG can exhibit strong absorption in the visible spectrum, including the green region as mentioned above. The xanthene-derived PPGs can be utilized for photo-activated release of bioactive molecules.

[0102] In some embodiments, the PPG can comprise a Cyanine-derived PPG. Cyanine dyes can exhibit intense absorption in the visible to near-infrared region. By incorporating photolabile linkers into cyanine structures, the resulting PPG can be activated by green light. In some cases, the Cyanine-derived PPG can be compatible with aqueous environment.

[0103] In some embodiments, the light comprises a laser light. Alternatively, or in addition to, the light can comprise a light-emitting diode (LED) light, an incandescent light, and / or a fluorescent light.

[0104] In some embodiments, the laser light has a wavelength range between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm. The laser light can have a wavelength of at least or at most about 300 nm, at least or at most about 350 nm, at least or at most about 400 nm, at least or at most about 450 nm, at least or at most about 500 nm, at least or at most about 550 nm, or at least or at most about 600 nm. The laser light can comprise a 405 nm and / or a 532 nm laser. The wavelength or wavelength range of the laser light can be determined based on a linker group coupled to one or more amino acids and / or peptide molecules.

[0105] In some embodiments, a target region of the light or the laser light within the surface of the substrate is controlled, thereby selectively effecting the conjugation within the target region. The target region can comprise a region of a pattern of the patterned array. The target region can correspond to one molecule of the plurality of candidate peptide molecules. Alternatively, or in addition to, the target region can correspond to two or more molecules of the plurality of candidate peptide molecules. The target region can comprise a square, a rectangular, a circular, a triangular, a polygonal, and / or an irregular shape. The target region can have a diameter and / or a side length of at least or at most about 0.05 mm, least or at most about 0.1 mm, least or at most about 0.25 mm, least or at most about 0.5 mm, least or at most about 1 mm, least or at most about 1.5 mm, least or at most about 2 mm, least or at most about 3 mm, least or at most about 4 mm, least or at most about 5 mm, least or at most about 10 mm. The target region can be determined by one or more processors.

[0106] In some embodiments, an optical pick-up unit (OPU) is instructed to control the target region of the light or the laser light. The OPU can be moved manually to control the target region. The movement of the OPU can comprise a rotation and / or a translation. The OPU can control the target region by changing an angle of the light or the laser light. The OPU can beinstructed to control the target region by one or more processors. The one or more processors can provide one or more spatial coordinates indicative of the target region to the OPU. The OPU can comprise two or more light or laser light emitting units. The OPU can selectively activate and / or deactivate the two or more light or laser light emitting units to control the target region. The target region can be controlled in a x axis, a y axis, and / or a z axis.

[0107] In some embodiments, the surface is contacted with a medium comprising at least one amino acid or the at least one peptide (e.g., for printing the patterned array of peptide molecules on the surface). The medium can comprise a mixture, a medium, a suspension, and / or a substrate. A composition of the medium can comprise a solvent. The solvent can be an organic solvent and / or an inorganic solvent. The solvent may or may not comprise dimethyl sulfoxide (DMSO). The medium can be stored at one or more locations separate from the surface. The surface can be contacted with the medium selectively. Alternatively, or in addition to, the surface can be contacted with the medium constantly. The surface can be contacted by two or more media. Each of the two or more media can have a different composition. The surface can be contacted by the two or media simultaneously and / or at different times. Each of the two or more media can contact a same and / or a different region of the surface.

[0108] In some embodiments, the at least one amino acid or the at least one peptide can be protected by a protecting group. The protecting group can protect an amine group and / or a carboxyl group of the at least one amino acid or the at least one peptide. The protecting group can comprise a photo-labile protecting group as provided herein, an acid-labile protecting group, and / or a base-labile protecting group. The protecting group can comprise 9- fluorenylmethyloxy carbonyl (Fmoc).

[0109] In some embodiments, the system comprises a container for holding the medium comprising at least one amino acid or the at least one peptide. A material of the container can comprise a metal, a crystal, and / or a polymer. The container can comprise a sealed reservoir. Alternatively, the container may not be sealed. The container can comprise one or more openings. The medium can be added and / or removed from the container via the one or more openings. The container can be configured to adjust a volume of the medium added and / or removed from the container. For example, a size of the one or more openings can be increased and / or decreased to increase and / or decrease the volume. Alternatively, or in addition to, the volume of the medium added and / or removed can remain constant. The container can comprise two or more chambers holding two or more media. Each chamber can comprise one or more corresponding openings for the respective medium to be added and / or removed from the chamber. The system can comprise two or more containers. Each of the two or more containerscan hold a medium comprising a different amino acid and / or peptide. The two or more containers can comprise 20 containers corresponding to each of the 20 naturally occurring amino acids. The two or more containers can comprise more than 20 containers corresponding to the 20 natural occurring amino acid and one or more synthetic amino acids. The container can hold at least or at most about 0.1 milliliters (mL), at least or at most about 0.2 mL, at least or at most about 0.5 mL, at least or at most about 1 mL, at least or at most about 2 mL, at least or at most about 3 mL, at least or at most about 4 mL, at least or at most about 5 mL, at least or at most about 6 mL, at least or at most about 7 mL, at least or at most about 8 mL, at least or at most about 9 mL, at least or at most about 10 mL, at least or at most about 20 mL, or at least or at most about 50 mL.

[0110] In some embodiments, the flow of the medium is directed from a source of the medium and towards the surface. The source of the medium can be the container. The flow of the medium can be further directed to a second location. The second location can comprise a waste container. The flow can have a constant rate. Alternatively, or in addition to, the rate can be adjusted. The flow of the medium can be conducted via one or more channels. The one or more channels can be adjusted to adjust a rate and / or a direction of the flow. Each of the one or more channels can correspond to a different rate and / or direction of the flow. The flow can comprise a one-way flow. Alternatively, or in addition to, the flow can comprise a two-way flow and / or a reversible flow. The respective flows of two or more media can be directed. The two or more media can be directed from a same source and / or two or more different sources. The two or more media can be directed to a same and / or a different region of the surface. The two or more media can be directed simultaneously and / or at different times. The flow rate can be at least or at most about at least or at most about 0.1 microliters per min (pL / min), at least or at most about 0.2 pL / min, at least or at most about 0.5 pL / min, at least or at most about 1 pL / min, at least or at most about 2 pL / min, at least or at most about 5 pL / min, at least or at most about 10 pL / min, at least or at most about 20 pL / min, at least or at most about 50 pL / min, at least or at most about 100 pL / min, at least or at most about 200 pL / min, at least or at most about 500 pL / min, at least or at most about 1 milliliter per min (mL / min), at least or at most about 2 mL / min, at least or at most about 5 mL / min, or at least or at most about 10 mL / min (e.g., for bulk operations).[OHl] In some embodiments, the system comprises a flow controller configured to direct the flow of at least a portion of the medium from the source of the medium and towards the surface. The flow controller can comprise a microfluidic system. The flow controller can be an automated flow controller. The flow controller can be controlled by one or more processors.The flow controller can adjust a rate and / or direction of the flow. The flow controller can direct the flow to one or more defined areas of the surface. The one or more defined areas can comprise one or more printing zones. The flow controller can adjust the rate and / or direction by adjusting one or more channels conducting the flow. Each of the one or more channels can be adjusted to provide a different flow rate and / or direction. The flow controller can be further configured to direct the flow of at least a portion of the medium from the surface to a second location. The second location can comprise a waste container. The flow controller can direct the flow of at least a portion of each of two or more media. The flow controller can direct the flow of the two or more media simultaneously or at different times. The flow controller can direct the flow of the two or more media to a same and / or a different region of the surface.

[0112] In some embodiments, the plurality of candidate peptide molecules is selected and / or designed from a library of candidate peptide molecules based on a type of the geological sample. The type of the geological sample can comprise one or more of: a target metal, a mining site, a chemical composition, a physical property, a chemical property, a size, and / or a purity. The plurality of candidate peptide molecules can be selected and / or designed by identifying one or more optimal peptide sequences for extracting the target metal from the geological sample. The plurality of candidate peptide molecules can be selected and / or designed based on the one or more optimal peptide sequences by providing one or more amino acid modifications. The plurality of candidate peptide molecules can be selected and / or designed based on analyzing collected and / or predicted metal binding performances of one or more amino acid sequences with the type of geological sample. The plurality of candidate peptide molecules can be selected and / or designed based on analyzing collected and / or predicted metal binding performances of one or more amino acid sequences with one or more different types of geological samples. The one or more different types of geological samples can have a similarity with the type of the geological sample, such as a same or similar target metal, a same or similar chemical composition, a same or similar physical property, a same or similar a chemical property, a same or similar size, a same or similar mining site, and / or a same or similar purity. The plurality of candidate peptide molecules can be selected and / or designed based on identifying optimal peptide sequences for the one or more different types of geological samples. The plurality of candidate peptide molecules can be selected and / or designed by one or more algorithms and / or one or more machine learning algorithms.

[0113] In some embodiments, the plurality of candidate peptide molecules is selected or designed from a library of candidate peptide molecules based on a geolocation of the geological sample. The geolocation can comprise a type of mine, a mine, a region of a mine, a depth of amine, an excavation site, a drilling site, a geographical region, a geological condition, or a country. The plurality of candidate peptide molecules can be selected and / or designed by identifying one or more optimal peptide sequences for extracting the target metal from a geological sample obtained at the geolocation. The plurality of candidate peptide molecules can be selected and / or designed based on the one or more optimal peptide sequences by providing one or more amino acid modifications. The plurality of candidate peptide molecules can be selected and / or designed based on analyzing collected and / or predicted metal binding performances of one or more amino acid sequences with geological samples obtained at the geolocation. The plurality of candidate peptide molecules can be selected and / or designed based on analyzing collected and / or predicted metal binding performances of one or more amino acid sequences with geological samples obtained at one or more different geolocations. The one or more different geolocations can have a similarity with the geolocation, such as a same or similar type of mine, a same or similar mine, a same or similar region of a mine, a same or similar depth of a mine, an same or similar excavation site, a same or similar drilling site, a same or similar geographical region, a same or similar geological condition, a same or similar country, or same or similar geological condition. The plurality of candidate peptide molecules can be selected and / or designed based on identifying optimal peptide sequences for the one or more different geolocations. The plurality of candidate peptide molecules can be selected and / or designed by one or more algorithms and / or one or more machine learning algorithms.

[0114] In some embodiments, the plurality of candidate peptide molecules is designed in silico based on data associated with a metal binding performance of another patterned array of peptide molecules comprising another plurality of candidate peptide molecules, wherein the plurality of candidate peptide molecules and the another plurality of candidate peptide molecules are different. The plurality of candidate molecules can be designed in silico by identifying one or more optimal peptide molecules of the another plurality of candidate peptide molecules. The one or more optimal peptide molecules of the another plurality of candidate peptide molecules can be identified based on having a high metal binding performance. The high metal binding performance can correspond to a high association constant and / or a moderate dissociation constant. The plurality of candidate peptide molecules can be designed to each comprise one or more amino acid modifications to the one or more optimal peptide molecules. The plurality of candidate peptide molecules can each have a different modification, such as a modification at a different position in the amino acid sequence, or a modification at a same position to a different amino acid. The plurality of candidate molecules can be designed in silico by one or more algorithms and / or one or more machine learning algorithms. The one or more algorithms and / orthe one or more machine learning algorithms can be trained on the data associated with the metal binding performance of the another patterned array of peptide molecules. The one or more algorithms and / or the one or more machine learning algorithms can be trained to predict metal binding performances of one or more candidate peptide molecules. The one or more algorithms and / or the one or more machine learning algorithms can be trained to generate one or more amino acid sequences likely to have high metal binding performances. The plurality of candidate molecules can be designed in silico based on having high predicted metal binding performances and / or based on being generated by the one or more algorithms and / or one or more machine learning algorithms.

[0115] In some embodiments, the metal binding performance of the another patterned array of peptide molecules is based on binding and extracting the target metal from (i) the same medium, (ii) another medium obtained from the same geological sample, and / or (iii) another medium obtained from another geological sample from a same geological site as the geological sample. Alternatively, or in addition to, the metal binding performance of the another patterned array of peptide molecules is based on binding and extracting the target metal from (i) a different medium, (ii) a medium obtained from a different geological sample, and / or (iii) a medium obtained from a geological sample from a different geological site. The same geological sample can comprise a same ore, a same type of geological sample, a sample obtained from a same region of a mine, and / or a sample obtained form a same mine. The same geological site can comprise a same type of mine, a same mine, a same region of a mine, a same depth of a mine, a same excavation site, a same drilling site, a same geographical region, a same geological condition, or a same country. When selecting and / or designing the plurality of candidate peptide molecules, the metal binding performance of the another patterned array of peptide molecules can be weighted based on the source of the target metal. A metal binding performance based on the same medium can be weighted more heavily than a metal binding performance based on another medium. A metal binding performance based on a medium obtained from the same geological site can be weighted more heavily than a metal binding performance based on a medium obtained from a different geological site.

[0116] FIG. 8 schematically illustrates an example system for printing a patterned array of peptide molecules. The system can comprise an optical lithography system. The optical lithography system can comprise an optical source 802. The optical lithography system can direct a light to a photo-sensitive material via the optical source 802. The optical lithography system can thereby selectively print a sequence of amino acids at a specified position of the photo-sensitive material to generate a patterned array. For example, the patterned array cancomprise a matrix of dots. The matrix of dots can have a resolution of about 2 mm, about 1 mm, about 0.5 mm, and / or about 0.25 mm.Additional details

[0117] In some embodiments, the geological sample comprises ores, concentrates, and / or waste materials. The geological sample can comprise a pure substance. Alternatively, or in addition to, the geological sample can comprise a mixture. The mixture can comprise a mixture of an ore and / or concentrate and a waste material. The methods and systems for biomining can be implemented to isolate the ore and / or the concentrate from the waste material. The mixture can comprise a mixture of two or more types of ores and / or concentrates. The methods and systems for biomining can be implementing to isolate one of the two or more types of ores and / or concentrates.

[0118] In some embodiments, the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, beryllium, sodium, magnesium, aluminum, potassium, calcium, gallium, rubidium, strontium, tin, cesium, barium, thallium, lead, bismuth, francium, radium, neodymium, and dysprosium. The target metal can comprise a rare earth metal. The rare earth metal can be selected from the group consisting of lanthanum, cerium, praseodymium, neodymium, promethium, europium, gadolinium, terbium, dysprosium, holmium, erbium, thulium, ytterbium, lutetium, scandium, and yttrium. The target metal can comprise a transition metal. The transition metal can be selected from the group consisting of scandium, titanium, vanadium, chromium, manganese, iron, cobalt, nickel, copper, zinc, yttrium, zirconium, niobium, molybdenum, technetium, ruthenium, rhodium, palladium, silver, cadmium, lutetium, hafnium, tantalum, tungsten, rhenium, osmium, iridium, platinum, gold, and mercury. The target metal can be a metalloid. The metalloid can be selected from the group consisting of boron, silicon, germanium, arsenic, selenium, antimony, selenium, polonium, and astatine. In some embodiments, the target metal comprises copper. In some embodiments, the target metal comprises lithium.

[0119] In some embodiments, at least one optimal peptide molecule for biomining the target metal is identified in absence of a living organism for extraction of the target metal. The at least one optimal peptide molecule can be identified from a biological sample obtained from a living organism. The biological sample can comprise an organ, a tissue, a cell, and / or an organelle. The at least one optimal peptide molecule can be identified based on a cell-free sample obtained from a living organism. Alternatively, or in addition to, the at least one optimal peptide molecule for biomining the target metal is identified within a living organism. The living organism cancomprise archaea, bacteria, and / or eukarya. The eukarya can comprise a plant, an animal, a protozoan, and / or a fungus.

[0120] In some embodiments, the substrate comprises a metal layer. The metal layer can have a thickness of at least or at most about 10 nm, at least or at most about 20 nm, at least or at most about 30 nm, at least or at most about 40 nm, at least or at most about 50 nm, at least or at most about 60 nm, at least or at most about 70 nm, at least or at most about 80 nm, at least or at most about 90 nm, or at least or at most about 100 nm. The metal layer can comprise a pure metal element. The metal layer can comprise an alloy of two or more metal elements.

[0121] In some embodiments, the metal layer comprises a noble metal. The noble metal can be resistant to chemical reactions and / or corrosion. The noble metal can be a naturally occurring metal. The noble metal can be selected from the group consisting of gold, platinum, ruthenium, rhodium, palladium, osmium, iridium, silver, copper, and mercury. The metal layer can comprise an alloy of two or more noble metals. Alternatively, or in addition to, the metal layer can comprise a non-noble metal (e.g, copper, lithium, cobalt, nickel, tin, zinc, aluminum, etc.). The metal layer can comprise an alloy of a noble metal and a non-noble metal.

[0122] In some embodiments, the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer. The transparent layer can comprise a plastic, a crystal, and / or glass. The transparent layer can provide a base for the substrate. A light (e.g., for measuring a metal binding performance and / or effecting conjugation of an amino acid or peptide) can pass through the transparent layer and reach another layer of the substrate. The substrate can be positioned so that the transparent layer is closest to a light source. The transparent layer can have a thickness of at least or at most about 1 nm, at least or at most about 2 nm, at least or at most about 5 nm, at least or at most about 10 nm, at least or at most about 20 nm, at least or at most about 50 nm, at least or at most about 75 nm, at least or at most about 100 nm, at least or at most about 200 nm, at least or at most about 500 nm, at least or at most about 750 nm, or at least or at most about 1000 nm.

[0123] In some embodiments, the substrate comprises (i) an adhesion layer disposed between the transparent layer and the metal layer, (ii) an immobilization layer configured to couple to the plurality of candidate peptide molecules, or both (i) and (ii). The adhesion layer can comprise a metal. The metal can be selected from the group consisting of titanium and chromium. The adhesion layer can have a thickness of at least or at most about 1 nm, at least or at most about 2 nm, at least or at most about 3 nm, at least or at most about 4 nm, at least or at most about 5 nm, at least or at most about 6 nm, at least or at most about 7 nm, at least or at most about 8 nm, atleast or at most about 9 nm, or at least or at most about 10 nm. The immobilization layer can be on an opposite side of the metal layer as the transparent layer and / or the adhesion layer. Alternatively, or in addition to, the immobilization layer can be on a same side of the metal layer as the transparent layer and / or the adhesion layer. The immobilization layer can comprise a photo-activated polymer. The photo-activated polymer can have a change in a property when exposed to light. The change in property can comprise an activation, a deactivation, a formation of a chemical bond, and / or a cleavage of a chemical bond. When the photo-activated polymer is exposed to light, an amino acid and / or a peptide molecule can be conjugated to the immobilization layer. The patterned array of peptide molecules can be conjugated to the immobilization layer. The immobilization layer can comprise one or more linkers. The one or more linkers can undergo a change in property when exposed to light. The change in property can comprise an activation, a deactivation, a formation of a chemical bond, and / or a cleavage of a chemical bond. The change in property of the one or more linkers can allow an amino acid and / or a peptide molecule to be conjugated to the immobilization layer.

[0124] In some embodiments, the metal layer further comprises a dielectric layer. The dielectric layer can be disposed between the metal layer and the immobilization layer. The dielectric layer can comprise a dielectric material. The dielectric material can be a material having a high electrical resistivity. The dielectric material can comprise a ceramic, a plastic, mica, and / or glass. The dielectric layer can act as an insulator for electrical energy. The dielectric layer can stop an electric charge of the metal layer from being conducted to the immobilization layer. The dielectric layer can improve stability and / or provide functionalization control. The dielectric layer can have a thickness of about 10 nm to about 30 nm. The dielectric layer can have a thickness of at least or at most about 5 nm, at least or at most about 10 nm, at least or at most about 20 nm, at least or at most about 40 nm, or at least or at most about 50 nm.

[0125] In some embodiments, a position of each of the plurality of candidate peptide molecules on the patterned array is spatially controlled. The position can be spatially controlled by directing the light from the optical source of the printer to the position. The patterned array can comprise a plurality of sites. Each of the plurality of candidate peptide molecules can be immobilized to the substrate at a different site of the patterned array. Each site can have a diameter of at least or at most about 0.1 mm, at least or at most about 0.25 mm, at least or at most about 0.5 mm, at least or at most about 1 mm, at least or at most about 2 mm, at least or at most about 5 mm, or at least or at most about 10 mm. The sites can be spatially disposed in a repeating pattern, a non-repeating pattern, or in a random arrangement.

[0126] In some embodiments, the plurality of candidate peptide molecules comprises at leastabout 10, at least about 20, at least about 50, or at least about 100 different peptide molecules. The plurality of candidate peptide molecules can comprise at least about 1, at least about 2, at least about 5, at least about 10, at least about 20, at least about 50, at least about 100, at least about 200, at least about 500, at least about 1000, at least about 2000, at least about 5000, and at least about 10000 different peptide molecules.

[0127] In some embodiments, each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues. The plurality of candidate peptide molecules can comprise at most about 10000, at most about 5000, at most about 2000, at most about 1000, at most about 500, at most about 400, at most about 300, at most about 200, at most about 100, at most about 40, at most about 30, at most about 20, at most about 10, at most about 5, at most about 4, at most about 3, or at most about 2 peptide molecules.

[0128] In some embodiments, each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues. Each peptide molecule of the plurality of candidate peptide molecules can have a length of at least or at most about 1, at least or at most 2, at least or at most 3, at least or at most 4, at least or at most 5, at least or at most 10, at least or at most 15, at least or at most 20, at least or at most 25, at least or at most 50, or at least or at most 100 amino acid residues.

[0129] In some embodiments, the plurality of candidate peptide molecules does not comprise an active enzyme. Alternatively, or in addition to, the plurality of candidate peptide molecules comprises an active enzyme. The active enzyme can comprise a metalloenzyme. The metalloenzyme can be a naturally occurring metalloenzyme, a derivative of a naturally occurring metalloenzyme, and / or a synthetic metalloenzyme. The naturally occurring metalloenzyme can be selected from the group consisting of: transferrin, ferritin, catalase, nitrogenase, cytochrome P-450, superoxide dismutase, carbonic anhydrase, carboxypeptidase, alcohol dehydrogenase, alkaline phosphatase, DNA polymerase, RNA polymerase, plastocyanin, glutathione peroxidase, and urease. The metalloenzyme can comprise a metal ion bound to the peptide molecule. The metal ion can comprise a zinc ion, an iron ion, a copper ion, a manganese ion, and a molybdenum ion. The plurality of candidate peptide molecules can have enzymatic activity. The enzymatic activity can comprise metal binding activity and / or catalytic activity.

[0130] In some embodiments, the printing of peptides on a substrate as provided herein may not and need not utilize RNA templates and translation thereof into peptides. Alternatively, or in addition to, the printing of peptides on a substrate as provided herein can comprise immobilizing RNA template molecules on the substrate and performing cell-free translation of the RNAtemplate molecules to generate the peptides.Example ecosystems

[0131] Certain aspects of the present disclosure provide systems and methods for detecting a presence of a sample on a microarray.

[0132] In an aspect, the present disclosure provides an ecosystem for performing both detection of peptide-metal binding and peptide array printing, as provided herein.

[0133] FIG. 5 shows an example system 500 for biomining a target metal from a medium (e.g., solution) obtained from a geological sample. The system can comprise a printer 510, a detector 520, an analysis module 530, and a database 540. The printer can comprise a 3D printer 512, a light source 514, and / or a substrate 516. The 3D printer 512 can print at least one amino acid and / or at least one peptide molecule onto the substrate 516. The light source 514 can direct a light toward the substrate to enable conjugation of the at least one amino acid and / or at least one peptide molecule with the substrate. The printer 510 can thereby print a patterned array of immobilized peptide molecules onto the substrate 516. The printed patterned array can be provided to the detector 520. The detector 520 can comprise the patterned array of immobilized peptide molecules 522, a light source 524, and / or a reader. The light source 524 can direct a light toward the patterned array of immobilized peptide molecules 522. The light can be reflected from the patterned array of immobilized peptide molecules 522 and detected by the reader 526. The detector 520 can provide data obtained by the reader 526 to the analysis module 530. The analysis module 530 can comprise a binding analysis module 532, a peptide analysis module 534, and / or a peptide recommendation module 536. The binding analysis module 532 can analyze a binding curve provided by the detector 520 to determine a metal binding performance. The peptide analysis module 534 can analyze one or more peptide sequences, based on corresponding metal binding performance. The peptide analysis 534 can predict a metal binding performance based on a peptide sequence. The peptide recommendation module 536 can recommend one or more peptide sequences predicted to have a high metal binding performance. The peptide recommendation module 536 can identify an optimal peptide of a plurality of candidate peptides and / or design in silico an optimal peptide. The binding analysis module 532, the peptide analysis module 534, and / or the peptide recommendation module 536 can communicate to enable reinforcement learning. The analysis module 530 can provide the one or more recommended peptide sequences to the printer 510 to be printed. The analysis module 530 can communicate with the database 540 to provide and / or receive data. The database 540 can comprise a binding performance database 542 and / or a peptide sequence database 544. The binding performancedatabase 542 can comprise binding data detected by the detector 520, or metal binding performance determined by the binding analysis module 532. The peptide sequence database 544 can comprise peptide sequence data corresponding to the binding data. Results of the analysis at the analysis module 530 can be stored at the database 540. The analysis at the analysis module 530 can be performed based on stored data retrieved from the database 540.

[0134] In some embodiments, the system as provided herein further comprises an adsorption column. FIGs. 12A and 12B show an adsorption column for biomining a target metal from a medium obtained from a geological sample. The adsorption column can comprise a biosorbent. The biosorbent can comprise a matrix comprising one or more biomolecules. The medium can be pumped into the adsorption column via a first opening. The medium can comprise a brine. The medium can comprise a plurality of types of ions, including an ion of the target metal (e.g., lithium ions, copper ions). The medium can contact the biosorbent in the adsorption column. The medium can be pumped out of the adsorption column at a second opening, after the ion of the target metal has been selectively adsorbed to the biosorbent. After the medium has been pumped out of the adsorption column, a second medium can be pumped into the adsorption column. The second medium can comprise a desorption medium, which can cause adsorbed ions to desorb from the biosorbent. The second medium can be pumped out of the adsorption column. The second medium, when pumped out, can comprise the ion of the target metal. A coupling efficiency between the target metal and the matrix can be at least about 95%.

[0135] In an aspect, the present disclosure provides a system for detecting a presence of a sample on a microarray. The system can comprise a 3D printer. The 3D printer can comprise a light engine and an optical pick-up unit. The system can further comprise a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate. The system can further comprise a reader to capture kinetic data for each of the one or more immobilized molecules. The system can further comprise an artificial intelligence (Al) algorithm to generate a model based on a codebase.

[0136] In another aspect, the present disclose provides a system for detecting a presence of a sample on a microarray. The system can comprise a 3D printer. The 3D printer can comprise a light engine and an optical pick-up unit. The system can further comprise a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate. The system can further comprise to capture endpoint measurements for each of the one or more immobilized molecules. The system can further comprise an artificial intelligence (Al) algorithm to generate a model based on a codebase.

[0137] In another aspect, the present disclosure provides a method of use of any one of thesystems provided herein. The method can be for detecting a presence of a sample on a microarray. The method can comprise providing a substrate. The method can further comprise positioning one or more molecule variants on a surface of the substrate. The method can further comprise performing analyte perfusion. The method can further comprise capturing kinetic data for the one or more molecule variants. The method can further comprise analyzing the captured kinetic data. The method can further comprise employing an artificial intelligence (Al) algorithm on the captured kinetic data. The method can further comprise generating one or more patterns from a plurality of molecular interactions of the one or more molecule variants. The method can further comprise generating a codebase based on one or more patterns.

[0138] In some embodiments, the surface of the first side of the substrate is a photosensitive surface.

[0139] In some embodiments, the reader is based on at least surface plasmon resonance.

[0140] In some embodiments, the molecule is a biomolecule.

[0141] In some embodiments, the reader is based on at least a fluorescent microarray scanner.

[0142] The fields of biology and medicine have transitioned into the genomic era, marked by the completion or imminent completion of genomes for numerous organisms. This includes not only humans, but also mice, fruit flies, rice, maize, soybean, and many other plant, animal and microbe species. This advancement indicates progress in basic biological research and medical technology. To harness the potential benefits of the vast amount of raw data obtained from genome sequencing endeavors, researchers focus on functional genomics and proteomics studies. For instance, the recent completion of the human genome project has provided the sequence of approximately 3 billion base pairs, representing the human genome in a four-letter alphabet. This wealth of sequence information opens avenues for detailed exploration of gene expression patterns, protein functions, and interactions between proteins and gene sequences. Such efforts are expected to yield extensive valuable data, which will be organized and exploited using emerging bioinformatics techniques.

[0143] The field of proteomics research typically currently generates vast amounts of data, driven by high-throughput methods such as protein expression assays, protein interaction assays, and clinical proteomic tests. Proteomics is configured for understanding the complex dynamics of protein functions and interactions in biological systems. Advanced technologies like mass spectrometry and protein microarrays have become essential tools for large-scale proteomics studies. Protein microarrays, for instance, consist of numerous small analysis sites arranged in a two-dimensional matrix on a substrate surface, where each site contains multiple identical proteinor peptide molecules that can interact with specific targets from a sample. The presence of detectable labels, such as fluorescent molecules, indicates the interaction of the target protein with the probe. Current production techniques for protein microarrays are continually evolving to enable larger arrays with higher density of probe elements, allowing for the simultaneous detection and quantification of thousands of proteins. However, the analysis of protein microarrays generates large volumes of data, posing significant challenges in terms of data processing and storage. Current fluorescent imaging devices must represent each microarray element with multiple pixels, resulting in substantial data processing overhead. Efficient systems and devices are needed to streamline the analysis of protein microarrays, minimizing data storage requirements, reducing costs, and simplifying the labor-intensive processes associated with working with these arrays.

[0144] The present innovation is directed to microfluidic biotechnology and computational biology, focusing on the integration of directed evolution platforms onto microchips. This advanced technology serves as a powerful tool for discovering novel molecules and / or novel functions in known molecules. The platform enables extensive testing by studying interactions between immobilized molecules on the chip surface and perfused analytes, generating valuable and unique experimental datasets. The chip's customizable design allows for the selection and immobilization of proteins, facilitating label-free detection through surface plasmon resonance technology, among other methods. The reader captures kinetics data for each immobilized molecule, producing association and dissociation curves that are analyzed by an Al algorithm. This dataset informs a reinforcing model, enhancing both chip design and biological assays.

[0145] As noted above, the Al algorithm integrated into the system is designed to analyze kinetics data captured for each immobilized molecule, particularly focusing on association and dissociation curves generated during biomolecular interactions. This algorithm employs machine learning techniques to process and interpret the complex kinetics data, extracting meaningful insights about the molecular interactions taking place on the chip's surface. The Al algorithm utilizes a combination of supervised and unsupervised learning approaches to identify patterns and relationships within the kinetics data. Through supervised learning, the algorithm is trained on a labeled dataset of known biomolecular interactions, allowing it to recognize and classify similar interactions in new datasets. Meanwhile, unsupervised learning enables the algorithm to detect hidden structures and clusters within the data, uncovering novel insights that may not be immediately apparent to human analysts. Moreover, the Al algorithm is capable of adapting and refining its models over time through continuous feedback and reinforcement learning. By iteratively analyzing new kinetics data and comparing its predictions with experimentaloutcomes, the algorithm can improve its accuracy and predictive capabilities, enhancing its ability to interpret and predict biomolecular interactions.

[0146] This system is configured to implement biomining, a process that utilizes biological agents to extract valuable metals or minerals from ores, concentrates, or waste materials. Leveraging its capability to synthesize and manipulate biomolecules such as peptides, proteins, and enzymes, the system can enhance the efficiency and effectiveness of metal extraction processes. Through the synthesis of metal-binding peptides or proteins, the system can create functionalized surfaces on the chip capable of selectively capturing and processing target metals from sol medium ution. This immobilization of metal-binding biomolecules facilitates higher yields and reduced processing times, optimizing metal recovery processes. Furthermore, the system's ability to design and modify biomolecules enables the customization of metal-binding properties, allowing for tailored mediums for extracting specific metals of interest from complex mineral matrices. Directed evolution techniques can further optimize the performance of metalbinding peptides or proteins, enhancing their affinity and specificity for target metals. Beyond metal extraction, the system can contribute to environmental remediation efforts by engineering biomolecules to selectively bind and sequester toxic metals from contaminated sites, thereby mitigating environmental pollution and improving ecosystem health.

[0147] This system holds potential for applications in both health and agriculture sectors. In health, the ability to synthesize peptides and proteins enables the development of novel therapeutics, diagnostic tools, drug delivery and drug discovery systems. In implementations, polymers, including, but not limited to, peptides, nucleotides, sugars, plastics and / or a combination of them may be synthesized on the chip surface and engineered to target disease biomarkers or pathogens, facilitating the development of precision medicines and rapid diagnostic tests. Additionally, the system's capability to immobilize proteins allows for the creation of high- throughput platforms for drug screening and protein-protein interaction studies, accelerating the drug discovery process. Furthermore, the system's integration with label-free detection technologies enhances the sensitivity and accuracy of diagnostic assays, enabling early detection and personalized treatment of diseases. In agriculture, the system's versatility can be leveraged for crop improvement and environmental monitoring. Peptides and proteins synthesized on the chip can be designed to enhance plant growth, increase disease resistance, or improve nutrient uptake efficiency. Additionally, the system's ability to monitor molecular interactions in real-time provides valuable insights into plant-microbe interactions, soil health, and environmental stress responses, facilitating informed decision-making in crop management and soil conservation efforts.

[0148] FIG. 1 is a block diagram that describes an example system 100, according to some embodiments of the present disclosure. In some embodiments, the system 100 may include a 3D printer 110, a light engine 120, a substrate 130 with a first side and a second side, one or more immobilized molecules 140 disposed at a surface of the first side of the substrate 130, a reader 150 to capture kinetic data for each of the one or more immobilized molecules 140, and an artificial intelligence 160 (Al) algorithm to generate a model based on codebase. An optical pickup unit is implemented in the system. In some embodiments, a surface of the first side of the substrate 130 may be a photo-sensitive surface. In some embodiments, the reader 150 may be based on surface plasmon resonance.

[0149] FIGs. 2A and 2B are flowcharts that describe an example method for detecting a presence of a sample on a microarray, according to some embodiments of the present disclosure. In some embodiments, at 202, the method may include providing a substrate. At 204, the method may include positioning one or more protein variants on a surface of the substrate. At 206, the method may include performing analyte perfusion. At 208, the method may include capturing kinetic data for the one or more protein variants. At 210, the method may include analyzing the captured kinetic data. At 212, the method may include employing an artificial intelligence (Al) algorithm on the captured kinetic data. At 214, the method may include generating one or more patterns from a plurality of molecular interactions of the one or more protein variants. At 216, the method may include generating a codebase based on the one or more patterns.

[0150] FIG. 3 illustrates an example flowchart of a technology platform to implement an example system for conducting molecular interaction studies, where immobilized molecules on a chip surface interact with perfused analytes to generate valuable experimental datasets. This system involves an instrument designed to print molecules onto photosensitive surfaces. Key innovations include the adaptation of an optical pick-up unit (OPU) from consumer electronics, modified to operate at a custom wavelength suited for biological applications. Integrated into a 3D printer, this modified OPU provides precise laser control across the x, y, and z axes, enabling the selective illumination of specific surface areas. The system utilizes photosensitive surfaces to assemble molecules such as peptides up to 15 amino acids, with the potential for future fabrication of longer proteins (Reference: "Protein Chip Fabrication by Capture of Nascent Polypeptides" by Sheng-Ce Tao & Heng Zhu, 2006). Proprietary software manages the entire printing process, ensuring the high- density immobilization of millions of proteins on each chip. This novel approach repurposes existing technology through wavelength adaptation, precise control integration, and real-time molecular interaction data, facilitating high-resolution, high- throughput molecule assembly.

[0151] In implementations, this system implements a low-cost optical lithography by leveraging its chip making instrument for molecule printing onto a photosensitive surface. By repurposing consumer electronics, particularly the optical pick-up unit (OPU), and integrating it into a 3D printer, the system achieves precise control over laser illumination on a photosensitive surface. This adaptation enables the selective printing of biomolecules, such as peptides and proteins, onto the chip's surface in discrete fields. Leveraging the photo-sensitive properties of the surface and the modified OPU, the system facilitates the assembly of biomolecules with high spatial resolution and accuracy. Moreover, the implementation of software streamlines the printing process, enabling high-density immobilization of biomolecules on each chip. By utilizing existing consumer electronics and innovative software solutions, this system reduces the cost associated with traditional optical lithography techniques, democratizing access to advanced fabrication methods for biological and nanotechnology applications.

[0152] In some implementations, this system revolutionizes biomolecule construction through its chemical and cell-free approach, offering unprecedented flexibility and efficiency in synthesizing peptides, proteins, and other biomolecules. By leveraging chemical synthesis techniques, it bypasses the need for cellular machinery, enabling precise control over biomolecule design and production. The system's capability to synthesize peptides up to 20 amino acids and proteins up to 30 kilodaltons (Kd) in size on the chip's surface demonstrates its versatility in constructing complex biomolecules. The system can synthesize peptides up to about 20, about 30, about 40, about 50, about 100, about 200, about 300, about 400, or about 500 amino acids. The system can synthesize proteins up to about 30 Kd, up to about 35 Kd, up to about 40 Kd, up to about 45 Kd, up to about 50 Kd, up to about 55 Kd, or up to about 60 Kd in size. To synthesize functional polypeptides exceeding 20 amino acids in length, the utilization of cellular machinery is important. This synthesis will be achieved through cell-free transcription and translation systems, leveraging in vitro reconstituted components of the transcriptional and translational machinery to facilitate the accurate and efficient assembly of these macromolecules. Additionally, the cell-free lysate optimized for proper protein folding facilitates protein expression, ensuring covalent and stable attachment to the chip, thus enhancing robustness and preservation of the molecules. This chemical and cell-free approach not only accelerates biomolecule synthesis but also eliminates the risk of cellular contamination and regulatory hurdles associated with traditional cell-based methods.

[0153] FIG. 9 shows translation of RNA molecules to peptide molecules on a patterned array. The RNA molecules can comprise full-length mRNA and / or short RNA. The RNA molecules can be immobilized to the patterned array. The immobilized RNA molecules can betranslated to peptide molecules, wherein the peptide molecules are immobilized to the patterned array upon being translated. FIG. 9 shows a representative figure of on-chip translation of immobilized full-length mRNA molecules coding for green fluorescent protein (GFP). As shown in FIG. 9, fluorescence corresponding to GFP is not detected at a first time point prior to the on- chip translation. At a second time point after the on-chip translation, fluorescence corresponding to GFP is detected, indicating that the chip comprises immobilized GFP proteins and the on-chip translation was successful. At the second time point, a negative control molecule is not expressed, indicating that GFP was selectively synthesized. FIG. 9 further shows a representative figure of on-chip translation of immobilized short RNA molecules coding for a FLAG protein tag. Fluorescence corresponding to the FLAG protein tag is not detected at a first time point prior to the on-chip translation. At a second time point after the on-chip translation, fluorescence corresponding to the FLAG protein tag is detected, indicating that the chip comprises immobilized FLAG protein tags and the on-chip translation was successful. In comparison, a negative control molecule is not detected at the second time point, indicating that the on-chip translation was performed for the FLAG protein tag selectively.

[0154] In some implementations, this system integrates a sophisticated reader configured for label-free Surface Plasmon Resonance (SPR) and / or fluorescence detection, enhancing its capability to analyze biomolecular interactions with high sensitivity and specificity. The SPR biosensor enables real-time measurement of biomolecule binding kinetics, allowing for the precise characterization of molecular interactions without the need for traditional labeling tags. By monitoring changes in refractive index as biomolecules bind to the chip surface, SPR provides valuable insights into association and dissociation rates, affinity constants, and binding kinetics. Additionally, the system incorporates fluorescence detection capabilities, among others, enabling the analysis of biomolecular interactions using fluorescently labeled molecules. This dual detection approach offers flexibility in experimental design and enables the system to accommodate a wide range of biomolecular assays and applications.

[0155] In some implementations, this system achieves a throughput ranging from 1 to 10,000, 000 biomolecular interactions, making it a highly efficient platform for high-throughput experimentation and analysis. The system can achieve a throughput of at least or at most about 1, at least or at most about 5, at least or at most about 10, at least or at most about 50, at least or at most about 100, at least or at most about 500, at least or at most about 1,000, at least or at most about 5,000, at least or at most about 10,000, at least or at most about 50,000, at least or at most about 100,000, at least or at most about 500,000, at least or at most about 1,000,000, at least or at most about 5,000,000, or at least or at most about 10,000,000 biomolecular interactions.

[0156] Through its advanced technology, the system enables the synthesis, immobilization, and analysis of biomolecules on a massive scale. By printing peptides, proteins, DNA, RNA, and other biomolecules onto the chip's surface, the system creates dense arrays capable of accommodating millions of interactions simultaneously. Furthermore, its ability to perform realtime measurements of biomolecular interactions using label -free Surface Plasmon Resonance (SPR) or labeled fluorescence, colorimetric, chemiluminescent detection technologies with ultra high throughput without compromising data quality or accuracy. This high-throughput capability allows researchers to rapidly screen large libraries of biomolecules, analyze complex molecular interactions, and generate vast datasets for comprehensive analysis.

[0157] In some implementations, this system implements affinity kinetics as a fundamental data type, providing valuable insights into the strength and specificity of biomolecular interactions. As noted above, through its SPR technology, the system enables real-time measurement of biomolecule binding kinetics, allowing researchers to characterize the affinity and kinetics of molecular interactions with high precision. By monitoring changes in refractive index as biomolecules bind to the chip's surface, the system generates association and dissociation curves for each immobilized molecule, providing detailed information on binding kinetics, including association rate constants (ka), dissociation rate constants (kd), and equilibrium dissociation constants (KD). This data type is critical for understanding the dynamic nature of biomolecular interactions, elucidating the strength of binding between molecules, and informing the design of novel therapeutics, diagnostics, and drug delivery systems.

[0158] FIG. 4 is a diagram illustrating a step-by-step description of an example flow of the technology platform of FIG. 3. The biological assay and reader system facilitate in vitro experiments by perfusing various analytes (such as molecules, proteins, and drugs) across the chip's surface to investigate their interactions with immobilized proteins. Key aspects of this system include the real-time measurement of interaction kinetics between surface-bound proteins and analytes, the adoption of a label-free system allowing direct observation of molecular interactions without traditional labeling tags (such as GFP, 6xHis, Flag), the capture of association and dissociation curves for each immobilized molecule to provide detailed insights into interaction strength, and the utilization of established reader technology (such as surface plasmon resonance) for precise and sensitive kinetic measurements. The novelty of this approach lies in the integration of label- free technologies with high-throughput experimentation to generate dense datasets on molecular interactions.

[0159] The platform integrates experimental data from chip readouts into a codebase for analysis. The platform is configured to generate one or more interaction curves to map sequence-specific affinities across multiple experiments. Additionally, the platform utilizes a codebase to correlate molecular sequences with interaction profiles across diverse substrates. Through constant reinforcement of learning models with experimental data, it leverages graph neural networks for predictive analytics. The system focuses on modeling based on empirical data, enhancing accuracy and enabling targeted understanding of interaction pathways. The novelty lies in the development of a robust, self-learning codebase capable of predictive modeling based on experimental measurements, contributing to a deeper understanding of molecular interactions.

[0160] Chip-based Protein Engineering for Industrial Applications utilizes a sophisticated approach to biomolecule synthesis, protein expression, protein activity detection, and directed evolution. Biomolecules such as peptides, DNA, or RNA are synthesized on the chip surface within discrete fields ranging from 1 to 100 pm, with the option to include barcodes for later DNA plasmid capture. The discrete fields can be at least or at most about 0.1 pm, at least or at most about 0.5 pm, at least or at most about 1 pm, at least or at most about 5 pm, at least or at most about 10 pm, at least or at most about 50 pm, at least or at most about 100 pm, at least or at most about 500 pm, or at least or at most about 1000 pm. A chip can comprise at least or at most about 1, at least or at most about 2, at least or at most about 5, at least or at most about 10, at least or at most about 50, at least or at most about 100, at least or at most about 500, at least or at most about 1000 discrete fields, at least or at most about 5000, or at least or at most about 10,000 discrete fields. Each biomolecule encodes a single peptide or protein, remaining tethered to the chip. Protein expression is initiated by adding a cell-free lysate optimized for proper protein folding, resulting in proteins that remain covalently and stably bound to the chip. Protein activity can be measured using either a SPR biosensor or a fluorescent microarray scanner, allowing for the synthesis and analysis of one to millions of biomolecules on a single chip. In cases where binding kinetics are implemented, an SPR assay against desired binders is performed, with top performers identified for further evolution through mutation of specific amino acids to achieve higher or lower affinity.

[0161] Unique features of the platform include the ability to independently optimize the association rate constant (Ka) and dissociation rate constant (Kd), as both processes are monitored. This capability is particularly advantageous for developing affinity resins or matrices where a high Ka and moderate Kd are desirable. The platform also goes beyond natural limits by “alienating” top performers through structural modifications, introducing non-natural synthetic compounds to enhance flexibility, improve affinity, and reduce cleavage sites. Applications include, but are not limited to, lithium extraction, copper processing, rare earth element separation, pharmaceuticals and agronomical. To select the most effective modifications,AlphaFold-like models will be built. Additionally, the platform can assist in the production of these “alienated” biomolecules by modifying the transcription / translation machinery within a cell-free system.

[0162] In implementations of this system, biomolecules can be synthesized and manipulated for industrial applications. Typically, many peptides are tools in drug development, enzyme inhibitors, and studies of protein-protein interactions. In implementations, this system is configured to utilize proteins with molecular weights of up to 30 kilodaltons (kDa). The system can be configured to utilize proteins with molecular weights of up to about 30 kDa, up to about 35 kDa, up to about 40 kDa, up to about 45 kDa, up to about 50 kDa, up to about 55 kDa, or up to about 60 kDa in size. Proteins, composed of one or more chains of amino acids folded into complex three-dimensional structures, play essential roles in the structure, function, and regulation of biological systems. Here, proteins can be expressed or synthesized, immobilized, and studied for their interactions with other molecules. While not explicitly mentioned, the system may also have the capability to work with sugars, carbohydrates, lipids or different kinds of biomolecules. Sugars are biomolecules involved in energy metabolism, cell signaling, and structural support. They can be modified or engineered for applications in pharmaceuticals, food additives, and materials science. Lastly, the system may also handle other types of biomolecules such as nucleic acids (DNA, RNA), lipids, or small molecules. Nucleic acids facilitate genetic information storage, transmission, and regulation, while lipids are involved in cell membrane structure, signaling, and energy storage. Further, small molecules encompass organic compounds with diverse functions, including enzyme substrates, inhibitors, and signaling molecules. This versatility allows for applications in industrial settings, including biopharmaceutical production, biochemical research, and biotechnology development.Examples of Machine Learning Methodologies

[0163] As used in this specification and the appended claims, the terms “artificial intelligence,” “artificial intelligence techniques,” “artificial intelligence operation,” and “artificial intelligence algorithm” generally refer to any system or computational procedure that may take one or more actions that simulate human intelligence processes for enhancing or maximizing a chance of achieving a goal. The term “artificial intelligence” may include “generative modeling,” “machine learning” (ML), or “reinforcement learning” (RL). As used in this specification and the appended claims, the terms “machine learning,” “machine learning techniques,” “machine learning operation,” and “machine learning model” generally refer to any system or analytical or statistical procedure that may progressively improve computer performance of a task.

[0164] In some cases, ML may generally involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. ML may include a ML model (which may include, for example, a ML algorithm). Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The ML model may be a trained model. ML techniques may comprise one or more supervised, semi-supervised, self-supervised, or unsupervised ML techniques. For example, an ML model may be a trained model that is trained through supervised learning (e.g., various parameters are determined as weights or scaling factors). ML may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deep learning, or ultra-deep learning. ML may comprise: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regression splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, non-negative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, auto-encoders, stacked autoencoders, perceptrons, multi-layer perceptrons, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, residual neural networks, physics-informed neural networks, long short-term memory, deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, large language models, transformer models, vision transformers, or generative adversarial networks.

[0165] Training the ML model may include, in some cases, selecting one or more untrained data models to train using a training data set. The selected untrained data models may include any type of untrained ML models for supervised, semi-supervised, self-supervised, or unsupervised machine learning. The selected untrained data models may be specified based upon input (e.g., user input) specifying relevant parameters to use as predicted variables or other variables to use as potential explanatory variables. For example, the selected untrained datamodels may be specified to generate an output (e.g., a prediction) based upon the input. Conditions for training the ML model from the selected untrained data models may likewise be selected, such as limits on the ML model complexity or limits on the ML model refinement past a certain point. The ML model may be trained (e.g., via a computer system such as a server) using the training data set. In some cases, a first subset of the training data set may be selected to train the ML model. The selected untrained data models may then be trained on the first subset of training data set using appropriate ML techniques, based upon the type of ML model selected and any conditions specified for training the ML model. In some cases, due to the processing power requirements of training the ML model, the selected untrained data models may be trained using additional computing resources (e.g., cloud computing resources). Such training may continue, in some cases, until at least one aspect of the ML model is validated and meets selection criteria to be used as a predictive model.

[0166] In some cases, one or more aspects of the ML model may be validated using a second subset of the training data set (e.g., distinct from the first subset of the training data set) to determine accuracy and robustness of the ML model. Such validation may include applying the ML model to the second subset of the training data set to make predictions derived from the second subset of the training data. The ML model may then be evaluated to determine whether performance is sufficient based upon the derived predictions. The sufficiency criteria applied to the ML model may vary depending upon the size of the training data set available for training, the performance of previous iterations of trained models, or user-specified performance requirements. If the ML model does not achieve sufficient performance, additional training may be performed. Additional training may include refinement of the ML model or retraining on a different first subset of the training dataset, after which the new ML model may again be validated and assessed. When the ML model has achieved sufficient performance, in some cases, the ML may be stored for present or future use. The ML model may be stored as sets of parameter values or weights for analysis of further input (e.g., further relevant parameters to use as further predicted variables, further explanatory variables, further user interaction data, etc.), which may also include analysis logic or indications of model validity in some instances. In some cases, a plurality of ML models may be stored for generating predictions under different sets of input data conditions. In some cases, the ML model may be stored in a database (e.g., associated with a server).Examples of Decision Trees and Random Forests

[0167] As described above, the machine learning model may implement a decision tree. Adecision tree may be a supervised ML algorithm that can be applied to both regression and classification problems. A decision tree may grow from a root (base condition), and when it meets a condition (internal node / feature), it may split into multiple branches. The end of the branch that does not split anymore may be an outcome (leaf). A decision tree can be generated using a training data set according to the following operations: (1) starting from a root node (the entire dataset), the algorithm may split the dataset in two branches using a decision rule or branching criterion; (2) each of these two branches may generate a new child node; (3) for each new child node, the branching process may be repeated until the dataset cannot be split any further; (4) each branching criterion may be chosen to maximize information gain (e.g., a quantification of how much a branching criterion reduces a quantification of how mixed the labels are in the children nodes). The labels may be the data or the classification that is predicted by the decision tree.

[0168] A random forest regression is an extension of the decision tree model that tends to yield more robust predictions by stretching the use of the training data partition. Whereas a decision tree may make a single pass through the data, a random forest regression may bootstrap 50% of the data (e.g., with replacement) and build many trees. Rather than using all explanatory variables as candidates for splitting, a random subset of candidate variables may be used for splitting, which may enable trees that have different data and different variables (hence the term random). The predictions from the trees, which may be collectively referred to as the “forest,” may then be averaged to produce a final prediction. Many trees (e.g., ten trees, fifty trees, one hundred trees, one thousand trees, etc.) may be included in a random forest model, with a number (e.g., 3, 6, 10, etc.) of terms sampled per split, a minimum of number (e.g., 1, 2, 4, 10, etc.) of splits per tree, and a minimum split size (e.g., 16, 32, 64, 128, 256, etc.). Random forests may be trained in a similar way as decision trees. Specifically, training a random forest may include the following operations: (1) randomly select k features from the total number of features; (2) create a decision tree from these k features using the same operations as for generating a decision tree; and (3) repeat the previous two operations until a target number of trees is created.

[0169] As disclosed, a random forest classifier, which may comprise a plurality of decision trees wherein the output prediction may be the mode of the predicted classifications of the individual trees, can be helpful in reducing overfitting to training data. In some cases, an ensemble of decision trees can be constructed using a random subset of features at each split or decision node. The Gini criterion may be employed, in some cases, to choose the best partition, wherein decision nodes having the lowest calculated Gini impurity index are selected. The Giniimpurity can be used, in some cases, as a criterion to find informative features based on which the splits in each decision tree may be constructed.

[0170] In some cases, each decision tree of a random forest may comprise one or more decision nodes, wherein each decision node specifies a predicate condition. For example, decision node may predicate the condition that, for a given dataset, the outcome to an question is a specific outcome. At each decision node, a decision tree can be split based on whether the predicate condition attached to the decision node holds true, leading to various prediction nodes. Each prediction node can comprise output values that represent “votes” for one or more of the classifications or conditions being evaluated by the assessment model. At prediction time, a “vote” can be taken over all of the decision trees, and the majority vote (or mode of the predicted classifications) can be output as the predicted classification.

[0171] In some cases, when the dataset being queried in the assessment model reaches a “leaf’, or a final prediction node with no further downstream splits, the output values of the leaf can be output as the votes for the particular decision tree. Since a random forest model comprises a plurality of decision trees, the final votes across all trees in the forest can be summed to yield the final votes and the corresponding classification of the subject. A large number of decision trees can help reduce overfitting of the assessment model to the training data, by reducing the variance of each individual decision tree. For example, an assessment model can comprise, for example, at least about 3 decision trees, at least about 5 decision trees, at least about 10 decision trees, at least about 20 decision trees, at least about 50 decision trees, at least about 100 decision trees, etc.Examples of Support Vector Machines

[0172] As also described above, the machine learning model may implement support vector machine learning techniques. In machine learning, support vector machines (SVMs) may be supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. SVMs may be a robust prediction method, being based on statistical learning. SVMs may be well-suited for domains characterized by the existence of large amounts of data, noisy patterns, or the absence of general theories.

[0173] In general terms, SVMs may map input vectors into high dimensional feature space through non-linear mapping function, chosen a priori. In this high dimensional feature space, an optimal separating hyperplane may be constructed. The optimal hyperplane may then be used to determine, for example, class separations, regression fit, accuracy in density estimation, etc. More formally, a SVM may construct a hyperplane or set of hyperplanes in a high or infinite-dimensional space, which can be used for classification, regression, or other tasks like outlier detection.

[0174] Support vectors may be defined as the data points that lie closest to the decision surface (or hyperplane). Support vectors may therefore be the data points that are most difficult to classify and may have direct bearing on an optimum location of the decision surface. Given a set of training examples, each marked as belonging to one of two categories, an SVM training algorithm may build a model that assigns new examples to one category or the other, making it a non-probabilistic binary linear classifier (although methods such as Platt scaling exist to use SVM in a probabilistic classification setting). In some cases, SVMs may map training examples to points in space so as increase (e.g., maximize) the width of the gap between the two categories. New examples may then be mapped into that same space and predicted to belong to a category based on which side of the gap the new examples fall. In addition to performing linear classification, SVMs can efficiently perform a non-linear classification using what may be referred to as a kernel trick, implicitly mapping their inputs into high-dimensional feature spaces.

[0175] Within a support vector machine, the dimensionally of the feature space may be large. For example, a fourth-degree polynomial mapping function may cause a 200-dimensional input space to be mapped into a 1.6 billionth dimensional feature space. The kernel trick and the Vapnik-Chervonenkis dimension may allow the SVM to thwart the “curse of dimensionality” limiting other methods and effectively derive generalizable answers from this very high dimensional feature space. Accordingly, SVMs may assist in discovering knowledge from vast amounts of input data.Examples of Long Short-Term Memory

[0176] Long short-term memory (LSTM) may be an artificial neural network used in the fields of artificial intelligence and deep learning. Unlike standard feedforward neural networks, LSTM may use feedback connections. The LSTM architecture may provide a short-term memory for a recurrent neural network (RNN). Such RNN can process not only single data points (such as images), but also entire sequences of data (such as audio or video). This characteristic may enable LSTM networks to be well-suited for processing and predicting data. The name of LSTM may refer to the analogy that a standard RNN has both “long-term memory” and “short-term memory.” The connection weights and biases in the RNN may change once per episode of training, analogous to how physiological changes in synaptic strengths store long-term memories; the activation patterns in the network may change once per time-step, analogous to how the moment-to-moment change in electric firing patterns in the brain store short-term memories. TheLSTM architecture may provide a short-term memory for an RNN that can last many (e.g., hundreds, thousands, tens of thousands, etc.) timesteps.

[0177] In some cases, a LSTM unit may comprise a cell, an input gate, an output gate, and a forget gate. The cell may remember values over arbitrary time intervals and the input gate, the output gate, and the forget gate may regulate the flow of information into and out of the cell. Forget gates may be used to decide what information to discard from a previous state by assigning a previous state, compared to a current input, a value between 0 and. For example, a (e.g., rounded) value of 1 may mean to keep the information, and a (e.g., rounded) value of 0 means to discard it). The input gate may decide which pieces of new information to store in the current state, using the same system as the forget gates. The output gate may control which pieces of information in the current state to output (e.g., by assigning a value from 0 to 1 to the information, considering the previous and current states). Selectively outputting relevant information from the current state may allow the LSTM network to maintain useful, long-term dependencies to make predictions, both in current and future time-steps. In some cases, LSTM networks may be well-suited to classifying, processing and making predictions based on time series data, since there can be lags of unknown duration between important events in a time series. LSTMs may resolve the vanishing gradient problem that can be encountered when training certain RNNs. Relative insensitivity to gap length may be an advantage of LSTM over RNNs, hidden Markov models, and other sequence learning methods in numerous applications.

[0178] In some cases, LSTMs may be used with one or more various types of neural networks (e.g., convolutional neural networks (CNNs), deep neural network (DNNs), RNNs, etc.). In some cases, CNNs, DNNs, and LTSMs are complementary in their modeling capabilities and may be combined a unified architecture. For example, in such unified architecture, CNNs may be well-suited at reducing frequency variations, LSTMs may be well-suited at temporal modeling, and DNNs may be well-suited for mapping features to a more separable space. For example, input features to a ML model using LSTM techniques in the unified architecture may include segment features for each of a plurality of segments. To process the input features for each of the plurality of segments, the segment features for the segment may be processed using one or more CNN layers to generate first features for the segment; the first features may be processed using one or more LSTM layers to generate second features for the segment; and the second features may be processed using one or more fully connected neural network layers to generate third features for the segments, where the third features may be used for classification operations. In some cases, to process the first features using the one or more LSTM layers to generate the second features, the first features may be processed using a linear layer to generatereduced features having a reduced dimension from a dimension of the first features; and the reduced features may be processed using the one or more LSTM layers to generate the second features. Short-term features having a first number of contextual frames may be generated based on the input features, where features generated using the one or more CNN layers may include long-term features having a second number of contextual frames that are more than the first number of contextual frames of the short-term features. In some cases, the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers may have been jointly trained to determine trained values of parameters of the one or more CNN layers, the one or more LSTM layers, and the one or more fully connected neural network layers. In some cases, the input features may include log-mel features having multiple dimensions. The input features may include one or more contextual frames indicating a temporal context of a signal (e.g., input data). Advantageously, implementations for such unified architecture may leverage complementary advantages associated with each of a CNN, a DNN, and a LTSM. For example, convolutional layers may reduce spectral variation in input, which may help the modeling of LSTM layers. Having DNN layers after LSTM layers may help reduce variation in the hidden states of the LSTM layers. Training the unified architecture jointly may provide a better overall performance. Training in the unified architecture may also remove the need to have separate CNN, LSTM and DNN architectures, which may be expensive (e.g., in computational resource, in network traffic, in financial resources, in energy consumption, etc.). By adding multi-scale information into the unified architecture, information may be captured at different time scales.Computer systems

[0179] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 13 shows an example computer system 1301 that is programmed or otherwise configured to execute one or more algorithms disclosed herein. The computer system 1301 can regulate various aspects of biomining a target metal from a medium (e.g., solution) obtained from a geological sample of the present disclosure, such as, for example, analyzing metal binding performance, identifying the at least one optimal peptide molecule, and / or selecting or designing the plurality of candidate peptide molecules. The computer system 1301 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.

[0180] The computer system 1301 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 1305, which can be a single core or multi core processor, or aplurality of processors for parallel processing. The computer system 1301 also includes memory or memory location 1310 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1315 (e.g., hard disk), communication interface 1320 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1325, such as cache, other memory, data storage and / or electronic display adapters. The memory 1310, storage unit 1315, interface 1320 and peripheral devices 1325 are in communication with the CPU 1305 through a communication bus (solid lines), such as a motherboard. The storage unit 1315 can be a data storage unit (or data repository) for storing data. The computer system 1301 can be operatively coupled to a computer network (“network”) 1330 with the aid of the communication interface 1320. The network 1330 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1330 in some cases is a telecommunication and / or data network. The network 1330 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1330, in some cases with the aid of the computer system 1301, can implement a peer-to-peer network, which may enable devices coupled to the computer system 1301 to behave as a client or a server.

[0181] The CPU 1305 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1310. The instructions can be directed to the CPU 1305, which can subsequently program or otherwise configure the CPU 1305 to implement methods of the present disclosure. Examples of operations performed by the CPU 1305 can include fetch, decode, execute, and writeback.

[0182] The CPU 1305 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1301 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0183] The storage unit 1315 can store files, such as drivers, libraries and saved programs. The storage unit 1315 can store user data, e.g., user preferences and user programs. The computer system 1301 in some cases can include one or more additional data storage units that are external to the computer system 1301, such as located on a remote server that is in communication with the computer system 1301 through an intranet or the Internet.

[0184] The computer system 1301 can communicate with one or more remote computer systems through the network 1330. For instance, the computer system 1301 can communicate with a remote computer system of a user (e.g., an instrument for measuring metal binding performance). Examples of remote computer systems include personal computers (e.g., portablePC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1301 via the network 1330.

[0185] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1301, such as, for example, on the memory 1310 or electronic storage unit 1315. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 1305. In some cases, the code can be retrieved from the storage unit 1315 and stored on the memory 1310 for ready access by the processor 1305. In some situations, the electronic storage unit 1315 can be precluded, and machine-executable instructions are stored on memory 1310.

[0186] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a precompiled or as-compiled fashion.

[0187] Aspects of the systems and methods provided herein, such as the computer system 1301, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk.“Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any mediumthat participates in providing instructions to a processor for execution.

[0188] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0189] The computer system 1301 can include or be in communication with an electronic display 1335 that comprises a user interface (UI) 1340 for providing, for example, the metal binding performance, the at least one optimal peptide molecule, or the plurality of candidate peptide molecules. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0190] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1305. The algorithm can, for example, analyze the metal binding performance, identify the at least one optimal peptide molecule, rank the plurality of candidate peptide molecules, design in silico the at least one peptide variant, or selecting or designing the plurality of candidate peptide molecules.

[0191] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to theaforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.EXAMPLES

[0192] Example 1: Medium for biomining

[0193] As provided herein, biomining (or detection of peptide-metal binding) can be performed in a medium (e.g., solution) comprising a target metal. In some embodiments, the medium comprises an acidic medium (e.g., a filtered acidic aqueous medium). The pH value of the acidic medium can be between about 3 and about 5. The acidic medium comprises disaggregated mineral slurry in the range of 1-10% weight per volume (w / v). The particle size of the minerals can be filtered below 10 pm to prevent channel clogging and ensure smooth perfusion across the substrate (e.g., the substrate of a SPR chip as provided herein). In some cases, chelating agents (e.g., EDTA or citrate) may be included in the medium, e.g., to aid in mobilizing loosely bound metals.

[0194] Example 2: Selection of optimal peptide molecules

[0195] As provided herein, a metal binding performance for a plurality of candidate peptide molecules arranged in a patterned array on a substrate (e.g., a SPR chip) can be analyzed, and such data can be utilized to select optimal peptide molecules (or optimal peptide molecule candidates) for biomining. In some embodiments, the selection of optimal peptide molecules can be performed via a multi-stage filtering and ranking process. The process comprises a primary screening step, in which initial kinetic parameters (e.g., Ka and / or Kd) are obtained and extracted (e.g., from the SPR data). The process further comprises utilizing a scoring model. In the scoring model, a weighted score is applied to the kinetic parameters such that Ka is prioritized (e.g., 70%) to favor strong initial binding, while Kd is weighted at 30% to maintain moderate off-rates, e.g., to improve material regeneration efficiency. The process further comprises generating AI- informed tiers of the candidate peptide molecules. For examples, top 30% scorers of thecandidate peptide molecules can be used to train graph neural networks (GNNs) programmed to suggest one or more substitutions for directed evolution of peptide molecules for biomining.

[0196] Example 3: Peptide variant design

[0197] As provided herein, given a peptide sequence (e.g., as identified or selected from a pool of candidate peptide molecules), variants of such peptide sequence can be generated then printed on a substrate (e.g., SPR chip) to further assessment for metal binding. A single-point substitution strategy can be implemented to target one or more residues in positions with high contribution to binding energy, e.g., derived from the SPR curve sensitivity analysis and amino acid residue hot spots annotated by a protein modeling software module (e.g., AlphaFold). In some embodiments, peptide variant design process utilizes in silico alanine scanning, comprising systematically replacing each non-alanine amino acid residue in a peptide with alanine and then analyzing the effects of these substitutions on the peptide’s property such as metal binding. The process further comprises GNN-guided residue prioritization, e.g., to identify and rank amino acid residues in a peptide that are most important for metal binding task. The process further comprises using a mutation library builder (e.g., Rosetta or other ML modules) to generate additional peptide sequences or peptide candidate sequences (e.g., for use in biomining, for testing for biomining, etc.). In some cases, the process is designed to avoid mutating core stabilizing residues (e.g., cysteines in loops) in early rounds of peptide variant design to preserve peptide fold integrity. FIG. 10A schematically illustrates an example process of peptide variant design as described herein.

[0198] Example 4: Pattern recognition via machine learning

[0199] As provided herein, a platform comprising a combination of machine learning algorithms can be utilized to identify patterns of molecule interaction (e.g., SPR patterns) between the target metal and the candidate peptide molecules, which patterns can be analyzed and utilized to select optimal peptide molecules and / or design peptide molecule variants for metal binding. In some embodiments, a combination of ML algorithms can collectively perform (i) peptide sequence-kinetics (e.g., Ka and / or Kd) graph embeddings to represent the metal binding interactions of peptide candidates, (ii) time-series clustering of the peptide sequencekinetics curves (e.g., SPR curves) for unsupervised grouping, and (iii) anomaly detection to flat or identify peptide candidates with unexpected (e.g., unexpectedly high) affinity profiles against the target metals. In addition, over time, a reinforcement learning updates the predictive space (e.g., comprising predicted metal binding kinetics performance of the peptide candidates) basedon the difference between predicted and experimental metal binding kinetics (e.g., SPR responses), thereby allowing autonomous refinement.

[0200] Example 5: Chip structure

[0201] As provided herein, a metal binding performance for a plurality of candidate peptide molecules arranged in a patterned array on a substrate (e.g., a sensor chip such as a SPR chip) can be analyzed, and such data can be utilized to select optimal peptide molecules (or optimal peptide molecule candidates) for biomining. Fabrication process of such chip can be compatible with scalable photolithographic and / or additive manufacturing techniques, thus enabling high- density array production for screening workflows (e.g., commercial screening workflows).

[0202] In some embodiments, the chip comprises (e.g., in the order as described) (i) a transparent or semi-transparent substrate (e.g., a glass substrate), (ii) an adhesion layer (e.g., Ti or Cr layer having a thickness ranging between about 2 and about 5 nanometers), (iii) a noble metal layer (e.g., Au layer having a thickness of about 50 nanometers), (iv) optional dielectric layer (e.g., a SiCh layer having a thickness ranging between about 10 nanometers and about 30 nanometers) for stability or peptide / amino acid functionalization control, and (v) a peptide immobilization layer comprising a photo-activated (or activatable) polymer and / or photo- activatable linker chemistry.

[0203] FIG. 14 schematically illustrates an example chip architecture. The chip architecture can comprise (i) a substrate base comprising glass or quartz. The substrate base can provide optical transparency allowing light to pass through during SPR. The chip can comprise, in contact with the substrate base, (ii) an adhesion layer comprising titanium or chromium. The adhesion layer can have a thickness ranging from about 2 nanometers (nm) to about 5 nm. The adhesion layer can promote adhesion of noble metals to the substrate base. The chip can comprise, in contact with the adhesion layer and opposite to the substrate base, (iii) a metal layer comprising gold or platinum. The metal layer can have a thickness of about 50 nm. The metal layer can comprise molecules that are active and detectable during SPR or chemically resilient during localized surface plasmon resonance (LSPR). The chip can optionally comprise, in contact with the metal layer and opposite to the adhesion layer, (iv) a dielectric layer comprising silicon dioxide. The dielectric layer can have a thickness of about 10 nm to about 30 nm. The dielectric layer can enhance stability of the chip and reduce nonspecific binding. The chip can further comprise, in contact with the dielectric layer or the metal layer, (v) a linker layer comprising aminopropyltri ethoxysilane (APTES) or aminothiolphenol. The linker layer can comprise APTES when in contact with the dielectric layer and aminothiolphenol when in contact with themetal layer. The linker layer can anchor peptide synthesis chemistry to the chip. As provided herein, the linker layer may form a self-assembled monolayer (SAM), such as APTES SAM or aminothiolphenol SAM. For example, APTES SAM may be provided on glass or oxide surfaces. In another example, aminothiolphenol SAM may be provided on gold surfaces.

[0204] Example 6: Printing chemistry

[0205] As provided herein, a sensor chip can be generated with peptide candidates printed on the sensor chip, to identify optimal peptide molecules for metal binding and biomining. In some embodiments, the peptide candidates can be printed using a solid-phase photolithography developed based on solid-phase peptide synthesis (SPPS) or liquid-phase peptide synthesis (LPPS). For example, the chip is pre-coated with stable linkers. Each stable linker may be a photo-labile linker that is protected at its free terminus (e.g., N-terminus or C-terminus) with a photo-labile protecting group (PPG), which PPG can be activated by light (e.g., laser) to deprotect the free terminus. The deprotected free terminus can then be utilized as a conjugation site for an amino acid or peptide molecule to be conjugated, to initiate printing of the patterned array of candidate peptide molecules. The conjugated amino acid or peptide molecule is thereby immobilized to the chip. The free terminus of the immobilized amino acid or peptide molecule can have a free terminus protected by a PPG. Similar to the PPG of the stable linker, the PPG of the immobilized amino acid or peptide can be activated by light to deprotect the free terminus, which can be utilized as the conjugation site for subsequent amino acids or peptide molecules. The conjugation of amino acids or peptide molecules to the free terminus of the immobilized amino acid or peptide can be repeated one or more times, thereby allowing the spatially controlled assembly of the patterned array of peptide molecules.

[0206] Alternatively, the stable linker, as provided herein, may not be protected at its free terminus. In such case, a PPG at a free terminus of a first immobilized amino acid or peptide can be utilized as a conjugation site for subsequent amino acids or peptide molecules, as provided above.

[0207] In some cases, a stable linker can comprise a PPG as provided herein (e.g., NVOC), to provide spatially resolved synthesis using light-directed deprotection.

[0208] Light (e.g., laser) is utilized to activate the PPGs. The laser can comprise the wavelength range between about 405 nanometers and about 532 nanometers, which wavelength range can depend on the specific type of the PPGs. Non-limiting examples of the PPGs includes o-nitrobenzyl derivatives (e.g., nitroveratryloxycarbonyl or “NVOC”, 2-(2-Nitrophenyl)- propyloxycarbonyl or “NPPOC”, 2,3-dimethyl-2,3-dinitrobutane or “DMNB”, etc.), p-Hydroxyphenacyl (pHP), or coumarin-4-ylmethyl groups. The PPGs can provide wavelengthspecific deprotection of immobilized amino acids or peptide molecules for conjugation of subsequent amino acid or peptide molecules. Upon activation of the PPG at a free terminus of an immobilized amino acid or peptide molecule, a desired amino acid is conjugated to the activated portion of the PPG. Each amino acid or peptide molecule is stored in individual sealed reservoirs as protected amino acids in a solvent (e.g., DMSO), delivered via automated microfluidic flow to defined printing zones. The protected amino acids can be protected with one or more members comprising an a-amino protecting group (e.g., Fmoc, Boc, Z), a side chain protecting group (e.g., Boc, tBu, OtBu, Trt, Pbf), and / or a C-terminal protecting group (e.g., a methoxy group such as OMe). In some cases, the C-terminus of the amino acid may not be protected (e.g., SPPS-like synthesis). Alternatively, the C-terminus of the amino acid may be protected by an activatable or active ester such as, for example, N-hydroxysuccinimide (NHS) esters or “Osu” (e.g., LPPS-like synthesis).

[0209] An SPPS-based protocol can comprise using a stable linker to immobilize peptides to the chip. For a glass substrate, aminopropyltriethoxysilane (APTES) can be added to generate stable amine S-Adenosyl methionines (SAMs) via siloxane bonding. Alternatively, for a metal substrate (e.g., gold), aminothiophenol can be added to form thiol-SAMs with terminal amines. The surface of the chip is thereby prepared to comprise an immobilized protected amine or carboxylic group. Light (e.g., laser) can be utilized to provide site-specific deprotection during one or more amino acid coupling cycles. During an amino acid coupling cycle, the surface can be selectively exposed to light in precise regions using an optical pickup unit (OPU)-based laser system. Exposure to light having a wavelength between about 365 and about 405 nm can remove a PPG (e.g., NVOC or pHP) at a free terminus, causing the free terminus to be deprotected. The amino acid coupling cycle can comprise adding an activated protected amino acid (e.g., an amino acid protected by Boc at the N-terminus and / or protected by Osu or NHS at the C-terminus). The activated amino acid can form a peptide bond at the deprotected free terminus, allowing for selective amino acid or peptide molecule conjugation at deprotected sites. The amino acid coupling cycle can be repeated one or more times, until the desired peptide molecules are synthesized. The methods for printing the patterned array of peptide molecules, as provided herein, eliminate the need for harsh reagents (e.g., piperidine for Fmoc) and are compatible with dry, digitally controlled peptide synthesis.

[0210] In some cases, peptide candidates can be immobilized (or conjugated) to a linker layer disposed on the substrate. The linker layer can comprise a self-assembled monolayer (SAM), e.g., comprised of aminopropyltriethoxysilane (APTES) or aminothiolphenol. Theimmobilization of the peptide candidates onto the chip surface can comprise forming an amide bond (or a peptide bond) between a carboxylic- or amine-terminated SAM and at least a portion of a peptide sequence (e.g., the initial amino acid residue of the peptide sequence) via a linker chemistry, e.g., NHS-ester or carbodiimide chemistry.

[0211] Example 7: Metal-binding peptide molecules

[0212] As provided herein, peptide molecules can be utilized for biomining (e.g., binding and extracting) target metals. In some cases, some peptide molecules can be utilized as basis to generate additional peptide variants to further test as candidates for metal biomining.

[0213] Non-limiting examples of gallium-binding polypeptide sequences can include TMHHAAIAHPPH (SEQ ID NO: 1), NYLPHQSSSPSR (SEQ ID NO: 2), SQALSTSRQDLR (SEQ ID NO: 3), HTQHIQSDDHLA (SEQ ID NO: 4), and NDLQRHRLTAG (SEQ ID NO: 5). Non-limiting examples of lithium -binding polypeptide sequences can include GPGDP (SEQ ID NO: 6), GPGAP (SEQ ID NO: 7), GPGNP (SEQ ID NO: 8), and repeats or combinations thereof such as GPGDPGPGDPGPGDP (SEQ ID NO: 9) or GPGDPEAAAKGPGDPEAAAKGPGDP (SEQ ID NO: 10). Non-limiting examples of molybdenite (M0S2) binding polypeptide sequences can include GVIHRNDQWTAPGGG (SEQ ID NO: 11) and DRWVARDPASIFGGG (SEQ ID NO: 12). Additional examples of metal-binding peptides can include TNTLSNN (SEQ ID NO: 13) for lead, CTQMLGQLC (SEQ ID NO: 14) for cobalt, and CNAKHHPRC (SEQ ID NO: 15) for nickel.

[0214] Example 8: In silica modeling methodology

[0215] As provided herein, the machine learning algorithm can be configured to compare (i) the measured metal binding performance (e.g., by a sensor as described throughout) and (ii) an in silico predicted metal binding performance. The in silico predicted metal binding performance can be provided (e.g., generated) via in silico modeling.

[0216] In some embodiments, such in silico modeling can utilize (e.g., combine) classical simulation, quantum-level simulation, or both. In some cases, the in silico modeling can combine both classical and quantum-level simulations. For example, molecular dynamics (MD) simulations of peptides and related biomolecules (e.g., for metal binding) can be performed using AMBER and OpenMM. Modeling their interactions with mineral surfaces (or metals) can be achieved via a hybrid ab initiol classical approach that integrates Density Functional Theory (DFT) calculations via VASP with classical MD simulations in LAMMPS. This multiscale method may provide accurate prediction of binding energetics and conformational behavior atcomplex bio-inorganic interfaces.

[0217] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.EMBODIMENTS

[0218] The following non-limiting embodiments provide illustrative examples of the invention, but do not limit the scope of the invention.

[0219] Embodiment 1. A system for biomining a target metal from a medium obtained from a geological sample, the system comprising: a substrate comprising a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and a sensor configured to measure a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or another geological sample, optionally wherein:(1) the geological sample and the another geological sample are from a same geological resource; and / or(2) the geological sample comprises ores, concentrates, or waste materials; and / or(3) the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium, further optionally wherein:(i) the target metal comprises copper; and / or(ii) the target metal comprises lithium; and / or(4) the at least one optimal peptide molecule is identified for biomining the target metal in absence of a living organism for extraction of the target metal; and / or(5) the medium has a pH that is less than about 7, ranges between about 2 and about6, or ranges between about 3 and about 5; and / or(6) the medium comprises one or more metal chelating agents; and / or(7) the medium comprises a plurality of mineral particles comprising the target metal, wherein (i) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (ii) the plurality of minerals has an average particle size of at least about 1 micrometer, at least about 5 micrometers, or at least about 10 micrometers; and / or(8) the measurement of the mental binding performance is non-fluorescence-based; and / or(9) the sensor comprises a surface plasmon resonance (SPR) sensor; and / or(10) the sensor comprises an optical source configured to direct a light comprising a single wavelength or a range of wavelengths to another surface of the substrate that is opposite of the surface, further optionally wherein:(i) the light comprises a laser light; and / or(ii) the sensor further comprises a detector configured to detect at least a portion of the light that is reflected from the another surface, to measure the metal binding performance; and / or(11) the metal binding performance comprises association between the target metal and each of the plurality of candidate peptide molecules; and / or(12) the metal binding performance comprises dissociation between the target metal and each of the plurality of candidate peptide molecules; and / or(13) the metal binding performance comprises the association and the dissociation, wherein the association is assigned a higher weight value than the dissociation, further optionally wherein a ratio of the assigned weight values to the association and the dissociation is between about 60:40 and about 90: 10, between about 60:40 and about 80:20, or about 70:30; and / or(14) the system further comprises a processor configured to execute an algorithm to identify the at least one optimal peptide molecule from the plurality of candidate peptide molecules based at least in part on the metal binding performance; and / or(15) the processor is configured to execute the algorithm or another algorithm to rank the plurality of candidate peptide molecules based at least in part on their respective metal binding performance, further optionally wherein the processor is configured execute the algorithm or another algorithm to select the at least one optimal peptide molecule from the top 50%, 40%, 30%, 20%, or 10% of the ranked candidate peptide molecules; and / or(16) the processor is configured execute the algorithm or another algorithm to design in silico at least one peptide variant comprising at least one amino acid modification as compared to the amino acid sequence of the at least one optimal peptide molecule, further optionally wherein:(i) the in silico design of the at least one peptide variant is based at least in part on a site-directed mutagenesis; and / or(ii) the at least one amino acid modification comprises at least 2, at least 3, at least 4, or at least 5 amino acid modifications; and / or(iii) the at least one amino acid modification comprises at most 5, at most 4, at most 3, at most 2, or at most 1 amino acid modification(s); and / or(17) the processor is configured to execute the algorithm or another algorithm to (i) determine a rate and / or yield of extraction of the target metal of one or more candidate peptide molecules and (ii) identify the at least one optimal peptide molecule based on the determined rate and / or yield of extraction; and / or(18) the processor is configured to execute a machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to identify the at least one optimal peptide molecule; and / or(19) the processor is configured to execute the machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to design in silico the at least one peptide variant, further optionally wherein the machine learning algorithm is trained to (1) generate a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule, (2) utilize graph neural network (GNN) to analyze the plurality of kinetic graph embeddings and identify one or more amino acid residues or positions thereof that are associated with the association and dissociation between the target metal and a candidatepeptide molecule, and (3) select the identified one or more amino acid residues or positions thereof for mutation to generate the at least one peptide variant; and / or(20) the machine learning algorithm is configured to analyze the association and / or dissociation between the target metal and each of the plurality of candidate peptide molecules, to identify a pattern of a molecular interaction between the target metal and the plurality of candidate peptide molecules, further optionally wherein the machine learning algorithm is trained to identify the pattern via (1) generation of a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule and (2) comparison of the plurality of kinetic graph embeddings, further optionally wherein the machine learning algorithm is further trained to identify the pattern via (2) clustering of two or more kinetic graph embeddings from the plurality of kinetic graph embeddings based on similarities and / or differences between the plurality of kinetic graph embeddings and / or (3) identification of one or more kinetic graph embeddings exhibiting an anomaly; and / or(21) the machine learning algorithm is trained based at least in part on supervised learning using a labeled dataset of known interactions between the target metal and the plurality of candidate peptide molecules or other control peptide molecules; and / or(22) the machine learning algorithm is trained based at least in part on unsupervised learning to detect a hidden or unknown trend from the interactions between the target metal and the plurality of candidate peptide molecules; and / or(23) the machine learning algorithm comprises reinforcement learning to dynamically adjust one or more performance characteristics in the analysis of the collection of the metal binding performance, the identification of the at least one optimal peptide molecule, and / or the in silico design of the least one peptide variant; and / or(24) the machine learning algorithm is configured to compare (i) the measured metal binding performance by the sensor and (ii) an in silico predicted metal binding performance, to improve its capability in predicting a metal binding performance of a candidate peptide molecule sequence; and / or(25) the machine learning algorithm is part of a self-learning codebase that is configured to improve its performance in (i) the identification of the at least one optimal peptide molecule and / or (ii) the design of the at least one peptide variant; and / or(26) the substrate comprises a metal layer, further optionally wherein:(A) the metal layer comprises a noble metal; and / or(B) the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer, further optionally wherein the substrate further comprises (i) an adhesion layer disposed between the transparent layer and the metal layer, (ii) an immobilization layer configured to couple to the plurality of candidate peptide molecules, or both (i) and (ii), further optionally wherein the metal layer further comprises a dielectric layer; and / or(27) a position of each of the plurality of candidate peptide molecules on the patterned array is spatially controlled; and / or(28) the plurality of candidate peptide molecules comprises at least about 10, at least about 20, at least about 50, or at least about 100 different peptide molecules; and / or(29) each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues; and / or(30) each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues; and / or(31) the plurality of candidate peptide molecules does not comprise an active enzyme; and / or(32) the medium comprises a solution.

[0220] Embodiment 2. A method for biomining a target metal from a medium obtained from a geological sample, the method comprising:(a) contacting the medium with a substrate, wherein the substrate comprises a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and(b) measuring, via a sensor, a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or an additional geological sample, optionally wherein:(1) the contacting and the measuring are performed at or near a same mining site asthe geological sample; and / or(2) the geological sample and the another geological sample are from a same geological resource; and / or(3) the geological sample comprises ores, concentrates, or waste materials; and / or(4) the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium; and / or(5) the method comprises identifying at least one optimal peptide molecule for biomining the target metal in absence of a living organism for extraction of the target metal; and / or(6) the method comprises, prior to (a), preparing the medium comprising the target metal at a pH that is less than about 7, ranges between about 2 and about 6, or ranges between about 3 and about 5; and / or(7) the method comprises, prior to (a), adding one or more metal chelating agents; and / or(8) the method comprises adding a plurality of mineral particles comprising the target metal to the medium, wherein (1) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (2) the plurality of minerals has an average particle size of at least about 1 micrometer, at least about 5 micrometers, or at least about 10 micrometers; and / or(9) the measuring of the mental binding performance is non-fluorescence-based; and / or(10) the measuring comprises detecting surface plasmon resonance (SPR) of a respective region of the surface adjacent to a peptide molecule of the array of peptide molecules; and / or(11) the measuring comprises directing a light comprising a single wavelength or a range of wavelengths to another surface of the substrate that is opposite of the surface, further optionally wherein:(i) the light comprises a laser light; and / or(ii) the measuring further comprises detecting at least a portion of the light that is reflected from the another surface, to measure the metal binding performance; and / or(12) the metal binding performance comprises association between the target metal and each of the plurality of candidate peptide molecules; and / or(13) the metal binding performance comprises dissociation between the target metal and each of the plurality of candidate peptide molecules; and / or(14) the method further comprises identifying the at least one optimal peptide molecule from the plurality of candidate peptide molecules based at least in part on the metal binding performance; and / or(15) the method further comprises ranking the plurality of candidate peptide molecules based at least in part on their respective metal binding performance, further optionally wherein the method comprises selecting the at least one optimal peptide molecule from the top 50%, 40%, 30%, 20%, or 10% of the ranked candidate peptide molecules; and / or(16) the method further comprises designing in silico at least one peptide variant comprising at least one amino acid modification as compared to the amino acid sequence of the at least one optimal peptide molecule; and / or further optionally wherein the in silico design of the at lease tone peptide variant is based at least in part on a site-directed mutagenesis; and / or(17) the method further comprises (i) determining a rate and / or yield of extraction of the target metal of one or more candidate peptide molecules and (ii) identifying the at least one optimal peptide molecule based on the determined rate and / or yield of extraction; and / or(18) the method further comprises analyzing, via a machine learning algorithm, a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to identify the at least one optimal peptide molecule; and / or(19) the method further comprises analyzing, via the machine learning algorithm, a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to design in silico the at least one peptide variant, further optionally wherein the method comprises utilizing the machine learning algorithm to (1) generate a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule, (2) utilize graph neural network (GNN) to analyze the plurality of kinetic graph embeddings and identify one or more amino acid residues or positions thereof that are associated with the association and dissociation between the target metal and a candidate peptide molecule, and (3) select the identified one or more amino acid residues or positions thereof for mutation to generate the at least one peptide variant;and / or(20) the method further comprises analyzing, via the machine learning algorithm, the association and / or dissociation between the target metal and each of the plurality of candidate peptide molecules, to identify a pattern of a molecular interaction between the target metal and the plurality of candidate peptide molecules, further optionally wherein the method comprises utilizing the machine learning algorithm to identify the pattern via (1) generating a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule and (2) comparing the plurality of kinetic graph embeddings, further optionally wherein the method comprises utilizing the machine learning algorithm to identify the pattern via one or more members comprising (2) clustering of two or more kinetic graph embeddings from the plurality of kinetic graph embeddings based on similarities and / or differences between the plurality of kinetic graph embeddings, and / or (3) identification of one or more kinetic graph embeddings exhibiting an anomaly; and / or(21) the method further comprises training the machine learning algorithm based at least in part on supervised learning using a labeled dataset of known interactions between the target metal and the plurality of candidate peptide molecules or other control peptide molecules; and / or(22) the method further comprises training the machine learning algorithm based at least in part on unsupervised learning to detect a hidden or unknown trend from the interactions between the target metal and the plurality of candidate peptide molecules; and / or(23) the method further comprises using reinforcement learning to dynamically adjust one or more performance characteristics of the machine learning algorithm in the analysis of the collection of the metal binding performance, the identification of the at least one optimal peptide molecule, and / or the in silico design of the least one peptide variant; and / or(24) the method further comprises comparing (i) the measured metal binding performance by the sensor and (ii) an in silico predicted metal binding performance, to improve its capability in predicting a metal binding performance of a candidate peptide molecule sequence; and / or(25) wherein the machine learning algorithm is part of a self-learning codebase that isconfigured to improve its performance in (i) the identification of the at least one optimal peptide molecule and / or (ii) the design of the at least one peptide variant; and / or(26) the substrate comprises a metal layer, further optionally wherein:(A) the metal layer comprises a noble metal; and / or(B) the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer, further optionally wherein the substrate further comprises (i) an adhesion layer disposed between the transparent layer and the metal layer, (ii) an immobilization layer configured to couple to the plurality of candidate peptide molecules, or both (i) and (ii), further optionally wherein the metal layer further comprises a dielectric layer; and / or(27) a position of each of the plurality of candidate peptide molecules on the patterned array is spatially controlled; and / or(28) the plurality of candidate peptide molecules does not comprise an active enzyme; and / or(29) the medium comprises a solution.

[0221] Embodiment 3. A method for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the method comprising:(a) providing at least one amino acid or at least one peptide; and(b) directing, via an optical source of a printer, a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of the at least one amino acid or the at least one peptide on or adjacent to the surface of the substrate, thereby printing the patterned array of peptide molecules on the surface, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample, optionally wherein:(1) the wavelength or the range of wavelengths of the light is sufficient to activate the surface or an amino acid bound to the surface for the conjugation, further optionally wherein the surface or the amino acid comprises a photo-labile protecting group, wherein the light cleaves the photo-labile protecting group to activatethe surface or the amino acid; and / or(2) the light comprises a laser light, further optionally wherein the laser light has a wavelength range between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm; and / or(3) the method further comprises controlling a target region of the light or the laser light within the surface of the substrate, thereby selectively effecting the conjugation within the target region, further optionally wherein the method comprises instructing an optical pick-up unit (OPU) to control the target region of the light or the laser light; and / or(4) the method further comprises contacting the surface with a medium comprising at least one amino acid or the at least one peptide, further optionally wherein the method comprises directing flow of the medium from a source of the medium and towards the surface; and / or(5) the method further comprises selecting or designing the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a type of the geological sample; and / or(6) the method further comprises selecting or designing the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a geolocation of the geological sample; and / or(7) the method further comprises designing in silico the plurality of candidate peptide molecules based on data associated with a metal binding performance of another patterned array of peptide molecules comprising another plurality of candidate peptide molecules, wherein the plurality of candidate peptide molecules and the another plurality of candidate peptide molecules are different, further optionally wherein the metal binding performance of the another patterned array of peptide molecules is based on binding and extracting the target metal from (i) the same medium, (ii) another medium obtained from the same geological sample, and / or (iii) another medium obtained from another geological sample from a same geological site as the geological sample; and / or(8) the providing, the directing, and / or the printing is performed at or near a same mining site as the geological sample; and / or(9) the geological sample comprises ores, concentrates, or waste materials; and / or(10) the target metal is selected from the group consisting of copper, lithium, cobalt,-n-tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium, further optionally wherein:(i) the target metal comprises copper; and / or(ii) the target metal comprises lithium; and / or(11) the method further comprises identifying at least one optimal peptide molecule for biomining the target metal in absence of a living organism for extraction of the target metal; and / or(12) the substrate comprises a metal layer, further optionally wherein:(i) the metal layer comprises a noble metal; and / or(ii) the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer; and / or(13) the method further comprises spatially controlling a position of each of the plurality of candidate peptide molecules on the patterned array during the printing; and / or(14) the plurality of candidate peptide molecules comprises at least about 10, at least about 20, at least about 50, or at least about 100 different peptide molecules; and / or(15) each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues; and / or(16) each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues; and / or(17) the plurality of candidate peptide molecules does not comprise an active enzyme; and / or(18) the medium comprises a solution.

[0222] Embodiment 4. A system for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the system comprising: a printer comprising: an optical source configured to direct a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of at least one amino acid or at least one peptide one or adjacent to the surface of the substrate; and a controller configured to instruct the optical source to direct the light towards the surface to print, via the conjugation, an array of peptide molecules on the surface,wherein the array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample, optionally wherein:(1) the wavelength or the range of wavelengths of the light is sufficient to activate the surface or an amino acid bound to the surface for the conjugation, further optionally wherein the surface or the amino acid comprises a photo-labile protecting group that is cleavable by the light; and / or(2) the light comprises a laser light, further optionally wherein the laser light has a wavelength range between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm; and / or(3) the controller is configured to control a target region of the light or the laser light within the surface of the substrate, to selectively effect the conjugation within the target region, further optionally wherein the optical source comprises an optical pick-up unit (OPU) to control the target region of the light or the laser light; and / or(4) the system further comprises a container for holding a medium comprising at least one amino acid or the at least one peptide, further optionally wherein the system further comprises a flow controller configured to direct flow of at least a portion of the medium comprising at least one amino acid or the at least one peptide towards the surface; and / or(5) the processor is further configured execute an algorithm to select or design the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a type of the geological sample; and / or(6) the processor is further configured to execute the algorithm or another algorithm to select or design the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a geolocation of the geological sample; and / or(7) the processor is further configured to execute the algorithm or another algorithm to design in silico the plurality of candidate peptide molecules based on data associated with a metal binding performance of another patterned array of peptide molecules comprising another plurality of candidate peptide molecules, wherein the plurality of candidate peptide molecules and the another plurality of candidate peptide molecules are different; and / or(8) the metal binding performance of the another patterned array of peptide molecules is based on binding and extracting the target metal from (i) the same medium, (ii) another medium obtained from the same geological sample, and / or (iii) another medium obtained from another geological sample from a same geological site as the geological sample; and / or(9) the geological sample comprises ores, concentrates, or waste materials; and / or(10) the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium; and / or(11) the substrate is substantially free of a living organism for extraction of the target metal; and / or(12) the substrate comprises a metal layer, further optionally wherein the metal layer comprises a noble metal, further optionally wherein the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer; and / or(13) the plurality of candidate peptide molecules comprises at least about 10, at least about 20, at least about 50, or at least about 100 different peptide molecules; and / or(14) each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues; and / or(15) each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues; and / or(16) the plurality of candidate peptide molecules does not comprise an active enzyme; and / or(17) the medium comprises a solution.

[0223] Embodiment 5. A system for detecting a presence of a sample on a microarray, comprising: a 3D printer, comprising: a light engine; and an optical pick-up unit; a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate;a reader to capture kinetic data for each of the one or more immobilized molecules; and an artificial intelligence (Al) algorithm to generate a model based on codebase, optionally wherein:(1) the surface of the first side of the substrate is a photosensitive surface; and / or(2) the reader is based on at least surface plasmon resonance; and / or(3) the molecule is a biomolecule.

[0224] Embodiment 6. A system for detecting a presence of a sample on a microarray, comprising: a 3D printer, comprising: a light engine; an optical pick-up unit; and a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate; a reader to capture endpoint measurements for each of the one or more immobilized molecules; and an artificial intelligence (Al) algorithm to generate a model based on codebase, optionally wherein:(1) the reader is based on at least a fluorescent microarray scanner; and / or(2) the molecule is a biomolecule.

[0225] Embodiment 7. A method for detecting a presence of a sample on a microarray, comprising: providing a substrate; positioning one or more molecule variants on a surface of the substrate; performing analyte perfusion; capturing kinetic data for the one or more molecule variants; analyzing the captured kinetic data; employing an artificial intelligence (Al) algorithm on the captured kinetic data; generating one or more patterns from a plurality of molecular interactions of the one or more molecule variants; and generating a codebase based on one or more patterns, optionally wherein the molecule is a biomolecule.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A system for biomining a target metal from a medium obtained from a geological sample, the system comprising: a substrate comprising a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and a sensor configured to measure a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or another geological sample.

2. The system of claim 1, wherein the geological sample and the another geological sample are from a same geological resource.

3. The system of any one of the preceding claims, wherein the geological sample comprises ores, concentrates, or waste materials.

4. The system of any one of the preceding claims, wherein the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium.

5. The system of claim 4, wherein the target metal comprises copper.

6. The system of claim 4, wherein the target metal comprises lithium.

7. The system of any one of the preceding claims, wherein the at least one optimal peptide molecule is identified for biomining the target metal in absence of a living organism for extraction of the target metal.

8. The system of any one of the preceding claims, wherein:(i) the medium has a pH that is less than about 7, ranges between about 2 and about 6, or ranges between about 3 and about 5;(ii) the medium comprises one or more metal chelating agents; and / or(iii) the medium comprises a plurality of mineral particles comprising the target metal, wherein (1) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (2) the plurality of minerals has an average particle size of at least about 1 micrometer, at least about 5 micrometers, or at least about 10 micrometers.

9. The system of any one of the preceding claims, wherein the measurement of the mental binding performance is non-fluorescence-based.

10. The system of any one of the preceding claims, wherein the sensor comprises a surface plasmon resonance (SPR) sensor.

11. The system of any one of preceding claims, wherein the sensor comprises an optical source configured to direct a light comprising a single wavelength or a range of wavelengths to another surface of the substrate that is opposite of the surface.

12. The system of claim 11, wherein the light comprises a laser light.

13. The system of claim 11, wherein the sensor further comprises a detector configured to detect at least a portion of the light that is reflected from the another surface, to measure the metal binding performance.

14. The system of any one of the preceding claims, wherein the metal binding performance comprises association between the target metal and each of the plurality of candidate peptide molecules.

15. The system of any one of the preceding claims, wherein the metal binding performance comprises dissociation between the target metal and each of the plurality of candidate peptide molecules.

16. The system of any one of the preceding claims, wherein the metal binding performance comprises the association and the dissociation, wherein the association is assigned a higher weight value than the dissociation.

17. The system of claim 16, wherein a ratio of the assigned weight values to the association and the dissociation is between about 60:40 and about 90: 10, between about 60:40 and about 80:20, or about 70:30.

18. The system of any one of the preceding claims, further comprising a processor configured to execute an algorithm to identify the at least one optimal peptide molecule from the plurality of candidate peptide molecules based at least in part on the metal binding performance.

19. The system of any one of the preceding claims, wherein the processor is configured to execute the algorithm or another algorithm to rank the plurality of candidate peptide molecules based at least in part on their respective metal binding performance.

20. The system of claim 19, wherein the processor is configured execute the algorithm or another algorithm to select the at least one optimal peptide molecule from the top 50%, 40%, 30%, 20%, or 10% of the ranked candidate peptide molecules.

21. The system of any one of the preceding claims, wherein the processor is configured execute the algorithm or another algorithm to design in silico at least one peptide variant comprising at least one amino acid modification as compared to the amino acid sequence of the at least one optimal peptide molecule.

22. The system of claim 21, wherein the in silico design of the at least one peptide variant is based at least in part on a site-directed mutagenesis.

23. The system of claim 21, wherein the at least one amino acid modification comprises at least 2, at least 3, at least 4, or at least 5 amino acid modifications.

24. The system of claim 21, wherein the at least one amino acid modification comprises at most 5, at most 4, at most 3, at most 2, or at most 1 amino acid modification(s).

25. The system of any one of the preceding claims, wherein the processor is configured to execute the algorithm or another algorithm to (i) determine a rate and / or yield of extraction of the target metal of one or more candidate peptide molecules and (ii) identify the at least one optimal peptide molecule based on the determined rate and / or yield of extraction.

26. The system of any one of the preceding claims, wherein the processor is configured to execute a machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to identify the at least one optimal peptide molecule.

27. The system of any one of the preceding claims, wherein the processor is configured to execute the machine learning algorithm to analyze a collection of the metal binding performance and the respective amino acid sequences of the candidate peptide molecules, to design in silico the at least one peptide variant.

28. The system of claim 27, wherein the machine learning algorithm is trained to (1) generate a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule, (2) utilize graph neural network (GNN) to analyze the plurality of kinetic graph embeddings and identify one or more amino acid residues or positions thereof that are associated with the association and dissociation between the target metal and a candidate peptide molecule, and (3) select the identified one or more amino acid residues or positions thereof for mutation to generate the at least one peptide variant.

29. The system of any one of the preceding claims, wherein the machine learning algorithm is configured to analyze the association and / or dissociation between the target metal and each of the plurality of candidate peptide molecules, to identify a pattern of a molecular interaction between the target metal and the plurality of candidate peptide molecules.

30. The system of claim 29, wherein the machine learning algorithm is trained to identify the pattern via (1) generation of a plurality of kinetic graph embeddings, each graph embedding associated with association and / or dissociation between the target metal and a candidate peptide molecule and (2) comparison of the plurality of kinetic graph embeddings.

31. The system of claim 30, wherein the machine learning algorithm is further trained to identify the pattern via (2) clustering of two or more kinetic graph embeddings from the plurality of kinetic graph embeddings based on similarities and / or differences between the plurality of kinetic graph embeddings and / or (3) identification of one or more kinetic graph embeddings exhibiting an anomaly.

32. The system of any one of the preceding claims, wherein the machine learning algorithm is trained based at least in part on supervised learning using a labeled dataset of known interactions between the target metal and the plurality of candidate peptide molecules or other control peptide molecules.

33. The system of any one of the preceding claims, wherein the machine learning algorithm is trained based at least in part on unsupervised learning to detect a hidden or unknown trend from the interactions between the target metal and the plurality of candidate peptide molecules.

34. The system of any one of the preceding claims, wherein the machine learning algorithm comprises reinforcement learning to dynamically adjust one or more performance characteristics in the analysis of the collection of the metal binding performance, the identification of the at least one optimal peptide molecule, and / or the in silico design of the least one peptide variant.

35. The system of any one of the preceding claims, wherein the machine learning algorithm is configured to compare (i) the measured metal binding performance by the sensor and (ii) an in silico predicted metal binding performance, to improve its capability in predicting a metal binding performance of a candidate peptide molecule sequence.

36. The system of any one of the preceding claims, wherein the machine learning algorithm is part of a self-learning codebase that is configured to improve its performance in (i) the identification of the at least one optimal peptide molecule and / or (ii) the design of the at least one peptide variant.

37. The system of any one of the preceding claims, wherein the substrate comprises a metal layer.

38. The system of claim 37, wherein the metal layer comprises a noble metal.

39. The system of claim 37, wherein the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer.

40. The system of claim 39, wherein the substrate further comprises (i) an adhesion layer disposed between the transparent layer and the metal layer, (ii) an immobilization layer configured to couple to the plurality of candidate peptide molecules, or both (i) and (ii).

41. The system of claim 40, wherein the metal layer further comprises a dielectric layer.

42. The system of any one of the preceding claims, wherein a position of each of the plurality of candidate peptide molecules on the patterned array is spatially controlled.

43. The system of any one of the preceding claims, wherein the plurality of candidate peptide molecules comprises at least about 10, at least about 20, at least about 50, or at least about 100 different peptide molecules.

44. The system of any one of the preceding claims, wherein each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues.

45. The system of any one of the preceding claims, wherein each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues.

46. The system of any one of the preceding claims, wherein the plurality of candidate peptide molecules does not comprise an active enzyme.

47. A method for biomining a target metal from a medium obtained from a geological sample, the method comprising:(a) contacting the medium with a substrate, wherein the substrate comprises a patterned array of peptide molecules on a surface of the substrate, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the medium; and(b) measuring, via a sensor, a metal binding performance for each of the plurality of candidate peptide molecules, wherein the metal binding performance is usable to identify at least one optimal peptide molecule from the plurality of candidate peptide molecules for biomining the target metal from the geological sample or an additional geological sample.

48. The method of claim 47, wherein the contacting and the measuring are performed at or near a same mining site as the geological sample.

49. The method of any one of the preceding claims, wherein:(i) the method comprises, prior to (a), preparing the medium comprising the target metal at a pH that is less than about 7, ranges between about 2 and about 6, or ranges between about 3 and about 5;(ii) the method comprises, prior to (a), adding one or more metal chelating agents; and / or(iii) the method comprises adding a plurality of mineral particles comprising the target metal to the medium, wherein (1) the plurality of mineral particles is present in the medium at a range of between about 2 and about 20% or a range of between about 1 and about 10% (weight by volume) and / or (2) the plurality of minerals has an average particle size of at least about 1micrometer, at least about 5 micrometers, or at least about 10 micrometers.

50. A method for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the method comprising:(a) providing at least one amino acid or at least one peptide; and(b) directing, via an optical source of a printer, a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of the at least one amino acid or the at least one peptide on or adjacent to the surface of the substrate, thereby printing the patterned array of peptide molecules on the surface, wherein the patterned array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

51. The method of claim 50, wherein the wavelength or the range of wavelengths of the light is sufficient to activate the surface or an amino acid bound to the surface for the conjugation.

52. The method of claim 51, wherein the surface or the amino acid comprises a photo-labile protecting group, wherein the light cleaves the photo-labile protecting group to activate the surface or the amino acid.

53. The method of any one of claims 50-52, wherein the light comprises a laser light.

54. The method of claim 53, wherein the laser light has a wavelength range between about 350 nanometers (nm) and about 550 nm or between about 400 nm and about 550 nm.

55. The method of any one of claims 50-53, further comprising controlling a target region of the light or the laser light within the surface of the substrate, thereby selectively effecting the conjugation within the target region.

56. The method of claim 55, further comprising instructing an optical pick-up unit (OPU) to control the target region of the light or the laser light.

57. The method of any one of claims 50-56, further comprising contacting the surface with a medium comprising at least one amino acid or the at least one peptide.

58. The method of claim 57, further comprising directing flow of the medium from a source of the medium and towards the surface.

59. The method of any one of claims 50-58, further comprising selecting or designing the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a type of the geological sample.

60. The method of any one of claims 50-59, further comprising selecting or designing the plurality of candidate peptide molecules from a library of candidate peptide molecules based on a geolocation of the geological sample.

61. The method of any one of claims 50-60, further comprising designing in silico the plurality of candidate peptide molecules based on data associated with a metal binding performance of another patterned array of peptide molecules comprising another plurality of candidate peptide molecules, wherein the plurality of candidate peptide molecules and the another plurality of candidate peptide molecules are different.

62. The method of claim 59, wherein the metal binding performance of the another patterned array of peptide molecules is based on binding and extracting the target metal from (i) the same medium, (ii) another medium obtained from the same geological sample, and / or (iii) another medium obtained from another geological sample from a same geological site as the geological sample.

63. The method of any one of claims 50-62, wherein the providing, the directing, and / or the printing is performed at or near a same mining site as the geological sample.

64. The method of any one of claims 50-63, wherein the geological sample comprises ores, concentrates, or waste materials.

65. The method of any one of claims 50-64, wherein the target metal is selected from the group consisting of copper, lithium, cobalt, tantalum, indium, rhodium, platinum, palladium, gold, silver, neodymium, and dysprosium.

66. The method of claim 65, wherein the target metal comprises copper.

67. The method of claim 65, wherein the target metal comprises lithium.

68. The method of any one of claims 50-67, comprising identifying at least one optimal peptide molecule for biomining the target metal in absence of a living organism for extraction of the target metal.

69. The method of any one of claims 50-68, wherein the substrate comprises a metal layer.

70. The method of claim 69, wherein the metal layer comprises a noble metal.

71. The method of claim 69, wherein the substrate further comprises a transparent layer adjacent to the metal layer, wherein the transparent layer and the patterned array of peptide molecules are disposed on opposite sides of the metal layer.

72. The method of any one of claims 50-71, further comprising spatially controlling a position of each of the plurality of candidate peptide molecules on the patterned array during the printing.

73. The method of any one of claims 50-72, wherein the plurality of candidate peptide molecules comprises at least about 10, at least about 20, at least about 50, or at least about 100 different peptide molecules.

74. The method of any one of claims 50-73, wherein each peptide molecule of the plurality of candidate peptide molecules has a length of at most about 100, at most about 50, at most about 40, at most about 30, or at most about 20 amino acid residues.

75. The method of any one of claims 50-74, wherein each peptide molecule of the plurality of candidate peptide molecules has a length of at least about 5, at least about 10, or at least about 15 amino acid residues.

76. The method of any one of claims 50-75, wherein the plurality of candidate peptide molecules does not comprise an active enzyme.

77. A system for printing a patterned array of peptide molecules for biomining a target metal from a medium obtained from a geological sample, the system comprising: a printer comprising: an optical source configured to direct a light towards a surface of a substrate, wherein the light comprises a single wavelength or a range of wavelengths sufficient to effect conjugation of at least one amino acid or at least one peptide one or adjacent to the surface of the substrate; and a controller configured to instruct the optical source to direct the light towards the surface to print, via the conjugation, an array of peptide molecules on the surface, wherein the array of peptide molecules comprises a plurality of candidate peptide molecules for binding and extracting the target metal from the geological sample or another geological sample.

78. The system of claim 77, further comprising a container for holding a medium comprising at least one amino acid or the at least one peptide.

79. The system of claim 78, further comprising a flow controller configured to direct flow of at least a portion of the medium comprising at least one amino acid or the at least one peptide towards the surface.

80. A system for detecting a presence of a sample on a microarray, comprising: a 3D printer, comprising: a light engine; and an optical pick-up unit; a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate; a reader to capture kinetic data for each of the one or more immobilized molecules; and an artificial intelligence (Al) algorithm to generate a model based on codebase.

81. The system of claim 80, wherein the surface of the first side of the substrate is aphotosensitive surface.

82. The system of claim 80, wherein the reader is based on at least surface plasmon resonance.

83. The system of claim 80, wherein the molecule is a biomolecule.

84. A system for detecting a presence of a sample on a microarray, comprising: a 3D printer, comprising: a light engine; an optical pick-up unit; and a substrate with a first side and a second side, wherein one or more immobilized molecules disposed at a surface of the first side of the substrate; a reader to capture endpoint measurements for each of the one or more immobilized molecules; and an artificial intelligence (Al) algorithm to generate a model based on codebase.

85. The system of claim 84, wherein the reader is based on at least a fluorescent microarray scanner.

86. The system of any one of claims 80-85, wherein the molecule is a biomolecule87. A method for detecting a presence of a sample on a microarray, comprising: providing a substrate; positioning one or more molecule variants on a surface of the substrate; performing analyte perfusion; capturing kinetic data for the one or more molecule variants; analyzing the captured kinetic data; employing an artificial intelligence (Al) algorithm on the captured kinetic data; generating one or more patterns from a plurality of molecular interactions of the one or more molecule variants; and generating a codebase based on one or more patterns.

88. The method of claim 87, wherein the molecule is a biomolecule.

Citation Information

Patent Citations

  • Systems and methods for artificial intelligence-based prediction of amino acid sequences

    US20240038337A1

  • Function guided in silico protein design

    US20240087674A1

  • Method for screening peptides for metal coordinating properties and fluorescent chemosensors derived therefrom

    US6083758A

  • Peptide domains that bind small molecules of industrial significance

    US9447150B2