Excipients for self-assembling nanomaterials and methods of making and using same
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-04-09
AI Technical Summary
Existing PEGylated nanoparticles exhibit low drug loading, typically around 10% or less, due to reliance on nonspecific hydrophobic interactions, limiting their versatility and requiring multi-step synthesis.
Integration of drug-specific recognition sites into the hydrophobic part of the carrier, combined with machine learning models to predict physicochemical properties of nanoparticles, enabling improved drug loading and PEGylation.
Enhances drug loading to 50% to 95%, improves nanoparticle stability, and reduces macrophage uptake, with extended systemic circulation and reduced protein binding.
Smart Images

Figure US2025036200_09042026_PF_FP_ABST
Abstract
Description
[0001] EXCIPIENTS FOR SELF-ASSEMBLING NANOMATERIALS AND METHODS OF MAKING AND USING SAME CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No.63 / 667,200, filed on July 3, 2024, which is incorporated by reference herein in its entirety. FEDERALLY SPONSORED RESEARCH This invention was made with government support under grant number GM151255 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND The formulation of poorly water-soluble drugs into nanoparticles offers multiple advantages over their free drug counterparts, including increased bioavailability, reduced toxic side effects, and enhanced delivery efficiency. The incorporation of polyethylene glycol (PEG) into nanoparticles, widely known as PEGylation, is extensively used in both US Food and Drug Administration (FDA)-approved drug nanoparticles and numerous preclinical studies. PEGylation significantly affects tissue or cellular uptake routes, serum retention, organ distribution, and various pharmacokinetic and pharmacodynamic parameters. Nanoparticles with high drug loading provide multiple benefits, including reduced need for carrier materials, controlled drug release, and improved efficacy and safety. However, a significant challenge today is that most PEGylated nanoparticles exhibit low drug loading, around 10% or less. This limitation is thought to be due to the reliance on nonspecific hydrophobic interactions between the carrier and the drug. A proposed solution is the integration of drug-specific recognition sites into the hydrophobic part of the carrier, which aims to achieve both high drug loading and effective PEGylation. Nonetheless, this method is limited to specific drugs and requires a multi-step synthesis, underscoring the need for a more versatile design approach. What is needed are methods for identifying excipients for self-assembling nanomaterials, excipient, and pharmaceutical compositions thereof, modified excipients with improved protein binding and drug loading, and methods for treating subjects using such excipients and pharmaceutical compositions. SUMMARY One embodiment described herein is a computer-implemented method for training a machine learning model for predicting one or more physicochemical properties of a nanoparticle, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular representation and a property value of each nanoparticle, wherein each nanoparticle includes a drug-excipient pair including a drug molecule and an excipient molecule; generate, using the molecular representation, a descriptor and fingerprint for each nanoparticle of the set of data; creating, by using cross-validation method and the set of data, a plurality of sets of training data and a plurality of sets of test data, wherein the set of data includes the descriptor and fingerprint for each nanoparticle of the set of data; training a machine learning model of an artificial intelligence (AI) system using the plurality of sets of training data, wherein the plurality of sets of training data includes drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data; and for a drug-excipient pair of the plurality of sets of test data, predicting a property of the drug-excipient pair of the plurality of sets of test data using the machine learning model as trained based on drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data. In one aspect, the one or more physicochemical properties comprise nanoparticle formation, nanoparticle radius, or drug loading. In another aspect, the machine learning model is selected from a classification model or a regression model. In another aspect, the method further comprises: assessing a performance of the classification model, using classification metrics, and the regression model, using regression metrics, wherein classification metrics include at least one of an accuracy, a precision, a recall, an area under the Receiver Operating Characteristic curve (“AUC” ROC), a F1 score, and a Mathews Correlation Coefficient (“MCC”), and wherein the regression metrics include at least one of a root mean squared error (“RMSE”), a mean absolute error (“MAE”), and a coefficient of determination (R2). In another aspect, the predicted property of the drug-excipient pair of the plurality of sets of test data is selected from a nanoparticle formation, a nanoparticle radius, or a drug loading. In another aspect, the property value of each nanoparticle is a nanoparticle measurement selected from a radius, a polydispersity, or a ratio of normalized intensity. In another aspect, the cross-validation method is selected from a ten (10) fold cross- validation and a Leave-One-Out (“LOO”) cross-validation, wherein the LOO cross-validation includes a Leave-One-Excipient-Out (“LOEO”) cross-validation, a Leave-One-Drug-Out (“LODO”) cross-validation, or a Leave-One-Pair-Out (“LOPO”) cross-validation. Another embodiment described herein is a computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute a method described herein. Another embodiment described herein is a computer-implemented method for training a machine learning model for predicting a formulation improvement of new molecular component modifications of nanoparticles, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular component representation and a value selected from a known exact absolute size value or a known bound absolute size value, wherein the known exact absolute size value and the known bound absolute size value are related to a radius of each nanoparticle of the set of data, and wherein each nanoparticle of the set of data comprises a drug and an excipient; creating a set of training data with the set of data; creating a set of nanoparticle pairs using each nanoparticle of the set of training data; filtering the set of training data based on a set of rules, wherein the filtered set of training data includes nanoparticle pairs of the set of nanoparticle pairs where a nanoparticle of the nanoparticle pairs with a larger size is known; training a machine learning model of an AI system using datapoints of the filtered set of training data, wherein the datapoints include the nanoparticle pairs of the filtered set of training data with shared representations, and wherein the datapoints include physicochemical properties of the shared representations; and for datapoints of a pair of nanoparticles, predicting a classification of a formulation improvement of new molecular component modifications of nanoparticles using the machine learning model as trained based on size of the pair of nanoparticles for the datapoints, wherein the formulation improvement indicates at least one selected from a nanoparticle formation, a nanoparticle radius, or a drug loading for the pair of nanoparticles. In one aspect, filtering the set of training data based on the set of rules, further comprises: removing nanoparticle pairs of the set of training data from the set of training data where whether a first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is greater than the second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is less than the second nanoparticle, wherein the first nanoparticle has a known bound absolute size value a known exact absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle and the second nanoparticle have a known bound absolute size value, wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle and a second nanoparticle of the nanoparticle pairs of the set of training data have a known bound absolute size value. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is less than a second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is greater than a second nanoparticle, wherein the first nanoparticle has a known bound absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: splitting the set of training data into a training set and a test set. In another aspect, creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: generating a pair of nanoparticles for at least one of the training set and the test set by cross merging a first nanoparticle and a second nanoparticle of the training set and the test set, wherein all possible nanoparticle pairs of the training set and the test set are generated. In another aspect, the predicted classification is an output indicating whether a first nanoparticle or a second nanoparticle in the pair nanoparticles has a larger radius. Another embodiment described herein is a computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute a method described herein. Another embodiment described herein is an excipient conjugated to a polyethylene glycol moiety, or pharmaceutically acceptable salt thereof, selected from: Cholesterol-PEG,
[0002] Stearic acid-PEG, or
[0003] Congo Red-PEG. In one aspect, n is about 22 to about 225. Another embodiment described herein is a nanoparticle formulation comprising: an excipient described herein, or pharmaceutically acceptable salt thereof; and a therapeutic compound. In one aspect, the therapeutic compound is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4- hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, or analogues or combinations thereof. In another aspect, the nanoparticles have a hydrodynamic radius of about 10 nm to about 500 nm. In another aspect, the nanoparticles have a therapeutic compound loading value of about 50% to about 95%. In another aspect, the
[0004] or pharmaceutically acceptable salt thereof, and the nanoparticles have a therapeutic compound loading value of 77%. In another aspect, the nanoparticles have a stability greater than 12 hours. Another embodiment described herein is a method for treating a disease or disorder or prophylaxis thereof in a subject in need thereof, the method comprising administering a therapeutically effective amount of a nanoparticle formulation described herein to a subject in need thereof. Another embodiment described herein is a method for reducing macrophage uptake, the method comprising administering a therapeutically effective amount of a nanoparticle formulation described herein to a subject in need thereof. Another embodiment described herein is a method for reducing protein binding of a therapeutic compound, the method comprising administering an effective amount of a nanoparticle formulation described herein. Another embodiment described herein is the use of an excipient described herein, or pharmaceutically acceptable salt thereof, to modulate systemic circulation of a therapeutic compound. DESCRIPTION OF THE DRAWINGS The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. FIG.1A–D outline the machine learning model and methods described herein as well as traditional approaches. FIG. 1A schematically illustrates a system according to some embodiments. FIG.1B shows a traditional approach for a difference comparison. FIG.1C shows a pairing of nanoparticles for a size difference comparison approach as described herein. FIG 1D shows a basic workflow for implementing the prediction method described herein. FIG. 2 shows experimental data for nanoparticle forming excipient-drug combinations predicted using the method described herein. FIG.3A–B show the performance for pairing nanoparticles for size difference comparison. Performance using 2-fold cross validation with 5 repeats. FIG.3A shows the true label vs. the predicted label. FIG.3B shows a pie chart representing the data of FIG.3A. FIG.4 shows the results from predicted drug loading from random forest model trained onRDKit descriptors and assessed using Leave-One-Excipient-Out cross-validation for 18 -glycyrrhetinic acid-PEG, IR783-PEG, rhodamine (TAMRA)-PEG, and verteporfin-PEG. FIG. 5 shows a transmission electron microscope (TEM) imaging of carfilzomib / 18 -glycyrrhetinic acid-PEG nanoparticles. FIG. 6A–B shows PEGylated excipients can improve the stability of the nanoparticlescompared to non-PEGylated excipients. A representative example is 18 -glycyrrhetinic acid as aPEGylated and as a non-PEGylated excipient forming nanoparticles with fulvestrant (FIG. 6A) and celecoxib (FIG.6B). FIG.7 shows drug loading data for nanoparticles formed with 18 -glycyrrhetinic acid-PEG.FIG.8 shows the UV / Vis absorbance spectroscopy of IR783-PEG. FIG.9 shows an HPLC chromatogram of IR783-PEG. FIG. 10 shows a transmission electron microscope image of fulvestrant / IR783-PEG nanoparticles. FIG.11 shows the drug loading of IR783-PEG nanoparticles. FIG.12 shows serum protein binding for a fulvestrant / verteporfin-PEG nanoparticle versus non-functionalized fulvestrant / verteporfin nanoparticles. The fulvestrant / verteporfin nanoparticles binds to BSA, while the nanoparticles formulated with fulvestrant / verteporfin-PEG exhibit significantly reduced BSA binding and opsonization. FIG.13A shows bright field and fluorescent microscopy of macrophages incubated with a fulvestrant / verteporfin-PEG nanoparticle versus non-functionalized fulvestrant / verteporfin nanoparticles. The fulvestrant / verteporfin nanoparticles are taken up by the macrophages, where the uptake of the PEG-functionalized fulvestrant / verteporfin-PEG is significantly reduced. FIG. 13B shows the fluorescent intensity at 700 nm. FIG. 14 shows pharmacokinetic studies for fulvestrant co-administered with verteporfin- PEG or verteporfin in mice. Fulvestrant levels were measured at multiple time points post- administration. Data are presented as individual points per animal (n = 4). The fulvestrant / verteporfin-PEG formulation demonstrated a slower decline in plasma concentrations, consistent with an extended terminal half-life and reduced clearance compared to the non- PEGylated formulation. DETAILED DESCRIPTION Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For example, any nomenclatures used in connection with, and techniques of biochemistry, molecular biology, immunology, microbiology, genetics, cell and tissue culture, and protein and nucleic acid chemistry described herein are well known and commonly used in the art. In case of conflict, the present disclosure, including definitions, will control. Exemplary methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in practice or testing of the embodiments and aspects described herein. The embodiments described herein are not limited in application to the details of the configurations and arrangements of components set forth in the following description or illustrated in the accompanying drawings. The embodiments are capable of being practiced or of being carried out in various ways. Also, the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Embodiments described herein may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more processing units, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, “servers,” “computing devices,” “controllers,” “processors,” etc., described in the specification can include one or more processing units, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the components. Although certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes only. Functionality described herein as being performed by one component may be performed by multiple components in a distributed manner. Likewise, functionality performed by multiple components may be consolidated and performed by a single component. In some embodiments, the illustrated components may be combined or divided into separate software, firmware, and / or hardware. For example, instead of being located within and performed by a single electronic processor, logic and processing may be distributed among multiple electronic processors. In some embodiments, a cloud-based computing infrastructure may be used in which the various components may be geographically separated between one or more locations. Regardless of how they are combined or divided, hardware and software components may be located on the same computing device or may be distributed among different computing devices connected by one or more networks or other suitable communication links. Similarly, a component described as performing particular functionality may also perform additional functionality not described herein. For example, a device or structure that is “configured” in a certain way is configured in at least that way but may also be configured in ways that are not explicitly listed. As used herein, the terms “amino acid,” “nucleotide,” “polynucleotide,” “vector,” “polypeptide,” and “protein” have their common meanings as would be understood by a biochemist of ordinary skill in the art. Standard single letter nucleotides (A, C, G, T, U) and standard single letter amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, or Y) are used herein. As used herein, terms such as “include,” “including,” “contain,” “containing,” “having,” and the like mean “comprising.” The present disclosure also contemplates other embodiments “comprising,” “consisting essentially of,” and “consisting of” the embodiments or elements presented herein, whether explicitly set forth or not. As used herein, “comprising,” is an “open- ended” term that does not exclude additional, unrecited elements or method steps. As used herein, “consisting essentially of” limits the scope of a claim to the specified materials or steps and those that do not materially affect the basic and novel characteristics of the claimed invention. As used herein, “consisting of” excludes any element, step, or ingredient not specified in the claim. As used herein, the term “a,” “an,” “the” and similar terms used in the context of the disclosure (especially in the context of the claims) are to be construed to cover both the singular and plural unless otherwise indicated herein or clearly contradicted by the context. In addition, “a,” “an,” or “the” means “one or more” unless otherwise specified. As used herein, the term “or” can be conjunctive or disjunctive. As used herein, the term “and / or” refers to both the conjunctive and disjunctive. As used herein, the term “substantially” means to a great or significant extent, but not completely. As used herein, the term “about” or “approximately” as applied to one or more values of interest, refers to a value that is similar to a stated reference value, or within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, such as the limitations of the measurement system. In one aspect, the term “about” refers to any values, including both integers and fractional components that are within a variation of up to ± 10% of the value modified by the term “about.” Alternatively, “about” can mean within 3 or more standard deviations, per the practice in the art. Alternatively, such as with respect to biological systems or processes, the term “about” can mean within an order of magnitude, in some embodiments within 5-fold, and in some embodiments within 2-fold, of a value. As used herein, the symbol “~” means “about” or “approximately.” All ranges disclosed herein include both end points as discrete values as well as all integers and fractions specified within the range. For example, a range of 0.1–2.0 includes 0.1, 0.2, 0.3, 0.4. . . 2.0. If the end points are modified by the term “about,” the range specified is expanded by a variation of up to ±10% of any value within the range or within 3 or more standard deviations, including the end points, or as described above in the definition of “about.” As used herein, the terms “room temperature,” “RT,” or “ambient temperature” refer to the typical temperature in an indoor laboratory setting. In one aspect, the laboratory setting is climate controlled to maintain the temperature at a substantially uniform temperature or with a specific range of temperatures. In one aspect, “room temperature” refers a temperature of about 15–30 °C, including all integers and endpoints within the specified range. In another aspect, “room temperature” refers a temperature of about 15–30 °C; about 20–30 °C; about 22–30 °C; about 25–30 °C; about 27–30 °C; about 15–22 °C; about 15–25 °C; about 15–27 °C; about 20–22 °C; about 20–25 °C; about 20–27 °C; about 22–25 °C; about 22–27 °C; about 25–27 °C; about 15 °C ± 10%; about 20 °C ± 10%; about 22 °C ± 10%; about 25 °C ± 10%; about 27 °C ± 10%; ~20 °C, ~22 °C, ~25 °C, or ~27 °C, at standard atmospheric pressure. As used herein, the terms “active ingredient” or “active pharmaceutical ingredient” refer to a pharmaceutical agent, active ingredient, compound, or substance, compositions, or mixtures thereof, that provide a pharmacological, often beneficial, effect. As used herein, the terms “control,” or “reference” are used herein interchangeably. A “reference” or “control” level may be a predetermined value or range, which is employed as a baseline or benchmark against which to assess a measured result. “Control” also refers to control experiments or control cells. As used herein, the term “dose” denotes any form of an active ingredient formulation or composition, including cells, that contains an amount sufficient to initiate or produce a therapeutic effect with at least one or more administrations. “Formulation” and “composition” are used interchangeably herein. As used herein, the term “prophylaxis” refers to preventing or reducing the progression of a disorder, either to a statistically significant degree or to a degree detectable by a person of ordinary skill in the art. As used herein, the terms “effective amount” or “therapeutically effective amount,” refers to a substantially non-toxic, but sufficient amount of an action, agent, composition, or cell(s) being administered to a subject that will prevent, treat, or ameliorate to some extent one or more of the symptoms of the disease or condition being experienced or that the subject is susceptible to contracting. The result can be the reduction or alleviation of the signs, symptoms, or causes of a disease, or any other desired alteration of a biological system. An effective amount may be based on factors individual to each subject, including, but not limited to, the subject’s age, size, type or extent of disease, stage of the disease, route of administration, the type or extent of supplemental therapy used, ongoing disease process, and type of treatment desired. As used herein, the term “subject” refers to an animal. Typically, the subject is a mammal. A subject also refers to primates (e.g., humans, male or female; infant, adolescent, or adult), non- human primates, rats, mice, rabbits, pigs, cows, sheep, goats, horses, dogs, cats, fish, birds, and the like. In one embodiment, the subject is a primate. In one embodiment, the subject is a human. As used herein, a subject is “in need of treatment” if such subject would benefit biologically, medically, or in quality of life from such treatment. A subject in need of treatment does not necessarily present symptoms, particular in the case of preventative or prophylaxis treatments. As used herein, the terms “inhibit,” “inhibition,” or “inhibiting” refer to the reduction or suppression of a given biological process, condition, symptom, disorder, or disease, or a significant decrease in the baseline activity of a biological activity or process. As used herein, “treatment” or “treating” refers to prophylaxis of, preventing, suppressing, repressing, reversing, alleviating, ameliorating, or inhibiting the progress of biological process including a disorder or disease, or completely eliminating a disease. A treatment may be either performed in an acute or chronic way. The term “treatment” also refers to reducing the severity of a disease or symptoms associated with such disease prior to affliction with the disease. “Repressing” or “ameliorating” a disease, disorder, or the symptoms thereof involves administering a cell, composition, or compound described herein to a subject after clinical appearance of such disease, disorder, or its symptoms. “Prophylaxis of” or “preventing” a disease, disorder, or the symptoms thereof involves administering a cell, composition, or compound described herein to a subject prior to onset of the disease, disorder, or the symptoms thereof. “Suppressing” a disease or disorder involves administering a cell, composition, or compound described herein to a subject after induction of the disease or disorder thereof but before its clinical appearance or symptoms thereof have manifest. As used herein, the term “PG” or “Protecting Group” refers to those reversibly formed derivatives of an existing functional group in a molecule that is temporarily attached to decrease reactivity so that the protected functional group does not react under synthetic conditions to which the molecule is subjected in one or more subsequent steps. The term “OPG” or “Orthogonal Protecting Group” refers to the strategy of building a larger molecule from subunits in which similar or identical functional groups have been differently protected beforehand. For example, a Boc- protected amino group can be deprotected in acidic media, whereas a Fmoc-protected amino group can be deprotected under basic conditions. The presence of both protective groups in the same molecule therefore enables selective deprotection of one protected amino group for a further reaction while the second protected amino group remains untouched. Suitable examplesof protecting groups include, but are not limited to, 9-fluorenylmethyl carbamate (Fmoc-NRR ), t-butyl carbamate (Boc-NRR ), benzyl carbamate (Z-NRR , Cbz-NRR ), acetamide,trifluoroacetamide, phthalimide, benzylamine (Bn-NRR ), triphenylmethylamine (Tr-NRR ),benzylideneamine, p-toluenesulfonamide (Ts-NRR ), dimethyl acetal, 1,3-dioxane, 1,3-dithiane,N,N-dimethylhydrazone, methyl ester, t-butyl ester, benzyl ester, s-t-butyl ester, 2-alkyl-1,3- oxazoline, methoxymethyl ether (MOM-OR), tetrahydropyranyl ether (THP-OR), t-butyl ether, allyl ether, benzyl ether (Bn-OR), t-butyldimethylsilyl ether (TBDMS-OR), t-butyldiphenylsilyl ether (TBDPS-OR), acetic acid ester, pivalic acid ester, benzoic acid ester, and the like. As used herein, the term “administering” an agent, such as a therapeutic entity to an animal or cell, is intended to refer to dispensing, delivering, or applying the substance to the intended target. In terms of the therapeutic agent, the term “administering” is intended to refer to contacting or dispensing, delivering or applying the therapeutic agent to a subject by any suitable route for delivery of the therapeutic agent to the desired location in the animal, including delivery by either the parenteral or oral route, intramuscular injection, subcutaneous / intradermal injection, intravenous injection, intrathecal administration, buccal administration, transdermal delivery, topical administration, and administration by the intranasal or respiratory tract route. The term “disease, disorder and / or condition” as used herein includes, but is not limited to, any abnormal condition and / or disorder of a structure or a function that affects a part of an organism (e.g., pain, inflammation, etc.). It may be caused by an external factor, such as an infectious disease, or by internal dysfunctions, such as cancer, cancer metastasis, and the like. Suitable examples include, but are not limited to, fungal infections, cancer, menopause / menstruation, diarrhea, cholera, gout, bacterial infections, viral infections, autoimmune disorder, cardiovascular diseases / disorders, and the like. As is known in the art, cancer is considered uncontrolled cell growth. The methods of the present invention can be used to treat any cancer, and any metastases thereof, including, but not limited to, carcinoma, lymphoma, blastoma, sarcoma, and leukemia. More particular examples of such cancers include breast cancer, prostate cancer, colon cancer, squamous cell cancer, small-cell lung cancer, non-small cell lung cancer, ovarian cancer, cervical cancer, gastrointestinal cancer, pancreatic cancer, glioblastoma, liver cancer, bladder cancer, hepatoma, colorectal cancer, uterine cervical cancer, endometrial carcinoma, salivary gland carcinoma, mesothelioma, kidney cancer, vulval cancer, pancreatic cancer, thyroid cancer, hepatic carcinoma, skin cancer, melanoma, brain cancer, neuroblastoma, myeloma, various types of head and neck cancer, acute lymphoblastic leukemia, acute myeloid leukemia, Ewing sarcoma, and peripheral neuroepithelioma. The present disclosure is based, in part, on recent developments made by the inventors for the identification of a modified excipient that may improve one or more property of a nanoparticle containing the modified excipient and a therapeutic compound. The identification of modified excipients that may be used in a nanoparticle formulation can be accomplished with a computer-implemented method for training a machine learning model described herein which can predict one or more physicochemical properties of the nanoparticle. Other embodiments described herein include methods, systems, and computer readable mediums for predicting one or more physicochemical properties of a nanoparticle and predicting a formulation improvement of new molecular component modifications of nanoparticles. FIG.1A illustrates a system 100 according to some embodiments. As illustrated in FIG.1A, the system 100 includes a server 105, an information repository 110, and a workstation 120. The server 105, the information repository 110, and the workstation 120 communicate over one or more wired or wireless communication networks 115. Portions of the wireless communication networks 115 may be implemented using a wide area network, such as the Internet, a local area network, such as a Bluetooth™ network or Wi-Fi, and combinations or derivatives thereof. It should be understood that the system 100 may include more or fewer servers and the single server 105 illustrated in FIG.1A is purely for illustrative purposes. For example, in some embodiments, the functionality described herein is performed via a plurality of servers in a distributed or cloud-computing environment. Also, in some embodiments, the server 105 may communicate with multiple information repositories. Additionally, it should be understood that the system 100 may include more workstations and the single workstation 120 illustrated in FIG. 1A is purely for illustrative purposes. For example, in some embodiments, the system 100 includes a plurality of workstations 120, each workstation associated with a care provider. Also, in some embodiments, the components illustrated in system 100 may communicate through one or more intermediary devices (not shown). The information repository 110 stores datasets, including, for example, nanoparticle data. The nanoparticle data may comprise a molecular representation and measured nanoparticle data, such as, a radius, a polydispersity, a ratio of normalized intensity, or other property values. The nanoparticle data may also comprise a bounded and / or exact regression datapoint related to the nanoparticle. For example, a bounded datapoint of a nanoparticle can include a known exact absolute size value and / or a known bound absolute size value of the nanoparticle. The nanoparticle data may also comprise drug and excipient molecules data of the nanoparticle, such as molecular component representations. The nanoparticle data may also comprise datapoints related to physicochemical properties, nanoparticle formation, nanoparticle radius, and drug loading. The information repository 110 stores datasets, including, for example, classification metrics, regression metrics, and filtering rules. In some embodiments, the information repository 110 may also be included as part of the server 105. Also, in some embodiments, the information repository 110 may represent multiple servers or systems. Accordingly, the server 105 may be configured to communicate with multiple systems or servers to perform the functionality described herein. Alternatively, or in addition, the information repository 110 may represent an intermediary device configured to communicate with the server 105 and one or more additional systems or servers. As illustrated in FIG.1A, the server 105 includes an electronic processor 130, a memory 135, and a communication interface 140. The electronic processor 130, the memory 135, and the communication interface 140 communicate wirelessly, over wired communication channels or buses, or a combination thereof. The server 105 may include additional components other than those illustrated in FIG. 1A in various configurations. For example, in some embodiments, the server 105 includes multiple electronic processors, multiple memory modules, multiple communication interfaces, or a combination thereof. Also, it should be understood that the functionality described herein as being performed by the server 105 may be performed in a distributed nature by a plurality of computers located in various geographic locations. For example, the functionality described herein as being performed by the server 105 may be performed by a plurality of computers included in a cloud computing environment. The electronic processor 130 may be, for example, a microprocessor, an application- specific integrated circuit (ASIC), and the like. The electronic processor 130 is configured to execute software instructions to perform a set of functions, including the functions described herein. The memory 135 includes a non-transitory computer-readable medium and stores data, including instructions executable by the electronic processor 130. The communication interface 140 may be, for example, a wired or wireless transceiver or port, for communication over the communication network 115 and, optionally, one or more additional communication networks or connections. As illustrated in FIG. 1A, the memory 135 of the server 105 includes machine learning model(s) 145A-N, which may be part of a nanoparticle prediction engine executed via the server 105. The model 145 may be, for example, an artificial intelligence system. In some embodiment, as a pair of molecules, for example, a drug-excipient pair, is received, the server 105 uses the model 145 to determine one or more physicochemical properties of a nanoparticle for the pair of molecules. The one or more physicochemical properties may include, for example, nanoparticle formation, nanoparticle radius, and / or drug loading. In some embodiment, as a pair of nanoparticles including a drug-excipient pair is received, the server 105 uses the model 145 to determine formulation improvement of new molecular component modifications of nanoparticles. The formulation improvements may include, for example, which of the nanoparticles of the pair of nanoparticles has a larger radius. FIG. 1B is a flowchart illustrating an example method 200 for a traditional approach for size difference comparison. In the example illustrated, the method 200 includes providing training nanoparticle data with individual nanoparticle representations and only known exact size values (at block 205). For example, a computing system receives a set of data including individual nanoparticle representations and corresponding known exact size values. The method 200 includes training machine learning models on absolute size values (at block 210). For example, the computing system trains the machine learning models on a set of training data including the individual nanoparticle representations and corresponding known exact size values. In one aspect, the machine learning model is trained to predict property values (e.g., absolute size values) for nanoparticles. The method 200 includes predicting absolute size values for new nanoparticles (at block 215). For example, the computing system inputs a nanoparticle representation from a set of test data to the trained machine learning model. The trained machine learning model outputs an absolute size value for the nanoparticle representation. The method 200 includes manually subtracting predicted values to approximate size changes (at block 220). For example, a user obtains outputs, from the trained machine learning model, for a first individual nanoparticle representation and a second individual nanoparticle representation. The user subtracts the outputs to determine an approximate change in size between the individual nanoparticle representations. The method 200 includes classifying formulation improvements based on changes in size (at block 225). For example, the user determines whether the first individual nanoparticle representation or the second individual nanoparticle representation led to predictions of larger size based on predictive models (block 215). FIG.1C is a flowchart illustrating a method 300 for training a machine learning model for predicting a formulation improvement of new molecular component modifications between pairs of nanoparticles, according to some embodiments. The method 300 may be performed by the server 105 (i.e., the electronic processor 130 implementing the model 145). However, in other embodiments, the method 300 may be performed by multiple servers or systems in various configurations and distributions. The method 300 includes providing training nanoparticles with molecular component representations as well as both exact and bounded size values (at block 305). For example, a set of data including nanoparticles may be uploaded to the information repository 110, and the server 105 may receive the set of data including nanoparticles and use the set of data including nanoparticles to access or receive associated information regarding a molecular component representation, a known exact absolute size value of each molecule, and a known bound size value, in the set of data from (e.g., through a push or pull configuration) the information repository 110, other data sources, or a combination thereof as described above. The known exact absolute size value and the known bound size value are related to a radius of each nanoparticle of the set of data, and each nanoparticle of the set of data includes a drug and an excipient. The method 300 includes combinatorically expanding available data by cross-merging and pairing training data (at block 310). For example, the server 105 creates a set of training data with the set of data. The server 105 splits the set of training data into a plurality of training sets and a plurality of test sets. The server 105 creates, by cross-merging, a set of nanoparticle pairs using each nanoparticle of the set of training data. Put another way, the server 105 generates a pair of nanoparticles for the training set and the test set by cross merging a first nanoparticle and a second nanoparticle of the plurality of training sets and / or the plurality of test sets, wherein all possible nanoparticle pairs of the training sets and the test sets are generated. The method 300 includes removing nanoparticle pairs when it is unknown which is larger in size (at block 315). For example, the server 105 filters the set of training data based on a set of rules, wherein the filtered set of training data includes nanoparticle pairs of the set of nanoparticle pairs where a nanoparticle of the nanoparticle pairs with a larger size is known. The server 105 removes nanoparticle pairs of the set of training data from the set of training data where whether a first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown. The server 105 retains / keeps nanoparticle pairs of the set of training data in the set of training data where it is possible to determine whether the first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size. The set of rules may include: (1) if a first nanoparticle (“NP 1”) and a second nanoparticle (“NP 2”) of a nanoparticle pair have exact value, then the server 105 keeps the nanoparticle pair in the set of training data; (2) if NP 1 has an exact value and NP 2 is “>” a nanoparticle size threshold (e.g., 200 nm), then (a) if NP 1 is less than NP 2 then the server 105 keeps the nanoparticle pair and (b) if NP 1 is greater than NP 2, then the server 105 removes the nanoparticle pair; (3) if NP 1 is “>” the nanoparticle size threshold and NP 2 has an exact value, then (a) if NP 1 is greater than NP 2, then the server 105 keeps the nanoparticle pair or if NP 1 is less than NP 2, then the server 105 removes the nanoparticle pair; (4) if NP 1 is “>” the nanoparticle size threshold and NP 2 is “>” the nanoparticle size threshold, then the server 105 removes the nanoparticle pair. In some instances, if a nanoparticle is greater than 200 nm (which in some instances would make the nanoparticle an invalid nanoparticle by definition), then the server 105 assigns a comparator value “>” for the nanoparticle and changes the value of the radius associated with the nanoparticle to 200 nm. Furthermore, the server 105 keeps the nanoparticle value as is when exact and sets the comparator value to The server 105 considers the comparator value as a parameter to remove incomparable nanoparticles. For example, if NP 1 > 200 nm but NP 2 = 231 nm, according to the logic above, NP 1 is “>” and NP 2 is and NP 1’s value (200 nm) is less NP 2’s value (231 nm), so the server 105 removes the nanoparticle pair. The server discards the nanoparticle pair as it is not certain which nanoparticle is larger in value (radius). The method 300 includes training a machine learning model to classify whether there is a formulation improvement between each pair of nanoparticles based on size (at block 320). For example, the server 105 trains a machine learning model of an AI system using datapoints of the filtered set of training data, wherein the datapoints include the nanoparticle pairs of the filtered set of training data with shared representations, physicochemical properties of the shared representations, and respective property differences in size of each pair of nanoparticles in the set of training data to train the model 145. In this example, the server 105 uses the respective property differences as the objective variable instead of the absolute property size of a single nanoparticle. The property differences enable the model 145 to directly learn property differences from nanoparticle pairs instead of learning absolute size values from single nanoparticles. In other examples, the set of training data may include information, such as input-output pairs, in memory 135. The input-output pairs may include a set of features of a shared molecular representation (e.g., input) and a property difference (e.g., output) corresponding to the set features. As noted above, the labels may be defined manually by an expert or determined based on another methodology. The method 300 includes directly predicting formulation improvements of new molecular component modifications of nanoparticles (at block 325). For example, for datapoints of a pair of nanoparticles, the server 105 determines, using the machine learning model as trained based on size of the pair of nanoparticles for the datapoints, a classification of a formulation improvement of new molecular component modifications of nanoparticles. The formulation improvement indicates at least one of: a nanoparticle formation, a nanoparticle radius, and a drug loading for the pair of nanoparticles. The predicted classification is an output (e.g., binary output) indicating whether the first or second nanoparticle in the pair of nanoparticles has a larger radius and drug loading potential. FIG.1D is a flowchart illustrating a method 400 for training a machine learning model for predicting one or more physicochemical properties of a nanoparticle, according to some embodiments. The physicochemical properties include nanoparticle formation, nanoparticle radius, and / or drug loading. The method 400 may be performed by the server 105 (i.e., the electronic processor 130 implementing the model 145). However, in other embodiments, the method 400 may be performed by multiple servers or systems in various configurations and distributions. The method 400 includes calculating descriptors and fingerprints from molecular representations (at block 405). For example, the server 105 receives a set of data including nanoparticles, each nanoparticle includes a drug-excipient pair including a drug molecule and an excipient molecule. Each nanoparticle of the set of data includes a molecular representation and a property value of each nanoparticle. The property value of each nanoparticle is, for example, a nanoparticle measurement including a radius, a polydispersity, and a ratio of normalized intensity. The server 105 generates, using the molecular representation, a descriptor and fingerprint for each nanoparticle of the set of data. The method 400 includes splitting a dataset into a training set and a test set using different splitting strategies (at block 410). For example, the server 105 creates, by using a cross-validation method and the set of data, a plurality of sets of training data and a plurality of sets of test data. The set of data includes the descriptor and the fingerprint for each nanoparticle of the set of data. The cross-validation methods for splitting the dataset includes a ten (10) fold cross-validation and a Leave-One-Out (“LOO”) cross-validation, which includes a Leave-One-Excipient-Out (“LOEO”) cross-validation, a Leave-One-Drug-Out (“LODO”) cross-validation, and a Leave-One-Pair-Out (“LOPO”) cross-validation. The method 400 includes training separate machine learning models on each training set and making predictions for drug-excipient combinations in the corresponding test set (at block 415). For example, the server 105 trains a machine learning model of an artificial intelligence (AI) system using the plurality of sets of training data, wherein the plurality of sets of training data includes drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data. In some embodiments, the machine learning model is either a classification model or a regression model. The server 105 predicts, using the machine learning model as trained based on drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data, a property of a drug-excipient pair of the plurality of sets of test data. The predicted property of the drug-excipient pair of the plurality of sets of test data includes nanoparticle formation, nanoparticle radius, and / or drug loading. The method 400 includes assessing model performance using classification or regression metrics (at block 420). For example, the server 105 assesses a performance of the classification model, using classification metrics. The classification metrics include accuracy, precision, recall, area under the Receiver Operating Characteristic curve (“AUC” ROC), F1 score, and Mathews Correlation Coefficient (“MCC”). The server assesses the regression model, using regression metrics. The regression metrics include a root mean squared error (“RMSE”), a mean absolute error (“MAE”), and a coefficient of determination (R2). Another embodiment described herein is a computing system configured to carry out the foregoing methods. The system can comprise any suitable components, which will be evident to a person of skill in the art. The components can include, but are not limited to, a processor, a memory, a computing platform, and a software algorithm. The systems and methods described herein can be implemented in hardware, software, firmware, or combinations of hardware, software, and / or firmware. In some examples, the systems and methods described herein may be implemented using a non-transitory computer readable medium storing computer executable instructions that when executed by one or more processors of a computer cause the computer to perform operations. Computer readable media suitable for implementing the systems and methods described in this specification include non- transitory computer-readable media, such as disk memory devices, chip memory devices, programmable logic devices, random access memory (RAM), read only memory (ROM), optical read / write memory, cache memory, magnetic read / write memory, flash memory, and application- specific integrated circuits. In addition, a computer readable medium that implements a system or method described herein may be located on a single device or computing platform or may be distributed across multiple devices or computing platforms. Although the method described herein may be used for any modification of an excipient, PEGylated small molecular excipients can serve as powerful tools to create highly loaded PEGylated nanoparticles. Specifically, the methods described herein represent notable progress in the co-assembly of a drug / therapeutic compound with a PEGylated excipient, leading to the creation of nanoparticles with high drug loading. The results demonstrate that the co-assembly method described herein, recently acknowledged for its potential in achieving high drug loading, is also effective with PEGylated excipients and offers more flexibility than traditional methods. This discovery expands the applicability of the drug-excipient co-assembly method in nanoparticle production. Both in vitro and in vivo studies were performed to validate new co-aggregated PEGylated nanoparticles and demonstrate their improved drug loading and bioavailability. The present disclosure is based, in part, on recent developments made by the inventors for the design and creation of novel excipients that include modifying available small molecular excipients through the chemical attachment of polyethylene glycol (PEG). The vast design space of these excipients, created by selecting various established excipients, modification sites, and PEG, has created a massive opportunity for the discovery of millions of novel excipients. Importantly, by chemically combining various safe materials, it is expected that these excipients will have a large potential to be safe for clinical translation. Accordingly, one embodiment described herein is a modified excipient, the excipient comprising, consisting of, or consisting essentially of a first compound linked to a PEG. Another embodiment described herein is an excipient conjugated to a polyethylene glycol (PEG) moiety consisting of, or consisting essentially of the general formula selected from: Folate-PEG,
[0005]
[0006] Congo Red-PEG, and any salt, ester, derivative, or variant thereof. In some embodiments, the excipient conjugated to a PEG, may be a pharmaceutically acceptable salt of the excipient. In some embodiments, the excipient is conjugated to a PEG repeating unit having astructure of: . In some embodiments, the number of PEG repeating units (n) isabout 22 to about 225. In some embodiments, the number of PEG repeating units (n) is about 22–50, 22–75, 22–100, 22–125, 22–150, 22–200, 22–225, 50–75, 50–100, 50–125, 50–150, 50– 200, 50–225, 75–100, 75–125, 75–150, 75–200, 75–225, 100–125, 100–150, 100–200, 100–225, 125–150, 125–200, 125–225, 150–200, 150–225, or 200–225, including all integers and endpoints within the specified range. “About” may apply to each individual range specified within the list. Another embodiment described herein is a nanoparticle or nanoparticle formulation comprising, consisting of, or consisting essentially of a PEG excipient as described herein and a therapeutic compound. In some embodiments, the therapeutic compound is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4- hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, analogues thereof, or combinations thereof. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a hydrodynamic radius of about 10 nanometers (nm) to about 500 nanometers (nm). In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a hydrodynamic radius of about 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 110 nm, 120 nm, 130 nm, 140 nm, 150 nm, 160 nm, 170 nm, 180 nm, 190 nm, 200 nm, 210 nm, 220 nm, 230 nm, 240 nm, 250 nm, 260 nm, 270 nm, 280 nm, 290 nm, 300 nm, 310 nm, 320 nm, 330 nm, 340 nm, 350 nm, 360 nm, 370 nm, 380 nm, 390 nm, 400 nm, 410 nm, 420 nm, 430 nm, 440 nm, 450 nm, 460 nm, 470 nm, 480 nm, 490 nm, or 500 nm. “About” may apply to each individual value in the list. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a hydrodynamic radius of about 10–50 nm, 10–100 nm, 10–150 nm, 10–200 nm, 10–250 nm, 10–300 nm, 10–350 nm, 10–400 nm, 10–450 nm, 10–500 nm, 50– 100 nm, 50–150 nm, 50–200 nm, 50–250 nm, 50–300 nm, 50–350 nm, 50–400 nm, 50–450 nm, 50–500 nm, 100–150 nm, 100–200 nm, 100–250 nm, 100–300 nm, 100–350 nm, 100–400 nm, 100–450 nm, 100–500 nm, 150–200 nm, 150–250 nm, 150–300 nm, 150–350 nm, 150–400 nm, 150–450 nm, 150–500 nm, 200–250 nm, 200–300 nm, 200–350 nm, 200–400 nm, 200–450 nm, 200–500 nm, 250–300 nm, 250–350 nm, 250–400 nm, 250–450 nm, 250–500 nm, 300–350 nm, 300–400 nm, 300–450 nm, 300–500 nm, 350–400 nm, 350–450 nm, 350–500 nm, 400–450 nm, 400–500 nm, or 450–500 nm, including all integers and endpoints within the specified range. “About” may apply to each individual range specified within the list. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a stability of 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours or more. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a stability of greater than 1 hour, greater than 2 hours, greater than 3 hours, greater than 4 hours, greater than 5 hours, greater than 6 hours, greater than 7 hours, greater than 8 hours, greater than 9 hours, greater than 10 hours, greater than 11 hours, or greater than 12 hours. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein have a loading value of about 50% to about 95%. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a loading value of about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%. “About” may apply to each individual value in the list. In some embodiments, the nanoparticle(s) or nanoparticle formulation described herein has a loading value of about 50–55%, 50–60%, 50–65%, 50–70%, 50–75%, 50–80%, 50–85%, 50– 90%, 50–95%, 55–60%, 55–65%, 55–70%, 55–75%, 55–80%, 55–85%, 55–90%, 55–95%, 60– 65%, 60–70%, 60–75%, 60–80%, 60–85%, 60–90%, 60–95%, 65–70%, 65–75%, 65–80%, 65– 85%, 65–90%, 65–95%, 70–75%, 70–80%, 70–85%, 70–90%, 70–95%, 75–80%, 75–85%, 75– 90%, 75–95%, 80–85%, 80–90%, 80–95%, 85–90%, 85–95%, or 90–95%, including all integers and endpoints within the specified range. “About” may apply to each individual range specified within the list. In some embodiments, the nanoparticle(s) or nanoparticle formulation includes an excipient, or a pharmaceutically acceptable salt thereof, having a structure shown below: , and has a loading value of about 50% to about 95%. In some embodiments, the nanoparticle(s) or nanoparticle formulation includes fulvestrant and an excipient, or pharmaceutically acceptable salt thereof, having a structure shown below: , and has a loading value of 77%. Another embodiment described herein is a pharmaceutical composition or formulation comprising, consisting of, or consisting essentially of a modified excipient as described herein or a nanoparticle as described herein and a pharmaceutically acceptable carrier, diluent, and / or excipient. Another embodiment described herein are methods of making and using modified excipients, nanoparticles, and pharmaceutical compositions or formulations thereof, as described herein. Another embodiment described herein is an approach to develop novel excipients by modifying available small molecular excipients through chemically attaching polyethylene glycol (PEG). By chemically combining various safe materials, novel excipients are created that have a large potential to be safe for clinical translation. Other embodiments described herein include novel excipients that can be easily synthesized using established syntheses and chemical reaction allowing for deployment and chemical scale-up. These excipients are capable of forming nanoparticles with drugs due to their nanoparticle-forming ability and have been characterized for other translationally relevant properties. These excipients have been shown to: be easily synthesizable following established chemical reactions with large potential for deployment and chemical scale-up for production; able to form nanoparticles with various clinically highly relevant drugs, showcasing the potential of these novel chemical entities to serve as excipients for our nanoparticle platforms; reduce the size and enhance stability of the nanoparticles formed with several drugs compared to their unmodified counterparts – indicating that several of these novel excipients can be superior compared to their unmodified counterparts. The size reduction of these nanoparticles is particularly important as smaller particles can more easily cross several biological barriers (e.g., blood-brain barrier), enhancing drug delivery to target tissues. The stability enhancement is equally important since nanoparticles need to be stable for a sufficient duration and in various conditions to enable their clinical translation; encapsulate challenging drugs thereby showcasing the potential of these modified excipients to vastly expand the design space of current nanoparticles and treat life- threatening diseases with nanoparticles that encapsulate drugs previously inaccessible to our nanoparticle formulations. Several of the modified excipients demonstrate reduced protein binding. Reduced protein binding is advantageous as it minimizes the interaction between the drug and serum proteins, thereby increasing the bioavailability and therapeutic potential of the drug. Several of the modified excipients demonstrate reduced macrophage uptake. By evading this uptake, particles circulate longer and are less likely to trigger immune responses, which can enhance their therapeutic efficacy and reduce toxicity. Several of the modified excipients demonstrate an ability to create nanoparticles that significantly extend the plasma half-life of the drug compared to unmodified excipients. Nanoparticles that increase the half-life of a drug help maintain therapeutic levels in the bloodstream for longer periods, thereby enhancing efficacy and reducing the frequency of dosing. Compositions Modified Excipients Another embodiment described herein is a modified excipient, the excipient comprising, consisting of, or consisting essentially of a first compound covalently linked with PEG. Another embodiment described herein is a modified excipient comprising, consisting of, or consisting essentially of the general formula selected from: Rhodamine B-PEG-2, Congo Red-PEG, and any salt, ester, derivative, or variant thereof. Nanoparticles The modified excipients described herein have been shown to be able to form nanoparticles with various clinically highly relevant drugs, showcasing the potential of these novel chemical entities to serve as excipients for nanoparticle platforms. Indeed, one of the novel features of the modified excipients described herein is the ability to encapsulate challenging drugs (e.g., drugs that are hard to, or never have been, encapsulated) into nanoparticles for delivery to a subject in need thereof. Any drug can be used in the methods described herein, for example, analgesics, NSAIDS, anti-inflammatory drugs, chemotherapeutic drugs, antiarrhythmic drugs, inhalants, opioids, antidepressants, anti-diarrheal drugs, antibacterial; (e.g., antibiotics), antiviral drugs, anti-fungal drugs, menstruation drugs, antipsychotics, and the like. Another embodiment described herein is a nanoparticle comprising of, consisting of, or consisting essentially of a modified excipient as described herein and a drug. In some embodiments, the drug is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4-hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, analogues thereof, or combinations thereof. Pharmaceutical Compositions Other embodiments described herein are compositions comprising one or more of the modified excipients and / or nanoparticles as described herein and an appropriate carrier, excipient, or diluent. The exact nature of the carrier, excipient or diluent will depend upon the desired use for the composition and may range from being suitable or acceptable for veterinary uses to being suitable or acceptable for human use. The composition may optionally include one or more additional compounds. When used to treat or prevent a disease, disorder, and or condition, such as a pain, cancer, inflammation, and the like, the compounds described herein may be administered singly, as mixtures of one or more compounds or in mixture or combination with other agents (e.g., therapeutic agents) useful for treating such diseases, disorders and / or conditions and / or the symptoms associated with such diseases, disorders and / or conditions. Such agents may include, but are not limited to, antibiotics, NSAIDS, anti-inflammatory compounds, chemotherapeutic agents, anti-cancer agents (e.g., immunotherapy, anti-mitotic compounds, etc.), opiates, steroids, to name a few. The compounds may be administered in the form of compounds per se, or as pharmaceutical compositions comprising a compound. Pharmaceutical compositions comprising the compounds may be manufactured by means of conventional mixing, dissolving, granulating, dragee-making, levigating, emulsifying, encapsulating, entrapping or lyophilization processes. The compositions may be formulated in a conventional manner using one or more physiologically acceptable carriers, diluents, excipients, or auxiliaries which facilitate processing of the compounds into preparations which can be used pharmaceutically. The compounds may be formulated in the pharmaceutical composition per se, or in the form of a hydrate, solvate, N-oxide or pharmaceutically acceptable salt, as previously described. Typically, such salts are more soluble in aqueous solutions than the corresponding free acids and bases, but salts having lower solubility than the corresponding free acids and bases may also be formed. Pharmaceutical compositions may take a form suitable for any mode of administration, including, for example, topical, ocular, oral, buccal, systemic, nasal, injection, transdermal, rectal, vaginal, etc., or a form suitable for administration by inhalation or insufflation. For topical administration, the compounds may be formulated as solutions, gels, ointments, creams, suspensions, etc. as are well-known in the art. Systemic formulations include those designed for administration by injection, e.g., subcutaneous, intravenous, intramuscular, intrathecal, or intraperitoneal injection, as well as those designed for transdermal, transmucosal oral, or pulmonary administration. Useful injectable preparations include sterile suspensions, solutions, or emulsions of the active compounds in aqueous or oily vehicles. The compositions may also contain formulating agents, such as suspending, stabilizing and / or dispersing agents. The formulations for injection may be presented in unit dosage form, e.g., in ampules or in multidose containers, and may contain added preservatives. Alternatively, the injectable formulation may be provided in powder form for reconstitution with a suitable vehicle, including but not limited to sterile pyrogen free water, buffer, dextrose solution, etc., before use. To this end, the active compounds may be dried by any art-known technique, such as lyophilization, and reconstituted prior to use. For transmucosal administration, penetrants appropriate to the barrier to be permeated are used in the formulation. Such penetrants are known in the art. For oral administration, the pharmaceutical compositions may take the form of, for example, lozenges, tablets or capsules prepared by conventional means with pharmaceutically acceptable excipients such as binding agents (e.g., pregelatinized maize starch, polyvinylpyrrolidone or hydroxypropyl methylcellulose); fillers (e.g., lactose, microcrystalline cellulose or calcium hydrogen phosphate); lubricants (e.g., magnesium stearate, talc or silica); disintegrants (e.g., potato starch or sodium starch glycolate); or wetting agents (e.g., sodium lauryl sulfate). The tablets may be coated by methods well known in the art with, for example, sugars, films, or enteric coatings. Liquid preparations for oral administration may take the form of elixirs, solutions, syrups, or suspensions, or they may be presented as a dry product to be reconstituted with water or another suitable vehicle before use. Such liquid preparations may be prepared by conventional means with pharmaceutically acceptable additives such as suspending agents (e.g., sorbitol syrup, cellulose derivatives or hydrogenated edible fats); emulsifying agents (e.g., lecithin or acacia); non-aqueous vehicles (e.g., almond oil, oily esters, ethyl alcohol, Cremophor™ or fractionated vegetable oils); and preservatives (e.g., methyl or propyl-p-hydroxybenzoates or sorbic acid). The preparations may also contain buffer salts, preservatives, flavoring, coloring, and sweetening agents as appropriate. Preparations for oral administration may be suitably formulated to provide controlled release of the compound, as is well established. For buccal administration, the compositions may take the form of tablets or lozenges formulated in conventional manner. For rectal and vaginal routes of administration, the compounds may be formulated as solutions (for retention enemas) suppositories or ointments containing conventional suppository bases such as cocoa butter or other glycerides. For nasal administration or administration by inhalation or insufflation, the compounds can be conveniently delivered in the form of an aerosol spray from pressurized packs or a nebulizer with the use of a suitable propellant, e.g., dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoroethane, fluorocarbons, carbon dioxide or other suitable gas. In the case of a pressurized aerosol, the dosage unit may be determined by providing a valve to deliver a metered amount. Capsules and cartridges for use in an inhaler or insufflator (for example, capsules and cartridges comprised of gelatin) may be formulated containing a powder mix of the compound and a suitable powder base such as lactose or starch. For ocular administration, the compounds may be formulated as a solution, emulsion, suspension, etc. suitable for administration to the eye. A variety of vehicles suitable for administering compounds to the eye are known in the art. For prolonged delivery, the compounds can be formulated as a depot preparation for administration by implantation or intramuscular injection. The compounds may be formulated with suitable polymeric or hydrophobic materials (e.g., as an emulsion in an acceptable oil) or ion exchange resins, or as sparingly soluble derivatives, e.g., as a sparingly soluble salt. Alternatively, transdermal delivery systems manufactured as an adhesive disc or patch which slowly releases the compounds for percutaneous absorption may be used. To this end, permeation enhancers may be used to facilitate transdermal penetration of the compounds. Alternatively, other pharmaceutical delivery systems may be employed. Liposomes and emulsions are well-known examples of delivery vehicles that may be used to deliver compounds. Certain organic solvents such as dimethyl sulfoxide (DMSO) may also be employed, although usually at the cost of greater toxicity. The pharmaceutical compositions may, if desired, be presented in a pack or dispenser device which may contain one or more unit dosage forms containing the compounds. The pack may, for example, comprise metal or plastic foil, such as a blister pack. The pack or dispenser device may be accompanied by instructions for administration. The compounds described herein, or compositions thereof, will be used in an amount effective to achieve the intended result, for example in an amount effective to treat or prevent the particular disease being treated. By therapeutic benefit is meant eradication or amelioration of the underlying disorder being treated and / or eradication or amelioration of one or more of the symptoms associated with the underlying disorder such that the patient reports an improvement in feeling or condition, notwithstanding that the patient may still be afflicted with the underlying disorder. Therapeutic benefit also generally includes halting or slowing the progression of the disease, regardless of whether improvement is realized. The dosage of compounds administered will depend upon a variety of factors, including, for example, the particular indication being treated, the mode of administration, whether the desired benefit is prophylactic or therapeutic, the severity of the indication being treated and the age and weight of the patient, the bioavailability of the particular compounds the conversation rate and efficiency into active drug compound under the selected route of administration, etc. Determination of an effective dosage of compounds for a particular use and mode of administration is well within the capabilities of those skilled in the art. Effective dosages may be estimated initially from in vitro activity and metabolism assays. For example, an initial dosage of compound for use in animals may be formulated to achieve a circulating blood or serum concentration of the metabolite active compound that is at or above an IC50of the particular compound as measured in as in vitro assay. Calculating dosages to achieve such circulating blood or serum concentrations considering the bioavailability of the particular compound via the desired route of administration is well within the capabilities of skilled artisans. Initial dosages of a particular compound can also be estimated from in vivo data, such as animal models. Animal models useful for testing the efficacy of the active metabolites to treat or prevent the various diseases described above are well-known in the art. Animal models suitable for testing the bioavailability and / or metabolism of compounds into active metabolites are also well-known. Ordinarily skilled artisans can routinely adapt such information to determine dosages of particular compounds suitable for human administration. Dosage amounts will typically range from about 0.0001 mg / kg / day, 0.001 mg / kg / day or 0.01 mg / kg / day to about 100 mg / kg / day, but may be higher or lower, depending upon various factors including the activity and bioavailability of the active compound, its metabolism kinetics and other pharmacokinetic properties, the mode of administration and various other factors, discussed above. Dosage amount and interval may be adjusted individually to provide plasma levels of the compounds and / or active metabolite compounds which are sufficient to maintain therapeutic or prophylactic effect. For example, the compounds may be administered once per week, several times per week (e.g., every other day), once per day or multiple times per day, depending upon, among other things, the mode of administration, the specific indication being treated, and the judgment of the prescribing physician. In cases of local administration or selective uptake, such as local topical administration, the effective local concentration of compounds and / or active metabolite compounds may not be related to plasma concentration. Skilled artisans will be able to optimize effective dosages without undue experimentation. Methods of Use The modified excipients, nanoparticles, and pharmaceutical compositions or formulations thereof can be used for the treatment and / or prevention of a disease, disorder, and / or condition in a subject. Another embodiment described herein is a method for treating and / or preventing a disease, disorder, or condition in a subject, the method comprising, consisting of, or consisting essentially of administering to the subject a therapeutically effective amount of a nanoparticle as described herein, or a pharmaceutical composition thereof, such that the disease, disorder and / or condition is prevented and / or treated in the subject. In some embodiments, the disease, disorder, and / or condition is selected from cancer, fungal infection, menopause, diarrhea, gout, or cholera. In some embodiments, the methods provide administering at least one additional therapeutic agent. In one embodiment, at least one additional therapeutic agent is administered prior to the nanoparticle, or composition thereof. In another embodiment, the at least one additional therapeutic agent is administered concurrently with the nanoparticle, or composition thereof. In another embodiment, the at least one additional therapeutic agent is administered after the nanoparticle, or composition thereof. The modified excipients, nanoparticles, and pharmaceutical compositions thereof can be used to reduce macrophage uptake. In some embodiments, macrophage uptake may be reduced by administering a therapeutically effective amount of nanoparticles or a nanoparticle formulation. In some embodiments, the nanoparticles or the nanoparticle formulation may be administered in vivo or in vitro. The modified excipients, nanoparticles, and pharmaceutical compositions thereof can be used to reduce the protein binding of a therapeutic compound. In some embodiments, protein binding of a therapeutic compound may be reduced by administering an effective amount of the nanoparticles or a nanoparticle formulation. In some embodiments, the nanoparticles or the nanoparticle formulation may be administered in vivo or in vitro. Kits Other embodiments described herein are kits comprising the compositions described herein and for carrying out the subject methods as described herein. For example, in one embodiment, a subject kit may comprise, consist of, or consist essentially of a modified excipient as described herein, a nanoparticle as described herein and / or pharmaceutical compositions as described herein. In other embodiments, a kit may further include other components. Such components may be provided individually or in combinations and may provide in any suitable container such as a vial, a bottle, or a tube. Examples of such components include, but are not limited to, one or more additional reagents, such as one or more dilution buffers; one or more reconstitution solutions; one or more wash buffers; one or more storage buffers, one or more control reagents and the like. Components (e.g., reagents) may also be provided in a form that is usable in a particular assay, or in a form that requires addition of one or more other components before use (e.g., in concentrate or lyophilized form). Suitable buffers include, but are not limited to, phosphate buffered saline, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In addition to the above-mentioned components, a kit can further include instructions for using the components of the kit to practice the subject methods. The instructions for practicing the subject methods are recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or subpackaging) etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g., CD-ROM, diskette, flash drive, etc. In yet other embodiments, the actual instructions are not present in the kit but means for obtaining the instructions from a remote source, e.g., via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate. One embodiment described herein is a computer-implemented method for training a machine learning model for predicting one or more physicochemical properties of a nanoparticle, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular representation and a property value of each nanoparticle, wherein each nanoparticle includes a drug-excipient pair including a drug molecule and an excipient molecule; generate, using the molecular representation, a descriptor and fingerprint for each nanoparticle of the set of data; creating, by using cross-validation method and the set of data, a plurality of sets of training data and a plurality of sets of test data, wherein the set of data includes the descriptor and fingerprint for each nanoparticle of the set of data; training a machine learning model of an artificial intelligence (AI) system using the plurality of sets of training data, wherein the plurality of sets of training data includes drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data; and for a drug-excipient pair of the plurality of sets of test data, predicting a property of the drug-excipient pair of the plurality of sets of test data using the machine learning model as trained based on drug-excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data. In one aspect, the one or more physicochemical properties comprise nanoparticle formation, nanoparticle radius, or drug loading. In another aspect, the machine learning model is selected from a classification model or a regression model. In another aspect, the method further comprises: assessing a performance of the classification model, using classification metrics, and the regression model, using regression metrics, wherein classification metrics include at least one of an accuracy, a precision, a recall, an area under the Receiver Operating Characteristic curve (“AUC” ROC), a F1 score, and a Mathews Correlation Coefficient (“MCC”), and wherein the regression metrics include at least one of a root mean squared error (“RMSE”), a mean absolute error (“MAE”), and a coefficient of determination (R2). In another aspect, the predicted property of the drug-excipient pair of the plurality of sets of test data is selected from a nanoparticle formation, a nanoparticle radius, or a drug loading. In another aspect, the property value of each nanoparticle is a nanoparticle measurement selected from a radius, a polydispersity, or a ratio of normalized intensity. In another aspect, the cross-validation method is selected from a ten (10) fold cross- validation and a Leave-One-Out (“LOO”) cross-validation, wherein the LOO cross-validation includes a Leave-One-Excipient-Out (“LOEO”) cross-validation, a Leave-One-Drug-Out (“LODO”) cross-validation, or a Leave-One-Pair-Out (“LOPO”) cross-validation. Another embodiment described herein is a computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute a method described herein. Another embodiment described herein is a computer-implemented method for training a machine learning model for predicting a formulation improvement of new molecular component modifications of nanoparticles, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular component representation and a value selected from a known exact absolute size value or a known bound absolute size value, wherein the known exact absolute size value and the known bound absolute size value are related to a radius of each nanoparticle of the set of data, and wherein each nanoparticle of the set of data comprises a drug and an excipient; creating a set of training data with the set of data; creating a set of nanoparticle pairs using each nanoparticle of the set of training data; filtering the set of training data based on a set of rules, wherein the filtered set of training data includes nanoparticle pairs of the set of nanoparticle pairs where a nanoparticle of the nanoparticle pairs with a larger size is known; training a machine learning model of an AI system using datapoints of the filtered set of training data, wherein the datapoints include the nanoparticle pairs of the filtered set of training data with shared representations, and wherein the datapoints include physicochemical properties of the shared representations; and for datapoints of a pair of nanoparticles, predicting a classification of a formulation improvement of new molecular component modifications of nanoparticles using the machine learning model as trained based on size of the pair of nanoparticles for the datapoints, wherein the formulation improvement indicates at least one selected from a nanoparticle formation, a nanoparticle radius, or a drug loading for the pair of nanoparticles. In one aspect, filtering the set of training data based on the set of rules, further comprises: removing nanoparticle pairs of the set of training data from the set of training data where whether a first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is greater than the second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is less than the second nanoparticle, wherein the first nanoparticle has a known bound absolute size value a known exact absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle and the second nanoparticle have a known bound absolute size value, wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle and a second nanoparticle of the nanoparticle pairs of the set of training data have a known bound absolute size value. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is less than a second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is greater than a second nanoparticle, wherein the first nanoparticle has a known bound absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. In another aspect, creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: splitting the set of training data into a training set and a test set. In another aspect, creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: generating a pair of nanoparticles for at least one of the training set and the test set by cross merging a first nanoparticle and a second nanoparticle of the training set and the test set, wherein all possible nanoparticle pairs of the training set and the test set are generated. In another aspect, the predicted classification is an output indicating whether a first nanoparticle or a second nanoparticle in the pair nanoparticles has a larger radius. Another embodiment described herein is a computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute a method described herein. Another embodiment described herein is an excipient conjugated to a polyethylene glycol moiety, or pharmaceutically acceptable salt thereof, selected from: Rhodamine B-PEG-2, Folate-PEG, Cholesterol-PEG,
[0007] Stearic acid-PEG, or
[0008] Congo Red-PEG. In one aspect, n is about 22 to about 225. Another embodiment described herein is a nanoparticle formulation comprising: an excipient described herein, or pharmaceutically acceptable salt thereof; and a therapeutic compound. In one aspect, the therapeutic compound is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4- hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, or analogues or combinations thereof. In another aspect, the nanoparticles have a hydrodynamic radius of about 10 nm to about 500 nm. In another aspect, the nanoparticles have a therapeutic compound loading value of about 50% to about 95%. In another aspect, the
[0009] or pharmaceutically acceptable salt thereof, and the nanoparticles have a therapeutic compound loading value of 77%. In another aspect, the nanoparticles have a stability greater than 12 hours. Another embodiment described herein is a method for treating a disease or disorder or prophylaxis thereof in a subject in need thereof, the method comprising administering a therapeutically effective amount of a nanoparticle formulation described herein to a subject in need thereof. Another embodiment described herein is a method for reducing macrophage uptake, the method comprising administering a therapeutically effective amount of a nanoparticle formulation described herein to a subject in need thereof. Another embodiment described herein is a method for reducing protein binding of a therapeutic compound, the method comprising administering an effective amount of a nanoparticle formulation described herein. Another embodiment described herein is the use of an excipient described herein, or pharmaceutically acceptable salt thereof, to modulate systemic circulation of a therapeutic compound. It will be apparent to one of ordinary skill in the relevant art that suitable modifications and adaptations to the compositions, formulations, methods, processes, and applications described herein can be made without departing from the scope of any embodiments or aspects thereof. The compositions and methods described are exemplary and are not intended to limit the scope of any of the specified embodiments. All of the various embodiments, aspects, and options disclosed herein can be combined in any variations or iterations. The scope of the compositions, formulations, methods, and processes described herein include all actual or potential combinations of embodiments, aspects, options, examples, and preferences herein described. The exemplary compositions and formulations described herein may omit any component, substitute any component disclosed herein, or include any component disclosed elsewhere herein. The ratios of the mass of any component of any of the compositions or formulations disclosed herein to the mass of any other component in the formulation or to the total mass of the other components in the formulation are hereby disclosed as if they were expressly disclosed. Should the meaning of any terms in any of the patents or publications incorporated by reference conflict with the meaning of the terms used in this disclosure, the meanings of the terms or phrases in this disclosure are controlling. Furthermore, the foregoing discussion discloses and describes merely exemplary embodiments. All patents and publications cited herein are incorporated by reference herein for the specific teachings thereof. Various embodiments and aspects of the inventions described herein are summarized by the following clauses: Clause 1. A computer-implemented method for training a machine learning model for predicting one or more physicochemical properties of a nanoparticle, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular representation and a property value of each nanoparticle, wherein each nanoparticle includes a drug-excipient pair including a drug molecule and an excipient molecule; generate, using the molecular representation, a descriptor and fingerprint for each nanoparticle of the set of data; creating, by using cross-validation method and the set of data, a plurality of sets of training data and a plurality of sets of test data, wherein the set of data includes the descriptor and fingerprint for each nanoparticle of the set of data; training a machine learning model of an artificial intelligence (AI) system using the plurality of sets of training data, wherein the plurality of sets of training data includes drug- excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data; and for a drug-excipient pair of the plurality of sets of test data, predicting a property of the drug-excipient pair of the plurality of sets of test data using the machine learning model as trained based on drug-excipient pairs and property values of each drug- excipient pair of the plurality of sets of training data. Clause 2. The method of clause 1, wherein the one or more physicochemical properties comprise nanoparticle formation, nanoparticle radius, or drug loading. Clause 3. The method of clause 1 or 2, wherein the machine learning model is selected from a classification model or a regression model. Clause 4. The method of any one of clauses 1–3, further comprising: assessing a performance of the classification model, using classification metrics, and the regression model, using regression metrics, wherein classification metrics include at least one of an accuracy, a precision, a recall, an area under the Receiver Operating Characteristic curve (“AUC” ROC), a F1 score, and a Mathews Correlation Coefficient (“MCC”), and wherein the regression metrics include at least one of a root mean squared error (“RMSE”), a mean absolute error (“MAE”), and a coefficient of determination (R2). Clause 5. The method of any one of clauses 1–4, wherein the predicted property of the drug- excipient pair of the plurality of sets of test data is selected from a nanoparticle formation, a nanoparticle radius, or a drug loading. Clause 6. The method of any one of clauses 1–5, wherein the property value of each nanoparticle is a nanoparticle measurement selected from a radius, a polydispersity, or a ratio of normalized intensity. Clause 7. The method of any one of clauses 1–6, wherein the cross-validation method is selected from a ten (10) fold cross-validation or a Leave-One-Out (“LOO”) cross- validation, wherein the LOO cross-validation includes a Leave-One-Excipient-Out (“LOEO”) cross-validation, a Leave-One-Drug-Out (“LODO”) cross-validation, and a Leave-One-Pair-Out (“LOPO”) cross-validation. Clause 8. A computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute the method of any one of any one of clauses 1–7. Clause 9. A computer-implemented method for training a machine learning model for predicting a formulation improvement of new molecular component modifications of nanoparticles, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular component representation and a value selected from a known exact absolute size value, or a known bound absolute size value, wherein the known exact absolute size value and the known bound absolute size value are related to a radius of each nanoparticle of the set of data, and wherein each nanoparticle of the set of data comprises a drug and an excipient; creating a set of training data with the set of data; creating a set of nanoparticle pairs using each nanoparticle of the set of training data; filtering the set of training data based on a set of rules, wherein the filtered set of training data includes nanoparticle pairs of the set of nanoparticle pairs where a nanoparticle of the nanoparticle pairs with a larger size is known; training a machine learning model of an AI system using datapoints of the filtered set of training data, wherein the datapoints include the nanoparticle pairs of the filtered set of training data with shared representations, and wherein the datapoints include physicochemical properties of the shared representations; and for datapoints of a pair of nanoparticles, predicting a classification of a formulation improvement of new molecular component modifications of nanoparticles using the machine learning model as trained based on size of the pair of nanoparticles for the datapoints, wherein the formulation improvement indicates at least one selected from a nanoparticle formation, a nanoparticle radius, or a drug loading for the pair of nanoparticles. Clause 10. The method of clause 9, wherein filtering the set of training data based on the set of rules, further comprises: removing nanoparticle pairs of the set of training data from the set of training data where whether a first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown. Clause 11. The method of clause 9 or 10, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is greater than the second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. Clause 12. The method of any one of clauses 9–11, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is less than the second nanoparticle, wherein the first nanoparticle has a known bound absolute size value a known exact absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. Clause 13. The method of any one of clauses 9–12, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle and the second nanoparticle have a known bound absolute size value, wherein the known bound absolute size value is greater than a nanoparticle size threshold. Clause 14. The method of any one of clauses 9–13, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle and a second nanoparticle of the nanoparticle pairs of the set of training data have a known bound absolute size value. Clause 15. The method of any one of clauses 9–14, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is less than a second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. Clause 16. The method of any one of clauses 9–15, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is greater than a second nanoparticle, wherein the first nanoparticle has a known bound absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold. Clause 17. The method of any one of clauses 9–16, wherein creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: splitting the set of training data into a training set and a test set. Clause 18. The method of any one of clauses 9–17, wherein creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: generating a pair of nanoparticles for at least one of the training set and the test set by cross merging a first nanoparticle and a second nanoparticle of the training set and the test set, wherein all possible nanoparticle pairs of the training set and the test set are generated. Clause 19. The method of any one of clauses 9–18, wherein the predicted classification is an output indicating whether a first nanoparticle or a second nanoparticle in the pair nanoparticles has a larger radius. Clause 20. A computer program product comprising program instructions stored on a machine-readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute the method of any one of clauses 9–19. Clause 21. An excipient conjugated to a polyethylene glycol moiety, or pharmaceutically acceptable salt thereof, selected from: Rhodamine B-PEG-1, Congo Red-PEG. Clause 22. The excipient of clause 21, wherein n is about 22 to about 225. Clause 23. A nanoparticle formulation comprising: an excipient of clause 21 or 22, or pharmaceutically acceptable salt thereof; and a therapeutic compound. Clause 24. The nanoparticle formulation of clause 23, wherein the therapeutic compound is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4- hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, or analogues or combinations thereof. Clause 25. The nanoparticle formulation of clause 23 or 24, wherein the nanoparticles have a hydrodynamic radius of about 10 nm to about 500 nm. Clause 26. The nanoparticle formulation of any one of clauses 23–25, wherein the nanoparticles have a therapeutic compound loading value of about 50% to about 95%. Clause 27. The nanoparticle formulation of any one of clauses 23–26, wherein the therapeutic compound is fulvestrant and the excipient pharmaceutically acceptable salt thereof, and the nanoparticles have a therapeutic compound loading value of 77%. Clause 28. The nanoparticle formulation of any one of clauses 23–27, wherein the nanoparticles have a stability greater than 12 hours. Clause 29. A method for treating a disease or disorder or prophylaxis thereof in a subject in need thereof, the method comprising administering a therapeutically effective amount of the nanoparticle formulation of any one of clauses 23–28 to a subject in need thereof. Clause 30. A method for reducing macrophage uptake, the method comprising administering a therapeutically effective amount of nanoparticle formulation of any one of clauses 23– 28 to a subject in need thereof. Clause 31. A method for reducing protein binding of a therapeutic compound, the method comprising administering an effective amount of the nanoparticle formulation of any one of clauses 23–28. Clause 32. Use of the excipient of clause 21 or 22, or pharmaceutically acceptable salt thereof, to modulate systemic circulation of a therapeutic compound. EXAMPLES Example 1 Data Generation and Machine Learning Model Training – Leave-One-Out Approach Drug-excipient combinations which successfully form nanoparticles were identified using Dynamic Light Scattering (DLS). Each drug-excipient combination was measured 3 times and classified as a nanoparticle if at least 2 / 3 of the measurements had the following values: Radius 200 nm Polydispersity 40% Ratio of Normalized intensity / Intensity > 10 Data quality: “Good” (based on Wyatt Dynamics software analysis) The resulting dataset of 19 excipients and 50 drugs containing 231 valid nanoparticles was used for training machine learning models to predict the probability of nanoparticle formation, nanoparticle radius, and drug loading. Drug and excipient molecules were defined using SMILES and then converted to RDKit descriptors and Morgan fingerprints (2048 bit) using RDKit (version: 2024.3.5). The predictive performance of machine learning models was assessed using different cross-validation (CV) strategies including 10-fold CV and three different versions of Leave-One- Out (LOO) CV: Leave-One-Excipient-Out (LOEO) CV, Leave-One-Drug-Out (LODO) CV and Leave-One-Pair-Out (LOPO) CV, in which a single excipient, single drug or a single pair of excipient and drug molecules was excluded from the training set. Model performance was compared using metrics including accuracy, precision, recall, area under the Receiver Operating Characteristic curve (AUC ROC), F1 score and Mathews Correlation Coefficient (MCC) for classification models, and root mean squared error (RMSE), mean absolute error (MAE) and coefficient of determination (R2) for regression models. Separate models were trained to predict the probability of nanoparticle formation (classification), nanoparticle radius (regression, both absolute radius and radius difference between two nanoparticles using a delta approach) and drug-loading (regression). The basic workflow is summarized in FIG 1D. Pairing Nanoparticles for Size Difference Comparison Approach Pairing nanoparticles for size difference comparison is a pairing approach that processes nanoparticle pairs for the application of property classification problems. This modality is specifically designed to be used for predicting which of the two paired nanoparticles has a larger radius. A unique feature of this approach is its capability of handling bounded data, allowing us to characterize large nanoparticles greater than the maximum satisfactory size of 200 nm or otherwise invalid nanoparticles as negative data. Providing this approach to a classification algorithm allows the algorithm to directly contrast molecular components of nanoparticles to guide optimization in formulation while incorporating all the available training data, even those that are bounded or too large to be accurately determined. Notably, this approach can be applied to any validated molecular machine learning process, including tree-based and deep learning models. Pairing nanoparticles for size difference comparison is especially well-suited to train machine learning models with small datasets that contain bounded or noisy data. To develop this approach, the following were performed: The pairing nanoparticles for size difference comparison approach was applied on an experimentally curated in-house data described above with 1583 unique drug-excipient valid and invalid nanoparticles for the training of a classification model. Pairing nanoparticles for size difference comparison merges training nanoparticle data comprising the components of each nanoparticle and its associated radius in nanometers. The approach then pairs the data and retains only the pairs for which the larger nanoparticle in each pair is definitively known. The inputs for this approach are paired nanoparticles where each nanoparticle consists of a drug and an excipient, totaling four molecules per input. This approach then allows models to train on nanoparticle pairs with shared representations of its components rather than individual representations. A unique feature of this modality is the additional incorporation of RDKit physicochemical properties concatenated with each of the components’ representation. The model then produced a binary output indicating whether the first or second nanoparticle in the pair has a larger radius. This modality allows classification models to learn formulation improvements between paired nanoparticles rather than learning absolute radius values from single nanoparticles. By providing a direct classification of formulation improvements, this method removes the intermediate step of subtracting predictions in order to classify formulation improvements. Example 2 Synthesis and Characterization of DSPE-PEG Congo Red Congo Red and DSPE-PEG-NHS (Congo red / DSPE-PEG-NHS: 6 / 1, mol / mol) were dissolved in 5 mL DMF, with stirring for 48 h at room temperature. Then the mixture was dialyzed extensively (dialysis bag MW cutoff 3.5 kDa) against distilled water for 48 h to remove all impurities and then lyophilized. NMR spectra were recorded on a Varian INOVA (800 MHz) spectrometer with a cryogenic probe at room temperature. All data were processed and analyzed using MNova (v 14.2.1) and 1 Hz line broadening. Chemical shifts are reported in ppm from tetramethylsilane (TMS) as the internal reference.1H NMR (600 MHz, DMSO) 8.77 (d, J = 8.5 Hz, 1H), 8.45 (d, J = 8.3 Hz, 1H), 8.32 (s, 1H), 8.12 (d, J = 7.4 Hz, 2H), 8.01–7.94 (m, 2H), 7.73 (s, 2H), 7.60 (t, J = 7.7 Hz, 1H), 7.50 (t, J = 7.6 Hz, 1H). Synthesis of IR-783 PEG IR783 (0.13 mmol), Thiol-PEG (0.067 mmol) and triethylamine (18 μL) were dissolved inanhydrous DMSO (2.25 mL), stirring continuously at room temperature for 48 h. The result crude product was dialyzed extensively (dialysis bag MW cutoff 3.5 kDa) against 70% ethanol and distilled water to remove all impurities, and then lyophilized. NMR spectra were recorded on a Varian INOVA (800 MHz) spectrometer with a cryogenic probe at room temperature. All data were processed and analyzed using MNova (v 14.2.1) and 1 Hz line broadening. Chemical shifts arereported in ppm from tetramethylsilane (TMS) as the internal reference. 1H NMR (800 MHz,DMSO-d6) 8.60–8.75 (m, 2H), 7.10–7.50 (m, 8H), 6.10–6.25 (m, 2H), 4.00–4.25 (m, 4H), 3.25– 4.00 (m, 203H), 2.75–3.00 (m, 4H), 2.50–2.70 (m, 4H), 1.75–2.00 (m, 10H), 1.50–1.70 (m, 12H). Example 3 Nanoparticle Formulations Drug solution (40 mM, 1 μL) in DMSO and excipient solution (10 mM, 1 μL) in DMSO were mixed in a microtube, followed by rapid addition of PBS (198 μL). The final drug concentration was 200 M, with a total of 1% DMSO (v / v). Triplicate 50 L samples were then transferred to a 384-well plate for high-throughput Dynamic light scattering (DLS) evaluation on a Wyatt DynaPro Plate Reader III (Wyatt Technology, USA) at 25 °C using five independent acquisitions of 5 s duration. Example 4 Drug Loading and Encapsulation Efficiency The drug loading of nanoparticles was determined using HPLC (Agilent 1200 Infinity Series, USA). The nanoparticles, prepared as previously described, were subjected to two rounds of centrifugation. After each round, the supernatant was removed, and the particles were redispersed in PBS. The particles were then dissolved in fixed volumes of acetonitrile to ensure complete dissolution. The amount of drug or excipient in the particles was calculated by subtracting the expected amount (in μg, assuming full solubility in PBS) from the actual measured concentrations. The encapsulation efficiency was determined as the ratio of the amount of drug found in the particles (in μg) to the amount of drug initially added from the stock solution during particle preparation (see Eq 1). Drug loading was calculated as the ratio of the amount of drug in the particles (in μg) to the total mass of the particles, which includes the cumulative amount of the drug (in μg) and the amount of excipient (in μg) in the particles (see Eq 2). Example 5 MH-S macrophage cell line was obtained from ATCC. MH-S cells were cultured in RPMI- 1640 medium supplemented with 10% fetal bovine serum, 0.05 mM 2-mercaptoethanol and 1% penicillin-streptomycin at 37 °C in a humidified 5% CO2 / 95% air atmosphere. Example 6 In vivo PK Study Male CD-1 mice (n = 4 per drug formulation; Avg.38 g body weight; range 35–44 g) were administered intravenously (retro-orbital sinus, under mild isoflurane anesthesia) with 2.9 mg / kg (2 mL / g body weight) of either standard nano-formulation (fulvestrant + verteporfin) or (fulvestrant + verteporfin-PEG). At 5 min, 15 min, 30 min, 1 h, 3 h, 8 h, and 24 h post-injection, ~30 mL blood was collected from tail into vials containing 1 L of 75 mg / mL K2EDTA in water, plasma was separated by centrifugation (1300 × g, 5 min) and frozen until analysis. All animal procedures were performed according to approved animal protocol (Duke University IACUC #: A107-23-04). A liquid chromatography / tandem-mass spectrometry (LC / MS / MS) assay to measure fulvestrant in mouse plasma was performed on Agilent 1200 LC, AB / Sciex API 5500 QTrap MS / MS instrument. In brief, 10 L of plasma was mixed with 30 L of 50 ng / mL fulvestrant-2H3 (internal standard) in methanol / acetonitrile (1:1), followed by vigorous agitation in Fast Prep 120 (Thermo- Savant) at speed 4.0 for 20 s, 2 cycles, and incubation for 20 min at 20 °C. After centrifugation (14,000 × g, 5 min), 30 L of supernatant was transferred to autosampler and 20 L injected into LC system. The LC / MS / MS conditions were as follows. Column: Eclipse Plus C18, 4.6 × 50 mm, 1.8 m, at 40 °C. Mobile phase A: 0.1%, formic acid, 6.5 mM ammonium formate, and 3% methanol, in water; mobile phase B: acetonitrile; flow: 1 mL / min; elution gradient (linear): 0–1 min 10–95% B, 1–1.5 min 95% B, 1.5–1.6 min 95–10% B. Run time 5 min. The analyte and internal standard were measured in positive electrospray ion mode. The following MS / MS transitions for quantification [and identity confirmation] of the respective [M+H]+ions were used. Fulvestrant: m / z 607.2 / 589.3 [607.2 / 467.1], and fulvestrant-2H3 (int. std): m / z 610.3 / 592.3 [610.3 / 467.7]. Calibration standards in 1.63 – 200 ng / mL range were prepared in drug-free mouse plasma. Lowest limit of quantification (LOQ) was 1.63 ng / mL. From the conc. / time data, non- compartmental approach within WinNonlin software (v. 2.1, Pharsight) was utilized to calculate the pharmacokinetic (PK) parameters shown in Table 1 (e.g., Tmax, Cmax, half-life, AUC, clearance, mean residence time, volume of distribution). Example 7 PEGylation of Reported Nanoparticle-Forming Excipients Building on previous work demonstrating that pentacyclic triterpenes such as 18 -glycyrrhetinic acid and glycyrrhizin can co-aggregate with drugs to form stable nanoparticles, theability of PEGylated 18 -glycyrrhetinic acid (18 -glycyrrhetinic acid-PEG) to facilitate nanoparticleformation was initially validated. To assess the nanoparticle-forming efficacy, dynamic light scattering (DLS) was employed to analyze its interaction with 50 drugs, selected based on their FDA approval status and prior inclusion in drug-excipient co-aggregation studies. DLS results indicated that 18 -glycyrrhetinic acid-PEG facilitated nanoparticle formationin 26 out of the 50 tested drugs, surpassing the performance of 18 -glycyrrhetinic acid, whichenabled nanoparticle formation for only 10 drugs. TEM imaging confirmed nanoparticlemorphology, for example for carfilzomib / 18 -glycyrrhetinic acid-PEG nanoparticle (FIG. 5).Notably, several nanoparticles that formed using either PEGylated or non-PEGylatedexcipient showed that PEGylation of 18 -glycyrrhetinic acid not only improved nanoparticleformation capabilities in terms of the numbers of drugs it can form nanoparticles with but also increased the nanoparticle stability compared to non-PEGylated excipients – suggesting the power of the PEGylated excipients to form more and more stable nanoparticles (FIG.6A–B). A key question is whether the obtained PEGylated nanoparticles exhibit high drug loading, similar to nanoparticles formed via previously established drug-excipient co-aggregation. To evaluate this, each nanoparticle was centrifuged under identical conditions, dissolved the resulting pellet in acetonitrile, and determined the drug loading, defined as the proportion of the drug incorporated into the nanoparticles, using high-performance liquid chromatography (HPLC) analysis. FIG. 7 illustrates that nanoparticles formed with 18 -glycyrrhetinic acid-PEG exhibitedexceptionally high drug loading, typically exceeding 50%. This is believed to be the first report demonstrating that a PEGylated small molecule can form nanoparticles with a drug loading capacity exceeding 50% across multiple drug types, without requiring functional group modifications tailored to specific polymers or drugs. To further evaluate the nanoparticle-forming ability of the PEGylated excipient, other excipients such as IR783, a water-soluble heptamethine cyanine dye notable for its ability to co- aggregate with over 20 drugs, was used and formed stable nanoparticles with significant drug loading. Capitalizing on IR783’s capacity to directly conjugate with thiol molecules via its meso- chloro moiety, Thiol-PEG was conjugated to IR783, yielding PEGylated IR783 (IR783-PEG). The expected structure of IR783-PEG was confirmed by1H-NMR. Absorbance spectroscopy revealed an increase in the IR783-PEGmaxat 780 nm and a decrease atmax640 nm, suggesting that PEGylation of IR783 impedes the formation of indocyanine H-type aggregates (FIG.8). HPLC analysis also confirmed the presence of IR783-PEG as a single peak (FIG.9). DLS results indicated that IR783-PEG could form nanoparticles with 24 of the 50 tested drugs, closely mirroring the behavior of IR783. FIG.10 shows the nanoparticle size distribution and a TEM image of an example of the obtained nanoparticle, fulvestrant / IR783-PEG nanoparticle. IR783-PEG nanoparticles also exhibited a high drug loading capacity exceeding 50% (FIG.11). While nanoparticles formed with IR783 and IR783-PEG demonstrated similar stability in PBS overall, notable differences were observed for specific drugs. Specifically, IR783 formed more stable nanoparticles with Dutasteride, Rilpivirine, and Tacrolimus, whereas Ospemifene exhibited greater stability with IR783-PEG. These differences may be attributable to the drug-stabilizing mechanism of planar IR783, which is primarily driven by hydrogen bonding and -interactions. PEGylation at the meso-chloro moiety significantly altered these interactions, leading to changes in nanoparticle stability. This finding underscores the importance of considering the modification site when designing PEGylated excipients. Conventionally, PEGylating agents capable of encapsulating drugs feature a hydrophobic moiety at the PEG terminus. However, these results demonstrate that even a highly water-soluble excipient like IR783, when PEGylated, can still form nanoparticles with a high drug loading capacity. Additionally, these results suggest the versatility of the nanoparticle design concept and highlight its potential applicability beyond traditional hydrophobic PEGylation strategies. Enhancing Nanoparticle Formation of Ineffective Excipients by PEGylation The ability of PEGylation to enhance the nanoparticle-forming ability of excipients previously deemed ineffective in drug interactions for nanoparticle formation was evaluated. Four excipients—stearic acid, Rhodamine B, biotin, and succinic acid — were selected based on the commercial availability of their PEGylated derivatives. PEGylated Rhodamine B was examined in two different molecules with distinct PEG modification sites. While PEGylation did not enhance nanoparticle formation for biotin and succinic acid, PEGylated stearic acid exhibited an increased capacity to stabilize drugs and facilitate nanoparticle formation. Notably, the two PEGylated Rhodamine B derivatives demonstrated different nanoparticle-forming abilities, underscoring the critical role of the PEG modification site and the nuanced molecular recognition between the drug and excipient in nanoparticle formation. Although numerous studies have reported nanoparticles incorporating Rhodamine B within their hydrophobic core or utilizing its PEG-end modification for fluorescent labeling, this study provides a novel example in which Rhodamine B itself contributes to nanoparticle formation. Overall, these findings suggest that PEGylation can broaden the scope of nanoparticle design by enabling the co-assembly of drugs and excipients, thereby overcoming the constraints imposed by the limited diversity of excipients capable of stabilizing drugs. Large Scale Generation of PEGylated Excipient Nanoparticles Encouraged by these results, a large dataset of drug-excipient nanoparticles was created. The dataset comprised 29 excipients in total, with 20 PEGylated and 9 non-PEGylated excipients. These excipients were screened against 57 small molecule drugs, resulting in a total of 1583 unique drug-excipient nanoparticle formulations (FIG. 2). Each formulation was characterized using Dynamic Light Scattering (DLS) to evaluate nanoparticle formation. The resulting dataset of drug-excipient combinations was used to develop a classification model. From this library, a subset of 19 PEGylated excipients and 50 drugs, yielding 231 valid nanoparticles based on DLS measurements, was used for training the regression models. Machine Learning to Predict Nanoparticle Properties Although the synthesis of the here described nanoparticles is comparatively simple, the combinatorial synthesis and characterization of millions of materials can still be restrictively laborious and costly. To support this process, machine learning is now an established tool to triage experiments and identify the most promising materials. The ability to train predictive algorithms to anticipate nanoparticle properties such as (1) nanoparticle stability, (2) nanoparticle size, and (3) nanoparticle drug loading. A range of machine learning models based on RDKit descriptors and / or Morgan fingerprints were trained to predict whether drug-excipient combinations could successfully form nanoparticles. A number of models showed strong predictive performance and could be used to identify novel nanoparticle forming drug-excipient pairs. The best performing model, a random forest trained on RDKit descriptors, achieved an AUC ROC of 0.96 ± 0.02 from 10-fold cross- validation. In addition, a similar approach was used to produce a number of regression models to predict nanoparticle radius. Many models showed satisfactory performance, with the best performing model, also a random forest, obtaining an RMSD of 22 ± 1 nm. To further investigate which drug-excipient pairs influence nanoparticle size for optimized formulation outcomes, a pair-wise machine learning approach to classify which of two nanoparticles is larger was constructed and subsequently augmented and derived from an approach previously called “DeltaClassifier.” Applying this approach with Chemprop, a direct message-passing neural network algorithm known for its robust performance, to the complete dataset of PEGylated and non-PEGylated excipients generated solid predictive performance. The approach achieved an AUC ROC of 0.84 ± 0.01 and average accuracy of 0.76 ± 0.01 (FIG.3A and 3B). These favorable results highlight the effectiveness of this approach to accurately identify size differences between nanoparticle pairs and guide formulation optimization. The role of PEGylated nanoparticle composition in maximizing drug loading was also evaluated. Specifically, a model to predict drug loading from a given nanoparticle was generated. Because drug loading was investigated in a subset of nanoparticles, Random Forest was selected to handle the limited regression data. The results showed promising performance of such an approach with RMSE < 0.14 when performing leave-one-excipient-out (LOEO) validations (FIG. 4), suggesting that the predictive models can accurately predict the drug encapsulation potential of a novel excipient across the tested drugs. Characterizing Biointerfaces of PEGylated Nanoparticles PEGylation is known to reduce protein binding, which can be important to impact the physiological behavior and properties of nanoparticles. To assess whether the PEGylated nanoparticles impact protein absorption, the PEGylated nanoparticles were assessed in vitro. To this end, nanoparticles that were formed with either the PEGylated excipient or the equivalent non-PEGylated excipient were created. Then the nanoparticles were incubated with BSA and subsequently performed centrifugation-based purification to extract our nanoparticles. The protein content of the nanoparticle solution was assessed either using a BCA assay or measuring fluorescence when using FITC-labeled protein. This analysis revealed that our PEGylated excipients significantly reduce opsonization (FIG.12). Similarly, PEGylation was expected to reduce immune recognition. Macrophage uptake of the PEGylated nanoparticles was compared to the equivalent non-PEGylated nanoparticles. The PEGylated nanoparticles completely halted macrophage uptake of the nanoparticles in vitro. This immune escape has again important implications for the in vivo behavior of these nanoparticles (FIG.13). In vivo Pharmacokinetic Analysis of PEGylated Nanoparticles Pharmacokinetic (PK) analysis revealed that the fulvestrant + verteporfin-PEG nanoparticle formulation exhibited a markedly higher terminal half-life (t½= 6.7 h vs.3.07 h) and mean residence time (MRT = 5.73 h vs. 1.04 h) compared to the non-PEGylated equivalent nanoparticle, indicating that prolonged systemic circulation can be achieved with the PEGylated formulations compared to non-PEGylated nanoparticles. Notably, the Cmaxwas more than two- fold lower with verteporfin-PEG (2.37 μg / mL vs.5.76 μg / mL), and the clearance (Cl) was reduced (2177 vs.2966 mL / h·kg), suggesting slower elimination. The concentration vs time profile (FIG. 14) further supports this observation, showing a slower decline in Fulvestrant levels when co- administered with Verteporfin-PEG. Together, these data suggest that PEGylation may enhance nanoparticle pharmacokinetics by reducing clearance and extending systemic exposure.
Claims
1. CLAIMS What is claimed:
1. A computer-implemented method for training a machine learning model for predicting one or more physicochemical properties of a nanoparticle, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular representation and a property value of each nanoparticle, wherein each nanoparticle includes a drug-excipient pair including a drug molecule and an excipient molecule; generate, using the molecular representation, a descriptor and fingerprint for each nanoparticle of the set of data; creating, by using cross-validation method and the set of data, a plurality of sets of training data and a plurality of sets of test data, wherein the set of data includes the descriptor and fingerprint for each nanoparticle of the set of data; training a machine learning model of an artificial intelligence (AI) system using the plurality of sets of training data, wherein the plurality of sets of training data includes drug- excipient pairs and property values of each drug-excipient pair of the plurality of sets of training data; and for a drug-excipient pair of the plurality of sets of test data, predicting a property of the drug-excipient pair of the plurality of sets of test data using the machine learning model as trained based on drug-excipient pairs and property values of each drug- excipient pair of the plurality of sets of training data.
2. The method of claim 1, wherein the one or more physicochemical properties comprise nanoparticle formation, nanoparticle radius, or drug loading.
3. The method of claim 1, wherein the machine learning model is selected from a classification model or a regression model.
4. The method of claim 3, further comprising: assessing a performance of the classification model, using classification metrics, and the regression model, using regression metrics, wherein classification metrics include at least one of an accuracy, a precision, a recall, an area under the Receiver Operating Characteristic curve (“AUC” ROC), a F1 score, and a MathewsCorrelation Coefficient (“MCC”), and wherein the regression metrics include at least one of a root mean squared error (“RMSE”), a mean absolute error (“MAE”), and a coefficient of determination (R2).
5. The method of claim 1, wherein the predicted property of the drug-excipient pair of the plurality of sets of test data is selected from a nanoparticle formation, a nanoparticle radius, or a drug loading.
6. The method of claim 1, wherein the property value of each nanoparticle is a nanoparticle measurement selected from a radius, a polydispersity, or a ratio of normalized intensity.
7. The method of claim 1, wherein the cross-validation method is selected from a ten (10) fold cross-validation and a Leave-One-Out (“LOO”) cross-validation, wherein the LOO cross-validation includes a Leave-One-Excipient-Out (“LOEO”) cross-validation, a Leave- One-Drug-Out (“LODO”) cross-validation, or a Leave-One-Pair-Out (“LOPO”) cross- validation.
8. A computer program product comprising program instructions stored on a machine- readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute the method of claim 1.
9. A computer-implemented method for training a machine learning model for predicting a formulation improvement of new molecular component modifications of nanoparticles, the method comprising: receiving a set of data including nanoparticles, wherein each nanoparticle of the set of data includes a molecular component representation and a value selected from a known exact absolute size value or a known bound absolute size value, wherein the known exact absolute size value and the known bound absolute size value are related to a radius of each nanoparticle of the set of data, and wherein each nanoparticle of the set of data comprises a drug and an excipient; creating a set of training data with the set of data; creating a set of nanoparticle pairs using each nanoparticle of the set of training data;filtering the set of training data based on a set of rules, wherein the filtered set of training data includes nanoparticle pairs of the set of nanoparticle pairs where a nanoparticle of the nanoparticle pairs with a larger size is known; training a machine learning model of an AI system using datapoints of the filtered set of training data, wherein the datapoints include the nanoparticle pairs of the filtered set of training data with shared representations, and wherein the datapoints include physicochemical properties of the shared representations; and for datapoints of a pair of nanoparticles, predicting a classification of a formulation improvement of new molecular component modifications of nanoparticles using the machine learning model as trained based on size of the pair of nanoparticles for the datapoints, wherein the formulation improvement indicates at least one selected from a nanoparticle formation, a nanoparticle radius, or a drug loading for the pair of nanoparticles.
10. The method of claim 9, wherein filtering the set of training data based on the set of rules, further comprises: removing nanoparticle pairs of the set of training data from the set of training data where whether a first nanoparticle or a second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown.
11. The method of claim 10, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is greater than the second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold.
12. The method of claim 10, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises:removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle is less than the second nanoparticle, wherein the first nanoparticle has a known bound absolute size value a known exact absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold.
13. The method of claim 10, wherein removing the nanoparticle pairs of the set of training data from the set of training data where whether the first nanoparticle or the second nanoparticle of the nanoparticle pairs of the set of training data is larger in size is unknown, further comprises: removing the nanoparticle pairs of the set of training data from the set of training data when the first nanoparticle and the second nanoparticle have a known bound absolute size value, wherein the known bound absolute size value is greater than a nanoparticle size threshold.
14. The method of claim 9, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle and a second nanoparticle of the nanoparticle pairs of the set of training data have a known bound absolute size value.
15. The method of claim 9, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is less than a second nanoparticle, wherein the first nanoparticle has a known exact absolute size value and the second nanoparticle has a known bound absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold.
16. The method of claim 9, wherein filtering the set of training data based on the set of rules, further comprises: keeping the nanoparticle pairs of the set of training data in the set of training data when a first nanoparticle is greater than a second nanoparticle, wherein the firstnanoparticle has a known bound absolute size value and the second nanoparticle has a known exact absolute size value, and wherein the known bound absolute size value is greater than a nanoparticle size threshold.
17. The method of claim 9, wherein creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: splitting the set of training data into a training set and a test set.
18. The method of claim 17, wherein creating the set of nanoparticle pairs using each nanoparticle of the set of training data, further comprises: generating a pair of nanoparticles for at least one of the training set and the test set by cross merging a first nanoparticle and a second nanoparticle of the training set and the test set, wherein all possible nanoparticle pairs of the training set and the test set are generated.
19. The method of claim 9, wherein the predicted classification is an output indicating whether a first nanoparticle or a second nanoparticle in the pair nanoparticles has a larger radius.
20. A computer program product comprising program instructions stored on a machine- readable storage medium, wherein when the program instructions are executed by a computer processor, the program instructions cause the computer processor to execute the method of claim 9. . An excipient conjugated to a polyethylene glycol moiety, or pharmaceutically acceptable salt thereof, selected from:Folate-PEG,Congo Red-PEG.
22. The excipient of claim 21, wherein n is about 22 to about 225.
23. A nanoparticle formulation comprising: an excipient of claim 21, or pharmaceutically acceptable salt thereof; and a therapeutic compound.
24. The nanoparticle formulation of claim 23, wherein the therapeutic compound is selected from afatinib, alpelisib, bicalutamide, binimetinib, bithionol, cabozantinib, carfilzomib, celecoxib, curcumin, cyclosporin, dabrafenib, dasatinib, docetaxel, dutasteride, duvelisib, econazole, entrectinib, enzalutamide, erlotinib, etravirine, everolimus, fulvestrant, gefinitib, glibenclamide, glimepiride, lapatinib, lumefantrine, midostaurin, nelfinavir, nilotinib, ospemifene, paclitaxel, panobinostat, pazopanib, ponatinib, probucol, rapamycin, regorafenib, rilpivirine, selumetinib, sorafenib, sunitinib, tacrolimus, talazoparib, tazarotene, terbinafine, valdecoxib, valruybicin, vemurafenib, 4-hydroxytamoxifen, boscalid, chlorotrianisene, IOWH-032, OSI-930, serdemetan, vatalanib, venetoclax, or analogues or combinations thereof.
25. The nanoparticle formulation of claim 23, wherein the nanoparticles have a hydrodynamic radius of about 10 nm to about 500 nm.
26. The nanoparticle formulation of claim 23, wherein the nanoparticles have a therapeutic compound loading value of about 50% to about 95%.
27. The nanoparticle formulation of claim 23, wherein the therapeutic compound is fulvestrant and the excipientor pharmaceuticallyacceptable salt thereof, and the nanoparticles have a therapeutic compound loading value of 77%.
28. The nanoparticle formulation of claim 23, wherein the nanoparticles have a stability greater than 12 hours.
29. A method for treating a disease or disorder or prophylaxis thereof in a subject in need thereof, the method comprising administering a therapeutically effective amount of the nanoparticle formulation of claim 23 to a subject in need thereof.
30. A method for reducing macrophage uptake, the method comprising administering a therapeutically effective amount of nanoparticle formulation of claim 23 to a subject in need thereof.
31. A method for reducing protein binding of a therapeutic compound, the method comprising administering an effective amount of the nanoparticle formulation of claim 23.
32. Use of the excipient of claim 21, or pharmaceutically acceptable salt thereof, to modulate systemic circulation of a therapeutic compound.
Citation Information
Patent Citations
Methods for drug screening and compositions useful for the inhibition of cell proliferation and / or cell survival
US20230093878A1
Formulation of API's and excipient's in an icell via hydrophobic ion pairing
WO2023227517A1
Sprayed multi adsorbed-droplet reposing technology (SMART)
WO2023283320A2
Methods for characterisation of nanocarriers
WO2024047246A1