Machine learning guided preparation of albumin nanosensors and applications thereof
Patent Information
- Application Number
- CN202611000594.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
第一类是机器学习辅助纳米材料合成,已有研究采用随机森林或神经网络模型预测纳米粒的粒径或产率,以减少实验试错次数,但这些模型的预测目标仅局限于物理化学性质,从未将降解传感性能作为优化对象,且多数模型作为“黑箱”使用,缺乏对合成参数间非线性相互作用的可解释性分析,更未与自指示功能相整合
制备方法绿色高效: 采用完全的水相体系,避免了有机溶剂、化学交联剂等破坏蛋白构象的有害物质。通过机器学习优化,系统化地解决了传统试错法效率低、难以平衡多目标性能的问题,将研发周期显著缩短。
Smart Images

Figure CN122814554A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nanomaterials technology, and in particular to a machine learning-guided method for preparing albumin nanosensors and their applications. Background Technology
[0002] Albumin nanoparticles (ANPs) are currently recognized as one of the "gold standard" delivery carriers in the field of nanomedicine. Human serum albumin and bovine serum albumin possess excellent biocompatibility, low immunogenicity, biodegradability, and natural liver targeting properties. Albumin-based nanomedicines, such as Abraxane®, have been successfully used in clinical cancer treatment. However, existing ANP preparation and functionalization technologies have long faced two interrelated challenges. One is the destructive nature of the preparation methods. Traditional processes rely on chemical cross-linking (such as glutaraldehyde), organic solvent denaturation, or thermally induced polymerization. These operations irreversibly alter the natural tertiary structure of albumin, masking its surface functional motifs such as protease recognition sites. This makes it difficult for the cross-linked nanoparticles to be efficiently degraded by lysosomal proteases, and may also trigger unnecessary immune responses. Secondly, there is a lack of real-time tracking technology for intracellular degradation. Conventional methods such as electron microscopy can only provide static images, radioactive isotope labeling cannot distinguish between intact carriers and degraded fragments, and exogenous fluorescent dye labeling suffers from aggregation-induced quenching, meaning that the signal is weak when the fluorophore is densely packed, but strengthens after dispersion, resulting in false positives due to the "dark when intact, bright after degradation." Therefore, existing technologies cannot provide a direct, quantitative, and indicator-free self-monitoring method that shows "strong fluorescence when the carrier structure is intact and weak fluorescence when the structure is damaged."
[0003] Among existing literature and patents, the implementation schemes most similar to this application can be mainly divided into three categories. The first category is machine learning-assisted nanomaterial synthesis. Existing studies have used random forest or neural network models to predict the particle size or yield of nanoparticles to reduce the number of experimental trials. However, the prediction targets of these models are limited to physicochemical properties and have never taken degradation sensing performance as an optimization object. Moreover, most models are used as "black boxes" and lack interpretable analysis of the nonlinear interactions between synthesis parameters, and are not integrated with self-indication functions. The second category is aggregation-induced emission (AIE) molecule-albumin complex system. AIE molecules exhibit significantly enhanced fluorescence in aggregated or confined environments, which precisely overcomes ACQ defects. Existing studies have mixed AIE molecules with albumin to form complexes for detecting albumin concentration or conformational changes. However, these schemes usually only yield unstable aggregates with a size of less than 10 nm, rather than sensors with a well-defined nanostructure (such as 85 nm, low PDI, and scalable). Furthermore, they have not revealed the precise binding sites and hierarchical assembly mechanisms between AIE molecules and albumin, and have not utilized the fluorescence-structure coupling characteristics of AIE to track the enzymatic degradation of nanoparticles in real time. The third category is the degradation study of chemically cross-linked albumin nanoparticles, which is currently the most classic method—locking nanoparticles by forming a covalent Schiff base between glutaraldehyde and lysine on the albumin surface. However, cross-linking rigidifies the albumin backbone, masking the recognition sites of natural proteases, resulting in slow degradation in lysosomes. More importantly, quantitative proteomics has shown that the plasma proteins adsorbed by cross-linked albumin nanoparticles are mainly enriched in complement and immunoglobulins (forming pro-inflammatory protein crowns), rather than proteases, which further weakens its degradation sensing ability. If tracing is required, additional covalently coupled fluorescent dyes must be used, but dye labeling still faces problems such as ACQ, leakage, and photobleaching.
[0004] In summary, existing technologies have not yet integrated the five key elements of "machine learning-based precise synthesis, non-covalent assembly of native conformations, multi-site driven hierarchical structures based on amphiphilicity, intrinsic fluorescence-degradation linear coupling via AIE, and enhancement of protein crowns by protease enrichment," thus failing to meet the urgent need for in-situ, real-time, and self-indicating monitoring of protein degradation processes within living cells. This invention aims to fill this long-standing technological gap. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a machine learning-guided method for preparing albumin nanosensors.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: The machine learning-guided albumin nanosensor preparation method is a systematic closed loop that covers the entire process of "data-driven optimization → green synthesis → performance characterization → functional verification".
[0007] (1) Multi-objective optimization and green synthesis method based on machine learning a. Data Acquisition: This invention is based on a non-covalent self-assembly reaction. Six key synthetic parameters were selected for systematic experimental design, and 1013 sets of experimental data were collected. Specific parameters were: BSA concentration (10-100 μmol / L), AIE molecule (TEZ-TPE-1) concentration (10-120 μmol / L), reaction temperature (15-40 ℃), stirring rate (300-800 rpm), reaction time (20-120 min), and pH value (7.0-8.0). For each synthesized product, four core performance indicators were measured: hydrated particle size, synthesis yield, fluorescence intensity at characteristic wavelength, and long-term (30 days) colloidal stability.
[0008] b. Feature engineering and model building: After cleaning the data, 31 input features were innovatively constructed. In addition to the original 6 parameters, second-order polynomial interaction features (such as BSA concentration × TEZ-TPE-1 concentration) and ratio features (such as BSA / TEZ-TPE-1 concentration ratio) were introduced to capture complex nonlinear relationships. After normalization using a standard Scaler, advanced machine learning models such as Gradient Boosting Regression Tree (GBRT) and Extreme Gradient Boosting (XGBoost) were used for training and five-fold cross-validation.
[0009] c. Model interpretation and parameter optimization: SHAP (SHapley Additive exPlanations) was used to perform interpretability analysis on the model, revealing the specific direction and intensity of the influence of each synthesis parameter on performance indicators. For example, the analysis found that the combination of low pH (7.0-7.2) and high stirring rate (600-800 rpm) in the synthesis reaction produces a negative synergistic effect (negative SHAP value), which is significantly beneficial to the synthesis of nanoparticles with smaller particle sizes. This key insight is difficult to discover using traditional single-factor rotation methods. Based on the optimization model, Pareto front analysis was used to find a balance among four conflicting objectives: minimizing particle size, maximizing yield, maximizing fluorescence intensity, and optimizing stability. Finally, a set of Pareto optimal synthesis parameters was determined: BSA concentration 60 μmol / L, TEZ-TPE-1 concentration 50 μmol / L, pH 7.4, temperature 25℃, stirring rate 500 rpm, and reaction time 120 min.
[0010] d. Green non-covalent self-assembly synthesis steps: Based on the above optimal parameters, the specific synthesis is carried out as follows: Step A: Prepare stock solutions of BSA and TEZ-TPE-1 (both at a concentration of 1 mmol / L).
[0011] Step B: Take the specified amount of BSA solution, dilute it to near the final volume with pure water, and adjust the pH to the optimal value with HCl or NaOH.
[0012] Step C: Preheat and stir the system in a thermostatic magnetic stirrer.
[0013] Step D: Slowly add the AIE molecule (TEZ-TPE-1) stock solution to the BSA solution, controlling the addition rate (e.g., 10 μL / min) to ensure slow and uniform mixing and drive non-covalent self-assembly.
[0014] Step E: Continue stirring the reaction under optimal conditions for the specified time to promote the stable formation of nanoparticles.
[0015] Step F: The reaction solution was purified by ultrafiltration and centrifugation, washed multiple times with pure water, and then resuspended in physiological buffer to obtain the final albumin nanosensor (ANPs) product.
[0016] (2) Characteristics and working mechanism of self-indicating nanosensors The ANPs synthesized in this invention have the following characteristics: Morphology and particle size: Transmission electron microscopy showed that the obtained nanoparticles were spherical with uniform particle size distribution and an average particle size of about 85 nm (based on machine learning prediction and experimental verification). The PDI was less than 0.2, indicating good colloidal stability.
[0017] Signal Principle: The core innovative mechanism of this invention lies in utilizing the rotationally restricted (RIR) effect of AIE molecules. TEZ-TPE-1 molecules are loaded into the hydrophobic pockets of BSA or adsorbed onto BSA molecules through hydrophobic interactions and π-π stacking, forming nanoassemblies. When the nanoparticle structure is intact, the movement (internal rotation) of TEZ-TPE-1 molecules is highly restricted by the tight assembly structure of BSA, non-radiative transition channels are suppressed, and fluorescence intensity is high. When the nanoparticles are degraded under the action of enzymes (such as lysosomal cathepsins), their assembly structure disintegrates, the degrees of freedom of movement of TEZ-TPE-1 molecules increase, non-radiative transitions are enhanced, leading to proportional quenching of fluorescence intensity.
[0018] Performance Validation: In a buffer simulating a lysosomal environment, the fluorescence intensity of ANPs decreased in a time-dependent manner with the addition of protease, and the rate of decrease was highly linearly correlated with the degradation rate of BSA (R² up to 0.9828). This "bright when intact, dark when degraded" signal pattern is the complete opposite of the ACQ effect of traditional fluorescent dyes ("dark when aggregated, bright when degraded"), achieving true self-indication and fundamentally avoiding false positives.
[0019] (3) Application methods in living cells Taking liver cancer cells (HepG2) and normal liver cells (THLE-2) as examples: Cellular uptake: When ANPs were co-incubated with cells for different times, confocal microscopy revealed that the ANPs were endocytosed and primarily localized in lysosomes (confirmed by LysoTracker-Green colocalization, with a high Pearson correlation coefficient). Flow cytometry quantification showed that cellular uptake was time-dependent.
[0020] Real-time tracking of intracellular degradation: During long-term co-culture of ANPs with cells (up to 168 hours), the overall intracellular fluorescence intensity was measured periodically and non-destructively using a fluorescence microplate reader or real-time fluorescence imaging system. In control cells, the fluorescence signal gradually weakened over time, demonstrating the gradual degradation of nanoparticles within lysosomes. This process could be plotted as a precise degradation kinetic curve.
[0021] Validation and Expansion: Treatment can be achieved by adding specific protease inhibitors (such as E64d targeting cysteine proteases and Pepstatin A targeting aspartic proteases). The attenuation of cellular fluorescence signals in the inhibitor-treated group was significantly suppressed, further demonstrating that this attenuation of fluorescence signals originates from protease-catalyzed nanoparticle degradation, and can be used to study the degradation contributions of different types of lysosomal proteases.
[0022] Compared with the prior art, the present invention has the following advantages: The preparation method is green and efficient: It adopts a completely aqueous system, avoiding harmful substances such as organic solvents and chemical cross-linking agents that damage protein conformation. Through machine learning optimization, it systematically solves the problems of low efficiency and difficulty in balancing multi-objective performance in traditional trial-and-error methods, significantly shortening the research and development cycle.
[0023] The sensor exhibits stable structure and intelligent function: The synthesized nanosensors are uniform in size (~85nm), exhibit good colloidal stability (change less than 5% within 30 days), and completely retain the natural conformation of albumin. This is the physical basis for its ability to be efficiently recognized and degraded by intracellular natural proteases.
[0024] Unique "degradation-quenching" self-indication mechanism: This invention cleverly utilizes the RIR effect of AIE molecules to linearly couple fluorescence intensity with the physical structural integrity of nanoparticles, avoiding signal interference and false positive problems caused by the ACQ effect of traditional labeling, and realizing truly accurate in-situ, real-time, and quantitative monitoring of the degradation process.
[0025] Providing powerful research and translational tools: This invention offers a general-purpose technology platform integrating machine learning-aided design, self-assembly preparation, and self-indication tracking. This platform can be used not only for intracellular pharmacokinetic and pharmacodynamic studies of albumin nanomedicines (such as generic Abraxane and albumin-bound novel drugs), but also provides innovative ideas for the development of other intelligent responsive biomaterials. Attached Figure Description
[0026] Figure 1 Contour plots showing the predictive performance of three machine learning models (random forest, XGBoost, and MLP) on particle size and yield.
[0027] Figure 2 The fluid dynamics dimensional changes of BSA mixed with different concentrations of TEZ-TPE-1 were measured for dynamic light scattering.
[0028] Figure 3 (a) is a transmission electron microscope image of ANPs; (b) is a histogram of the particle size distribution of ANPs.
[0029] Figure 4 The cellular uptake kinetics curves of ANPs from normal hepatocytes and liver cancer cells are shown.
[0030] Figure 5 This is a confocal image showing the colocalization of ANPs in living cells, demonstrating the colocalization of ANPs with lysosomes.
[0031] Figure 6 Figure (a) shows the intracellular degradation kinetics curves of ANPs in normal hepatocytes and hepatocellular carcinoma cells; Figure (b) shows the heatmap analysis of the degradation status of ANPs in cells treated with different lysosomal protease inhibitors. Detailed Implementation
[0032] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0033] (a) Data collection and preprocessing Step 1: Data Acquisition for Synthesis Experiment Following single-factor rotation or orthogonal experimental design, a total of 1013 experimental groups were set up, including BSA (bovine serum albumin) concentrations (10, 20, 30, 40, 50, 60, 70, 80, 90, 100 μmol / L), TEZ-TPE-1 (tetraazole functionalized tetraphenylethylene derivative) concentrations (10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120 μmol / L), reaction temperatures (15, 20, 25, 30, 35, 40 ℃), stirring rates (300, 400, 500, 600, 700, 800 rpm), reaction times (20, 40, 60, 80, 100, 120 min), and pH values (7.0, 7.2, 7.4, 7.6, 7.8, 8.0). ANPs were prepared in each group of experiments according to the green synthesis method described below, and the following four performance indicators were measured: particle size (nm, dynamic light scattering), synthesis yield (mg, lyophilized weight), fluorescence intensity (normalized value, λ_ex=365 nm, λ_em=501 nm), and colloidal stability (percentage change in particle size after 30 days of storage, with PDI (polydispersity index) <0.3 at 4℃ as the stability criterion).
[0034] Step 2: Data Cleaning and Feature Engineering Outlier samples with particle sizes greater than 300 nm or yields less than 0.1 mg were removed using the interquartile range method. For the remaining data, six original parameters were used as basic features, and second-order polynomial interaction features (e.g., BSA concentration × TEZ-TPE-1 concentration, pH × stirring rate) and ratio features (e.g., BSA / TEZ-TPE-1 concentration ratio) were generated, resulting in 31 input features. The features were standardized using StandardScaler in scikit-learn, ensuring a mean of 0 and a standard deviation of 1 for each feature.
[0035] (II) Machine Learning Model Training and Optimization Step 3: Dataset Partitioning and Model Selection The preprocessed dataset was randomly divided into training and test sets in an 8:2 ratio. Five-fold cross-validation was used on the training set to evaluate the performance of the following regression models: linear regression, random forest (100 trees, maximum depth 10), gradient boosting regression tree (GBRT, learning rate 0.1, 100 trees), extreme gradient boosting (XGBoost, learning rate 0.1, maximum depth 6), and multilayer perceptron (MLP, hidden layer (100, 50), activation function ReLU).
[0036] like Figure 1As shown, contour plots of granular size and yield are provided between experimental values and machine learning predictions. Among them, (a) is the Random Forest model, (b) is the XGBoost model, and (c) is the MLP model. In (a), (b), and (c), the left side represents the performance in terms of granular size, and the right side of the graph represents the performance in terms of yield.
[0037] Step 4: Model Training and Hyperparameter Tuning For each model, hyperparameter tuning was performed using grid search combined with five-fold cross-validation. Taking particle size prediction as an example, the optimal hyperparameters for GBRT were: 150 trees, a maximum depth of 5, and a minimum number of splits of samples of 4. After training, R², root mean square error (RMSE), and mean absolute error (MAE) were calculated on the test set. The best-performing model (using GBRT for particle size and fluorescence / stability, and XGBoost for yield) was retained for subsequent predictions.
[0038] Step 5: SHAP Interpretability Analysis The contribution of each feature to the model output was calculated using the SHAP library. For the particle size model, a SHAP summary plot was plotted to identify TEZ-TPE-1 concentration, BSA concentration, pH, and stirring rate as the four most important features. The interaction SHAP values between pH and stirring rate were analyzed, confirming that the combined effect of low pH (7.0-7.2) and high stirring rate (600-800 rpm) resulted in a negative SHAP value (which is beneficial for reducing particle size), thus revealing a nonlinear synergistic effect.
[0039] (III) Experimental verification of optimal synthesis parameters and determination of the Pareto front Step 6: Multi-objective Pareto optimization The four optimization objectives were particle size (minimization), yield (maximization), fluorescence intensity (maximization), and colloidal stability (constraint: particle size change <5% over 30 days). A trained GBRT and XGBoost model was used to perform grid prediction in the parameter space (step size: BSA concentration 10 μmol / L, TEZ-TPE-1 concentration 10 μmol / L, pH 0.2, temperature 2℃, stirring speed 100 rpm, time 20 min) to screen for non-dominated solutions. An equilibrium solution was selected as the Pareto optimal formulation: BSA concentration 60 μmol / L, TEZ-TPE-1 concentration 50 μmol / L, pH 7.4, temperature 25℃, stirring speed 500 rpm, time 120 min.
[0040] Step 7: Experimental Verification Three independent batches (200 mL each) were synthesized according to the parameters determined in step 6. The actual particle size, yield, fluorescence intensity, and colloidal stability were measured and compared with the model predictions. A relative error of less than 5% confirms the reliability of the model.
[0041] (iv) Green non-covalent self-assembly preparation of ANPs Step 8: Preparation of Mother Liquor Weigh out BSA powder (purity ≥98%) and dissolve it in pure water to prepare a 1 mmol / L stock solution. Store at 4°C for later use. Weigh out TEZ-TPE-1 powder and dissolve it in a small amount of pure water (ultrasonic dissolution for 5 min) to prepare a 1 mmol / L stock solution. Store at 4°C protected from light.
[0042] Step 9: Assembly reaction Take 60 μL of BSA stock solution and add pure water to a total volume of approximately 900 μL. Adjust the pH to 7.4 with 0.1 mol / L HCl or NaOH. Place in a thermostatic magnetic stirrer and preheat at 25°C and 500 rpm for 30 min. Take 50 μL of TEZ-TPE-1 stock solution, dilute with pure water to 100 μL, and add it dropwise to the BSA solution at a rate of 10 μL / min (a micro-injection pump can be used for control). After the addition is complete, continue stirring at 25°C and 500 rpm for 120 min. The synthesis route is shown in Figure 2.
[0043] Step 10: Purification Transfer the reaction solution to a 3 kDa ultrafiltration centrifuge tube (Millipore, 15 mL) and centrifuge at 3000 × g for 15 min at 4 °C. Discard the filtrate, resuspend the precipitate in 10 mL of pure water, and centrifuge again. Repeat three times. Finally, resuspend the purified ANPs in 1 mL of physiological saline or PBS (phosphate-buffered saline) buffer, lyophilize, or store at 4 °C for later use.
[0044] (V) Characterization of the physicochemical properties of ANPs Step 12: Particle size determination The purified ANPs were diluted with pure water to 0.1 mg / mL and injected into disposable cuvettes. The hydration kinetic diameter and PDI were measured at 25 °C using a dynamic light scattering instrument (Malvern Zetasizer Nano ZS). Each sample was measured three times, and each measurement consisted of 15 runs.
[0045] like Figure 2 As shown, it illustrates the dynamic size changes of bovine serum albumin (BSA, 60 μM) mixed with different concentrations (10–120 μM) of TZE-TPE-1.
[0046] Step 13: Morphological observation ANPs were diluted to 0.05 mg / mL, and 10 μL was added to a copper mesh (supported by a carbon film). After air drying, the mesh was observed and photographed using a transmission electron microscope (accelerating voltage 80 kV). The diameters of at least 100 particles were counted, and a particle size distribution histogram was plotted, as shown below. Figure 3 As shown.
[0047] Step 14: Fluorescence Spectroscopy Measurement The lyophilized ANPs powder was redissolved in pure water to prepare a series of concentrations of 0, 0.2, 0.5, 1.0, 1.5, and 2.0 mg / L. Emission spectra were recorded using a fluorescence spectrophotometer (excitation wavelength 365 nm, slit width 5 nm, emission scan range 400-700 nm). A standard curve and detection limit (LOD = 3.3 × blank standard deviation / slope) were obtained by plotting the fluorescence intensity at 501 nm against the ANPs concentration and performing linear regression.
[0048] Step 15: Long-term colloidal stability assessment The optimized and control formulations (TEZ-TPE-1 concentrations of 80 μmol / L and 120 μmol / L) of ANPs were stored at 4℃ and 25℃ for 30 days, respectively. Particle size and PDI were measured every 5 days. A particle size change exceeding 20% of the initial value or a PDI > 0.3 was considered unstable.
[0049] (vi) Cellular-level degradation self-indication monitoring Step 16: Cell Culture Human hepatocellular carcinoma cells (HepG2) and normal liver epithelial cells (THLE-2) were cultured in DMEM medium containing 10% fetal bovine serum and passaged in an incubator at 37°C and 5% CO2.
[0050] Step 17: Cell viability assessment The two types of cells were respectively planted at 1×10 per well. 4 Cells were seeded into 96-well plates and cultured for 24 h. Different concentrations (0, 10, 25, 50, 100 mg / L) of ANPs were added, and the cells were cultured for another 24 h. 20 μL of MTT solution (5 mg / mL) was added to each well, and after 4 h of culture, the supernatant was aspirated, and 150 μL of DMSO was added. The plates were shaken for 10 min, and the absorbance at 570 nm was measured using a microplate reader. Cell viability (%) = (experimental group absorbance / control group absorbance) × 100%.
[0051] Step 18: Uptake Kinetics HepG2 and THLE-2 cells were seeded into 6-well plates (5 × 10⁶ cells per well). 5Cells were cultured overnight. ANPs (final concentration 50 mg / L) were added, and cells were cultured for 0.5, 1, 2, 4, 8, 12, and 24 h. Cells were collected, washed three times with cold PBS, and resuspended in PBS containing 2% FBS. The mean fluorescence intensity (MFI) of each sample was detected using flow cytometry (excitation 488 nm, emission 525 ± 20 nm). Three replicates were measured at each time point, with untreated cells used as a background control. Figure 4 As shown, the uptake kinetics of ANP in normal hepatocytes (a) and hepatocellular carcinoma cells (b) are disclosed.
[0052] Step 19: Intracellular tracking Cells were seeded in confocal culture dishes (1×10⁶ cells per dish). 5 After culturing for 24 h, ANPs (50 mg / L) were added and incubated for 12 h. The culture medium was removed, and 50 nM LysoTracker Green (lysosomal marker) and 100 nM MitoTracker Red (mitochondrial marker) were added, followed by incubation for another 30 min. The cells were washed three times with cold PBS and fixed with 4% paraformaldehyde for 15 min. Images were acquired using a confocal laser scanning microscope (60× oil immersion): AIE channel (excitation 405 nm, emission 500-550 nm), LysoTracker channel (excitation 488 nm, emission 510-530 nm), and MitoTracker channel (excitation 561 nm, emission 580-620 nm). Pearson correlation coefficients were calculated using ImageJ software. Figure 5 As shown, confocal laser imaging of ANPs in normal hepatocytes (a) and hepatocellular carcinoma cells (b) is presented.
[0053] Step 20: Degradation kinetics tracking Cells were seeded in 96-well plates (1 × 10⁶ cells per well). 4 Cells were cultured for 24 h, and then ANPs (50 mg / L) were added. Cells were cultured for 0, 24, 48, 72, 96, 120, 144, and 168 h. At each time point, after washing the cells, the overall fluorescence intensity of each well was measured using a fluorescence microplate reader (excitation 365 nm, emission 501 nm). The residual fluorescence percentage at each time point was calculated with the fluorescence intensity at 0 h as 100%. To confirm protease dependence, E64d (cysteine protease inhibitor, 10 μmol / L) or Pepstatin A (aspartic protease inhibitor, 10 μmol / L) was added to some wells, and the cells were co-incubated with ANPs for 168 h before fluorescence intensity was measured. Figure 6As shown, (a) AIE fluorescence long-term tracking degradation kinetics curves of ANP in normal hepatocytes and hepatocellular carcinoma cells are disclosed, and (b) ANP degradation heatmap in hepatocellular carcinoma cells treated with lysosomal inhibitors (E64d: cysteine protease inhibitor; pepsin A: aspartic protease inhibitor) is disclosed. The key points of this invention include the following four aspects: First, a non-covalent self-assembly method without organic solvents, chemical crosslinking agents, or additives is adopted. Nanoparticle formation is driven by the hydrophobic interaction and π-π stacking between BSA and TEZ-TPE-1, completely preserving the natural tertiary conformation of albumin. Second, a multi-objective optimization framework based on GBRT and XGBoost machine learning is established. Combined with SHAP interpretability analysis, particle size, yield, fluorescence intensity, and colloidal stability are jointly predicted, and the nonlinear synergistic effect of pH and stirring rate is revealed, thereby determining the Pareto optimal formulation. Third, based on the principle of intramolecular rotational confinement of aggregation-induced emission, fluorescence intensity is linearly coupled with the integrity of the nanoparticle structure (R²=0.9828), realizing self-indicating monitoring of "strong fluorescence when intact and fluorescence quenching when degraded".
[0054] Corresponding to the key points mentioned above, the specific implementation schemes of existing technologies differ fundamentally from the technical means of this invention. Regarding the preparation method, existing chemical cross-linking schemes (such as glutaraldehyde cross-linking) involve forming a covalent Schiff base between glutaraldehyde and the ε-amino group of lysine on the surface of BSA, thereby locking the nanoparticle structure. In contrast, this invention does not use any cross-linking agent, relying solely on the non-covalent interaction between BSA and TEZ-TPE-1 for spontaneous assembly in a pure aqueous phase. Regarding parameter optimization, existing machine learning-assisted synthesis schemes typically only use random forests or linear regression to predict a single objective (such as particle size), without performing feature interaction analysis and SHAP interpretation. This invention, however, constructs a 31-dimensional input space containing second-order interaction features and ratio features, uses GBRT and XGBoost for four-objective joint optimization, and quantitatively reveals the nonlinear synergistic effect of pH and stirring rate through SHAP analysis. Regarding the molecular assembly mechanism, existing AIE-albumin composite systems (such as Nie et al.) simply mix AIE molecules with albumin, resulting in loose aggregates smaller than 10 nm, without forming stable nanoparticles. In terms of degradation monitoring methods, existing technologies rely on exogenous fluorescent labels (such as FITC), whose signals are affected by aggregation quenching. Intact nanoparticles have weak signals due to the dense accumulation of fluorophores, and the dispersion of fluorophores after degradation actually enhances the signal, resulting in false positives. This invention utilizes the rotation-restricted characteristics of AIE molecules to make the fluorescence intensity positively correlated with the integrity of nanoparticles, and the degree of degradation can be directly reported without the need for external labeling.
[0055] In summary, this invention differs clearly and fundamentally from existing technologies in three dimensions: preparation principle (non-covalent vs. covalent crosslinking), optimization method (multi-objective machine learning + SHAP vs. single-factor trial and error or single-objective prediction), and signal logic (linear coupling of intact bright / degraded dark vs. ACQ false positives). These differences are not merely functional improvements, but represent fundamental technological innovations at the levels of molecular interaction modes, data-driven strategies, and photophysical mechanisms.
[0056] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for preparing a machine learning-guided albumin nanosensor, characterized in that, Includes the following steps: (1) By designing multivariate experiments, a dataset was collected on the relationship between albumin, aggregation-induced emission (AIE) molecules, reaction condition parameters and nanosensor performance indicators; (2) Perform feature engineering on the dataset obtained in step (1), and use a machine learning regression model to predict and optimize at least one performance index of the nanosensor; (3) Use the machine learning model to determine one or more optimized combinations of synthesis parameters; (4) Based on the optimized parameters determined in step (3), AIE molecules and albumin are subjected to a non-covalent self-assembly reaction in an aqueous solution to prepare the albumin nanosensor; The specific steps for preparing the albumin nanosensor include: (i) adjusting the albumin solution to the optimal pH value and preheating it at the optimal temperature and stirring rate; (ii) slowly adding the AIE molecule solution dropwise to the albumin solution treated in step (i) and continuously stirring under optimal conditions to carry out a non-covalent self-assembly reaction; (iii) purifying the reaction solution obtained in step (ii) by ultrafiltration and centrifugation to obtain the albumin nanosensor. The sensor has an average particle size in the range of 20 nm to 200 nm and a polydispersity index (PDI) of less than 0.
2. The fluorescence intensity of the nanosensor is positively linearly correlated with its structural integrity.
2. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, The reaction conditions parameters mentioned in step (1) include at least the concentration of albumin, the concentration of AIE molecules, the pH value of the reaction system, the temperature, the stirring rate, and the reaction time.
3. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, The performance indicators mentioned in step (1) include at least particle size, yield, fluorescence intensity and colloidal stability.
4. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, The feature engineering process described in step (2) includes converting the original parameters into second-order polynomial interactive features or ratio features; the machine learning regression model is a gradient boosting regression tree (GBRT) or extreme gradient boosting (XGBoost) model.
5. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, Step (2) also includes using interpretability modeling methods to analyze the trained machine learning model and reveal the specific impact of each synthetic parameter on the performance of the nanosensor.
6. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, The interpretability model method is the SHAP method, which reveals the nonlinear synergistic effect of the interaction between the pH value and stirring rate of the reaction system on the particle size of the nanosensor through SHAP analysis.
7. The method for fabricating a machine learning-guided albumin nanosensor according to claim 1, characterized in that, Step (3) involves using the Pareto front optimization method to obtain a Pareto optimal combination of synthesis parameters by minimizing the particle size, maximizing the yield, and maximizing the fluorescence intensity of the nanosensor, combined with colloidal stability as a constraint.
8. The method for preparing a machine learning-guided albumin nanosensor according to claim 1, characterized in that, The AIE molecule is a tetraphenylethylene derivative functionalized with a tetrazolium group.
9. An application of the albumin nanosensor as described in any one of claims 1-9, characterized in that, The applications include: (i) Exposing the albumin nanosensor to living cells; (ii) Obtain the fluorescence decay kinetic curve by measuring the fluorescence signal intensity derived from the nanosensor in the living cells in real time or periodically; (iii) Evaluate the degradation rate and extent of the nanosensor or its loaded functional molecules in cells based on the fluorescence decay kinetic curves.