A method for predicting cell endocytosis based on endocytic gene expression profiling and nanoparameters

CN122575481APending Publication Date: 2026-08-14NANJING MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

该方法存在两大局限:第一,内吞量受细胞异质性(如细胞类型、分子表型)和纳米颗粒理化参数(如粒径、表面电荷和蛋白冠组成等)共同影响

Benefits of technology

[0058]将细胞内吞相关基因表达特征与纳米颗粒参数统一纳入机器学习预测框架,并基于连续内吞量数据或分类内吞量标签训练模型,使模型在训练完成后能够对新的“细胞-纳米颗粒”组合输出内吞量预测结果,从而降低重复筛选负担,并为模型选择与特征贡献度分析提供支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575481A_ABST
    Figure CN122575481A_ABST
Patent Text Reader

Abstract

This invention discloses a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters, specifically relating to the field of nanoparticle delivery and cellular endocytosis mechanism prediction. The method involves synthesizing and characterizing nanoparticles to obtain particle size and charge parameters; establishing a multi-cell model and obtaining cell type data; co-incubating nanoparticles with cells and detecting the cellular endocytosis of nanoparticles; acquiring expression level data of the set of genes related to cellular endocytosis; preprocessing and feature screening of the expression level data to remove highly collinear features and selecting a set of key gene features related to endocytosis; using nanoparticle size, charge parameters, and the set of key gene features as input features, and endocytosis data as labels, establishing a machine learning prediction model; and evaluating the performance of the machine learning prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of nanoparticle delivery and prediction of cellular endocytosis mechanisms, specifically to a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters. Background Technology

[0002] Nanomedicines (such as nanovaccines and antibody-drug conjugates) have shown great potential in personalized cancer treatment, and their efficacy depends on the effective internalization of nanoparticles by tumor cells, thereby achieving intracellular delivery. Therefore, accurately assessing and predicting the internalization amount of different nanoparticles by different cells is crucial for the rational design and personalized application of nanomedicines.

[0003] Currently, endocytosis assessment mainly relies on in vitro co-incubation and quantitative detection (such as inductively coupled plasma mass spectrometry, ICP-MS). This method has two major limitations: First, endocytosis is influenced by both cellular heterogeneity (e.g., cell type, molecular phenotype) and the physicochemical parameters of nanoparticles (e.g., particle size, surface charge, and protein crown composition). When the research system expands (e.g., by adding cell lines or nanoparticle types), numerous cell-nanoparticle combinations need to be tested individually, leading to a significant increase in cost and time. Second, existing studies are mostly limited to empirical comparisons of tested combinations, lacking a universal model framework that integrates cellular characteristics and nanoparticle parameters for quantitative prediction of unknown combinations. Therefore, it is difficult to achieve pre-experimental prediction and parameter optimization guidance.

[0004] In recent years, technologies such as transcriptomics have made it possible to analyze cellular heterogeneity at the molecular level. In particular, focusing on the expression characteristics of endocytosis-related gene sets (such as genes covering clathrin-mediated endocytosis and non-clathrin endocytosis pathways) can reduce data dimensionality while preserving key biological information. Furthermore, combining machine learning methods to screen for key genes most strongly associated with endocytosis and establishing a joint prediction model of these genes and nanoparticle parameters holds promise for pre-experimental quantitative prediction of endocytosis amounts in different tumor cell-nanoparticle combinations.

[0005] However, to date, there is no systematic study on predicting cellular endocytosis based on endocytosis-related gene expression profiles and nanoparticle parameters. Therefore, developing a tool that can integrate multi-dimensional features for pre-experimental prediction is an urgent issue to be addressed in this field. This tool can be used for early-stage cell screening, gene target discovery, and nanoparameter optimization in nanomedicine development, thereby improving development efficiency and reducing experimental costs. Summary of the Invention

[0006] Therefore, this invention provides a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters to address the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters, comprising the following steps:

[0008] Step 1: Synthesize and characterize nanoparticles to obtain particle size and charge parameters;

[0009] Step 2: Establish a multi-cell model and obtain cell type data;

[0010] Step 3: Co-incubate nanoparticles with cells and detect the amount of nanoparticles endocytosed by cells;

[0011] Step 4: Obtain the expression level data of the set of genes related to endocytosis;

[0012] Step 5: Preprocess and feature screening of the expression level data obtained in Step 4, remove highly collinear features, and screen to obtain a set of key gene features related to the endocytosis amount described in Step 3;

[0013] Step 6: Using the nanoparticle size, charge parameters, and the key gene feature set obtained in Step 5 as input features, and the endocytosis data obtained in Step 3 as labels, a machine learning prediction model is established.

[0014] Step 7: Evaluate the performance of the machine learning prediction model.

[0015] Preferably, step 1 specifically includes synthesizing and characterizing nanoparticles; the nanoparticle material includes one or more combinations of gold nanoparticles, iron oxide nanoparticles, silica nanoparticles, polymer nanoparticles, liposome nanoparticles, or nanomicelles; the hydrated particle size of the nanoparticles is 10-300 nm, and the morphology includes spherical, rod-shaped, or multi-branched shapes; and different surface charge parameters are obtained through surface modification or formulation control, wherein the surface modification includes one, two, or more combinations of polyethylene glycol modification, charged group modification, and mixed charge modification.

[0016] Preferably, step 2 specifically includes selecting and establishing multiple cell models, as follows:

[0017] The selected cell models include mature cell lines, primary cells, or tissue-derived cell populations from different sources. When using mature cell lines, multiple cell lines are selected and cultured and expanded at 37°C and 5% CO2 to obtain cell samples for co-incubation experiments.

[0018] For each cell model, cell type data is retrieved, extracted, and recorded from cell bank annotation information, supplier information, or database annotation information based on its name or number. The cell type data includes cell name or number, source tissue type, and molecular typing information to construct multiple cell models.

[0019] Preferably, in step 3, the co-incubation conditions between the nanoparticles and cells are determined and fixed through preliminary experiments, and the endocytosis of nanoparticles by cells is measured, as follows:

[0020] The co-incubation conditions included: incubation time of 0.5-24 h, nanoparticle concentration of 0.1-200 μg / mL, and cell seeding density of 1 × 10⁻⁶ cells / mL. 4 –5×10 5 The standard washing process is 2–5 times with PBS, followed by weak acid or heparin washing to reduce surface adsorption.

[0021] Preliminary experiments were used to determine the combination of incubation time and nanoparticle concentration that kept the detection signal in a measurable, unsaturated range and maintained stable comparability in repeated experiments. By standardizing co-incubation conditions, batch differences and measurement noise were reduced, and the differences in endocytosis were made to reflect cell heterogeneity and parameters such as nanoparticle size and surface charge more significantly, thereby improving the comparability of data from different cell-nanoparticle combinations and the stability of model extrapolation.

[0022] The process for measuring the amount of nanoparticles endocytosed by cells is as follows:

[0023] After co-incubation, the cells were washed to remove residual extracellular nanoparticles, and cell samples were collected. For detection using inductively coupled plasma mass spectrometry (ICP-MS), the cell samples were acid-digested to determine the target element content of the nth sample. And record the corresponding number of cells. , The continuous numerical value of the amount of endocytosis per unit cell for sample A is defined as:

[0024] ;

[0025] The internalization tag vector is as follows:

[0026] ;

[0027] When using fluorescence detection, flow cytometry, or confocal microscopy, the intracellular detection signal of the nth sample is obtained. And based on background signals and reference signal Calculate relative internalization:

[0028] ;

[0029] The relative internalization of sample A is used, and based on the distribution of detection signals in the same experimental batch, an adaptive thresholding algorithm is employed to determine the threshold. Convert to hierarchical or category labels When performing binary classification, according to the threshold The samples are divided into two classes; when performing multi-level classification, the threshold sequence is used. The sample is divided into K levels.

[0030] Preferably, step 4 specifically includes acquiring cellular gene expression data and forming a whole-genome expression matrix:

[0031] ;

[0032] Where, 𝑛 represents the number of cell-nano composite samples, and 𝑝 represents the number of genes. For the genes of sample 𝑖 The expression levels; and the expression matrix of the whole genome. Normalization or standardization was performed to mitigate the impact of sequencing depth, batch, or platform differences on expression levels. Normalization or standardization included any one or a combination of two or more of logarithmic transformation, quantile normalization, and Z-score normalization. Gene expression data sources included RNA sequencing, microarray detection, real-time quantitative PCR, or database retrieval. The data was based on a predefined set of endocytosis-related genes (F). , The number of genes, from the whole genome expression matrix Extract the expression entries corresponding to 𝐺 to obtain the endocytosis-related gene expression profile feature matrix. :

[0033] ;

[0034] in, To obtain from the whole genome expression matrix The expression profile matrix of endocytosis-related genes was obtained after extracting the expression entries corresponding to 𝐺. For sample 𝑖 The expression levels of endocytosis-related genes, where 𝑛 represents the number of cell-nano composite samples. The number of genes in the set G of genes related to endocytosis;

[0035] The pathways or functional modules covered by the endocytosis-related gene set 𝐺 include vesicle-mediated transport, lysosomal pathway, extracellular matrix binding and adhesion-related pathways, transmembrane transport-related families and processes (SLC family and ABC family), and cholesterol transport and efflux.

[0036] Preferably, step 5 specifically includes using the expression profile feature matrix of endocytosis-related genes. As input features, the internal throughput label vector As a supervisory label, the expression profile feature matrix of endocytosis-related genes was analyzed. Feature selection is performed to obtain a key gene feature set 𝐹, where |𝐹|=𝑚, and 𝑚 is the number of genes; based on the set 𝐹, the matrix is... Extract the corresponding gene features to form a dimensionality-reduced gene feature matrix. :

[0037] ;

[0038] in, This is a feature matrix of expression profiles of genes related to endocytosis. Let be the set of key gene features obtained through feature selection, and be the set of . The number of genes in For the set From the matrix The resulting dimensionality-reduced gene feature matrix is ​​formed after extracting the corresponding features. For sample 𝑖 The expression levels of endocytosis-related genes, where 𝑛 represents the number of cell-nano composite samples.

[0039] The feature selection methods include, but are not limited to, feature selection based on conventional techniques such as LASSO and / or importance selection based on tree models.

[0040] Preferably, step 6 includes using the dimensionality-reduced gene feature matrix. As a genetic characteristic, a nanoparticle feature matrix is ​​formed by feature encoding nanoparticle parameters. As a nano-parameter feature, and denoted by the endocytosis tag vector As a supervisory label;

[0041] in, For sample 𝑖 There are 1 nanoparticle characteristic value, where 𝑛 represents the number of cell-nanoparticle composite samples. The number of features obtained after encoding the nanoparticle parameters;

[0042] The dimensionality-reduced gene feature matrix With nanoparticle feature matrix The samples are concatenated according to their alignment to form a joint input matrix:

[0043] ;

[0044] Based on joint input matrix and internalization tag vector Train a machine learning prediction model and output prediction results. ;

[0045] For sample 𝑖 A joint eigenvalue, Here, is the predicted result, and is the machine learning prediction model obtained through training;

[0046] The prediction model includes, but is not limited to, linear models, support vector machines, random forests, gradient boosting trees, or extreme gradient boosting tree models.

[0047] Preferably, step (7) includes using the trained machine learning model on the validation and test sets that were not used for training. For the joint input matrix Prediction results are obtained. And based on real label vectors The performance index S of the computational model;

[0048] Among them, when When the value is continuous, the performance index S includes the coefficient of determination R. 2 One or more of the following: mean square error (MSE), mean absolute error (MAE), and / or root mean square error (RMSE); when When used for grading or classification labels, the performance metric S includes one or more of the following: accuracy (Acc), area under the receiver operating characteristic (AUC), precision (Precision), recall (Recall), and / or F1 score.

[0049] Preferably, its application stages include:

[0050] (1) Synthesis process: used for the synthesis of nanoparticles, characterization of physicochemical properties such as particle size and charge, and quality control;

[0051] (2) Modeling stage: used to construct various cell models and obtain tumor cell type data;

[0052] (3) Quantitative step: used to co-incubate cells with nanoparticles and detect the amount of cells endocytosis;

[0053] (4) Sequencing step: used to obtain gene expression profiles related to endocytosis;

[0054] (5) Feature screening step: used to preprocess gene expression profile data, remove highly collinear features and screen a set of key gene features closely related to endocytosis;

[0055] (6) Prediction step: used to integrate nanoparticle size, surface charge and key gene features and establish a machine learning prediction model to output endocytosis prediction data;

[0056] (7) Evaluation stage: used to evaluate the performance of the prediction model.

[0057] The present invention has the following advantages:

[0058] By integrating the expression characteristics of endocytosis-related genes and nanoparticle parameters into a unified machine learning prediction framework, and training the model based on continuous endocytosis data or categorical endocytosis labels, the model can output endocytosis prediction results for new "cell-nanoparticle" combinations after training, thereby reducing the burden of repeated screening and providing support for model selection and feature contribution analysis. Attached Figure Description

[0059] Figure 1 The flowchart illustrates a method for predicting cellular endocytosis based on endocytosis-related gene expression profiles and nanoparticle parameters, as provided in this embodiment of the invention.

[0060] Figure 2 This is a graph showing the endocytosis data of 22 types of tumor cells on gold nanoparticles of different sizes and surface charges, provided in embodiments of the present invention. Figure 2 (Above) shows two particle sizes. Figure 2 (Below) are 7 types of surface charges.

[0061] Figure 3 The image shows the results of screening key gene features based on LASSO regression, as provided in an embodiment of the present invention.

[0062] Figure 4 This is a comparison chart of the prediction performance of different machine learning regression models provided in the embodiments of the present invention. Detailed Implementation

[0063] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] This invention proposes a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters. The method combines the expression profile characteristics of cellular endocytosis-related genes with nanoparticle parameters (including particle size and surface charge) to predict the amount of nanoparticles endocytosed by cells, as well as a prediction tool to implement the method.

[0065] This invention proposes a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters, comprising the following steps:

[0066] Step 1. Prepare nanoparticles and obtain their particle size and surface charge data;

[0067] Step 2. Establish multiple cell models and obtain cell type data;

[0068] Step 3. Incubate the cells with nanoparticles and detect the amount of nanoparticles endocytosed by the cells;

[0069] Step 4. Obtain the expression profile data of the set of genes related to endocytosis;

[0070] Step 5. Perform preprocessing and feature screening on the gene expression profile data obtained in Step 4, remove highly collinear features, and screen to obtain a set of key gene features closely related to the endocytosis amount described in Step 3;

[0071] Step 6. Integrate the nanoparticle size, surface charge parameters, and key gene feature set obtained in Step 5 as input features, and use the endocytosis data obtained in Step 3 as labels to establish a machine learning prediction model.

[0072] Step 7. Evaluate the performance of the prediction model.

[0073] Based on the above technical solution, the following improvements can be made:

[0074] As a preferred embodiment, step 1 includes:

[0075] Step 1.1: Prepare nanoparticles and characterize and quality control their physicochemical properties. The nanoparticle materials include, but are not limited to, metal nanoparticles, silica nanoparticles, polymer nanoparticles, nanoliposomes, and nanomicelles; the characterization includes at least hydrated particle size, morphology, and surface charge parameters; the hydrated particle size of the nanoparticles is 10 to 300 nm, and the morphology of the nanoparticles includes spherical, rod-shaped, and multi-branched shapes.

[0076] Step 1.2: To obtain nanoparticles with different surface charge parameters, their surface charge is controlled through surface modification or formulation. The surface modification includes, but is not limited to, polyethylene glycol modification, charged group modification, mixed charge modification, or a combination thereof; the surface charge parameter is expressed as zeta potential.

[0077] As a preferred option, step 2 includes:

[0078] Step 2.1: Select and establish multiple cell models. The cell models include, but are not limited to, cell lines from different sources, primary cells, or tissue-derived cell populations; wherein, when using mature cell lines, the establishment includes obtaining multiple cell lines and culturing and expanding them at 37°C and 5% CO2 to obtain cell samples for co-incubation experiments; when using primary cells or tissue-derived cell populations, the establishment includes isolating cells from tissue samples and culturing them to form cell samples for co-incubation experiments.

[0079] Step 2.2: Acquire and record cell type data. The cell type data includes, but is not limited to, cell name or number, source tissue type and / or molecular typing information; the cell type data can be retrieved and extracted based on annotation information from cell banks or suppliers.

[0080] As a preferred option, step 3 includes:

[0081] Step 3.1: Determine and fix the co-incubation conditions. First, select and fix the co-incubation time *h* and nanoparticle concentration *l* through preliminary experiments to ensure the detection signal is within the measurable and unsaturated range, and maintain stability in repeated experiments. The co-incubation conditions include: incubation time 0.5–24 h, nanoparticle concentration 0.1–200 μg / mL, and cell seeding density 1 × 10⁻⁶ cells / mL. 4 –5×10 5 Each well is selected based on its specific characteristics, and a standardized washing process is employed (2-5 washes with PBS, followed by a weak acid wash or heparin wash to reduce surface adsorption).

[0082] Step 3.2: Measure endocytosis and generate endocytosis data. After co-incubation, wash cells to remove extracellular residual nanoparticles and collect cell samples. The detection methods include, but are not limited to, inductively coupled plasma mass spectrometry (ICP-MS), fluorescence detection, flow cytometry, confocal microscopy, or combinations thereof. Specifically, when using ICP-MS, the content of the target element in the cell sample is detected and combined with cell counting to calculate a continuous numerical value of endocytosis per unit cell. When using fluorescence detection, flow cytometry, or confocal microscopy, the intracellular detection signal is obtained as the relative endocytosis, and based on the detection signal distribution of the same experimental batch, the relative endocytosis is converted into a graded or classified label using an adaptive threshold determination algorithm. The adaptive threshold determination algorithm includes the Otsu threshold method or bimodal Gaussian mixture model (GMM) fitting.

[0083] Step 3.3: Represent the internalization result obtained in Step 3.2 as an internalization label vector.

[0084] ;

[0085] in, This is the label value for the endocytosis amount of sample A; when using ICP-MS detection... This represents a continuous numerical value of the amount of cells endocytosis per unit of sample A; when using fluorescence detection, flow cytometry, or confocal microscopy, the value may vary. This refers to the hierarchical or classification label corresponding to sample A. The internalization volume label vector... Used for subsequent feature selection and training and validation of machine learning prediction models.

[0086] As a preferred option, step 4 includes:

[0087] Step 4.1: Obtain whole-genome expression data of cells and construct a whole-genome expression matrix. The whole-genome expression data sources include, but are not limited to, RNA sequencing, microarray detection, real-time quantitative PCR, or database retrieval; and the whole-genome expression data will be represented as a whole-genome expression matrix.

[0088] ;

[0089] Where, 𝑛 represents the number of cell-nano composite samples, and 𝑝 represents the number of genes. For the genes of sample 𝑖 The expression level. After obtaining the whole-genome expression matrix. ,right Normalization or standardization is performed to improve the comparability of expression levels between different samples or batches, resulting in a normalized expression matrix for subsequent feature extraction.

[0090] Step 4.2: Construct the endocytosis-related gene set 𝐺. The construction of the endocytosis-related gene set includes, but is not limited to: summarizing candidate genes based on Gene Ontology (GO) entries, KEGG pathways, endocytosis-related literature, or databases, and then removing duplicates, merging synonymous genes, and unifying gene identifiers to obtain the endocytosis-related gene set 𝐺. , (The number of genes); wherein, the candidate genes preferably cover pathways or functional modules related to endocytosis and intracellular transport, including but not limited to vesicle-mediated transport, lysosomal pathway, extracellular matrix binding and adhesion-related pathways, transmembrane transport-related families and processes (SLC family-mediated transmembrane transport, ABC family-mediated transmembrane transport), and cholesterol transport and efflux, etc.

[0091] Step 4.3: Based on the aforementioned 𝐺 from An expression profile (gene feature matrix) of endocytosis-related genes was constructed. Based on the endocytosis-related gene set K, the expression level of the whole genes was analyzed from the gene expression matrix. The expression entries corresponding to 𝐺 are extracted to obtain the endocytosis-related gene expression profile (gene feature matrix).

[0092] ;

[0093] in, This is a feature matrix of expression profiles of genes related to endocytosis. For sample 𝑖 The expression levels of endocytosis-related genes, where 𝑛 represents the number of cell-nano composite samples. For the set of genes related to endocytosis The number of genes.

[0094] As a preferred embodiment, step 5 includes:

[0095] Step 5.1: Using the endocytosis-related gene expression profile feature matrix obtained in Step 4 As input features, the endocytosis label vector obtained in step 3 As a supervisory label, for the matrix Feature screening is performed to obtain a key gene feature set F, where |F| = Σ, and Σ is the number of genes in the set F.

[0096] Step 5.2: Based on the key gene feature set 𝐹, from the matrix Extract the corresponding gene features to form a dimension-reduced gene feature matrix:

[0097] ;

[0098] in, This is a set of key gene features obtained through feature screening. For the set From the matrix The dimensionality-reduced gene feature matrix extracted from it. For sample 𝑖 The expression level of each selected gene, where 𝑛 represents the number of cell-nano composite samples.

[0099] Step 5.3: The feature screening method may include LASSO-based feature screening and / or tree model-based importance screening to obtain key gene features that are highly correlated with endocytosis prediction.

[0100] As a preferred embodiment, step 6 includes:

[0101] Step 6.1: Construct the nanoparticle feature matrix. The nanoparticle size and surface charge are encoded as classification features to obtain the nanoparticle feature matrix.

[0102] ;

[0103] Among them, row 1 𝑖This represents the nanoparticle parameter encoding result corresponding to the nth "cell-nanoparticle" sample.

[0104] Step 6.2: Construct the joint input matrix. Using the dimensionality-reduced gene feature matrix obtained in Step 5.2... Compared with the nanoparticle feature matrix obtained in step 6.1 As input, samples are concatenated one-to-one to form a joint input matrix.

[0105] ;

[0106] in, For sample 𝑖 A joint eigenvalue.

[0107] Step 6.3: Using the internal throughput label vector As a supervision label, based on the joint input matrix Train a machine learning prediction model and output the prediction results:

[0108] ;

[0109] in, For the predicted results, This is the machine learning prediction model that has been trained.

[0110] The prediction models include, but are not limited to, linear models, support vector machines, random forests, gradient boosting trees, or extreme gradient boosting tree models.

[0111] As a preferred embodiment, step 7 includes:

[0112] Step 7.1: Construct evaluation data and generate prediction results. Divide the samples into a training set and a test set (and / or validation set) according to a preset ratio. The test set samples are denoted as \ ,in The joint input features of the nth sample (as described above) (assembled) Labels corresponding to the internal throughput. The trained model... Applying this to the test set input yields the prediction result:

[0113] ;

[0114] When the task is classified, It can be a category probability or a confidence score; when the task is regression. For continuous numerical prediction.

[0115] Step 7.2: Calculate performance metrics and output evaluation results. When 𝑦 is a continuous value (regression), based on and Calculate error-related and goodness-of-fit indices, for example:

[0116] ;

[0117] , ;

[0118] in The number of samples in the test set. This is a label representing the actual internal throughput. These are the model's predicted values. This represents the average of the actual labels.

[0119] When 𝑦 is a hierarchical or categorical label, the classification performance metric is calculated based on the predicted category (or probability) and the true label; for example, Mapping to predicted categories Post-calculation accuracy:

[0120] ;

[0121] When 𝑦 is a hierarchical or categorical label, the model outputs... If it is a category probability or confidence score, then a threshold needs to be applied. Convert to prediction category Then, with real labels Comparisons are made. For example, in a binary classification problem, a threshold θ is set (usually 0.5). If... θ, then 1 (High internal swallowing category), otherwise 0 (low ingestion category). Then calculate the accuracy:

[0122] ;

[0123] Where 1(⋅) is the indicator function (1 if the condition is true, 0 otherwise). Simultaneously, AUC can be calculated based on the predicted probability to evaluate the model's discriminative ability, or precision, recall, and F1 score can be calculated based on the confusion matrix. The final output is a set of performance metrics, 𝑆(𝑆={ , 2 , , (a subset of {Precision, Recall, F1}) is used to characterize the predictive reliability of the model for untested cell-nanoparticle combinations.

[0124] On the other hand, this predictive tool for the amount of nanoparticles endocytosis by cells, based on the expression profiles of endocytosis-related genes and nanoparticle parameters, includes the following technical steps: synthesis, modeling, quantification, sequencing, screening, prediction, and evaluation.

[0125] These steps are used for the following applications:

[0126] The application of the synthesis process is used to prepare or obtain nanoparticles and perform quality control.

[0127] The application modeling stage is used to construct cell models and record cell information;

[0128] The quantitative step is applied to co-incubate and to obtain endocytosis data using methods such as ICP-MS, fluorescence detection, flow cytometry, or confocal imaging.

[0129] Sequencing is used to obtain endocytosis-related gene expression data;

[0130] The application screening process is used to select key gene features for modeling.

[0131] The application prediction stage is used to build machine learning models and output internal throughput prediction results;

[0132] The application evaluation phase is used to assess predictive performance.

[0133] As used in this invention, the term "nanoparticle" refers to a particle or self-assembled nanostructure with at least one dimension at the nanoscale, typically with a characteristic size in the range of about 10–1000 nm; in a preferred embodiment of this invention, the hydrated particle size of the nanoparticle is preferably 10–300 nm. The nanoparticle may be composed of inorganic materials, organic polymers, or lipids / surfactants, and different interfacial properties (e.g., hydrated particle size and surface charge parameters) can be obtained through surface modification or formulation control to form different combinations of nanoparameters for endocytosis modeling and prediction. For ease of understanding, some types of nanoparticles covered by this invention can be described as follows (but are not limited thereto):

[0134] Nanoparticle 1: As used in this invention, the term "gold nanoparticle" refers to a nanoscale particle system formed with gold as the main component; its particle size, morphology and surface ligand / polymer coating are tunable, thereby obtaining different hydrated particle sizes and surface charge parameters.

[0135] Nanoparticle 2: As used in this invention, the term "iron oxide nanoparticles" refers to a nanoscale particle system formed mainly of iron oxide, whose surface is easily modified with charged groups or polymers to control dispersion stability and surface charge parameters.

[0136] Nanoparticle 3: As used in this invention, the term "silica nanoparticle" refers to a nanoparticle system with silica as the main component, which may be dense or porous; its surface silanol groups are easy to functionalize, thereby achieving the regulation of surface charge and surface chemical properties.

[0137] Nanoparticles 4: As used in this invention, the term "polymer nanoparticles" refers to a nanoparticle system formed from natural or synthetic polymer materials, which can be prepared by polymerization / precipitation / self-assembly, etc.; its composition and surface modification layer (e.g., PEGylation or charged groups) can be used to control hydrated particle size, surface charge and protein adsorption behavior.

[0138] Nanoparticle 5: As used in this invention, the term "liposome nanoparticle" refers to a nanoscale vesicle structure formed by lipid molecules, typically in the form of a bilayer-encapsulated cavity system; its lipid composition, cholesterol ratio and surface modification can be adjusted to obtain different size and surface charge parameters.

[0139] Nanoparticles 6: As used in this invention, the term “nanomimus” refers to nanoscale aggregates formed by the self-assembly of amphiphilic molecules in an aqueous phase, typically having a hydrophobic core and a hydrophilic shell; the composition ratio and end-group electrical properties are adjustable to form different combinations of hydrated particle size and surface charge parameters.

[0140] The term "endocytosis" as used in this invention refers to characterizing data on the intracellular accumulation level formed after nanoparticles enter cells; the endocytosis data can be a continuous numerical value or a hierarchical or classification label based on a relative quantity obtained from a detection signal and further converted. Continuous numerical values ​​can be obtained by methods such as ICP-MS; relative quantities can be obtained by methods such as fluorescence detection, flow cytometry detection, and confocal microscopy.

[0141] As used in this invention, the term "endocytosis-related gene expression profile data" refers to the expression profile information of a set of genes related to the cellular endocytosis process in cells, characterized as an expression vector or matrix of the gene set; the data sources include, but are not limited to, RNA sequencing, microarray detection, real-time quantitative PCR, or database retrieval.

[0142] As used in this invention, the term "key gene feature" refers to a set of features obtained from endocytosis-related gene expression profile data through preprocessing and feature selection, used to construct machine learning prediction models.

[0143] As used in this invention, the term "prediction" refers to the prediction of endocytosis amount for a new "cell-nanoparticle" combination based on the expression characteristics of endocytosis-related genes and nanoparticle parameters after the model training is completed; the prediction result may be a continuous prediction value or a grade / category and its probability.

[0144] As used in this invention, the term "hydrated particle size" refers to the equivalent particle size of nanoparticles in a liquid-phase dispersion system due to the presence of a solvation layer and an adsorption layer (e.g., a protein / polymer layer). Hydrated particle size is typically expressed as the average particle size and distribution parameters measured by dynamic light scattering (DLS), including but not limited to the polydispersity index (PDI).

[0145] As used in this invention, the term "surface charge" refers to the interfacial electrical characteristics of nanoparticles in a liquid phase, typically characterized by zeta potential. Zeta potential can be measured in a specific dispersion medium and can vary with pH, ​​ionic strength, and surface modification; surface charge can also be expressed as positive, negative, or near-neutral, etc., according to its electrical class. The surface charge parameter described in this invention can be a numerical value of zeta potential or an equivalent charge class parameter.

[0146] As used in this invention, the term "polyethylene glycol (PEG)" refers to polymers and their derivatives consisting of repeating ethoxy units; PEG molecules may have different molecular weights and end group types.

[0147] As used in this invention, the term "surface modification" refers to the process of modifying the surface of nanoparticles through chemical or physical means to alter their surface composition and interfacial properties. Surface modification methods include, but are not limited to, PEGylation, charged group modification, mixed charge modification, ligand modification, or combinations thereof; the results of surface modification may manifest as changes in hydrated particle size, zeta potential, stability, or biological interfacial interactions.

[0148] Example:

[0149] To achieve the objectives of this invention, this embodiment provides a method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters, such as... Figure 1 As shown, the implementation process includes the following seven steps: synthesis, modeling, quantification, sequencing, screening, prediction, and evaluation.

[0150] The synthesis step is used to obtain nanoparticles with different particle sizes and different surface charges;

[0151] The modeling step is used to construct various tumor cell models;

[0152] The quantitative step is used to obtain the amount of nanoparticles endocytosed by cells.

[0153] The sequencing step is used to obtain endocytosis-related gene expression profile data;

[0154] The screening step is used to screen for key gene features;

[0155] The prediction step is used to construct various prediction models;

[0156] The evaluation step is used to evaluate the prediction model.

[0157] The synthesis process, taking the preparation of gold nanoparticles (GNPs) as an example, constructs a set of nanoparticle parameters comprising two particle sizes and seven surface charges, totaling 14 GNPs. First, approximately 15 nm GNPs are prepared using the sodium citrate reduction method, and these are used as seeds to obtain approximately 70 nm GNPs via seed-mediated growth. Subsequently, to obtain seven types of surface modification covering positive, neutral, and negative charges, a mixed-charge polyethylene glycol (PEG) modification strategy is applied. This involves co-modification with amino-PEG-thiols (NH2-PEG-SH) and carboxyl-PEG-thiols (COOH-PEG-SH) in different proportions, or the use of single ligands (NH2-PEG-SH, COOH-PEG-SH, mPEG-SH) to achieve systematic adjustment of surface charges. The modification system uses an equal surface area input (the surface area of ​​both 15 nm and 70 nm GNPs is 40 cm²). 2 The process was carried out to ensure consistent ligand coverage across different particle sizes. The resulting 14 GNPs were characterized by their core size, hydrated size, polydispersity, and surface charge using transmission electron microscopy, dynamic light scattering, and zeta potential analysis. These characteristics were then used as nanoparticle parameters for subsequent model inputs.

[0158] The modeling process employed mature cell lines of known origin, stability, and reproducibility to construct multi-cell models. Twenty-two human solid tumor cell lines from the NCI-60 panel were selected, including ACHN, MCF7, A549, HS578T, A498, NCI-H460, MDA-MB-231, M14, HCT-15, SK-MEL-5, HCT-116, T47D, DU-145, SN12C, SK-MEL-2, SF-295, 786-0, NCI-H226, SK-MEL-28, BT-549, PC-3, and NCI-H522. These cell lines were expanded under standard culture conditions: RPMI-1640 containing 10% fetal bovine serum, cultured at 37°C in a 5% CO2 incubator, and used in subsequent experiments when the cells were in the logarithmic growth phase. Cell names, tissue origins, and typing information are recorded based on cell bank annotation information to form cell type data, which is then used as part of the input features for subsequent models to construct a cross-cell type nanoparticle endocytosis response database.

[0159] The quantitative step is used to obtain the amount of nanoparticles endocytosed by cells. First, preliminary experiments were conducted to evaluate the linearity of the detection signal, batch consistency, and cell viability under different combinations of incubation times (2, 4, 8, 12, 24 h) and concentrations (0.01–0.20 mg / mL). Finally, the unified co-incubation conditions were determined to be: 4 h incubation, GNP concentration of 0.10 mg / mL, and a cell seeding density of approximately 1 × 10⁶ cells per well. 5 Cells were collected. During standard co-incubation, GNP was added to 24-well cell plates to the final concentration mentioned above, and incubated at 37°C and 5% CO2 for 4 h. After incubation, the cells were washed four times with PBS, digested, and collected for ICP-MS analysis. The cell pellet was digested with aqua regia, and the target element content of the nth sample was determined. and combined with cell number Calculate the amount of endocytosis per unit cell: The result Constituting a continuous internalization volume label A total of 308 endocytosis data points were obtained based on 14 GNPs and 22 cell model combinations, which were used for training and testing of subsequent prediction models. This dataset systematically presents the differences in endocytosis of nanoparameters among different cells, such as... Figure 2 As shown: Figure 2 (Above) Shows the endocytosis distribution of 15 nm and 70 nm GNPs in 22 cell types, demonstrating the particle size effect; Figure 2 (Below) This section shows the differences in endocytosis of seven different surface charge GNPs by 22 cell types, reflecting the influence of surface charge regulation on cellular uptake.

[0160] The sequencing step used the CellMiner public database to obtain the full genome expression profile of the NCI-60 cell line, which contained 23,808 genes, denoted as:

[0161] ,

[0162] The obtained raw expression data were standardized by log2(FPKM+1) transformation. To obtain gene characteristics related to nanoparticle endocytosis, this embodiment screened 15 endocytosis-related functional pathways (including clathrin-mediated endocytosis, non-clathrin-dependent endocytosis, macropinocytosis, vesicle transport, SLC / ABC transport, etc.) based on literature and database data. The pathway genes were summarized and deduplicated to construct an endocytosis-related gene set F, totaling 1405 genes. Subsequently, the standardized whole-genome expression matrix was analyzed. The gene expression profiles of endocytosis-related genes were extracted according to the column index corresponding to gene set F, forming a 22×1,405-gene expression profile.

[0163] .

[0164] The screening process uses endocytosis-related gene expression profiles. As input, and with the internalization label vector As labels, they are used for key gene feature selection. First, the entire dataset is randomly divided into a training set (n = 216) and a test set (n = 96). Then, feature selection is performed only based on the training set data to avoid information leakage. On the training set, using 1405 gene expression features as independent variables and their corresponding endocytosis as dependent variables, the LASSO method with an K1 regularization term is used, combined with ten-fold cross-validation to determine the optimal regularization parameter, shrinking the coefficients of unimportant features to zero. Finally, 21 genes with non-zero coefficients are selected from the 1405 candidate genes, constituting the key feature set K (e.g., ...). Figure 3 Based on this, a dimensionality-reduced gene feature matrix is ​​constructed for subsequent prediction.

[0165] The prediction step uses a gene feature matrix composed of the expression characteristics of 21 screened genes. As part of the input, the particle size and surface charge of the nanoparticles are encoded using a classification feature method to form a nanoparticle feature matrix. The two are concatenated according to the one-to-one correspondence of samples to obtain the joint input matrix: Each row corresponds to all input features of a "cell-nanoparticle" combination sample. The endocytosis label vector... As a supervisory signal, the machine learning model is trained using 𝑍 as input. and output the prediction results. This embodiment trains and compares linear models, support vector machines, random forests, gradient boosting trees, and extreme gradient boosting trees (XGBoost). During training, grid search combined with 10-fold cross-validation is used to optimize the hyperparameters of each model to obtain the optimal configuration for each model.

[0166] The evaluation phase uses independent test sets to assess the performance of each trained optimized model. By calculating metrics such as correlation coefficient, coefficient of determination, and mean squared error, the predictive performance of each model is compared. The performance results of each model are shown below. Figure 4 Among them, the XGBoost model performed best on the test set, with the highest correlation coefficient and determination coefficient between the predicted and measured values ​​(0.83 and 0.70, respectively) and the lowest mean squared error (0.11). Therefore, it was determined to be the best performing prediction model.

[0167] In summary, the method and tool for predicting cellular endocytosis based on endocytosis-related gene expression profiles and nanoparticle parameters provided in this embodiment have the following technical advantages:

[0168] This embodiment provides a novel approach for predicting the endocytosis of gold nanoparticles in tumor cells based on transcriptomics technology, including synthesis, modeling, quantification, sequencing, screening, prediction, and evaluation.

[0169] Through these steps, a quantitative dataset of endocytosis can be constructed covering multi-cell models and multiple nano-parameter combinations. Expression characteristics of endocytosis-related genes can be extracted and regression models trained, enabling the prediction of endocytosis for untested "cell-parameter combinations" before experiments. This provides a basis for optimizing parameters such as nanoparticle size, surface charge, and surface modification, reduces repetitive in vitro screening during cell model expansion or parameter adjustment, accelerates nanomedicine development, and saves research and development costs.

[0170] It should be understood that the aforementioned technical means can be implemented individually in hardware or software form, or through an integration of both. Therefore, the methods and apparatus involved in this invention, as well as some of their constituent elements or aspects, can be embedded in a physical medium, such as a floppy disk, CD-ROM, hard disk, or any other machine-readable storage medium carrying program code (i.e., a series of instructions). When this program is loaded onto a device such as a computer and executed, the device becomes a tool for implementing this invention.

[0171] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters, characterized in that: Includes the following steps: Step 1: Synthesize and characterize nanoparticles to obtain particle size and charge parameters; Step 2: Establish a multi-cell model and obtain cell type data; Step 3: Co-incubate nanoparticles with cells and detect the amount of nanoparticles endocytosed by cells; Step 4: Obtain the expression level data of the set of genes related to endocytosis; Step 5: Preprocess and feature screening of the expression level data obtained in Step 4, remove highly collinear features, and screen to obtain a set of key gene features related to the endocytosis amount described in Step 3; Step 6: Using the nanoparticle size, charge parameters, and the key gene feature set obtained in Step 5 as input features, and the endocytosis data obtained in Step 3 as labels, a machine learning prediction model is established. Step 7: Evaluate the performance of the machine learning prediction model.

2. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Step 1 specifically includes synthesizing and characterizing nanoparticles; the nanoparticle material includes one or more combinations of gold nanoparticles, iron oxide nanoparticles, silica nanoparticles, polymer nanoparticles, liposome nanoparticles, or nanomicelles; the hydrated particle size of the nanoparticles is 10-300 nm, and the morphology includes spherical, rod-shaped, or multi-branched shapes; and different surface charge parameters are obtained through surface modification or formulation control, wherein the surface modification includes one, two, or more combinations of polyethylene glycol modification, charged group modification, and mixed charge modification.

3. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Step 2 specifically includes selecting and establishing multiple cell models, as follows: The selected cell models include mature cell lines, primary cells, or tissue-derived cell populations from different sources. When using mature cell lines, multiple cell lines are selected and cultured and expanded at 37°C and 5% CO2 to obtain cell samples for co-incubation experiments. For each cell model, cell type data is retrieved, extracted, and recorded from cell bank annotation information, supplier information, or database annotation information based on its name or number. The cell type data includes cell name or number, source tissue type, and molecular typing information to construct multiple cell models.

4. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: In step 3, the co-incubation conditions between nanoparticles and cells were determined and fixed through preliminary experiments, and the amount of nanoparticles endocytosed by cells was measured, as follows: The co-incubation conditions included: incubation time of 0.5-24 h, nanoparticle concentration of 0.1-200 μg / mL, and cell seeding density of 1 × 10⁻⁶ cells / mL. 4 –5×10 5 The standard washing procedure is 2–5 washes with PBS per well; The process for measuring the amount of nanoparticles endocytosed by cells is as follows: After co-incubation, the cells were washed to remove residual extracellular nanoparticles, and cell samples were collected. For detection using inductively coupled plasma mass spectrometry (ICP-MS), the cell samples were acid-digested to determine the target element content of the nth sample. And record the corresponding number of cells. , The continuous numerical value of the amount of endocytosis per unit cell for sample A is defined as: ; The internalization tag vector is as follows: ; When using fluorescence detection, flow cytometry, or confocal microscopy, the intracellular detection signal of the nth sample is obtained. And based on background signals and reference signal Calculate relative internalization: ; in, The relative internalization of sample A is used, and based on the distribution of detection signals in the same experimental batch, an adaptive thresholding algorithm is employed to determine the threshold. Convert to hierarchical or category labels When performing binary classification, according to the threshold The samples are divided into two classes; when performing multi-level classification, the threshold sequence is used. The sample is divided into K levels.

5. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Step 4 specifically includes acquiring cellular gene expression data and forming a whole-genome expression matrix: ; Where, 𝑛 represents the number of cell-nano composite samples, and 𝑝 represents the number of genes. For the genes of sample 𝑖 The expression levels; and the expression matrix of the whole genome. Normalization or standardization was performed to mitigate the impact of sequencing depth, batch, or platform differences on expression levels. Normalization or standardization included any one or a combination of two or more of logarithmic transformation, quantile normalization, and Z-score normalization. Gene expression data sources included RNA sequencing, microarray detection, real-time quantitative PCR, or database retrieval. The data was based on a predefined set of endocytosis-related genes (F). , The number of genes, from the whole genome expression matrix Extract the expression entries corresponding to 𝐺 to obtain the endocytosis-related gene expression profile feature matrix. : ; in, To obtain from the whole genome expression matrix The expression profile matrix of endocytosis-related genes was obtained after extracting the expression entries corresponding to 𝐺. For sample 𝑖 The expression levels of endocytosis-related genes, where 𝑛 represents the number of cell-nano composite samples. denoted as the number of genes in the endocytosis-related gene set G.

6. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 5, characterized in that: Step 5 specifically includes using the expression profile feature matrix of endocytosis-related genes. As input features, the internal throughput label vector As a supervisory label, the expression profile feature matrix of endocytosis-related genes was analyzed. Feature selection is performed to obtain a key gene feature set 𝐹, where |𝐹|=𝑚, and 𝑚 is the number of genes; based on the set 𝐹, the matrix is... Extract the corresponding gene features to form a dimensionality-reduced gene feature matrix. : ; in, This is a feature matrix of expression profiles of genes related to endocytosis. Let be the set of key gene features obtained through feature selection, and be the set of . The number of genes in For the set From the matrix The resulting dimensionality-reduced gene feature matrix is ​​formed after extracting the corresponding features. For sample 𝑖 The expression levels of endocytosis-related genes, where 𝑛 represents the number of cell-nano composite samples. The feature selection methods include feature screening based on the conventional technique LASSO and / or importance screening based on tree models.

7. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Step 6 includes using the dimensionality-reduced gene feature matrix. As a genetic characteristic, a nanoparticle feature matrix is ​​formed by feature encoding nanoparticle parameters. As a nano-parameter feature, and denoted by the endocytosis tag vector As a supervisory label; in, For sample 𝑖 There are 1 nanoparticle characteristic value, where 𝑛 represents the number of cell-nanoparticle composite samples. The number of features obtained after encoding the nanoparticle parameters; The dimensionality-reduced gene feature matrix With nanoparticle feature matrix The samples are concatenated according to their alignment to form a joint input matrix: , Based on joint input matrix and internalization tag vector Train a machine learning prediction model and output prediction results. ; in, For sample 𝑖 A joint eigenvalue, Here, is the predicted result, and is the machine learning prediction model obtained through training; The prediction model includes a conventional linear model, a support vector machine, a random forest, a gradient boosting tree, or an extreme gradient boosting tree model.

8. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Step (7) includes using the trained machine learning model on the validation and test sets that were not used for training. For the joint input matrix Prediction results are obtained. And based on real label vectors With the prediction results The performance index S of the computational model; Among them, when When the value is continuous, the performance index S includes the coefficient of determination R. 2 One or more of the following: mean square error (MSE), mean absolute error (MAE), and / or root mean square error (RMSE); when When used for grading or classification labels, the performance metric S includes one or more of the following: accuracy (Acc), area under the receiver operating characteristic (AUC), precision (Precision), recall (Recall), and / or F1 score.

9. The method for predicting cellular endocytosis based on endocytic gene expression profiles and nanoparameters according to claim 1, characterized in that: Its application stages include: (1) Synthesis process: used for the synthesis of nanoparticles, characterization of physicochemical properties such as particle size and charge, and quality control; (2) Modeling stage: used to construct various cell models and obtain tumor cell type data; (3) Quantitative step: used to co-incubate cells with nanoparticles and detect the amount of cells endocytosis; (4) Sequencing step: used to obtain gene expression profiles related to endocytosis; (5) Feature screening step: used to preprocess gene expression profile data, remove highly collinear features and screen a set of key gene features closely related to endocytosis; (6) Prediction step: used to integrate nanoparticle size, surface charge and key gene features and establish a machine learning prediction model to output endocytosis prediction data; (7) Evaluation stage: used to evaluate the performance of the prediction model.