A method and system for identifying and classifying single-vesicle particle electrochemical signals

By fabricating ultra-carbon fiber microelectrodes and patch-clamp systems, and combining signal processing and various machine learning models, rapid and intelligent identification and detection of vesicle particle size can be achieved. This solves the problem of signal classification and analysis difficulties in traditional methods, and improves the identification accuracy of electrochemical signals and the intelligence level of the system.

CN120992425BActive Publication Date: 2026-01-27NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511511521.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-27
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Traditional methods struggle to achieve rapid and intelligent identification and detection of vesicle size, especially in electrochemical signals where noise and unstructured issues make signal classification and analysis difficult.

Method used

By fabricating ultra-carbon fiber microelectrodes and combining them with a patch-clamp system to collect vesicle current signals, signal processing and feature extraction are performed. Principal component analysis, oversampling algorithms and various machine learning models are then used for automatic classification and identification.

Benefits of technology

It significantly reduces noise interference in current signals, improves the accuracy of signal peak identification, enhances the interpretability and accuracy of classification models, and improves the accuracy of vesicle size classification and the intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120992425B_ABST
    Figure CN120992425B_ABST
Patent Text Reader

Abstract

The application discloses a single-vacuole particle electrochemical signal recognition classification method and a recognition system, and the method comprises the following steps: S1, a carbon fiber ultramicroelectrode is prepared, and a voltage is applied in a membrane patch clamp system constant potential mode, and a picoampere-level current response generated by oxidation of single-vacuole content on the surface of the electrode is collected; S2, the collected original current signal is subjected to smoothing filtering, and a multi-parameter peak recognition method based on a threshold value is used to realize feature peak recognition and multi-dimensional physical / statistical feature extraction; S3, the extracted peak feature data is subjected to a preprocessing operation including missing value processing, standardization, principal component analysis dimension reduction and class balancing; and S4, at least one machine learning model is constructed and trained based on the preprocessed feature data, and is used for automatically recognizing the particle size class of the vacuole particle. The application combines the advantages of single-vacuole electrochemical high-sensitivity measurement and intelligent algorithm classification, and has the advantages of high recognition efficiency, good repeatability and wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electrochemical detection technology, specifically to a method and identification system for automatically identifying and quantitatively analyzing vesicle particle size distribution based on patch-clamp technology to acquire current signals and combined with machine learning algorithms. Background Technology

[0002] Vesicles are widely used in biomarker detection, disease diagnosis, and drug delivery, and their particle size distribution directly affects their biological functions and effects. Traditional particle size analysis methods, such as dynamic light scattering (DLS) and transmission electron microscopy, can effectively characterize particle size, but these methods are usually complex to operate, costly, and lack real-time performance and intelligent identification capabilities. On the other hand, electrochemical techniques such as patch-clamp have millisecond-level time resolution and can detect redox reactions.

[0003] However, electrochemical signals often suffer from strong noise, unstructured nature, and large event fluctuations, making it difficult to classify and analyze them using traditional methods. Therefore, combining signal processing, statistical feature extraction, and machine learning to achieve rapid and intelligent identification and detection of vesicle size has become an urgent technical challenge.

[0004] To address this, this application proposes a method and system for identifying and classifying electrochemical signals from single vesicle particles. By combining membrane particle size control, electrode microstructure design, current characteristic modeling, and machine learning training processes, an efficient and reusable identification system is constructed, which can improve the measurement and identification level of vesicles and other nanostructures, thereby solving the current technical problems. Summary of the Invention

[0005] The main objective of this invention is to provide a method and system for identifying and classifying electrochemical signals of single vesicle particles. By preparing ultra-carbon fiber microelectrodes and combining them with a patch-clamp system, current signals are collected from vesicle samples with different particle size distributions obtained by extrusion and centrifugation purification. Subsequently, multi-dimensional feature parameters are obtained using signal processing and feature extraction algorithms. Furthermore, principal component analysis (PCA), oversampling algorithm (SMOTE), and various machine learning models are used to achieve automatic classification and identification of vesicle particle size and quantitative calculation of the number of molecules contained in the vesicles, thereby solving the technical problems mentioned in the background art.

[0006] The present invention solves the above-mentioned technical problems by adopting the following technical solutions:

[0007] A method for identifying and classifying electrochemical signals from vesicle particles, specifically including:

[0008] S1. Fabricate carbon fiber microelectrodes and connect them to a patch clamp system, and apply a set voltage to collect the oxidation current signal caused by vesicles on the electrode surface;

[0009] S2. The raw data of the acquired current signal is smoothed and filtered, and a threshold-based multi-parameter peak identification method is used to realize feature peak identification and multi-dimensional physical / statistical feature extraction;

[0010] S3. Perform preprocessing operations on the peak feature data in the extracted feature data, including missing value handling, standardization, principal component analysis (PCA) dimensionality reduction (preserving 95% of the variance contribution rate), and class balance;

[0011] S4. Based on the preprocessed feature data, construct and train at least one machine learning model for automatic identification of vesicle particle size categories.

[0012] Preferably, the preparation process of the carbon fiber microelectrode in step S1 includes: connecting a specified micron-sized carbon fiber filament with a copper wire using conductive silver paste and embedding it into a glass capillary; filling it with epoxy resin and ensuring that the carbon fiber filament is exposed; and shearing and polishing the front end of the carbon fiber microelectrode to form a smooth microporous structure for use as a working electrode of the patch clamp system.

[0013] Preferably, in step S1, the extracted vesicles are continuously extruded through an extruder and purified by centrifugation. During the extrusion process, a specified nm series filter membrane is used for multiple extrusions, with the pore size gradually decreasing: 800nm, 400nm, 200nm, and 100nm. Each sample undergoes at least 10 rounds of extrusion through the filter to obtain a sample population with significant differences in particle size distribution, and corresponding labels are prepared.

[0014] Preferably, the current signal is acquired using a carbon fiber microelectrode as the working electrode, an Axon MultiClamp 700B patch clamp system, under a +600mV bias voltage, at a sampling frequency of 100kHz, and the data is exported in CSV format.

[0015] Preferably, the specific operation process of step S2 includes:

[0016] S21. The original current signal is smoothed by Gaussian filtering. Based on multiple parameters including peak height (>mean + 20pA), significance (≥30pA), width (≥3 time points), and spacing (≥10 time points), the characteristic peaks are jointly identified and the top N peaks are selected according to significance.

[0017] That is, there are conditions and restrictions: the peak height exceeds the baseline mean plus a specified threshold, the significance is greater than 30 pA, the peak width is not less than 3 time points, and the peak-to-peak distance is not less than 10 time points;

[0018] S22. Perform physical and statistical characteristic calculations, including charge integration (unit conversion to nC / C) and molecule number conversion, calculate peak height, peak width, and current statistics (mean, standard deviation, skewness, kurtosis), and then calculate the ratio of peak area time normalized value to baseline current value.

[0019] That is, feature extraction includes at least one of the following: charge integral, peak height, peak width, average current, standard deviation, skewness, kurtosis, number of converted molecules and normalized area;

[0020] S23. Generate a signal overview diagram and a single-peak analysis diagram with peak position markers.

[0021] Preferably, in the S21 characteristic peak identification process, a fixed width signal interval is truncated to the left and right of each group of characteristic peaks, and the baseline average current is estimated by extracting a short interval outside the segment.

[0022] If there is no data to the left or right of the characteristic peak, then the relative side estimate is used.

[0023] Preferably, in the physical and statistical characteristic calculation process of step S22:

[0024] The formula for the integral of charge is:

[0025]

[0026] in, It is the total charge. For time Changing current, and Each peak segment is assigned a start and end time.

[0027] The number of molecules to be converted based on Faraday's constant and Avogadro's constant is calculated using the following formula:

[0028]

[0029] Where N is the number of molecules encapsulated in a single current peak, and Q is the charge. is Avogadro's constant, F is Faraday's constant, and n is the electron transfer number;

[0030] Calculate the time-normalized value of peak area based on residence time;

[0031] The baseline current ratio is the ratio of the peak-to-peak current difference ΔI to the baseline.

[0032] Preferably, in step S3, at least 95% of the cumulative variance contribution rate is retained during the principal component analysis dimensionality reduction process for feature dimension compression and noise removal, thereby further improving data quality.

[0033] Preferably, the machine learning model in step S4 includes any one or more of the following five types of models: decision tree, random forest, support vector machine, K-nearest neighbors or XGBoost, and is validated using a multi-model ensemble classification scheme.

[0034] Preferably, the following optimization strategies are adopted in the construction and training of the machine learning model: stratified sampling to divide the training set / test set (80%:20%), grid search to fine-tune key hyperparameters, and 5-fold cross-validation to ensure generalization ability.

[0035] Preferably, during the model training process in step S4, the SMOTE algorithm is used to perform minority class oversampling on the training data during the cross-validation process to ensure the balance of the training data.

[0036] Preferably, the model evaluation in step S4 includes constructing a confusion matrix, a classification report, multi-class ROC curves, and calculating and visualizing AUC values ​​to evaluate the recognition and classification performance of the machine learning model.

[0037] Preferably, during the execution of step S4, a set of thresholded correlation formulas are constructed, which include a ratio parameter. for:

[0038] When the ratio parameter is in different threshold ranges, different machine learning models are invoked to classify the specified sample.

[0039] Extract physical and statistical features, including peak height and peak width, from a portion of the training set;

[0040] The ratio parameter k of the peak height to the peak width is calculated to characterize the peak shape features of the detected signal;

[0041] Based on the thresholding correlation formula, the classification results of each machine learning model are compared under different threshold conditions, and the optimal model for the ratio range is determined.

[0042] Output the prediction results of the optimal model to achieve optimal discrimination of vesicle particle size category.

[0043] A system for identifying single-vesicle electrochemical signals, used to perform any of the above-described methods for identifying and classifying vesicle electrochemical signals, specifically including:

[0044] The current acquisition module, including carbon fiber electrodes, patch-clamp amplifiers, and a data acquisition card, is used to record the electrical signals caused by vesicle oxidation reactions.

[0045] The signal processing and feature extraction module is used to perform signal smoothing, peak detection and multi-dimensional feature construction operations based on the electrical signal data from the current acquisition module in order to obtain the processed feature extraction data.

[0046] The data preprocessing module is used to perform missing value processing, feature standardization, PCA dimensionality reduction and class balance processing operations on the feature extraction data based on the signal processing and feature extraction module.

[0047] The classification and prediction module loads the training model and is used to perform particle size prediction and classification on the data processed by the data preprocessing module using the loaded training model.

[0048] Preferably, the classification and prediction module supports a graphical interface for displaying ROC analysis charts and uses Python to automatically adjust and optimize parameters through multi-model grid search of the data to improve classification performance.

[0049] As can be seen from the above technical solution, the present invention provides a method and system for identifying and classifying vesicle electrochemical signals. Compared with the prior art, the present invention has the following advantages:

[0050] 1. By combining multi-parameter dynamic thresholding and Gaussian filtering in the data feature extraction process, this invention can significantly reduce noise interference in current signals, improve the accuracy of signal peak identification, and ensure that the features extracted from electrochemical signals are more accurate and reliable.

[0051] 2. This invention enhances the interpretability and accuracy of the classification model by integrating physical features such as charge integral and molecular number conversion, and combining them with machine learning models for deep learning and classification. In particular, in complex signal analysis tasks, it can effectively improve the classification accuracy of vesicle size by combining machine learning technology.

[0052] 3. By combining PCA dimensionality reduction and SMOTE data augmentation strategies, this invention can significantly improve the robustness of small sample class identification, especially when facing the problem of imbalanced data. Furthermore, combining machine learning methods can effectively improve the adaptability and generalization ability of the model, thereby enabling it to adapt to the diversity of data in different experimental environments.

[0053] 4. This invention integrates multiple machine learning models for classification tasks and significantly improves the classification accuracy and robustness of the models through optimization methods such as cross-validation and hyperparameter tuning. It not only enhances the accuracy of signal classification but also enables the processing of complex electrochemical signal data, thereby improving the intelligence and automation level of the system. It is suitable for large-scale vesicle size classification tasks.

[0054] 5. This invention combines the advantages of high-sensitivity electrochemical measurement and intelligent algorithm classification, and has the advantages of high identification efficiency, good repeatability and strong applicability. It can be applied to fields such as nanocarriers, membrane channel mechanisms and biomedical material analysis.

[0055] It should be understood that the descriptions in this section are not intended to identify key or essential features of embodiments of the invention, nor are they intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Of course, implementing any product of the invention does not necessarily require achieving all of the advantages described above simultaneously. Attached Figure Description

[0056] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0057] Figure 1 This is a schematic diagram of the overall operation process of the present invention;

[0058] Figure 2 This is an overview of the peak distribution in the feature extraction step of the present invention;

[0059] Figure 3 This is a visualization diagram of the single-peak analysis of the current peak extracted in the feature extraction step of the present invention;

[0060] Figure 4 This is a schematic diagram of the feature correlation thermodynamic analysis in the data preprocessing step of the present invention;

[0061] Figure 5 This is a diagram illustrating the cumulative variance of PCA in the data preprocessing step of this invention.

[0062] Figure 6 This is a schematic diagram comparing the accuracy of different models in the model training and prediction steps of this invention.

[0063] Figure 7 This is the confusion probability matrix of the decision tree model in the model training and prediction steps of this invention;

[0064] Figure 8 This is the confusion probability matrix of the random forest model in the model training and prediction steps of this invention;

[0065] Figure 9 This is the confusion probability matrix of the SVM model in the model training and prediction steps of this invention;

[0066] Figure 10 This is the confusion probability matrix of the XGboost model in the model training and prediction steps of this invention;

[0067] Figure 11 This is the confusion probability matrix of the KNN model in the model training and prediction steps of this invention;

[0068] Figure 12 This is the ROC curve of the decision tree model in the model training and prediction steps of this invention;

[0069] Figure 13 This is the ROC curve of the random forest model in the model training and prediction steps of this invention;

[0070] Figure 14 This is the ROC curve of the SVM model in the model training and prediction steps of this invention;

[0071] Figure 15 This is the ROC curve of the XGboost model in the model training and prediction steps of this invention;

[0072] Figure 16 This is the ROC curve of the KNN model in the model training and prediction steps of this invention. Detailed Implementation

[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] For details in the embodiments, please refer to Figures 1 to 16 .

[0075] like Figure 1 As shown, the method for identifying and classifying the electrochemical signals of single vesicle particles proposed in this embodiment of the invention includes the following operational steps in its specific implementation:

[0076] S1. Fabricate carbon fiber microelectrodes and connect them to a patch clamp system. Apply a set voltage to collect the oxidation current signal caused by vesicles on the electrode surface.

[0077] In this specific implementation, a carbon fiber filament with a diameter of approximately 4 μm is used to fabricate a microelectrode via glass capillary encapsulation. The specific steps include:

[0078] (a1) Select commercially available carbon fiber filaments and cut approximately 1 cm;

[0079] (a2) One end of it is connected to a copper wire for conduction with conductive silver paste, and then dried in a high-temperature drying oven at 80°C for 20 minutes;

[0080] (a3) Insert the copper wire with carbon fiber filaments into the capillary glass tube, fill the glass tube with epoxy resin until it completely covers the copper wire part, and ensure that part of the carbon fiber filaments protrudes from the glass tube, and then dry it in a high-temperature drying oven at 80°C for 40 minutes.

[0081] (a4) The front carbon fiber is sanded flat with sandpaper to form a smooth electrode surface structure;

[0082] (a5) The electrode may be electrochemically activated (acid or alkali) to improve sensitivity;

[0083] (a6) The electrode is connected to a patch clamp system and used as a working electrode.

[0084] Furthermore, this application uses vesicles as the identified object. Here, the vesicles are separated by extrusion using a generator. The specific sample acquisition process includes:

[0085] (b1) The vesicle samples were continuously extruded by an extruder (Genizer) using a series of filter membranes of different specifications with pore sizes of 100nm, 200nm, 400nm and 800nm. Each sample was extruded through the filter for at least 10 rounds and then purified by gradient centrifugation to obtain four samples with significant differences in the main peak particle size, which served as the classification basis for subsequent machine learning model training.

[0086] In addition, it should be noted that the extruded components of each stage of the vesicle sample are preserved in independent samples. The particle size distribution is mainly concentrated in small, small-medium, large-medium, and large particles, respectively. Labels are also made as target variable columns, such as VC_50, VC_100, VC_150, and VC_200, so that they correspond to vesicle sample groups of 100nm, 200nm, 400nm, and 800nm, respectively.

[0087] In addition, during signal acquisition, the corresponding electrical signals were acquired by the Axon MultiClamp 700B patch-clamp system, and the main settings parameters are as follows:

[0088] (c1) Use 4μm± carbon fiber microelectrodes;

[0089] (c2) Connects to the amplifier system and is equipped with a data acquisition card;

[0090] (c3) Add the vesicle sample to a centrifuge tube at room temperature;

[0091] (c4) The system applies a bias voltage of +600mV and records the current signal generated by the oxidation reaction of vesicles on the surface of the microelectrode.

[0092] (c5) Set the sampling frequency to 100kHz;

[0093] (c6) The collected signals are exported as text format (.txt / .csv) for subsequent data analysis.

[0094] S2. The raw data of the acquired current signal is smoothed and filtered, and a threshold-based multi-parameter peak identification method is used to achieve characteristic peak identification and multi-dimensional physical / statistical feature extraction. The single-peak analysis of the characteristic peak identification (extracted current peaks) is referenced. Figure 2 .

[0095] In a preferred embodiment of this application, the raw current signal measured by patch clamp or grinding line is processed by the following steps for peak identification and feature extraction:

[0096] (S21) Import the raw current-time series data from the electrochemical signal acquisition device. The data format is dual-channel text (time / ms, current / pA). To reduce noise interference, the raw signal is smoothed using a Gaussian filter (σ=3).

[0097] (S22) Peak identification is performed using the dynamic threshold method, where the following conditions are used for multi-peak identification: (1) Peak height is greater than the mean of the smoothed signal plus a threshold (e.g., +20pA); (2) Peak prominence is not less than 30pA; (3) Peak width is not less than 3 time points; (4) The distance between adjacent peaks is not less than 10 time points; (5) When the number of candidate peaks is greater than the specified upper limit, the most significant N peaks are selected based on prominence.

[0098] (S23) For each peak, extract a fixed-width signal interval to the left and right, and extract a short interval outside the interval to estimate the baseline current (take the average). If there is no data on the left and right, use the estimated value on the opposite side to ensure baseline stability.

[0099] (S24) Eigenvalue Calculation: Perform the following physical and statistical characteristic calculations for each peak segment:

[0100] (a) Thresholding correlation formula, with ratio parameter for:

[0101] When the ratio parameter is in different threshold ranges, different machine learning models are called to classify the samples respectively;

[0102] The formula for converting the number of molecules by charge transfer is as follows:

[0103]

[0104] Q is the total charge (unit: pA·s);

[0105] The formula for converting the number of molecules by charge transfer is as follows:

[0106]

[0107] Where N is the number of molecules encapsulated in a single current peak, and Q is the charge (in C). Where is Avogadro's constant, F is Faraday's constant, and n is the electron transfer number. It is the current that changes over time (unit: pA). and These are the start and end times of the integration (in seconds).

[0108] (b) Peak height, peak width, average current, and standard deviation;

[0109] (c) Skewness and kurtosis;

[0110] (d) Current peak area normalization (by residence time);

[0111] (e) The ratio of peak-to-peak current difference (ΔI) to baseline;

[0112] (f) Number of charge-transformed molecules (calculated based on Faraday's constant and Avogadro's constant).

[0113] (S25) Plot each peak independently, mark the peak start and end times and current curve, and save it as an image, such as... Figure 3 The diagram shows a visualization of single-peak analysis. The feature data of all peaks are then saved as a structured CSV file to provide training input for subsequent machine learning models.

[0114] (S26) Draw an overlay diagram of the original signal and the smoothed signal, mark each peak position and number, and realize a visual evaluation of the overall signal recognition effect.

[0115] S3. Perform preprocessing operations on the peak feature data extracted from the feature data, including missing value handling, standardization, principal component analysis for dimensionality reduction, and class balance.

[0116] In another embodiment of this application, the extracted vesicle current signal feature data undergoes further preprocessing and feature engineering to improve the robustness and discriminative ability of the subsequent machine learning model. This includes the following steps:

[0117] (S31) Missing values ​​are detected in the collected vesicle peak feature data, and the average value of the numerical column is used for filling to eliminate the impact of incomplete samples on model performance.

[0118] (S32) The Pearson correlation coefficient matrix was used to analyze the linear correlation between the features, and the following methods were employed: Figure 4 The heatmap shown illustrates highly correlated feature pairs, providing a reference for subsequent redundant feature processing and dimensionality reduction.

[0119] (S33) All input features are numerical and Z-score normalization (StandardScaler) is used to ensure that features of different scales have a uniform distribution, which is beneficial for training distance-sensitive models (such as SVM).

[0120] (S34) Divide the dataset into training and test sets at a ratio of 80% / 20%, and use stratified split sampling with balanced label distribution to prevent bias caused by uneven class distribution.

[0121] (S35) Principal component analysis (PCA) is used to retain 95% of the cumulative variance contribution rate, mapping the original high-dimensional feature data to a lower-dimensional space, reducing noise while compressing the feature dimension. The results are referenced. Figure 5 This demonstrates the contribution of different principal components to the cumulative explained variance. The process automatically selects the appropriate dimension, eliminating the need for manual setting.

[0122] (S36) Using an absolute correlation coefficient greater than 0.85 as a threshold to identify highly redundant feature pairs can serve as a basis for selective feature simplification, reducing feature redundancy while maintaining model performance.

[0123] (S37) Save the preprocessed and dimensionality-reduced training dataset as a structured CSV file for use by the model training and deployment module.

[0124] S4. Based on the preprocessed feature data, construct and train at least one machine learning model for automatic identification of vesicle particle size categories.

[0125] This further describes a multi-model ensemble training and classification method for vesicle electrochemical signal feature data. Its core is to construct multiple mainstream machine learning models and use a unified interface for training, optimization, evaluation, and deployment. Specifically, it includes the following steps:

[0126] (S41) Data loading and preprocessing interface, including:

[0127] (S411) Data loading: Load the CSV format data generated by the aforementioned feature engineering stage.

[0128] (S412) Feature selection: Select several feature components after principal component analysis (PCA) as input features, retain useful information and remove redundancy.

[0129] (S413) Label extraction: Extract target classification labels, such as VC_50, VC_100, VC_150, and VC_200, which are categorized according to vesicle size.

[0130] (S414) Label Encoding: Use LabelEncoder to convert category labels from strings to integers to ensure data compatibility.

[0131] (S415) Missing value imputation: Use SimpleImputer’s mean strategy to imput missing values ​​in the data to ensure data integrity.

[0132] (S416) Standardization: StandardScaler is used to standardize all feature dimensions to eliminate the influence of different feature scales.

[0133] (S42) Dataset partitioning strategy

[0134] (S421) Training and test set splitting: The dataset is split into training and test sets in a ratio of 80%:20% using the train_test_split method.

[0135] (S422) Stratified sampling: Enable the stratify parameter to ensure that the vesicle size category distribution is consistent between the training set and the test set, avoiding model bias.

[0136] (S423) Oversampling enhancement: To solve the problem of imbalanced samples, the SMOTE method is used to oversample and enhance the training set.

[0137] (S43) Model Selection and Training Framework: To improve classification accuracy and enhance the system's adaptability, this embodiment constructs and trains the following five types of machine learning models. Each model is clearly defined and implemented specifically for particular scenarios.

[0138] (S431) A set of correlation formulas based on physical and statistical characteristics was designed and constructed to achieve intelligent discrimination of vesicle particle size categories. Specifically, the ratio of signal peak height to peak width was used as a key discrimination feature, and a thresholding formula was constructed for model splitting. When the ratio parameter k is in different threshold ranges, different machine learning models are called to classify the samples respectively.

[0139] At this point, it should be added that, for example Figure 6 As shown, different machine learning methods achieve different levels of accuracy during the training and prediction processes of a model.

[0140] Therefore, multiple sets of binarization formulas can be constructed by combining different thresholds, thereby establishing a dynamic correlation between physical and statistical features and machine learning models. Taking experimental results with k values ​​ranging from 0.25 to 2.25 as an example, such as... Figure 6 As shown, when k=0.25, the accuracy of Support Vector Machine, XGBoost, and Random Forest all reach approximately 0.905, while the accuracy of K-Nearest Neighbors is slightly lower at 0.895. As the threshold increases, the accuracy of K-Nearest Neighbors gradually decreases, while Random Forest and XGBoost remain stable at around 0.905. These results demonstrate that by constructing a reasonable threshold formula and calling different models under different conditions, the robustness and generalization ability of the model can be improved while maintaining classification accuracy.

[0141] Based on the experimental results k Based on the pattern in the interval [0.25, 2.25], the following flow division and priority determination strategy is established:

[0142]

[0143] When k≤1.0, five types of models, including K-nearest neighbors and decision trees, are allowed to be used to balance the flexibility of the low threshold range.

[0144] When k > 1.0, K-nearest neighbors are disabled to prevent a rapid drop in accuracy.

[0145] When k>1.25, decision trees are further disabled, and only random forest, support vector machine and XGBoost are retained to ensure that the classification accuracy is stable above 0.90.

[0146] Within each threshold range where k>1.0, Random Forest and XGBoost are preferentially used because they exhibit the highest stability and robustness within this range.

[0147] (S432) Hyperparameter Tuning

[0148] Hyperparameter tuning is an indispensable part of model optimization. GridSearchCV is used to adjust the hyperparameters of each model to improve its performance. The goal of hyperparameter tuning is to find the optimal set of hyperparameters that improves model performance by searching for different combinations of parameters.

[0149] (1) Decision tree model: The decision tree structure is optimized by adjusting parameters such as the maximum depth of the tree (max_depth), the minimum number of split samples (min_samples_split), and the minimum number of leaf samples (min_samples_leaf).

[0150] (2) Random forest model: optimize the number of trees (n_estimators), the maximum depth of the trees (max_depth) and the minimum number of split samples for each tree (min_samples_split).

[0151] (3) Support Vector Machine (SVM) model: Adjust the kernel type (kernel), penalty coefficient (C) and kernel width (gamma) parameters.

[0152] (4) XGBoost model: optimize learning rate, maximum tree depth and number of trees (n_estimators).

[0153] (5) K-Nearest Neighbors (KNN): Optimize the number of neighbors (n_neighbors), weight strategy (weights) and distance metric (p).

[0154] (S433) Cross-Validation

[0155] To ensure stable performance of each model across different data subsets and avoid overfitting or performance bias caused by the randomness of data partitioning, this invention employs K-fold cross-validation (e.g., 5-fold cross-validation). The specific steps are as follows:

[0156] L1. Divide the dataset into K subsets, use each subset as a validation set once, and use the remaining K-1 subsets as the training set.

[0157] L2. Train and evaluate each model, and ensure the generalization ability of model (3) by calculating the average performance on multiple training and validation sets.

[0158] Cross-validation can effectively assess the stability of the model and reduce random errors caused by data partitioning.

[0159] To address the issue of uneven distribution of vesicle size categories, SMOTE (Synthetic Minority Oversampling) was introduced during cross-validation. This technique involves interpolating and synthesizing new samples in the feature space to enhance the representativeness of minor categories and improve the model's generalization ability.

[0160] During the model training phase, the original dataset is first divided into multiple folds for cross-validation. In each round of cross-validation, only the training fold data of the current round is oversampled using the SMOTE algorithm to generate a balanced training sample set; the validation fold data is not oversampled to maintain the objectivity of the evaluation. Subsequently, the model is trained using the SMOTE-processed training sample set, and the model performance is evaluated using the corresponding validation fold data.

[0161] (S434) Model Training and Evaluation

[0162] In the model selection and training phases, this invention constructs five machine learning models and trains and evaluates each model. The training process for each model combines hyperparameter tuning and cross-validation to ensure optimal performance.

[0163] (1) Decision Tree Model

[0164] Model definition: A decision tree is a tree-structured classifier that constructs multiple decision paths by splitting feature conditions and outputs the classification result at the leaf nodes.

[0165] Implementation: DecisionTreeClassifier is used, combined with grid search to optimize hyperparameters (such as max_depth, min_samples_split, and min_samples_leaf) to select the optimal decision tree structure. This model has strong interpretability and computational efficiency, making it suitable for vesicle size classification tasks.

[0166] (2) Random Forest Model

[0167] Model definition: Random forest is an ensemble learning model composed of multiple decision trees, which improves the model's generalization ability through the "Bagging" technique.

[0168] Implementation: RandomForestClassifier is used to optimize hyperparameters such as the number of trees, maximum depth, and minimum number of split samples through grid search, thereby enhancing the model's robustness. This model can effectively cope with noise and data imbalance issues.

[0169] (3) Support Vector Machine (SVM) model

[0170] Model definition: SVM aims to construct a hyperplane that maximizes the inter-class margin. For nonlinear problems, a kernel function is used to map the input to a high-dimensional space.

[0171] Implementation: Use the SVC model and enable the probabilistic prediction option to adapt to multi-class classification tasks. Optimize classification performance by adjusting the kernel function type (linear, RBF), penalty coefficient C, and kernel width parameter gamma.

[0172] (4) K-Nearest Neighbors (KNN) model

[0173] Model definition: KNN is an instance-based learning method that classifies data based on the distance between samples and the training set.

[0174] Implementation: KNeighborsClassifier is used, and the robustness of the model is enhanced by optimizing the number of neighbors n_neighbors and the weighting method (equal weight / distance weighting). It is suitable for classification tasks with clear feature distributions.

[0175] (5) XGBoost model (eXtreme Gradient Boosting)

[0176] Model definition: XGBoost is a tree-based ensemble learning method that combines regularization terms with efficient feature selection.

[0177] Implementation: The model is trained using XGBClassifier, and hyperparameters such as learning rate, maximum tree depth, and number of trees are tuned to ensure superior performance in high-dimensional, sparse data.

[0178] Furthermore, based on the five machine learning models mentioned above, a multi-dimensional evaluation of the vesicle particle size classification method was conducted. Specific metrics included accuracy, confusion matrix, classification report, ROC curve, and AUC value, among others, to comprehensively assess the model's performance.

[0179] (a) Accuracy refers to the final output results of each model at the terminal. The accuracy of each model on the test set is as follows:

[0180] Decision tree model: 0.8875

[0181] Random Forest model: 0.9062

[0182] SVM model: 0.9234

[0183] XGBoost model: 0.9187

[0184] KNN model: 0.9142

[0185] The above accuracy rate demonstrates that the vesicle particle size classification method of the present invention can effectively distinguish different particle size categories.

[0186] (b) A confusion matrix is ​​a tool used to evaluate the performance of a classification model. It is a square matrix where rows represent the actual class (true label) and columns represent the predicted class (predicted label). Each element represents the comparison between the predicted result for a certain class and the true label. In binary classification problems, the confusion matrix typically contains four values: true positive (TP), false positive (FP), true negative (TN), and false negative (FN).

[0187] For multi-class classification problems, the confusion matrix is ​​expanded into a square matrix containing multiple classes. Each element represents the match between the actual class and the predicted class. The elements on the diagonal represent the proportion of correct classifications, while the elements off-diagonal represent the proportion of incorrect classifications.

[0188] The confusion matrix not only provides classification accuracy but also helps evaluate the model's performance across classes, especially in cases of class imbalance. It can be used for a more comprehensive evaluation through other metrics such as F1 score, precision, and recall.

[0189] The confusion matrix for each model displays the classification results. The confusion probability matrix is ​​generated by normalizing the confusion magnitude moments. See details... Figures 7-11 , used to represent the specific actual result of performing model training and prediction operations in a specific embodiment of the multi-model ensemble classification of vesicle electrochemical signal feature data in this application, wherein, Figure 7 This is the confusion probability matrix of the decision tree model. Figure 8 This is the confusion probability matrix of the random forest model. Figure 9 This is the confusion probability matrix of the SVM model. Figure 10 This is the confusion probability matrix of the XGboost model. Figure 11 This is the confusion probability matrix of the KNN model. By comparing the prediction results of different categories with the actual categories, the classification performance of the model can be clearly seen, providing the prediction accuracy of each category under each model.

[0190] (c) Classification Report, which displays the precision, recall, and f1-score for each category of each model. In a specific embodiment, for the vesicle size classification operation of this application, the classification report for each model is as follows:

[0191] Decision Tree Classification Report:

[0192]

[0193] Random Forest Classification Report:

[0194]

[0195] SVM Classification Report:

[0196]

[0197] XGboost Category Report:

[0198]

[0199] KNN Classification Report:

[0200]

[0201] In summary, the decision tree model has a macro-average f1-score of 0.89, indicating balanced classification performance; the random forest model has a macro-average f1-score of 0.91, demonstrating strong classification ability; the SVM model has a macro-average f1-score of 0.92, showcasing excellent classification performance; the XGBoost model has a macro-average f1-score of 0.92, proving the excellent performance of this ensemble model in vesicle size classification; and the KNN model has a macro-average f1-score of 0.92, demonstrating good stability and accuracy.

[0202] (d) ROC curve and AUC index (multiple types)

[0203] The ROC curve is a tool used to evaluate the performance of classification models. It shows the relationship between the false positive rate (FPR) and the true positive rate (TPR) of a classification model at different decision thresholds. The closer the ROC curve is to the upper left corner, the better the model's classification performance, meaning a high true positive rate and a low false positive rate.

[0204] False positive rate (FPR): The percentage of samples that are incorrectly predicted as positive when they are actually negative. FP represents the number of false positives, and TN represents the number of true negatives.

[0205] True Positive Rate (TPR): The ratio of correctly predicted positive samples out of all actual positive samples; also known as sensitivity or recall. , where TP represents the number of true positives and FN represents the number of false negatives.

[0206] AUC (Area Under the Curve): The value is the area under the ROC curve, representing a comprehensive indicator of the model's classification performance. The AUC value ranges from 0 to 1; the closer the value is to 1, the stronger the model's classification ability. An AUC value of 0.5 indicates that the model's classification ability is comparable to random guessing. Macro-average AUC is calculated by averaging the AUC for each class, reflecting the model's overall performance in multi-class classification tasks.

[0207] The ROC curves for each model show the AUC value for each class, as well as the macro-average AUC value. See details. Figures 12-16 The comparison is used to represent the specific actual results of performing model training and prediction operations in a specific embodiment of the multi-model ensemble classification of vesicle electrochemical signal feature data in this application, wherein, Figure 12 The ROC curve of the decision tree model is shown. Figure 13 This is the ROC curve for the random forest model. Figure 14 Here is the ROC curve of the SVM model. Figure 15 This is the ROC curve for the XGboost model. Figure 16 The ROC curve of the KNN model is shown below. Based on this, we can conclude that:

[0208] Decision tree model: The macro average AUC is 0.96.

[0209] Random forest model: macro-average AUC is 0.98.

[0210] SVM model: Macro average AUC is 0.97.

[0211] XGBoost model: Macro-average AUC is 0.98.

[0212] KNN model: The macro-average AUC is 0.96.

[0213] In summary, after training multiple models, the accuracy and macro-average AUC of each model were compared horizontally. Based on the evaluation results on the test set, the Random Forest model and the XGBoost model performed exceptionally well, both achieving a macro-average AUC of 0.98. The XGBoost model also achieved a high accuracy of 0.9187 on the test set. The SVM model followed closely behind, with a macro-average AUC of 0.97 and a test set accuracy of 0.9234, demonstrating its excellent classification performance. The KNN model and the decision tree model had a macro-average AUC of 0.96, slightly lower than the aforementioned models, but still exhibiting good classification stability and accuracy.

[0214] In summary, all models performed well in the vesicle size classification task, especially in terms of macro-average AUC and accuracy, demonstrating that the method of this invention has broad applicability and superior classification ability under various machine learning models.

[0215] On the other hand, the present invention also discloses a single-vesicle electrochemical signal recognition system for performing the vesicle electrochemical signal recognition and classification method in the above embodiments, specifically including:

[0216] (1) Current acquisition module, including carbon fiber electrode, patch clamp amplifier and data acquisition card, used to record electrical signals caused by vesicle oxidation reaction;

[0217] (2) Signal processing and feature extraction module, which is used to perform signal smoothing, peak detection and multi-dimensional feature construction operations based on the electrical signal data of the current acquisition module, so as to obtain the processed feature extraction data;

[0218] (3) Data preprocessing module, used to perform missing value processing, feature standardization, PCA dimensionality reduction and class balance processing operations on feature extraction data based on signal processing and feature extraction module;

[0219] (4) Classification and prediction module: Load training model. The loaded training model is used to perform particle size prediction and classification on the data processed by the data preprocessing module. In addition, the classification and prediction module also supports the display of ROC analysis chart in a graphical interface, and automatically adjusts parameters and optimizes the data through multi-model grid search using Python language to improve the classification effect.

[0220] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the methods for identifying and classifying electrochemical signals of single vesicle particles in the above embodiments.

[0221] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above method.

[0222] This application also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus.

[0223] Memory, used to store computer programs;

[0224] The processor, when executing the program stored in the memory, implements the above-mentioned method for identifying and classifying the electrochemical signals of single vesicle particles.

[0225] The communication bus mentioned in the above-mentioned electronic devices can be a standard bus for interconnecting peripheral components or an extended industrial standard structure bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.

[0226] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0227] The memory may include random access memory or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0228] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0229] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0230] Furthermore, it should be noted that if any directional indication (such as up, down, left, right, front, back, etc.) is involved in the embodiments of the present invention, the directional indication is only used to explain the relative positional relationship and movement of each component in a specific posture. If the specific posture changes, the directional indication will also change accordingly.

[0231] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, in the embodiments of this invention, "multiple" refers to two or more. Moreover, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

Claims

1. A method for identifying and classifying electrochemical signals of single vesicle particles, characterized in that, include: S1. Fabricate carbon fiber microelectrodes and apply voltage in constant potential mode of patch clamp system to collect the picoampere-level current response generated by the oxidation of the contents of a single vesicle on the electrode surface; S2. The acquired raw current signal is smoothed and filtered, and a threshold-based multi-parameter peak identification method is used to realize feature peak identification and multi-dimensional physical / statistical feature extraction; S3. Perform preprocessing operations on the peak feature data in the extracted feature data, including missing value handling, standardization, principal component analysis dimensionality reduction, and class balance; S4. Based on the preprocessed feature data, construct and train at least one machine learning model for automatic identification of the particle size category of vesicle particles; The specific operation process of step S2 includes: S21. The original current signal is smoothed by Gaussian filtering. Feature peaks are identified based on multiple parameters including peak height, significance, width, and spacing. The top N peaks are selected according to their significance. S22. Perform physical and statistical characteristic calculations, including charge integral and molecule number conversion, calculate peak height, peak width, and current statistics, and then calculate the peak area time normalized value to the baseline current ratio. S23. Generate a signal overview diagram and a single-peak analysis diagram with peak position markers.

2. The method for identifying and classifying electrochemical signals of single vesicle particles as described in claim 1, characterized in that, The preparation process of the carbon fiber microelectrode in step S1 includes: connecting a specified micron-sized carbon fiber filament with a copper wire using conductive silver paste and embedding it into a glass capillary; filling it with epoxy resin and ensuring that the carbon fiber filament is exposed; and shearing and polishing the front end of the carbon fiber microelectrode to form a smooth disk structure, which serves as the working electrode of the single vesicle electrochemical measurement system.

3. The method for identifying and classifying electrochemical signals of single vesicle particles as described in claim 1, characterized in that, In the S21 characteristic peak identification process, a fixed width signal interval is truncated to the left and right of each group of characteristic peaks, and the baseline average current is estimated by extracting a short interval outside the segment. If there is no data to the left or right of the characteristic peak, then the relative side estimate is used.

4. The method for identifying and classifying electrochemical signals of single vesicle particles as described in claim 3, characterized in that, During the calculation of physical and statistical characteristics in step S22: The formula for the integral of charge is: in, It is the total charge. For time Changing current, and Each peak segment is assigned a start and end time. The number of molecules to be converted based on Faraday's constant and Avogadro's constant is calculated using the following formula: Where N is the number of molecules encapsulated in a single current peak, and Q is the charge. is Avogadro's constant, F is Faraday's constant, and n is the electron transfer number; Calculate the time-normalized value of peak area based on residence time; The baseline current ratio is the ratio of the peak-to-peak current difference ΔI to the baseline.

5. The method for identifying and classifying electrochemical signals of single vesicle particles as described in claim 1, characterized in that, The machine learning models in step S4 include five types: parallel training decision tree, random forest, SVM, XGBoost, and KNN. The following optimization strategies are adopted in the construction and training of the machine learning model: the training set and the test set are divided by stratified sampling at 80%:20%, key hyperparameters are tuned by grid search, and 5-fold cross-validation is used to ensure generalization ability. In step S4, a confusion matrix is ​​used to evaluate the recognition and classification performance of the machine learning model.

6. The method for identifying and classifying electrochemical signals of single vesicle particles as described in claim 5, characterized in that, During the execution of step S4, a set of thresholded correlation formulas are constructed, which include ratio parameters. for: When the ratio parameter k is in different threshold ranges, different machine learning models are called to classify the specified samples to achieve optimal discrimination of vesicle particle size category.

7. A system for recognizing single-vesicle electrochemical signals, characterized in that, The method for identifying and classifying vesicle electrochemical signals according to any one of claims 1-6 specifically includes: The current acquisition module, including a current amplifier and an analog-to-digital converter, uses carbon fiber microelectrodes as working electrodes to record the electrical signals generated by the oxidation of vesicle contents. The signal processing and feature extraction module is used to perform signal smoothing, peak detection and multi-dimensional feature construction operations based on the electrical signal data from the current acquisition module in order to obtain the processed feature extraction data. The data preprocessing module is used to perform missing value processing, feature standardization, PCA dimensionality reduction and class balance processing operations on the feature extraction data based on the signal processing and feature extraction module. The classification and prediction module loads the training model and is used to perform particle size prediction and classification on the data processed by the data preprocessing module using the loaded training model.

8. The single-vesicle electrochemical signal recognition system as described in claim 7, characterized in that, The classification and prediction module supports a graphical interface for displaying ROC analysis charts and uses Python to automatically adjust and optimize parameters through multi-model grid search of data to improve classification performance.

Citation Information

Patent Citations

  • Detection device and detection method for cell in-situ single vesicle inclusions

    CN112595842A

  • Recognition system and recognition method of nano vesicles

    CN118430656A