Method for identifying breast cancer subtype by exosome surface marker
Multi-dimensional fluorescence intensity matrix is constructed through multi-color fluorescence labeling and flow cytometry, combined with fluorescence resonance energy transfer technology and machine learning algorithms, and the accuracy problem of exosome surface markers is solved, achieving support for non-invasive diagnosis and individualized treatment.
Patent Information
- Application Number
- CN202510580047.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art is difficult to accurately identify breast cancer molecular subtypes through the characteristics of exosome surface markers in blood samples, and cannot meet the needs of clinical non-invasive diagnosis and individualized treatment.
Multicolor fluorescence labeling exosome surface markers were constructed, and the interaction between markers was analyzed using fluorescence resonance energy transfer technology, combined with K-mean clustering and entropy weighting method to screen the marker combination, and a support vector machine classification model was constructed to achieve accurate identification of breast cancer subtypes.
It improves the accuracy of exosome markers to identify breast cancer subtypes, realizes support for non-invasive diagnosis and individualized treatment, and solves the limitations of traditional liquid biopsy.
Smart Images

Figure CN120445960A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of exosome surface marker detection, and specifically relates to a method for identifying breast cancer subtypes using exosome surface markers. Background Art
[0002] Breast cancer is one of the most common malignancies in women. Based on molecular phenotype, it can be divided into multiple subtypes, including Luminal A, Luminal B, HER2-positive, and triple-negative. Treatment options and prognosis vary significantly among these subtypes. Traditionally, breast cancer subtyping relies primarily on tissue biopsy, where molecular subtypes are determined through immunohistochemistry or gene expression analysis. However, tissue biopsy is invasive and lacks real-time monitoring, limiting its effectiveness for early screening and dynamic evaluation of treatment efficacy.
[0003] In recent years, circulating tumor exosomes have garnered widespread attention as a key component of "liquid biopsy." Existing exosome detection technologies primarily focus on analyzing the total amount of exosomes or the expression levels of single markers, such as ultracentrifugation combined with Western blotting or ELISA for specific proteins. However, these methods struggle to fully capture the systematic differences in exosome surface marker expression patterns.
[0004] The existing technology lacks an analytical method that can comprehensively consider the interactions and expression patterns of multiple markers on the surface of exosomes. This results in insufficient accuracy in accurately identifying breast cancer molecular subtypes using blood exosomes, making it difficult to meet the clinical needs of non-invasive diagnosis and personalized treatment. In other words, the existing technology has a technical problem of making it difficult to accurately identify breast cancer molecular subtypes based on the characteristics of exosome surface markers in blood samples. Summary of the Invention
[0005] In view of this, the present invention provides a method for identifying breast cancer subtypes using exosome surface markers, which can solve the technical problem in the prior art that it is difficult to accurately identify breast cancer molecular subtypes through the characteristics of exosome surface markers in blood samples.
[0006] The present invention is implemented as follows: The present invention provides a method for identifying breast cancer subtypes using exosome surface markers, comprising: collecting blood samples from breast cancer patients, isolating and extracting exosomes; performing multi-color fluorescent labeling on exosome surface markers; measuring the fluorescence intensity of surface markers by flow cytometry and constructing a multidimensional fluorescence intensity matrix; calculating the interaction strength of surface markers using the fluorescence resonance energy transfer equation; performing K-means cluster analysis on samples of known subtypes to determine the fluorescence characteristic pattern; calculating the Mahalanobis distance between cluster centers; determining the cluster silhouette coefficient and cohesion to screen the marker combination with the highest discrimination; applying the entropy weight method to optimize the function to evaluate the contribution of markers; constructing a support vector machine classification model; inputting the sample to be identified into the model and calculating the distance score to the cluster center of each subtype; determining the breast cancer subtype according to the minimum distance principle and calculating the classification reliability index.
[0007] Among them, the step of multi-color fluorescent labeling of exosome surface markers is to use fluorescently labeled antibodies to perform multi-color fluorescent labeling of exosome surface markers, including CD9, CD63, CD81 and breast cancer subtype-related surface proteins.
[0008] Among them, the breast cancer subtype-related surface proteins include human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, and matrix metalloproteinases.
[0009] The multidimensional fluorescence intensity matrix refers to a data structure composed of the fluorescence signal intensities of multiple surface markers in each exosome sample, where each row represents a sample and each column represents the fluorescence intensity value of a surface marker.
[0010] Among them, the silhouette coefficient is an indicator to measure the quality of clustering. It calculates the standardized value of the difference between the average similarity of each sample with other samples in the same cluster and the average similarity of samples in other clusters. The closer the value is to 1, the better the clustering effect. The cohesion refers to the average distance from all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification.
[0011] Among them, the K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence characteristic patterns. It minimizes the sum of the squared distances from sample points to their cluster centers through iterative optimization. The input is a multidimensional fluorescence intensity matrix, and the output is the fluorescence characteristic patterns of different breast cancer subtypes. The Mahalanobis distance is a distance measurement method that takes into account the correlation between features. The input is the cluster centers of different subtypes, and the output is a quantitative indicator of the difference in exosome surface marker expression between subtypes.
[0012] Among them, the fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the surface of exosomes. The input includes the fluorescence donor quantum yield, the distance between fluorescence donor and acceptor molecules, the fluorescence donor and acceptor spectral overlap integral, the refractive index, and the fluorescence resonance energy transfer direction factor. The output is the interaction strength between surface markers.
[0013] The distance score is the Mahalanobis distance between the sample to be identified and the cluster center of each subtype, which is calculated by the support vector machine classification model and is used for breast cancer subtype determination.
[0014] The classification reliability index is obtained by calculating the ratio of the distance score between the sample to be tested and the nearest subtype cluster center to the distance score between the sample to be tested and the nearest subtype cluster center. The smaller the value, the more reliable the classification result.
[0015] The step of constructing a support vector machine classification model is to construct a support vector machine classification model based on the screened surface marker combination and weight coefficient, and use known subtype samples to train the model.
[0016] The present invention establishes a high-precision breast cancer subtype classification model by constructing a multicolor fluorescent labeling system and flow cytometry measurement to obtain a multidimensional fluorescence intensity matrix of exosome surface markers, and combining fluorescence resonance energy transfer technology to analyze the interactions between markers.
[0017] This method uses a K-means clustering algorithm to automatically identify the fluorescence signature patterns of different breast cancer subtypes, quantifies differences between subtypes using the Mahalanobis distance, and selects the optimal marker combination by combining the silhouette coefficient and cohesion index. Furthermore, the entropy weighting method is used to optimize and assign appropriate weights to each marker. Ultimately, a support vector machine classification model is constructed to accurately identify subtypes. This systematic analytical approach significantly improves the accuracy of exosome markers in identifying breast cancer subtypes.
[0018] This invention solves the technical problem that traditional liquid biopsy is difficult to accurately identify breast cancer molecular subtypes, and realizes the accurate judgment of breast cancer subtypes through the multidimensional characteristics of exosome surface markers, providing important technical support for non-invasive diagnosis, prognosis assessment and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0021] like Figure 1FIG. 1 is a flow chart of a method for identifying breast cancer subtypes using exosome surface markers provided by the present invention. The method comprises the following steps:
[0022] S01. Collect blood samples from patients with known breast cancer subtypes and patients with subtypes to be identified, and separate and extract exosomes by ultracentrifugation;
[0023] S02. Use fluorescently labeled antibodies to perform multi-color fluorescent labeling of exosome surface markers, including CD9, CD63, CD81, and surface proteins related to breast cancer subtypes;
[0024] S03. Measure the fluorescence intensity of all surface markers in each exosome sample using flow cytometry and construct a multidimensional fluorescence intensity matrix for each sample;
[0025] S04. Calculate the interaction strength between exosome surface markers using the fluorescence resonance energy transfer equation and analyze the conformational characteristics of surface proteins of different exosome subtypes;
[0026] S05. Perform K-means clustering algorithm analysis on the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes to determine the fluorescence characteristic patterns of different breast cancer subtypes;
[0027] S06. Calculate the Mahalanobis distance between cluster centers of different subtypes and establish a quantitative index for the difference in expression of exosome surface markers between subtypes;
[0028] S07. Measure the silhouette coefficient and cohesion of each known subtype cluster to screen for the surface marker combination with the highest discrimination;
[0029] S08. Apply the entropy weight method to optimize the function to quantitatively evaluate the contribution of surface markers and determine the weight coefficient of each marker in subtype identification;
[0030] S09. Construct a support vector machine classification model based on the screened surface marker combination and weight coefficients, and train the model using samples of known subtypes;
[0031] S10, inputting the multidimensional fluorescence intensity matrix of the exosome samples of the patient with the subtype to be identified into the trained support vector machine classification model, and calculating the distance score between the matrix and the cluster center of each subtype;
[0032] S11. Determine the breast cancer subtype to which the patient sample to be identified belongs based on the minimum distance principle, and calculate the classification reliability index.
[0033] The multidimensional fluorescence intensity matrix refers to a data structure composed of the fluorescence signal intensities of multiple surface markers in each exosome sample. Each row represents a sample, and each column represents the fluorescence intensity value of a surface marker. It is obtained by flow cytometry measurement in step S03 and is used for K-means clustering algorithm analysis in step S05 and support vector machine classification model input in step S10.
[0034] The fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the surface of exosomes. The input includes the fluorescence donor quantum yield measured by step S02, the distance between the fluorescence donor and acceptor molecules measured by step S03, the fluorescence donor and acceptor spectral overlap integral obtained by step S02, the refractive index measured by step S02, and the fluorescence resonance energy transfer direction factor calculated by step S03. The output is the interaction strength between surface markers, which is used for surface marker combination screening in step S07.
[0035] The K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence characteristic patterns. It minimizes the sum of the squared distances from sample points to their cluster centers through iterative optimization. The input is the multidimensional fluorescence intensity matrix obtained in step S03, and the output is the fluorescence characteristic patterns of different breast cancer subtypes, which are used for expression difference analysis between subtypes in step S06.
[0036] Among them, Mahalanobis distance is a distance measurement method that takes into account the correlation between features. Compared with Euclidean distance, it is more suitable for measuring the degree of difference between subtype cluster centers in a multidimensional feature space. The input is the different subtype cluster centers obtained in step S05, and the output is a quantitative index of the difference in expression of exosome surface markers between subtypes, which is used for discrimination evaluation in step S07.
[0037] Among them, the silhouette coefficient is an indicator to measure the clustering quality. The normalized value of the difference between the average similarity of each sample with other samples in the same cluster and the average similarity of samples in other clusters is calculated. The closer the value is to 1, the better the clustering effect. The input is the clustering result obtained in step S05 and the multidimensional fluorescence intensity matrix obtained in step S03. The output is used for surface marker combination screening in step S07.
[0038] Among them, cohesion refers to the average distance from all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification. The input is the clustering result obtained in step S05 and the multidimensional fluorescence intensity matrix obtained in step S03. The output is used for surface marker combination screening in step S07.
[0039] Among them, the entropy weight optimization function is used to objectively evaluate the contribution of different surface markers in the identification of breast cancer subtypes. The input includes the coefficient of variation of each marker between different subtypes calculated in step S06, the marker expression correlation matrix obtained in step S04, the marker information entropy value calculated in step S05, the marker expression abundance measured in step S03, and the marker discrimination ability index in known subtype samples obtained in step S07. The output is the weight coefficient of each surface marker, which is used to construct the support vector machine classification model in step S09.
[0040] Among them, the support vector machine classification model is a supervised learning algorithm that classifies data by constructing a hyperplane in a high-dimensional space. The input is the surface marker combination screened in step S07 and the weight coefficient determined in step S08. The output is a breast cancer subtype classifier, which is used to distinguish the subtype of the sample to be identified in step S10.
[0041] The distance score is the Mahalanobis distance between the sample to be identified and the cluster center of each subtype, which is calculated by the support vector machine classification model in step S10 and is used for breast cancer subtype determination in step S11.
[0042] The classification reliability index is obtained by calculating the ratio of the distance score between the sample to be tested and the nearest subtype cluster center obtained in step S10 to the distance score between the sample to be tested and the nearest subtype cluster center. The smaller the value, the more reliable the classification result is, which is used as a reference for clinical diagnosis.
[0043] Among them, breast cancer subtype-related surface proteins refer to protein markers specifically expressed on the surface of exosomes secreted by breast cancer cells of different molecular subtypes, including human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, matrix metalloproteinases and other molecules. They are identified and labeled by the fluorescently labeled antibodies in step S02. The difference in their expression patterns is the key basis for the K-means clustering algorithm in step S05 to distinguish different breast cancer subtypes.
[0044] The specific implementation of the above steps is described in detail below.
[0045] The specific implementation method of step S01 is to first collect 5-10 ml of peripheral venous blood from patients with known molecular subtypes of breast cancer and patients with subtypes to be identified, use blood collection tubes containing EDTA anticoagulant, and store at 4°C for no more than 2 hours. After collecting the blood, exosomes are extracted by differential centrifugation. First, centrifuge at 300×g for 10 minutes to remove cellular components. After taking the supernatant, centrifuge at 2000×g for 20 minutes to remove cell debris, and then centrifuge at 10000×g for 30 minutes to remove large particles. Finally, use an ultracentrifuge to precipitate the exosomes at 100,000×g for 70 minutes. The precipitate is resuspended with sterile phosphate buffer and washed again by ultracentrifugation at 100,000×g for 60 minutes. The exosome precipitate finally obtained is resuspended with 200μl sterile phosphate buffer for use. Nanoparticle tracking analysis technology is used to verify that the size distribution of exosomes is in the range of 30-150nm, and the concentration reaches 10 10 To ensure the accuracy of subsequent analysis.
[0046] In step S02, the extracted exosome sample is divided into several equal portions. Each portion is then labeled with a different fluorophore-labeled antibody. These antibodies include common exosome markers such as anti-CD9-FITC, anti-CD63-PE, and anti-CD81-APC, as well as surface proteins associated with breast cancer subtypes such as anti-human epidermal growth factor receptor 2-Cy5, anti-estrogen receptor-Cy3, anti-progesterone receptor-Cy7, anti-Ki-67-PerCP, anti-epithelial cell adhesion molecule-AF647, anti-CD44-PE-Cy7, and anti-matrix metalloproteinase-BV421. The antibody dilution ratio is 1:100 to 1:500. The samples are incubated at 4°C for 2 hours, gently mixing every 30 minutes. After incubation, the samples are washed three times by centrifugation at 10,000 × g for 10 minutes to remove unbound antibodies. The washed samples are resuspended in 100 μl of sterile phosphate buffered saline to ensure sufficient binding of the fluorescently labeled antibodies to the exosome surface markers. At the same time, a negative control group was set up, and the specific antibody was replaced with an isotype control antibody to correct nonspecific fluorescence signals.
[0047] A specific implementation of step S03 involves measuring the fluorescence intensity of each surface marker in each exosome sample using a high-resolution flow cytometer. The flow cytometer excitation light source is set to three wavelengths: 488 nm, 561 nm, and 640 nm, corresponding to the fluorescence signals emitted by different fluorophores. Acquisition parameters are set as follows: the forward scatter detector threshold is set to 200 to eliminate background noise, the acquisition rate is controlled at 1000-2000 events / second, and at least 50,000 valid events are acquired for each sample. The acquired raw data is processed using a fluorescence compensation matrix within the flow cytometry software to eliminate spectral overlap interference between different fluorophores. A multidimensional fluorescence intensity matrix is constructed for each sample, with each row representing an exosome sample and each column representing the median fluorescence intensity value of a surface marker. All sample data are normalized using the flow cytometry software to eliminate batch effects and instrument errors, resulting in a standardized multidimensional fluorescence intensity matrix for subsequent analysis.
[0048] The specific implementation of step S04 is to calculate the interaction strength between different markers on the surface of exosomes based on the aforementioned flow cytometry data. The fluorescence resonance energy transfer equation is used: The calculation is performed, where E is the energy transfer efficiency, R0 is the distance at which 50% energy transfer occurs (Förster distance), and r is the actual distance between the fluorescent donor and the acceptor molecules. The Förster distance R0 is calculated by the formula R0 = 0.211 [κ 2 ·n -4 Q D ·J(λ)] 1 / 6 Calculate, where κ is the direction factor (usually 2 / 3), n is the refractive index of the medium (about 1.35 in plasma), Q D is the quantum yield of the fluorescence donor, and J(λ) is the overlap integral of the donor emission spectrum and the acceptor absorption spectrum. Based on the energy transfer efficiency between different markers, a surface protein interaction network was constructed to analyze the conformational characteristics and spatial arrangement patterns of exosome surface proteins from different breast cancer subtypes. When the energy transfer efficiency E is greater than 0.1, the two markers are considered to have an effective interaction. The exosome surface marker interaction strength data obtained in this step will be used for subsequent marker combination screening.
[0049] The specific implementation of step S05 is to apply the K-means clustering algorithm to analyze the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes. First, the optimal number of clusters K is determined. The elbow rule and silhouette coefficient are used to evaluate the clustering effect of different K values (2 to 6). The parameter with the highest silhouette coefficient and the smallest K value is selected as the final number of clusters. Then, K cluster centers are initialized, and the K-means++ algorithm is used to optimize the selection of initial centers to avoid local optimal solutions caused by random initialization. The K-means iterative process is performed to calculate the Euclidean distance of each sample to each cluster center, assign the sample to the cluster with the closest distance, and recalculate the coordinates of each cluster center until the change in the position of the cluster center is less than a preset threshold (usually set to 0.001) or the maximum number of iterations is reached (usually 100 times). Finally, K clusters are obtained, each cluster represents the fluorescence characteristic pattern of a breast cancer subtype, and the coordinates of the cluster center reflect the typical expression pattern of the exosome surface markers of that subtype.
[0050] The specific implementation of step S06 is to calculate the Mahalanobis distance between the cluster centers of different subtypes to quantitatively evaluate the difference in the expression of surface markers of exosomes of each subtype. The Mahalanobis distance calculation formula is:
[0051] Where x and y are the cluster center vectors of two different subtypes, and S is the covariance matrix of all samples. First, the covariance matrix S of all samples obtained in step S05 is calculated, and the singular value decomposition is performed on it to ensure that the matrix is invertible. Then, the Mahalanobis distance between any two subtype cluster centers is calculated to construct an inter-subtype distance matrix. The larger the value in the distance matrix, the more significant the difference in the expression pattern of exosome surface markers between the two subtypes. The threshold is set to 3.0. When the Mahalanobis distance between two subtypes is greater than this threshold, the two subtypes are considered to have significant discrimination. The distance matrix data is used to screen the surface marker combination with the highest discrimination in the subsequent steps.
[0052] The specific implementation of step S07 is to determine the silhouette coefficient and cohesion of each known subtype cluster, so as to screen the surface marker combination with the highest discrimination. The silhouette coefficient calculation formula is: s(i) = (b(i) - a(i)) /
[0053] max(a(i), b(i)), where a(i) is the average distance between sample i and other samples in the same cluster, and b(i) is the average distance between sample i and the closest sample in the other cluster. Cohesion is calculated as the average distance from all samples in a cluster to the cluster center. First, the cluster silhouette coefficient and cohesion of all surface marker combinations are calculated. Then, recursive feature elimination is used to remove each surface marker one at a time, and K-means clustering is re-performed to recalculate the new silhouette coefficient and cohesion. When the silhouette coefficient exceeds 0.7 and the cohesion is less than 30% of the average distance between cluster centers, the marker combination is considered to have good discrimination. By comparing the performance indicators of different marker combinations, the optimal surface marker combination is determined, which should achieve the highest subtype discrimination ability with the minimum number of features.
[0054] The specific implementation of step S08 is to use the entropy weight optimization function to quantitatively evaluate the contribution of surface markers. First, the coefficient of variation CV of each surface marker between different subtypes is calculated. j =σ j / μ j , where σ j is the standard deviation of marker j among subtypes, μ j Then calculate the information entropy of each marker where p ij is the normalized expression value of marker j in subtype i, and m is the number of subtypes. The difference coefficient d of each marker is calculated based on information entropy j =1-e j Combined with the marker expression correlation matrix obtained in step S04, the marker discrimination ability index obtained in step S07, and the marker expression abundance, a comprehensive evaluation function F is constructed. j =w1·CV j +w2·(1-r j )+w3·d j +w4·A j +w5·D j , where w1~w5 are the weight coefficients of each factor (all set to 0.2), r j is the average correlation coefficient between marker j and other markers, A j is the marker expression abundance, D j The weight coefficient of each surface marker in subtype identification is finally determined. Where n is the total number of markers.
[0055] The specific implementation of step S09 is to construct a support vector machine classification model based on the screened surface marker combination and weight coefficient. First, the sample data set is randomly divided into a training set (accounting for 75%) and a validation set (accounting for 25%). The radial basis function is selected as the kernel function: K(x, y) = exp(-γ||xy|| 2 ), where the optimal value of parameter γ is determined by the grid search method (the range is set to 0.001~10). In order to deal with multi-class problems, a one-to-one strategy is adopted to construct multiple binary support vector machines. The support vector machine penalty parameter C (the range is set to 0.1~100) and kernel function parameters are optimized, and the performance of different parameter combinations is evaluated by five-fold cross validation. The marker weight coefficient determined in step S08 is introduced in the training process, and the contribution of key markers is enhanced by the weighted support vector machine algorithm. The classification accuracy of the final model on the validation set should reach more than 85%, and the sensitivity and specificity should both exceed 80% before it can be used for subtype discrimination of subsequent samples to be identified.
[0056] The specific implementation method of step S10 is to input the multidimensional fluorescence intensity matrix of the exosome sample of the patient with the subtype to be identified into the trained support vector machine classification model. First, the sample to be identified is subjected to the same standardization processing as the training set to ensure the consistency of the data distribution. Then, the fluorescence intensity data corresponding to the optimal surface marker combination determined in step S07 is extracted to form a feature vector. The feature vector is input into the support vector machine classification model, and the distance from the sample to each hyperplane is calculated. Based on these distance values, the Mahalanobis distance score between the sample and the cluster center of each subtype is calculated. The distance score calculation takes into account the marker weight coefficient determined in step S08, so that important markers play a greater role in classification judgment. The score calculation formula is: Where W j is the weight coefficient of marker j, d ij is the standardized distance difference between the sample to be tested and the cluster center of subtype i on marker j.
[0057] The specific implementation of step S11 is to determine the breast cancer subtype to which the patient sample to be identified belongs based on the minimum distance principle. First, the distance scores of each subtype calculated in step S10 are compared, and the sample is classified into the subtype with the smallest distance score. Then, the classification reliability index R=d is calculated. min / d sec , where d min is the distance score between the sample and the nearest subtype cluster center, d secThe R value represents the distance between the sample and the next closest subtype cluster center. A smaller R value indicates a more reliable classification result. Generally, an R value < 0.6 is considered highly reliable, 0.6 ≤ R < 0.8 is moderately reliable, and R ≥ 0.8 is considered lowly reliable. The sample's position within the distribution of the closest subtype is also calculated, and the Mahalanobis distance P value is used to assess whether the sample is an outlier. A P value < 0.05 indicates an atypical presentation or a new subtype. A subtype identification report is generated, including the sample subtype determination result, classification reliability index, and key surface marker expression levels, providing an accurate basis for clinical diagnosis.
[0058] The mathematical model or calculation process involved in the present invention is described in detail below.
[0059] In step S03, a multidimensional fluorescence intensity matrix is constructed, which is specifically represented as follows:
[0060]
[0061] Where M is the multidimensional fluorescence intensity matrix; I ij represents the median fluorescence intensity value of the jth surface marker in the i-th sample; m is the number of samples; n is the number of surface marker types.
[0062] The Z-score method is used for data standardization, which is expressed as follows:
[0063]
[0064] Where Z ij is the normalized fluorescence intensity value; I ij is the original fluorescence intensity value; μ j is the average fluorescence intensity of the jth surface marker in all samples; σ j is the standard deviation of the jth surface marker in all samples.
[0065] In step S04, the fluorescence resonance energy transfer equation is calculated as follows:
[0066]
[0067] Where E is the energy transfer efficiency, ranging from 0 to 1; R0 is the distance at which 50% energy transfer occurs (Förster distance), in nanometers; r is the actual distance between the fluorescent donor and acceptor molecules, in nanometers.
[0068] The calculation formula of Foster distance R0 is:
[0069] R0=0.211×[κ 2 ×n -4 ×Q D ×J(λ)] 1 / 6;
[0070] Where κ is the orientation factor, which describes the spatial orientation relationship between the donor emission dipole and the acceptor absorption dipole. Its theoretical value range is 0 to 4, and it is usually 2 / 3 under random orientation conditions. n is the refractive index of the medium, which is about 1.35 in plasma. Q D is the quantum yield of the fluorescence donor, ranging from 0 to 1; J(λ) is the overlap integral of the donor emission spectrum and the acceptor absorption spectrum, with the unit of M -1 cm 3 .
[0071] The calculation formula of spectral overlap integral J(λ) is:
[0072] J(λ)=∫F D (λ)×ε A (λ)×λ 4 ×dλ;
[0073] Where, F D (λ) is the normalized donor emission spectrum; ε A (λ) is the molar absorption coefficient of the receptor, in M -1 cm -1 ;λ is the wavelength, in nanometers.
[0074] Considering the spatial mobility of exosome surface proteins, the modified energy transfer efficiency calculation formula is:
[0075] E actual =E×(1-α×D);
[0076] Where, E actual is the actual energy transfer efficiency; E is the theoretical energy transfer efficiency; α is the mobility correction coefficient, ranging from 0 to 0.5; D is the diffusion coefficient of exosome surface proteins, in μm 2 / s.
[0077] In step S05, the objective function of the K-means clustering algorithm is:
[0078]
[0079] Where J is the objective function value that needs to be minimized; k is the number of clusters; C i is the i-th cluster; x is the sample point (multidimensional fluorescence intensity vector); μ i is the center point of the i-th cluster; ||x-μ i || represents the distance from the sample point x to the cluster center μ i The Euclidean distance of .
[0080] The formula for calculating Euclidean distance is:
[0081]
[0082] Where x j is the coordinate of the sample point x in the jth dimension (fluorescence intensity value of the jth surface marker); μ ij is the coordinate of the i-th cluster center in the j-th dimension; n is the number of surface marker types.
[0083] The cluster center update formula is:
[0084]
[0085] In the formula, |C i | is the number of samples in the i-th cluster.
[0086] In the K-means++ algorithm, the probability of initial cluster center selection is calculated as:
[0087]
[0088] Where P(x) is the probability of selecting sample point x as the next cluster center; D(x) is the distance from sample point x to the nearest selected cluster center; X is the set of all sample points.
[0089] In step S06, the Mahalanobis distance calculation formula is:
[0090]
[0091] Where d(x, y) is the Mahalanobis distance between the two cluster centers x and y; x and y are the cluster center vectors of two different subtypes; S is the covariance matrix of all samples; S -1 is the inverse of the covariance matrix.
[0092] The calculation formula of the covariance matrix S is:
[0093]
[0094] Where m is the total number of samples; x i is the i-th sample point; is the mean vector of all samples.
[0095] Considering that the covariance matrix may not be invertible, singular value decomposition plus regularization is used:
[0096] S′=S+λI;
[0097] Where S′ is the regularized covariance matrix; λ is the regularization parameter, which is usually between 0.01 and 0.1; I is the identity matrix.
[0098] In step S07, the silhouette coefficient calculation formula is:
[0099]
[0100] Where s(i) is the silhouette coefficient of sample i, ranging from -1 to 1. The closer it is to 1, the better the clustering effect. a(i) is the average distance between sample i and other samples in the same cluster. b(i) is the average distance between sample i and the samples in the closest cluster.
[0101] The calculation formulas for a(i) and b(i) are:
[0102]
[0103] Where C i is the cluster to which sample i belongs; |C i | is the number of samples in the cluster; d(i, j) is the distance between sample i and sample j; C k are other clusters to which i does not belong.
[0104] The formula for calculating cohesion is:
[0105]
[0106] Where Cohesion(C i ) is the cohesion of the i-th cluster; |C i | is the number of samples in the i-th cluster; x is the number of clusters C i The sample points in μ i is the center point of the i-th cluster; ||x-μ i || is the distance from the sample point x to the cluster center μ i The Euclidean distance of .
[0107] The overall clustering quality evaluation function is:
[0108]
[0109] Where Q is the clustering quality score, the higher the value, the better the clustering effect; m is the total number of samples; k is the number of clusters; α and β are weight coefficients, usually α = 0.7 and β = 0.3.
[0110] In step S08, the coefficient of variation is calculated as follows:
[0111]
[0112] Where, CV j is the coefficient of variation of the jth surface marker; σ j is the standard deviation of the marker among each subtype; μ j is the average expression value of the marker.
[0113] The formula for calculating marker information entropy is:
[0114]
[0115] Where, e j is the information entropy of the jth surface marker, ranging from 0 to 1, and the closer to 0, the higher the discrimination; k is the number of breast cancer subtypes; p ij is the normalized expression value of the j-th marker in the i-th subtype, calculated as: where x ij is the average expression value of the j-th marker in the i-th subtype.
[0116] The formula for calculating the coefficient of difference is:
[0117] d j =1-e j ;
[0118] Where, d j is the difference coefficient of the jth surface marker, ranging from 0 to 1. The closer it is to 1, the higher the discrimination.
[0119] The formula for calculating the correlation of marker expression is:
[0120]
[0121] Where r jl is the correlation coefficient between the jth marker and the lth marker, ranging from -1 to 1; x ij is the expression value of the jth marker in the i-th sample; is the average expression value of the jth marker; m is the total number of samples.
[0122] The formula for calculating the average correlation coefficient is:
[0123]
[0124] Where r j is the absolute value of the average correlation coefficient between the jth marker and all other markers; n is the total number of markers.
[0125] The calculation formula of the discrimination ability index is:
[0126]
[0127] Where D j is the discrimination ability index of the jth marker; μ ij is the average expression value of the jth marker in the i-th subtype; σ ijis the standard deviation of the jth marker in the ith subtype; k is the number of subtypes.
[0128] The calculation formula of the comprehensive evaluation function is:
[0129] F j =w1×CV j +w2×(1-r j )+w3×d j +w4×A j +w5×D j ;
[0130] Where, F j is the comprehensive score of the jth surface marker; w1~w5 are the weight coefficients of each factor, all set to 0.2; A j The abundance of marker expression was measured by flow cytometry and normalized to a value range of 0 to 1.
[0131] The final weight coefficient calculation formula of surface markers is:
[0132]
[0133] Where W j is the weight coefficient of the jth surface marker in subtype identification, satisfying
[0134] In step S09, the radial basis function kernel calculation formula in the support vector machine is:
[0135] K(x, y) = exp(-γ||xy|| 2 );
[0136] Where K(x, y) is the kernel function value; x and y are the eigenvectors of the two samples; γ is the kernel function parameter, and the optimal value is determined by grid search, ranging from 0.001 to 10; ||xy|| is the Euclidean distance between the eigenvectors of the two samples.
[0137] The calculation formula of weighted Euclidean distance considering the weight of the marker is:
[0138]
[0139] Where, ||xy|| w is the weighted Euclidean distance; W j is the weight coefficient of the jth surface marker; x j and y j are the eigenvalues of the two samples in the jth dimension; n is the number of feature dimensions.
[0140] The optimization objective function of the weighted support vector machine is:
[0141]
[0142] Constraint: y i (w T φ(x i )+b)≥1-ξ i ,ξ i ≥0, i=1, 2, ..., m;
[0143] Where w is the normal vector of the hyperplane; b is the intercept; ξ i is a slack variable; C is a penalty parameter, ranging from 0.1 to 100; y i is the category label of sample i; φ(x i ) is the feature vector of sample i after being mapped by the kernel function; m is the number of training samples.
[0144] In step S10, the distance score calculation formula between the sample to be identified and the cluster center of each subtype is:
[0145]
[0146] Where Score(i) is the distance score between the sample to be tested and the center of the ith subtype cluster; W j is the weight coefficient of the jth surface marker; d ij is the standardized distance difference between the sample to be tested and the i-th subtype cluster center on the j-th marker; n is the feature dimension.
[0147] Normalized distance difference d ij The calculation formula is:
[0148]
[0149] Where x j is the expression value of the sample to be tested on the jth marker; μ ij is the coordinate of the i-th subtype cluster center on the j-th marker; σ ij is the standard deviation of the i-th subtype on the j-th marker.
[0150] In step S11, the classification reliability index calculation formula is:
[0151]
[0152] Where R is the classification reliability index, ranging from 0 to 1, and the smaller the value, the more reliable the classification result; d min is the distance score between the sample and the nearest subtype cluster center; d sec is the distance score between the sample and the next closest subtype cluster center.
[0153] The Mahalanobis distance P-value calculation for outlier assessment is based on the chi-square distribution:
[0154]
[0155] Where P is the probability that the sample is an outlier; is the cumulative distribution function of the chi-square distribution with n degrees of freedom; D 2 is the square of the Mahalanobis distance from the sample to the cluster center; n is the feature dimension.
[0156] For error compensation in practical applications, a noise correction term is introduced:
[0157] Score'(i)=Score(i)×(1+ε×β i );
[0158] Where Score′(i) is the corrected distance score; Score(i) is the original distance score; ε is random noise, which obeys the normal distribution with mean 0 and standard deviation 0.05; β i is the noise sensitivity coefficient of the ith subtype, ranging from 0.5 to 1.5, and is determined by cross-validation.
[0159] Specifically, the principle of the present invention is: by integrating the expression patterns of multiple surface markers and considering their interactions, the present invention has established a new method for identifying breast cancer subtypes. The core principles of the present invention are as follows:
[0160] First, exosomes, as nanoscale membrane vesicles secreted by cells, inherit the molecular characteristics of their source cells through their surface marker composition. Exosomes secreted by different molecular subtypes of breast cancer cells carry specific surface protein profiles. For example, luminal exosomes are enriched in estrogen receptor-related proteins, HER2-positive exosomes highly express human epidermal growth factor receptor 2, and triple-negative exosomes express specific cancer stem cell markers. This molecular subtype specificity provides a theoretical basis for the use of exosomal markers to distinguish breast cancer subtypes.
[0161] Secondly, this method breaks through the limitations of traditional single-marker analysis by employing multicolor fluorescent labeling combined with flow cytometry to construct a multidimensional fluorescence intensity matrix. This not only considers the absolute expression levels of each marker, but also analyzes the interaction strength and spatial conformational characteristics between markers using the fluorescence resonance energy transfer equation. This systematic analysis method can capture the complex patterns of protein expression on the exosome surface and provide more comprehensive information on the characteristics of molecular subtypes.
[0162] Furthermore, the present invention incorporates multiple machine learning algorithms to optimize classification. The K-means clustering algorithm automatically identifies fluorescence signature patterns of different subtypes; the Mahalanobis distance considers inter-feature correlations to more accurately quantify differences between subtypes; the silhouette coefficient and cohesion assessment ensure clustering quality; the entropy weight optimization function objectively determines the contribution of each marker in subtype identification; and the support vector machine constructs a high-dimensional classification model, ultimately achieving accurate subtype determination.
[0163] Finally, the present invention calculates the distance scores between the test sample and the cluster centers of each subtype, determines the subtype based on the minimum distance principle, and introduces a classification reliability index to assess the accuracy of the results, providing a reference for clinical decision-making. The entire technical solution follows a complete logical chain from sample collection, data acquisition, model construction, to result verification, and effectively solves the technical problem of accurately identifying breast cancer subtypes using exosomes from blood samples.
[0164] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.
[0165] The specific implementation method of step S01 is to first collect 5-10 ml of peripheral venous blood from patients with known molecular subtypes of breast cancer and patients with subtypes to be identified, use blood collection tubes containing EDTA anticoagulant, and store at 4°C for no more than 2 hours. After collecting the blood, exosomes are extracted by differential centrifugation. First, centrifuge at 300×g for 10 minutes to remove cellular components. After taking the supernatant, centrifuge at 2000×g for 20 minutes to remove cell debris, and then centrifuge at 10000×g for 30 minutes to remove large particles. Finally, use an ultracentrifuge to precipitate the exosomes at 100,000×g for 70 minutes. The precipitate is resuspended with sterile phosphate buffer and washed again by ultracentrifugation at 100,000×g for 60 minutes. The exosome precipitate finally obtained is resuspended with 200μl sterile phosphate buffer for use. Nanoparticle tracking analysis technology is used to verify that the size distribution of exosomes is in the range of 30-150nm, and the concentration reaches 10 10 This step utilizes the differences in size and density between exosomes and other components in the blood to achieve high-purity separation of exosomes through multiple differential centrifugation steps. It is a fundamental step in establishing a method for identifying breast cancer subtypes using exosome surface markers.
[0166] In step S02, the extracted exosome sample is divided into several equal portions. Each portion is then labeled with a different fluorophore-labeled antibody. These antibodies include common exosome markers such as anti-CD9-FITC, anti-CD63-PE, and anti-CD81-APC, as well as surface proteins associated with breast cancer subtypes such as anti-human epidermal growth factor receptor 2-Cy5, anti-estrogen receptor-Cy3, anti-progesterone receptor-Cy7, anti-Ki-67-PerCP, anti-epithelial cell adhesion molecule-AF647, anti-CD44-PE-Cy7, and anti-matrix metalloproteinase-BV421. The antibody dilution ratio is 1:100 to 1:500. The samples are incubated at 4°C for 2 hours, gently mixing every 30 minutes. After incubation, the samples are washed three times by centrifugation at 10,000 × g for 10 minutes to remove unbound antibodies. The washed samples are resuspended in 100 μl of sterile phosphate buffered saline to ensure sufficient binding of the fluorescently labeled antibodies to the exosome surface markers. This step uses the principle of antigen-antibody specific binding to label exosome surface proteins and uses fluorescent groups with different excitation and emission wavelengths to achieve multicolor fluorescent labeling, laying the foundation for subsequent flow cytometry detection and multidimensional feature analysis.
[0167] The specific implementation of step S03 is to use a high-resolution flow cytometer to measure the fluorescence intensity of each surface marker in each exosome sample, and construct a multidimensional fluorescence intensity matrix based on the multi-color fluorescent labeling in step S02. The flow cytometer excitation light source is set to three wavelengths of 488nm, 561nm and 640nm, corresponding to the detection of fluorescence signals emitted by different fluorescent groups. The forward scatter detector threshold is set to 200 to eliminate background noise, the acquisition rate is controlled at 1000-2000 events / second, and at least 50,000 valid events are collected for each sample. The multidimensional fluorescence intensity matrix is expressed as:
[0168]
[0169] Where M is the multidimensional fluorescence intensity matrix; I ij represents the median fluorescence intensity of the jth surface marker in the i-th sample; m is the number of samples; and n is the number of surface marker types. The acquired raw data were processed using the fluorescence compensation matrix in the flow cytometry software to eliminate spectral overlap between different fluorophores. To eliminate batch effects and instrument errors, the Z-score method was used for normalization:
[0170]
[0171] Where Z ij is the normalized fluorescence intensity value; I ij is the original fluorescence intensity value; μ j is the average fluorescence intensity of the jth surface marker in all samples; σj is the standard deviation of the jth surface marker across all samples. This step quantitatively detects exosome surface markers using flow cytometry, constructing a multidimensional data structure reflecting the surface protein expression characteristics of each sample, laying the data foundation for subsequent analysis.
[0172] The specific implementation of step S04 is to calculate the interaction strength between different markers on the surface of exosomes based on flow cytometry data, and analyze the conformational characteristics and spatial arrangement patterns of exosome surface proteins of different breast cancer subtypes. The calculation is performed using the fluorescence resonance energy transfer equation:
[0173]
[0174] Where E is the energy transfer efficiency, ranging from 0 to 1; R0 is the distance at which 50% energy transfer occurs (Förster distance), in nanometers; r is the actual distance between the fluorescent donor and acceptor molecules, in nanometers. The Förster distance R0 is calculated using the formula:
[0175] R0=0.211×[κ 2 ×n -4 ×Q D ×J(λ)] 1 / 6 ;
[0176] Where κ is the orientation factor, which describes the spatial orientation relationship between the donor emission dipole and the acceptor absorption dipole. Its theoretical value range is 0 to 4, and it is usually 2 / 3 under random orientation conditions. n is the refractive index of the medium, which is about 1.35 in plasma. Q D is the quantum yield of the fluorescence donor, ranging from 0 to 1; J(λ) is the overlap integral of the donor emission spectrum and the acceptor absorption spectrum, with the unit of M -1 cm 3 , calculated by the formula:
[0177] J(λ)=∫F D (λ)×ε A (λ)×λ 4 ×dλ;
[0178] Where, F D (λ) is the normalized donor emission spectrum; ε A (λ) is the molar absorption coefficient of the receptor, in M -1 cm -1 ; λ is the wavelength, in nanometers. Considering the spatial mobility of exosome surface proteins, the modified energy transfer efficiency calculation formula is introduced:
[0179] E actual =E×(1-α×D);
[0180] Where, Eactual is the actual energy transfer efficiency; E is the theoretical energy transfer efficiency; α is the mobility correction coefficient, ranging from 0 to 0.5; D is the diffusion coefficient of exosome surface proteins, in μm 2 / s. When the energy transfer efficiency E actual When the value is greater than 0.1, it is considered that the two markers have an effective interaction. This step uses the principle of fluorescence resonance energy transfer to analyze the spatial relationship and interaction between exosome surface markers, revealing the molecular arrangement characteristics of exosome surface proteins in different breast cancer subtypes, and providing conformational information for subtype identification.
[0181] The specific implementation of step S05 is to apply the K-means clustering algorithm to analyze the multidimensional fluorescence intensity matrix of the exosome samples of patients with known subtypes to determine the fluorescence characteristic patterns of different breast cancer subtypes. The objective function of the K-means clustering algorithm is:
[0182]
[0183] Where J is the objective function value that needs to be minimized; k is the number of clusters; C i is the i-th cluster; x is the sample point (multidimensional fluorescence intensity vector); μ i is the center point of the i-th cluster; ||x-μ i || represents the distance from the sample point x to the cluster center μ i The Euclidean distance is calculated as follows:
[0184]
[0185] Where x j is the coordinate of the sample point x in the jth dimension (fluorescence intensity value of the jth surface marker); μ ij is the coordinate of the i-th cluster center in the j-th dimension; n is the number of surface marker types. The cluster center update formula is:
[0186]
[0187] In the formula, |C i | is the number of samples in the i-th cluster. Specifically, we first determine the optimal number of clusters, K. We then use the elbow rule and silhouette coefficient to evaluate the clustering effects of different K values (2 to 6). The parameter with the highest silhouette coefficient and the smallest K value is selected as the final number of clusters. We then use the K-means++ algorithm to optimize the selection of the initial center points. The probability of selecting the initial cluster center is calculated as:
[0188]
[0189] Where P(x) is the probability of selecting sample point x as the next cluster center; D(x) is the distance from sample point x to the nearest selected cluster center; and X is the set of all sample points. The K-means iteration process is performed until the change in the position of the cluster center is less than a preset threshold (typically set to 0.001) or the maximum number of iterations (typically 100) is reached. This step uses the K-means clustering algorithm to automatically identify and separate the expression patterns of exosome surface markers in different breast cancer subtypes, providing a clustering basis for subsequent differentiation of different subtypes.
[0190] The specific implementation of step S06 is to calculate the Mahalanobis distance between the cluster centers of different subtypes to quantitatively evaluate the difference in the expression of surface markers of exosomes of each subtype. The Mahalanobis distance calculation formula is:
[0191]
[0192] Where d(x, y) is the Mahalanobis distance between the two cluster centers x and y; x and y are the cluster center vectors of two different subtypes; S is the covariance matrix of all samples; S -1 is the inverse matrix of the covariance matrix. The calculation formula of the covariance matrix S is:
[0193]
[0194] Where m is the total number of samples; x i is the i-th sample point; is the mean vector of all samples. Considering that the covariance matrix may not be invertible, singular value decomposition plus regularization is used:
[0195] S′=S+λI;
[0196] Where S′ is the regularized covariance matrix; λ is the regularization parameter, typically ranging from 0.01 to 0.1; and I is the identity matrix. First, the covariance matrix S of all samples is calculated. Then, the Mahalanobis distance between any two subtype cluster centers is calculated to construct the inter-subtype distance matrix. A threshold of 3.0 is set. When the Mahalanobis distance between two subtypes exceeds this threshold, the two subtypes are considered to be significantly distinguishable. This step uses the Mahalanobis distance to account for inter-feature correlations and is more suitable than the Euclidean distance for assessing the degree of difference between subtypes in a multidimensional feature space, providing a quantitative indicator for subsequent marker combination screening.
[0197] The specific implementation of step S07 is to measure the silhouette coefficient and cohesion of each known subtype cluster and screen the surface marker combination with the highest discrimination. The silhouette coefficient calculation formula is:
[0198]
[0199] Where s(i) is the silhouette coefficient of sample i, ranging from -1 to 1. The closer it is to 1, the better the clustering effect. a(i) is the average distance between sample i and other samples in the same cluster. b(i) is the average distance between sample i and the closest samples in the other cluster. The calculation formulas for a(i) and b(i) are:
[0200]
[0201] Where C i is the cluster to which sample i belongs; |C i | is the number of samples in the cluster; d(i, j) is the distance between sample i and sample j; C k is the other clusters to which i does not belong. The formula for calculating cohesion is:
[0202]
[0203] Where Cohesion(C i ) is the cohesion of the i-th cluster; |C i | is the number of samples in the i-th cluster; x is the number of clusters C i The sample points in μ i is the center point of the i-th cluster; ||x-μ i || is the distance from the sample point x to the cluster center μ i The overall clustering quality evaluation function is:
[0204]
[0205] Where Q is the clustering quality score, with higher values indicating better clustering; m is the total number of samples; k is the number of clusters; and α and β are weighting coefficients, typically set to α = 0.7 and β = 0.3. First, the cluster silhouette coefficient and cohesion of all surface marker combinations are calculated. Then, recursive feature elimination is used to remove each surface marker one at a time, and K-means clustering is re-performed to recalculate the new silhouette coefficient and cohesion. When the silhouette coefficient exceeds 0.7 and the cohesion is less than 30% of the average distance between cluster centers, the marker combination is considered to have good discriminatory power. This step, by evaluating clustering quality, selects the optimal surface marker combination that achieves the highest subtype discrimination with the minimum number of features.
[0206] The specific implementation of step S08 is to use the entropy weight optimization function to quantitatively evaluate the contribution of surface markers and determine the weight coefficient of each marker in subtype identification. First, calculate the coefficient of variation:
[0207]
[0208] Where, CV jis the coefficient of variation of the jth surface marker; σ j is the standard deviation of the marker among each subtype; μ j is the average expression value of the marker. Then calculate the information entropy:
[0209]
[0210] Where, e j is the information entropy of the jth surface marker, ranging from 0 to 1; k is the number of breast cancer subtypes; p ij is the normalized expression value of the j-th marker in the i-th subtype, calculated as: Calculate the difference coefficient based on information entropy:
[0211] d j =1-e j ;
[0212] Where, d j is the difference coefficient of the jth surface marker, ranging from 0 to 1, with values closer to 1 indicating higher discrimination. The correlation of marker expression is calculated as:
[0213]
[0214] The average correlation coefficient is calculated as:
[0215]
[0216] The discrimination ability index is calculated as:
[0217]
[0218] The comprehensive evaluation function is calculated as:
[0219] F j =w1×CV j +w2×(1-r j )+w3×d j +w4×A j +w5×D j ;
[0220] Where, F j is the comprehensive score of the jth surface marker; w1~w5 are the weight coefficients of each factor, all set to 0.2; A j is the marker expression abundance. The final weight coefficient is calculated as:
[0221]
[0222] Where W j is the weight coefficient of the jth surface marker in subtype identification, satisfying This step comprehensively considers the variability, correlation, information entropy, expression abundance and discrimination ability of markers among different subtypes through the entropy weight method, objectively evaluates the contribution of each marker in subtype identification, and provides weight parameters for the construction of the support vector machine classification model.
[0223] The specific implementation of step S09 is to construct a support vector machine classification model based on the screened surface marker combination and weight coefficient for subsequent subtype discrimination of samples to be identified. The radial basis function kernel calculation formula in the support vector machine is:
[0224] K(x, y) = exp(-γ||xy|| 2 );
[0225] Where K(x, y) is the kernel function value; x and y are the feature vectors of the two samples; γ is the kernel function parameter, the optimal value of which is determined by grid search and ranges from 0.001 to 10; ||xy|| is the Euclidean distance between the feature vectors of the two samples. The weighted Euclidean distance considering the marker weight is calculated as:
[0226]
[0227] Where, ||xy|| w is the weighted Euclidean distance; W j is the weight coefficient of the jth surface marker; x j and y j are the eigenvalues of the two samples in the jth dimension; n is the number of characteristic dimensions. The optimization objective function of the weighted support vector machine is:
[0228]
[0229] Constraint: y i (w T φ(x i )+b)≥1-ξ i ,ξ i ≥0, i=1, 2, ..., m;
[0230] Where w is the normal vector of the hyperplane; b is the intercept; ξ i is a slack variable; C is a penalty parameter, ranging from 0.1 to 100; y i is the category label of sample i; φ(x i) is the feature vector of sample i after kernel function mapping; m is the number of training samples. First, the sample dataset is randomly divided into a training set (75%) and a validation set (25%). The performance of different parameter combinations is evaluated through five-fold cross-validation. The final model should achieve a classification accuracy of at least 85% on the validation set. This step utilizes the nonlinear classification capabilities of the support vector machine algorithm, combined with the marker weight coefficients determined in step S08, to construct a classification model capable of accurately distinguishing different breast cancer subtypes.
[0231] The specific implementation of step S10 is to input the multidimensional fluorescence intensity matrix of the exosome sample of the patient with the subtype to be identified into the trained support vector machine classification model and calculate the distance score between the sample and the cluster center of each subtype. First, the sample to be identified is normalized in the same way as the training set. Then, the fluorescence intensity data corresponding to the optimal surface marker combination determined in step S07 is extracted to form a feature vector. The feature vector is input into the support vector machine classification model, and the Mahalanobis distance score between the sample and the cluster center of each subtype is calculated. The distance score calculation formula is:
[0232]
[0233] Where Score(i) is the distance score between the sample to be tested and the center of the ith subtype cluster; W j is the weight coefficient of the jth surface marker; d ij is the standardized distance difference between the sample to be tested and the i-th subtype cluster center on the j-th marker, and the calculation formula is:
[0234]
[0235] Where x j is the expression value of the sample to be tested on the jth marker; μ ij is the coordinate of the i-th subtype cluster center on the j-th marker; σ ij is the standard deviation of the i-th subtype on the j-th marker. This step uses the support vector machine classification model to analyze the expression characteristics of exosome surface markers of the samples to be identified, calculates the distance score between the sample and the cluster center of each subtype, and provides a quantitative basis for the final subtype determination.
[0236] The specific implementation of step S11 is to determine the breast cancer subtype to which the patient sample to be identified belongs based on the minimum distance principle and calculate the classification reliability index. First, the distance scores of each subtype calculated in step S10 are compared, and the sample is classified as the subtype with the smallest distance score. Then, the classification reliability index is calculated:
[0237]
[0238] Where R is the classification reliability index, ranging from 0 to 1, and the smaller the value, the more reliable the classification result; d min is the distance score between the sample and the nearest subtype cluster center; d sec The distance score between the sample and the next closest subtype cluster center. Generally, when R<0.6, the classification result is considered highly reliable, 0.6≤R<0.8 is considered moderately reliable, and R≥0.8 is considered lowly reliable. The position of the sample in the distribution of the closest subtype is also calculated, and the P value of the Mahalanobis distance is used to assess whether the sample is an outlier:
[0239]
[0240] Where P is the probability that the sample is an outlier; is the cumulative distribution function of the chi-square distribution with n degrees of freedom; D 2 is the square of the Mahalanobis distance from the sample to the cluster center; n is the feature dimension. Considering the error in practical applications, the noise correction term is introduced:
[0241] Score'(i)=Score(i)×(1+ε×β i );
[0242] Where Score′(i) is the corrected distance score; Score(i) is the original distance score; ε is random noise, which obeys the normal distribution with mean 0 and standard deviation 0.05; β i is the noise sensitivity coefficient for the i-th subtype, ranging from 0.5 to 1.5. A subtype identification report is generated, including the sample subtype determination result, classification reliability index, and key surface marker expression levels, providing an accurate basis for clinical diagnosis. This step evaluates the reliability of the classification results by calculating the classification reliability index and outlier probability, thereby improving the accuracy of subtype identification and its clinical reference value.
[0243] To better understand and implement the present invention, Example 2, a specific application scenario of the present invention, is provided below: Researchers collected peripheral venous blood samples from 120 patients diagnosed with breast cancer at a medical center, including 30 patients with Luminal A, 30 with Luminal B, 30 with HER2-overexpressing breast cancer, and 30 with triple-negative breast cancer. Blood samples were also collected from 30 healthy women as controls. All blood samples were collected in EDTA-anticoagulant tubes and stored at 4°C for no more than 2 hours. Cellular components were removed by differential centrifugation at 300×g for 10 minutes. The supernatant was then centrifuged at 2000×g for 20 minutes to remove cellular debris and then at 10,000×g for 30 minutes to remove large particles. Finally, exosomes were pelleted using an ultracentrifuge at 100,000×g for 70 minutes. The pellet was resuspended in sterile PBS and washed again by ultracentrifugation at 100,000×g for 60 minutes. The resulting exosome pellet was resuspended in 200μl of PBS for later use.
[0244] The extracted exosomes were identified using nanoparticle tracking analysis technology. The results showed that the size of exosomes was mainly distributed in the range of 30 to 150 nm, and the concentration reached 5.3×10 10 The extracted exosome samples were added with antibodies labeled with different fluorescent groups, including anti-CD9-FITC, anti-CD63-PE, anti-CD81-APC, and antibodies against surface proteins related to breast cancer subtypes, and incubated at 4°C for 2 hours. The average fluorescence intensity of the main exosome markers in each type of sample is shown in Table 1:
[0245] Table 1 Mean fluorescence intensity (MFI) of major exosome markers in various types of samples
[0246] Sample type CD9-FITC CD63-PE CD81-APC Luminal A 8654 7892 9123 Luminal B 8976 8234 9356 HER2 overexpression 8432 7756 8975 triple negative 7986 7432 8543 Healthy controls 7654 6897 8132
[0247] Flow cytometry was used to measure the fluorescence intensity of each surface marker in each exosome sample and construct a multidimensional fluorescence intensity matrix. The fluorescence intensity of breast cancer-related markers on the surface of exosomes of different breast cancer subtypes is shown in Table 2:
[0248] Table 2 Fluorescence intensity (MFI) of breast cancer-related markers on the surface of exosomes of different breast cancer subtypes
[0249]
[0250] Based on flow cytometry data, the interaction strength between different markers on the exosome surface was calculated. Using the fluorescence resonance energy transfer equation, the energy transfer efficiency of each group of fluorescent donor-acceptor pairs was obtained as shown in Table 3:
[0251] Table 3 Energy transfer efficiency of each group of fluorescent donor-acceptor pairs
[0252]
[0253] The K-means clustering algorithm was applied to the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes. Using the elbow rule and silhouette coefficient evaluation, the optimal number of clusters, K = 4, was determined, corresponding to the four breast cancer subtypes. The silhouette coefficient and cohesion of the K-means clustering results are shown in Table 4:
[0254] Table 4 Silhouette coefficient and cohesion of K-means clustering results
[0255] Evaluation indicators Luminal A Luminal B HER2 overexpression triple negative Silhouette coefficient 0.78 0.75 0.82 0.81 Cohesion 0.32 0.37 0.29 0.28
[0256] The Mahalanobis distance between the cluster centers of different subtypes was calculated to quantitatively evaluate the differences in the expression of exosome surface markers of each subtype. The results are shown in Table 5:
[0257] Table 5 Mahalanobis distances between cluster centers of different subtypes
[0258] Subtype comparison Mahalanobis distance Luminal A vs Luminal B 3.56 Luminal A vs HER2 overexpression 6.87 Luminal A vs triple negative 8.32 Luminal B vs HER2 overexpression 5.43 Luminal B vs triple negative 7.21 HER2 overexpression vs triple negative 6.98
[0259] The surface marker combinations with the highest discrimination were screened by recursive feature elimination combined with silhouette coefficient and cohesion evaluation. The clustering quality scores of different marker combinations are shown in Table 6:
[0260] Table 6 Clustering quality scores of different marker combinations
[0261] Marker combination Clustering quality score All 7 markers 0.75 HER2, ER, PR, Ki-67, CD44, MMP 0.79 HER2,ER,PR,Ki-67,CD44 0.82 HER2, ER, PR, CD44 0.78 HER2, ER, CD44 0.71
[0262] The entropy weight method was applied to optimize the function to quantitatively evaluate the contribution of surface markers and determine the weight coefficient of each marker in subtype identification. The results are shown in Table 7:
[0263] Table 7 Weight coefficients of surface markers in subtype identification
[0264]
[0265] A support vector machine classification model was constructed based on the combination of the five selected markers (HER2, ER, PR, Ki-67, and CD44) and their weight coefficients. Through five-fold cross-validation, the optimal parameters were determined to be γ = 0.05 and C = 10. Ninety breast cancer patient samples were randomly selected as the training set, and 30 patient samples were used as the validation set. The model performance on the validation set is shown in Table 8:
[0266] Table 8 Performance of the support vector machine classification model on the validation set
[0267] Performance indicators Luminal A Luminal B HER2 overexpression triple negative average value Accuracy 0.91 0.87 0.92 0.89 0.90 Sensitivity 0.88 0.84 0.90 0.85 0.87 Specificity 0.93 0.89 0.95 0.92 0.92 F1 score 0.90 0.86 0.92 0.88 0.89
[0268] The constructed model was used to identify subtypes of 20 newly diagnosed breast cancer patients with undefined molecular subtypes. The multidimensional fluorescence intensity matrix of the exosome samples from these patients was input into the trained support vector machine classification model, and the distance scores between the matrix and the cluster center of each subtype were calculated. The results were compared with the immunohistochemistry results, as shown in Table 9:
[0269] Table 9 Comparison of exosome model prediction results and immunohistochemistry results
[0270]
[0271]
[0272] Traditional molecular subtyping of breast cancer relies primarily on immunohistochemistry or genetic testing based on tissue biopsies. These methods are not only highly invasive but also fail to reflect tumor heterogeneity and dynamic changes. The proposed method for identifying breast cancer subtypes using exosome surface markers requires only peripheral venous blood collection from the patient. By analyzing the expression patterns and interaction characteristics of tumor-derived exosome surface markers in the blood, this method enables non-invasive identification of breast cancer subtypes. Compared with traditional methods, this method offers the following advantages: First, non-invasive sampling significantly reduces patient burden and is suitable for repeated monitoring and dynamic assessment. Second, exosomes, as information carriers secreted by tumor cells, can more comprehensively reflect the heterogeneity of the primary tumor. Third, the analysis method, combining multidimensional fluorescence signatures with fluorescence resonance energy transfer, considers not only the expression levels of surface markers but also the spatial arrangement and interactions between proteins, providing richer molecular information. Fourth, a machine learning-based classification model and weight optimization strategy achieve high-precision subtype identification, with a prediction accuracy of 90%, approaching that of traditional gold standard methods. In actual validation, the method achieved a 95% concordance rate between its predictions and immunohistochemistry results for 20 newly diagnosed patients, demonstrating its clinical value. Furthermore, the method also provides a classification reliability index, providing a reference for clinical decision-making.
[0273] It should be noted that the variables involved in the present invention are explained in detail as shown in Tables 10 and 11 below.
[0274] Table 10 Variable Explanation Table (Part 1)
[0275]
[0276]
[0277] Table 11 Variable Explanation Table (Part 2)
[0278]
[0279] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A method for identifying breast cancer subtypes using exosome surface markers, characterized in that: include: Collect blood samples from breast cancer patients, isolate and extract exosomes; perform multi-color fluorescent labeling on exosome surface markers; Flow cytometry was used to measure the fluorescence intensity of surface markers and construct a multidimensional fluorescence intensity matrix. The fluorescence resonance energy transfer equation was used to calculate the interaction strength of surface markers. K-means cluster analysis was performed on samples with known subtypes to determine the fluorescence characteristic patterns; Calculate the Mahalanobis distance between cluster centers; Determine the cluster silhouette coefficient and cohesion to screen the marker combination with the highest discrimination; apply the entropy weight method to optimize the function to evaluate the marker contribution; A support vector machine classification model was constructed; samples to be identified were input into the model, and the distance scores to the cluster centers of each subtype were calculated; breast cancer subtypes were determined based on the minimum distance principle, and the classification reliability index was calculated.
2. The method according to claim 1, characterized in that The step of performing multi-color fluorescent labeling on the surface markers of exosomes is to perform multi-color fluorescent labeling on the surface markers of exosomes using fluorescent labeled antibodies, including CD9, CD63, CD81 and surface proteins related to breast cancer subtypes.
3. The method according to claim 2, characterized in that The breast cancer subtype-related surface proteins include human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, and matrix metalloproteinase.
4. The method according to claim 3, characterized in that The multidimensional fluorescence intensity matrix refers to a data structure composed of the fluorescence signal intensities of multiple surface markers in each exosome sample, where each row represents a sample and each column represents the fluorescence intensity value of a surface marker.
5. The method according to claim 4, characterized in that The silhouette coefficient is an indicator to measure the quality of clustering. It calculates the normalized value of the difference between the average similarity of each sample with other samples in the same cluster and the average similarity of samples in other clusters. The closer the value is to 1, the better the clustering effect. The cohesion refers to the average distance from all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification.
6. The method according to claim 5, characterized in that The K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence characteristic patterns. It minimizes the sum of the squared distances from sample points to their cluster centers through iterative optimization. The input is a multidimensional fluorescence intensity matrix, and the output is the fluorescence characteristic patterns of different breast cancer subtypes. The Mahalanobis distance is a distance metric method that takes into account the correlation between features. The input is the cluster centers of different subtypes, and the output is a quantitative indicator of the difference in exosome surface marker expression between subtypes.
7. The method according to claim 6, characterized in that The fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the surface of exosomes. The input includes the fluorescence donor quantum yield, the distance between fluorescence donor and acceptor molecules, the fluorescence donor and acceptor spectral overlap integral, the refractive index, and the fluorescence resonance energy transfer direction factor. The output is the interaction strength between surface markers.
8. The method according to claim 7, characterized in that The distance score is the Mahalanobis distance between the sample to be identified and the cluster center of each subtype, which is calculated by the support vector machine classification model and is used for breast cancer subtype determination.
9. The method according to claim 8, characterized in that The classification reliability index is obtained by calculating the ratio of the distance score between the sample to be tested and the nearest subtype cluster center to the distance score between the sample to be tested and the nearest subtype cluster center. The smaller the value, the more reliable the classification result.
10. The method according to claim 9, characterized in that The step of constructing a support vector machine classification model is to construct a support vector machine classification model based on the screened surface marker combination and weight coefficient, and use known subtype samples to train the model.
Citation Information
Patent Citations
Clustering method based on network for disease subtype problem
CN105160208A
Method and system for identifying rare macrophage subgroups and disease markers in idiopathic pulmonary fibrosis
CN114708918A
Breast cancer subtype classification method and system based on graph convolutional neural network
CN116959572A
Method for determination and identification of cell signatures and cell markers
US20180127823A1
Methods for subtyping of lung adenocarcinoma
US20190338365A1