A method for identifying breast cancer subtypes using exosome surface markers

By measuring the fluorescence intensity of exosome surface markers using multicolor fluorescent labeling and flow cytometry, and combining fluorescence resonance energy transfer and cluster analysis, a support vector machine model was constructed. This solved the accuracy problem of identifying breast cancer subtypes using exosome surface markers, and achieved high-precision identification of breast cancer subtypes.

CN120445960BActive Publication Date: 2026-01-27THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510580047.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2026-01-27
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Current technologies struggle to accurately identify molecular subtypes of breast cancer using exosome surface markers in blood samples, resulting in insufficient accuracy in non-invasive diagnosis and personalized treatment.

Method used

By collecting blood samples from breast cancer patients, exosomes were isolated and extracted. The fluorescence intensity of surface markers was measured using multicolor fluorescent labeling and flow cytometry to construct a multidimensional fluorescence intensity matrix. The interaction between markers was analyzed using fluorescence resonance energy transfer technology. K-means clustering analysis and support vector machine classification models were applied to screen the combination of markers with the highest discriminative power and construct a breast cancer subtype classification model.

Benefits of technology

This technology enables precise identification of breast cancer subtypes based on the multidimensional characteristics of exosome surface markers, improving the accuracy of non-invasive diagnosis and personalized treatment, and providing important support for clinical diagnosis and prognostic assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120445960B_ABST
    Figure CN120445960B_ABST
Patent Text Reader

Abstract

The application provides a method for identifying breast cancer subtypes by using exosome surface markers, and belongs to the technical field of exosome surface marker detection. The method uses multi-color fluorescent markers to target CD9, CD63, CD81 and breast cancer subtype related surface proteins, constructs a multi-dimensional fluorescence intensity matrix by using flow cytometry, analyzes the interaction strength between markers by using fluorescence resonance energy transfer technology, determines the fluorescence characteristic mode of different subtypes by using K-means clustering, quantifies the differences between subtypes by calculating Mahalanobis distance, screens the optimal marker combination by evaluating the contour coefficient and cohesion degree, determines the marker weight coefficient by using entropy weight method, constructs a support vector machine classification model, and finally accurately determines the breast cancer subtype based on the minimum distance principle by calculating the distance score between the to-be-detected sample and the cluster center of each subtype, thereby providing an important basis for individualized diagnosis and treatment in the clinic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of exosome surface marker detection technology, and more specifically, relates to a method for identifying breast cancer subtypes using exosome surface markers. Background Technology

[0002] Breast cancer is one of the most common malignant tumors in women. Based on molecular phenotype, it can be classified into several subtypes, including Luminal A, Luminal B, HER2-positive, and triple-negative, with significant differences in treatment plans and prognosis among the different subtypes. Traditional breast cancer subtype identification mainly relies on tissue biopsy, using immunohistochemistry or gene expression analysis to determine the molecular subtype. However, tissue biopsy is invasive and cannot provide real-time monitoring, limiting its effectiveness for early screening and dynamic assessment of treatment efficacy.

[0003] In recent years, circulating tumor exosomes have received widespread attention as an important component of "liquid biopsy". Existing exosome detection technologies mainly focus on the analysis of total exosome count or single biomarker expression levels, such as ultracentrifugation combined with Western blotting or ELISA to detect specific proteins. However, these methods are difficult to comprehensively capture the systematic differences in the expression patterns of exosome surface biomarkers.

[0004] Current technologies lack an analytical method that can comprehensively consider the interactions and differences in expression patterns among multiple markers on the surface of exosomes. This results in insufficient accuracy in identifying molecular subtypes of breast cancer through blood exosomes, making it difficult to meet the needs of non-invasive clinical diagnosis and personalized treatment. In other words, existing technologies suffer from the technical problem of accurately identifying molecular subtypes of breast cancer based on the characteristics of exosome surface markers in blood samples. Summary of the Invention

[0005] In view of this, the present invention provides a method for identifying breast cancer subtypes using exosome surface markers, which can solve the technical problem in the prior art that it is difficult to accurately identify the molecular subtypes of breast cancer through the characteristics of exosome surface markers in blood samples.

[0006] This invention is implemented as follows: A method for identifying breast cancer subtypes using exosome surface markers includes: collecting blood samples from breast cancer patients and separating and extracting exosomes; labeling exosome surface markers with multicolor fluorescence; measuring the fluorescence intensity of surface markers using flow cytometry and constructing a multidimensional fluorescence intensity matrix; calculating the interaction intensity of surface markers using the fluorescence resonance energy transfer equation; performing K-means clustering analysis on known subtype samples to determine fluorescence characteristic patterns; calculating the Mahalanobis distance between cluster centers; determining the cluster profile coefficient and cohesion, and screening the marker combination with the highest discriminative power; applying the entropy weight method to optimize the function and evaluate the marker contribution; constructing a support vector machine classification model; inputting the sample to be identified into the model and calculating the distance score to each subtype cluster center; determining the breast cancer subtype according to the minimum distance principle and calculating the classification reliability index.

[0007] The step of multicolor fluorescent labeling of exosome surface markers involves using fluorescently labeled antibodies to perform multicolor fluorescent labeling of exosome surface markers, including CD9, CD63, CD81, and breast cancer subtype-related surface proteins.

[0008] The breast cancer subtype-related surface proteins include human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, and matrix metalloproteinase.

[0009] The multidimensional fluorescence intensity matrix refers to a data structure composed of the fluorescence signal intensity of multiple surface markers in each exosome sample, where each row represents a sample and each column represents the fluorescence intensity value of a surface marker.

[0010] The silhouette coefficient is an indicator of clustering quality. It is a standardized value calculated as the difference between the average similarity of each sample with other samples in the same cluster and the average similarity of samples in other clusters. The closer the value is to 1, the better the clustering effect. The cohesion refers to the average distance of all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification.

[0011] The K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence feature patterns. It iteratively optimizes and minimizes the sum of squared distances from sample points to their respective cluster centers. The input is a multidimensional fluorescence intensity matrix, and the output is the fluorescence feature patterns of different breast cancer subtypes. The Mahalanobis distance is a distance metric that considers the correlation between features. The input is the cluster centers of different subtypes, and the output is a quantitative index of the difference in expression of exosome surface markers between subtypes.

[0012] The fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the surface of exosomes. The inputs include the quantum yield of the fluorescent donor, the distance between the fluorescent donor and acceptor molecules, the spectral overlap integral of the fluorescent donor and acceptor, the refractive index, and the fluorescence resonance energy transfer direction factor. The output is the interaction strength between surface markers.

[0013] The distance score is the Mahalanobis distance between the sample to be identified and the cluster centers of each subtype, calculated by the support vector machine classification model, and used for breast cancer subtype determination.

[0014] The classification reliability index is calculated by dividing the distance score between the sample and the nearest subtype cluster center by the distance score between the sample and the second nearest subtype cluster center. The smaller the value, the more reliable the classification result.

[0015] The step of constructing the support vector machine classification model involves building the support vector machine classification model based on the selected combination of surface markers and weight coefficients, and training the model using known subtype samples.

[0016] This invention establishes a high-precision breast cancer subtype classification model by constructing a multicolor fluorescent labeling system and measuring it by flow cytometry to obtain the multidimensional fluorescence intensity matrix of exosome surface markers, and by combining fluorescence resonance energy transfer technology to analyze the interaction between markers.

[0017] This method automatically identifies fluorescence feature patterns of different breast cancer subtypes using the K-means clustering algorithm, quantifies differences between subtypes using Mahalanobis distance, selects the optimal biomarker combination by combining silhouette coefficient and cohesion index, and optimizes the assignment of reasonable weight coefficients to each biomarker using the entropy weight method. Finally, a support vector machine classification model is constructed to achieve accurate subtype identification. This systematic analysis method significantly improves the accuracy of exosome biomarkers in identifying breast cancer subtypes.

[0018] This invention solves the technical problem that traditional liquid biopsy is difficult to accurately identify molecular subtypes of breast cancer, and realizes the accurate determination of breast cancer subtypes through the multidimensional characteristics of exosome surface markers, providing important technical support for non-invasive diagnosis, prognostic assessment and personalized treatment. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0021] like Figure 1The diagram shown is a flowchart of a method for identifying breast cancer subtypes using exosome surface markers provided by the present invention. This method includes the following steps:

[0022] S01. Collect blood samples from patients with known subtypes of breast cancer and blood samples from patients with subtypes to be identified, and extract exosomes by ultracentrifugation.

[0023] S02. Multicolor fluorescent labeling of exosome surface markers using fluorescently labeled antibodies, including CD9, CD63, CD81, and breast cancer subtype-related surface proteins;

[0024] S03. Flow cytometry was used to measure the fluorescence intensity of all surface markers in each exosome sample, and a multidimensional fluorescence intensity matrix was constructed for each sample.

[0025] S04. Calculate the interaction strength between exosome surface markers using the fluorescence resonance energy transfer equation, and analyze the conformational characteristics of proteins on the surface of different subtypes of exosomes.

[0026] S05. Perform K-means clustering analysis on the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes to determine the fluorescence characteristic patterns of different breast cancer subtypes;

[0027] S06. Calculate the Mahalanobis distance between cluster centers of different subtypes and establish a quantitative index for the differential expression of exosome surface markers between subtypes;

[0028] S07. Measure the silhouette coefficient and cohesion of each known subtype cluster and screen the combination of surface markers with the highest discrimination.

[0029] S08. Apply the entropy weight method to optimize the function and quantitatively evaluate the contribution of surface markers, and determine the weight coefficient of each marker in subtype identification.

[0030] S09. Construct a support vector machine classification model based on the selected surface marker combinations and weight coefficients, and train the model using known subtype samples.

[0031] S10. Input the multidimensional fluorescence intensity matrix of the exosome sample of the patient to be identified into the trained support vector machine classification model, and calculate the distance score between it and the cluster center of each subtype.

[0032] S11. Determine the breast cancer subtype of the patient sample to be identified based on the minimum distance principle, and calculate the classification reliability index.

[0033] The multidimensional fluorescence intensity matrix refers to the data structure composed of the fluorescence signal intensity of multiple surface markers in each exosome sample. Each row represents a sample, and each column represents the fluorescence intensity value of a surface marker. It is obtained by flow cytometry measurement in step S03 and is used for K-means clustering algorithm analysis in step S05 and support vector machine classification model input in step S10.

[0034] The fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the exosome surface. The inputs include the fluorescence donor quantum yield measured in step S02, the fluorescence donor-acceptor intermolecular distance measured in step S03, the fluorescence donor-acceptor spectral overlap integral obtained in step S02, the refractive index measured in step S02, and the fluorescence resonance energy transfer direction factor calculated in step S03. The output is the interaction strength between surface markers, which is used for screening surface marker combinations in step S07.

[0035] Among them, the K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence feature patterns. It minimizes the sum of squared distances from sample points to their respective cluster centers through iterative optimization. The input is the multidimensional fluorescence intensity matrix obtained in step S03, and the output is the fluorescence feature patterns of different breast cancer subtypes, which are used for subtype expression difference analysis in step S06.

[0036] Mahalanobis distance is a distance metric that considers the correlation between features. Compared with Euclidean distance, it is more suitable for measuring the degree of difference between subtype cluster centers in a multidimensional feature space. The input is the different subtype cluster centers obtained in step S05, and the output is a quantitative index of the difference in expression of exosome surface markers between subtypes, which is used for discrimination evaluation in step S07.

[0037] The silhouette coefficient is an indicator of clustering quality. It is a standardized value that calculates the difference between the average similarity of each sample and other samples in the same cluster and the average similarity of samples in other clusters. The closer the value is to 1, the better the clustering effect. The input is the clustering result obtained in step S05 and the multidimensional fluorescence intensity matrix obtained in step S03. The output is used for surface marker combination screening in step S07.

[0038] Wherein, cohesion refers to the average distance from all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification. The input is the clustering result obtained in step S05 and the multidimensional fluorescence intensity matrix obtained in step S03. The output is used for surface marker combination screening in step S07.

[0039] The entropy weight optimization function is used to objectively evaluate the contribution of different surface biomarkers in the identification of breast cancer subtypes. The inputs include the coefficient of variation of each biomarker among different subtypes calculated in step S06, the biomarker expression correlation matrix obtained in step S04, the biomarker information entropy value calculated in step S05, the biomarker expression abundance measured in step S03, and the biomarker discrimination index in known subtype samples obtained in step S07. The output is the weight coefficient of each surface biomarker, which is used to construct the support vector machine classification model in step S09.

[0040] Among them, the support vector machine classification model is a supervised learning algorithm that classifies data by constructing a hyperplane in a high-dimensional space. The input is the combination of surface markers selected in step S07 and the weight coefficients determined in step S08. The output is a breast cancer subtype classifier, which is used to identify the subtype of the sample to be identified in step S10.

[0041] The distance score is the Mahalanobis distance between the sample to be identified and the cluster center of each subtype, which is calculated by the support vector machine classification model in step S10 and used for breast cancer subtype determination in step S11.

[0042] The classification reliability index is calculated by taking the ratio of the distance score between the test sample and the nearest subtype cluster center to the distance score between the second nearest subtype cluster center obtained in step S10. The smaller the value, the more reliable the classification result, which is used for clinical diagnostic reference.

[0043] Among them, breast cancer subtype-related surface proteins refer to protein markers specifically expressed on the surface of exosomes secreted by breast cancer cells of different molecular subtypes, including human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, matrix metalloproteinase, etc. These are identified and labeled by fluorescently labeled antibodies in step S02, and the differences in their expression patterns are the key basis for distinguishing different breast cancer subtypes by the K-means clustering algorithm in step S05.

[0044] The specific implementation methods of the above steps are described in detail below.

[0045] The specific implementation of step S01 is as follows: First, collect 5-10 ml of peripheral venous blood from patients with known molecular subtypes of breast cancer and patients with subtypes to be identified, using blood collection tubes containing EDTA anticoagulant, and store at 4°C for no more than 2 hours. After blood collection, exosomes are extracted using differential centrifugation. First, centrifuge at 300×g for 10 minutes to remove cellular components. After collecting the supernatant, centrifuge at 2000×g for 20 minutes to remove cell debris, then centrifuge at 10000×g for 30 minutes to remove large particles. Finally, centrifuge at 100000×g for 70 minutes using an ultracentrifuge to precipitate the exosomes. The precipitate is resuspended in sterile phosphate buffer and washed again by ultracentrifugation at 100000×g for 60 minutes. The final exosome precipitate is resuspended in 200 μl of sterile phosphate buffer for later use. Nanoparticle tracking analysis technology verifies that the exosome size distribution is in the range of 30-150 nm, and the concentration reaches 10. 10 To ensure the accuracy of subsequent analyses, samples must be at least 1 per ml.

[0046] The specific implementation of step S02 involves dividing the extracted exosome sample into several equal portions, and adding antibodies labeled with different fluorescent groups to each portion. These antibodies include common exosome marker antibodies such as anti-CD9-FITC, anti-CD63-PE, and anti-CD81-APC, as well as antibodies against breast cancer subtype-related surface proteins such as anti-human epidermal growth factor receptor 2-Cy5, anti-estrogen receptor-Cy3, anti-progesterone receptor-Cy7, anti-Ki-67-PerCP, anti-epithelial cell adhesion molecule-AF647, anti-CD44-PE-Cy7, and anti-matrix metalloproteinase-BV421. The antibody dilution ratio is 1:100 to 1:500, and the samples are incubated at 4°C for 2 hours, gently mixed every 30 minutes. After incubation, the samples are washed three times by centrifugation at 10000×g for 10 minutes to remove unbound antibodies. The washed samples are then resuspended in 100 μl of sterile phosphate buffer to ensure sufficient binding of the fluorescently labeled antibodies to the exosome surface markers. A negative control group was set up, with isotype control antibodies used to replace specific antibodies to correct non-specific fluorescence signals.

[0047] The specific implementation of step S03 involves using a high-resolution flow cytometer to measure the fluorescence intensity of each surface marker in each exosome sample. The flow cytometer excitation source is set to three wavelengths: 488 nm, 561 nm, and 640 nm, corresponding to the fluorescence signals emitted by different fluorophores. The acquisition parameters are set as follows: the forward scatter detector threshold is set to 200 to eliminate background noise; the acquisition rate is controlled at 1000–2000 events / second; and at least 50,000 valid events are acquired for each sample. The acquired raw data is processed using a fluorescence compensation matrix in the flow cytometry software to eliminate spectral overlap interference between different fluorophores. A multidimensional fluorescence intensity matrix is ​​constructed for each sample, where each row represents an exosome sample and each column represents the median fluorescence intensity value of a surface marker. All sample data are standardized using the flow cytometry software to eliminate batch effects and instrument errors, resulting in standardized multidimensional fluorescence intensity matrix data, ready for subsequent analysis.

[0048] The specific implementation of step S04 is based on the aforementioned flow cytometry data, calculating the interaction strength between different markers on the exosome surface. The fluorescence resonance energy transfer equation is used: The calculation is performed, where E is the energy transfer efficiency, R0 is the distance at which 50% energy transfer occurs (Foster distance), and r is the actual distance between the fluorescent donor and acceptor molecules. The Foster distance R0 is calculated using the formula R0 = 0.211[κ]. 2 ·n -4 ·Q D ·J(λ)] 1 / 6 Calculate, where κ is the orientation factor (usually taken as 2 / 3), n is the refractive index of the medium (approximately 1.35 in blood plasma), and Q... D Let E be the quantum yield of the fluorescent donor, and J(λ) be the integral of the overlap between the donor emission spectrum and the acceptor absorption spectrum. Based on the energy transfer efficiency between different biomarkers, a surface protein interaction network was constructed to analyze the conformational characteristics and spatial arrangement patterns of exosome surface proteins in different breast cancer subtypes. When the energy transfer efficiency E is greater than 0.1, the two biomarkers are considered to have an effective interaction. The exosome surface biomarker interaction strength data obtained in this step are used for subsequent biomarker combination screening.

[0049] The specific implementation of step S05 involves analyzing the multidimensional fluorescence intensity matrix of exosome samples from known subtype patients using the K-means clustering algorithm. First, the optimal number of clusters K is determined. The elbow rule and silhouette coefficient are used to evaluate the clustering effect of different K values ​​(2-6), and the parameter with the highest silhouette coefficient and the smallest K value is selected as the final number of clusters. Then, K cluster centers are initialized, and the K-means++ algorithm is used to optimize the selection of initial centers, avoiding local optima caused by random initialization. The K-means iteration process is executed, calculating the Euclidean distance from each sample to each cluster center, assigning the sample to the nearest cluster, and recalculating the coordinates of each cluster center until the change in the cluster center position is less than a preset threshold (usually set to 0.001) or the maximum number of iterations (usually 100) is reached. Finally, K clusters are obtained, each representing a fluorescence characteristic pattern of a breast cancer subtype, and the coordinates of the cluster centers reflect the typical expression pattern of exosome surface markers for that subtype.

[0050] The specific implementation of step S06 involves calculating the Mahalanobis distance between cluster centers of different subtypes to quantitatively assess the degree of difference in the expression of exosome surface markers among different subtypes. The formula for calculating Mahalanobis distance is:

[0051] Where x and y are the cluster center vectors of two different subtypes, and S is the covariance matrix of all samples. First, the covariance matrix S of all samples obtained in step S05 is calculated, and singular value decomposition is performed to ensure matrix invertibility. Then, the Mahalanobis distance between the cluster centers of any two subtypes is calculated, constructing an inter-subtype distance matrix. Larger values ​​in the distance matrix indicate more significant differences in the expression patterns of exosome surface markers between the two subtypes. A threshold of 3.0 is set; when the Mahalanobis distance between two subtypes is greater than this threshold, the two subtypes are considered to have significant discriminative power. The distance matrix data is used in subsequent steps to screen for the surface marker combinations with the highest discriminative power.

[0052] The specific implementation of step S07 involves determining the silhouette coefficient and cohesion of each known subtype cluster, thereby selecting the surface marker combination with the highest discriminative power. The silhouette coefficient calculation formula is: s(i)=(b(i)-a(i)) /

[0053] The function `max(a(i), b(i))` is used, where `a(i)` is the average distance between sample `i` and other samples in the same cluster, and `b(i)` is the average distance between sample `i` and the nearest sample in another cluster. Cohesion is calculated as the average distance from all samples within a cluster to the cluster center. First, the silhouette coefficient and cohesion of all surface marker combinations are calculated. Then, a recursive feature elimination method is used to remove one surface marker at a time, and K-means clustering is performed again to calculate the new silhouette coefficient and cohesion. When the silhouette coefficient exceeds 0.7 and the cohesion is less than 30% of the average distance between cluster centers, the marker combination is considered to have good discriminative power. By comparing the performance indicators of different marker combinations, the optimal surface marker combination is determined; this combination should achieve the highest subtype discrimination ability with the minimum number of features.

[0054] The specific implementation of step S08 involves applying the entropy weight method to optimize the function and quantitatively evaluate the contribution of surface markers. First, the coefficient of variation (CV) of each surface marker across different subtypes is calculated. j =σ j / μ j , where σ j μ represents the standard deviation of marker j among the subtypes. j The average value is then calculated. Then, the information entropy of each marker is calculated. Where p ij Let be the normalized expression value of marker j in subtype i, and m be the number of subtypes. The dissimilarity coefficient d for each marker is calculated based on information entropy. j =1-e j Then, combining the biomarker expression correlation matrix obtained in step S04, the biomarker discrimination index obtained in step S07, and the biomarker expression abundance, a comprehensive evaluation function F is constructed. j =w1·CV j +w2·(1-r j )+w3·d j +w4·A j +w5·D j Where w1 to w5 are the weight coefficients of each factor (all set to 0.2), r j Let A be the average correlation coefficient between marker j and other markers. j To express abundance as a marker, D j To differentiate capabilities, the weighting coefficients of each surface marker in subtype identification were ultimately determined. Where n is the total number of markers.

[0055] The specific implementation of step S09 is to construct a support vector machine classification model based on the selected surface marker combinations and weight coefficients. First, the sample dataset is randomly divided into a training set (75%) and a validation set (25%). A radial basis function is selected as the kernel function: K(x, y) = exp(-γ||xy|| 2 The optimal value of parameter γ is determined using a grid search method (range: 0.001–10). To handle multi-class problems, a one-to-one strategy is used to construct multiple binary support vector machines (SVMs). The SVM penalty parameter C (range: 0.1–100) and kernel function parameters are optimized, and the performance of different parameter combinations is evaluated using five-fold cross-validation. During training, the marker weight coefficients determined in step S08 are introduced to enhance the contribution of key markers through a weighted SVM algorithm. The final model should achieve a classification accuracy of over 85% on the validation set, and both sensitivity and specificity should exceed 80% before it can be used for subtype discrimination of subsequent samples.

[0056] The specific implementation of step S10 involves inputting the multidimensional fluorescence intensity matrix of the exosome sample from the patient to be identified into the trained support vector machine classification model. First, the sample to be identified undergoes the same standardization process as the training set to ensure data distribution consistency. Then, the fluorescence intensity data corresponding to the optimal surface marker combination determined in step S07 is extracted to form a feature vector. The feature vector is input into the support vector machine classification model to calculate the distance from the sample to each hyperplane. Based on these distance values, the Mahalanobis distance score between the sample and the cluster center of each subtype is calculated. The distance score calculation considers the marker weight coefficients determined in step S08, allowing important markers to play a greater role in classification. The score calculation formula is: Among them W j d is the weighting coefficient of marker j. ij denoted as the standardized distance difference between the sample to be tested and the cluster center of subtype i on marker j.

[0057] The specific implementation of step S11 involves determining the breast cancer subtype of the patient sample to be identified based on the minimum distance principle. First, the distance scores of each subtype calculated in step S10 are compared, and the sample is classified into the subtype with the smallest distance score. Then, the classification reliability index R = d is calculated. min / d sec , where d min d is the distance score between the sample and the nearest subtype cluster center. secThe R-value represents the distance score between the sample and the next nearest subtype cluster center. A smaller R-value indicates a more reliable classification result; generally, R < 0.6 is considered highly reliable, 0.6 ≤ R < 0.8 is moderately reliable, and R ≥ 0.8 is low reliable. Simultaneously, the sample's position within the nearest subtype distribution is calculated, and the p-value of the Mahalanobis distance is used to assess whether the sample is an outlier; p < 0.05 may indicate atypical presentation or a novel subtype. Finally, a subtype identification report is generated, including the sample subtype determination result, classification reliability index, and expression levels of key surface markers, providing accurate evidence for clinical diagnosis.

[0058] The mathematical model or calculation process involved in this invention will be described in detail below.

[0059] In step S03, a multidimensional fluorescence intensity matrix is ​​constructed, which is specifically represented as follows:

[0060]

[0061] In the formula, M is the multidimensional fluorescence intensity matrix; I ij The median fluorescence intensity of the j-th surface marker in the i-th sample is represented by m, where m is the number of samples and n is the number of surface marker types.

[0062] Data standardization uses the Z-score method, specifically as follows:

[0063]

[0064] In the formula, Z ij This is the standardized fluorescence intensity value; I ij This represents the original fluorescence intensity value; μ j σ is the average fluorescence intensity of the j-th surface marker in all samples; j Let be the standard deviation of the j-th surface marker across all samples.

[0065] In step S04, the fluorescence resonance energy transfer equation is calculated as follows:

[0066]

[0067] In the formula, E is the energy transfer efficiency, ranging from 0 to 1; R0 is the distance at which 50% energy transfer occurs (Foster distance), in nanometers; and r is the actual distance between the fluorescent donor and acceptor molecules, in nanometers.

[0068] The formula for calculating the Foster distance R0 is:

[0069] R0 = 0.211 × [κ] 2 ×n -4 ×Q D ×J(λ)] 1 / 6;

[0070] In the formula, κ is the orientation factor, describing the spatial orientation relationship between the donor-emitted dipole and the acceptor-absorbed dipole, with a theoretical value ranging from 0 to 4, and typically taking a value of 2 / 3 under random orientation conditions; n is the refractive index of the medium, approximately 1.35 in plasma; Q D λ represents the quantum yield of the fluorescent donor, ranging from 0 to 1; J(λ) is the overlap integral of the donor emission spectrum and the acceptor absorption spectrum, in megohms. -1 cm 3 .

[0071] The formula for calculating the spectral overlap integral J(λ) is:

[0072] J(λ)=∫F D (λ)×ε A (λ)×λ 4 ×dλ;

[0073] In the formula, F D (λ) represents the normalized donor emission spectrum; ε A (λ) is the molar absorptivity of the receptor, in units of M. -1 cm -1 λ represents wavelength, measured in nanometers.

[0074] Considering the spatial fluidity of exosome surface proteins, the formula for calculating energy transfer efficiency is modified as follows:

[0075] E actual =E×(1-α×D);

[0076] In the formula, E actual E represents the actual energy transfer efficiency; E represents the theoretical energy transfer efficiency; α represents the mobility correction coefficient, ranging from 0 to 0.5; and D represents the diffusion coefficient of exosome surface proteins, in μm. 2 / s.

[0077] In step S05, the objective function of the K-means clustering algorithm is:

[0078]

[0079] In the formula, J is the objective function value, which needs to be minimized; k is the number of clusters; C i Let x be the i-th cluster; x be the sample point (multidimensional fluorescence intensity vector); μ i Let x be the center point of the i-th cluster; ||x-μ i || represents the distance from sample point x to cluster center μ. i The Euclidean distance.

[0080] The Euclidean distance calculation formula is:

[0081]

[0082] In the formula, x j Let x be the coordinate of the sample point x in the j-th dimension (the fluorescence intensity value of the j-th surface marker); μ ij Let be the coordinates of the i-th cluster center in the j-th dimension; n is the number of surface marker types.

[0083] The formula for updating cluster centers is:

[0084]

[0085] In the formula, |C i | represents the number of samples in the i-th cluster.

[0086] In the K-means++ algorithm, the probability of selecting the initial cluster centers is calculated as follows:

[0087]

[0088] In the formula, P(x) is the probability of selecting sample point x as the next cluster center; D(x) is the distance from sample point x to the nearest selected cluster center; and X is the set of all sample points.

[0089] In step S06, the formula for calculating Mahalanobis distance is:

[0090]

[0091] In the formula, d(x, y) is the Mahalanobis distance between two cluster centers x and y; x and y are the cluster center vectors of two different subtypes; S is the covariance matrix of all samples; S -1 is the inverse of the covariance matrix.

[0092] The formula for calculating the covariance matrix S is:

[0093]

[0094] In the formula, m is the total number of samples; x i Let i be the i-th sample point; This is the mean vector of all samples.

[0095] Considering that the covariance matrix may be non-invertible, singular value decomposition and regularization are used:

[0096] S′=S+λI;

[0097] In the formula, S′ is the regularized covariance matrix; λ is the regularization parameter, which usually takes the value of 0.01 to 0.1; and I is the identity matrix.

[0098] In step S07, the formula for calculating the profile coefficient is:

[0099]

[0100] In the formula, s(i) is the silhouette coefficient of sample i, which ranges from -1 to 1. The closer it is to 1, the better the clustering effect. a(i) is the average distance between sample i and other samples in the same cluster. b(i) is the average distance between sample i and the nearest other sample in the cluster.

[0101] The formulas for calculating a(i) and b(i) are as follows:

[0102]

[0103] In the formula, C i The cluster to which sample i belongs; |C i | represents the number of samples in this cluster; d(i,j) is the distance between sample i and sample j; C k Other clusters that do not belong to i.

[0104] The formula for calculating cohesion is:

[0105]

[0106] In the formula, Cohesion(C i ) represents the cohesion of the i-th cluster; |C i | represents the number of samples in the i-th cluster; x represents the number of clusters C. i Sample points in; μ i Let x be the center point of the i-th cluster; ||x-μ i || represents the distance from sample point x to cluster center μ. i The Euclidean distance.

[0107] The overall clustering quality evaluation function is:

[0108]

[0109] In the formula, Q is the clustering quality score, and the higher the score, the better the clustering effect; m is the total number of samples; k is the number of clusters; α and β are weight coefficients, usually α = 0.7 and β = 0.3.

[0110] In step S08, the formula for calculating the coefficient of variation is:

[0111]

[0112] In the formula, CV j σ is the coefficient of variation for the j-th surface marker; j The standard deviation of the marker among its subtypes; μ j This represents the average expression value of the marker.

[0113] The formula for calculating the information entropy of a marker is:

[0114]

[0115] In the formula, e j p represents the information entropy of the j-th surface marker, ranging from 0 to 1, with values ​​closer to 0 indicating higher discrimination; k represents the number of breast cancer subtypes; p ij The normalized expression value of the j-th marker in the i-th subtype is calculated as follows: Where x ij is the average expression value of the j-th marker in the i-th subtype.

[0116] The formula for calculating the coefficient of difference is:

[0117] d j =1-e j ;

[0118] In the formula, d j is the difference coefficient of the j-th surface marker, with a value ranging from 0 to 1. The closer it is to 1, the higher the discrimination.

[0119] The formula for calculating the correlation of biomarkers is:

[0120]

[0121] In the formula, r jl x is the correlation coefficient between the j-th marker and the l-th marker, with a value ranging from -1 to 1; ij Let be the expression value of the j-th marker in the i-th sample; is the average expression value of the j-th marker; m is the total number of samples.

[0122] The formula for calculating the average correlation coefficient is:

[0123]

[0124] In the formula, r j is the absolute value of the average correlation coefficient between the j-th marker and all other markers; n is the total number of markers.

[0125] The formula for calculating the differentiation ability index is:

[0126]

[0127] In the formula, D j μ is the discrimination index of the j-th marker; ij σ is the average expression value of the j-th marker in the i-th subtype; ijLet be the standard deviation of the j-th marker in the i-th subtype; k is the number of subtypes.

[0128] The formula for calculating the comprehensive evaluation function is:

[0129] F j =w1×CV j +w2×(1-r j )+w3×d j +w4×A j +w5×D j ;

[0130] In the formula, F j The comprehensive score for the j-th surface marker; w1 to w5 are the weight coefficients of each factor, all set to 0.2; A j The abundance of biomarker expression was obtained by flow cytometry and normalized to a range of 0–1.

[0131] The formula for calculating the final weighting coefficient of surface markers is:

[0132]

[0133] In the formula, W j Let be the weighting coefficient of the j-th surface marker in subtype identification, satisfying...

[0134] In step S09, the formula for calculating the radial basis function kernel in the support vector machine is:

[0135] K(x, y) = exp(-γ||xy|| 2 );

[0136] In the formula, K(x, y) is the kernel function value; x and y are the feature vectors of two samples; γ is the kernel function parameter, the optimal value is determined by grid search, and the value range is 0.001 to 10; ||xy|| is the Euclidean distance between the feature vectors of the two samples.

[0137] The formula for calculating the weighted Euclidean distance considering the weights of the markers is:

[0138]

[0139] In the formula, ||xy|| w W is the weighted Euclidean distance; j x is the weighting coefficient for the j-th surface marker; j and y j denoted as , where are the feature values ​​of the two samples in the j-th dimension; n is the feature dimension.

[0140] The objective function for the weighted support vector machine is:

[0141]

[0142] Constraints: y i (w T φ(x i )+b)≥1-ξ i ξ i ≥0, i=1,2,...,m;

[0143] In the formula, w is the normal vector of the hyperplane; b is the intercept; ξ i y is a slack variable; C is a penalty parameter, ranging from 0.1 to 100; i Let φ(x) be the class label for sample i; i ) is the feature vector of sample i after being mapped by the kernel function; m is the number of training samples.

[0144] In step S10, the distance score between the sample to be identified and the cluster centers of each subtype is calculated using the following formula:

[0145]

[0146] In the formula, Score(i) is the distance score between the sample to be tested and the i-th subtype cluster center; W j d represents the weighting coefficient of the j-th surface marker; ij is the standardized distance difference between the sample to be tested and the i-th subtype cluster center on the j-th marker; n is the feature dimension.

[0147] Standardized distance difference d ij The calculation formula is:

[0148]

[0149] In the formula, x j μ is the expression value of the sample to be tested on the j-th biomarker; ij Let σ be the coordinates of the i-th subtype cluster center on the j-th marker; ij Let be the standard deviation of the i-th subtype on the j-th marker.

[0150] In step S11, the formula for calculating the classification reliability index is:

[0151]

[0152] In the formula, R is the classification reliability index, ranging from 0 to 1; the smaller the value, the more reliable the classification result. min d represents the distance score between the sample and the nearest subtype cluster center. sec The distance score is the distance score between the sample and the next nearest subtype cluster center.

[0153] The Mahalanobis distance P-value for outlier assessment is calculated based on the chi-square distribution:

[0154]

[0155] In the formula, P is the probability that the sample is an outlier; D is the cumulative distribution function of a chi-square distribution with n degrees of freedom; 2 is the square of the Mahalanobis distance from the sample to its cluster center; n is the feature dimension.

[0156] For error compensation in practical applications, a noise correction term is introduced:

[0157] Score'(i)=Score(i)×(1+ε×β i );

[0158] In the formula, Score′(i) is the corrected distance score; Score(i) is the original distance score; ε is random noise, following a normal distribution with a mean of 0 and a standard deviation of 0.05; β i is the noise sensitivity coefficient for the i-th subtype, with a value ranging from 0.5 to 1.5, determined through cross-validation.

[0159] Specifically, the principle of this invention is as follows: By integrating the expression modes of multiple surface markers and considering their interactions, this invention establishes a novel method for identifying breast cancer subtypes. Its core principles lie in the following aspects:

[0160] First, exosomes, as nanoscale membrane vesicles secreted by cells, inherit the molecular characteristics of their source cells in terms of surface biomarker composition. Exosomes secreted by different molecular subtypes of breast cancer cells carry specific surface protein profiles. For example, Luminal-type exosomes are rich in estrogen receptor-related proteins, HER2-positive exosomes highly express human epidermal growth factor receptor 2, and triple-negative exosomes express specific tumor stem cell markers. This molecular subtype specificity provides a theoretical basis for using exosomal biomarkers to identify breast cancer subtypes.

[0161] Secondly, this invention overcomes the limitations of traditional single-marker analysis by employing multicolor fluorescent labeling combined with flow cytometry to construct a multidimensional fluorescence intensity matrix. This not only considers the absolute expression levels of each marker but also analyzes the interaction strength and spatial conformation characteristics between markers through fluorescence resonance energy transfer equations. This systematic analytical method can capture the complex patterns of exosome surface protein expression, providing more comprehensive information on molecular subtype characteristics.

[0162] Furthermore, this invention introduces multiple machine learning algorithms to optimize classification performance. K-means clustering enables automatic identification of fluorescence feature patterns in different subtypes; Mahalanobis distance considers the correlation between features, more accurately quantifying differences between subtypes; silhouette coefficient and cohesion evaluation ensure clustering quality; entropy weighting optimizes the function to objectively determine the contribution weight of each marker in subtype identification; support vector machine constructs a high-dimensional classification model, ultimately achieving accurate subtype determination.

[0163] Finally, this invention determines the subtype of a sample by calculating the distance score between the sample and the cluster centers of each subtype, based on the principle of minimum distance. It also introduces a classification reliability index to assess the accuracy of the results, providing a reference for clinical decision-making. The entire technical solution conforms to a complete logical chain from sample collection, data acquisition, model construction to result verification, effectively solving the technical problem of accurately identifying breast cancer subtypes using exosomes from blood samples.

[0164] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.

[0165] The specific implementation of step S01 is as follows: First, collect 5-10 ml of peripheral venous blood from patients with known molecular subtypes of breast cancer and patients with subtypes to be identified, using blood collection tubes containing EDTA anticoagulant, and store at 4°C for no more than 2 hours. After blood collection, exosomes are extracted using differential centrifugation. First, centrifuge at 300×g for 10 minutes to remove cellular components. After collecting the supernatant, centrifuge at 2000×g for 20 minutes to remove cell debris, then centrifuge at 10000×g for 30 minutes to remove large particles. Finally, centrifuge at 100000×g for 70 minutes using an ultracentrifuge to precipitate the exosomes. The precipitate is resuspended in sterile phosphate buffer and washed again by ultracentrifugation at 100000×g for 60 minutes. The final exosome precipitate is resuspended in 200 μl of sterile phosphate buffer for later use. Nanoparticle tracking analysis technology verifies that the exosome size distribution is in the range of 30-150 nm, and the concentration reaches 10. 10 To ensure the accuracy of subsequent analyses, the number of exosomes per ml must be above a certain level. This step utilizes the differences in size and density between exosomes and other components in the blood, achieving high-purity separation of exosomes through multiple differential centrifugations. This is a fundamental step in establishing a method for identifying breast cancer subtypes using exosome surface markers.

[0166] The specific implementation of step S02 involves dividing the extracted exosome sample into several equal portions, and adding antibodies labeled with different fluorescent groups to each portion. These antibodies include common exosome marker antibodies such as anti-CD9-FITC, anti-CD63-PE, and anti-CD81-APC, as well as antibodies against breast cancer subtype-related surface proteins such as anti-human epidermal growth factor receptor 2-Cy5, anti-estrogen receptor-Cy3, anti-progesterone receptor-Cy7, anti-Ki-67-PerCP, anti-epithelial cell adhesion molecule-AF647, anti-CD44-PE-Cy7, and anti-matrix metalloproteinase-BV421. The antibody dilution ratio is 1:100 to 1:500, and the samples are incubated at 4°C for 2 hours, gently mixed every 30 minutes. After incubation, the samples are washed three times by centrifugation at 10000×g for 10 minutes to remove unbound antibodies. The washed samples are then resuspended in 100 μl of sterile phosphate buffer to ensure sufficient binding of the fluorescently labeled antibodies to the exosome surface markers. This step utilizes the principle of specific binding between antigen and antibody to label exosome surface proteins, and uses fluorescent groups with different excitation and emission wavelengths to achieve multicolor fluorescent labeling, laying the foundation for subsequent flow cytometry detection and multidimensional feature analysis.

[0167] The specific implementation of step S03 involves using a high-resolution flow cytometer to measure the fluorescence intensity of each surface marker in each exosome sample, and constructing a multidimensional fluorescence intensity matrix based on the multicolor fluorescent labeling in step S02. The flow cytometer excitation source is set to three wavelengths: 488 nm, 561 nm, and 640 nm, corresponding to the fluorescence signals emitted by different fluorophores. The forward scattering detector threshold is set to 200 to eliminate background noise, and the acquisition rate is controlled at 1000–2000 events / second, with at least 50,000 valid events acquired for each sample. The multidimensional fluorescence intensity matrix is ​​represented as follows:

[0168]

[0169] In the formula, M is the multidimensional fluorescence intensity matrix; I ij This represents the median fluorescence intensity of the j-th surface marker in the i-th sample; m is the number of samples; and n is the number of surface marker species. The acquired raw data were processed using a fluorescence compensation matrix in the flow cytometry software to eliminate spectral overlap interference between different fluorescent groups. To eliminate batch effects and instrument errors, Z-score standardization was performed.

[0170]

[0171] In the formula, Z ij This is the standardized fluorescence intensity value; I ij This represents the original fluorescence intensity value; μ j σ is the average fluorescence intensity of the j-th surface marker in all samples;j Let be the standard deviation of the j-th surface biomarker across all samples. This step uses flow cytometry to quantitatively detect exosome surface biomarkers, constructing a multidimensional data structure reflecting the expression characteristics of surface proteins in each sample, thus laying the data foundation for subsequent analysis.

[0172] The specific implementation of step S04 involves calculating the interaction strength between different markers on the exosome surface based on flow cytometry data, and analyzing the conformational characteristics and spatial arrangement patterns of exosome surface proteins in different breast cancer subtypes. The fluorescence resonance energy transfer equation is used for calculation.

[0173]

[0174] In the formula, E is the energy transfer efficiency, ranging from 0 to 1; R0 is the distance at which 50% energy transfer occurs (Foster distance), in nanometers; and r is the actual distance between the fluorescent donor and acceptor molecules, in nanometers. The Foster distance R0 is calculated using the formula:

[0175] R0 = 0.211 × [κ] 2 ×n -4 ×Q D ×J(λ)] 1 / 6 ;

[0176] In the formula, κ is the orientation factor, describing the spatial orientation relationship between the donor-emitted dipole and the acceptor-absorbed dipole, with a theoretical value ranging from 0 to 4, and typically taking a value of 2 / 3 under random orientation conditions; n is the refractive index of the medium, approximately 1.35 in plasma; Q D λ represents the quantum yield of the fluorescent donor, ranging from 0 to 1; J(λ) is the overlap integral of the donor emission spectrum and the acceptor absorption spectrum, in megohms. -1 cm 3 Calculated using the formula:

[0177] J(λ)=∫F D (λ)×ε A (λ)×λ 4 ×dλ;

[0178] In the formula, F D (λ) represents the normalized donor emission spectrum; ε A (λ) is the molar absorptivity of the receptor, in units of M. -1 cm -1 λ represents wavelength in nanometers. Considering the spatial fluidity of proteins on the exosome surface, a modified formula for calculating energy transfer efficiency is introduced:

[0179] E actual =E×(1-α×D);

[0180] In the formula, Eactual E represents the actual energy transfer efficiency; E represents the theoretical energy transfer efficiency; α represents the mobility correction coefficient, ranging from 0 to 0.5; and D represents the diffusion coefficient of exosome surface proteins, in μm. 2 / s. When the energy transfer efficiency E actual A value greater than 0.1 indicates an effective interaction between the two biomarkers. This step utilizes the fluorescence resonance energy transfer principle to analyze the spatial relationships and interactions between exosome surface biomarkers, revealing the molecular arrangement characteristics of exosome surface proteins in different breast cancer subtypes and providing conformational information for subtype identification.

[0181] The specific implementation of step S05 involves applying the K-means clustering algorithm to analyze the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes to determine the fluorescence characteristic patterns of different breast cancer subtypes. The objective function of the K-means clustering algorithm is:

[0182]

[0183] In the formula, J is the objective function value, which needs to be minimized; k is the number of clusters; C i Let x be the i-th cluster; x be the sample point (multidimensional fluorescence intensity vector); μ i Let x be the center point of the i-th cluster; ||x-μ i || represents the distance from sample point x to cluster center μ. i The Euclidean distance is calculated using the following formula:

[0184]

[0185] In the formula, x j Let x be the coordinate of the sample point x in the j-th dimension (the fluorescence intensity value of the j-th surface marker); μ ij Let be the coordinates of the i-th cluster center in the j-th dimension; n is the number of surface marker species. The cluster center update formula is:

[0186]

[0187] In the formula, |C i | represents the number of samples in the i-th cluster. In practice, the optimal number of clusters K is first determined. The elbow rule and silhouette coefficient are used to evaluate the clustering effect of different K values ​​(2-6), and the parameter with the highest silhouette coefficient and the smallest K value is selected as the final number of clusters. Then, the K-means++ algorithm is used to optimize the selection of initial cluster centers. The probability of selecting the initial cluster centers is calculated as follows:

[0188]

[0189] In the formula, P(x) is the probability of selecting sample point x as the next cluster center; D(x) is the distance from sample point x to the nearest selected cluster center; and X is the set of all sample points. The K-means iterative process is performed until the change in the cluster center position is less than a preset threshold (usually set to 0.001) or the maximum number of iterations (usually 100) is reached. This step automatically identifies and separates the expression patterns of exosome surface markers for different breast cancer subtypes using the K-means clustering algorithm, providing a clustering basis for subsequent differentiation of different subtypes.

[0190] The specific implementation of step S06 involves calculating the Mahalanobis distance between cluster centers of different subtypes to quantitatively assess the degree of difference in the expression of exosome surface markers among different subtypes. The formula for calculating Mahalanobis distance is:

[0191]

[0192] In the formula, d(x, y) is the Mahalanobis distance between two cluster centers x and y; x and y are the cluster center vectors of two different subtypes; S is the covariance matrix of all samples; S -1 Let S be the inverse of the covariance matrix. The formula for calculating the covariance matrix S is:

[0193]

[0194] In the formula, m is the total number of samples; x i Let i be the i-th sample point; Let be the mean vector of all samples. Considering that the covariance matrix may not be invertible, singular value decomposition and regularization are used:

[0195] S′=S+λI;

[0196] In the formula, S′ is the regularized covariance matrix; λ is the regularization parameter, typically ranging from 0.01 to 0.1; and I is the identity matrix. First, the covariance matrix S of all samples is calculated. Then, the Mahalanobis distance between any two subtype cluster centers is calculated, constructing an inter-subtype distance matrix. A threshold of 3.0 is set; when the Mahalanobis distance between two subtypes is greater than this threshold, the two subtypes are considered to have significant discriminative power. This step utilizes Mahalanobis distance to consider the correlation between features, making it more suitable than Euclidean distance for assessing the degree of difference between subtypes in a multidimensional feature space, providing a quantitative indicator for subsequent biomarker combination screening.

[0197] The specific implementation of step S07 involves determining the silhouette coefficient and cohesion of each known subtype cluster, and then selecting the surface marker combination with the highest discriminative power. The silhouette coefficient calculation formula is:

[0198]

[0199] In the formula, s(i) is the silhouette coefficient of sample i, ranging from -1 to 1, with values ​​closer to 1 indicating better clustering; a(i) is the average distance between sample i and other samples in the same cluster; b(i) is the average distance between sample i and the nearest sample in another cluster. The formulas for calculating a(i) and b(i) are as follows:

[0200]

[0201] In the formula, C i The cluster to which sample i belongs; |C i | represents the number of samples in this cluster; d(i,j) is the distance between sample i and sample j; C k For other clusters not belonging to i, the formula for calculating cohesion is:

[0202]

[0203] In the formula, Cohesion(C i ) represents the cohesion of the i-th cluster; |C i | represents the number of samples in the i-th cluster; x represents the number of clusters C. i Sample points in; μ i Let x be the center point of the i-th cluster; ||x-μ i || represents the distance from sample point x to cluster center μ. i The Euclidean distance. The overall clustering quality evaluation function is:

[0204]

[0205] In the formula, Q represents the clustering quality score, with a higher score indicating better clustering performance; m is the total number of samples; k is the number of clusters; and α and β are weighting coefficients, typically α = 0.7 and β = 0.3. First, the cluster profile coefficient and cohesion of all surface marker combinations are calculated. Then, a recursive feature elimination method is used to remove one surface marker at a time, and K-means clustering is performed again to calculate the new profile coefficient and cohesion. When the profile coefficient exceeds 0.7 and the cohesion is less than 30% of the average distance between cluster centers, the marker combination is considered to have good discriminative power. This step, by evaluating clustering quality, selects the optimal surface marker combination, achieving the highest subtype discrimination ability with the minimum number of features.

[0206] The specific implementation of step S08 involves applying the entropy weight method to optimize the function and quantitatively evaluate the contribution of surface markers, determining the weight coefficient of each marker in subtype identification. First, the coefficient of variation is calculated:

[0207]

[0208] In the formula, CV jσ is the coefficient of variation for the j-th surface marker; j The standard deviation of the marker among its subtypes; μ j This represents the average expression value of the marker. Then, the information entropy is calculated:

[0209]

[0210] In the formula, e j p represents the information entropy of the j-th surface marker, ranging from 0 to 1; k represents the number of breast cancer subtypes; p ij The normalized expression value of the j-th marker in the i-th subtype is calculated as follows: Calculate the difference coefficient based on information entropy:

[0211] d j =1-e j ;

[0212] In the formula, d j Let be the difference coefficient for the j-th surface marker, ranging from 0 to 1, with values ​​closer to 1 indicating higher discriminative power. The correlation coefficient between markers is calculated as follows:

[0213]

[0214] The average correlation coefficient is calculated as follows:

[0215]

[0216] The differentiation ability index is calculated as follows:

[0217]

[0218] The comprehensive evaluation function is calculated as follows:

[0219] F j =w1×CV j +w2×(1-r j )+w3×d j +w4×A j +w5×D j ;

[0220] In the formula, F j The comprehensive score for the j-th surface marker; w1 to w5 are the weight coefficients of each factor, all set to 0.2; A j The abundance of the biomarkers is expressed. The final weighting coefficients are calculated as follows:

[0221]

[0222] In the formula, W j Let be the weighting coefficient of the j-th surface marker in subtype identification, satisfying... This step uses the entropy weighting method to comprehensively consider the variability, correlation, information entropy, expression abundance, and discriminative ability of biomarkers among different subtypes, objectively assessing the contribution of each biomarker in subtype identification and providing weight parameters for the construction of support vector machine classification models.

[0223] The specific implementation of step S09 is to construct a support vector machine (SVM) classification model based on the selected surface marker combinations and weight coefficients, which is used for subtype discrimination of subsequent samples to be identified. The formula for calculating the radial basis function kernel in the SVM is:

[0224] K(x, y) = exp(-γ||xy|| 2 );

[0225] In the formula, K(x, y) is the kernel function value; x and y are the feature vectors of the two samples; γ is the kernel function parameter, the optimal value is determined through grid search, and its value ranges from 0.001 to 10; ||xy|| is the Euclidean distance between the feature vectors of the two samples. The weighted Euclidean distance considering the marker weights is calculated as follows:

[0226]

[0227] In the formula, ||xy|| w W is the weighted Euclidean distance; j x is the weighting coefficient for the j-th surface marker; j and y j Let be the feature values ​​of the two samples in the j-th dimension, and n be the feature dimension. The optimization objective function of the weighted support vector machine is:

[0228]

[0229] Constraints: y i (w T φ(x i )+b)≥1-ξ i ξ i ≥0, i=1,2,...,m;

[0230] In the formula, w is the normal vector of the hyperplane; b is the intercept; ξ i y is a slack variable; C is a penalty parameter, ranging from 0.1 to 100; i Let φ(x) be the class label for sample i; iLet be the feature vector of sample i after kernel function mapping; m be the number of training samples. First, the sample dataset is randomly divided into a training set (75%) and a validation set (25%). Five-fold cross-validation is used to evaluate the performance of different parameter combinations. The final model should achieve a classification accuracy of over 85% on the validation set. This step utilizes the non-linear classification capability of the support vector machine algorithm, combined with the marker weight coefficients determined in step S08, to construct a classification model capable of accurately distinguishing different breast cancer subtypes.

[0231] The specific implementation of step S10 involves inputting the multidimensional fluorescence intensity matrix of the exosome sample from the patient to be identified into the trained support vector machine classification model, and calculating the distance score between the sample and the cluster center of each subtype. First, the sample to be identified undergoes the same standardization process as the training set. Then, the fluorescence intensity data corresponding to the optimal surface marker combination determined in step S07 is extracted to form a feature vector. The feature vector is input into the support vector machine classification model to calculate the Mahalanobis distance score between the sample and the cluster center of each subtype. The distance score calculation formula is:

[0232]

[0233] In the formula, Score(i) is the distance score between the sample to be tested and the i-th subtype cluster center; W j d represents the weighting coefficient of the j-th surface marker; ij The standardized distance difference between the sample to be tested and the i-th subtype cluster center on the j-th marker is calculated using the following formula:

[0234]

[0235] In the formula, x j μ is the expression value of the sample to be tested on the j-th biomarker; ij Let σ be the coordinates of the i-th subtype cluster center on the j-th marker; ij Let be the standard deviation of the i-th subtype on the j-th biomarker. This step analyzes the expression characteristics of exosome surface biomarkers in the samples to be identified using a support vector machine classification model, calculates the distance score between the sample and the cluster centers of each subtype, and provides a quantitative basis for the final subtype determination.

[0236] The specific implementation of step S11 involves determining the breast cancer subtype of the patient sample to be identified based on the minimum distance principle and calculating the classification reliability index. First, the distance scores of each subtype calculated in step S10 are compared, and the sample is classified into the subtype with the smallest distance score. Then, the classification reliability index is calculated:

[0237]

[0238] In the formula, R is the classification reliability index, ranging from 0 to 1; the smaller the value, the more reliable the classification result. min d represents the distance score between the sample and the nearest subtype cluster center. sec This represents the distance score between the sample and the next nearest subtype cluster center. Generally, R < 0.6 indicates high reliability, 0.6 ≤ R < 0.8 indicates moderate reliability, and R ≥ 0.8 indicates low reliability. Simultaneously, the sample's position within the nearest subtype distribution is calculated, and the p-value of the Mahalanobis distance is used to assess whether the sample is an outlier.

[0239]

[0240] In the formula, P is the probability that the sample is an outlier; D is the cumulative distribution function of a chi-square distribution with n degrees of freedom; 2 Let be the squared Mahalanobis distance from the sample to its cluster center; n is the feature dimension. Considering errors in practical applications, a noise correction term is introduced:

[0241] Score'(i)=Score(i)×(1+ε×β i );

[0242] In the formula, Score′(i) is the corrected distance score; Score(i) is the original distance score; ε is random noise, following a normal distribution with a mean of 0 and a standard deviation of 0.05; β i The noise sensitivity coefficient for the i-th subtype ranges from 0.5 to 1.5. A subtype identification report is ultimately generated, including the sample subtype determination result, classification reliability index, and expression levels of key surface markers, providing accurate evidence for clinical diagnosis. This step assesses the reliability of the classification results by calculating the classification reliability index and outlier probability, improving the accuracy of subtype identification and its clinical reference value.

[0243] To better understand and implement this invention, the following is an example 2 of a specific application scenario: Researchers collected peripheral venous blood samples from 120 patients diagnosed with breast cancer at a medical center, including 30 patients with Luminal A, 30 with Luminal B, 30 with HER2 overexpression, and 30 with triple-negative breast cancer. Blood samples from 30 healthy women were also collected as controls. All blood samples were collected in EDTA anticoagulant tubes and stored at 4°C for no more than 2 hours. Using differential centrifugation, cellular components were first removed by centrifugation at 300×g for 10 minutes. The supernatant was then centrifuged at 2000×g for 20 minutes to remove cell debris, followed by centrifugation at 10000×g for 30 minutes to remove large particles. Finally, exosomes were precipitated by ultracentrifugation at 100000×g for 70 minutes. The precipitate was resuspended in sterile PBS and washed again by ultracentrifugation at 100000×g for 60 minutes. The final exosome precipitate was resuspended in 200 μl of PBS for later use.

[0244] The extracted exosomes were identified using nanoparticle tracking analysis. The results showed that the exosomes were mainly distributed in the range of 30–150 nm in size, with a concentration of 5.3 × 10⁻⁶. 10 The sample size was [number] cells / ml, consistent with exosome characteristics. Extracted exosome samples were incubated with antibodies labeled with different fluorescent groups, including anti-CD9-FITC, anti-CD63-PE, anti-CD81-APC, and antibodies against breast cancer subtype-related surface proteins, at 4°C for 2 hours. The average fluorescence intensity of the main exosome markers in each type of sample is shown in Table 1.

[0245] Table 1. Mean fluorescence intensity (MFI) of major exosome markers in different sample types.

[0246] Sample type CD9-FITC CD63-PE CD81-APC Luminal A 8654 7892 9123 Luminal B 8976 8234 9356 HER2 overexpression 8432 7756 8975 Triple Negative 7986 7432 8543 Health comparison 7654 6897 8132

[0247] Fluorescence intensities of surface markers in each exosome sample were measured using flow cytometry, and a multidimensional fluorescence intensity matrix was constructed. Table 2 shows the fluorescence intensities of breast cancer-related markers on the surface of exosomes from different breast cancer subtypes.

[0248] Table 2. Fluorescence Intensities (MFI) of Breast Cancer-Related Markers on Exosome Surfaces of Different Breast Cancer Subtypes

[0249]

[0250] Based on flow cytometry data, the interaction strengths between different markers on the exosome surface were calculated. The energy transfer efficiencies of each fluorescent donor-acceptor pair were obtained using the fluorescence resonance energy transfer equation, as shown in Table 3.

[0251] Table 3 Energy transfer efficiency of each fluorescent donor-acceptor pair

[0252]

[0253] The multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes was analyzed using the K-means clustering algorithm. Based on the elbow rule and silhouette coefficient evaluation, the optimal cluster number K=4 was determined, corresponding to the four breast cancer subtypes. The silhouette coefficient and cohesion of the K-means clustering results are shown in Table 4.

[0254] Table 4 shows the silhouette coefficient and cohesion of the K-means clustering results.

[0255] Evaluation indicators Luminal A Luminal B HER2 overexpression Triple Negative Profile coefficient 0.78 0.75 0.82 0.81 cohesion 0.32 0.37 0.29 0.28

[0256] Mahalanobis distances between cluster centers of different subtypes were calculated to quantitatively assess the differences in the expression of exosome surface markers among different subtypes. The results are shown in Table 5.

[0257] Table 5 Mahalanobis distances between cluster centers of different subtypes

[0258] Subtype comparison Mahalanobis distance Luminal A vs Luminal B 3.56 Luminal A vs HER2 overexpression 6.87 Luminal A vs Triple Negative 8.32 Luminal B vs HER2 overexpression 5.43 Luminal B vs Triple Negative 7.21 HER2 overexpression vs. triple negative 6.98

[0259] By combining recursive feature elimination with silhouette coefficient and cohesion evaluation, the surface marker combinations with the highest discriminative power were selected. The clustering quality scores of different marker combinations are shown in Table 6.

[0260] Table 6 Cluster quality scores for different biomarker combinations

[0261] Logo combination Cluster quality score All 7 types of markers 0.75 HER2, ER, PR, Ki-67, CD44, MMP 0.79 HER2, ER, PR, Ki-67, CD44 0.82 HER2,ER,PR,CD44 0.78 HER2,ER,CD44 0.71

[0262] The entropy weight method was used to optimize the function and quantitatively evaluate the contribution of surface markers, determining the weight coefficient of each marker in subtype identification. The results are shown in Table 7.

[0263] Table 7. Weighting coefficients of various surface markers in subtype identification.

[0264]

[0265] A support vector machine (SVM) classification model was constructed based on the selected combinations of five biomarkers (HER2, ER, PR, Ki-67, and CD44) and their weighting coefficients. Through five-fold cross-validation, the optimal parameters were determined to be γ = 0.05 and C = 10. A sample of 90 breast cancer patients was randomly selected as the training set, and a sample of 30 patients was selected as the validation set. The model's performance on the validation set is shown in Table 8.

[0266] Table 8 shows the performance of the Support Vector Machine classification model on the validation set.

[0267] Performance indicators Luminal A Luminal B HER2 overexpression Triple Negative average value accuracy 0.91 0.87 0.92 0.89 0.90 Sensitivity 0.88 0.84 0.90 0.85 0.87 Specificity 0.93 0.89 0.95 0.92 0.92 F1 score 0.90 0.86 0.92 0.88 0.89

[0268] The constructed model was used to identify the subtype of 20 newly diagnosed breast cancer patients with undetermined molecular subtypes. The multidimensional fluorescence intensity matrix of exosome samples from these patients was input into the trained support vector machine classification model, and the distance score between the sample and the cluster center of each subtype was calculated. The results are compared with the immunohistochemical results, as shown in Table 9.

[0269] Table 9 Comparison of exosome model prediction results and immunohistochemical results

[0270]

[0271]

[0272] Traditional molecular subtype identification of breast cancer mainly relies on immunohistochemistry or gene detection via tissue biopsy. These methods are not only highly invasive but also fail to reflect the heterogeneity and dynamic changes of tumors. The method for identifying breast cancer subtypes using exosome surface markers proposed in this invention requires only peripheral venous blood collection from the patient. By analyzing the expression patterns and interaction characteristics of tumor-derived exosome surface markers in the blood, non-invasive identification of breast cancer subtypes is achieved. Compared with traditional methods, this invention has the following advancements: First, non-invasive sampling significantly reduces the burden on patients and is suitable for repeated monitoring and dynamic assessment; second, exosomes, as information carriers secreted by tumor cells, can more comprehensively reflect the heterogeneous characteristics of the primary tumor; third, the analytical method combining multidimensional fluorescence characteristics and fluorescence resonance energy transfer considers not only the expression level of surface markers but also the spatial arrangement and interactions between proteins, providing richer molecular information; fourth, the machine learning-based classification model and weight optimization strategy achieve high-precision subtype identification, with a prediction accuracy of 90%, approaching the traditional gold standard method. In practical validation, the predictive results of this method for 20 newly diagnosed patients showed a 95% concordance rate with immunohistochemical results, demonstrating the clinical application value of this method. Furthermore, this method also provides a classification reliability index, offering a reference for clinical decision-making.

[0273] It should be noted that the variables involved in this invention are explained in detail in Tables 10 and 11 below.

[0274] Table 10 Variable Explanation Table (Part 1)

[0275]

[0276]

[0277] Table 11 Variable Explanation Table (Part Two)

[0278]

[0279] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying breast cancer subtypes using exosome surface markers, characterized in that, include: A multidimensional fluorescence intensity matrix for each exosome sample was constructed using the fluorescence intensities of all surface markers obtained from each sample. The interaction strength between exosome surface markers was calculated using the fluorescence resonance energy transfer equation, and the conformational characteristics of exosome surface proteins in different subtypes were analyzed. K-means clustering was used to analyze the multidimensional fluorescence intensity matrix of exosome samples from patients with known subtypes to determine the fluorescence characteristic patterns of different breast cancer subtypes. Mahalanobis distance between cluster centers of different subtypes was calculated to establish a quantitative index for the expression differences of exosome surface markers among subtypes. The silhouette coefficient and cohesion of each known subtype cluster were measured to screen the surface marker combination with the highest discriminative power. The entropy weighting optimization function was applied to quantitatively evaluate the contribution of surface markers and determine the weight coefficient of each marker in subtype identification. A support vector machine (SVM) classification model was constructed based on the screened surface marker combinations and weight coefficients, and the model was trained using known subtype samples. The multidimensional fluorescence intensity matrix of exosome samples from patients to be identified was input into the trained SVM model. In the support vector machine (SVM) classification model, the distance score between the sample and the cluster center of each subtype is calculated; the breast cancer subtype to which the patient sample to be identified belongs is determined according to the minimum distance principle, and the classification reliability index is calculated; wherein, the multidimensional fluorescence intensity matrix refers to the data structure composed of the fluorescence signal intensity of multiple surface markers in each exosome sample, with each row representing a sample and each column representing the fluorescence intensity value of a surface marker; the fluorescence resonance energy transfer equation is used to calculate the energy transfer efficiency between different marker molecules on the exosome surface, the inputs include the fluorescence donor quantum yield, the distance between fluorescence donor and acceptor molecules, the fluorescence donor and acceptor spectral overlap integral, the refractive index, and the fluorescence resonance energy transfer direction factor, and the output is the interaction strength between surface markers; the construction of the SVM classification model is based on the selected combination of surface markers and weight coefficients to construct the SVM classification model, and the model is trained using known subtype samples; The step of multicolor fluorescent labeling of exosome surface markers involves using fluorescently labeled antibodies to perform multicolor fluorescent labeling of exosome surface markers, including CD9, CD63, CD81, and breast cancer subtype-related surface proteins. The breast cancer subtype-related surface proteins include human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, epithelial cell adhesion molecule, tumor stem cell marker CD44, and matrix metalloproteinase. The breast cancer subtypes include Luminal A, Luminal B, HER2 overexpression, and triple negative. The fluorescence resonance energy transfer equation is expressed as: E = In the formula, E represents the energy transfer efficiency. The Foster distance is the distance between the fluorescent donor and acceptor molecules, and r is the actual distance between the fluorescent donor and acceptor molecules. Through formula Calculate, where κ is the orientation factor and n is the refractive index of the medium. The quantum yield of the fluorescent donor is given by λ, and J(λ) is the integral of the overlap between the donor emission spectrum and the acceptor absorption spectrum. When the energy transfer efficiency E is greater than 0.1, the two markers are considered to have an effective interaction. In the step of calculating the Mahalanobis distance between cluster centers of different subtypes and establishing a quantitative index for differential expression of exosome surface markers among subtypes, the Mahalanobis distance calculation formula is: d(x, y) = Where x and y are the cluster center vectors of two different subtypes, and S is the covariance matrix of all samples. Specifically, firstly, the covariance matrix S of all samples is calculated, and then the Mahalanobis distance between the cluster centers of any two subtypes is calculated to construct the inter-subtype distance matrix. The larger the value in the distance matrix, the more significant the difference in the expression patterns of exosome surface markers between the two subtypes. The threshold is set to 3.

0. When the Mahalanobis distance between two subtypes is greater than this threshold, the two subtypes are considered to have significant distinguishability. In the step of determining the silhouette coefficient and cohesion of each known subtype cluster and screening the surface marker combination with the highest discriminative power, the silhouette coefficient is calculated using the following formula: ,in For the sample The average distance to other samples in the same cluster. For the sample The average distance to the nearest other cluster sample; cohesion is calculated as the average distance of all samples within a cluster to the cluster center. When the silhouette coefficient exceeds 0.7 and the cohesion is less than 30% of the average distance between cluster centers, the combination of markers is considered to have good discriminative power. The steps for determining the weighting coefficients of each marker in subtype identification specifically include: firstly, calculating the coefficient of variation of each surface marker among different subtypes. ,in Let j be the standard deviation of marker j among the subtypes. It is the average value; then the information entropy of each marker is calculated. ,in For marker j in subtype Normalized expression values ​​in The number of subtypes; the dissimilarity coefficient of each marker is calculated based on information entropy. Construct a comprehensive evaluation function ,in ~ These are the weighting coefficients for each factor. Let be the average correlation coefficient between marker j and other markers. To use markers to express abundance, To differentiate the ability index, the weighting coefficients of each surface marker in subtype identification were ultimately determined. ,in The total number of markers; The steps for calculating the distance score between the sample and the cluster centers of each subtype are as follows: First, the sample to be identified is standardized in the same way as the training set to ensure data distribution consistency; then, the fluorescence intensity data corresponding to the determined optimal combination of surface markers is extracted to form a feature vector; the feature vector is input into the support vector machine classification model to calculate the distance from the sample to each hyperplane; based on the calculated distance values, the Mahalanobis distance score between the sample and the cluster centers of each subtype is calculated, using the formula: Score(i) = ,in The weighting coefficient for marker j. The standardized distance difference between the sample to be tested and the cluster center of subtype i on marker j; The selected combination of surface markers includes human epidermal growth factor receptor 2, estrogen receptor, progesterone receptor, cell proliferation antigen Ki-67, and tumor stem cell marker CD44.

2. The method according to claim 1, characterized in that, The silhouette coefficient is an indicator of clustering quality. It is a standardized value calculated as the difference between the average similarity of each sample with other samples in the same cluster and the average similarity of samples in other clusters. The closer the value is to 1, the better the clustering effect. The cohesion refers to the average distance of all sample points in the same cluster to the cluster center. The smaller the value, the better the clustering effect and the more accurate the sample subtype classification.

3. The method according to claim 2, characterized in that, The K-means clustering algorithm is an unsupervised machine learning method that automatically groups samples with similar exosome fluorescence feature patterns. It iteratively optimizes and minimizes the sum of squared distances from sample points to their respective cluster centers. The input is a multidimensional fluorescence intensity matrix, and the output is the fluorescence feature patterns of different breast cancer subtypes. The Mahalanobis distance is a distance metric that considers the correlation between features. The input is the cluster centers of different subtypes, and the output is a quantitative index of the difference in expression of exosome surface markers among subtypes.

4. The method according to claim 3, characterized in that, The distance score is the Mahalanobis distance between the sample to be identified and the cluster centers of each subtype, calculated by the support vector machine classification model, and used for breast cancer subtype determination.

5. The method according to claim 4, characterized in that, The classification reliability index is calculated by dividing the distance score between the sample and the nearest subtype cluster center by the distance score between the next nearest subtype cluster center. The smaller the value, the more reliable the classification result.

Citation Information

Patent Citations

  • Method and system for identifying rare macrophage subgroups and disease markers in idiopathic pulmonary fibrosis

    CN114708918A

  • Breast cancer subtype classification method and system based on graph convolutional neural network

    CN116959572A