A method for analyzing exosome proteins in breast cancer tissue fluid
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]有鉴于此,本发明提供一种乳腺癌组织液外泌体蛋白分析的方法,能够解决现有技术中存在乳腺癌组织液外泌体蛋白质谱数据因高噪声临床样品导致蛋白丰度估计偏离真实生物状态、差异蛋白假阳性率高的技术问题
[0027]本发明提出一种乳腺癌组织液外泌体蛋白分析方法,通过构建热力学势能面蛋白丰度优化算法,将蛋白互作网络拓扑结构以谱图拉普拉斯矩阵形式内嵌于自由能损失函数的二次型系数矩阵中,同时引入香农信息熵作为系统熵项防止丰度分布极端稀疏化,并以朗之万动力学热噪声机制赋予优化轨迹跳出局部极小值的能力。传统归一化方法将各蛋白丰度独立处理,无法利用蛋白互作网络的社区结构对丰度估计施加协同约束,因而在噪声驱动的势能面局部极小值处收敛,导致假阳性差异蛋白大量出现。本发明通过拓扑约束使同一互作网络模块内的蛋白丰度协调收敛,通过退火调度控制的朗之万动力学使优化轨迹越过噪声形成的浅层势垒,最终在更接近真实生物状态的全局极小值区域输出热力学平衡态蛋白丰度矩阵。综上所述,本发明解决了背景技术中提到的乳腺癌组织液外泌体蛋白质谱数据因高噪声临床样品导致蛋白丰度估计偏离真实生物状态、差异蛋白假阳性率高的技术问题。
Smart Images

Figure CN122575468A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of breast cancer tissue fluid exosome protein analysis technology, specifically, it relates to a method for analyzing breast cancer tissue fluid exosome proteins. Background Technology
[0002] Early diagnosis and molecular subtyping of breast cancer rely on the accurate quantification of low-abundance protein biomarkers in the tumor microenvironment. Tissue fluid exosomes are rich in tumor-derived proteins and are an important source for liquid biopsies. Exosomal proteomics analysis based on data-independent acquisition mode mass spectrometry has been widely applied in breast cancer biomarker discovery research. In actual clinical sample processing, methods such as high-abundance protein depletion, multidimensional peptide fractionation, and spectral library searching are typically used to improve the identification coverage of low-abundance proteins. The protein abundance matrix is then corrected using a normalization algorithm before subsequent statistical analysis.
[0003] However, when processing batches of clinical breast cancer tissue fluid samples, the high heterogeneity of protein expression among individuals, the complexity of the sample matrix, and batch-to-batch signal drift in mass spectrometry instruments result in a large number of noise-driven anomalous abundance estimates remaining in the normalized protein abundance matrix. This noise creates numerous local minima in the high-dimensional protein abundance space, causing subsequent statistical models to tend to capture noise-driven spurious differences rather than genuine biological signals. Furthermore, existing normalization methods do not incorporate protein interaction network topological constraints into the abundance optimization process, leading to the neglect of protein abundance synergies within the same protein interaction network module, further exacerbating the false-positive problem in differential protein identification.
[0004] In other words, existing technologies have technical problems such as the estimation of protein abundance in breast cancer tissue fluid exosome protein profile data deviating from the true biological state due to high-noise clinical samples, and a high false positive rate for differentially expressed proteins. Summary of the Invention
[0005] In view of this, the present invention provides a method for protein analysis of exosomes in breast cancer tissue fluid, which can solve the technical problems in the prior art where protein abundance estimation of exosome protein profile data in breast cancer tissue fluid deviates from the true biological state due to high noise clinical samples, and the false positive rate of differential proteins is high.
[0006] This invention is achieved as follows: This invention provides a method for analyzing exosome proteins in breast cancer tissue fluid, comprising the following steps:
[0007] The tissue fluid exosome samples were subjected to high-abundance protein depletion treatment and multidimensional peptide fractionation treatment in sequence to obtain low-abundance peptide enrichment samples.
[0008] The exosome protein spectrum feature matrix was obtained by data-independent acquisition mode mass spectrometry detection of samples enriched with low-abundance peptides.
[0009] The exosomal protein spectrum feature matrix is input into the thermodynamic potential energy surface protein abundance optimization algorithm. The protein interaction network topology is embedded in the quadratic coefficient matrix of the free energy loss function using the spectrum Laplacian matrix. Shannon information entropy is introduced as the system entropy term. The Langevin dynamic annealing mechanism is used to make the optimization trajectory jump out of the local minimum, and the thermodynamic equilibrium protein abundance matrix is obtained.
[0010] The thermodynamic equilibrium protein abundance matrix, transcriptome quantification matrix, and lipidome quantification matrix are input into a multi-view spectral fusion dimensionality reduction process to obtain a low-dimensional latent factor matrix.
[0011] The exosome proteomic feature matrix, protein-protein interaction network diagram, and patient clinical multi-omics labels are input into a dynamic graph attention cross-modal recursive fusion model to output a breast cancer proteomics risk score.
[0012] During mass spectrometry detection in data-independent acquisition mode, real-time quality control is performed on the total ion current curve, the retention time drift value of the internal standard peptide, and the mass-to-charge ratio calibration drift value. When an anomaly is triggered, data-independent acquisition mode mass spectrometry detection is re-executed on samples enriched with low-abundance peptides.
[0013] The high-abundance protein depletion treatment uses a depletion kit based on the principle of isoelectric focusing enrichment. The incubation time and ionic strength of the elution buffer are determined iteratively by repeatedly experimenting on breast cancer tissue fluid standards at multiple concentration gradients, with the number of proteins and the recovery rate of low-abundance proteins as evaluation indicators.
[0014] The multidimensional peptide fractionation process involves separating the peptides after proteolytic digestion through two dimensions: strong cation exchange chromatography and reversed-phase liquid chromatography. The gradient duration and salt concentration switching nodes are determined with the goal of ensuring that the overlap rate of peptides between adjacent components is lower than the overlap rate threshold.
[0015] The data-independent acquisition mode divides the entire mass number range into fixed mass number windows, and simultaneously fragments all ions within each mass number window and acquires secondary spectra. The width and overlap of the mass number windows are determined by the number of peptides identified and the quantitative coefficient of variation as dual indicators.
[0016] The thermodynamic potential energy surface protein abundance optimization algorithm defines the system internal energy by the sum of squared abundance deviations weighted by the edge weights of the protein interaction network, defines the system entropy by the Shannon information entropy of the protein abundance distribution, constructs the free energy loss function by the difference between the system internal energy and the system entropy, and the temperature parameter decreases linearly from the highest temperature to the lowest temperature according to the annealing schedule.
[0017] The Laplacian matrix of the spectrum is the difference between the degree matrix and the adjacency matrix of the protein interaction network graph. In the thermodynamic potential energy surface protein abundance optimization algorithm, it serves as the quadratic coefficient matrix of the system's internal energy, enabling highly interconnected protein nodes within the same protein interaction network community to converge collaboratively during the abundance optimization process.
[0018] The multi-view genealogy fusion and dimensionality reduction process adopts a multi-factor analysis framework algorithm, which maps the three high-dimensional data views to low-dimensional latent factor spaces and then fuses them. Missing values are filled using a multiple imputation algorithm based on random forest, and cross-omics high correlation features are screened through sparse canonical correlation analysis.
[0019] The dynamic graph attention cross-modal recursive fusion model includes a dynamic graph attention layer, a cross-modal recursive module, a conditional jump mechanism, a topology data analysis module, and a final fusion layer. The dynamic graph attention layer consists of multiple parallel attention heads, and the attention coefficients between nodes are dynamically calculated by the cosine similarity of the protein node feature vectors.
[0020] The cross-modal recursive module uses a bidirectional long short-term memory network as its backbone. At each time step, it performs a Hadamard product-gated fusion of the current hidden state vector with the embedding vector of the number of kilobase fragments per million fragments in the transcriptome. The fusion result is then normalized and passed to the next time step.
[0021] The conditional jump mechanism calculates a dynamic threshold using a graph density adaptive attention threshold adjustment function. This function takes the current batch graph density, protein abundance Shannon information entropy, and average edge activation intensity as inputs and outputs a comprehensive adjustment index.
[0022] The graph density adaptive attention threshold adjustment function is implemented by a two-layer fully connected sub-network, with the input being the global graph density of the current batch. When the comprehensive adjustment index is in the first range, the conditional jump threshold is the 75th percentile of the attention weight distribution. When the comprehensive adjustment index is in the second range, the conditional jump threshold is the 85th percentile. When the comprehensive adjustment index reaches the upper limit threshold, the conditional jump threshold is the 95th percentile.
[0023] The topology data analysis module performs persistent cohomology calculations on the protein interaction network graph, extracts the Betty number vector composed of the 0th-order Betty number and the 1st-order Betty number as the graph-level global topology feature, and concatenates the Betty number vector with the sequence-level embedding and the conditional jump node embedding, and outputs the breast cancer proteomics risk score through the fully connected classification head.
[0024] The real-time quality control employs a real-time monitoring system based on a streaming data processing framework. This system performs sliding window statistical process control on the total ion current curve, the internal standard peptide retention time drift value, and the mass-to-charge ratio calibration drift value. The control chart types include Shewhart control charts and cumulative sum control charts.
[0025] Among them, the real-time quality control is set with multiple warning thresholds. When the first-level warning is triggered, a warning is issued. When the second-level warning is triggered, the machine is automatically stopped and the data-independent acquisition mode mass spectrometry detection is re-executed for samples enriched with low-abundance peptides. The allowable drift window for the retention time of the internal standard peptide is set within the retention time drift threshold range.
[0026] In the multidimensional peptide fractionation process, the peptide overlap threshold between adjacent components was set at 15%; in the data-independent acquisition mode mass spectrometry detection, the mass number window width was set between 4 and 25. Within the specified range; the allowable drift window for the retention time of the internal standard peptide is set between 0.1 and 1.5. Within the specified range; the conditional jump trigger rate is constrained to within the range of 20% to 40%; the protein interaction network confidence threshold is determined by grid search within the range of 0.4 to 0.9; and the multi-view phylogenetic fusion dimensionality reduction output dimension is within the range of 20 to 50 dimensions.
[0027] This invention proposes a method for analyzing exosome proteins in breast cancer tissue fluid. It constructs a thermodynamic potential surface protein abundance optimization algorithm, embedding the protein interaction network topology as a spectral Laplace matrix within the quadratic coefficient matrix of the free energy loss function. Simultaneously, it introduces Shannon information entropy as a system entropy term to prevent extreme sparsity of abundance distribution and employs Langevin dynamics thermal noise mechanism to enable the optimized trajectory to escape local minima. Traditional normalization methods treat each protein abundance independently, failing to utilize the community structure of the protein interaction network to impose cooperative constraints on abundance estimation. Consequently, they converge at noise-driven local minima, leading to a large number of false positive differential proteins. This invention uses topological constraints to coordinate the convergence of protein abundance within the same interaction network module and employs annealing-controlled Langevin dynamics to allow the optimized trajectory to overcome shallow potential barriers formed by noise, ultimately outputting a thermodynamic equilibrium protein abundance matrix in a global minima region that more closely approximates the true biological state. In summary, this invention solves the technical problems mentioned in the background art, such as the deviation of protein abundance estimation from the true biological state and the high false positive rate of differentially expressed proteins in breast cancer tissue fluid exosome protein profile data due to high-noise clinical samples. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method of the present invention.
[0029] Figure 2 This is a distribution diagram of peptides in each component.
[0030] Figure 3 Line graph for real-time monitoring of retention time drift of internal standard peptides.
[0031] Figure 4 The graph shows the change of normalized free energy during the annealing process with the number of training rounds.
[0032] Figure 5 This is a scatter plot of the spatial distribution of low-dimensional latent factors in the sample.
[0033] Figure 6 This is a dynamic graph showing the distribution of attention weights among nodes in the attention layer. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.
[0035] like Figure 1 The diagram shown is a flowchart of a method for analyzing exosome proteins in breast cancer tissue fluid provided by the present invention. This method includes the following steps:
[0036] S01. The tissue fluid exosome samples were subjected to high-abundance protein depletion treatment and multidimensional peptide fractionation treatment in sequence to obtain low-abundance peptide enrichment samples.
[0037] S02. Mass spectrometry in data-independent acquisition mode was used to detect the low-abundance peptide-enriched samples to obtain the exosome protein spectrum feature matrix.
[0038] S03. Input the exosome protein spectrum feature matrix into the thermodynamic potential energy surface protein abundance optimization algorithm to obtain the thermodynamic equilibrium protein abundance matrix.
[0039] S04. Input the thermodynamic equilibrium protein abundance matrix, transcriptome quantitative matrix and lipidome quantitative matrix into the multi-view spectral fusion dimensionality reduction process to obtain a low-dimensional potential factor matrix.
[0040] S05. Input the exosome proteomic feature matrix, protein-protein interaction network diagram and patient clinical multi-omics labels into the dynamic graph attention cross-modal recursive fusion model to output a breast cancer proteomics risk score.
[0041] S06. During the data-independent acquisition mode mass spectrometry detection process in step S02, real-time quality control is performed on the total ion current curve, the internal standard peptide retention time drift value, and the mass-to-charge ratio calibration drift value. If an abnormality is triggered, step S02 is re-executed on the low-abundance peptide enrichment sample.
[0042] The high-abundance protein depletion treatment refers to using a depletion kit based on the principle of isoelectric focusing enrichment to reduce the concentration of high-abundance proteins in tissue fluid exosome samples, thereby releasing the signals of low-abundance tumor marker protein peptides from mass spectrometry noise. The incubation time and ionic strength of the elution buffer for the depletion treatment are determined by repeated experiments on 5 to 20 breast cancer tissue fluid standards with different concentration gradients, with the final number of identified proteins and the recovery rate of low-abundance proteins as evaluation indicators, after no less than 3 rounds of iterative optimization.
[0043] The multidimensional peptide fractionation process refers to separating the peptides after proteolytic digestion through two dimensions: strong cation exchange chromatography and reversed-phase liquid chromatography. This disperses the peptides into complementary two-dimensional retention spaces based on their charge and hydrophobicity. The gradient duration and salt concentration switching points for strong cation exchange chromatography and reversed-phase liquid chromatography were determined through preliminary experiments on 6–12 representative breast cancer tissue fluid samples, with the optimization target being a peptide overlap rate of less than 15% between adjacent components. Strong cation exchange chromatography is a liquid chromatography technique that separates peptides based on the difference in the amount of net positive charge they carry under acidic conditions. Reversed-phase liquid chromatography is a liquid chromatography technique that separates peptides on a nonpolar stationary phase based on the difference in their hydrophobicity.
[0044] The data-independent acquisition mode refers to a mass spectrometry acquisition mode in which the entire mass number range is divided into fixed mass number windows during the acquisition process. All ions within each mass number window are simultaneously fragmented and acquired via secondary spectrometry, and peptide identification is subsequently completed through library retrieval. The width and overlap of the mass number windows were determined through method development experiments on 30–50 exosome samples, using the number of peptides identified and the coefficient of variation for quantification as dual indicators. After repeated experiments, the values were set between 4 and 25. Within the range.
[0045] The thermodynamic potential energy surface protein abundance optimization algorithm treats the normalized exosome protein spectrum feature matrix as a high-dimensional potential energy surface system, where the abundance vector of each protein corresponds to a particle coordinate on the potential energy surface. The algorithm defines the system's internal energy as the sum of squared abundance deviations weighted by the edge weights of the protein interaction network, and the system entropy as the Shannon information entropy of the protein abundance distribution. A differentiable free energy loss function is constructed based on the difference between the system's internal energy and its entropy. This free energy loss function is expressed as follows: ;in For the system free energy, For the internal energy of the system, For temperature parameters, For system entropy, The reference normalized quantity of the system free energy obtained from the statistics of the training dataset. The temperature reference normalization quantity is obtained statistically from the training dataset. The system entropy reference normalization quantity is obtained from the statistical analysis of the training dataset; the topology of the protein interaction network is encoded as a quadratic coefficient matrix of the system's internal energy through the spectral Laplacian matrix; a Langevin dynamic thermal noise term is introduced during the gradient descent process, with temperature parameters... According to the annealing schedule linearly decreasing to and Annealing curve scanning experiments were conducted on no fewer than 20 breast cancer exosome samples. The false positive rate of differentially expressed proteins and the optimization convergence speed were used as dual evaluation indicators, determined after multiple rounds of comparative experiments. The particle coordinate output after optimization convergence is a thermodynamic equilibrium protein abundance matrix. The thermodynamic potential energy surface protein abundance optimization algorithm embeds the protein interaction network topology into the quadratic coefficient matrix of the free energy loss function, ensuring that the protein abundance within the same protein interaction network module is topologically constrained during optimization, thus coordinating convergence. The introduction of the system entropy term prevents the optimization from tending towards extremely sparse solutions, preserving the biologically reasonable diversity of protein abundance distribution. The Langevin dynamics thermal noise mechanism endows the optimization trajectory with random perturbation capability, thereby finding a global minimum region closer to the real biological state in the local minimum terrain of the potential energy surface caused by high-noise clinical samples, reducing noise-driven false positives of differentially expressed proteins, and improving the signal reliability of the subsequent multi-view spectral fusion and dimensionality reduction processing.
[0046] The spectral Laplacian matrix is the difference between the degree matrix and the adjacency matrix of the protein interaction network graph. Its eigenvalues and eigenvectors encode the global connectivity and community structure of the protein interaction network. In the thermodynamic potential surface protein abundance optimization algorithm, the spectral Laplacian matrix is used as the quadratic coefficient matrix to construct the system's internal energy term. This allows highly interconnected protein nodes within the same protein interaction network community to be mutually constrained during the abundance optimization process, and to converge collaboratively to a thermodynamic equilibrium state.
[0047] The Shannon information entropy is a statistic that measures the dispersion of the protein abundance probability distribution. The more uniform the protein abundance distribution, the higher the Shannon information entropy. The more concentrated the protein abundance is in a few proteins, the lower the Shannon information entropy. In the thermodynamic potential energy surface protein abundance optimization algorithm, the Shannon information entropy is introduced as a system entropy term into the free energy loss function to prevent the optimization process from compressing the protein abundance vector into an extremely sparse distribution, thereby preserving the reasonable diversity of protein abundance in biological samples.
[0048] The Langevin dynamics described herein involves superimposing a Gaussian-distributed random noise term onto the parameter updates during gradient descent iterations. The noise intensity is related to the temperature parameter. It is proportional to the annealing schedule and decreases with the annealing schedule. In the high-temperature stage, the large random noise enables the optimization trajectory to cross the shallow potential barrier and escape the local minima caused by batch effect or biological noise. In the low-temperature stage, the noise intensity decreases and the optimization trajectory converges to the minimum region in the current basin, and finally outputs the thermodynamic equilibrium protein abundance matrix.
[0049] The multi-view genealogy fusion and dimensionality reduction process employs a multi-factor analysis framework algorithm. It maps three high-dimensional data views—the thermodynamic equilibrium protein abundance matrix, the transcriptome quantification matrix, and the lipidome quantification matrix—to low-dimensional latent factor spaces of 20–50 dimensions, then fuses them to output a low-dimensional latent factor matrix. The transcriptome quantification matrix uses the number of kilobase fragments per million fragments as the quantification unit. Missing values in the three high-dimensional data views are filled using a random forest-based multiple imputation algorithm. Cross-omics high-association features are further screened using sparse canonical correlation analysis. The matrix factorization process is accelerated using a graphics processor. The number of trees and the maximum number of iterations in the random forest-based multiple imputation algorithm are determined by performing 5-fold cross-validation on a complete dataset with artificially missing data, with the goal of minimizing the root mean square error. The regularization penalty coefficient of the sparse canonical correlation analysis is determined by applying the penalty coefficient to... ~ A grid search is performed within the range, and the result is determined by combining it with 10-fold cross-validation.
[0050] The specific structure of the dynamic graph attention cross-modal recursive fusion model is as follows: the model input layer receives three data streams: an exosome protein profile feature matrix, a protein interaction network graph, and patient clinical multi-omics labels. The protein interaction network graph is input in the form of an adjacency matrix, with node features being the normalized abundance vectors of each protein. The patient clinical multi-omics labels include gene mutation status vectors and the embedding vector of the number of kilobase fragments per million fragments in the transcriptome. The protein interaction network graph first enters the dynamic graph attention layer, which consists of four parallel attention heads. The attention coefficients between nodes in each attention head are dynamically calculated based on the cosine similarity of the protein node feature vectors in the current batch, and the edge weights are co-expressed with the proteins. The correlation is adaptively updated in each forward propagation; the outputs of the four parallel attention heads are concatenated and linearly transformed to obtain the protein node embedding sequence, which has a 128-dimensional dimension; the protein node embedding sequence is arranged according to the pathway modules in the protein interaction network and then sent to the cross-modal recursive module; the cross-modal recursive module is based on a bidirectional long short-term memory network and has three layers, each with a hidden state dimension of 256 dimensions; at each time step, the current hidden state vector is fused with the synchronously input embedding vector of the number of kilobase fragments per million fragments in the transcriptome using a Hadamard product gate, and the fusion result is normalized and passed to the next time step; the dynamic graph attention cross-modal recursive fusion model has a conditional jump mechanism, which allows a certain protein node to jump when a condition is met. When the attention weight of a node exceeds a dynamic threshold, the protein node embedding bypasses the current recursive layer and jumps directly to the final fusion layer. The dynamic threshold is calculated and output by a graph density adaptive attention threshold adjustment function. This function is implemented by a two-layer fully connected subnetwork with an input dimension of 1, an output dimension of 1, and a hidden layer dimension of 16. The input is the global graph density of the current batch. The topology data analysis module independently performs persistent cohomology calculations on the protein interaction network graph, extracting the Betti number vector as a graph-level global topology feature. The Betti number vector has a length of 64 dimensions and is composed of 0th-order and 1st-order Betti numbers. The final fusion layer integrates the sequence-level data output from the cross-modal recursive module. After embedding, conditional jump node embedding, and concatenation of Betty number vectors, a breast cancer proteomics risk score is output through a 2-layer fully connected classification head. The dimensions of the 2-layer fully connected classification head are 512 and 128, respectively. In the memory allocation strategy of the dynamic graph attention cross-modal recursive fusion model, the attention coefficient matrix of the dynamic graph attention layer is stored in an independent CUDA stream A, the hidden states of each layer of the bidirectional long short-term memory network are allocated to CUDA stream B, and the forward propagation of the 2-layer fully connected sub-network is allocated to CUDA stream C. The three CUDA streams are executed in parallel. The intermediate activation values of the cross-modal recursive module are stored in mixed precision, the forward propagation uses a 16-bit floating-point format, and the backpropagation gradient uses a 32-bit floating-point format.Persistent co-homology computation is allocated to the central processing unit (CPU) memory and executed asynchronously with the graphics processing unit (GPU), accelerating data transfer through a fixed memory mechanism. The adjacency matrices of different samples within a batch are stored in GPU memory using a sparse tensor format. The thread block size of each CUDA stream is determined to be within the range of 128–256 threads per thread block after benchmark testing of the target GPU's stream multiprocessor utilization. The dynamic graph attention cross-modal recursive fusion model uses a dynamic graph attention mechanism to adaptively adjust the edge weights of the protein interaction network with the sample batch during training, overcoming the limitation of static graph convolution in capturing dynamic changes in protein co-expression. Hadamard product... The gated fusion mechanism enables fine-grained interaction between protein abundance information and transcriptional regulatory information in the temporal dimension, allowing protein abundance temporal features and transcriptional regulatory signals to mutually constrain each other within the same feature dimension space. The conditional jump mechanism shortens the gradient propagation path for high-attention-weighted protein nodes, accelerating convergence while preserving accurate gradient information for key protein nodes. The introduction of Betty number vectors encodes the global algebraic topology of the protein-protein interaction network as a supplementary basis for classification decisions, enabling the model to perceive changes in the overall connectivity of the protein-protein interaction network, thus improving the overall biological interpretability and clinical predictive accuracy of breast cancer proteomics risk scoring.
[0051] The steps for establishing the training dataset for the dynamic graph attention cross-modal recursive fusion model specifically include: collecting tissue fluid exosome mass spectrometry data from no less than 200 pathologically confirmed breast cancer patients and no less than 100 healthy controls as training samples for the exosome protein spectrum feature matrix; extracting the transcriptome quantification matrix from the same batch of RNA sequencing data; downloading the protein interaction network graph from the STRING database and filtering it according to a confidence threshold, which is determined by performing a grid search within the range of 0.4 to 0.9 on the pre-experimental dataset using the area under the curve of the model validation set as an indicator; combining the patient's clinical multi-omics labels with the gene mutation status in the clinical gene testing report and RNA sequencing data; dividing all data into training, validation, and test sets in a 7:2:1 ratio after unified Z-score normalization; stratified random sampling of the training set according to breast cancer subtypes to ensure that the sample proportion of each subtype is consistent with the clinical epidemiological proportion.
[0052] The specific steps for training the dynamic graph attention cross-modal recursive fusion model include: using the AdamW optimizer, with the initial learning rate set at... ~ Within the specified range, cosine annealing is used for learning rate scheduling, with each annealing cycle consisting of 50 training epochs. The batch size is set to 16–32 samples, with the specific value determined after benchmark testing of memory usage on the target graphics processor. The loss function is the sum of the binary cross-entropy loss and the graph attention weight regularization term. The regularization coefficient of the graph attention weight regularization term is determined by… ~ The area under the curve (AUC) is determined after scanning the validation set within the specified range. During training, the AUC is evaluated on the validation set every 5 rounds. If the AUC does not increase for 20 consecutive rounds, early stopping is triggered. The adjacency matrix of the dynamic graph attention layer is recalculated based on the protein co-expression matrix of the current batch samples at the start of the forward propagation of each batch. The global graph density of the current batch is calculated by the proportion of non-zero elements in the adjacency matrix of the current batch and is used as the input to the graph density adaptive attention threshold adjustment function. The model ultimately uses the checkpoint weight with the highest AUC on the validation set as the optimal weight.
[0053] The graph density adaptive attention threshold adjustment function is used to adjust the dynamic threshold of the conditional jump mechanism in the dynamic graph attention cross-modal recursive fusion model; the input of the graph density adaptive attention threshold adjustment function includes the graph density of the current batch of protein interaction networks. Current batch protein abundance Shannon information entropy The average edge activation intensity of the current batch of dynamic graph attention layer output The output is the comprehensive adjustment index. The graph density adaptive attention threshold adjustment function is expressed as follows: ;in The normalized value is used as a reference for the comprehensive adjustment index. The average graph density of the training set, The average Shannon information entropy of the training set. The average edge activation strength of the training set. , , The weighting coefficients are weighted and their sum is 1. These weighting coefficients are determined by performing a grid search on the validation set with a conditional jump trigger rate constrained to be between 20% and 40%. When, the conditional jump threshold is set to the 75th percentile of the attention weight distribution; when When, the conditional jump threshold is set to the 85th percentile of the attention weight distribution; when When the conditional jump threshold is set to the 95th percentile of the attention weight distribution, the 75th percentile, 85th percentile, and 95th percentile are determined by conducting multiple rounds of ablation experiments on no less than 30 batches of validation set data, with the harmonic mean of differential protein recognition recall and precision as the optimization target.
[0054] The real-time quality control employs a real-time monitoring system based on a streaming data processing framework, performing sliding window statistical process control on the total ion current curve, internal standard peptide retention time drift, and mass-to-charge ratio calibration drift. The control chart types for this sliding window statistical process control include Shewhart control charts and cumulative sum control charts. Among the multi-level warning thresholds, the first and second level warning thresholds are determined by performing normality tests and quantile statistics on each quality control indicator from at least 50 batches of historical mass spectrometry operation data, using three times the standard deviation as a reference and correcting for these thresholds by combining the linear fitting results of the intra-batch sensitivity decay rate. The allowable drift window for the internal standard peptide retention time is set within the range of 0.1–1.5 min after statistical analysis of at least 20 repeated injection experiments on the chromatographic column. A warning is issued when a first-level warning is triggered, and the system automatically shuts down and re-executes step S02 for the low-abundance peptide enrichment sample when a second-level warning is triggered.
[0055] The Hadamard product gated fusion refers to the operation of multiplying the corresponding elements of two vectors with the same dimension one by one to obtain a fusion vector with the same dimension as the input vector. Compared with vector concatenation, Hadamard product gated fusion realizes the direct modulation of protein abundance temporal features and transcriptional regulatory signals in each dimension, so that the two types of signals interact in the same feature dimension space, rather than simply expanding the feature dimension.
[0056] The Betti number vector is an invariant in algebraic topology used to describe the spatial connectivity structure; the 0th-order Betti number reflects the number of connected components in the protein-protein interaction network, and the 1st-order Betti number reflects the number of independent loops in the protein-protein interaction network; the Betti number vector is output by persistent cohomology analysis after calculating the protein-protein interaction network graph at different filtering scales, and can capture the global structural change information of the protein-protein interaction network at different filtering scales.
[0057] Optionally, the present invention also provides a computer-based method for forming a breast cancer tissue fluid exosome protein analysis system, wherein the computer is provided with a readable storage medium storing program instructions, which, when executed in the computer, are used to perform the above-described method.
[0058] The specific implementation of step S01 is as follows: First, the collected breast cancer patient tissue fluid is pretreated by centrifugation to remove cell debris, and then exosome particles are extracted. After protein quantification, the obtained tissue fluid exosome sample is incubated using a high-abundance protein depletion kit based on the isoelectric focusing enrichment principle. This utilizes the difference in isoelectric point to selectively retain low-abundance protein components, thereby significantly reducing the concentration of high-abundance proteins such as albumin and immunoglobulins, releasing the signals of low-abundance tumor marker protein peptides from the mass spectrometry noise background. The incubation time and ion strength of the elution buffer are determined through at least three rounds of iterative optimization experiments on 5–20 breast cancer tissue fluid standards of different concentration gradients. The evaluation indicators are the final number of identified proteins and the recovery rate of low-abundance proteins. After depletion, the sample was enzymatically digested and then fractionated into peptides using both strong cation exchange chromatography and reversed-phase liquid chromatography. Strong cation exchange chromatography separated peptides based on the difference in the amount of net positive charge they carried under acidic conditions, while reversed-phase liquid chromatography separated peptides based on the difference in their hydrophobicity on a nonpolar stationary phase. The complementary retention mechanisms of these two dimensions ensured that the peptides were fully dispersed in a two-dimensional space, with an overlap rate of less than 15% between adjacent components, ultimately yielding a low-abundance peptide-enriched sample.
[0059] The specific implementation of step S02 is as follows: For samples enriched with low-abundance peptides, mass spectrometry detection is performed using a data-independent acquisition mode. This acquisition mode divides the entire mass number range into several fixed-width mass number windows. All ions within each window are simultaneously subjected to fragmentation and secondary spectrum acquisition. Peptide identification is then completed through library retrieval. The width and overlap of the mass number windows were determined through method development experiments on 30–50 exosome samples. Using the number of identified peptides and the coefficient of variation for quantification as dual indicators, the values were repeatedly tested and set between 4 and 25. Within the specified range. During the detection process, the real-time quality control system operates synchronously, performing sliding window statistical process control on the total ion current curve, internal standard peptide retention time drift value, and mass-to-charge ratio calibration drift value. Control chart types include Shewhart control charts and cumulative sum control charts. Multi-level warning thresholds are determined by normality testing and quantile statistics from no fewer than 50 batches of historical mass spectrometry data, using 3 times the standard deviation as a reference benchmark and corrected by combining the linear fitting results of the intra-batch sensitivity decay rate. The allowable drift window for the internal standard peptide retention time is set between 0.1 and 1.5. Within the specified range, a warning is issued when a Level 1 warning is triggered, and the system automatically stops and re-executes this step when a Level 2 warning is triggered, ultimately outputting the exosome protein profile feature matrix.
[0060] The specific implementation of step S03 is as follows: The normalized exosome protein spectrum feature matrix is regarded as a high-dimensional potential energy surface system, and the abundance vector of each protein corresponds to a particle coordinate on the potential energy surface. The topology of the protein interaction network is encoded by the spectral Laplacian matrix, which is the difference between the network degree matrix and the adjacency matrix. This difference is used as the quadratic coefficient matrix to construct the internal energy term of the system, so that highly interconnected protein nodes within the same protein interaction network community are mutually constrained in abundance optimization, resulting in coordinated convergence. The system entropy is defined by the Shannon information entropy of the protein abundance distribution to prevent the optimization process from compressing the abundance vector into an extremely sparse distribution. The free energy loss function is expressed as... ,in For the system free energy, For the internal energy of the system, For temperature parameters, For system entropy, , , This is a reference normalization quantity obtained statistically from the training dataset. A Langevin dynamic thermal noise term is introduced during gradient descent, with temperature parameters... According to the annealing schedule linearly decreasing to During the high-temperature stage, the large random noise causes the optimized trajectory to jump out of the local minimum by crossing the shallow potential barrier. During the low-temperature stage, the noise intensity decreases and converges to the minimum region, finally outputting the thermodynamic equilibrium protein abundance matrix.
[0061] The specific implementation of step S04 is as follows: Three high-dimensional data views—the thermodynamic equilibrium protein abundance matrix, the transcriptome quantitative matrix, and the lipidome quantitative matrix—are input into a multi-factor analysis framework algorithm, mapped to a 20-50 dimension low-dimensional latent factor space, and then fused. The transcriptome quantitative matrix uses the number of kilobase fragments per million fragments as the quantitative unit. Missing values in the three views are filled using a random forest-based multiple imputation algorithm. The number of trees and the maximum number of iterations are determined through 5-fold cross-validation with the goal of minimizing the root mean square error. Highly correlated features across omics are further screened using sparse canonical correlation analysis, and the regularization penalty coefficient is... ~ The latent factor matrix is determined by performing a grid search within the specified range combined with 10-fold cross-validation. The matrix factorization process is accelerated using a graphics processor, ultimately outputting a low-dimensional latent factor matrix.
[0062] The specific implementation of step S05 is as follows: The input layer of the dynamic graph attention cross-modal recursive fusion model receives three data streams: exosome proteomic feature matrix, protein interaction network graph, and patient clinical multi-omics labels. The protein interaction network graph is input in the form of an adjacency matrix, and the node features are the normalized abundance vectors of each protein. The dynamic graph attention layer consists of four parallel attention heads. The attention coefficients between nodes are dynamically calculated by the cosine similarity of the feature vectors of the current batch of protein nodes. The edge weights are adaptively updated in each forward propagation based on the correlation of protein co-expression. The outputs of the four attention heads are concatenated and linearly transformed to obtain a 128-dimensional protein node embedding sequence. The protein node embedding sequence is arranged in the order of pathway modules and then fed into a cross-modal recursive module with a bidirectional long short-term memory network as its backbone. There are three layers in total, and the hidden state dimension of each layer is 256 dimensions. At each time step, the current hidden state vector is fused with the embedding vector of the number of kilobase fragments per million fragments in the transcriptome using a Hadamard product gate. The fusion result is normalized by the layer and then passed to the next time step. When the attention weight of a protein node exceeds the dynamic threshold output by the graph density adaptive attention threshold adjustment function, the node embedding bypasses the current recursive layer and jumps directly to the final fusion layer. The topology data analysis module performs persistent cohomology calculation on the protein interaction network graph and extracts a 64-dimensional Betty number vector. The final fusion layer concatenates the sequence-level embeddings, conditional jump node embeddings, and Betty number vectors, and outputs a breast cancer proteomics risk score through a two-layer fully connected classification head with dimensions of 512 and 128, respectively. The model training uses the AdamW optimizer, the loss function is the sum of the binary cross-entropy loss and the graph attention weight regularization term, and the learning rate is scheduled using a cosine annealing strategy.
[0063] It should be noted that the key technologies of this invention include: the thermodynamic potential surface protein abundance optimization algorithm integrates the topological constraints of the protein interaction network, Shannon information entropy regularization, and Langevin dynamic annealing mechanism into a unified free energy loss function framework. The three interact during the optimization process: topological constraints prevent protein abundance within the same community from drifting independently; Shannon information entropy prevents the abundance distribution from converging to an extremely sparse state that is biologically unreasonable; and Langevin dynamics provides random perturbations to the optimization trajectory at high temperatures to overcome the shallow potential barrier formed by noise, and converges to a reasonable region jointly defined by topological constraints and entropy regularization at low temperatures. The synergy of the three improves the optimization results simultaneously in three dimensions: biological rationality, network consistency, and noise robustness. This provides a low false positive and high reliability input protein abundance matrix for the subsequent dynamic graph attention cross-modal recursive fusion model, making the breast cancer proteomics risk score output by the model more biologically interpretable.
[0064] It is important to note that in breast cancer tissue fluid exosomal proteomics studies, when samples originate from multiple collection batches and exhibit significant inter-individual molecular heterogeneity, the abundance of proteins within the same functional module of the protein interaction network often shows a synergistic variation across different samples. However, existing mass spectrometry data processing methods independently normalize the abundance of each protein, failing to detect this synergy within the network. This results in contradictory anomalies in the abundance estimates of different proteins within the same functional module in the normalized abundance matrix, leading to numerous false positive results in subsequent differential protein analyses. The reason for this technical problem is that the mass spectrometry signal itself is subject to the combined interference of instrument batch drift, sample matrix effects, and individual heterogeneity. The objective function of traditional normalization methods only minimizes the signal intensity deviation of each protein itself, without incorporating the physical interactions between proteins into the optimization constraints. Therefore, the optimization results are insensitive to the synergy within the network and are prone to converge to local minima constructed by instrument noise or batch effects in high-noise clinical samples, rather than global minima that reflect the true biological state. To address the aforementioned technical issues, a common solution is to introduce batch effect correction algorithms, such as `removeBatchEffect` in Combat or Limma, which use a linear mixture model to regress batch variables from the protein abundance matrix. However, these methods assume that batch effects are uniformly distributed among all proteins and linearly separable from biological signals. Given the nonlinear topology of protein-protein interaction networks, linear batch correction cannot distinguish between co-abundance variations within the network and batch noise. The correction process may simultaneously eliminate some genuine network co-abundance signals, thereby reducing the consistency of protein abundance within network modules, exacerbating the fragmentation of network functional modules in subsequent analyses, and further generating false positives. This invention effectively solves this technical problem. The thermodynamic potential energy surface protein abundance optimization algorithm uses the Laplace matrix of the protein interaction network's spectral graph as the quadratic coefficient matrix of the system's internal energy, directly encoding the network topology into the optimization objective. This minimizes the system's internal energy while simultaneously requiring coordinated adjustment of protein abundance vectors within the same network community along the direction of topological constraints, rather than independent drift. Shannon information entropy, as the system entropy term, competes with the system's internal energy in the free energy loss function, preventing excessive compression of the abundance distribution to an extremely sparse state by topological constraints. The Langevin dynamic thermal noise mechanism, under annealing scheduling control, ensures that the optimization trajectory has sufficient random perturbation capability at high temperatures to overcome shallow barriers constructed by instrument noise or batch effects, and accurately converges to the global minimum region jointly defined by network topology and entropy regularization at low temperatures. This region physically corresponds to a thermodynamic equilibrium state and biologically corresponds to the real biological state of consistent abundance within the protein interaction network community, fundamentally avoiding the destruction of network cooperative signals by the linear assumption in traditional batch correction methods.
[0065] Specifically, the principle of this invention is:
[0066] The core reason why this invention can solve the above-mentioned technical problems is that it combines the thermodynamic statistical physics framework with the topological constraints of protein interaction networks, fundamentally changing the design logic of the objective function for protein abundance optimization.
[0067] In traditional proteomics data processing, protein abundance normalization is typically performed independently for each protein, with no coupling between the abundance estimates of different proteins. This design is acceptable in low-noise samples, but in high-noise clinical samples, isolated normalization cannot distinguish between real biological changes and noise perturbations. This causes the optimization process to converge to a large number of local minima constructed by batch effects or biological noise, ultimately outputting an abundance matrix that deviates from the true biological state.
[0068] This invention treats the normalized protein spectrum feature matrix as a high-dimensional potential energy surface system, with the abundance vector of each protein corresponding to a particle coordinate on the potential energy surface. The internal energy of the system is defined by the sum of squared abundance deviations weighted by the edge weights of the protein interaction network, where the topology of the protein interaction network is encoded as a quadratic coefficient matrix of the internal energy through the spectral Laplacian matrix. The spectral Laplacian matrix is the difference between the network degree matrix and the adjacency matrix, and its eigenvalues and eigenvectors encode the global connectivity and community structure of the network. Using it as a quadratic coefficient matrix implies that highly interconnected protein nodes are mutually constrained in abundance optimization, and the protein abundance within the same community is forced to converge in a coordinated manner. This is physically equivalent to the collective behavior of interacting particles tending towards energy minimization, and biologically corresponds to the synergistic regulation of protein expression within the same pathway, thus possessing sufficient biological rationale.
[0069] The system entropy is defined by the Shannon information entropy of the protein abundance distribution. A higher Shannon information entropy indicates a more uniform abundance distribution. After introducing the entropy term, the free energy loss function needs to minimize the internal energy while maximizing the entropy, thereby preventing the optimization result from compressing the abundance vector into an extremely sparse distribution. This preserves the reasonable diversity of protein abundance in biological samples and avoids biologically unreasonable extreme solutions caused by simply minimizing the internal energy.
[0070] Langevin dynamics superimposes a Gaussian-distributed random noise term during gradient descent iterations, with noise intensity proportional to the temperature parameter. Annealing scheduling linearly decreases the temperature from high to low: the larger random noise at high temperatures allows the optimized trajectory to overcome shallow barriers and escape local minima caused by batch effects or biological noise; at low temperatures, the noise intensity decreases, and the optimized trajectory converges to the minimum region within the current basin. This mechanism physically simulates the dynamic behavior of atoms relaxing from a high-energy disordered state to a low-energy ordered state during solid annealing, making the final thermodynamic equilibrium protein abundance matrix closer to the real biological state. This effectively reduces the false positive rate of noise-driven differential proteins, providing a reliable input signal for subsequent multi-view phylogenetic fusion dimensionality reduction and dynamic graph attention cross-modal recursive fusion models.
[0071] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0072] The specific implementation of step S01 is as follows: The tissue fluid exosome sample is sequentially subjected to high-abundance protein depletion treatment and multidimensional peptide fractionation treatment to obtain a low-abundance peptide-enriched sample. The high-abundance protein depletion treatment uses a depletion kit based on the isoelectric focusing enrichment principle to reduce the concentration of high-abundance proteins such as albumin and immunoglobulins in the sample, thereby releasing the signals of low-abundance tumor marker protein peptides from mass spectrometry noise. The incubation time and elution buffer ionic strength are determined through repeated experiments on 5–20 breast cancer tissue fluid standards of different concentration gradients, with the final number of identified proteins and the recovery rate of low-abundance proteins as evaluation indicators, after at least three rounds of iterative optimization. Multidimensional peptide fractionation involves separating the proteolytic peptides sequentially using strong cation exchange chromatography (SCAC) and reversed-phase liquid chromatography (RPLC). SCAC separates peptides based on the difference in the net positive charge carried by the peptides under acidic conditions, while RLC separates them based on the difference in hydrophobicity of the peptides on a nonpolar stationary phase. The gradient duration and salt concentration switching points for the two chromatographic methods were determined through preliminary experiments on 6–12 representative breast cancer tissue fluid samples, with the optimization target being a peptide overlap rate of less than 15% between adjacent components.
[0073] The specific implementation of step S02 is as follows: Data-independent acquisition mode mass spectrometry is used to detect low-abundance peptide-enriched samples to obtain the exosome protein spectral feature matrix. The data-independent acquisition mode divides the entire mass number range into fixed mass number windows, and all ions within each window are simultaneously subjected to fragmentation and secondary spectral acquisition. Peptide identification is then completed through library retrieval. The width and overlap of the mass number window were determined through method development experiments on 30–50 exosome samples, using the number of identified peptides and the coefficient of variation for quantification as dual indicators. After repeated experiments, the range was set between 4 and 25. Within the range. Exosomal protein profile feature matrix middle, Number of samples To detect protein count, matrix elements For the first The first sample Quantitative values of mass spectrometry peak area for each protein. The value range is 1~ , The value range is 1~ .
[0074] The specific implementation of step S03 is as follows: The exosome protein spectrum feature matrix is input into the thermodynamic potential surface protein abundance optimization algorithm to obtain the thermodynamic equilibrium protein abundance matrix. Firstly, the... Proceed by column Fraction normalization yields a normalized matrix. Each element is:
[0075] ;
[0076] In the formula, Normalized matrix The Middle The first sample Dimensionless normalized abundance values of each protein For the first Each protein in all The mean of the quantitative values of mass spectrometry peak areas in each sample. For the first Each protein in all The standard deviation of the quantitative values of mass spectrometry peak areas in each sample , , If the units are the same, after division It is a dimensionless quantity. The protein-protein interaction network is represented as a weighted undirected graph. , A set of protein nodes. Let be the set of edges. This is the edge weight matrix. The spectral graph Laplacian matrix. Defined as:
[0077] ;
[0078] In the formula, This is the adjacency matrix of a protein-protein interaction network, with elements... For protein With protein The interaction confidence weights are dimensionless. The value range is 1~ For a degree matrix, its diagonal elements Off-diagonal elements are 0. It is a dimensionless matrix. The internal energy of the system. Defined as the weighted sum of squared abundance deviations with the Laplace matrix of the spectrum as the coefficient matrix of the quadratic form, the formula is as follows:
[0079] ;
[0080] In the formula, For the first The normalized protein abundance column vector of each sample, i.e. The Line transpose, superscript Indicates transpose. On the training dataset The statistical mean is used as a reference normalized quantity for the system's internal energy, making This is a dimensionless ratio. System entropy. Using the Shannon information entropy definition, the protein abundance probability distribution for each sample is as follows:
[0081] ;
[0082] In the formula, For the first The first sample Normalized abundance probability of each protein, dimensionless. For the first The first sample The normalized abundance values of each protein. The Shannon information entropy of each sample is The system entropy of the entire dataset is taken as the mean. , On the training dataset The statistical mean is used as the reference normalization quantity for system entropy, making This is a dimensionless ratio. The formula for the free energy loss function is as follows:
[0083] ;
[0084] In the formula, For the system free energy, For temperature parameters, The statistical mean of the system free energy on the training dataset is used as the reference normalization quantity for the system free energy. The statistical mean of temperature parameters on the training dataset is used as a temperature reference normalization measure. , , All are dimensionless ratios, with consistent dimensions for each term. The free energy loss function is related to... The partial derivatives are expanded as follows:
[0085] ;
[0086] In the formula, Shannon's information entropy about The gradient vector, whose th... Each component is ,in For the first The sum of the normalized abundance of each sample For the first The first sample Normalized abundance probability of each protein. For protein indexing. A Langevin kinetic thermal noise term is introduced during gradient descent. The step parameter update rule is as follows:
[0087] ;
[0088] In the formula, The learning rate is an empirical value. , For a random noise vector that follows a standard Gaussian distribution, It is the identity matrix. For the first Step temperature parameters, according to annealing schedule from linearly decreasing to The annealing rules are as follows:
[0089] ;
[0090] In the formula, This is the annealing initiation temperature. This is the annealing termination temperature. This represents the total number of iterations. This represents the current iteration step. and Annealing curve scanning experiments were performed on no fewer than 20 breast cancer exosome samples. The false positive rate of differentially expressed proteins and the optimized convergence speed were used as dual evaluation indicators, determined after multiple rounds of comparative experiments. The particle coordinate output after optimized convergence is a thermodynamic equilibrium protein abundance matrix. .
[0091] The specific implementation of step S04 is as follows: Thermodynamic equilibrium protein abundance matrix... Transcriptome Quantitative Matrix ( (This refers to the number of genes, with the quantitative unit being the number of kilobase fragments per million fragments) and the lipidome quantitative matrix. ( The number of lipid types was used as input for multi-factor analysis, and dimensionality reduction was performed through multi-view phylogenetic fusion. Missing values in the three data views were first filled using a random forest-based multiple imputation algorithm. The number of trees and the maximum number of iterations were determined by performing 5-fold cross-validation on a complete dataset with artificially missing values, with the goal of minimizing the root mean square error. Highly correlated features across omics were screened using sparse canonical correlation analysis, and the regularization penalty coefficient was... ~ A grid search is performed within the specified range, followed by 10-fold cross-validation to determine the optimal approach. The multi-factor analysis framework algorithm maps the three data views to a low-dimensional latent factor space, outputting a low-dimensional latent factor matrix. , The potential factor dimension has a value range of 20 to 50, and the matrix factorization process is accelerated using a graphics processor.
[0092] The specific implementation of step S05 is as follows: The exosome proteomic feature matrix, protein-protein interaction network graph, and patient clinical multi-omics labels are input into a dynamic graph attention cross-modal recursive fusion model, which outputs a breast cancer proteomics risk score. The model input layer receives three data streams: a protein-protein interaction network graph and an adjacency matrix. The input format is as follows: node features are normalized abundance vectors of each protein; patient clinical multi-omics labels include gene mutation status vectors. ( This represents the number of mutation sites. Each component is either 0 or 1, indicating whether a mutation exists at the corresponding site, and the embedding vector is the number of kilobase fragments per million fragments in the transcriptome. ( (For the embedding dimension). The dynamic graph attention layer consists of 4 parallel attention heads, the first of which is the first one. Nodes in each attention head With nodes Attention coefficient between The cosine similarity is dynamically calculated, and the formula is expressed as follows:
[0093] ;
[0094] In the formula, This is the attention head index, with values of 1, 2, 3, and 4. For the first Nodes in each attention head eigenvectors, For attention head feature dimension, For nodes The set of neighboring nodes, For the neighborhood node index, The cosine similarity function is used. through After normalization, the values are dimensionless probability values. The edge weights are adaptively updated in each forward propagation based on protein co-expression correlation, with the update rule being:
[0095] ;
[0096] In the formula, For the forward propagation count index, For the first During the next forward propagation The weights are dimensionless. For learnable weight row vectors, For learnable bias scalars, for Activation function For Hadama accumulation, For the first Second forward propagation node eigenvectors, For the first Second forward propagation node The feature vectors. The outputs of the four attention heads are concatenated and then linearly transformed. Obtain the protein node embedding vector:
[0097] ;
[0098] In the formula, For nodes protein node embedding vectors, This is a concatenated vector of the outputs of the four attention heads. The transformation matrix is a learnable linear transformation matrix, with all terms being dimensionless. The cross-modal recursive module employs a bidirectional long short-term memory network with three layers, each with a hidden state dimension of 256. At each time step, the current hidden state vector With transcriptome embedding vector ( Transcriptome embedding vector In the The representation of a time step, and Dimensions are consistent, here Hadamard product gating fusion is performed, and the formula is expressed as follows:
[0099] ;
[0100] In the formula, For the first The hidden state vector after time step fusion and layer normalization The time step index here is a layer normalization function. To distinguish it from Langevin dynamics iteration steps and forward propagation times All terms are dimensionless. The comprehensive adjustment index of the graph density adaptive attention threshold adjustment function. From graph density Protein abundance Shannon information entropy and average edge activation strength Weighted calculation, the formula is expressed as follows:
[0101] ;
[0102] In the formula, The statistical mean of the comprehensive regulation index on the training set is used as a reference normalization factor for the comprehensive regulation index. The density of the protein-protein interaction network graph for the current batch is the proportion of non-zero elements in the adjacency matrix for the current batch; it is dimensionless. The average graph density of the training set, The Shannon information entropy represents the protein abundance of the current batch. The average Shannon information entropy of the training set. The average edge activation intensity of the current batch of dynamic graph attention layer output. The average edge activation strength of the training set. , , The weighted coefficients are and satisfy the following conditions: The determination was made by performing a grid search on the validation set with a conditional jump trigger rate constrained to be between 20% and 40%. , , , All are dimensionless ratios, with consistent dimensions for each value. When, the conditional jump threshold is set to the 75th percentile of the attention weight distribution; when When set to the 85th percentile; when The time frame was set at the 95th percentile. The topology data analysis module performed persistent cohomology calculations on the protein-protein interaction network graph and extracted the Betty number vector. , by the 0th order Betty number (Reflecting the number of connected components) and the first-order Betti number (Reflecting the number of independent loops) together constitute the final fusion layer, which embeds the output sequence of the cross-modal recursive module at the level. Embedded conditional jump nodes and Betty number vector After concatenation, a risk score is output via a two-layer fully connected classification head. The two layers have 512 dimensions and 128 dimensions respectively, as expressed in the following formula:
[0103] ;
[0104] In the formula, For breast cancer proteomics risk scoring, dimensionless. For the first fully connected layer transformation, the concatenated vectors will be... (Overall dimensions) Mapped to 512 dimensions, The second fully connected layer transforms the 512-dimensional vector to 128 dimensions, which is then mapped to a scalar via a linear output layer. It is a linear rectified activation function. This represents a vector concatenation operation, where all terms are dimensionless. Model training uses... Optimizer, initial learning rate at ~ Within the specified range, the learning rate is scheduled using a cosine annealing strategy, with each annealing cycle consisting of 50 training epochs. The batch size ranges from 16 to 32 samples. The loss function is the binary cross-entropy loss. With graph attention weight regularization term The sum, expressed by the formula is as follows:
[0105] ;
[0106] In the formula, The regularization coefficient is determined by... ~ After scanning the validation set within the specified range, the binary cross-entropy loss is determined as follows:
[0107] ;
[0108] In the formula, For the first The actual label of the sample For the first Predicted risk score for each sample, It is a dimensionless quantity. The graph attention weight regularization term is defined as follows:
[0109] ;
[0110] In the formula, This represents the total number of edges in the protein-protein interaction network. For the current forward propagation time edge The square of the weight, As a dimensionless quantity, the regularization term applies to all edge weights. Constraints are imposed to prevent excessive concentration of attention weight.
[0111] The specific implementation of step S06 is as follows: During the data-independent acquisition mode mass spectrometry detection process in step S02, a real-time monitoring system based on a streaming data processing framework is used to perform sliding window statistical process control on the total ion current curve, the internal standard peptide retention time drift value, and the mass-to-charge ratio calibration drift value. The control chart types include... Control charts and cumulative control charts were used. The primary and secondary warning thresholds were determined by performing normality tests and quantile statistics on each quality control indicator from at least 50 batches of historical mass spectrometry operational data, using three standard deviations as a reference and adjusting for linearity in the intra-batch sensitivity decay rate. The allowable drift window for the internal standard peptide retention time was set between 0.1 and 1.5 after statistical analysis of at least 20 replicate injection experiments on the column. Within the specified range. A warning is issued when a Level 1 warning is triggered, and the system automatically shuts down and re-executes step S02 for samples enriched with low-abundance peptides when a Level 2 warning is triggered.
[0112] To better understand and implement this invention, the following is a specific application scenario of the invention, Example 2: Technicians set up a test environment and collected tissue fluid samples from 240 pathologically confirmed breast cancer patients and 110 healthy controls. The samples were processed and analyzed according to the complete process of the method of this invention to verify the feasibility of each step and the overall operational effect of the scheme.
[0113] After extracting exosome particles from tissue fluid samples by differential centrifugation, the samples were processed using a high-abundance protein depletion kit, with an incubation time set to 90 minutes. The sodium chloride concentration of the elution buffer was set to 150 g / L. This parameter combination was determined after five rounds of iterative optimization on 15 concentration gradient standards, achieving an average recovery rate of 78% for low-abundance proteins. After depletion, the sample was digested with trypsin, and the peptides were fractionated sequentially by strong cation exchange chromatography and reversed-phase liquid chromatography. The gradient duration for strong cation exchange chromatography was set to 60 seconds. The gradient time for reversed-phase liquid chromatography was set to 90. The overlap rate of peptides between adjacent components was controlled to be within 12%, lower than the optimization target of 15%. Ultimately, 24 low-abundance peptide enrichment components were obtained for each sample. The identified peptide components reliably covered the low-abundance protein space, such as... Figure 2 The diagram shows the peptide distribution of each component, with the horizontal axis representing chromatographic retention time and the vertical axis representing the number of peptides identified.
[0114] Mass spectrometry detection in data-independent acquisition mode uses a mass number window width of 8. Adjacent windows overlap by 1 The parameter settings were configured, and the spectral library search was used to identify peptides based on a project-specific spectral library. The real-time quality control system continuously monitored the total ion current curve, the internal standard peptide retention time drift, and the mass-to-charge ratio calibration drift during the detection process. Statistical analysis results of 350 batches of historical mass spectrometry data are shown in Table 1. Control chart settings and warning thresholds were determined after correction based on three times the standard deviation, and the allowable drift window for the internal standard peptide retention time was set to 0.4. .
[0115] Table 1 Statistical Analysis Results of Real-time Quality Control Indicators
[0116]
[0117] During the testing of this batch of 350 samples, a Level 1 warning was triggered 12 times and a Level 2 warning was triggered 3 times. After a Level 2 warning was triggered, the system automatically stopped and re-executed the mass spectrometry analysis. Ultimately, all samples passed quality control, yielding a reliable exosome protein profile matrix covering a total of 3847 proteins. Figure 3 The image shows a line graph illustrating the real-time monitoring of the retention time drift of the internal standard peptide during the quality control process.
[0118] The exosome protein profile feature matrix was normalized and then input into a thermodynamic potential surface protein abundance optimization algorithm. The protein interaction network was downloaded from the STRING database, with a confidence threshold set to 0.7 after grid search. The network contains 3124 protein nodes and 41836 interaction edges. The Laplacian matrix of the spectrum was constructed in a sparse matrix format with dimensions of 3124×3124. Temperature parameters... Set to 5.0. The annealing value was set to 0.05, and the annealing curves were determined through scanning experiments on 25 breast cancer exosome samples. After convergence, the algorithm outputs a thermodynamic equilibrium protein abundance matrix. The cooperative consistency of protein abundance within the same community in the protein interaction network is significantly improved, such as... Figure 4 The graph shown is a curve of the free energy loss function during annealing as a function of training rounds. The horizontal axis represents the training rounds, and the vertical axis represents the normalized free energy. .
[0119] The multi-view phylogenetic fusion dimensionality reduction process merges three high-dimensional data views: the thermodynamic equilibrium protein abundance matrix, the transcriptome quantitative matrix, and the lipidome quantitative matrix. The transcriptome quantitative matrix is measured in units of kilobases per million fragments, and the lipidome quantitative matrix contains 876 lipid features. Missing values from the three views are filled using a random forest-based multiple imputation algorithm with 100 trees and a maximum of 10 iterations. The root mean square error of missing values is optimally achieved in 5-fold cross-validation. The regularization penalty coefficient for sparse canonical correlation analysis is set to 0.05 after grid search. The final output is a 30-dimensional low-dimensional latent factor matrix, and the distribution of each sample in the latent factor space is shown below. Figure 5 As shown, the horizontal axis represents the first latent factor, and the vertical axis represents the second latent factor. Breast cancer patients and healthy controls show a clear separation trend in the low-dimensional space.
[0120] The dynamic graph attention cross-modal recursive fusion model was trained using 350 samples, divided into a 7:2:1 ratio: 245 samples for training, 70 samples for validation, and 35 samples for testing. The training set was stratified and randomly sampled according to breast cancer subtype, including luminal A, luminal B, HGF receptor 2 overexpression, and triple-negative subtypes, with the proportions of each subtype consistent with clinical epidemiological proportions. The model was trained using the AdamW optimizer, with an initial learning rate set to [value missing]. The batch size is 24, the cosine annealing cycle is 50 rounds, and the early stopping condition is that the area under the validation set curve no longer increases after 20 consecutive rounds. The weighting coefficients of the graph density adaptive attention threshold adjustment function are... , , After grid search, the threshold values were set to 0.4, 0.3, and 0.3, and the conditional jump trigger rate stabilized at 28%, within the constraint range of 20% to 40%. The model's breast cancer proteomics risk score output on the test set is shown in Table 2. Figure 3 The real-time quality control shown ensures the reliability of the input data. Figure 6 The image shows the distribution of node attention weights in the dynamic graph attention layer on a typical sample. The horizontal axis represents the protein node number, and the vertical axis represents the attention weight value. Nodes with high attention weights correspond to known key proteins in breast cancer.
[0121] Table 2. Results of the Grouping of Breast Cancer Proteomics Risk Score in the Test Set
[0122]
[0123] Compared with traditional methods, this invention brings the following technological advancements: In traditional proteomics analysis, protein abundance normalization and protein interaction network analysis are two completely independent steps. The former cannot perceive the topological constraints of the latter, resulting in a large number of inconsistent false positive signals within the network in the normalization results. This invention embeds the network topological constraints into the free energy loss function, enabling the normalization optimization process itself to have network awareness capabilities. It eliminates inconsistent noise within the network from the signal processing stage, rather than filtering false positives ex-post in the statistical analysis stage. In principle, this is a more proactive and fundamental noise suppression strategy. Furthermore, traditional models typically input classifiers by simply concatenating proteomic, transcriptomic, and lipidomic data. Cross-modal interactions between different omics data are achieved only through dimensional concatenation in the feature space, failing to capture the fine-grained modulation relationship between temporal protein abundance features and transcriptional regulatory signals on the same feature dimension. This invention achieves direct modulation of the two types of signals on each dimension through Hadamard product gating fusion, enabling cross-modal information interaction to occur within the feature dimension rather than through external concatenation. This is superior to simple concatenation strategies in terms of expressive power and theoretically can capture more fine-grained protein-transcriptome co-regulatory patterns, thereby improving the biological interpretability of risk scores.
[0124] It should be noted that the variables involved in this invention are explained in detail in Tables 3, 4, and 5.
[0125] Table 3. Variable Explanation Table (Part 1)
[0126]
[0127] Table 4. Variable Explanation Table (Part Two)
[0128]
[0129] Table 5. Variable Explanation Table (Part 3)
[0130]
[0131] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for analyzing exosome proteins in breast cancer tissue fluid, characterized in that, Includes the following steps: The tissue fluid exosome samples were subjected to high-abundance protein depletion treatment and multidimensional peptide fractionation treatment in sequence to obtain low-abundance peptide enrichment samples. The exosome protein spectrum feature matrix was obtained by data-independent acquisition mode mass spectrometry detection of samples enriched with low-abundance peptides. The exosomal protein spectrum feature matrix is input into the thermodynamic potential energy surface protein abundance optimization algorithm. The protein interaction network topology is embedded in the quadratic coefficient matrix of the free energy loss function using the spectrum Laplacian matrix. Shannon information entropy is introduced as the system entropy term. The Langevin dynamic annealing mechanism is used to make the optimization trajectory jump out of the local minimum, and the thermodynamic equilibrium protein abundance matrix is obtained. The thermodynamic equilibrium protein abundance matrix, transcriptome quantification matrix, and lipidome quantification matrix are input into a multi-view spectral fusion dimensionality reduction process to obtain a low-dimensional latent factor matrix. The exosome proteomic feature matrix, protein-protein interaction network diagram, and patient clinical multi-omics labels are input into a dynamic graph attention cross-modal recursive fusion model to output a breast cancer proteomics risk score. During mass spectrometry detection in data-independent acquisition mode, real-time quality control is performed on the total ion current curve, the retention time drift value of the internal standard peptide, and the mass-to-charge ratio calibration drift value. When an anomaly is triggered, data-independent acquisition mode mass spectrometry detection is re-executed on samples enriched with low-abundance peptides.
2. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 1, characterized in that, The high-abundance protein depletion treatment uses a depletion kit based on the principle of isoelectric focusing enrichment. The incubation time and ionic strength of the elution buffer are determined iteratively by repeatedly experimenting on breast cancer tissue fluid standards at multiple concentration gradients, with the number of proteins and the recovery rate of low-abundance proteins as evaluation indicators.
3. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 2, characterized in that, The multidimensional peptide fractionation process separates the peptides after proteolytic digestion through two dimensions: strong cation exchange chromatography and reversed-phase liquid chromatography. The gradient duration and salt concentration switching nodes are determined with the goal of ensuring that the overlap rate of peptides between adjacent components is lower than the overlap rate threshold.
4. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 3, characterized in that, The data-independent acquisition mode divides the entire mass number range into fixed mass number windows, and simultaneously fragments all ions within each mass number window and acquires secondary spectra. The width and overlap of the mass number windows are determined by dual indicators: the number of peptides identified and the quantitative coefficient of variation.
5. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 4, characterized in that, The thermodynamic potential energy surface protein abundance optimization algorithm defines the system internal energy as the sum of squared abundance deviations weighted by the edge weights of the protein interaction network, defines the system entropy as the Shannon information entropy of the protein abundance distribution, constructs the free energy loss function as the difference between the system internal energy and the system entropy, and the temperature parameter decreases linearly from the highest temperature to the lowest temperature according to the annealing schedule.
6. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 5, characterized in that, The Laplacian matrix of the spectrum is the difference between the degree matrix and the adjacency matrix of the protein interaction network graph. In the thermodynamic potential energy surface protein abundance optimization algorithm, it serves as the quadratic coefficient matrix of the system's internal energy, enabling highly interconnected protein nodes within the same protein interaction network community to converge collaboratively during the abundance optimization process.
7. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 6, characterized in that, The multi-view phylogenetic fusion and dimensionality reduction process adopts a multi-factor analysis framework algorithm, which maps the three high-dimensional data views to low-dimensional latent factor spaces and then fuses them. Missing values are filled using a multiple imputation algorithm based on random forest, and cross-omics high correlation features are screened through sparse canonical correlation analysis.
8. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 7, characterized in that, The dynamic graph attention cross-modal recursive fusion model includes a dynamic graph attention layer, a cross-modal recursive module, a conditional jump mechanism, a topology data analysis module, and a final fusion layer. The dynamic graph attention layer consists of multiple parallel attention heads, and the attention coefficients between nodes are dynamically calculated by the cosine similarity of the protein node feature vectors.
9. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 8, characterized in that, The cross-modal recursive module uses a bidirectional long short-term memory network as its backbone. At each time step, it performs a Hadamard product-gated fusion of the current hidden state vector with the embedding vector of the number of kilobase fragments per million fragments in the transcriptome. The fusion result is then normalized and passed to the next time step.
10. The method for analyzing exosome proteins in breast cancer tissue fluid according to claim 9, characterized in that, The conditional jump mechanism calculates a dynamic threshold using a graph density adaptive attention threshold adjustment function. This function takes the current batch graph density, protein abundance Shannon information entropy, and average edge activation intensity as inputs and outputs a comprehensive adjustment index.