Method and system for predicting the effect of immunotherapy for liver cancer
By employing multiple immunofluorescence staining techniques and deep learning methods, a multi-dimensional feature space was constructed and principal component analysis was performed. Combined with Bayesian networks and neural networks for anomaly detection, the problem of neglecting the prediction of immune cytokine interactions in existing technologies was solved, enabling refined grading and quantitative risk assessment of the efficacy of liver cancer immunotherapy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIVERSITY SHENZHEN HOSPITAL
- Filing Date
- 2025-06-05
- Publication Date
- 2026-04-17
AI Technical Summary
Current clinical predictions of the efficacy of immunotherapy for liver cancer mainly rely on the analysis of the expression levels of single immune cytokines, neglecting the complex interactions between immune cytokines. Furthermore, traditional data analysis methods cannot effectively capture abnormal features in high-dimensional, non-linear immune cytokine expression data.
Data from mononuclear cells were collected using multiplex immunofluorescence staining techniques. A multidimensional feature space was constructed and principal component analysis was performed. The conditional probability relationships and mutual information values between immune factors were analyzed using Bayesian networks. Anomaly detection was performed using a multilayer neural network model. Quantitative risk assessment was conducted using minimum tail bound analysis and a semidefinite programming model.
This approach enables refined grading and quantitative risk assessment of the efficacy of immunotherapy for liver cancer, improves the efficiency and accuracy of data processing, reveals the complex relationships between immune cytokines, and enhances the ability to identify complex immune expression patterns.
Smart Images

Figure CN120636543B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for predicting the efficacy of immunotherapy for liver cancer. Background Technology
[0002] Current clinical predictions of the efficacy of immunotherapy for liver cancer mainly rely on the analysis of the expression levels of single immune cytokines. This method ignores the complex interactions between immune cytokines. At the same time, traditional data analysis methods have limitations in processing high-dimensional, non-linear immune cytokine expression data and cannot effectively capture abnormal features in immune cytokine expression patterns. Summary of the Invention
[0003] This invention provides a method and system for predicting the efficacy of immunotherapy for liver cancer. This invention enables quantitative risk assessment of the prediction results and refined grading of the prediction results.
[0004] In a first aspect, the present invention provides a method for predicting the efficacy of immunotherapy for liver cancer, the method comprising:
[0005] Clinical samples were pretreated with multiplex immunofluorescence staining, and forward scattering data, side scattering data, and multichannel fluorescence data of mononuclear cells were collected to obtain an initial data matrix of cell expression.
[0006] The initial data matrix of cell expression was standardized to construct an n-dimensional feature space, and principal component analysis was performed to obtain a multi-dimensional feature matrix of immune factor expression.
[0007] Bayesian network analysis is performed based on the multi-dimensional feature matrix of the immune factor expression to calculate the conditional probability relationship and mutual information value between immune cell factors, and an immune factor association network is constructed according to a preset association strength threshold.
[0008] Extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into a multi-layer neural network model for anomaly detection to obtain the immune factor expression anomaly detection result;
[0009] The results of abnormal immune factor expression were analyzed using minimum tail bound analysis, and the optimal risk boundary was solved using a semidefinite programming model to output quantitative assessment data of immune factor abnormalities.
[0010] Secondly, the present invention provides a prediction system for the efficacy of immunotherapy for liver cancer, the prediction system comprising:
[0011] The preprocessing module is used to perform multiplex immunofluorescence staining preprocessing on clinical samples and to collect forward scattering data, side scattering data and multi-channel fluorescence data of mononuclear cells to obtain an initial data matrix of cell expression.
[0012] The analysis module is used to standardize the initial data matrix of cell expression, construct an n-dimensional feature space, and perform principal component analysis to obtain a multi-dimensional feature matrix of immune factor expression.
[0013] The construction module is used to perform Bayesian network analysis based on the multi-dimensional feature matrix of the immune factor expression, calculate the conditional probability relationship and mutual information value between immune cell factors, and construct an immune factor association network according to a preset association strength threshold.
[0014] The detection module is used to extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into a multi-layer neural network model for anomaly detection to obtain the immune factor expression anomaly detection result.
[0015] The output module is used to perform minimum tail boundary analysis on the abnormal expression detection results of the immune factors, and solve the optimal risk boundary through a semidefinite programming model to output quantitative assessment data of immune factor abnormalities.
[0016] A third aspect of the present invention provides a device for predicting the efficacy of immunotherapy for liver cancer, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the device for predicting the efficacy of immunotherapy for liver cancer to perform the above-described method for predicting the efficacy of immunotherapy for liver cancer.
[0017] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described method for predicting the efficacy of liver cancer immunotherapy.
[0018] The technical solution provided by this invention achieves effective dimensionality reduction and feature extraction of high-dimensional immune cytokine expression data by constructing a multi-dimensional feature space and performing principal component analysis, thereby improving the efficiency and accuracy of data processing. A Bayesian network analysis method is used to construct an immune factor association network, revealing the conditional probability relationships and mutual information characteristics among immune cytokines, overcoming the limitations of traditional single-factor analysis methods. A multi-layer neural network model is introduced for anomaly detection, combining network centrality and expression features to enhance the ability to identify complex immune expression patterns. Minimum tail bound analysis and semidefinite programming models are used to achieve quantitative risk assessment of prediction results. Multiplex immunofluorescence staining and flow cytometry are employed to ensure the accuracy and reproducibility of sample data collection. A three-level risk assessment mechanism is used to achieve refined grading of prediction results. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating the method for predicting the efficacy of liver cancer immunotherapy provided in this application embodiment;
[0021] Figure 2 A schematic block diagram of the structure of a system for predicting the efficacy of liver cancer immunotherapy provided in an embodiment of this application;
[0022] Figure 3 A schematic block diagram of the structure of a device for predicting the efficacy of liver cancer immunotherapy provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change based on the actual situation.
[0025] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0026] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features described herein can be combined with each other.
[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating the method for predicting the efficacy of liver cancer immunotherapy provided in an embodiment of this application, as shown below. Figure 1 As shown, the method for predicting the efficacy of liver cancer immunotherapy provided in this application includes steps S100 to S600.
[0029] Step S100: Perform multiple immunofluorescence staining pretreatment on clinical samples, and collect forward scattering data, side scattering data and multi-channel fluorescence data of mononuclear cells to obtain the initial cell expression data matrix;
[0030] It is understood that the executing entity of this invention can be a prediction system for the efficacy of liver cancer immunotherapy, or it can be a terminal or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.
[0031] Specifically, clinical samples were processed using density gradient centrifugation to separate a mesolayer rich in mononuclear cells. Utilizing the density differences of different cell components, the gradient centrifugation medium (such as Ficoll or Percoll) effectively removed impurities such as red blood cells and plasma, retaining only mononuclear cells in the mesolayer. After centrifugation, the cell suspension from the mesolayer was extracted and processed. To ensure sample purity, the mesolayer sample was washed with phosphate-buffered saline (PBS) to remove residual gradient medium and other contaminants. After multiple washes, the purified nucleated cell sample was used for subsequent experimental steps. To screen and focus on viable nucleated cells, flow cytometry was performed on the purified nucleated cell sample. Forward scattering (FSC-A) and side scattering (SSC-A) data were acquired using flow cytometry and a two-parameter scatter plot was generated, effectively distinguishing cell population size and granularity, thus differentiating viable cells from debris or dead cells. During this process, appropriate threshold settings were used to precisely select viable nucleated cells, which were used for subsequent immunofluorescence labeling and data acquisition steps. The cell density of the screened viable nucleated cell sample was adjusted to 1 × 10⁻⁶. 6The sample density was controlled at 1 / mL to ensure sample homogeneity and data acquisition accuracy. Fluorescently labeled monoclonal antibodies, including CD3, CD4, CD8, CD19, CD25, and CD56, were added to the adjusted cell suspension. These antibodies label specific types of immune cells and related factors, respectively. This step, through the specific binding of antibodies to cell surface antigens, endows cells with specific fluorescent characteristics, providing information for subsequent multichannel fluorescence detection. After antibody addition, the reaction system was mixed and incubated at a suitable temperature to ensure sufficient binding of the fluorescent antibodies. After labeling, flow cytometry was used to acquire data from the immunofluorescently labeled cells. The power of the 488nm laser was set to 20mW for excitation of the fluorescent labeling. Forward scattering (FSC) signals were acquired as forward scattering data of the cells, reflecting cell size and refractive properties. Simultaneously, side scattering (SSC) signals were acquired as side scattering data, reflecting intracellular granularity and structural complexity. The fluorescence signals of the fluorescently labeled cells were detected, and the fluorescence intensity signals of six fluorescence channels—FITC, PE, PerCP, APC, PE-Cy7, and APC-Cy7—were acquired. These fluorescence data correspond to the expression levels of different antigens on the cell surface, including detailed information on immune cell subsets and related factors. After fluorescence signal acquisition, a compensation matrix was calculated for the acquired cell morphology parameters and fluorescence data. Due to spectral overlap between multi-channel fluorescence signals, to avoid cross-interference, compensation coefficients were calculated using data from single-stained control samples, and the original multi-channel data were spectrally compensated to obtain corrected multi-parameter data. The compensated multi-parameter data were arranged in a specific order, including FSC, SSC, FITC-CD3, PE-CD4, PerCP-CD8, APC-CD19, PE-Cy7-CD25, and APC-Cy7-CD56, to generate an initial cell expression data matrix. This matrix contains cell morphology parameters and covers multi-channel fluorescence data of immune cell surface factors.
[0032] Step S200: Standardize the initial cell expression data matrix, construct an n-dimensional feature space, and perform principal component analysis to obtain a multi-dimensional feature matrix of immune factor expression.
[0033] Specifically, the initial cell expression data matrix, including forward scattering (FSC), side scattering (SSC), and data from the six fluorescence channels, underwent maximum and minimum value normalization. Each data feature was adjusted to the same numerical range to eliminate bias caused by differences in feature dimensions. The maximum and minimum values of FSC, SSC, and each fluorescence channel were used as benchmarks, and the corresponding data were transformed and remapped to the normalized range. After normalization, a normalized eigenvalue matrix was obtained. The values of each feature were then centered by subtracting the mean of that feature, making the data distribution centered at zero. The goal of centering is to eliminate eigenvalue bias, make the data more symmetrical, and enhance the feasibility of relative comparisons. After centering, a normalized data matrix was obtained. Based on the expression of various immune factors in the normalized data matrix, an n-dimensional feature space was constructed using CD3, CD4, CD8, CD19, CD25, and CD56 as primary markers. The dimensions of the n-dimensional feature space correspond to the number of immune factors studied, with each dimension representing a specific factor. In this process, each feature in the standardized data is mapped onto an n-dimensional space, completing the projection of all data features into the high-dimensional space. The covariance matrix between features is calculated based on the n-dimensional projected data. The covariance relationship between each feature is calculated, and a correlation coefficient matrix is constructed based on the covariance relationship. This matrix is used to quantify the strength and direction of the association between features. Based on the correlation coefficient matrix, the set of principal component eigenvalues and their corresponding eigenvectors are extracted using eigenvalue decomposition. Each eigenvalue reflects the importance of the principal component in explaining data variability. Based on the eigenvalues, the contribution rate and cumulative contribution rate of each principal component are calculated to determine the number of principal components to select. Selecting several principal components with high contribution rates retains most of the information while effectively reducing the dimensionality of the data. The extracted principal component eigenvector set is subjected to Schmitt orthogonalization to construct a fully orthogonal eigenvector matrix, where each vector is independent and there is no redundancy or linear correlation. This orthogonalization matrix is used to perform a linear transformation on the n-dimensional projected data to generate a multi-dimensional feature matrix of immune factor expression.
[0034] Step S300: Perform Bayesian network analysis based on the multi-dimensional feature matrix of immune factor expression, calculate the conditional probability relationship and mutual information value between immune cell factors, and construct an immune factor association network according to the preset association strength threshold.
[0035] Specifically, based on the multi-dimensional feature matrix of immune factor expression, six immune cytokines—CD3, CD4, CD8, CD19, CD25, and CD56—are defined as nodes in a Bayesian network, forming an immune factor node set. These nodes represent the expression characteristics of immune factors and serve as the basic units for network analysis. When constructing the Bayesian network structure, probabilistic statistical analysis is used to confirm the possible connections between each node, calculating the conditional probability relationships and mutual information values between nodes. To accurately estimate the probability distribution characteristics of each immune factor node, kernel density estimation is used to fit the probability density of each node. Kernel density estimation is a non-parametric statistical method that uses a kernel function to smoothly fit the sample distribution of the nodes, obtaining the probability density distribution curve for each node. After obtaining the probability distribution data of the nodes, the conditional probability distribution between each pair of nodes is calculated, and a conditional independence test is performed on the conditional probability data. The conditional independence test uses statistical methods to verify whether there is a direct dependency between node pairs. The G-value for each pair of nodes is calculated. 2 A statistical value is used to quantify the conditional dependencies between nodes, and the result is stored in the conditional dependency matrix. Based on the conditional dependency matrix, the mutual information value between each pair of nodes is calculated. The mutual information value is an indicator used to measure the degree of information sharing between random variables, and is calculated by the entropy difference between the joint probability distribution and the marginal probability distribution. Through mutual information analysis, the association strength between nodes is quantified, and the potential relationship structure between immune factors is revealed. After organizing the mutual information values of all node pairs into mutual information strength data, strongly correlated node pairs are selected by setting an association strength threshold. Only node pairs with mutual information values greater than the threshold are retained. These selected node pairs constitute the initial connection relationship and are represented as an adjacency matrix. Based on the initial connection structure, to ensure the acyclicity of the constructed network and optimize its connection method, the maximum weight spanning tree algorithm is executed. Using the mutual information value as the edge weight, the maximum weight spanning tree algorithm generates an acyclic tree topology. This structure connects all nodes in a way that minimizes cycles and retains the most important association information. To determine the directionality of edges in the tree topology, conditional probabilities are calculated for each edge, and causal relationships are determined using the local Markov property. The core of the local Markov property lies in determining the directionality of dependency between two nodes through contextual conditional probability analysis. For example, if the conditional probability change of node A significantly affects the distribution of node B, but the converse is not true, then it is inferred that A is a causal precursor of B. Through analysis, each edge in the tree topology is assigned a direction, generating a set of directed edges. All directed edges are then integrated back into the tree topology to complete the construction of the immune factor association network.
[0036] Step S400: Extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into the multi-layer neural network model for anomaly detection to obtain the abnormal detection result of immune factor expression.
[0037] Specifically, the association feature sets of each node are extracted from the constructed immune factor association network. To quantify the importance and structural attributes of each node in the network, the degree centrality of the nodes is calculated. Degree centrality measures the connectivity of each node in the network by counting the number of edges directly connected to it. To ensure the comparability of features between different networks, the degree centrality is normalized to obtain the node connectivity feature, which is calculated by dividing the number of edges connected to the node by the total number of edges in the network. The node connectivity feature reflects the degree of direct association between a particular immune factor and other factors. Betweenness centrality is calculated for each node in the network, counting the number of paths passing through that node in the shortest paths between all pairs of nodes. Betweenness centrality measures the bridging role of a node in the network, that is, its role in the global information flow. The statistically obtained betweenness centrality data is normalized to form the node betweenness feature. Simultaneously, the proximity centrality of each node is calculated; this feature evaluates the reachability of a node by measuring the sum of the reciprocals of the shortest distances from the node to all other nodes in the network. Proximity centrality, after normalization, forms the reachability feature of a node, reflecting its proximity to other nodes in the network's global position. The node connectivity feature, node intermediary feature, and node reachability feature are combined to form an immune factor association feature set. This immune factor association feature set is concatenated with the multi-dimensional feature matrix of immune factor expression along its feature dimensions, combining network structure information with expression data to form a fused feature vector containing rich contextual information. This fused feature vector is input into the input layer of a multi-layer neural network model, where the number of neurons matches the dimension of the fused feature vector. The input layer data undergoes a non-linear transformation using the Swish activation function. The Swish activation function's smooth and differentiable properties help improve the model's non-linear fitting ability. To enhance the model's robustness and prevent overfitting, an Alpha Dropout layer is added after the input layer, randomly masking the activation values of some neurons, allowing the model to generalize better. The feature vector processed by the input layer is then passed to the first convolutional block of the multi-layer neural network model. This first convolutional block contains two one-dimensional convolutional layers to extract local patterns and sequence correlations from the feature vector. Convolutional layers can progressively abstract features through a sliding window operation, transforming the input high-dimensional feature vector into first-level convolutional features. These first-level convolutional features are then input into a second convolutional block. The second convolutional block also contains two one-dimensional convolutional layers, with residual connections added between them. The residual structure effectively mitigates the vanishing gradient problem in deep networks by preserving input information at a deeper level, while simultaneously improving the expressive power of the features. The second-level convolutional features processed through residual connections are more robust and contain higher-level abstract features.The second convolutional features are processed by Global Average Pooling to reduce the feature vector to a low-dimensional representation of a fixed size, thereby obtaining the detection results of abnormal expression of immune factors and reflecting the potential abnormal patterns in multidimensional data.
[0038] Step S500: Perform minimum tail boundary analysis on the results of abnormal immune factor expression detection, and solve the optimal risk boundary using a semidefinite programming model to output quantitative assessment data of immune factor abnormalities.
[0039] Specifically, a risk distribution matrix is constructed based on the abnormal samples identified in the detection results of abnormal immune factor expression to capture the statistical characteristics of these abnormal samples. In the risk distribution matrix, the third and fourth moments are calculated for each abnormal sample. These higher-order moments reflect the skewness and tail behavior of the sample distribution. The third moment is used to calculate the skewness value, representing the degree of skewness of the distribution relative to the center of symmetry, while the fourth moment is used to calculate the kurtosis value, quantifying the concentration or dispersion of the distribution tails. The Jarque-Bera test is performed on the skewness and kurtosis values; this is a statistical method used to determine whether the data conforms to a normal distribution. Through this test, the tail distribution characteristics of the abnormal samples are identified, and a conditional risk value function is constructed accordingly. The conditional risk value function focuses on the tail characteristics of the abnormal samples, and calculates the minimum tail bound probability by combining expectation and variance, thereby determining the initial risk boundary. The initial risk boundary is a preliminary risk measure, reflecting the probability that abnormal samples will appear in the tails under the current data distribution. After the initial risk boundary is determined, a semidefinite programming model is constructed to optimize the risk boundary. The initial risk boundary is used as input to the semidefinite programming model. A constraint matrix is set to ensure it satisfies positive semidefiniteness. The diagonal elements of the constraint matrix are defined as risk confidences, quantifying the independent risk probability of each sample, while the off-diagonal elements represent the covariance coefficients between samples, describing their correlation. Based on the constraint matrix, an optimization objective function is defined to minimize the risk boundary. Linear inequality constraints and probability measure constraints are introduced, transforming the problem into a standardized semidefinite programming problem. The linear inequality constraints limit the shape and range of the risk boundary, while the probability measure constraints ensure the statistical reasonableness and interpretability of the optimization results. Through this standardization process, the entire problem is clearly formalized into a solvable optimization model. The standardized semidefinite programming problem is solved using the interior-point method. The interior-point method is an efficient convex optimization algorithm that iteratively finds the optimal value of the objective function within the constraints. After solving using the interior-point method, the resulting risk optimization boundary is a risk measure optimized under multiple conditions, which can more accurately describe the risk distribution of anomalous samples. Based on the risk optimization boundary, a three-level risk assessment interval is constructed to classify the risk level of abnormal samples. The three-level risk interval includes low risk, medium risk, and high risk. By setting boundary points on the optimization boundary, abnormal samples are assigned to the corresponding intervals according to their risk level. The risk level classification quantifies the risk level of each sample and visually displays the relative position of the abnormal sample in the risk distribution. The risk classification results are correlated and mapped with the detection results of abnormal immune factor expression, labeling each abnormal sample with its specific risk level and risk probability value. The correlation process combines the abnormal characteristics of the detection results and the risk information of the optimization boundary to generate accurate quantitative assessment data for each abnormal sample.The final output of quantitative assessment data on abnormal immune factors includes the risk level, risk probability value, and related tail boundary analysis information for each abnormal sample.
[0040] To apply the interior-point method to the standardized semidefinite programming problem, a logarithmic barrier function is constructed, transforming inequality constraints into penalty terms in the objective function. By introducing a penalty function in logarithmic form, the original objective function is extended to a central path equation containing constraints. The logarithmic barrier function, through continuously increasing penalty values, forces the solution to gradually approach the boundary of the feasible region, while ensuring the continuity and differentiability of the problem. The construction of the central path equation ensures that the interior-point method optimizes within the feasible region under the influence of the penalty function, thus guaranteeing the stability of the solution. Based on the central path equation, an augmented system matrix is constructed to linearize the KKT (Karush-Kuhn-Tucker) conditions for the original, dual, and slack variables. The KKT conditions are a set of necessary optimality conditions; by performing a first-order Taylor expansion on these conditions, a linearized system of equations describing the search direction of the variables is obtained. The augmented system matrix transforms the optimization problem into a matrix-solving problem about the search direction, forming the core mathematical framework of the interior-point method. The search direction system is numerically solved by calculating the Newton step sequence. The Newton step-size sequence optimizes the direction and magnitude of the current solution in each iteration, gradually approaching the optimal solution. After each iteration, the original and dual solutions are adjusted using the updated Newton step-size, and the current duality gap, i.e., the optimization error between the original and dual solutions, is calculated. The duality gap reflects the convergence degree of the solution; when the duality gap is less than a preset target value, the optimal solution to the standardized semidefinite programming problem, i.e., the optimal risk boundary solution, is considered to have been found. Conditional risk values (VaR and CVaR) are calculated for the optimal risk boundary solution. VaR (Value at Risk) represents the maximum possible loss at a given confidence level, while CVaR (Conditional Value at Risk) represents the average loss exceeding the VaR value. By combining the calculation results of VaR and CVaR, a comprehensive risk interval division standard is constructed. This standard divides the optimal risk boundary solution into three intervals: low risk, medium risk, and high risk, forming a three-level risk assessment interval. A risk metric is calculated for each abnormal sample in the abnormal immune factor expression detection results. The risk metric for each sample is calculated by combining the optimal risk boundary solution and the distribution characteristics of outlier samples, yielding a risk probability value for each sample. These risk probability values are then compared and matched against a three-tiered risk assessment interval, assigning a corresponding risk level label to each outlier sample. The matching process categorizes samples as low-risk, medium-risk, or high-risk based on the interval within which their risk probabilities fall, thus achieving precise risk grading. Combining the risk level labels and corresponding risk probability values of all outlier samples generates a complete risk grading result.
[0041] In this embodiment of the invention, by constructing a multi-dimensional feature space and performing principal component analysis, effective dimensionality reduction and feature extraction of high-dimensional immune cytokine expression data are achieved, improving the efficiency and accuracy of data processing. A Bayesian network analysis method is used to construct an immune factor association network, revealing the conditional probability relationships and mutual information characteristics among immune cytokines, overcoming the limitations of traditional single-factor analysis methods. A multi-layer neural network model is introduced for anomaly detection, combining network centrality and expression features to enhance the ability to identify complex immune expression patterns. Quantitative risk assessment of prediction results is achieved through minimum tail bound analysis and semidefinite programming models. Multiplex immunofluorescence staining and flow cytometry are used to ensure the accuracy and reproducibility of sample data collection. A three-level risk assessment mechanism enables refined grading of prediction results.
[0042] In one specific embodiment, the process of performing step S100 may specifically include the following steps:
[0043] Clinical samples were subjected to density gradient centrifugation to obtain an intermediate layer sample containing mononuclear cells. The intermediate layer sample was then washed by centrifugation with phosphate buffer to obtain a purified nucleated cell sample.
[0044] Flow cytometry was used to count purified nuclear cell samples, and viable nuclear cells were obtained by analyzing FSC-A and SSC-A dual-parameter scatter plots.
[0045] The density of viable nucleated cells was adjusted to 1×106 cells / mL, and monoclonal antibodies against CD3, CD4, CD8, CD19, CD25, and CD56 were added to obtain immunofluorescently labeled cells.
[0046] For immunofluorescence-labeled cells, the power of a 488nm laser was set to 20mW. The FSC signal was collected as forward scattering data, and the SSC signal was collected as side scattering data to obtain cell morphology parameters.
[0047] Fluorescence intensity signals from six fluorescence channels (FITC, PE, PerCP, APC, PE-Cy7, and APC-Cy7) were collected from immunofluorescently labeled cells to obtain multi-channel fluorescence data.
[0048] Compensation matrix calculations were performed on cell morphology parameters and multi-channel fluorescence data. Compensation coefficients were determined using single-stained control sample data to obtain compensated multi-parameter data.
[0049] The compensated multi-parameter data were arranged in the order of FSC, SSC, FITC-CD3, PE-CD4, PerCP-CD8, APC-CD19, PE-Cy7-CD25, and APC-Cy7-CD56 to construct the initial cell expression data matrix.
[0050] Specifically, mononuclear cells are isolated from clinical samples. Using density gradient centrifugation, clinical samples are placed in a gradient medium (e.g., Ficoll) and processed in a centrifuge at specific speeds and times. During centrifugation, cell components of different densities distribute along the gradient, eventually forming a multilayer structure, with the middle layer rich in mononuclear cells (PBMCs, peripheral blood mononuclear cells). The middle layer sample is carefully aspirated via pipette and transferred to phosphate-buffered saline (PBS), followed by low-speed centrifugation to wash away residual gradient medium and other impurities. After purification, the resulting cell suspension contains highly purified mononuclear cells. The purified cell samples are then counted and quality-assessed using flow cytometry. The optical system of the flow cytometer detects forward scattering (FSC-A) and side scattering (SSC-A) signals, two parameters reflecting cell size and internal structural complexity. Viable cells are screened using a gating strategy based on the generated two-parameter scatter plot. For example, let the forward scattering signal of the cell be x. FSC The side-scattered signal is x SSC The active nuclear cell population meets the following conditions:
[0051] x FSC ∈[F min ,F max ],x SSC ∈[S min ,S max ];
[0052] Where F min ,F max ,S min ,S max These represent the gating ranges for forward and lateral scattering signals, respectively. The concentration of the selected viable nucleated cell samples was adjusted to 1×10⁻⁶. 6Cells were incubated at a density of [number] cells / mL, and the adjusted cell suspension was mixed with fluorescently labeled antibodies. The monoclonal antibodies used included CD3 (T cell marker), CD4 (helper T cell marker), CD8 (cytotoxic T cell marker), CD19 (B cell marker), CD25 (activated T cell marker), and CD56 (natural killer cell marker). During incubation, the antibodies were ensured to fully bind to cell surface antigens, forming immunofluorescently labeled cells. The labeled cells were then analyzed using flow cytometry to collect various signal data. Cells were excited using a 488nm laser (20mW), and forward scatter (FSC) and side scatter (SSC) signals were collected as morphological parameters. In addition, fluorescence intensity signals from six channels were collected: FITC (green fluorescence), PE (orange fluorescence), PerCP (red fluorescence), APC (far-infrared fluorescence), PE-Cy7 (orange-red fluorescence), and APC-Cy7 (deep-infrared fluorescence). These fluorescence data corresponded one-to-one with the antibody-labeled immune factors, recording the expression of specific surface markers for each cell. Because there is spectral overlap between multi-channel fluorescence signals, a compensation matrix is calculated to eliminate cross-interference between different fluorescence signals. Let the signal of each fluorescence channel be x. i Where i = 1, 2, ..., 6 represents six fluorescence channels, and the correction model for cross-interference is:
[0053] x′ i =x i -∑ j≠i c ij ·x j ;
[0054] Where x′ i For the compensated signal, c ij This represents the compensation coefficient between channels i and j. The compensation coefficient was experimentally determined using data from single-stain control samples. For example, if the compensation coefficient c for FITC and PE channels... FITC,PE =0.02, then when compensating for the FITC signal, a 2% contribution from the PE signal needs to be subtracted. After compensation, the resulting multi-parameter data are arranged in a specific order, including FSC (forward scattering signal), SSC (side scattering signal), FITC-CD3 (green fluorescence signal), PE-CD4 (orange fluorescence signal), PerCP-CD8 (red fluorescence signal), APC-CD19 (far-infrared fluorescence signal), PE-Cy7-CD25 (orange-red fluorescence signal), and APC-Cy7-CD56 (deep-infrared fluorescence signal), thus constructing an initial data matrix for cell expression. Let the data matrix be X, where the i-th row represents the data of the i-th cell, and the matrix form is:
[0055]
[0056] Where n is the total number of cells, x i,k This represents the k-th feature of the i-th cell.
[0057] In one specific embodiment, the process of performing step S200 may specifically include the following steps:
[0058] The FSC, SSC, and six fluorescence channel data in the initial cell expression data matrix were normalized by maximum and minimum values to obtain the normalized eigenvalue matrix.
[0059] The features of each dimension of the normalized eigenvalue matrix are centered to obtain a standardized data matrix.
[0060] An n-dimensional feature space is constructed based on the standardized data matrix, labeled with CD3, CD4, CD8, CD19, CD25, and CD56. The data of each dimension is then projected onto the n-dimensional feature space to obtain n-dimensional projected data.
[0061] The covariance matrix between features is calculated based on n-dimensional projection data, and an n×n-dimensional correlation coefficient matrix is constructed using the correlation between n immune cytokines to obtain the feature correlation matrix.
[0062] Perform eigenvalue decomposition on the feature correlation matrix to obtain the principal component eigenvalue set, and calculate the contribution rate and cumulative contribution rate of each eigenvalue based on the principal component eigenvalue set to obtain the principal component eigenvector group;
[0063] The principal component eigenvector group is subjected to Schmitt orthogonalization to construct an orthogonalized eigenvector matrix. Then, the n-dimensional projection data is linearly transformed to obtain a multi-dimensional feature matrix of immune factor expression.
[0064] Specifically, the FSC, SSC, and six fluorescence channel data in the initial cell expression data matrix were normalized to their minimum and maximum values, scaling their range to [0,1]. Let x be the data of a certain feature column. j The maximum value is max(x) j The minimum value is min(x). j The normalization formula is:
[0065]
[0066] Where x′ i,j These are the normalized eigenvalues. After normalization, we obtain the normalized eigenvalue matrix X′, where all eigenvalues are between [0,1]. We then center the normalized eigenvalue matrix X′ so that the mean of each eigenvalue is 0. Let μ... j Let be the mean of the features in column j, then the centering formula is:
[0067] x″i,j =x′ i,j -μj;
[0068] Where x″ i,j These are the eigenvalues after centering. After centering, a standardized data matrix X″ is obtained. Using CD3, CD4, CD8, CD19, CD25, and CD56 as labels, an n-dimensional feature space is constructed. The n-dimensional feature space consists of the feature axes corresponding to the six immune factors. Each sample in the standardized data matrix X″ is projected onto the n-dimensional feature space to obtain the projected n-dimensional data Y. Assume the i-th row of X″ is [x″...]. i,CD3 ,x″ i,CD4 ,x″ i,CD8 ,…,x″ i,CD56 Let the i-th row of Y be the projection of the aforementioned features into n-dimensional space. Based on the projected data Y, calculate the covariance matrix between the features. Let the mean vector of Y be μ. y Then the formula for the covariance matrix C is:
[0069]
[0070] The covariance matrix C is an n×n symmetric matrix whose elements C ij Let represent the covariance between the i-th and j-th features. By performing eigenvalue decomposition on the covariance matrix, we obtain the set of principal component eigenvalues {λ1, λ2, ..., λj}. n} and its corresponding set of feature vectors {v1, v2, ..., v n Each eigenvalue λ i The eigenvector v represents the variance of the corresponding principal component. i This indicates the direction of the principal component. To measure the importance of each principal component, its contribution rate is calculated:
[0071]
[0072] The cumulative contribution rate is calculated by summing the contribution rates of each principal component. Principal components with a cumulative contribution rate of 95% or higher are selected to form a principal component eigenvector group. To ensure the orthogonality between the principal component eigenvectors, they are subjected to Schmitt orthogonalization to construct an orthogonalized eigenvector matrix V. Schmitt orthogonalization is achieved by progressively eliminating correlated components between vectors, resulting in an orthogonal matrix. The orthogonalized eigenvector matrix V is used to perform a linear transformation on the projected data Y to obtain the multidimensional feature matrix Z of immune factor expression.
[0073] Z = Y·V;
[0074] Each row of matrix Z represents the principal component features of a sample, containing key expression information of immune factors.
[0075] In one specific embodiment, the process of performing step S300 may specifically include the following steps:
[0076] A Bayesian network structure was constructed based on the multidimensional feature matrix of immune factor expression, and each immune cytokine, CD3, CD4, CD8, CD19, CD25, and CD56, was set as a network node to obtain the immune factor node set.
[0077] The conditional probability distribution between each pair of nodes is calculated based on the immune factor node set. The probability density of each node is fitted using the kernel density estimation method to obtain the node probability distribution data.
[0078] Perform a conditional independence test on the node probability distribution data and calculate G for each pair of nodes. 2 The statistical values are used to obtain the conditional dependency matrix. Based on the conditional dependency matrix, the mutual information values between nodes are calculated. The mutual information content is calculated using the entropy difference between the joint probability distribution and the marginal probability distribution, thus obtaining the mutual information intensity data.
[0079] A threshold for association strength is set for mutual information strength data. Node pairs with values greater than the association strength threshold are selected to construct an adjacency matrix, resulting in an initial connection structure. The maximum weight spanning tree algorithm is then executed based on the initial connection structure, using mutual information values as edge weights to construct an acyclic connection structure, thus obtaining a tree-like topology.
[0080] The directionality of edges in the tree topology is determined by conditional probability calculation, and causal relationships are determined by local Markov properties to obtain a set of directed edges. The set of directed edges is then integrated into the tree topology to obtain the immune factor association network.
[0081] Specifically, the multidimensional feature matrix X of immune factor expression contains n samples and m = 6 immune factors (CD3, CD4, CD8, CD19, CD25, CD56), and the matrix form is as follows:
[0082]
[0083] Where x i,j Let X represent the expression value of the j-th immune factor in the i-th sample. The six immune factors in X are set as nodes in a Bayesian network, forming a node set V = {CD3, CD4, CD8, CD19, CD25, CD56}. The connections between nodes represent the conditional probability relationships between immune factor components. To construct the conditional probability distribution between nodes, the kernel density estimation method is used to fit the probability density function of each node. For a given node v... j The set of observations {x 1,j ,x 2,j ,…,x n,j The expression for kernel density estimation is:
[0084]
[0085] in It is node v j The probability density estimation function is given by K(·), where K is the kernel function (e.g., a Gaussian kernel) and h is the bandwidth parameter. The kernel density estimation result is the probability density function for each node, providing the basis for subsequent conditional probability and mutual information calculations. Based on the fitted probability density function, the conditional probability distribution for each pair of nodes is calculated. Let P(A|B) represent the conditional probability of node A given node B, and its calculation formula is:
[0086]
[0087] Where P(A,B) is the joint probability distribution and P(B) is the marginal probability distribution. A conditional independence test is performed on the conditional probability distribution data by calculating G for each pair of nodes. 2 The conditional dependency is quantified using statistical values. Let n be the sample size, and p... ij e represents the observed frequency. ij For the desired frequency, then G 2 The calculation formula is:
[0088]
[0089] G 2 The larger the value, the stronger the conditional dependency between nodes. G represents the sum of all pairs of nodes. 2 The values are stored in the conditional dependency matrix C. Based on the conditional dependency matrix C, the mutual information values between nodes are calculated. The mutual information value is calculated using the entropy difference between the joint probability distribution and the marginal probability distribution, and its formula is:
[0090] I(A;B)=H(A)+H(B)-H(A,B);
[0091] Where H(A) = -∑P(A)lnP(A) is the marginal entropy, and H(A,B) is the joint entropy. The mutual information value I(A;B) reflects the degree of information sharing between nodes A and B, and the result is stored as mutual information strength data. By setting an association strength threshold T, node pairs with mutual information values greater than T are selected, and an adjacency matrix M is constructed. The elements M of the adjacency matrix... ij This represents the strength of the association between nodes i and j. Based on the adjacency matrix, the maximum weighted spanning tree algorithm is used to construct an acyclic connection structure with mutual information values as edge weights. Let the weight of the tree structure be W, then the objective function is:
[0092] maxW = ∑ (i,j)∈E I(i;j);
[0093] Here, E is the set of edges. The generated tree topology requires further determination of the directionality of the edges. Causal relationships are determined by calculating conditional probabilities, and the causal direction between nodes is determined using the local Markov property. For example, if P(A|B)>P(B|A), then it is inferred that A→B. After determining the directionality, the set of directed edges E will be... d Integrating back into a tree-like topology, this ultimately forms the immune factor association network. The immune factor association network is a directed acyclic graph that reflects the associations and potential causal relationships between immune factors.
[0094] In one specific embodiment, the process of performing step S400 may specifically include the following steps:
[0095] The degree centrality of each node in the immune factor association network is calculated, and the node connectivity feature is obtained by normalizing the number of edges directly connected to the node by the total number of edges in the network.
[0096] The betweenness centrality of each node in the immune factor association network is calculated by dividing the number of paths passing through that node in the shortest paths between all pairs of nodes by the total number of shortest paths, and then obtaining the node betweenness characteristics.
[0097] For each node in the immune factor association network, the proximity centrality is calculated, and the node reachability feature is obtained by normalizing the sum of the inverses of the shortest distances from the node to all other nodes.
[0098] The node connectivity feature, node mediation feature, and node reachability feature are combined to obtain the immune factor association feature set. The immune factor association feature set is then concatenated with the multi-dimensional feature matrix of immune factor expression along the feature dimension to obtain the fused feature vector.
[0099] The fused feature vector is input into the input layer of a multi-layer neural network model. The input layer contains the same number of neurons as the fused feature vector. After processing by the Swish activation function and the AlphaDropout layer, the output feature of the input layer is obtained.
[0100] The output features of the input layer are input into the first convolutional block of the multi-layer neural network model. The first convolutional block contains two one-dimensional convolutional layers to obtain the first convolutional features.
[0101] The first convolutional feature is input into the second convolutional block of the multi-layer neural network model. The second convolutional block contains two one-dimensional convolutional layers. A residual connection is added between the two one-dimensional convolutional layers to obtain the second convolutional feature.
[0102] The second convolutional features were processed using GlobalAveragePooling to obtain the results of abnormal immune factor expression detection.
[0103] Specifically, based on the immune factor association network, the degree centrality of each node is calculated. Degree centrality is a fundamental feature in network analysis, used to measure the degree of direct connection between a node and other nodes. Let the set of nodes in the network be V, and the set of edges be E. For a node v... i Degree centrality C D (v i Calculated using the following formula:
[0104]
[0105] Where deg(v i ) is node v i The degree of a node is the number of edges directly connected to it, where |E| is the total number of edges in the network. Degree centrality is normalized so that the result falls within the interval [0,1]. For example, in a network with 6 nodes and 10 edges, if node v1 is connected to 4 other nodes, its degree centrality is... Calculate the betweenness centrality for each node. Betweenness centrality is determined by finding the shortest path between any two nodes in the network. i Its role as a network "bridge" is measured by the proportion of paths it occupies. Let σ jk Represents node v j and v k The number of shortest paths between them, σ jk (v i ) indicates passing through node v i The number of shortest paths, then the betweenness centrality C B (v i ) is defined as:
[0106]
[0107] The denominator is the total number of shortest paths. For example, in a network, if node v1 is included on 5 shortest paths, and the total number of shortest paths in the network is 20, then... Simultaneously, proximity centrality is calculated. Proximity centrality reflects the global reachability of a node by evaluating its distances to all other nodes in the network. Let d(v i ,v j ) represents node v i to node v j The shortest distance is close to centrality C. C (v i ) is defined as:
[0108]
[0109] Where |V| is the total number of nodes. For example, in a network with 6 nodes, if the distances from node v1 to other nodes are {1, 2, 1, 3, 2}, then its closeness centrality is:
[0110]
[0111] After calculating the degree centrality, betweenness centrality, and proximity centrality, they are combined into an immune factor association feature set F. Assuming the association network has 6 nodes (corresponding to immune factors CD3, CD4, CD8, CD19, CD25, CD56), and each feature contains 6 values, the feature set matrix form is as follows:
[0112]
[0113] The immune factor association feature set and the multi-dimensional feature matrix X of immune factor expression are concatenated along their feature dimensions to form a fused feature vector Z. Assuming X has a dimension of n×m and F has a dimension of n×p, the dimension of the fused Z is n×(m+p), where n is the number of samples and m+p is the total feature dimension after fusion. The fused feature vector Z is input into the input layer of a multi-layer neural network model, containing the same number of neurons as Z. The Swish activation function is applied to the input layer data, with the following form:
[0114] f(x) = x·sigmoid(x);
[0115] in The Swish function, with its smoothing and non-linear properties, is suitable for feature extraction in deep models. An Alpha Dropout layer is added to enhance the model's generalization ability by randomly masking the output of a subset of neurons. The features processed by the input layer are then passed to the first convolutional block, which contains two one-dimensional convolutional layers. The convolution operation in each layer is as follows:
[0116]
[0117] in These are the input features of the l-th layer. b represents the kernel weights. (l) This is the bias term. After two convolutional layers, the first convolutional feature is obtained. This first convolutional feature is then input into the second convolutional block. The second convolutional block is similar to the first, but a residual connection is added between the two convolutional layers. The output of the residual connection is:
[0118]
[0119] Where f(z) (l)The output of the convolutional layer is denoted as . Residual connections can effectively alleviate the vanishing gradient problem and enhance the training capability of deep networks. After the second convolutional block, the second convolutional feature is obtained. Global average pooling is then applied to the second convolutional feature, using the following formula:
[0120]
[0121] Where n is the feature length. Global average pooling maps high-dimensional features to fixed-length low-dimensional representations, generating abnormal immune factor expression detection results. Through this process, the expression characteristics and associated structures of immune factors are accurately captured, enabling effective detection of abnormal samples.
[0122] In one specific embodiment, the process of executing step S500 may specifically include the following steps:
[0123] A risk distribution matrix is constructed based on abnormal samples from the results of abnormal immune factor expression detection. The third and fourth moment statistics are calculated for each abnormal sample to obtain the distribution skewness and kurtosis values.
[0124] Jarque-Bera tests were performed on the skewness and kurtosis values to obtain the tail distribution characteristics. Based on the tail distribution characteristics, a conditional risk value function was constructed. The minimum tail bound probability was calculated using the expected value and variance to obtain the initial risk boundary.
[0125] To construct the constraint matrix of a semidefinite programming model for the initial risk boundary, the diagonal elements are set as risk confidence levels and the off-diagonal elements are set as covariance coefficients, resulting in a positive semidefinite matrix.
[0126] The objective function is constructed based on the positive semidefinite matrix. By adding linear inequality constraints and probability measure constraints, a standardized semidefinite programming problem is obtained.
[0127] The interior point method is used to solve the standardized semidefinite programming problem to obtain the risk optimization boundary. Based on the risk optimization boundary, a three-level risk assessment interval is constructed, and the risk level of abnormal samples is classified to obtain the risk classification result.
[0128] The risk grading results are correlated and mapped with the results of abnormal immune factor expression detection, and the specific risk level and risk probability value of each abnormal sample are labeled, outputting quantitative assessment data of immune factor abnormalities.
[0129] Specifically, abnormal samples are extracted from the results of abnormal immune factor expression detection, and a risk distribution matrix R is constructed. Assuming that n abnormal samples are identified in the detection, and each sample has a feature dimension of m, the risk distribution matrix takes the form:
[0130]
[0131] Where r i,jThis represents the outlier value of the i-th outlier sample in the j-th feature. To characterize the distribution of each outlier sample, the third and fourth moment statistics are calculated. The third moment (skewness) reflects the asymmetry of the distribution and is defined as:
[0132]
[0133] Where μ i Let σ be the mean of the i-th sample. i The standard deviation is denoted by 4; the fourth moment (kurtosis) reflects the steepness of the tail of the distribution and is defined as:
[0134]
[0135] Subtracting 3 is done to zero out the kurtosis of the normal distribution. The skewness value γ1 and kurtosis value γ2 of each outlier sample are calculated, forming the tail statistics of the sample. A Jarque-Bera test is performed based on the skewness and kurtosis values to determine whether the distribution approximates a normal distribution. The test statistic is defined as:
[0136]
[0137] Where n is the sample size. The calculated JB value is compared with the critical value of the chi-square distribution. If JB > 0. This rejects the hypothesis that the sample distribution is normally distributed, thus confirming the tail distribution characteristics. Based on the tail distribution characteristics, a conditional hazard value function is constructed. The conditional hazard value is the maximum possible loss at a given confidence level, defined as:
[0138] VaR α =inf{x∈R:P(X≤x)≥α};
[0139] Where α is the confidence level, and P(X≤x) is the cumulative distribution function. The expected loss after calculating the conditional risk value is defined as:
[0140] CVaR α =E[X|X>VaR α ];
[0141] The minimum tail bound probability is calculated by combining expectation and variance, defining the initial risk boundary. To optimize the risk boundary, a constraint matrix Q for the semidefinite programming model is constructed based on the initial risk boundary. Let the diagonal elements of Q represent the risk confidence levels q. ii The off-diagonal elements are the covariance coefficients q. ij Defined as:
[0142] q ij =Cov(r i ,r j );
[0143] Where Cov(r) i ,r j Let $\mathbf{i}$ represent the covariance between the $i$-th and $j$-th samples. The constructed $Q$ is a positive semi-definite matrix satisfying $Q ≥ 0$. After constructing the constraint matrix, the optimization objective function $f(x)$ is defined, aiming to minimize the risk boundary. Linear inequality constraints $Ax ≤ b$ and probability measure constraints $x$ are added. T Qx≤c, forming a standardized semidefinite programming problem:
[0144] minf(x) st Ax≤b,x T Qx≤c, Q≥0;
[0145] The interior-point method is used to solve the above problem. This method transforms the inequality constraints into a smooth optimization problem by introducing a logarithmic barrier function φ(x) = -∑ln(b-Ax), and iterates the variable x using Newton's step size. In each iteration, the dual gap g(x) is calculated:
[0146] g(x) = f(x) - d(x);
[0147] Where d(x) is the dual solution. When g(x) is less than the preset tolerance, the optimized risk boundary is obtained. Based on the risk optimization boundary, a three-level risk assessment interval is constructed: low-risk interval [0, VaR]. α ], medium risk range [VaR] α ,CVaR α High-risk range [CVaR] α ,∞). Calculate the risk probability p for abnormal samples. i The risk level is then matched to the corresponding risk range and assigned a risk level label (low, medium, high). The risk grading results are then correlated with the results of abnormal immune factor expression detection. For example, for an abnormal sample r... i Risk probability p i =0.95, which falls within the high-risk range and is marked as "high risk". The output quantitative assessment data of abnormal immune factors includes the risk level and risk probability of each abnormal sample.
[0148] In one specific embodiment, the process of performing the steps of solving the standardized semidefinite programming problem using the interior point method to obtain the risk optimization boundary, constructing a three-level risk assessment interval based on the risk optimization boundary, classifying the risk level of abnormal samples, and obtaining the risk classification result can specifically include the following steps:
[0149] For the standardized semidefinite programming problem, a logarithmic barrier function is constructed, and the inequality constraints are introduced into the objective function through a penalty function term to obtain the central path equation;
[0150] An augmented system matrix is constructed based on the central path equation. The KKT conditions for the original variables, dual variables, and slack variables are linearized to obtain the search direction system.
[0151] Solve the search direction system to obtain the Newton step size sequence, and update the original solution and dual solution based on the Newton step size sequence. Calculate the dual gap after each iteration. When the dual gap is less than the target value, the optimal risk boundary solution is obtained.
[0152] The conditional risk value is calculated for the optimal risk boundary solution to obtain VaR and CVaR values. The VaR and CVaR values are then used to construct a risk interval division standard and a three-level risk assessment interval is constructed.
[0153] For each abnormal sample in the abnormal immune factor expression detection results, a risk metric is calculated to obtain the sample risk probability. The sample risk probability is then compared and matched with the three-level risk assessment interval. Each abnormal sample is assigned a corresponding risk level label to obtain the risk classification result.
[0154] Specifically, in the standardized semidefinite programming problem, assume the objective function is f(x), and the constraints include the linear inequality constraint Ax≤b and the positive semidefinite constraint Q(x)≥0. The goal is to minimize f(x) to satisfy these constraints. To solve this problem within the interior-point method framework, the inequality constraint Ax≤b is transformed into a logarithmic barrier function form. By introducing a penalty function term, the objective function becomes smooth and differentiable. The logarithmic barrier function is defined as:
[0155]
[0156] Where b i A represents the upper bound of the constraint. i represents the coefficients of the corresponding variables, and x is the vector of optimization variables. The new objective function is:
[0157] F(x; μ) = f(x) + μφ(x);
[0158] Here, μ>0 is the barrier parameter, controlling the weight of the penalty function. As μ gradually decreases, the solution gradually approaches the feasible region boundary, forming a trajectory called the central path. Based on the above logarithmic barrier function, an augmented system matrix of the central path equation is constructed to describe the interaction between the original variable x, the dual variable λ, and the relaxed variable s during the optimization process. The necessary conditions for the optimization problem are expressed using the KKT (Karush-Kuhn-Tucker) conditions:
[0159]
[0160] Where s is the slack variable and λ is the dual variable. A first-order Taylor expansion of the KKT conditions yields a linearized search direction system:
[0161]
[0162] Where Δx, Δλ, and Δs are the search directions for the original, dual, and slack variables, respectively; S and Λ are diagonal matrices, composed of s and λ, respectively; and e is a vector of all ones. Solving the above linear system yields the Newton step-size sequence, which is used to update the original and dual solutions:
[0163] x (k+1) =x (k) +αΔx,λ (k+1) =λ (k) +αΔλ,s (k+1) =s (k) +αΔs;
[0164] Where α is the step size parameter, determined through line search. After each iteration, the duality gap is calculated:
[0165] g(x,λ)=s T λ;
[0166] When the duality gap is less than the preset tolerance ∈, the iteration stops, and the optimal risk boundary solution x is obtained. * Based on the optimal risk boundary solution x * Calculate the conditional risk values (VaR and CVaR). VaR is the maximum possible loss given a confidence level α, defined as:
[0167] VaR α =inf{x∈R:P(X≤x)≥α};
[0168] Where P(X≤x) is the cumulative distribution function. CVaR represents the average loss exceeding VaR, defined as:
[0169] CVaR α =E[X|X>VaR α ];
[0170] Risk interval division criteria were constructed using VaR and CVaR, including a low-risk interval [0, VaR]. α Medium-risk range [VaR] α ,CVaR α ] and high-risk range [CVaR α For each abnormal sample in the abnormal immune factor expression detection results, calculate its risk metric p. i For example, it can be calculated using the cumulative probability of its anomalous distribution:
[0171] p i =P(X≤r) i );
[0172] For each abnormal sample in the abnormal immune factor expression detection results, calculate its risk metric p. i For example, it can be calculated using the cumulative probability of its anomalous distribution:
[0173] p i =P(X≤r) i );
[0174] Where r i This is a measurement of an anomaly sample. The risk probability p... i The risk level is assigned a label based on the risk level range matched with the three-tier risk assessment interval. For example, if the risk probability p of an abnormal sample... i =0.92, and falls within the medium-risk range [VaR] α ,CVaR α If the result is negative, its risk level is "medium risk". The risk classification results are then correlated and mapped with the detection results to output the specific risk level and risk probability value for each abnormal sample.
[0175] Please see Figure 2 , Figure 2 This is a schematic block diagram of the structure of the liver cancer immunotherapy efficacy prediction system 200 provided in the embodiments of this application, as shown below. Figure 2 As shown, the predictive system 200 for the efficacy of immunotherapy for liver cancer includes:
[0176] The preprocessing module 210 is used to perform multiple immunofluorescence staining preprocessing on clinical samples and to collect forward scattering data, side scattering data and multi-channel fluorescence data of mononuclear cells to obtain an initial data matrix of cell expression.
[0177] Analysis module 220 is used to standardize the initial cell expression data matrix, construct an n-dimensional feature space, and perform principal component analysis to obtain a multi-dimensional feature matrix of immune factor expression.
[0178] The module 230 is used to perform Bayesian network analysis based on the multi-dimensional feature matrix of immune factor expression, calculate the conditional probability relationship and mutual information value between immune cell factors, and construct an immune factor association network according to a preset association strength threshold.
[0179] The detection module 240 is used to extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into the multi-layer neural network model for anomaly detection to obtain the abnormal detection result of immune factor expression.
[0180] Output module 250 is used to perform minimum tail boundary analysis on the detection results of abnormal expression of immune factors, and solve the optimal risk boundary through a semidefinite programming model to output quantitative assessment data of abnormal immune factors.
[0181] Through the collaborative efforts of the aforementioned components, and by constructing a multi-dimensional feature space and performing principal component analysis, effective dimensionality reduction and feature extraction of high-dimensional immune cytokine expression data were achieved, improving the efficiency and accuracy of data processing. A Bayesian network analysis method was used to construct an immune factor association network, revealing the conditional probability relationships and mutual information characteristics among immune cytokines, overcoming the limitations of traditional single-factor analysis methods. A multi-layer neural network model was introduced for anomaly detection, combining network centrality and expression features to enhance the ability to identify complex immune expression patterns. Minimum tail bound analysis and semidefinite programming models were used to achieve quantitative risk assessment of the prediction results. Multiplex immunofluorescence staining and flow cytometry were employed to ensure the accuracy and reproducibility of sample data collection. A three-level risk assessment mechanism was used to achieve refined grading of the prediction results.
[0182] Please see Figure 3 , Figure 3 This is a schematic block diagram of the structure of a liver cancer immunotherapy efficacy prediction device 300 provided in an embodiment of this application. The liver cancer immunotherapy efficacy prediction device 300 includes a processor 301 and a memory 302, which are connected through a system bus 303. The memory 302 may include a non-volatile storage medium and internal memory.
[0183] The non-volatile storage medium can store a computer program. The computer program includes program instructions that, when executed by the processor 301, cause the processor 301 to perform any of the aforementioned methods for predicting the efficacy of liver cancer immunotherapy.
[0184] The processor 301 provides computing and control capabilities to support the operation of the prediction device 300 for the overall efficacy of liver cancer immunotherapy.
[0185] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor 301, the processor 301 can execute any of the above-mentioned methods for predicting the efficacy of liver cancer immunotherapy.
[0186] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the liver cancer immunotherapy efficacy prediction device 300 involved in the present application. The specific liver cancer immunotherapy efficacy prediction device 300 may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0187] It should be understood that processor 301 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0188] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the liver cancer immunotherapy efficacy prediction device 300 described above can be referred to the corresponding process of the aforementioned liver cancer immunotherapy efficacy prediction method, and will not be repeated here.
[0189] This application also provides a computer-readable storage medium storing a computer program that, when executed by one or more processors, causes the one or more processors to implement the method for predicting the efficacy of liver cancer immunotherapy as provided in this application.
[0190] The computer-readable storage medium can be an internal storage unit of the liver cancer immunotherapy efficacy prediction device 300 described in the foregoing embodiments, such as a hard disk or memory of the liver cancer immunotherapy efficacy prediction device 300. The computer-readable storage medium can also be an external storage device of the liver cancer immunotherapy efficacy prediction device 300, such as a pluggable hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped with the liver cancer immunotherapy efficacy prediction device 300.
[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0193] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for predicting the efficacy of immunotherapy for liver cancer, characterized in that, include: Clinical samples were pretreated with multiplex immunofluorescence staining, and forward scattering data, side scattering data, and multichannel fluorescence data of mononuclear cells were collected to obtain an initial data matrix of cell expression. The initial data matrix of cell expression was standardized to construct an n-dimensional feature space, and principal component analysis was performed to obtain a multi-dimensional feature matrix of immune factor expression. Bayesian network analysis is performed based on the multi-dimensional feature matrix of the immune factor expression to calculate the conditional probability relationship and mutual information value between immune cell factors, and an immune factor association network is constructed according to a preset association strength threshold. Extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into a multi-layer neural network model for anomaly detection to obtain the immune factor expression anomaly detection result; The results of abnormal immune factor expression were analyzed using minimum tail bound analysis, and the optimal risk boundary was solved using a semidefinite programming model to output quantitative assessment data of immune factor abnormalities.
2. The method for predicting the efficacy of liver cancer immunotherapy according to claim 1, characterized in that, The clinical samples underwent multiplex immunofluorescence staining pretreatment, and forward scattering data, side scattering data, and multichannel fluorescence data of mononuclear cells were collected to obtain an initial cell expression data matrix, including: Clinical samples were subjected to density gradient centrifugation to obtain an intermediate layer sample containing mononuclear cells, and the intermediate layer sample was washed by centrifugation with phosphate buffer to obtain a purified nucleated cell sample. The purified nuclear cell sample was subjected to flow cytometry cell counting, and the viable nuclear cells were obtained by FSC-A and SSC-A dual-parameter scatter plot analysis. The density of the active nucleated cells was adjusted to 1×10⁻⁶. 6 Cells were labeled with immunofluorescence by adding monoclonal antibodies against CD3, CD4, CD8, CD19, CD25, and CD56 at a concentration of 1 / mL. The immunofluorescence-labeled cells were excited by a 488nm laser from a flow cytometer. The power of the 488nm laser was set to 20mW. The FSC signal was collected as forward scattering data and the SSC signal was collected as side scattering data to obtain cell morphology parameters. The fluorescence intensity signals of six fluorescence channels (FITC, PE, PerCP, APC, PE-Cy7, and APC-Cy7) of the immunofluorescently labeled cells were collected to obtain multi-channel fluorescence data; Compensation matrix calculations were performed on the cell morphology parameters and the multi-channel fluorescence data. The compensation coefficients were determined using single-stain control sample data to obtain the compensated multi-parameter data. The compensated multi-parameter data were arranged in the order of FSC, SSC, FITC-CD3, PE-CD4, PerCP-CD8, APC-CD19, PE-Cy7-CD25, and APC-Cy7-CD56 to construct an initial cell expression data matrix.
3. The method for predicting the efficacy of liver cancer immunotherapy according to claim 2, characterized in that, The initial cell expression data matrix is standardized to construct an n-dimensional feature space, and principal component analysis is performed to obtain a multi-dimensional feature matrix of immune factor expression, including: The FSC, SSC, and six fluorescence channel data in the initial cell expression data matrix were normalized by maximum and minimum values to obtain a normalized eigenvalue matrix. The features of each dimension of the normalized eigenvalue matrix are centered to obtain a standardized data matrix. Based on the standardized data matrix, an n-dimensional feature space is constructed using CD3, CD4, CD8, CD19, CD25, and CD56 as labels. Data from each dimension is then projected onto the n-dimensional feature space to obtain n-dimensional projected data. Based on the n-dimensional projection data, the covariance matrix between features is calculated, and an n×n-dimensional correlation coefficient matrix is constructed using the correlation between n immune cell factors to obtain the feature correlation matrix. Perform eigenvalue decomposition on the feature correlation matrix to obtain a set of principal component eigenvalues, and calculate the contribution rate and cumulative contribution rate of each eigenvalue based on the set of principal component eigenvalues to obtain a set of principal component eigenvectors; The principal component feature vector group is subjected to Schmitt orthogonalization to construct an orthogonalized feature vector matrix, and the n-dimensional projection data is linearly transformed to obtain a multi-dimensional feature matrix of immune factor expression.
4. The method for predicting the efficacy of liver cancer immunotherapy according to claim 3, characterized in that, The step of performing Bayesian network analysis based on the multi-dimensional feature matrix of the immune factor expression to calculate the conditional probability relationships and mutual information values between immune cell factors, and constructing an immune factor association network according to a preset association strength threshold, includes: A Bayesian network structure was constructed based on the multidimensional feature matrix of the immune factor expression, and each immune cytokine, CD3, CD4, CD8, CD19, CD25, and CD56, was set as a network node to obtain an immune factor node set. The conditional probability distribution between each pair of nodes is calculated based on the immune factor node set, and the probability density of each node is fitted using the kernel density estimation method to obtain the node probability distribution data. The conditional independence test is performed on the node probability distribution data, the G² statistic value of each pair of nodes is calculated, the conditional dependency matrix is obtained, and the mutual information value between nodes is calculated based on the conditional dependency matrix. The mutual information content is calculated using the entropy difference between the joint probability distribution and the marginal probability distribution, and the mutual information intensity data is obtained. A correlation strength threshold is set for the mutual information strength data, and node pairs with a correlation strength greater than the correlation strength threshold are selected to construct an adjacency matrix to obtain an initial connection structure. The maximum weight spanning tree algorithm is then executed based on the initial connection structure, and the acyclic connection structure is constructed using the mutual information value as the edge weight to obtain a tree topology. The directionality of the edges in the tree topology is determined by conditional probability calculation, and the causal relationship is determined by the local Markov property to obtain a set of directed edges. The set of directed edges is then integrated into the tree topology to obtain the immune factor association network.
5. The method for predicting the efficacy of liver cancer immunotherapy according to claim 4, characterized in that, The process involves extracting the immune factor association feature set from the immune factor association network, and inputting the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into a multi-layer neural network model for anomaly detection, thereby obtaining anomaly detection results for immune factor expression, including: The degree centrality of each node in the immune factor association network is calculated, and the node connectivity feature is obtained by normalizing the number of edges directly connected to the node by the total number of edges in the network. The betweenness centrality of each node in the immune factor association network is calculated by dividing the number of paths passing through that node in the shortest paths between all pairs of nodes by the total number of shortest paths, and the node betweenness feature is obtained. For each node in the immune factor association network, the proximity centrality is calculated, and the node reachability feature is obtained by normalizing the sum of the inverses of the shortest distances from the node to all other nodes. The node connectivity feature, the node mediation feature, and the node reachability feature are combined to obtain an immune factor association feature set. The immune factor association feature set is then concatenated with the multi-dimensional feature matrix expressing the immune factor in terms of feature dimensions to obtain a fused feature vector. The fused feature vector is input into the input layer of a multi-layer neural network model. The input layer contains the same number of neurons as the fused feature vector. After processing by the Swish activation function and the Alpha Dropout layer, the output features of the input layer are obtained. The output features of the input layer are input into the first convolutional block of the multi-layer neural network model. The first convolutional block contains two one-dimensional convolutional layers to obtain the first convolutional features. The first convolutional feature is input into the second convolutional block of the multi-layer neural network model. The second convolutional block contains two one-dimensional convolutional layers. A residual connection is added between the two one-dimensional convolutional layers to obtain the second convolutional feature. The second convolutional feature is processed using Global Average Pooling to obtain the results of abnormal immune factor expression detection.
6. The method for predicting the efficacy of liver cancer immunotherapy according to claim 5, characterized in that, The method involves performing minimum tail boundary analysis on the abnormal expression detection results of the immune factors, and solving for the optimal risk boundary using a semidefinite programming model to output quantitative assessment data of immune factor abnormalities, including: Based on the abnormal samples in the abnormal immune factor expression detection results, a risk distribution matrix is constructed, and the third and fourth moment statistics are calculated for each abnormal sample to obtain the distribution skewness and kurtosis values. The Jarque-Bera test is performed on the skewness and kurtosis values to obtain the tail distribution characteristics. A conditional risk value function is constructed based on the tail distribution characteristics. The minimum tail bound probability is calculated by the expectation and variance to obtain the initial risk boundary. A semidefinite programming model constraint matrix is constructed for the initial risk boundary, with diagonal elements representing risk confidence and off-diagonal elements representing covariance coefficients, resulting in a positive semidefinite matrix. Based on the semidefinite matrix, construct the optimization objective function, add linear inequality constraints and probability measure constraints, and obtain the standardized semidefinite programming problem; The standardized semidefinite programming problem is solved using the interior point method to obtain the risk optimization boundary. Based on the risk optimization boundary, a three-level risk assessment interval is constructed, and the risk level of abnormal samples is classified to obtain the risk classification result. The risk grading results are correlated and mapped with the abnormal expression detection results of immune factors, and the specific risk level and risk probability value of each abnormal sample are labeled, and quantitative assessment data of immune factor abnormalities are output.
7. The method for predicting the efficacy of liver cancer immunotherapy according to claim 6, characterized in that, The standardized semidefinite programming problem is solved using the interior-point method to obtain the risk optimization boundary. Based on the risk optimization boundary, a three-level risk assessment interval is constructed to classify the risk level of abnormal samples, resulting in a risk classification result, including: A logarithmic obstacle function is constructed for the standardized semidefinite programming problem, and the inequality constraints are introduced into the objective function through a penalty function term to obtain the central path equation. Based on the central path equation, an augmented system matrix is constructed, and the KKT conditions for the original variables, dual variables, and slack variables are linearized to obtain the search direction system. The search direction system is solved to obtain the Newton step size sequence, and the original solution and dual solution are updated based on the Newton step size sequence. The dual gap is calculated after each iteration. When the dual gap is less than the target value, the optimal risk boundary solution is obtained. The optimal risk boundary solution is subjected to conditional risk value calculation to obtain VaR and CVaR values, and the VaR and CVaR values are used to construct risk interval division criteria to construct a three-level risk assessment interval. For each abnormal sample in the abnormal immune factor expression detection results, a risk metric is calculated to obtain the sample risk probability. The sample risk probability is then compared and matched with the three-level risk assessment interval. Each abnormal sample is assigned a corresponding risk level label to obtain the risk grading result.
8. A predictive system for the efficacy of immunotherapy for liver cancer, characterized in that, A method for predicting the efficacy of liver cancer immunotherapy as described in any one of claims 1-7, comprising: The preprocessing module is used to perform multiple immunofluorescence staining preprocessing on clinical samples and to collect forward scattering data, side scattering data and multi-channel fluorescence data of mononuclear cells to obtain an initial data matrix of cell expression. The analysis module is used to standardize the initial data matrix of cell expression, construct an n-dimensional feature space, and perform principal component analysis to obtain a multi-dimensional feature matrix of immune factor expression. The construction module is used to perform Bayesian network analysis based on the multi-dimensional feature matrix of the immune factor expression, calculate the conditional probability relationship and mutual information value between immune cell factors, and construct an immune factor association network according to a preset association strength threshold. The detection module is used to extract the immune factor association feature set of the immune factor association network, and input the immune factor association feature set and the multi-dimensional feature matrix of immune factor expression into a multi-layer neural network model for anomaly detection to obtain the immune factor expression anomaly detection result. The output module is used to perform minimum tail boundary analysis on the abnormal expression detection results of the immune factors, and solve the optimal risk boundary through a semidefinite programming model to output quantitative assessment data of immune factor abnormalities.
9. A device for predicting the efficacy of immunotherapy for liver cancer, characterized in that, The device for predicting the efficacy of liver cancer immunotherapy includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the liver cancer immunotherapy efficacy prediction device to execute the liver cancer immunotherapy efficacy prediction method as described in any one of claims 1-7.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the method for predicting the efficacy of liver cancer immunotherapy as described in any one of claims 1-7.
Citation Information
Patent Citations
Flow cytometry sample flow velocity control method, medium and system
CN117607009A
Ophthalmology data analysis method and device based on crowd queue analysis and medium
CN119742077A