PRAK signaling pathway analysis method and system
By constructing a multi-omics data matrix and graph neural network algorithm for the PRAK signaling pathway, the problem of the inability to intelligently process multi-omics data in existing technologies is solved, enabling accurate quantitative assessment of the PRAK signaling pathway and optimization of target combinations, thereby improving the scientific nature and effectiveness of treatment plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies cannot intelligently process multi-omics data, lack the ability to quantitatively predict the overall activity status of the PRAK signaling pathway, have difficulty identifying the dynamic response patterns of the pathway under different disease conditions, and lack systematic target optimization and selection methods.
Phosphorylation data of circRNA, miRNA, and PRAK pathway proteins were collected using high-throughput sequencing technology. Pathway connectivity matrices were constructed along the P38-PRAK-HSP27 signaling axis, the VEGFA-PRAK activation axis, and the PRAK-GTPBP4-RhoA regulatory axis. Graph neural network algorithms were used to predict pathway activity status. Combined with K-means clustering and synergistic effect prediction algorithms, pathway response characteristics of chemotherapy-resistant, apoptosis-sensitive, and proliferation-regulated pathways were identified to optimize the combination of therapeutic targets.
It enables precise quantitative assessment of the PRAK signaling pathway and dynamic identification of response patterns, providing a scientific basis for disease subtyping and personalized treatment strategies, and improving the scientific nature and effectiveness of treatment plans.
Smart Images

Figure CN121350496B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a PRAK signal pathway analysis method and system. BACKGROUND
[0002] In the prior art, PRAK signal pathway analysis mainly relies on traditional molecular biology experimental methods, including Western blot detection of protein phosphorylation state, qRT-PCR analysis of gene expression level, and cell function experiment verification of pathway activity. These methods can detect the expression changes of a single molecule or a small number of molecules, and preliminarily verify the interaction relationship between molecules through in vitro experiments, thereby laying a foundation for understanding the basic function of the PRAK pathway.
[0003] However, the prior art has significant deficiencies. Traditional experimental methods can only analyze the PRAK pathway in fragments, cannot systematically integrate multi-omics data such as circRNA, miRNA and protein phosphorylation, lack quantitative prediction ability of the overall activity state of the pathway, and the existing methods are difficult to identify the dynamic response mode of the pathway under different disease conditions, and cannot establish the precise correlation between the pathway activity and the disease phenotype.
[0004] Based on the above technical limitations, further analysis found that the prior art lacks intelligent pathway analysis methods to process multi-dimensional molecular data and predict pathway activity states, lacks disturbance response feature recognition technology based on artificial intelligence algorithms to distinguish the pathway response mode of different disease types, and lacks a systematic target optimization selection method to determine the best multi-target joint intervention strategy. SUMMARY
[0005] The present application provides a PRAK signal pathway analysis method and system, which solves the technical problem that the existing PRAK signal pathway analysis method cannot intelligently process multi-omics data and accurately predict the pathway activity state.
[0006] In a first aspect, the present application provides a PRAK signal pathway analysis method, comprising: performing multi-omics data acquisition and processing on a biological sample by high-throughput sequencing technology to obtain a comprehensive data set comprising circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data; performing molecular interaction relationship mining processing according to the comprehensive data set to obtain a pathway connection matrix comprising P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulation axis; performing pathway activity state prediction processing on the pathway connection matrix by graph neural network algorithm to obtain PRAK pathway activation intensity values under different conditions; performing perturbation response feature extraction processing according to the PRAK pathway activation intensity values to obtain pathway response feature identifiers comprising chemoresistance type, apoptosis sensitivity type and proliferation regulation type; and performing treatment target matching processing on the pathway response feature identifiers to obtain PRAK pathway intervention target combinations for specific disease phenotypes.
[0007] In a second aspect, the present application provides a PRAK signal pathway analysis system, comprising:
[0008] An acquisition module is configured to perform multi-omics data acquisition and processing on a biological sample by high-throughput sequencing technology to obtain a comprehensive data set comprising circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data;
[0009] A mining module is configured to perform molecular interaction relationship mining processing according to the comprehensive data set to obtain a pathway connection matrix comprising P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulation axis;
[0010] A prediction module is configured to perform pathway activity state prediction processing on the pathway connection matrix by graph neural network algorithm to obtain PRAK pathway activation intensity values under different conditions;
[0011] An extraction module is configured to perform perturbation response feature extraction processing according to the PRAK pathway activation intensity values to obtain pathway response feature identifiers comprising chemoresistance type, apoptosis sensitivity type and proliferation regulation type;
[0012] A matching module is configured to perform treatment target matching processing on the pathway response feature identifiers to obtain PRAK pathway intervention target combinations for specific disease phenotypes.
[0013] In a third aspect, a PRAK signaling pathway analysis device is provided, comprising a memory and at least one processor, the memory storing instructions; the at least one processor invoking the instructions in the memory to cause the PRAK signaling pathway analysis device to perform the PRAK signaling pathway analysis method described above.
[0014] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing instructions, which, when executed on a computer, cause the computer to perform the PRAK signaling pathway analysis method described above.
[0015] In the technical solutions provided in the present application, the circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data are systematically integrated through the integration of multi-omics data acquisition and processing technology, overcoming the limitation of traditional methods that can only analyze a single molecule type. Through molecular interaction relationship mining and processing, a complete pathway network containing the P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis and the PRAK-GTPBP4-RhoA regulation axis is constructed, solving the problem that the prior art cannot comprehensively describe the complex regulation relationship of the PRAK pathway. The application of graph neural network algorithm enables the pathway activity state prediction processing to accurately quantify the PRAK pathway activation intensity value under different conditions, realizing the accurate numerical evaluation of the pathway activity compared with the traditional qualitative analysis method. The disturbance response feature extraction processing identifies three response modes of chemotherapy resistance, apoptosis sensitivity and proliferation regulation, laying a scientific foundation for disease typing and personalized treatment strategy making. The treatment target matching processing systematically evaluates the intervention potential of key targets such as HSP27 and GTPBP4 and optimizes the target combination, realizing the transition from single target to multi-target joint intervention strategy, significantly improving the scientificity and effectiveness of the treatment scheme.
[0016] The node feature aggregation and graph attention mechanism weighting characteristics of the graph neural network algorithm enable the algorithm to adaptively learn the importance weight of different molecular nodes in the pathway, dynamically adjust the contribution of the P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis, and the PRAK-GTPBP4-RhoA regulation axis in the prediction of the overall pathway activity, and ensure the accuracy and biological rationality of the pathway analysis results. The application of the K-means clustering algorithm in the identification of perturbation response characteristics fully utilizes its unsupervised learning ability, can automatically discover hidden response patterns without prior labels, and the iterative optimization characteristics of the algorithm ensure the stability and reproducibility of the clustering results, providing reliable data support for establishing the classification standards of chemotherapy-resistant, apoptosis-sensitive, and proliferation-regulated types. The combination application of the synergistic effect prediction algorithm and the optimization selection algorithm reflects the unique advantages of artificial intelligence in multi-objective optimization problems, by simultaneously considering the maximization of therapeutic effect and the minimization of toxicity and side effects, the global optimization of the treatment scheme is realized. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative labor.
[0018] Figure 1 An embodiment schematic diagram of the PRAK signal pathway analysis method in the embodiments of the present application;
[0019] Figure 2 An embodiment schematic diagram of the PRAK signal pathway analysis system in the embodiments of the present application;
[0020] Figure 3 The structure schematic block diagram of the PRAK signal pathway analysis device in the embodiments of the present application. DETAILED DESCRIPTION
[0021] The embodiments of the present application provide a PRAK signal pathway analysis method and system. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0022] For ease of understanding, the specific processes of the embodiments of the present application are described below. Please refer to Figure 1 One embodiment of the PRAK signal pathway analysis method in the embodiments of the present application includes:
[0023] Step S101, performing multi-omics data acquisition and processing on a biological sample by high-throughput sequencing technology to obtain a comprehensive data set containing circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data;
[0024] Step S102, performing molecular interaction relationship mining processing according to the comprehensive data set to obtain a pathway connection matrix containing P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulation axis;
[0025] Step S103, performing pathway activity state prediction processing on the pathway connection matrix by a graph neural network algorithm to obtain PRAK pathway activation intensity values under different conditions;
[0026] Step S104, performing perturbation response feature extraction processing according to the PRAK pathway activation intensity values to obtain pathway response feature identifiers containing chemoresistance type, apoptosis sensitivity type and proliferation regulation type;
[0027] Step S105, performing treatment target matching processing on the pathway response feature identifiers to obtain a PRAK pathway intervention target combination for a specific disease phenotype.
[0028] It can be understood that the execution subject of the present application can be a PRAK signal pathway analysis system, and can also be a terminal or a server, which is not limited here. The embodiments of the present application take the server as the execution subject for example.
[0029] Specifically, high-throughput sequencing technology collects multi-dimensional biological data, RNA-seq sequencing obtains expression data of circular RNA, and forms a circRNA expression profile after quality control algorithm standardization, wherein the quality control algorithm eliminates batch effects by removing low-quality reads and performing Z-score normalization processing. Small RNA sequencing obtains microRNA data, and target gene prediction algorithm is used to analyze the binding sites of miRNA and mRNA to form miRNA regulation data. Proteomics technology detects the phosphorylation state of key proteins such as PRAK, HSP27 and GTPBP4, and mass spectrometry analysis identifies phosphorylation sites and quantifies modification levels. Data fusion algorithm matches and integrates three types of data according to sample ID to form a comprehensive data set.
[0030] Molecular interaction relationship mining is performed, Pearson correlation algorithm is used to calculate the linear correlation between any two molecules in the comprehensive data set, and a correlation coefficient matrix is generated. Threshold screening algorithm sets the standard that the absolute value of the correlation coefficient is greater than 0.6 and the P value is less than 0.05, and extracts significant associated molecule pairs from the correlation coefficient matrix. Based on the significant associated molecule pairs, three regulation axes are constructed: P38-PRAK-HSP27 signal axis is formed by identifying the upstream regulation of P38 kinase on PRAK and the phosphorylation regulation of PRAK on HSP27, VEGFA-PRAK activation axis is established by verifying the positive correlation between VEGFA and PRAK phosphorylation state, and PRAK-GTPBP4-RhoA regulation axis is constructed by analyzing the phosphorylation regulation of PRAK on GTPBP4 and the influence of GTPBP4 on RhoA-GTP hydrolysis activity. Matrix integration algorithm combines the connection weights of the three regulation axes to form a pathway connection matrix.
[0031] Graph neural network algorithm is used to predict pathway activity, and graph convolution layer takes P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulation axis in the pathway connection matrix as graph structure input, and calculates the feature representation of each molecular node through neighborhood aggregation function. Graph attention mechanism assigns different weights to the three regulation axes, and dynamically adjusts the importance of each axis based on molecular expression level and phosphorylation state. Importance score algorithm calculates the contribution of the three regulation axes according to the attention weight, and full connection layer fuses multi-axis information into pathway overall activity prediction value through nonlinear transformation, and Sigmoid activation function maps the prediction value to PRAK pathway activation intensity value between 0 and 1.
[0032] The disturbance response feature is extracted, and the time series analysis algorithm processes the time change of the PRAK pathway activation intensity value to identify the dynamic characteristics of the response mode. The parameter extraction algorithm calculates three key parameters of response amplitude, response time and recovery time from the pathway dynamic response curve to form a disturbance response feature vector. The K-means clustering algorithm divides the disturbance response feature vector into three clusters, and according to the characteristic mode of the cluster center, the chemotherapy resistance type, the apoptosis sensitivity type and the proliferation regulation type are corresponded respectively. The biological annotation algorithm combines the expression changes of Bax, Bcl-2 and other apoptosis markers to establish classification standards for the three cluster centers, and the pattern matching algorithm assigns the pathway response feature identifier to the input sample according to the classification standard.
[0033] The treatment target matching is performed, and the disease phenotype correlation analysis algorithm establishes a mapping relationship between the chemotherapy resistance type, the apoptosis sensitivity type, the proliferation regulation type and the specific disease state. The intervention assessable algorithm analyzes the operability of two key target points HSP27 and GTPBP4, calculates the intervention priority of each target point combined with the protein structure characteristics and drug binding site information. The drug-target affinity calculation algorithm predicts the binding strength of candidate drug molecules to three target points through molecular docking simulation, and the synergistic effect prediction algorithm evaluates the joint intervention effect of the HSP27-GTPBP4 target point combination. The optimization selection algorithm selects the PRAK pathway intervention target point combination for a specific disease phenotype from the candidate target point combination through five sub-steps of numerical standardization, weight calculation, multi-objective optimization function construction, genetic algorithm optimization and Pareto front screening.
[0034] In a specific embodiment, the process of performing step S101 can specifically include the following steps:
[0035] The biological sample is subjected to RNA-seq sequencing processing to obtain circular RNA raw expression data;
[0036] The circular RNA raw expression data is input into the quality control algorithm for standardization processing to obtain a circRNA expression profile;
[0037] The biological sample is subjected to small RNA sequencing processing to obtain small RNA raw expression data;
[0038] The small RNA raw expression data is subjected to target gene prediction analysis processing to obtain miRNA regulation data;
[0039] The biological sample is subjected to phosphorylated protein detection processing based on proteomics technology to obtain PRAK pathway protein phosphorylation data;
[0040] The circRNA expression profile, the miRNA regulation data and the PRAK pathway protein phosphorylation data are subjected to data fusion processing to obtain a comprehensive data set.
[0041] Specifically, the RNA-seq sequencing process adopts a high-throughput sequencing platform to construct and sequence the extracted total RNA library, generating raw sequencing data containing linear RNA and circular RNA, wherein the circular RNA raw expression data is obtained by identifying the reverse splicing site. The quality control algorithm first filters the reads of the circular RNA raw expression data, removing sequence fragments with a quality value less than 20, then counts the expression abundance of each circular RNA in different samples through the reads count algorithm, and then converts the original count value to a standardized expression value using the Z-score standardization algorithm, wherein the Z-score standardization achieves data normalization by calculating the difference between each circular RNA expression value and the average value divided by the standard deviation, and finally forms a circRNA expression profile matrix. The small RNA sequencing process specifically enriches and sequences small RNA molecules with a length of 18-30 nucleotides, generating raw expression data containing miRNA, siRNA and other small RNA molecules, and the microRNA is annotated and quantified by comparing with the miRBase database. The target gene prediction analysis process adopts a seed sequence matching algorithm, and based on the complementary pairing analysis of the seed region of miRNA and the 3'UTR region of target mRNA, the regulation relationship between miRNA and target gene is predicted by calculating the binding free energy and conservation score, forming miRNA regulation data containing miRNA-mRNA pairing information.
[0042] The proteomics technology adopts a mass spectrometry analysis platform to extract and digest proteins from biological samples, specifically captures phosphorylated peptides through phosphopeptide enrichment technology, and identifies the phosphorylation sites of PRAK, HSP27, GTPBP4 and other key proteins through mass spectrometry detection and quantifies the phosphorylation level, wherein the phosphorylation state of the Thr182 site of the PRAK protein is verified by a specific phosphorylation antibody, the phosphorylation of the Ser78 and Ser82 sites of the HSP27 protein is quantified by mass spectrometry intensity calculation, and the phosphorylation modification of the GTPBP4 protein is quantified by phosphorylation-specific mass spectrometry tag analysis, finally forming a PRAK pathway protein phosphorylation data matrix.
[0043] The data fusion processing adopts a sample ID matching algorithm to associate and integrate circRNA expression profiles, miRNA regulation data and PRAK pathway protein phosphorylation data according to the same sample identifier, unifies the numerical ranges of different omics data to the same order of magnitude through a data type standardization algorithm, fills in the missing data points in some samples using a missing value imputation algorithm, and finally generates a comprehensive data set matrix containing three types of molecular data. Taking triple-negative breast cancer patient samples as an example, specific circRNA expression was found to be significantly down-regulated in the patient's tumor tissue detected by RNA-seq, while small RNA sequencing showed that the expression of related miRNAs was up-regulated, and proteomic analysis detected that the phosphorylation level of PRAK protein Thr182 site was reduced while GGNBP2 protein expression was down-regulated. The data fusion algorithm integrates the circRNA expression value, miRNA regulation intensity and protein phosphorylation level of the same patient into a multi-dimensional molecular feature vector of the patient, and the feature vectors of multiple patient samples are combined to form a comprehensive data set for subsequent artificial intelligence analysis.
[0044] In a specific embodiment, the process of performing step S102 can specifically include the following steps:
[0045] Pearson correlation calculation and processing are performed on the circRNA expression profiles, miRNA regulation data and PRAK pathway protein phosphorylation data in the comprehensive data set to obtain a molecular correlation coefficient matrix;
[0046] Threshold screening processing is performed on the molecular correlation intensity according to the correlation coefficient matrix to obtain a significant correlation molecule pair with a correlation coefficient absolute value greater than 0.6 and a P value less than 0.05;
[0047] Based on the significant correlation molecule pair, P38-PRAK-HSP27 signal axis is constructed by performing phosphorylation cascade relationship construction processing on P38 and PRAK, and PRAK and HSP27;
[0048] VEGFA-PRAK activation axis is obtained by performing VEGFA and PRAK activation relationship verification analysis processing on the significant correlation molecule pair;
[0049] PRAK-GTPBP4-RhoA regulatory axis is obtained by performing PRAK and GTPBP4, GTPBP4 and RhoA regulatory relationship construction processing on the significant correlation molecule pair;
[0050] The P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulatory axis are integrated into a matrix to obtain a pathway connection matrix.
[0051] Specifically, the molecular interaction relationship mining process analyzes the linear correlation between different molecules in the integrated dataset by Pearson correlation calculation. The Pearson correlation algorithm calculates the covariance of any two molecular variables in the circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data divided by the product of the standard deviations of the two variables, generating a correlation coefficient with a value range of minus one to plus one. A positive value indicates a positive correlation, while a negative value indicates a negative correlation. The closer the absolute value of the correlation coefficient is to one, the stronger the linear correlation. The algorithm traverses all molecular pairs in the integrated dataset, calculates the correlation of the expression values of each pair of molecules in different samples, forms a symmetric matrix containing all inter-molecular correlation coefficients, and the matrix rows and columns correspond to different molecules, respectively. The matrix elements are the correlation coefficient values of the corresponding molecular pairs.
[0052] The threshold screening process performs statistical significance test and intensity filtering on the inter-molecular correlation coefficient matrix. First, calculate the P value of each correlation coefficient. The P value evaluates the statistical significance of the correlation through the t-test statistic. The t-test statistic is equal to the correlation coefficient multiplied by the square root of the sample size minus two divided by the square root of one minus the square of the correlation coefficient. The screening algorithm also checks whether the absolute value of the correlation coefficient is greater than the set threshold of zero point six and whether the P value is less than zero point zero five. Molecular pairs that meet both conditions are identified as significantly associated molecular pairs, and molecular pairs that do not meet the conditions are excluded. Finally, a set of molecular pairs with strong correlation and statistical significance is obtained.
[0053] The P38-PRAK-HSP27 signal axis construction process identifies the upstream regulation relationship between P38 kinase and PRAK kinase and the downstream phosphorylation relationship between PRAK kinase and HSP27 heat shock protein from the significantly associated molecular pairs. The algorithm retrieves the positive correlation pairs of P38 protein expression and PRAK protein phosphorylation state in the set of significantly associated molecular pairs, confirming that P38 activates the phosphorylation regulation of PRAK as an upstream kinase. At the same time, it retrieves the positive correlation pairs of PRAK protein phosphorylation state and HSP27 protein phosphorylation state, confirming that PRAK phosphorylates and modifies HSP27 as a downstream kinase. The cascade relationship construction algorithm connects the P38-PRAK regulation relationship and the PRAK-HSP27 regulation relationship according to the time sequence of signal transmission, forming a three-level phosphorylation cascade signal axis from P38 to PRAK to HSP27.
[0054] VEGFA-PRAK activation axis verification analysis is specifically to verify the activation regulation relationship between VEGFA and PRAK kinase. The verification algorithm searches for the positive correlation between the expression level of VEGFA protein and the phosphorylation level of PRAK protein in the significantly associated molecule pair, and verifies the biological reasonableness of the regulation relationship by calculating the consistency of the expression change trend of the two under different cell stimulation conditions. The algorithm also checks the matching degree of the VEGFA receptor signal and the PRAK activation time sequence, confirms that VEGFA activates PRAK upstream through its receptor signaling pathway, and finally establishes VEGFA as a secondary signal axis to activate PRAK kinase.
[0055] PRAK-GTPBP4-RhoA regulation axis construction process analyzes the phosphorylation regulation of PRAK kinase on GTPBP4 protein and the regulatory effect of GTPBP4 on RhoA GTPase activity. The algorithm extracts the positive correlation between the phosphorylation state of PRAK and the phosphorylation state of GTPBP4 from the significantly associated molecule pair, confirms that PRAK directly phosphorylates and modifies GTPBP4, and then extracts the negative correlation between the phosphorylation state of GTPBP4 and the activity state of RhoA-GTP, confirms that phosphorylated GTPBP4 inhibits the activation state of RhoA by enhancing GTP hydrolysis activity. The regulation relationship construction algorithm connects the positive phosphorylation regulation of PRAK on GTPBP4 and the negative activity regulation of GTPBP4 on RhoA in series to form a three-level regulation axis of PRAK indirectly inhibiting the activity of RhoA through phosphorylated GTPBP4.
[0056] The matrix integration process combines the three constructed signal regulation axes to form a unified pathway connection matrix. The integration algorithm first assigns a weight coefficient to each regulation axis, and the weight coefficient is calculated based on the average correlation coefficient strength of the molecule pairs in each axis, and then constructs an adjacency matrix containing all related molecules. Each element in the matrix represents the regulation strength between the corresponding molecule pair. The connection weights of the P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis and the PRAK-GTPBP4-RhoA regulation axis are filled into the corresponding matrix positions according to the respective weight coefficients, forming a pathway connection matrix reflecting the complete topological structure of the PRAK signaling pathway.
[0057] In a specific embodiment, the process of performing step S103 can specifically include the following steps:
[0058] The P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis and the PRAK-GTPBP4-RhoA regulation axis in the pathway connection matrix are input into the graph convolution layer for node feature aggregation processing to obtain node embedding vectors containing neighborhood information;
[0059] The node embedding vectors are weighted by a graph attention mechanism to obtain attention weight distributions of the P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis, and the PRAK-GTPBP4-RhoA regulation axis.
[0060] The P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis, and the PRAK-GTPBP4-RhoA regulation axis are scored based on the attention weight distributions to obtain an importance ranking list of the three regulation axes.
[0061] The importance ranking list is input into a fully connected layer for nonlinear transformation to obtain a pathway overall activity prediction value.
[0062] The pathway overall activity prediction value is mapped by a Sigmoid activation function to obtain a PRAK pathway activation intensity value.
[0063] Specifically, the node feature aggregation process uses a graph convolution layer to extract features from the three regulation axes in the pathway connection matrix. The graph convolution layer takes the P38-PRAK-HSP27 signal axis, the VEGFA-PRAK activation axis, and the PRAK-GTPBP4-RhoA regulation axis as input data of the graph structure, where each protein molecule is a node in the graph, and the regulation relationship between molecules is an edge connecting the nodes. The graph convolution layer includes 2 layers of graph convolution structure. The first layer of graph convolution transforms the initial node features from 3 dimensions to 64 dimensions, and the second layer of graph convolution further transforms the 64-dimensional features into 32-dimensional node embedding vectors. The initial 3-dimensional features include gene expression level, protein expression level, and phosphorylation modification level. The number of nodes in this embodiment is 7 molecular nodes including P38, PRAK, HSP27, VEGFA, GTPBP4, RhoA, and circRNA. A ReLU activation function is used after each layer of graph convolution to introduce nonlinear characteristics. The graph convolution algorithm calculates the feature representation of each node through a neighborhood aggregation function. The aggregation function combines the features of the target node itself and its directly adjacent nodes by weighting. The weights are determined based on the connection strength of the edges. The construction of the adjacency matrix is based on the correlation coefficient matrix. For molecular pairs with an absolute correlation coefficient greater than 0.6 and a P value less than 0.05, the absolute value of the correlation coefficient is filled into the adjacency matrix as the edge weight. Otherwise, the corresponding position is filled with 0 to indicate the absence of direct connection. The algorithm traverses each node in the graph, collects the feature information of its first-order neighbor nodes, processes the neighborhood features through linear transformation and nonlinear activation function, generates node embedding vectors that fuse local network topology information, and the node embedding vectors contain the position information and functional characteristics of the node in the entire pathway network.
[0064] The graph attention mechanism weights the importance weight distribution of the node embedding vector, and the attention mechanism dynamically adjusts the weight distribution of each axis by calculating the correlation strength between different regulatory axes. The attention calculation includes two steps of query vector and key vector generation and attention score calculation. The query vector and the key vector are obtained by multiplying the average node embedding vector of the regulatory axis with the learnable parameter matrix, and the dimension of the parameter matrix is 32x32. The attention algorithm first calculates the dot product of the P38-PRAK-HSP27 signal axis node embedding vector and the query vector as the original attention score. Specifically, the average value of the embedding vectors of the three nodes P38, PRAK and HSP27 is calculated to obtain the representation vector of the axis, and then the dot product of the representation vector and the query vector is calculated. Similarly, the original attention scores of the VEGFA-PRAK activation axis and the PRAK-GTPBP4-RhoA regulatory axis are calculated. For the VEGFA-PRAK activation axis, the average value of the embedding vectors of the two nodes VEGFA and PRAK is extracted, and for the PRAK-GTPBP4-RhoA regulatory axis, the average value of the embedding vectors of the three nodes PRAK, GTPBP4 and RhoA is extracted. Then, the three original scores are converted into attention weights in the form of probability distribution by the softmax normalization function. The softmax function calculates the exponential value of each regulatory axis divided by the sum of the exponential values of all regulatory axes, ensuring that the sum of the three attention weights is equal to one. The attention weight distribution reflects the relative contribution of each regulatory axis to the overall activity of the PRAK pathway under the current biological conditions. The greater the weight value, the more significant the influence of the regulatory axis on the pathway activity under the pathological state of the current patient sample.
[0065] The importance score processing calculates the importance values of the three regulatory axes based on the attention weight distribution. The scoring algorithm multiplies the attention weight of each regulatory axis with the average activity value of the molecular nodes in the axis to obtain the importance score. The importance score of the P38-PRAK-HSP27 signal axis is equal to the attention weight multiplied by the average of the P38, PRAK, and HSP27 node embedding vectors, where the node embedding vector is a 32-dimensional vector of the corresponding node extracted from the second layer graph convolution output, the average function calculates the element-wise average of the three vectors to obtain a 32-dimensional vector, and the sum of all elements of the 32-dimensional vector is obtained to obtain a scalar score. The importance score of the VEGFA-PRAK activation axis is equal to the attention weight multiplied by the average of the VEGFA and PRAK node embedding vectors. The importance score of the PRAK-GTPBP4-RhoA regulatory axis is equal to the attention weight multiplied by the average of the PRAK, GTPBP4, and RhoA node embedding vectors. The sorting algorithm sorts the three regulatory axes in descending order according to the size of the importance score to generate an importance ranking list. The regulatory axis with a higher ranking in the list has a more significant impact on the PRAK pathway activity under the current conditions. The importance ranking list is represented in the form of a three-dimensional vector as input features of the fully connected layer.
[0066] The nonlinear transformation process is to fuse and transform the dimension of the importance ranking list through a fully connected layer, which includes an input layer, a hidden layer and an output layer. The specific structure of the fully connected layer is that the input layer receives a 3-dimensional feature vector, the first hidden layer contains 16 neurons, the second hidden layer contains 8 neurons, and the output layer contains one neuron to output a scalar prediction value. The input layer receives the importance scores of the three regulatory axes in the importance ranking list as the input feature vector. The hidden layer performs matrix multiplication operation with the input feature vector through the weight matrix, then adds the bias vector and performs nonlinear transformation through the ReLU activation function. The weight matrix dimension of the first hidden layer is 16x3, and the weight matrix dimension of the second hidden layer is 8x16. The bias vector dimension of each layer is consistent with the number of neurons in the corresponding layer. The ReLU activation function sets negative values to zero while keeping positive values unchanged, introduces nonlinear characteristics to enable the neural network to learn complex nonlinear mapping relationships, which is crucial for capturing the synergistic and antagonistic effects between the three regulatory axes, because the pathway activity is not simply linearly superimposed but has complex mutual regulation relationships. The output layer maps the output of the hidden layer through another linear transformation to a single numerical value, which represents the overall activity prediction value of the PRAK pathway. The weight matrix dimension of the output layer is 1x8, and the prediction value ranges from negative infinity to positive infinity. To prevent model overfitting, dropout layers are added after the first hidden layer and the second hidden layer, with a dropout rate of 0.3, i.e., 30% of the neuron outputs are randomly discarded during training, and dropout is not used in the testing and inference stages.
[0067] The model training process includes four key steps: training data preparation, loss function definition, optimizer selection, and training iteration. In the training data preparation stage, multi-omics data from 300 triple-negative breast cancer patients were collected, including circRNA expression profiles, miRNA regulation data, and PRAK pathway protein phosphorylation data. The data of each patient was labeled according to the patient's clinical treatment response. Patients with tumor volume reduction greater than or equal to 50% or pathologic complete remission after neoadjuvant chemotherapy were labeled as chemotherapy-sensitive with a label value of 1. Patients with tumor volume reduction less than 30% or disease progression were labeled as chemotherapy-resistant with a label value of 0. The 300 samples were divided into a training set of 240 samples, a validation set of 30 samples, and a test set of 30 samples in a ratio of 8:1:1. Stratified sampling was used to ensure that the proportion of sensitive and resistant samples in each data set was consistent with the overall proportion. The binary cross-entropy loss function was used to measure the difference between the model's predicted values and the true labels. The true label takes a value of 0 or 1, and the predicted probability value is a value between 0 and 1 after being mapped by the Sigmoid function. This loss function penalizes the model for predicting a probability that is too small when the true label is 1, and for predicting a probability that is too large when the true label is 0. The Adam adaptive learning rate optimization algorithm was used to update the model parameters. The learning rate of the Adam optimizer was set to 0.001, the momentum parameter was set to 0.9, and the second-order momentum parameter was set to 0.999. The Adam optimizer can adaptively adjust the learning rate for each parameter to accelerate the convergence process. In the training iteration process, 32 samples were randomly selected from the training set each time to form a batch for forward propagation and backpropagation. A batch size of 32 strikes a balance between computational efficiency and gradient estimation accuracy. The total number of training rounds was set to 100, which means the training set was traversed 100 times. After each round of training, the model performance was evaluated using the validation set. The early stopping strategy was used to terminate training early when the validation set loss value did not decrease for 10 consecutive rounds to prevent overfitting. The convergence criterion for the model was a classification accuracy of more than 85% on the validation set and a validation set loss value of less than 0.3. All learnable parameters of the model, including the weight matrices of the graph convolution layers, the query matrices and key matrices of the attention mechanism, and the weight matrices and bias vectors of the fully connected layers, were updated during the training process through the backpropagation algorithm and the Adam optimizer. The Xavier uniform distribution initialization method was used to initialize the parameters to ensure that the variance of the gradients remained stable between layers at the beginning of training.
[0068] The sigmoid activation function mapping process converts the overall pathway activity prediction value into a standardized activation intensity value. The sigmoid function calculates the quotient of one divided by the negative exponent of the prediction value, mapping any real number to the interval between zero and one. When the prediction value is positive, the sigmoid output is greater than 0.5, indicating pathway activation. When the prediction value is negative, the output is less than 0.5, indicating pathway inhibition. The closer the output value is to one, the higher the PRAK pathway activation intensity. When the prediction value is 0, the sigmoid output is 0.5, indicating that the pathway is in a balanced state. When the prediction value tends to positive infinity, the sigmoid output tends to 1, indicating complete activation. When the prediction value tends to negative infinity, the sigmoid output tends to 0, indicating complete inhibition. Taking adriamycin-resistant triple-negative breast cancer cells as an example, the graph convolution layer extracts node features of specific circRNA down-regulation leading to related miRNA up-regulation. Specifically, the first layer of graph convolution aggregates the initial expression level features of circRNA nodes with the features of adjacent PRAK nodes to obtain intermediate representations. The second layer of graph convolution further aggregates the feature relationships between PRAK nodes and miRNA nodes. The attention mechanism finds that the PRAK-GTPBP4-RhoA regulatory axis has the highest weight because the GGNBP2 protein expression is significantly down-regulated. The attention weight calculation result shows that the weight of the PRAK-GTPBP4-RhoA regulatory axis is 0.65, which is much higher than that of the other two axes, indicating that this regulatory axis is the most important in the drug-resistant cell sample. The importance score shows that this regulatory axis ranks first. After the importance score vector is arranged in descending order, an importance ranking list is obtained. The fully connected layer fuses the importance scores of the three axes and outputs a negative prediction. After the 16-dimensional vector output by the first hidden layer is activated by ReLU, most of the elements are suppressed and close to 0. The second hidden layer further compresses the features, and the output layer calculates the negative prediction result. The sigmoid function maps the negative value to an activation intensity value less than 0.5, about 0.06, indicating that the PRAK pathway is in an inhibited state in the drug-resistant cells. The activation intensity value of 0.06 represents a pathway activation probability of only 6%, which is much lower than the threshold value of 0.5. Clinically, it is explained that the mechanism of drug resistance to chemotherapy drugs in this patient is related to the significant inhibition of the PRAK signaling pathway.
[0069] In a specific embodiment, the process of performing step S104 can specifically include the following steps:
[0070] Performing time series trend analysis on the PRAK pathway activation intensity value to obtain a pathway dynamic response curve;
[0071] Extracting response amplitude, response time, and recovery time parameters based on the pathway dynamic response curve to obtain a perturbation response feature vector;
[0072] The disturbance response feature vector is subjected to K-means clustering analysis processing to obtain a chemotherapy resistance type cluster center, an apoptosis sensitivity type cluster center and a proliferation regulation type cluster center;
[0073] According to the expression changes of Bax and Bcl-2 apoptosis markers, the chemotherapy resistance type cluster center, the apoptosis sensitivity type cluster center and the proliferation regulation type cluster center are subjected to biological annotation processing to obtain a chemotherapy resistance type classification standard, an apoptosis sensitivity type classification standard and a proliferation regulation type classification standard;
[0074] Based on the chemotherapy resistance type classification standard, the apoptosis sensitivity type classification standard and the proliferation regulation type classification standard, the input sample is subjected to pattern matching processing to obtain a pathway response feature identifier.
[0075] Specifically, the time series change trend analysis processing dynamically tracks the changes of the PRAK pathway activation intensity value at different time points. The algorithm arranges the activation intensity values in time sequence to form time series data, and calculates the change rate between adjacent time points through sliding window technology. The change rate is equal to the activation intensity of the next time point minus the activation intensity of the previous time point divided by the time interval. The trend analysis algorithm identifies the rising trend, the falling trend and the stable trend of the activation intensity value. The rising trend corresponds to consecutive multiple positive change rates, the falling trend corresponds to consecutive multiple negative change rates, and the stable trend corresponds to the interval with change rate close to zero. The algorithm draws a pathway dynamic response curve with time as the horizontal axis and PRAK pathway activation intensity as the vertical axis. The curve shape reflects the activation mode and time evolution characteristics of the pathway under specific stimulation conditions.
[0076] The parameter extraction processing extracts three key kinetic parameters from the pathway dynamic response curve to quantify the disturbance response features. The response amplitude calculation algorithm identifies the difference between the maximum value and the baseline value of the activation intensity in the curve. The baseline value is determined by calculating the average value of the activation intensity within a certain time window before stimulation. The maximum value is obtained by searching for the peak value of the activation intensity throughout the response process. The response time calculation algorithm measures the time length required from the start of stimulation to the peak value of the activation intensity. The algorithm identifies the starting time point of the activation intensity rising by detecting the change of the curve slope, and then calculates the time difference between the starting time point and the peak time point. The recovery time calculation algorithm measures the time required for the activation intensity to drop from the peak value to the baseline level. The algorithm sets the recovery threshold to one tenth of the difference between the peak value and the baseline value, and calculates the time length for the activation intensity to drop from the peak value to the recovery threshold. The combination of the three parameters forms a disturbance response feature vector that describes the disturbance response kinetic characteristics.
[0077] K-means clustering analysis process groups perturbation response feature vectors according to similarity, and the clustering algorithm sets the number of clusters to three, corresponding to different biological response modes. The algorithm randomly initializes three cluster center coordinates, each of which represents a typical response mode in a three-dimensional feature space. Then, the Euclidean distance between each perturbation response feature vector and the three cluster centers is calculated. The Euclidean distance is equal to the square root of the sum of the squares of the differences between the feature vector and the cluster center coordinates. The distance calculation algorithm assigns each feature vector to the nearest cluster center, forming three initial cluster clusters. Then, the algorithm updates and recalculates the center of gravity of all feature vectors in each cluster as the new cluster center coordinates. The iterative algorithm repeats the assignment and update process until the cluster center coordinates converge and no longer change significantly. Finally, the stable chemotherapy-resistant cluster center, apoptosis-sensitive cluster center, and proliferation-regulated cluster center are obtained.
[0078] Biological annotation processing gives biological significance to the three cluster centers according to the expression changes of Bax and Bcl-2 apoptosis markers. The annotation algorithm analyzes the Bax protein expression level and Bcl-2 protein expression level of the samples corresponding to each cluster center. High expression of Bax, as a pro-apoptotic protein, indicates that the cell tends to apoptosis, and high expression of Bcl-2, as an anti-apoptotic protein, indicates that the cell has strong anti-apoptotic ability. The algorithm calculates the Bax / Bcl-2 expression ratio of samples in each cluster center. High ratio indicates pro-apoptotic tendency, and low ratio indicates anti-apoptotic tendency. Combined with the finding in related research that overexpression of specific circRNA leads to upregulation of Bax and downregulation of Bcl-2 to promote apoptosis, the annotation algorithm annotates the cluster center with high Bax / Bcl-2 ratio and large response amplitude as the apoptosis-sensitive classification standard, the cluster center with low Bax / Bcl-2 ratio and long recovery time as the chemotherapy-resistant classification standard, and the cluster center with medium ratio and short response time as the proliferation-regulated classification standard.
[0079] The mode matching processing is based on the established classification standard to identify the response mode of the new input sample. The matching algorithm calculates the similarity of the perturbation response feature vector of the input sample and the three classification standards. The similarity calculation uses the cosine similarity algorithm, which calculates the dot product of the feature vector and the classification standard vector divided by the product of the two vector lengths. The cosine similarity value ranges from -1 to 1, and the closer the value is to 1, the higher the similarity. The algorithm compares the similarity values of the input sample and the three classification standards, and selects the classification standard with the highest similarity as the pathway response feature identifier of the sample. Taking adriamycin-treated triple-negative breast cancer cells as an example, the PRAK pathway activation intensity of the cells after drug treatment first rapidly decreased and then slowly recovered, forming a dynamic response curve with a large response amplitude and a long recovery time. At the same time, it was detected that the down-regulation of Bax protein expression and the up-regulation of Bcl-2 protein expression led to a low Bax / Bcl-2 ratio. K-means clustering classified this response mode into the chemotherapy-resistant cluster center, and biological annotation confirmed that it met the characteristics of the chemotherapy-resistant classification standard. The mode matching algorithm finally assigned the chemotherapy-resistant pathway response feature identifier to the sample.
[0080] In a specific embodiment, the process of performing step S105 can specifically include the following steps:
[0081] Performing disease phenotype correlation analysis processing on the chemotherapy-resistant, apoptosis-sensitive, and proliferation-regulated types in the pathway response feature identifier to obtain chemotherapy-resistant-disease mapping relationship, apoptosis-sensitive-disease mapping relationship, and proliferation-regulated-disease mapping relationship;
[0082] Based on the chemotherapy-resistant-disease mapping relationship, apoptosis-sensitive-disease mapping relationship, and proliferation-regulated-disease mapping relationship, performing intervention assess processing on the HSP27 and GTPBP4 key target points to obtain the HSP27 target point intervention priority and the GTPBP4 target point intervention priority;
[0083] According to the HSP27 target point intervention priority and the GTPBP4 target point intervention priority, performing drug-target affinity calculation processing to obtain the binding strength score of the candidate drug molecule and HSP27 and the binding strength score of the candidate drug molecule and GTPBP4;
[0084] Inputting the binding strength score of the candidate drug molecule and HSP27 and the binding strength score of the candidate drug molecule and GTPBP4 into the synergistic effect prediction algorithm to perform multi-target joint intervention evaluation processing to obtain the HSP27-GTPBP4 target point combination synergistic index;
[0085] Based on the HSP27-GTPBP4 target point combination synergistic index, performing optimization selection processing to obtain the PRAK pathway intervention target point combination.
[0086] Specifically, the mapping relationship between three response types and disease states is established based on clinical data of triple-negative breast cancer patients. The algorithm first collects clinical indicators of patients with chemotherapy resistance, including tumor volume change rate after neoadjuvant chemotherapy, pathological complete remission rate, progression-free survival, etc. The mean and variance of the clinical phenotype characteristics of these patient groups are calculated through statistical analysis. When the tumor volume reduction rate is lower than the average minus one standard deviation and the pathological complete remission rate is lower than 20%, the mapping relationship between chemotherapy resistance and treatment-resistant disease phenotype is established. This mapping relationship indicates that patients with chemotherapy resistance pathway response characteristics are clinically insensitive to standard chemotherapy regimens and have poor prognosis. In the data processing process of patients with apoptosis sensitivity, the algorithm analyzes the expression ratio of Bax pro-apoptotic protein and Bcl-2 anti-apoptotic protein in tumor tissue of patients. The protein expression level is obtained through quantitative analysis of immunohistochemical staining intensity. The Bax protein expression level is quantified by the optical density value of the brown staining area, and the Bcl-2 protein expression level is quantified by the same method. When the Bax and Bcl-2 expression ratio is greater than the median plus one quartile range of the sample, the mapping relationship between apoptosis sensitivity and treatment-sensitive disease phenotype is established. At the same time, the accuracy of the mapping relationship is verified by combining the clinical data such as significant tumor volume reduction and prolonged survival of patients. This mapping relationship indicates that tumor cells of patients with apoptosis sensitivity are more likely to be induced to apoptosis by chemotherapy drugs, thereby achieving better treatment effect. In the data processing of patients with proliferation regulation, the algorithm calculates the product of the percentage of Ki-67 positive cells and the expression intensity of cell cycle proteins as the proliferation activity score by analyzing the expression level of Ki-67 proliferation marker and the expression pattern of cell cycle proteins. The percentage of Ki-67 positive cells is obtained by counting the number of stained positive cells under a microscope and dividing by the total number of cells. The expression intensity of cell cycle proteins is quantified by the gray value of the protein immunoblotting band. When the score is at the median level of the sample distribution, the mapping relationship between proliferation regulation and abnormal tumor proliferation disease phenotype is established. This mapping relationship indicates that the proliferation rate of tumor cells of patients with proliferation regulation is at a medium level, neither highly proliferative nor in a static state.
[0087] Based on the established three disease mapping relationships, the two key target points of HSP27 and GTPBP4 are evaluated for intervention. The HSP27 target point evaluation algorithm analyzes the phosphorylation regulation in the P38-PRAK-HSP27 signal axis. The phosphorylation state of the HSP27 protein Ser78 and Ser82 sites is detected, and the correlation strength of the apoptosis-sensitive disease mapping relationship is analyzed. The phosphorylation of the Ser78 site and the Ser82 site of the HSP27 protein is detected by immunoblotting experiment with specific phospho-antibody. The algorithm calculates the Pearson correlation coefficient between the expression level of phosphorylated HSP27 and the Bax / Bcl-2 ratio. The calculation method of the Pearson correlation coefficient is the covariance of two variables divided by the product of the standard deviations of the two variables. When the correlation coefficient is greater than zero point five and the P value is less than zero point zero five, it is determined that the HSP27 target point has a high intervention priority. Because the phosphorylation activation of HSP27 is positively correlated with the sensitivity of cell apoptosis, regulating this target point can enhance the sensitivity of tumor cells to chemotherapeutic drugs. High intervention priority means that drug development targeting this target point has high clinical transformation value. The GTPBP4 target point evaluation is based on the functional role in the PRAK-GTPBP4-RhoA regulatory axis. The algorithm analyzes the correlation between the regulatory effect of GTPBP4 phosphorylation modification on RhoA-GTP hydrolysis activity and the proliferation-regulated disease phenotype. The intervention potential is evaluated by detecting the negative correlation between the GTPBP4 phosphorylation level and the expression of the cell proliferation marker Ki-67. The GTPBP4 phosphorylation level is identified by mass spectrometry analysis of the phosphorylation site and quantification of the modification degree. When the GTPBP4 phosphorylation level is negatively correlated with the expression of Ki-67 and the absolute value of the correlation coefficient is greater than zero point four, it is determined that the GTPBP4 target point has a medium intervention priority. Because enhancing GTPBP4 phosphorylation can inhibit RhoA activity and thus regulate cell proliferation, the medium intervention priority indicates that the intervention effect of this target point is relatively mild but still has clinical application potential.
[0088] The drug-target affinity calculation process uses a molecular docking simulation algorithm to predict the binding strength of candidate drug molecules to HSP27 and GTPBP4 proteins. The algorithm inputs the three-dimensional structure coordinate file of the candidate drug molecule and the crystal structure data of the target protein. The three-dimensional structure of the candidate drug molecule is constructed by chemical drawing software and optimized by energy minimization. The crystal structure data of the target protein is obtained from the protein database. First, the protein structure is preprocessed, including removing water molecules, adding hydrogen atoms, optimizing side chain conformation, etc. Water molecules are removed because most of them do not participate in drug binding. Hydrogen atoms are added for accurate calculation of hydrogen bond interactions. Optimizing the side chain conformation ensures that the protein structure is in a reasonable energy state. Then, potential drug binding pockets are identified on the protein surface. The binding pocket recognition algorithm is based on the concave region and hydrophobicity distribution on the protein surface. For HSP27 protein, the algorithm identifies the hydrophobic binding cavity around the phosphorylation site as the drug binding site. The identification of this binding cavity is based on the hydrophobicity score of the amino acids around the Ser78 and Ser82 sites. For GTPBP4 protein, the algorithm identifies the allosteric regulation site near the GTP binding domain. The identification of the allosteric regulation site is achieved by analyzing the conformational change region of the GTP binding domain. The search algorithm systematically tries different orientations and conformations of the drug molecule within the binding pocket. A large number of candidate binding poses are generated through rotation and translation transformations. The rotation transformation traverses the 360-degree space with a step size of 10 degrees. The translation transformation covers the three-dimensional space of the binding pocket with a step size of 1 angstrom. The binding strength scoring function calculates the interaction energy between the drug molecule and the protein in each binding pose, including van der Waals force, electrostatic interaction, hydrogen bond contribution, and desolvation effect. Van der Waals force is calculated using the Lennard-Jones potential energy function. Electrostatic interaction is calculated using Coulomb's law. Hydrogen bond contribution is scored based on the distance and angle between hydrogen bond donors and acceptors. Desolvation effect considers the contact area between the drug molecule and the hydrophobic region of the protein surface. The algorithm selects the lowest energy binding pose as the optimal binding mode and outputs the corresponding binding strength score. The more negative the score, the stronger the binding affinity. The unit of the binding strength score is kilocalorie per mole. A typical strong binding score is negative 8 to negative 12 kilocalories per mole.
[0089] The Bliss independence model was used to evaluate the therapeutic effect of the HSP27-GTPBP4 dual-target combination intervention. The Bliss independence model is based on the assumption of independent action, i.e., the combined effect of two targets is equal to the product of the probability of each individual effect. The theoretical basis of the Bliss independence model is that if two drugs act independently, the combined effect can be predicted by probability theory. The algorithm converts the binding strength score of the candidate drug molecule to HSP27 into the inhibition efficiency of the HSP27 target. The conversion method is to map the binding strength score through the exponential function to the inhibition efficiency value between 0 and 1. The specific formula is that the inhibition efficiency is equal to 1 minus the exponential of the ratio of the binding strength score of the natural exponential and the temperature constant. The binding strength score of the candidate drug molecule to GTPBP4 is converted into the activation efficiency of the GTPBP4 target. The conversion method is similar to that of HSP27 but in the opposite direction because GTPBP4 needs to be activated rather than inhibited. The change in cell survival rate when acting on the HSP27 target alone and the GTPBP4 target alone was determined by cell viability detection experiments. Cell viability detection was performed using the MTT colorimetric method or the CCK-8 kit. In the experiment, the tumor cells were divided into a control group, an HSP27-only treatment group, a GTPBP4-only treatment group, and a combined treatment group. Each group had 6 replicate wells. After 48 hours of culture, the absorbance value was measured. The cell survival rate was equal to the absorbance of the treatment group divided by the absorbance of the control group multiplied by 100%. Then the actual cell survival rate when acting on both targets simultaneously was determined. The Bliss model calculates the expected combined effect as the product of the probabilities of the two individual effects. The specific formula is that the cell survival rate under the expected combined effect is equal to the survival rate after HSP27-only treatment multiplied by the survival rate after GTPBP4-only treatment. The HSP27-GTPBP4 target combination synergy index is equal to the actual combined effect divided by the expected combined effect. The specific calculation is the cell survival rate after actual combined treatment divided by the survival rate under the expected combined effect. When the synergy index is greater than one, it indicates that there is synergy between the two targets, i.e., the combined intervention effect is better than the sum of the individual intervention effects. Synergy means that combined drug therapy can produce an effect of 1 plus 1 greater than 2. When the synergy index is less than one, it indicates that there is antagonism. Antagonism means that the combined intervention of the two targets actually weakens the therapeutic effect.
[0090] The optimal selection process screens the optimal HSP27-GTPBP4 target combination intervention scheme through a single-objective optimization algorithm. Since only one target combination synergy index needs to be optimized, an optimization function is constructed to maximize the synergy index. The optimization function is in the form of a target function equal to the synergy index minus a toxicity penalty term. Meanwhile, constraint conditions are set to ensure that the drug molecule's toxicity indicators are below the safety threshold. The toxicity indicators include cytotoxicity, hepatotoxicity, and nephrotoxicity. The safety threshold is set according to the preclinical safety evaluation standard. The algorithm first normalizes the HSP27-GTPBP4 target combination synergy index to map the numerical value to the interval of zero to one to eliminate the dimension effect. The normalization process uses the min-max normalization method, with the formula being normalized value equal to original value minus minimum value divided by maximum value minus minimum value. Then, based on the important role of HSP27 in apoptosis regulation and the key function of GTPBP4 in proliferation control, a weight coefficient is assigned to the target combination. The weight coefficient is determined based on the literature-reported target importance score and preclinical research data. The weight coefficient of HSP27 is set to 0.6, and the weight coefficient of GTPBP4 is set to 0.4. The weighted synergy index is equal to the normalized synergy index multiplied by 0.6 plus the GTPBP4-related effect multiplied by 0.4. The gradient ascent algorithm determines the search direction by calculating the partial derivative of the target function with respect to each variable. The partial derivative is calculated using the numerical differentiation method, with a step size of 0.01. The gradient ascent iteration formula is new parameter equal to old parameter plus learning rate multiplied by gradient, with a learning rate of 0.1. The candidate solution is updated iteratively until the drug dosage ratio and dosing schedule that maximize the synergy index are found. The convergence condition is that the target function changes by less than 0.001 for 10 consecutive iterations or reaches the maximum iteration number of 100. The optimal scheme includes specific drug molecule selection for HSP27 and GTPBP4 targets, dosage ratio, and joint dosing strategy. The specific drug molecule selection screens the molecule with the optimal binding strength score from the candidate drug library. The dosage ratio determines the dosage proportion of HSP27-targeted drugs and GTPBP4-targeted drugs. The joint dosing strategy includes simultaneous administration or sequential administration and dosing interval. The output format of the optimal scheme is HSP27-targeted drug dosage X mg / kg, GTPBP4-targeted drug dosage Y mg / kg, and dosing method simultaneous administration or HSP27 drug administration first with a 24-hour interval followed by GTPBP4 drug administration.
[0091] In a specific embodiment, the process of performing the optimal selection process based on the HSP27-GTPBP4 target combination synergy index can specifically include the following steps:
[0092] Numerical standardization is performed on the HSP27-GTPBP4 target combination synergy index to obtain a normalized synergy index matrix.
[0093] According to the normalized synergy index matrix, the weight of the target combination effect is calculated and processed to obtain the HSP27-GTPBP4 combination weight coefficient.
[0094] Based on the HSP27-GTPBP4 combination weight coefficient, a multi-objective optimization function is constructed and processed to obtain the target combination optimization objective function.
[0095] The target combination optimization objective function is input into the genetic algorithm for global optimization processing to obtain the optimal target combination solution set.
[0096] The optimal target combination solution set is screened by the Pareto front to obtain the PRAK pathway intervention target combination.
[0097] Specifically, the numerical standardization processing is for the single HSP27-GTPBP4 target combination synergy index normalization transformation, although there is only one synergy index value, but still needs to be standardized for comparison and subsequent calculation with other numerical values, the algorithm uses Z-score standardization method to convert the synergy index into standard normal distribution form, in the specific calculation process, the algorithm first collects the synergy index data of the same type of target combination in history as the reference sample, the reference sample includes 50 groups of reported drug target combination synergy index data, the mean and standard deviation of the reference sample are calculated, the mean calculation method is the sum of all sample values divided by the sample number, the standard deviation calculation method is the square sum of the difference between each sample value and the mean divided by the sample number and then taking the square root, then the current HSP27-GTPBP4 synergy index is subtracted from the sample mean and divided by the sample standard deviation to obtain the standardized value. The normalized synergy index matrix is simplified to a single element matrix in this case, which contains the standardized HSP27-GTPBP4 synergy index value, the standardization processing eliminates the dimension influence of the original value, making the subsequent weight calculation and optimization processing more accurate, at the same time, the standardized value has clear statistical significance, positive value indicates that the synergy index is higher than the historical average level, negative value indicates that it is lower than the average level, the absolute value size reflects the deviation degree, for example, the standardized value is 1.5, which means that the synergy index is 1.5 standard deviations higher than the historical average level.
[0098] The weight calculation process evaluates the importance of the HSP27-GTPBP4 target combination effect based on the normalized synergy index matrix. The algorithm first analyzes the functional role of HSP27 heat shock protein in the PRAK signaling pathway. HSP27, as a chaperone protein, plays a key role in cellular stress response and apoptosis regulation. Its phosphorylation status directly affects the sensitivity of cells to chemotherapy drugs. The algorithm assigns a functional importance score to HSP27 based on its central position in the P38-PRAK-HSP27 signaling axis. The determination of the functional importance score is based on literature analysis and statistical analysis of the frequency of citation and experimental verification of HSP27 in apoptosis regulation-related research. The score ranges from 0 to 10, and the functional importance score of HSP27 is 8.5, indicating its high importance in the pathway. GTPBP4, as a GTP-binding protein, controls the GTP hydrolysis activity of RhoA in the PRAK-GTPBP4-RhoA regulatory axis, thereby regulating cell proliferation and migration. The algorithm analyzes the regulatory effect of GTPBP4 phosphorylation modification on downstream RhoA signaling and assigns a functional score to GTPBP4. The functional importance score of GTPBP4 is 6.2, indicating that its importance is relatively lower than that of HSP27 but still has a significant role. The HSP27-GTPBP4 combination weight coefficient calculation uses the weighted geometric mean method, which multiplies the functional importance scores of the two targets and takes the square root. The specific calculation is that the combination weight is equal to the square root of the product of 8.5 and 6.2, which is 7.26. Then multiply by the normalized synergy index to get the final weight coefficient. For example, when the normalized synergy index is 1.5, the final weight coefficient is 7.26 multiplied by 1.5, which is equal to 10.89. This weight coefficient reflects the relative importance of the HSP27-GTPBP4 combination in the entire PRAK pathway intervention scheme. The larger the weight value, the more significant the contribution of the combination to the treatment effect.
[0099] The multi-objective optimization function construction process is based on the HSP27-GTPBP4 combination weight coefficient to establish an objective function that considers both the maximization of therapeutic effect and the minimization of toxicity and side effects. Since there is only one target point combination, the structure of the objective function is relatively simple, but multiple conflicting objectives still need to be balanced. The mathematical expression of the multi-objective optimization function is F = w × (E - T), where F represents the objective function value, w represents the safety weight factor, E represents the therapeutic effect term, and T represents the toxicity and side effect term. The optimization goal is to maximize the value of F. The calculation formula of the therapeutic effect term E is E = C × S × A, where C represents the HSP27-GTPBP4 combination weight coefficient, S represents the synergy index, and A represents the therapeutic effect amplification factor. The combination weight coefficient C reflects the importance of the target point combination in the pathway, and the value calculated earlier is 10.89. The synergy index S reflects the degree of synergy of the combined action of the two targets, and the value range is usually 0.8 to 2.0, where a value greater than 1 indicates the presence of synergy. The therapeutic effect amplification factor A is determined based on clinical trial data and reflects the actual therapeutic potential of the target point combination, with a value range of 1 to 5. The determination of the therapeutic effect amplification factor is based on tumor volume reduction and survival data in animal experiments. For example, when the HSP27-GTPBP4 combination reduces tumor volume by 80% and prolongs survival by 2 times in animal models, the therapeutic effect amplification factor is set to 4.5. The larger the value of the therapeutic effect term E, the better the expected therapeutic effect. The calculation formula of the toxicity and side effect term T is T = T1 + T2 + 0.5 × T1 × T2, where T1 represents the toxicity score of the HSP27 target drug, T2 represents the toxicity score of the GTPBP4 target drug, and 0.5 × T1 × T2 represents the additional interactive toxicity when the two targets are used together. The toxicity T1 of the HSP27 modulator is determined by normal cell viability detection experiments. When the survival rate of normal cells is 85% at the therapeutic concentration of the drug, the toxicity score T1 is equal to 1 minus 0.85, which is 0.15. The toxicity score T2 of the GTPBP4 inhibitor is determined by the same method and is 0.22. Substituting the formula, the combination toxicity T is equal to 0.15 plus 0.22 plus 0.5 times 0.15 times 0.22, which is equal to 0.3865. The smaller the value of the toxicity and side effect term T, the lower the toxicity of the drug combination. The safety weight factor w is dynamically adjusted according to patient tolerance and clinical safety requirements. For young patients with good tolerance, the safety weight factor w is set to 0.8, emphasizing the therapeutic effect. For elderly patients with poor tolerance, the safety weight factor w is set to 1.2, emphasizing safety. The weight factor w adjusts the relative importance of therapeutic effect and toxicity in the objective function.The parameters are substituted into the objective function F = w x (E - T) for calculation, for example, when C = 10.89, S = 1.5, A = 4.5, T1 = 0.15, T2 = 0.22, w = 0.8, first, the treatment effect term E = 10.89 x 1.5 x 4.5 = 73.5 is calculated, then the toxicity term T = 0.15 + 0.22 + 0.5 x 0.15 x 0.22 = 0.3865 is calculated, and finally the objective function value F = 0.8 x (73.5 - 0.3865) = 58.49 is calculated, the greater the objective function value indicates the higher the therapeutic benefit of the target point combination scheme under the premise of ensuring safety, and the core of multi-objective optimization is to find the best scheme that maximizes the objective function F by adjusting the drug dosage ratio and the drug administration time sequence and other parameters.
[0100] The genetic algorithm global optimization process takes the target point combination optimization objective function as the fitness function to search for the optimal solution. Since only the HSP27-GTPBP4 single target point combination is involved, the algorithm mainly optimizes the drug dosage ratio and administration timing parameters of this combination. The algorithm initializes a population containing multiple candidate solutions, with a population size of 50 individuals. Each individual represents a specific implementation of the HSP27-GTPBP4 target point combination. The individual encoding uses a real number vector form, with the first dimension representing the drug dosage for the HSP27 target point, ranging from 0 to 100 mg / kg body weight, the second dimension representing the drug dosage for the GTPBP4 target point, ranging from 0 to 50 mg / kg body weight, and the third dimension representing the administration time interval of the two drugs, ranging from 0 to 24 hours. The individual parameters of the initial population are generated by a random number generator uniformly sampling within the value range of each dimension. In the fitness evaluation stage, the algorithm substitutes the individual encoded parameters into the objective function F = w x (E-T) to calculate the fitness score. A higher fitness indicates better overall treatment effectiveness of the parameter combination. The selection operation uses the tournament selection method, where the algorithm randomly selects 5 individuals for comparison, and the individual with the highest fitness is selected as the parent. This selection method maintains population diversity while converging to high fitness individuals. The tournament size is set to 5 to balance selection pressure and diversity. The crossover operation exchanges parameters between selected parent individuals. The algorithm uses the arithmetic crossover method to linearly combine the parameter vectors of two parent individuals to generate new offspring individuals. The specific calculation is offspring parameters equal to parent 1 parameters multiplied by a crossover coefficient plus parent 2 parameters multiplied by 1 minus the crossover coefficient. The crossover coefficient is randomly taken from the range of 0.3 to 0.7, and the crossover probability is set to 0.8, indicating that 80% of individuals will undergo crossover operations. The mutation operation introduces new genetic variations by adding Gaussian noise to individual parameters. The standard deviation of Gaussian noise is initially set to 10% of the parameter value range, and the mutation strength decreases with the number of evolution generations to balance exploration and exploitation capabilities. The specific decay formula is the mutation standard deviation of the current generation equal to the initial standard deviation multiplied by 0.95 raised to the power of the current generation number. The mutation probability is set to 0.1, indicating that 10% of individuals will undergo mutation. The iterative evolution process continues until the fitness converges or the preset number of generations is reached. The convergence criterion is that the optimal fitness of the population changes by less than 0.01 for 20 consecutive generations, and the maximum evolution generation is set to 200. The final optimal target point combination solution set containing multiple high fitness individuals is obtained, with the top 10 individuals in terms of fitness as candidate optimal schemes.
[0101] The Pareto front screening process extracts non-dominated solutions from the optimal target point combination solution set to form a Pareto optimal front. The algorithm determines the multi-objective dominance relationship of each solution in the solution set. Solution A dominates solution B if and only if solution A is not worse than solution B in both therapeutic effect and safety, and is strictly better than solution B in at least one objective. The specific judgment criteria are that the therapeutic effect item E of solution A is greater than or equal to the therapeutic effect item of solution B, and the side effect item T of solution A is less than or equal to the side effect item of solution B, and at least one of the two inequalities is a strict inequality. The screening algorithm calculates the dominance level of each solution, which is determined by counting how many solutions dominate the current solution. The solution with a dominance level of zero is a non-dominated solution that forms the first layer of the Pareto front. The solution with a dominance level of one forms the second layer of the front, and so on. In this embodiment, since the solution set is relatively small, it usually only forms 1 to 2 layers of the Pareto front. Since the two objectives of maximizing therapeutic effect and minimizing side effects conflict with each other, the solutions on the Pareto front represent different effect and safety trade-off schemes. The solution with a larger therapeutic effect item E on the front has a higher therapeutic effect but a relatively larger side effect item T. The solution with a smaller side effect item T on the front has a smaller side effect but a corresponding decrease in the therapeutic effect item E. The algorithm selects the most suitable solution from the Pareto front as the final PRAK pathway intervention target point combination according to the patient's specific condition and tolerance. The selection criteria prioritize the combination scheme with the largest therapeutic effect within the acceptable safety range. The specific selection method is to first screen solutions with a side effect item T less than 0.4 as acceptable safety range, and then select the solution with the largest therapeutic effect item E from these solutions as the final scheme. The final output of the PRAK pathway intervention target point combination includes the HSP27 target drug dose, the GTPBP4 target drug dose, and the drug administration time interval.
[0102] The PRAK signal pathway analysis method in the embodiments of the present application is described above, and the PRAK signal pathway analysis system in the embodiments of the present application is described below. Please refer to Figure 2 The PRAK signal pathway analysis system in the embodiments of the present application includes one embodiment:
[0103] The acquisition module is configured to perform multi-omics data acquisition and processing on the biological sample by high-throughput sequencing technology to obtain a comprehensive data set containing circRNA expression profile, miRNA regulation data, and PRAK pathway protein phosphorylation data.
[0104] The mining module is configured to perform molecular interaction relationship mining processing according to the comprehensive data set to obtain a pathway connection matrix containing a P38-PRAK-HSP27 signal axis, a VEGFA-PRAK activation axis, and a PRAK-GTPBP4-RhoA regulation axis.
[0105] A prediction module is configured to perform path activity state prediction processing on the path connection matrix through a graph neural network algorithm to obtain PRAK path activation intensity values under different conditions.
[0106] An extraction module is configured to perform perturbation response feature extraction processing according to the PRAK path activation intensity values to obtain path response feature identifiers including chemoresistance, apoptosis sensitivity and proliferation regulation.
[0107] A matching module is configured to perform treatment target matching processing on the path response feature identifiers to obtain a PRAK path intervention target combination for a specific disease phenotype.
[0108] The above Figure 2 The PRAK signal path analysis system in the embodiment of the present application is described in detail from the perspective of a modular functional entity, and the PRAK signal path analysis device in the embodiment of the present application is described in detail from the perspective of hardware processing.
[0109] Referring to Figure 3 , the embodiment of the present application also provides a PRAK signal path analysis device, which can be a server, and the internal structure thereof can be as shown in Figure 3 The PRAK signal path analysis device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer is configured to provide computing and control capabilities. The memory of the PRAK signal path analysis device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the PRAK signal path analysis device is configured to store the corresponding data in the embodiment. The network interface of the PRAK signal path analysis device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0110] Those skilled in the art can understand Figure 3 the structure shown in the embodiment of the present application, which is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the PRAK signal path analysis device to which the scheme of the present application is applied.
[0111] The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium. The computer readable storage medium stores instructions, and when the instructions are run on a computer, the computer performs the steps of the PRAK signal path analysis method.
[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, system and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.
[0113] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a PRAK signal pathway analysis device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0114] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A PRAK signal path analysis method, characterized in that, The method comprises: The biological sample is processed by high-throughput sequencing technology to obtain a comprehensive data set comprising circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data; According to the comprehensive data set, molecular interaction relationship mining processing is performed to obtain a pathway connection matrix comprising P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulation axis; The pathway connection matrix is processed by a graph neural network algorithm to predict the activity state of the pathway to obtain the PRAK pathway activation intensity value under different conditions; According to the PRAK pathway activation intensity value, the perturbation response feature is extracted to obtain the pathway response feature identifier comprising the chemotherapy resistance type, the apoptosis sensitivity type and the proliferation regulation type; The pathway response feature identifier is matched with the treatment target to obtain the PRAK pathway intervention target combination for a specific disease phenotype.
2. The PRAK signaling pathway analysis method of claim 1, wherein, The biological sample is processed by high-throughput sequencing technology to obtain a comprehensive data set comprising circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data, comprising: The biological sample is processed by RNA-seq sequencing to obtain raw expression data of circular RNA; The raw expression data of circular RNA is input into a quality control algorithm for standardization processing to obtain the circRNA expression profile; The biological sample is processed by small RNA sequencing to obtain raw expression data of small RNA; The small RNA raw expression data is processed by target gene prediction analysis to obtain the miRNA regulation data; The biological sample is detected by proteomics technology to obtain the PRAK pathway protein phosphorylation data; The circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data are fused to obtain the comprehensive data set.
3. The PRAK signaling pathway analysis method of claim 1, wherein, The biological sample is processed by high-throughput sequencing technology to obtain a comprehensive data set comprising circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data, comprising: The circRNA expression profile, miRNA regulation data and PRAK pathway protein phosphorylation data in the comprehensive data set are processed by Pearson correlation calculation to obtain a molecular correlation coefficient matrix; According to the correlation coefficient matrix, the threshold value of the molecular correlation strength is screened to obtain a significant correlation molecule pair with an absolute value of the correlation coefficient greater than 0.6 and a P value less than 0.05; Based on the significant correlation molecule pair, the P38 and PRAK, PRAK and HSP27 phosphorylation cascade relationship construction processing is performed to obtain the P38-PRAK-HSP27 signal axis; The significant correlation molecule pair is analyzed by VEGFA and PRAK activation relationship verification to obtain the VEGFA-PRAK activation axis; The significant correlation molecule pair is subjected to PRAK and GTPBP4, GTPBP4 and RhoA regulatory relationship construction processing, and the PRAK-GTPBP4-RhoA regulatory axis is obtained; The P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulatory axis in the pathway connection matrix are input into the graph convolution layer for node feature aggregation processing, and the node embedding vector containing neighborhood information is obtained; 4. The PRAK signaling pathway analysis method of claim 1, wherein, The node embedding vector is subjected to graph attention mechanism weighting processing, and the attention weight distribution of the P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulatory axis is obtained; Based on the attention weight distribution, the importance score processing is performed on the P38-PRAK-HSP27 signal axis, VEGFA-PRAK activation axis and PRAK-GTPBP4-RhoA regulatory axis, and the importance ranking list of the three regulatory axes is obtained; The importance ranking list is input into the full connection layer for nonlinear transformation processing, and the overall activity prediction value of the pathway is obtained; The overall activity prediction value of the pathway is subjected to Sigmoid activation function mapping processing, and the PRAK pathway activation intensity value is obtained. The PRAK pathway activation intensity value is subjected to time series change trend analysis processing, and the pathway dynamic response curve is obtained; Based on the pathway dynamic response curve, the response amplitude, response time and recovery time parameters are extracted, and the disturbance response feature vector is obtained; 5. The PRAK signaling pathway analysis method of claim 1, wherein, The disturbance response feature vector is subjected to K-means clustering analysis processing, and the chemotherapy resistance type cluster center, apoptosis sensitivity type cluster center and proliferation regulation type cluster center are obtained; According to the expression changes of Bax and Bcl-2 apoptosis markers, the chemotherapy resistance type cluster center, apoptosis sensitivity type cluster center and proliferation regulation type cluster center are subjected to biological annotation processing, and the chemotherapy resistance type classification standard, apoptosis sensitivity type classification standard and proliferation regulation type classification standard are obtained; Based on the chemotherapy resistance type classification standard, apoptosis sensitivity type classification standard and proliferation regulation type classification standard, the input sample is subjected to pattern matching processing, and the pathway response feature identifier is obtained. The pathway response feature identifier is subjected to treatment target matching processing, and the PRAK pathway intervention target combination for specific disease phenotypes is obtained, including: 6. The PRAK signaling pathway analysis method of claim 1, wherein, Perform disease phenotype correlation analysis on the chemotherapy resistance, apoptosis sensitivity, and proliferation regulation in the pathway response signature, to obtain a chemotherapy resistance-disease mapping relationship, an apoptosis sensitivity-disease mapping relationship, and a proliferation regulation-disease mapping relationship; Perform intervention assess processing on the HSP27 and GTPBP4 key target points based on the chemotherapy resistance-disease mapping relationship, the apoptosis sensitivity-disease mapping relationship, and the proliferation regulation-disease mapping relationship, to obtain an HSP27 target point intervention priority and a GTPBP4 target point intervention priority; Perform drug-target affinity calculation processing according to the HSP27 target point intervention priority and the GTPBP4 target point intervention priority, to obtain a binding strength score of the candidate drug molecule and HSP27 and a binding strength score of the candidate drug molecule and GTPBP4; Input the binding strength score of the candidate drug molecule and HSP27 and the binding strength score of the candidate drug molecule and GTPBP4 into a synergistic effect prediction algorithm to perform multi-target joint intervention evaluation processing, to obtain an HSP27-GTPBP4 target point combination synergistic index; Perform optimal selection processing based on the HSP27-GTPBP4 target point combination synergistic index, to obtain the PRAK pathway intervention target point combination.
7. The PRAK signaling pathway analysis method of claim 6, wherein, The optimal selection processing based on the HSP27-GTPBP4 target point combination synergistic index to obtain the PRAK pathway intervention target point combination includes: Perform numerical standardization processing on the HSP27-GTPBP4 target point combination synergistic index, to obtain a normalized synergistic index matrix; Perform weight calculation processing on the target point combination effect according to the normalized synergistic index matrix, to obtain an HSP27-GTPBP4 combination weight coefficient; Perform multi-objective optimization function construction processing based on the HSP27-GTPBP4 combination weight coefficient, to obtain a target point combination optimization objective function; Input the target point combination optimization objective function into a genetic algorithm to perform global optimization processing, to obtain an optimal target point combination solution set; Perform Pareto front screening processing on the optimal target point combination solution set, to obtain the PRAK pathway intervention target point combination.
8. A PRAK signaling pathway analysis system, comprising: The PRAK signal pathway analysis system for implementing the PRAK signal pathway analysis method in any of claims 1-7 includes: A collection module for performing multi-omics data collection processing on biological samples through high-throughput sequencing technology, to obtain a comprehensive data set containing circRNA expression profiles, miRNA regulation data, and PRAK pathway protein phosphorylation data; A mining module for performing molecular interaction relationship mining processing on the comprehensive data set, to obtain a pathway connection matrix containing a P38-PRAK-HSP27 signal axis, a VEGFA-PRAK activation axis, and a PRAK-GTPBP4-RhoA regulation axis; A prediction module for performing pathway activity state prediction processing on the pathway connection matrix through a graph neural network algorithm, to obtain PRAK pathway activation intensity values under different conditions; An extraction module is configured to perform a perturbation response feature extraction process according to the PRAK pathway activation intensity value, to obtain a pathway response feature identifier including a chemotherapy-resistant type, an apoptosis-sensitive type, and a proliferation regulation type. A matching module is configured to perform a treatment target matching process on the pathway response feature identifier, to obtain a PRAK pathway intervention target combination for a specific disease phenotype.
9. A PRAK signaling pathway analysis device, comprising: The computer program, when executed by the processor, causes the processor to perform the PRAK signal pathway analysis method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, causes the processor to perform the PRAK signal pathway analysis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Drug-disease interaction prediction method, device, medium and product
CN120048332A
Mapkap kinase-2 as a specific target for blocking proliferation of P53-defective cells
US20090010927A1