Method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization
Through the improved multi-objective particle swarm optimization algorithm and graph attention mechanism, the poor model performance problem caused by excessive parameters in the existing technology is solved, and more efficient correlation recognition of circRNA with specific clinical manifestations is achieved, which significantly improves the prediction ability and accuracy of the model.
Patent Information
- Application Number
- CN202410883053.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-03
AI Technical Summary
The prior art in predicting that circRNA is associated with specific clinical manifestations, excessive parameters lead to poor model performance and lacks a clear approach when selecting metapathways.
The improved multi-objective particle swarm optimization algorithm and graph attention mechanism (GAT) are used to obtain topological context vectors through similarity calculation and random swimming, and combined with Pareto dominant sorting and extrusion distance comparison optimization parameters, and finally input the GAT model for correlation evaluation.
It improves the accuracy and efficiency of the association identification of circRNA with specific clinical manifestations, significantly improves the prediction ability and accuracy of the model, especially the performance performance under different thresholds.
Smart Images

Figure CN118762755B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of complex networks and artificial intelligence bioinformatics, and particularly relates to a technique for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization. Background Art
[0002] circRNA is a covalently closed RNA molecule. In the early research stage, circRNA was considered as spliced error cellular waste and had no significance in biochemical processes. However, with the rapid development of high-throughput sequencing technology, a large number of wet experiments have shown that the structure of circRNA provides the ability to resist exonuclease degradation, some circRNAs are highly expressed in tissues and cells, and are involved in important biological processes. However, it is time-consuming and resource-consuming to establish the circRNA-specific clinical manifestation relationship by biological experiments. Therefore, it is crucial to develop a computational method for predicting the circRNA-specific clinical manifestation connection.
[0003] In the past few decades, many computational methods have been proposed, mainly divided into three types: network-based methods, machine learning-based methods, and deep learning-based methods. Among them, network-based methods are the most commonly used by researchers. In 2019, the restart random walk strategy was implemented on the circRNA similarity network, and the features of each circRNA-specific clinical manifestation pair were extracted according to the random walk results and the known association matrix. Machine learning-based methods mainly focus on classification algorithms and feature extraction methods to predict associations. In 2020, it was proposed that using metapaths and matrix factorization on multiple data sources could predict potential circRNA-specific clinical manifestation pairs. The third category is based on deep learning. A method called AE-DNN was proposed, which uses the initial features to input into a deep autoencoder to extract more high-level features and uses a deep neural network for prediction. In addition, in 2023, it was proposed to use an autoencoder to enhance multiple biological features, and then use collaborative message propagation on the central network of different specific clinical manifestations and circRNAs to extract high-order features.
[0004] Although the above methods have been widely used, too many parameters are difficult to optimize the performance of the model. In addition, although the prior art performs metapath-guided random walks to learn node embeddings, it is still unclear how to select metapaths from a given heterogeneous graph. The above problems inspired the key idea of the invention. This invention proposes a new method for predicting the circRNA-specific clinical manifestation association based on an improved multi-objective particle swarm optimization algorithm and balancing the sequence features of node neighbor types. Summary of the Invention
[0005] The object of the present invention is to improve the method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization.
[0006] The present invention is a method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization, including the following steps:
[0007] S1. Similarity calculation: Calculate various similarities between circRNA and specific clinical manifestations, and construct a heterogeneous graph network of specific clinical manifestations and circRNA.
[0008] S2. Since only using similarity may lead to the lack of graph structure, therefore, apply random walk with jump and stay strategies to obtain the topological context vectors of circRNA and specific clinical manifestations, and fuse the various similarities between circRNA and specific clinical manifestations with the topological vectors to obtain the node feature embedding of the heterogeneous graph.
[0009] S3. Use an improved multi-objective particle swarm optimization algorithm to optimize the parameters. This algorithm adopts new particle update and elimination strategies, and determines pbest and gbest through Pareto dominance sorting, crowding distance comparison, and tournament selection, reaches the Pareto optimal solution set area, and obtains the optimal parameters.
[0010] S4. Input the optimal parameters into the graph attention mechanism (GAT), which extracts more expressive features by learning the weights between nodes, uses the GAT model to evaluate the association between circRNA and specific clinical manifestations, and conducts comparative analysis with classical models to verify the accuracy and effectiveness of the present invention.
[0011] Compared with the prior art, the present invention has the following beneficial effects:
[0012] 1. Better prediction ability under different thresholds: The present method realizes the improvement of the identification of the association between circRNA and specific clinical manifestations. Under five-fold cross-validation, the higher the percentage (AUC) of the area under the curve generated by ROC in the total possible area, the better the performance of the classifier.
[0013] 2. Higher model accuracy: The present method realizes the improvement of the accuracy of association identification and obtains a higher F1 score. The F1 score indicates the performance of the model by counting each predicted answer in the entire sample, and is calculated as: where Pre is the precision of the sample, representing the percentage of all predicted results of the model that are correct, and is calculated as Rec is the regression value of the sample, indicating how many of all standard label answers the model correctly identified, and is calculated as Acc accuracy rate is one of the most commonly used evaluation indicators, and is calculated as
[0014] 3. More accurate prediction of the top n results of specific clinical manifestations: This method conducted a case study on the data of three common clinical manifestations, namely gastric cancer, osteosarcoma, and colorectal cancer, and their related circRNAs. In the generated prediction results, these prediction scores were sorted in descending order, and according to the literature in PubMed, it was verified that these circRNAs were all related to these three clinical manifestations. Brief Description of the Drawings
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 It is a flowchart of a method for identifying the association between circRNA and specific clinical manifestations based on improved multi-objective particle swarm optimization. Figure 2 It is a schematic diagram of the parameters based on multi-objective particle swarm optimization. Tables 1 and 2 show the comparison results of the method of the present invention with 8 methods for identifying the association between circRNA and specific clinical manifestations. Embodiment
[0017] As Figure 1 shown, the present invention is a method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization. This method first uses the existing information of circRNA and specific clinical manifestations to obtain various similarities, and then obtains the topological context vectors of circRNA and specific clinical manifestations from the heterogeneous graph to obtain the node embedding feature vectors of the heterogeneous graph. Then, an improved multi-objective particle swarm optimization algorithm is used to optimize the parameters, and finally, the obtained parameters are input into the GAT to complete the identification task. The method includes a data preprocessing and training stage, a multi-objective particle swarm training stage, and a node feature embedding stage, and its steps are as follows:
[0018] S1. Similarity calculation: Calculate various similarities between circRNA and specific clinical manifestations, and construct a heterogeneous graph network of specific clinical manifestations and circRNA;
[0019] S2. Since only using similarities may lead to the lack of graph structure, therefore, a random walk with a jump and stay strategy is applied to obtain the topological context vectors of circRNA and specific clinical manifestations, and the various similarities between circRNA and specific clinical manifestations are fused with the topological vectors to obtain the node feature embedding of the heterogeneous graph;
[0020] S3. Use the improved multi-objective particle swarm optimization algorithm to optimize the parameters. This algorithm adopts a new particle update and elimination strategy, and determines pbest and gbest through Pareto domination sorting, crowding distance comparison, and tournament selection, reaches the Pareto optimal solution set region, and obtains the optimal parameters.
[0021] S4. Input the optimal parameters into the graph attention mechanism (GAT). It extracts more expressive features by learning the weights between nodes, uses the GAT model to evaluate the association between circRNA and specific clinical manifestations, and conducts a comparative analysis with the classical model to verify the accuracy and effectiveness of the present invention.
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0024] Such as Figure 1 The embodiments of the present invention provide a method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization;
[0025] The method for identifying the association between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization of the present invention includes:
[0026] S1. Similarity calculation, calculate various similarities between circRNA and specific clinical manifestations, and construct a heterogeneous graph network of specific clinical manifestations and circRNA;
[0027] Specifically, in the embodiments of the present invention, step S1 includes the following steps:
[0028] S11. circRNA sequence similarity: The similarity between any two circRNA sequences is calculated according to the Levenshtein distance. The greater the similarity, the circRNA r i and circRNA r j The similarity calculation is as follows:
[0029]
[0030] Among them, LevDis(r i , r j ) represents r iConvert to r j The minimum number of operations, len(r i ) represents the sequence length of ri;
[0031] S12, circRNARBP similarity: circRNA-related RBPs are used to measure the interactions between circRNAs. i and circRNA j The similarities between them are described as follows:
[0032]
[0033] Where R i Representative and circRNA i The relevant RBP set, R = R i ∩R j ;
[0034] S13. Similarity of specific clinical manifestations: Query specific clinical manifestations through symptom information. Specific clinical manifestation i is represented as a group of symptoms, denoted as in Calculated by TF-IDF:
[0035]
[0036] tf(i, j) represents the strength of association between specific clinical manifestation i and symptom j. The similarity matrix of specific clinical manifestations is calculated by cosine similarity as follows:
[0037]
[0038] S14. Semantic similarity of specific clinical manifestations: For a specific clinical manifestation c i , which can be expressed as in Indicates a specific clinical manifestation node, Represents the relationship between various specific clinical manifestation nodes. i and c j The similarity between is calculated as:
[0039]
[0040] in Represents specific clinical manifestationsc i and its ancestors in the DAG, express All nodes in the specific clinical manifestations c i The contribution value of
[0041]
[0042] Among them, the child node (t) represents the node belonging to the child nodes of t. Represents the semantic contribution score of a specific clinical manifestation, which is set to 0.5.
[0043] S15. Gaussian Interaction Spectrum Kernel Similarity: To complete the similarity matrix, the Gaussian Interaction Spectrum Kernel (GIP) similarity is introduced. In the known circRNA-specific clinical manifestation association matrix A, its j-th column Col(j) represents the association between a specific clinical manifestation c j and all circRNAs. The GIP similarity is calculated as follows:
[0044] GIPSim RNA (r i ,r j ) = exp(-ρ r |||Row(i i ) - Row(i j )|| 2 )
[0045] GIPSim dis (c i ,c j ) = exp(-ρ d ||Col(c i ) - Col(c j )|| 2 )
[0046] Among them, ρ r and ρ d are introduced to control the kernel bandwidth, and the calculation formula is as follows:
[0047]
[0048] Among them, n and m respectively represent the number of circRNAs and specific clinical manifestations in the above dataset;
[0049] S16. Integration of Different Similarities: To solve the problem that the single similarity between circRNAs and specific clinical manifestations is too sparse, the sequence similarity, RBP similarity, and GIP similarity of circRNAs are integrated, and their combination formula is as follows:
[0050]
[0051] Similarly, the kernel similarity is also a supplement to the semantic similarity of specific clinical manifestations and the representation of specific clinical manifestations. The comprehensive similarity between specific clinical manifestations c i and c j is calculated as follows:
[0052]
[0053] S2. Since using only similarity may lead to the lack of graph structure, a random walk with jump and stay strategies is applied to obtain the topological context vectors of circRNAs and specific clinical manifestations, and multiple similarities of circRNAs and specific clinical manifestations are fused with the topological vectors to obtain the node feature embeddings of the heterogeneous graph;
[0054] Specifically, in the embodiment of the present invention, step S2 includes the following steps:
[0055] S21. The heterogeneous graph network of specific clinical manifestations and circRNAs is represented as G=(V, E), where V includes circRNA and specific clinical manifestation nodes and E includes specific clinical manifestation-specific clinical manifestation, circRNA-circRNA, and specific clinical manifestation-circRNA links. Based on the similarity matrix of specific clinical manifestations and circRNAs, a specific clinical manifestation subgraph MD and a circRNA subgraph MC are respectively constructed.
[0056]
[0057] Next, MD, MC, and A are combined into an adjacency matrix H to represent the association relationship of nodes, as follows:
[0058]
[0059] Therefore, the number of nodes in H is the sum of the number of circRNA and specific clinical manifestation nodes, and the nodes Edges Then, the heterogeneous graph is composed of circRNAs and specific clinical manifestations and is implemented through the Deep Graph Library. The initial node feature matrix of the heterogeneous graph is defined as:
[0060]
[0061] In the formula, k represents the number of features of each node, and x i ∈R 1×k represents the feature of the i-th node in the heterogeneous graph. At the same time, x i is also the i-th row in the matrix x, where it is further represented as:
[0062]
[0063] Specific transformation matrices and are designed to put circRNA and specific clinical manifestation nodes into the same feature space respectively;
[0064] S22. Initially, the method performs a random walk on the input heterogeneous graph by probabilistically balancing the jump and stay options when selecting the next node. It treats the random walk on the heterogeneous graph as a two-step calculation. That is, the first step is to decide whether to jump or stay, and the second step is, if jumping, to select the position to jump to. The node sequence obtained from the above two steps is sent into the Word2Vec model to output node embeddings. In the problem of the present invention, for a node v ∈ V collected during the random walk process, the selection strategy for the next node vnext is expressed as:
[0065] (I.) Jump to nodes of other types, and use the method of uniform sampling to select a node vnext that is heterogeneous with v;
[0066] (II.) Stay with the same type of v, and use uniform sampling to select a node vnext that is homogeneous with v. For the current node v, select P stay (v) to represent the following probability of staying, through:
[0067]
[0068] where and represent the isomorphic and heterogeneous neighbors of v, is the initial probability of Stay, and l is equal to the number of consecutive nodes of the same type as node v selected during the random walk process;
[0069] Specifically, use the above jump and stay strategies for random walks on the heterogeneous graph, and for each node v ∈ V, initialize a random walk w starting from v i until the maximum length Lmax is reached, and generate a set of random walk sets W by performing r random walks on each node v;
[0070] Based on the generated set of random walks W, use the Word2Vec model to learn node embeddings. Word2Vec can maximize the co-occurrence probability of nodes appearing in a specific context window in the random walk sequence to learn node embeddings. For a pair of nodes v i and v j appearing in the context window in W, the co-occurrence probability is defined as:
[0071]
[0072] In the formula, is the s-type function, v i and v j respectively represent the embeddings of v i and v j .
[0073] S3. Use the improved multi-objective particle swarm optimization algorithm to optimize the parameters. This algorithm adopts a new particle update and elimination strategy, and determines the best positions pbest and gbest through Pareto dominance sorting, crowding distance comparison, and tournament selection, reaches the Pareto optimal solution set region, and obtains the optimal parameters;
[0074] Specifically, in the embodiment of the present invention, step S3 includes the following steps:
[0075] S31. In the multi-objective particle swarm optimization algorithm (MOPSO), each possible solution is called a "particle", and the positions and velocities of these particles are continuously adjusted through an iterative process to finally find the optimal solution. The particles update their velocities and positions through the following formulas:
[0076] V i,j (t + 1) = ωV i,j (t) + c1 × rand() × [pbest i,j -x i,j (t)] + c2 × rand() × [gbest i,j -x i,j (t)]
[0077] x i,j (t + 1) = x i,j (t) + V i,j (t + 1), j = 1, 2, … d
[0078] In the formula, i = 1, 2, … M, M is the total number of particles in the population, V i is the velocity of the particle, pbest and gbest respectively define the best position of the particle and the best position of the population, rand() is a random number between (0, 1), x i is the current position of the particle, c1 and c2 are learning factors, and ω is the compression factor;
[0079] This invention uses a new particle update strategy. Based on Pareto dominance sorting and crowding distance comparison, it adopts a tournament selection strategy to select the individual historical best positions pbest and gbest, and during the process of searching for the Pareto optimal solution, it performs Pareto non-dominated sorting on the population particles. In the update process of each generation, the following steps are used to eliminate and update the particles. The implementation steps are as follows:
[0080] Step 1: Randomly initialize the population, that is, the positions and velocities of each particle are initialized to random values;
[0081] Step 2: Solve the values of each objective function, sort based on the Pareto dominance relationship according to fitness, and use the tournament selection strategy to select pbest and gbest;
[0082] Step 3: Update the population according to the above formula;
[0083] Step 4: Evaluate the values of each objective function, sort based on the Pareto dominance relationship according to fitness, and use the tournament selection strategy to compare the current individual position with the individual's historical best position, and select pbest and gbest;
[0084] Step 5: Use the number of iterations as the termination condition, output the individuals of the last generation of the population, otherwise execute Step 3.
[0085] In this invention, an improved multi-objective particle swarm optimization algorithm is adopted to optimize the parameters in the multi-head attention mechanism. Input the parameters to obtain the fitness surface representing the parameters, that is, the optimal solution. The particle swarm optimization algorithm better adapts to the characteristics of the data on the association between circRNA expression and specific clinical manifestations through an improved update formula, a weighted objective function, and a new particle elimination strategy.
[0086] S4. Input the obtained optimal parameters into the graph attention mechanism (GAT). It extracts more expressive features by learning the weights between nodes, uses the GAT model to evaluate the association between circRNA and specific clinical manifestations, and conducts a comparative analysis with the classical model to verify the accuracy and effectiveness of this invention.
[0087] Specifically, in the embodiment of this invention, step S4 includes the following steps:
[0088] S41. Use GAT to further extract features. Its optimal parameters are from Section S31. The graph attention network is a graph neural network constructed based on the attention mechanism, which can effectively capture the complex relationships between nodes and perform node-level tasks. The graph attention network dynamically adjusts the importance of information transmission by learning the weights between each node and its adjacent nodes, so as to more accurately capture the detailed information of the structure. Through the graph attention network, this model further extracts more representative and expressive features on the basis of the original features, providing better input for the prediction task;
[0089] For the network G, the update of its nodes is expressed as:
[0090]
[0091] Among them, is the new feature representation of node i, p k is the weight matrix of the Kth layer attention mechanism, and K represents the total number of heads in the multi-head attention mechanism. is the input feature vector of node i, represents the attention coefficient of the K-th layer between node i and node j;
[0092] Using the graph attention network can effectively aggregate information from neighboring nodes, while considering the different importance between nodes, thereby generating more accurate and expressive node features, making this model have stronger expressive ability and prediction performance when dealing with complex graph structures;
[0093] Finally, the extracted features are input into a multi-layer perceptron (MLP) for prediction. The multi-layer perceptron is a deep learning model that can perform non-linear transformations on input features. It uses the sigmoid function as the activation function to map the output value to the interval (0,1), thereby generating a prediction score;
[0094] S42, AUC: AUC is the area under the curve generated by ROC, which is the percentage of the total possible area;
[0095] S43, Model evaluation metrics: The F1 score can indicate the performance of the model by counting each predicted answer in the entire sample, and is calculated as: where Pre is the precision of the sample, representing the percentage of all predicted results of the model that are correct, and is calculated as Rec is the regression value of the sample, indicating how many of all the standard label answers the model correctly identified, and is calculated as The Acc accuracy rate is one of the commonly used evaluation metrics, and is calculated as
[0096] S44, Case study: This method sets all circRNA-specific clinical manifestation samples as unknown, and then uses this method to generate association scores for three common specific clinical manifestations of gastric cancer, osteosarcoma, and colorectal cancer. In the generated prediction results, these scores are sorted in descending order, and it is proven according to the literature of PubMed that these circRNAs are all related to specific clinical manifestations.
[0097] Table 1 shows the data comparison between the method of the present invention and 8 circRNA-specific clinical manifestation association recognition methods on the CircR2Disease v1.0 dataset:
[0098] Table 1:
[0099]
[0100] Table 2 shows the data comparison between the method of the present invention and 8 circRNA-specific clinical manifestation association recognition methods on the CircR2Disease v2.0 dataset:
[0101] Table 2:
[0102]
[0103] Tables 1 and 2 show the comparison results of the method of the present invention and eight methods for identifying the association between circRNAs and specific clinical manifestations.
[0104] The various embodiments of the present invention are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for related parts.
[0105] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A method for identifying associations between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization, characterized in that: First, data preprocessing is performed, and multiple similarities are obtained using existing information about circRNA and specific clinical manifestations. Next, topological context vectors of circRNA and specific clinical manifestations are obtained from heterogeneous graphs, and heterogeneous graph node embedding feature vectors are obtained. Then, an improved multi-objective particle swarm optimization algorithm is used to optimize parameters. Finally, the obtained parameters are input into the graph attention mechanism to complete the recognition task. The method includes data preprocessing and training stages, multi-objective particle swarm training stages, and node feature embedding stages, and the steps are as follows: S1. Similarity calculation: calculate multiple similarities between circRNA and specific clinical manifestations, and construct a heterogeneous graph network of specific clinical manifestations and circRNA; S2, applying random walks with jump and stay strategies to obtain topological context vectors of circRNAs and specific clinical manifestations, fusing multiple similarities of circRNAs and specific clinical manifestations with the topological vectors to obtain node feature embeddings of heterogeneous graphs; S3, use the improved multi-objective particle swarm optimization algorithm to optimize the parameters. The algorithm adopts a new particle update and elimination strategy, determines pbest and gbest through Pareto dominance sorting, exclusion distance comparison and tournament selection, reaches the Pareto optimal solution set area, and obtains the optimal parameters; S4. The optimal parameters are input into the graph attention mechanism, which extracts more expressive features by learning the weights between nodes. The graph attention mechanism is used to evaluate the association between circRNA and specific clinical manifestations, and a comparative analysis is performed with the classic model to verify the accuracy and effectiveness of the above method.
2. The method for identifying associations between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization according to claim 1, characterized in that: The specific sub-steps of step S2 are as follows: S21. The heterogeneous graph network of specific clinical manifestations and circRNAs is represented as G = (V, E), where V includes circRNA and specific clinical manifestation nodes and E includes specific clinical manifestation-specific clinical manifestation, circRNA-circRNA and specific clinical manifestation-circRNA links. Based on the similarity matrix of specific clinical manifestations and circRNAs, the specific clinical manifestation subgraph MD and the circRNA subgraph MC are constructed respectively, and the corresponding graph adjacency matrix is constructed. Next, MD, MC, and A are combined into an adjacency matrix H to represent the relationship between nodes, as follows: Therefore, the number of nodes in H is the sum of the number of circRNA and specific clinical manifestation nodes. Edge (h i ,h j )∈E={h i,j =1∈(A∪MD∪MC)}, Then the heterogeneous graph consists of circRNA and specific clinical manifestations, which is implemented through the deep graph library, and the initial node feature matrix of the heterogeneous graph is defined as: In the formula, k represents the number of features of each node, x i ∈R 1×k represents the feature of the i-th node in the heterogeneous graph. At the same time, x i is also the i-th row in the matrix x, which is further represented as: Designing a specific type of transformation matrix and The circRNA and specific clinical manifestation nodes were placed into the same feature space; S22. At the beginning, the method performs a random walk on the input heterogeneous graph by probabilistically balancing the jump and stay options when selecting the next node. It regards the random walk on the heterogeneous graph as a two-step calculation. Simply put, the first step is to decide whether to jump or stay, and the second step is to choose the jump location when choosing the jump in the first step. The node sequence obtained in the above two steps is sent to the Word2Vec model and the node embedding is output. For a node v∈V collected during the random walk, the selection strategy of the next node vnext is expressed as: (I.) Jump to other types of nodes and use uniform sampling to select vnext, a node that is heterogeneous with vnext; (II.) Keep v of the same type, use uniform sampling to select vnext and v's homogeneous node vnext, and for the current node v, select P stay (v) is used to express the probability of the following stays, by: in and represents the homogeneous and heterogeneous neighbors of v, ∝∈[0,1] is the initial probability of Stay, and l is equal to the number of consecutive nodes of the same type as node v that are selected during the random walk; Specifically, we perform random walks using the above jump and hold strategy to sample from the input heterogeneous graph. More precisely, for each node v∈V, we initialize a random walk w from v i Starting from, until the maximum length Lmax is reached, similar to the existing random walk-based graph embedding techniques, a set of random walks W is generated by performing r random walks on each node v; Based on the generated random walk set W, the Word2Vec model is used to learn node embedding. Word2Vec can maximize the co-occurrence probability of nodes appearing in a specific context window in the random walk sequence to learn node embedding. Formally, for a pair of nodes v appearing in the context window in W i and v j , the co-occurrence probability is defined as: In the formula, is a s-type function, v i and v j Respectively represent v i and v j Embedding.
3. The method for identifying associations between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization according to claim 1, characterized in that: The specific sub-steps of step S3 are as follows: S31. In the multi-objective particle swarm optimization algorithm, each possible solution is called a "particle", and the position and speed of these particles are continuously adjusted through an iterative process to eventually find the optimal solution. The particles update their speed and position through the following formula: V i,j (t+1)=ωV i,j (t)+c1×rand()×[pbest i,j -x i,j (t)]+c2×rand()×[gbest i,j -x i,j (t)] x i,j (t+1)=x i,j (t)+V i,j (t+1),j=1,2,…d In the formula, i = 1, 2, ... M, where M is the total number of particles in the group, V i is the speed of the particle, pbest and gbest define the best position of the particle and the best position of the population respectively, rand() is a random number between (0, 1), x i is the current position of the particle, c1 and c2 are learning factors, and ω is the compression factor; An improved multi-objective particle swarm optimization algorithm is used to optimize parameters. A new particle update strategy is used. Based on Pareto dominance sorting and exclusion distance comparison, a tournament selection strategy is used to select the individual historical best positions pbest and gbest. In the process of searching for the Pareto optimal solution, the algorithm sorts the population particles based on Pareto non-inferiority. In the update process of each generation, the following steps are used to eliminate and update the particles, and finally all the particles in the population reach the Pareto optimal solution set area and obtain the optimal parameters. The implementation steps of the multi-objective particle swarm optimization algorithm based on the Pareto optimal solution are as follows: Step 1: Randomly initialize the population, that is, the position and velocity of each particle are initialized to random values; Step 2: Solve the values of each objective function, sort them based on the Pareto dominance relationship according to fitness, and use the tournament selection strategy to select pbest and gbest; Step 3: Update the population according to the above formula; Step 4: Evaluate the values of each objective function, sort them based on the Pareto dominance relationship according to their fitness, use the tournament selection strategy to compare the current individual position with the individual's best historical position, select pbest, and use this method to select the best position gbest of the population; Step 5 uses the number of iterations as the termination condition and outputs the last generation of population individuals, otherwise, execute step 3; 4. The method for identifying associations between circRNA and specific clinical manifestations based on multi-objective particle swarm optimization according to claim 1, characterized in that: The specific sub-steps of step S4 are as follows: S41, use the multi-head attention mechanism to further extract features, and its optimal parameters come from Section S31; For network F, its node update is expressed as: in, is the new feature representation of node i, p k is the weight matrix of the K-th layer attention mechanism, K represents the total number of heads in the multi-head attention mechanism, is the input feature vector of node i, represents the K-th layer attention coefficient between node $i$ and node j; Finally, the extracted features were input into a multilayer perceptron for prediction. The multilayer perceptron is a deep learning model that can perform nonlinear transformations on input features. It uses a sigmoid function as an activation function to map output values to the (0,1) interval to generate prediction scores. These scores are then used to predict the potential association between circRNAs and specific clinical manifestations.
Citation Information
Patent Citations
Ring RNA-disease association prediction method and device based on weighted graph attention and heterogeneous graph neural network, and medium
CN115798730A
CircRNA-disease association prediction method based on scale map convolution network and feature convolution
CN118053574A