Web API recommendation method based on graph diffusion reconstruction and graph contrast learning
By using graph diffusion reconstruction and graph comparison learning methods in Web API recommendations, the Mashup-API call relationship diagram is expanded and reconstructed, which solves the problems of sparse data and insufficient mining of relationship features, and significantly improves the accuracy of recommendations.
Patent Information
- Application Number
- CN202411927582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing Web API recommendation methods face the problems of sparse data and insufficient mining of relationship characteristics, resulting in low recommendation performance.
Using a method based on graph diffusion reconstruction and graph comparison learning, the Mashup-API call relationship diagram is expanded and reconstructed, deep interactive information is explored, and more accurate node feature representation is generated.
The problem of sparse historical interactive data is solved through graph diffusion and adaptive graph reconstruction models, which improves the accuracy of recommendations, and extracts the relationship characteristics between Web APIs more accurately through graph comparison learning and graph convolution neural network.
Smart Images

Figure CN120067432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer technology, and more particularly, to a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning. Background Art
[0002] A Web API (Web Application Programming Interface) is an API service provided over a network. By providing a unified interface and a flexible design style, it enables communication and interaction between different applications. Mashup is a web application development technology that can aggregate different types of Web APIs to build new applications. The Web API recommendation technology can recommend high-quality Web APIs for Mashup developers from a vast amount of Web APIs, which is a research hotspot in the current software development field.
[0003] Existing Web API recommendation methods are mainly divided into Web API recommendation methods based on functional description documents and Web API recommendation methods based on service historical call information.
[0004] Web API recommendation methods based on functional description documents. Such methods usually use natural language processing technology to solve the similarity between functional description documents to achieve service recommendation. Functional description documents are usually manually written by developers, and it is very likely that the functions of the service are not fully and accurately described. Some Web APIs even do not have functional description documents, so the functional description information is often insufficient; in addition, due to the lack of a unified specification for the functional description documents of services, there may be multiple descriptions for the same function, resulting in the fact that text similarity cannot well reflect the similarity in terms of functions; and such methods do not consider the interaction relationship and collaboration information between services. Therefore, the method based on functional description documents may make poor recommendations.
[0005] Web API Recommendation Method Based on Service History Call Information. Such methods usually utilize the idea of collaborative filtering. First, Mashups and APIs are mapped into vector representations. Based on these vectors, the interactions between Mashups and APIs are reconstructed, and the probability of a Mashup calling an API is predicted. This method takes into account the interaction information and collaboration information between services and achieves relatively good results. However, in the process of mapping services into vectors, many methods only focus on other services that have interacted with the service, extract features using the direct interaction information between services, and lack the mining of deep interaction information, resulting in inaccurate feature extraction and affecting the recommendation effect; moreover, this method depends on historical call data. However, some Mashups often only call a very small number of APIs, resulting in the problem of data sparsity and affecting the final recommendation effect; when extracting node features through neighborhood aggregation, the vector representations of nodes in the same neighborhood tend to be the same, resulting in the problem of data smoothing, causing the loss of diversity of node feature vectors and affecting the recommendation effect.
[0006] In summary, most service recommendation methods first map Mashups and APIs into vector representations, predict the probability of a Mashup calling an API based on the vector representations, and then perform API recommendation. They have the following disadvantages:
[0007] Seriously affected by the data sparsity problem. The method based on functional description information depends on the functional description documents of services to estimate the relevance between Mashups and Web APIs, while the functional description documents are often limited or even missing; the method based on interaction information depends on the historical interaction data between services to estimate the relevance between Mashups and Web APIs. Due to the sparsity of historical interaction data, the final recommendation effect is not good.
[0008] Insufficient mining of relationship features. Most Web API recommendation methods only focus on the direct interaction relationships between services, lack the mining of deep interaction information, and there is a problem of data smoothing during feature extraction, resulting in inaccurate feature extraction and affecting the recommendation effect. Summary of the Invention
[0009] Aiming at the problems of data sparsity faced by existing Web API recommendation methods and insufficient mining of relationship features between services leading to low recommendation performance, the present invention proposes a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning. The present invention uses a graph diffusion and adaptive graph reconstruction model to expand the Mashup-API call relationship graph, solves the problem of sparse historical interaction data between Mashups and APIs, and improves the accuracy of recommendation at the same time.
[0010] The technical means adopted by the present invention are as follows:
[0011] A Web API recommendation method based on graph diffusion reconstruction and graph contrastive learning, comprising the following steps:
[0012] S1. Obtain the historical call information of Mashup and Web API, and construct a Mashup-API call relationship graph based on the historical call information;
[0013] S2. Perform data augmentation based on graph diffusion on the Mashup-API call relationship graph to generate a diffusion matrix of the Mashup-API call relationship;
[0014] S3. Use adaptive graph reconstruction to reconstruct the diffusion matrix to generate a contrastive view;
[0015] S4. Obtain the Mashup-API call relationship graph and the contrastive view, use the trained Light GCN as a basic graph encoder to process the Mashup-API call relationship graph and the contrastive view, and generate the final feature representation vectors of the nodes;
[0016] S5. Sort the APIs that meet the Mashup requirements based on the inner product between the feature representation vectors of the Mashup nodes and the feature representation vectors of the API nodes, so as to generate a recommendation list.
[0017] Further, constructing a Mashup-API call relationship graph based on the historical call information includes:
[0018] S101. Store each Mashup and its called API list in the form of a dictionary to obtain structured Mashup-API call information;
[0019] S102. Based on the structured Mashup-API information, regard Mashup and API as nodes, regard the call relationship between Mashup and API as edges, and construct a bipartite graph of the call relationship between Mashup and API.
[0020] Further, performing data augmentation based on graph diffusion on the Mashup-API call relationship graph to generate a diffusion matrix of the Mashup-API call relationship includes:
[0021] S201. Represent the Mashup-API call relationship graph as a sparse graph G:
[0022] G=(V,E),
[0023] where V represents the node set, including the Mashup node set M and the API node set N; E represents the edge set;
[0024] S202. Convert the sparse graph G into an adjacency matrix A for representation. The adjacency matrix A represents the connection relationship between nodes, where A[i][j] being 1 indicates that there is an edge between node i and node j, otherwise it is 0;
[0025] S203. Add self-loops to the adjacency matrix A, and calculate the symmetric transition matrix based on the adjacency matrix after adding self-loops;
[0026] S204. Perform random walks on the symmetric transition matrix to generate a dense diffusion matrix, and sparsify the dense diffusion matrix based on a preset threshold to obtain the final diffusion matrix.
[0027] Further, perform random walks on the symmetric transition matrix to generate a dense diffusion matrix according to the following formula:
[0028] S = α(I - (1 - α)T sym ) -1
[0029] where α is the probability of random walk, I is the identity matrix, and T sym represents the symmetric transition matrix, and:
[0030]
[0031] where A loop represents the adjacency matrix after adding self-loops, D loop represents the self-loop degree matrix, and:
[0032] D loop = diag(A loop ·1)
[0033] A loop = I + A
[0034] where 1 represents the all-ones vector and diag represents the diagonal matrix.
[0035] Further, perform reconstruction on the diffusion matrix using adaptive graph reconstruction to generate a contrast view, including:
[0036] S301. Use a graph convolutional neural network to extract the relationship features in the diffusion matrix;
[0037] S302. Calculate the distance between each feature Mashup i and each API j in the feature space according to the relationship features of the nodes;
[0038] S303. Normalize the node feature space distance using the Sigmod function and map it between 0 and 1 to obtain the distance factor θ ij ;
[0039] S304. Screen the distance factor based on a preset threshold for reconstructing the Mashup-API relationship graph.
[0040] Furthermore, use the trained Light GCN as the basic graph encoder to process the Mashup-API call relationship graph and the contrast view, and generate the final feature representation vectors of the nodes, including:
[0041] S401. For the nodes in the graph, their representation at the l+1 layer:
[0042]
[0043] where and represent the feature vector of API i and the feature vector of Mashup j respectively, l+1 represents the number of propagation layers, N j represents the set of APIs that have interacted with Mashup j, N i represents the set of Mashups that have interacted with API i, |N j | represents the number of APIs in N j | represents the number of Mashups in N i | represents the number of Mashups in N i ;
[0044] S402. Increase the depth of the Light GCN, repeat S401, and obtain multiple feature representations {h 1 , h 2 ,..., h L} of the nodes at each layer;
[0045] S403. Perform layer combination of features, stack the multiple representations of the nodes at different layers, and obtain the final feature representation:
[0046]
[0047] where, h i and h j represent the final feature representation vectors of API i and Mashup j respectively, L represents the depth of the Light GCN; α l represents the importance parameter of the embedding at the l-th layer in constructing the final embedding.
[0048] Furthermore, the inner product between the feature representation vector of the Mashup node and the feature representation vector of the API node is calculated according to the following formula:
[0049]
[0050] where, h i and hj respectively represent the final feature representation vectors of API i and Mashup j.
[0051] Furthermore, when training the Light GCN, a joint training strategy is adopted, and three subtasks of BPR loss, regularization loss, and contrastive loss are trained respectively.
[0052] The objective function of the BPR loss subtask is:
[0053]
[0054] where j represents the corresponding Mashup, and i pos represents the API that has interacted with j, and i neg represents the API that has not interacted with j, Θ = {(j, i pos , i neg )} represents the paired training data. The symbol σ(·) represents the Sigmod function, and λ bpr is the regularization strength parameter of the BPR loss, which is used to control the trade-off between the model fitting the training data and maintaining the generalization ability;
[0055] The objective function of the regularization loss subtask is:
[0056]
[0057] where λ reg represents the regularization strength parameter of the regularization loss, n is the number of parameters of the Light GCN model, and ω i is the i-th parameter of the Light GCN model;
[0058] The objective function of the contrastive loss subtask is:
[0059]
[0060] where h and h c respectively represent the feature representation vectors of the node in the original graph and the contrast view, sim(x, y) represents the cosine similarity between two vectors x and y, and τ represents the temperature coefficient in contrastive learning.
[0061] Compared with the prior art, the present invention has the following advantages:
[0062] The present invention proposes a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning. This method uses graph diffusion and an adaptive graph reconstruction model to expand the Mashup-API call relationship graph, solving the problem of sparse historical interaction data between Mashup and API. It uses graph contrast learning and graph convolutional neural network to mine the deep interaction information between Web APIs, improving the accuracy of recommendations. Experimental results on real datasets show that compared with the baseline method, the method of the present invention has a significant improvement in various indicators. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0064] Figure 1 It is a flowchart of a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0066] As Figure 1 shown, the present invention provides a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning, including the following steps:
[0067] S1. Obtain the historical call information of Mashup and Web APIs, and construct a Mashup-API call relationship graph based on the historical call information. Specifically, it includes:
[0068] S101. Store each Mashup and its called API list in the form of a dictionary to obtain structured Mashup-API call information.
[0069] The information of Mashup and API is unstructured. First, the present invention preprocesses the original data information, extracts the interaction information of Mashup and API, and stores the list of each Mashup and the APIs it calls in the form of a dictionary to obtain structured Mashup-API call information. The dictionary obtained after preprocessing is D = {Mashup 1 :{API 1 ,...,API k},...,Mashup m :{API i ,...,API j}}, and the corresponding Mashup set M: {Mashup 1 ,Mashup 2 ,...,Mashup m} and API set N: {API 1 ,API 2 ,...,API n}, where m is the number of Mashups, n is the number of APIs, and each API may appear in the call lists corresponding to different Mashups.
[0070] S102. Based on the structured Mashup-API information, regarding Mashup and API as nodes and the call relationship between Mashup and API as edges, construct a bipartite graph of the call relationship between Mashup and API.
[0071] Extract the interaction relationship from the structured Mashup-API information. The present invention constructs a bipartite graph of the call relationship between Mashup and API by regarding Mashup and API as nodes and the call relationship between Mashup and API as edges. In this way, construct a Mashup-API relationship graph represented as G = (V, E), where V represents the set of nodes in the graph, including Mashup and API, and E represents the set of edges in the graph.
[0072] S2. Perform data augmentation based on graph diffusion on the Mashup-API call relationship graph, so as to generate a diffusion matrix of the Mashup-API call relationship.
[0073] To improve the embedding quality of node feature extraction, graph contrastive learning is a very effective solution. By enhancing image data to generate contrastive views as auxiliary training signals, they participate in the training of the model together with the original data information. Considering the data sparsity problem in the Web API recommendation field, this application designs a contrastive view generation method based on a graph diffusion model and adaptive graph reconstruction. This step is mainly used for relationship data augmentation based on the graph diffusion model. Specifically, it includes:
[0074] S201. Represent the Mashup-API call relationship graph as a sparse graph G:
[0075] G = (V, E),
[0076] where V represents the set of nodes, including the Mashup node set M and the API node set N; E represents the set of edges.
[0077] S202. Convert the sparse graph G into an adjacency matrix A representation. The adjacency matrix A represents the connection relationship between nodes, where A[i][j] being 1 indicates that there is an edge between nodes i and j, otherwise it is 0. The underlying facts are more complex than what the graph captures. For example, a molecule can be described by a graph of atoms and bonds, but its underlying interactions are much more complex. The graph diffusion convolution model aims to capture the potential relationship between nodes by simulating the propagation of information and obtaining the influence degree of global nodes on the target node. The present invention uses the graph diffusion model as the encoder for graph reconstruction to expand the receptive field of nodes, enabling them to consider information from more distant neighbor nodes, which helps with more complex and global learning on the Mashup-API relationship graph to address the problem of data sparsity. Intuitively, first focus all attention on the target node, and then through the global propagation of attention, reconstruct the relationship between the target node and other nodes in the global scope. The final attention distribution defines the edges from the starting node to other nodes.
[0078] The core calculations of the graph diffusion model are shown in formulas (1) and (2).
[0079]
[0080] T = AD -1 (2)
[0081] where T defines the transition matrix, D is the diagonal matrix, and d ii = ∑ j a ij ; θ k is a coefficient defined by an optional specific diffusion variable. In this application, personalized PageRank is used to define the diffusion variable.
[0082] S203. Add self-loops to the adjacency matrix A, and calculate the symmetric transition matrix based on the adjacency matrix with self-loops added.
[0083] The addition of self-loops is calculated as shown in Equation (3).
[0084] A loop = I + A (3)
[0085] where I is the identity matrix.
[0086] The calculation of the symmetric transition matrix is as shown in Equation (4) and Equation (5).
[0087] D loop = diag(A loop · 1) (4)
[0088]
[0089] where 1 represents the all-ones vector, diag represents the diagonal matrix, D loop represents the self-loop degree matrix, and T sym represents the symmetric transition matrix.
[0090] S204. Perform random walks on the symmetric transition matrix to generate a dense diffusion matrix, and sparsify the dense diffusion matrix based on a preset threshold to obtain the final diffusion matrix.
[0091] The core calculation for generating the diffusion matrix by random walks is as shown in Equation (6).
[0092] S = α(I - (1 - α)T sym ) -1 (6)
[0093] where α is the probability of random walks. The calculated diffusion matrix S is a dense graph. Set the threshold ∈ to sparsify S and obtain the final diffusion matrix, which includes the reconstruction of Mashup-API long-distance relationships and underlying complex relationships.
[0094] S3. Reconstruct the diffusion matrix using adaptive graph reconstruction to generate a contrast view. This step is mainly used to reconstruct the diffusion matrix based on the adaptive graph reconstruction model. The focus of adaptive graph reconstruction mainly lies in the feature similarity between nodes, that is, the degree of proximity of node attributes. By analyzing and comparing node features, the relationship of the graph can be reshaped so that the graph structure can more accurately reflect the actual similarity between nodes. This application uses adaptive graph reconstruction to reconstruct the diffusion matrix S, making the generated augmented graph more in line with the original data distribution law. The traditional adaptive graph reconstruction model is mainly designed for homogeneous graphs. Drawing on previous research on graph generation models, this application designs an adaptive graph generation model for bipartite graphs for the recommendation model proposed in the present invention, which specifically includes the following steps:
[0095] S301. Use a graph convolutional neural network to extract the relationship features in the diffusion matrix.
[0096] S302. Calculate the distance between each feature Mashup i and each API j in the feature space according to the relationship features of the nodes.
[0097] S303. Use the Sigmod function to normalize the node feature space distance and map it between 0 and 1 to obtain the distance factor θ. ij 。
[0098] S304. Screen the distance factor based on a preset threshold for reconstructing the Mashup-API relationship graph. The calculation is shown in formulas (7), (8) and (9).
[0099]
[0100] Wherein, and respectively represent the feature vectors of API i and Mashup j, K represents the dimension of the feature vector, represents the k-th dimensional value of, d ij is the distance between Mashup and API in the feature space, and e ij is the edge between API i and Mashup j.
[0101] In summary, the contrast view generation algorithm based on graph diffusion reconstruction is shown in Algorithm 1.
[0102]
[0103]
[0104] S4. Obtain the Mashup-API call relationship graph and the comparison view, and use the trained Light GCN as the basic graph encoder to process the Mashup-API call relationship graph and the comparison view to generate the final feature representation vector of the nodes.
[0105] The adjacency matrix of the Mashup-API relationship graph and the comparison view contains the connection relationships between nodes. The node features are obtained by random initialization as the initial node representations. For the Mashup-API bipartite graph designed in this application, Light GCN is used as the basic graph encoder. The message passing method of Light GCN uses a simple weighted sum of the adjacency matrix without introducing additional parameters, mainly focusing on the message passing between nodes without using node-specific weight parameters to focus on mining the deep interaction information between Mashup and API and more accurately extract the relationship features between nodes. The same operation is performed on the two input views. Specifically, it includes:
[0106] S401. For the nodes in the graph, the calculation of their representations at the l+1 layer is shown in Formulas (10) and (11).
[0107]
[0108] where and represent the feature vectors of API i and Mashup j respectively, l+1 represents the number of propagation layers, N j represents the set of APIs that have interacted with Mashup j, N i represents the set of Mashups that have interacted with API i, |N j | represents the number of APIs in N j | represents the number of Mashups in N i | represents N i the number of Mashups in.
[0109] S402. Increase the depth of Light GCN, and repeat S401 to obtain multiple feature representations {h 1 , h 2 ,..., h L} of the nodes at each layer;
[0110] S403. Perform layer combination of features, stack the multiple representations of the nodes at different layers to obtain the final feature representation, and the calculation is shown in Formulas (12) and (13).
[0111]
[0112] where, h i and h jrespectively represent the final feature representation vectors of API i and Mashup j, L represents the depth of Light GCN; α l represents the importance parameter of the embedding of the l-th layer in constructing the final embedding. It can be regarded as a hyperparameter to be manually adjusted or as a model parameter to be automatically optimized. In the experiments of this application, setting it uniformly to 1 / (L + 1) can obtain good performance.
[0113] Furthermore, this application also presents a solution for jointly training the Light GCN model.
[0114] According to the joint training strategy, this application designs different subtasks: BPR loss, regularization loss, and contrastive loss.
[0115] In the BPR loss, it is assumed that the observed interactions (compared to the unobserved interactions) should reflect higher predicted values, and thus can better reflect the correlation relationship between services. The objective function is defined as shown in formula (14).
[0116]
[0117] Among them, j represents the corresponding Mashup, and i pos represents the API that has interacted with j, and i neg represents the API that has not interacted with j, Θ = {(j, i pos , i neg )} represents the pairwise training data. The symbol σ(·) represents the Sigmod function, and λ bpr is the regularization intensity parameter of the BPR loss, which is used to control the trade-off between the model fitting the training data and maintaining the generalization ability.
[0118] The regularization loss is a loss term used to control the model complexity and prevent overfitting. In machine learning and deep learning, the regularization loss is usually used together with the main task loss of the model to balance the relationship between fitting the training data and restricting the model complexity. By penalizing the magnitude or distribution of the model parameters, it is avoided that the model performs too well on the training data and poorly on the unseen data. The calculation of the regularization loss function is shown in formula (15).
[0119]
[0120] Among them, λ reg represents the regularization intensity parameter of the regularization loss, n is the number of parameters of the Light GCN model, and ω i is the i-th parameter of the Light GCN model.
[0121] To train the encoder end-to-end and learn rich node and graph-level representations that are agnostic to downstream tasks, after obtaining the feature representations of the same node from the original graph and the contrast view, the present invention uses a contrastive learning loss to maximize the consistency between them. The representations of the same node under different views are regarded as positive sample pairs, and the representations of other nodes are regarded as negative sample pairs. During training, the consistency between positive sample pairs is maximized, and the consistency between negative sample pairs is minimized. The calculation of the contrastive loss function is shown in Equation (16).
[0122]
[0123] where h and h c respectively represent the feature representation vectors of the node in the original graph and the contrast view, sim(x, y) represents the cosine similarity between two vectors x and y, τ represents the temperature coefficient in contrastive learning, which is an important parameter in contrastive learning and affects the ability of the model to learn features and distinguish positive and negative samples. The temperature coefficient is a tuning parameter in contrastive learning that controls the sensitivity of the model to positive and negative samples. When the temperature coefficient is large, dividing by the temperature coefficient can reduce the distinguishability between positive and negative samples, making the model less sensitive to positive and negative samples. When the temperature coefficient is small, the distinguishability between positive and negative samples is amplified, making the model more sensitive to positive and negative samples. Selecting an appropriate temperature coefficient is crucial for the effect of contrastive learning.
[0124] The present invention uses the Adam optimizer to adjust the model and update the model parameters.
[0125] In summary, the joint training process of the Web API recommendation model is shown in Algorithm 2.
[0126]
[0127]
[0128] S5. Sort the APIs that meet the Mashup requirements based on the inner product between the feature representation vectors of Mashup nodes and API nodes, so as to generate a recommendation list.
[0129] After obtaining the final node features of Mashup and API, they will be used for recommendation score prediction. The present invention uses the inner product to predict the score of the candidate API and Mashup, and the calculation is shown in Equation (17).
[0130]
[0131] where j and i respectively represent Mashup and API, h j and h i respectively represent the feature vectors of j and i.
[0132] To obtain the final recommendation results, it is necessary to sort the matching score lists of Mashup requirements and each Web API in descending order of the matching scores, and select the top K Web APIs to form the final recommendation list, where K is the number of recommendations that can be set arbitrarily.
[0133] To verify the effectiveness of the proposed method, the present invention selects MF (Matrix Factorization), NGCF, GFormer, and AdaGCL as baseline methods, conducts experimental comparisons on publicly available datasets, and demonstrates the experimental results through two widely used evaluation metrics.
[0134] (1) Dataset
[0135] The experiments of the present invention use the most popular online Web API repository, ProgrammableWeb (abbreviated as PW), to evaluate the model performance. PW is a website that collects metadata about Web APIs and the corresponding applications (such as Mashups) that use them. All Web APIs and Mashups are crawled from PW, and the interaction information between Web APIs and Mashups is analyzed. This dataset includes 21,900 APIs, 6,435 Mashups, and 13,340 interactions between Mashups and APIs. For evaluation, the present invention deletes the Mashups that have no interaction with any API and the APIs that have no interaction with any Mashup, leaving 6,298 Mashups and 1,069 APIs. Finally, 70% of the interaction records are used as the training set, and the remaining 30% are used as the test set.
[0136] (2) Baseline methods
[0137] To verify the effectiveness of the method proposed by the present invention, the present invention conducts comparative experiments using four advanced methods in the field of Web API recommendation, including MF, NGCF, GForme, and AdaGCL.
[0138] MF: It is a ranking-oriented recommendation algorithm proposed for the scenario of implicit feedback. This method decomposes a high-dimensional rating matrix into the product of two low-dimensional matrices to capture the latent features between Mashups and APIs. This method explores the direct interactions (first-order connectivity) of nodes on the Mashup-API bipartite graph and makes API recommendations based on interaction features.
[0139] NGCF: It is a neural network-based collaborative filtering algorithm that models the interaction relationship between Mashup and API through a graph neural network. The model uses hidden layers to map the IDs of Mashup and API to obtain initialization vectors, and utilizes the Mashup-API interaction matrix to achieve high-level information interaction and information transmission to obtain the final feature vectors, and conducts API recommendation based on the feature vectors.
[0140] GFormer: It is a recommendation model that combines a graph neural network and a Transformer structure. The main idea of this model is to apply the attention mechanism of Transformer to the graph neural network to capture long-range dependencies in graph-structured data. And it conducts API recommendation by automating the self-supervised enhancement process based on SSL and extracting the interaction patterns of Mashup-API.
[0141] AdaGCL: It is a graph representation learning and recommendation method that combines contrastive learning and graph generation techniques. The model uses VGAE and Gaussian denoising models as contrastive view generators to generate contrastive views for graph contrastive learning, and utilizes GCN to extract features and conduct API recommendation.
[0142] (3) Evaluation Metrics
[0143] In the experiment, the present invention uses the widely used Recall@N and NDCG@N to evaluate the effects of the proposed method and the baseline methods. These metrics will be introduced separately below.
[0144] Recall@N: It represents the proportion of marked items listed in the top N recommendation lists, and is calculated as shown in formula (18).
[0145]
[0146] NDCG@N: It is a standard for measuring the ranking quality of the recommendation list, which considers the graded correlation between positive and negative items within the top N of the ranking list. It is calculated as shown in formula (19).
[0147]
[0148] where S m represents the ideal maximum DCG score that can be achieved for m.
[0149] (4) Experimental Environment
[0150] The experimental hardware environment of the present invention is a server with NVIDIA GeForce RTX3090; the software environment is Python 3.7, and Pytorch is used as the deep learning framework to build a neural network. The embedding size of all models is fixed at 64, and the model depth is set to 5. In the model training stage, the present invention uses the Adam optimizer to optimize the model, where the batch size is fixed at 512 and the learning rate is set to 0.001.
[0151] (5) Experimental results and comparative analysis
[0152] The present invention conducts experimental comparisons between the proposed service recommendation method (Ours) and four baseline methods to verify the effectiveness of the method proposed in the present invention. Table 1 shows the comparative experimental results of the method proposed in the present invention and other baseline methods on the PW dataset.
[0153] Table 1 Comparative experimental results on the PW dataset
[0154]
[0155] It can be seen from Table 1 that the method proposed in the present invention is always superior to the baseline in all cases. More specifically, in the Recall@N metric, the method of the present invention improves the best baseline by 18.14% compared with others; in the NDCG@K metric, the method proposed in the present invention improves the best baseline by 14.28% compared with others.
[0156] Among them, the performance of MF is the worst, and this model only explores the first-order interaction relationship of nodes on the Mashup-API bipartite graph. Compared with NGCL, GFormer and AdaGCL show better performance. These two models are both recommendation models based on information network embedding and contrast learning ideas, and use the designed contrast view generation method and specific graph neural network models to obtain the node embeddings of Mashup and Web API entities. However, compared with the model proposed in the present invention, these two models do not work well because they have deficiencies in data augmentation and feature extraction. GFormer uses the global attention mechanism of graph Transformer for feature extraction, but its network structure is too complex and more noise will be introduced during the feature extraction process. When AdaGCL performs data augmentation and feature extraction, it lacks high-order information interaction. Experiments show that the method of the present invention can perform reasonable data augmentation, can accurately extract the relationship features between Mashup and Web API, and improves the Web API recommendation effect.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Web API recommendation method based on graph diffusion reconstruction and graph contrast learning, characterized in that: The following steps are involved: S1. Obtain historical call information of Mashup and Web API, and construct a Mashup-API call relationship graph based on the historical call information; S2, performing data expansion based on graph diffusion on the Mashup-API call relationship graph, thereby generating a diffusion matrix of the Mashup-API call relationship; S3, reconstructing the diffusion matrix using adaptive graph reconstruction to generate a contrast view; S4. Obtain the Mashup-API call relationship graph and comparison view, use the trained Light GCN as the basic graph encoder to process the Mashup-API call relationship graph and comparison view, and generate the final feature representation vector of the node; S5. Sort the APIs that meet the Mashup requirements based on the inner product between the feature representation vector of the Mashup node and the feature representation vector of the API node, thereby generating a recommendation list.
2. According to claim 1, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: Constructing a Mashup-API call relationship graph based on the historical call information includes: S101, storing each Mashup and its called API list in the form of a dictionary to obtain structured Mashup-API calling information; S102. Based on the structured Mashup-API information, the Mashup and the API are regarded as nodes, and the calling relationship between the Mashup and the API is regarded as an edge, so as to construct a bipartite graph of the calling relationship between the Mashup and the API.
3. According to claim 1, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: The Mashup-API call relationship graph is subjected to data expansion based on graph diffusion, thereby generating a diffusion matrix of the Mashup-API call relationship, including: S201. Represent the Mashup-API call relationship graph as a sparse graph G: G=(V,E), Where V represents the node set, including the Mashup node set M and the API node set N; E represents the edge set; S202, converting the sparse graph G into an adjacency matrix A, where the adjacency matrix A represents the connection relationship between nodes, where A[i][j] is 1 to indicate that there is an edge between node i and node j, otherwise it is 0; S203, adding self-loops to the adjacency matrix A, and calculating a symmetric transfer matrix according to the adjacency matrix after adding the self-loops; S204: Perform random walk on the symmetric transfer matrix to generate a dense diffusion matrix, and perform sparseness on the dense diffusion matrix based on a preset threshold, so as to obtain a final diffusion matrix.
4. According to claim 3, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: A dense diffusion matrix is generated by performing a random walk on the symmetric transfer matrix according to the following formula: S=α(I-(1-α)T sym ) -1 Where α is the probability of random walk, I is the identity matrix, and T sym represents the symmetric transfer matrix, and: Among them, A loop represents the adjacency matrix after adding the self-loop, D loop represents the self-loop matrix, and: D loop =diag(A loop ·1) A loop =I+A Among them, 1 represents a vector of all 1s, and diag represents a diagonal matrix.
5. According to claim 3, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: The diffusion matrix is reconstructed using adaptive graph reconstruction to generate contrastive views, including: S301, using graph convolutional neural network to extract relational features in diffusion matrix; S302, calculating the distance between each feature Mashup i and each API j in the feature space according to the relationship features of the nodes; S303, using the Sigmod function to normalize the node feature space distance, mapping it between 0 and 1, and obtaining the distance factor θ ij ; S304: Filter the distance factor based on the preset threshold value to reconstruct the Mashup-API relationship diagram.
6. According to claim 1, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: Use the trained Light GCN as the basic graph encoder to process the Mashup-API call relationship graph and the comparison view to generate the final feature representation vector of the node, including: S401. For a node in the graph, its representation at the l+1 layer is: in and They represent the feature vector of API i and the feature vector of Mashup j respectively, l+1 represents the number of propagation layers, N j Represents the set of APIs that have interacted with Mashup j, N i Represents the Mashup collection that has interacted with API i, |N j | indicates N j The number of APIs in |N i | indicates N i The number of mashups in the S402, increase the depth of Light GCN, repeat S401, and obtain multiple feature representations of nodes at each layer {h 1 ,h 2 , ..., h L }; S403, perform layer combination of features, superimpose multiple representations of nodes at different layers, and obtain the final feature representation: Among them, h i and h j They represent the final feature representation vectors of API i and Mashup j respectively, L represents the depth of Light GCN; α l A parameter that represents the importance of the embedding at layer l in forming the final embedding.
7. According to claim 1, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: The inner product between the feature representation vector of the Mashup node and the feature representation vector of the API node is calculated according to the following formula: Among them, h i and h j They represent the final feature representation vectors of API i and Mashup j respectively.
8. According to claim 1, a Web API recommendation method based on graph diffusion reconstruction and graph contrast learning is characterized in that: The Light GCN adopts a joint training strategy to train the three subtasks of BPR loss, regularization loss and contrast loss respectively. The objective function of the BPR loss subtask is: Among them, j represents the corresponding Mashup, i pos Indicates the API that has interacted with j, i neg represents the API that has never interacted with j, Θ = {(j, i pos ,i neg )} represents paired training data. The symbol σ(·) represents the Sigmod function, λ bpr is the regularization strength parameter of the BPR loss, which is used to control the trade-off between the model fitting the training data and maintaining the generalization ability; The objective function of the regularization loss subtask is: Among them, λ reg represents the regularization strength parameter of the regularization loss, n is the number of parameters of the Light GCN model, ω i is the i-th parameter of the Light GCN model; The objective function of the contrast loss subtask is: Among them, h and h c They represent the feature representation vectors of the node in the original image and the contrast view respectively, sim(x,y) represents the cosine similarity between two vectors x and y, and τ represents the temperature coefficient in contrastive learning.
Citation Information
Patent Citations
Web API recommendation method and device based on functional semantics and structure interaction
CN116628328A
Recommendation method based on adaptive graph contrast learning
CN117390271A
Web API recommendation method based on correlation and compatibility fusion
CN117743678A
Project recommendation method based on graph contrast learning
CN118133882A
Restful-type web service clustering method fusing service cooperation relationships
WO2022156328A1
Cited By
Traditional Chinese medicine text classification method and system based on mixed graph diffusion model and heterogeneous graph convolutional neural network
CN121030000A