A Web API Recommendation Method Based on Factorization Machines

Through the Sentence-BERT and AFMHN models, the Mashup and API text are processed, and the API popularity and compatibility are learned, which solves the problem of insufficient modeling of feature interaction relationships in the existing technology, and realizes Web API recommendations that are more in line with the needs of developers, improving the accuracy of recommendations and user satisfaction.

CN116107619BActive Publication Date: 2025-07-18HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211534754.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-07-18
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively model implicit features such as API popularity and compatibility, and simple DNN networks cannot learn low-order and high-order feature interactions, and lack attention mechanisms, resulting in Web API recommendations that do not meet developers' needs.

Method used

The Sentence-BERT model is used to process Mashup and API text descriptions, combined with the AFMHN model, and learn feature interactions using linear components, DNN components, CIN components and attention components. By calculating API popularity and compatibility, Top-k Web APIs that meet developers' needs are output.

Benefits of technology

Improves the accuracy of Web API recommendations, reduces developer search costs, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116107619B_ABST
    Figure CN116107619B_ABST
Patent Text Reader

Abstract

The present invention provides a Web API recommendation method based on factorization machines, comprising the following steps: crawling Mashup and Web API metadata from the ProgrammableWeb website to construct a service library dataset; preprocessing the obtained Mashup and API functional text descriptions; inputting the preprocessed text into a Sentence-BERT model to obtain vector representations of sentences, and calculating the similarity between APIs and Mashups through the obtained sentence vectors and a multi-feature extraction component; calculating the popularity of APIs and the compatibility of API combinations according to the interaction records between Mashups and APIs; obtaining a feature matrix of Mashup-API by complete concatenation as the input of the AFMHN model; and outputting the top-k Web APIs with the highest probabilities by the AFMHN model. This method can make the recommended Web APIs meet the requirements of developers for developing Mashups by extracting different features of Web API metadata, using a deep neural network to capture the interaction between any low-order and high-order non-linear features, and using an attention mechanism to capture the different importance between features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining and recommendation, and particularly relates to a Web API recommendation method based on a factorization machine. Background Art

[0002] With the in-depth research in the field of service computing, the value of Internet service resources has been increasingly recognized, and more and more enterprises have published their business functions as remotely accessible APIs (Application Programming Interfaces). In the past decade, driven by the microservices architecture (MSA), more and more enterprises have started to use reusable APIs to create Mashup applications that meet complex business requirements, rather than coding from scratch, which has greatly shortened the development cycle.

[0003] With the continuous prosperity of the API economy, many API sharing libraries have emerged. Among them, ProgrammableWeb is one of the largest online repositories. According to the statistics of the website, as of April 2022, the number of available APIs has reached 24,000, including more than 400 categories. The gradual increase in the number of APIs makes it difficult for developers to select suitable APIs from a large number of candidate APIs when composing services. Therefore, there is an urgent need to develop better recommendation technologies to enable developers to find APIs suitable for Mashup development.

[0004] Recently, the hybrid model based on factorization machines combined with deep neural networks has been proven to be a relatively successful recommendation model, achieving high scalability. However, in practical applications, implicit features such as API popularity and compatibility between APIs are not well modeled, and these features play a very important role in effective recommendation. For long text information such as service function descriptions, simple word models cannot fully reflect content associations. On the other hand, a pure DNN network cannot learn the feature interaction relationships between arbitrary low-order and high-order features. Worse still, it lacks a corresponding attention mechanism and cannot assign different weights to feature interactions to reduce the impact of noise. Summary of the Invention

[0005] Aiming at the above technical problems existing in the prior art, the present invention provides a Web API recommendation method based on a factorization machine, which can enable the recommended Web APIs to meet the requirements of developers for developing Mashups.

[0006] A Web API recommendation method based on a factorization machine includes the following steps:

[0007] (1) Crawl Mashup and Web API metadata from the ProgrammableWeb website to build a service library dataset, including category information of Mashup and APIs, functional description text information, historical interaction records, etc.

[0008] (2) Preprocess the obtained functional text descriptions of Mashup and APIs, including deduplication and normalization of statements, that is, deleting invalid words and restoring abbreviations in English statements; standardization, that is, removing word affixes and unifying statement tenses.

[0009] (3) Input the preprocessed text into the Sentence - BERT model to obtain vector representations of sentences, and calculate the similarity between APIs and Mashup through the obtained sentence vectors and multi - feature extraction components.

[0010] (4) Calculate the popularity of APIs and the compatibility of API combinations according to the interaction records between Mashup and APIs. The definition formula for the popularity of API is as follows:

[0011]

[0012] Where: M is the number of all Mashup, N is the number of all APIs, and p i,j represents the number of calls between Mashup i and API j. If Mashup i calls API j, p i,j = 1; otherwise, p i,j = 0. The definition formula for the compatibility of API j is as follows:

[0013]

[0014] Where: Let G be an API co - call graph containing API nodes. If APIs i and j are co - called by the same Mashup, add an edge between them. Let d(i,j)≥1 represent the shortest distance between i and j. The compatibility of i and j is defined as com(i,j)=e 1-d(i,j) .

[0015] (5) Obtain the feature matrix of Mashup - API by complete concatenation as the input of the AFMHN model. The AFMHN model has four training components, namely, a linear component, a DNN component, a CIN component, and an attention component, as well as two output components, namely, a prediction component and an evaluation component. The formula for the AFMHN model to learn the interaction of complex features is defined as follows:

[0016]

[0017] Among them: The formula can be divided into four parts, namely, the linear regression part for learning the contributions of basic features the attention mechanism part for learning feature interactions with different importance the DNN network part for capturing implicit high-order feature interactions and the CIN network part for capturing explicit high-order feature interactions

[0018] (6) The output of the AFMHN model, that is, the top-k Web APIs with the highest probabilities given according to the Mashup text requirement description of the developer

[0019] Preferably, in the step (5), the linear regression part for learning the contributions of basic features is calculated by the following formula

[0020]

[0021] Among them: is the global bias is the strength of the i-th variable A row v in i represents the embedding vector of feature i, where is a hyperparameter that defines the dimension of factorization

[0022] Preferably, in the step (5), the attention mechanism part for learning feature interactions with different importance is calculated by the following formula

[0023] a i ′ j =h T ReLU((v i ⊙v j )x i x j +b)

[0024]

[0025]

[0026] Among them: are the parameters of the model, t represents the size of the hidden layer of the attention network, called the attention factor. v i represents the embedding vector corresponding to the feature domain, x i represents the feature value, ⊙ represents the element-wise product of two vectors, ReLU represents the activation function of the network. a ij represents the attention score representing the weight of the cross term, p represents the neural weight of the prediction layer

[0027] Preferably, in step (5), the DNN network part for capturing implicit high-order feature interactions is calculated by the following formula

[0028]

[0029] where: the vector h represents the neural weights of the prediction layer, L represents the number of hidden layers, W L , b L and σ L represent the weight matrix, bias vector, and activation function of the L-th layer respectively.

[0030] Preferably, in step (5), the CIN network part for capturing explicit high-order feature interactions is calculated by the following formula

[0031]

[0032]

[0033]

[0034]

[0035] where: is the input feature vector, m is the number of features, and D is the dimension of the embedding vector. represents the output of the k-th layer of the CIN network, where H k represents the number of features in the k-th layer, which can also be understood as the number of neurons. In addition, ° represents the Hadamard product, that is, the multiplication of corresponding dimension elements between vectors. represents the h-th feature vector of the k-th layer, which is added from the D dimension to obtain Let T represent the depth of the network, and a pooling vector with a length of H k of the k-th layer can be obtained By concatenating all the vectors obtained from different layers, we generate the final output of the CIN, and w0 is the regression parameter.

[0036] The beneficial effects of the present invention are as follows:

[0037] The present invention proposes a novel hybrid network factorization machine model for WebAPI recommendation when developing Mashups. The Sentence-BERT model is used to better learn the functional text description features of Mashups and Web APIs; the AFMHN model is proposed to integrate the DNN and CIN networks to learn the explicit and implicit feature interactions between low-order and high-order, and integrate the attention network (i.e., the attention component) to capture the specific importance of different feature interactions; the present invention can make the recommended Web APIs more in line with the development needs of developers for Mashups, thereby reducing the search cost of developers and improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of the system architecture of the Web API recommendation method of the present invention.

[0039] Figure 2 It is a schematic diagram of the neural network structure of the AFMHN model in the Web API recommendation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0041] This embodiment provides a Web API recommendation method based on a factorization machine, including the following steps:

[0042] (1) Crawl Mashup and Web API metadata from the ProgrammableWeb website to construct a service library dataset, including category information, functional description text information, historical interaction records, etc. of Mashups and APIs.

[0043] (2) Preprocess the obtained Mashup and API functional text descriptions, including deduplication, normalization, and standardization of sentences.

[0044] (3) Input the preprocessed text into the Sentence-BERT model to obtain the vector representation of the sentences, and calculate the similarity between APIs and Mashups through the obtained sentence vectors and the multi-feature extraction component.

[0045] (4) According to the interaction records between Mashups and APIs, calculate the popularity of APIs and the compatibility of API combinations. The definition formula of the popularity of APIs is as follows:

[0046]

[0047] Where: M is the number of all Mashups, N is the number of all APIs, pi,j Denotes the number of calls between Mashup i and API j. If Mashup i calls API j, p i,j = 1; otherwise, p i,j = 0. The compatibility formula for API j is defined as follows:

[0048]

[0049] Where: Let G be an API co-call graph containing API nodes. If APIs i and j are co-called by the same Mashup, an edge is added between them. Let d(i, j) ≥ 1 denote the shortest distance between i and j. The compatibility of i and j is defined as com(i, j) = e 1-d(i,j) .

[0050] (5) Obtain the feature matrix of Mashup-API by complete concatenation as the input of the AFMHN model. The AFMHN model has four training components, namely the linear component, the DNN component, the CIN component, and the attention component, as well as two output components, namely the prediction component and the evaluation component. The interaction formula for the AFMHN model to learn complex features is defined as follows:

[0051]

[0052] Where: The formula can be divided into four parts, namely the linear regression part for learning the contribution of basic features The attention mechanism part for learning the interaction of features with different importance The DNN network part for capturing implicit high-order feature interactions And the CIN network part for capturing explicit high-order feature interactions

[0053] (6) The output of the AFMHN model, that is, the Top-k Web APIs with the highest probabilities given according to the Mashup text requirement description of the developer.

[0054] Figure 1The architecture of the Web API recommendation method based on the factorization machine in this embodiment is shown. This framework consists of a data training part and an attention factorization machine based on a hybrid network model (AFMHN), which combines the factorization machine with a deep neural network and an attention mechanism. The data training part mainly includes preprocessing the obtained Mashup and API text descriptions and sentence vector embedding. According to the interaction records between Mashup and API, the popularity of the API and the compatibility of the combined API are calculated, and the obtained Mashup-API feature matrix is used as the input of the AFMHN model. The AFMHN model has four training components, namely a linear component, a DNN component, a CIN component, and an attention component, as well as two output components, namely a prediction component and an evaluation component. Finally, the output of the AFMHN model is to give the top-k WebAPIs with the highest probability according to the Mashup text requirement description of the developer.

[0055] Figure 2 The neural network structure of the proposed AFMHN model is shown. It has four training components, namely a linear component, a DNN component, a CIN component, and an attention component, as well as two output components, namely a prediction component and an evaluation component. Among them, the linear part, the attention mechanism part, the DNN part, and the CIN part are represented by red, green, blue, and purple connections respectively. The prediction component adds up all the outputs of these training components to predict the probability of the Mashup calling the network API.

[0056] The above description of the embodiments is to facilitate the understanding and application of the present invention by those of ordinary skill in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A Web API recommendation method based on factorization machines, comprising the following steps: (1) Crawl Mashup and Web API metadata to construct a service library dataset, where the service library dataset includes category information, functional description text information, and historical interaction records of Mashup and APIs; (2) Preprocess the obtained service library dataset; (3) Input the preprocessed functional description text information into the Sentence-BERT model to obtain the vector representation of sentences, and calculate the similarity between APIs and Mashup through the obtained sentence vectors and the multi-feature extraction component; (4) Calculate the popularity of APIs and the compatibility of API combinations according to the historical interaction records between Mashup and APIs; (5) Obtain the feature matrix of Mashup-API by full concatenation as the input of the AFMHN model. The AFMHN model has four training components, namely a linear component, a DNN component, a CIN component, and an attention component, and two output components, namely a prediction component and an evaluation component. The interaction formula for the AFMHN model to learn complex features is defined as follows: Among them, the formula can be divided into four parts, namely, the linear regression part for learning the contribution of basic features the attention mechanism part for learning feature interactions of different importance the DNN network part for capturing implicit high-order feature interactions and the CIN network part for capturing explicit high-order feature interactions (6) The output of the AFMHN model, that is, the Top-k Web APIs with the highest probability according to the Mashup text requirement description of the developer.

2. The Web API recommendation method based on a factorization machine according to claim 1, wherein In the step (1), the service library dataset includes category information, functional description text information, and historical interaction records of Mashup and APIs.

3. The Web API recommendation method based on a factorization machine according to claim 2, wherein The preprocessing method of the service library dataset: delete invalid words, restore the abbreviations of English sentences, remove word affixes, and unify the tenses of sentences.

4. The Web API recommendation method based on a factorization machine according to claim 1, wherein The calculation method of the popularity of the API is as follows: The definition formula of the popularity of the API is as follows: Where M is the number of all Mashups, N is the number of all APIs, and p i,j represents the number of calls between Mashup i and API j. If Mashup i calls API j, p i,j = 1; otherwise, p i,j = 0.

5. The Web API recommendation method based on a factorization machine according to claim 4, characterized in that The calculation method of the compatibility of the API combination is as follows: The compatibility formula of API j is defined as follows: Among them, let G be an API co-call graph containing API nodes. If APIs i and j are co-called by the same Mashup, then an edge is added between them. Let d(i, j) ≥ 1 denote the shortest distance between i and j. The compatibility of i and j is defined as com(i, j) = e 1-d(i,j) .

6. The Web API recommendation method based on a factorization machine according to claim 1, wherein In the said step (5), the linear regression part for calculating the contribution of the learning basic features is calculated through the following formula Among them, is the global bias, is the strength of the i-th variable, a row v in i represents the embedding vector of feature i, where is a hyperparameter that defines the dimension of factorization.

7. The Web API recommendation method based on a factorization machine according to claim 1, wherein In the step (5), the attention mechanism part for learning feature interactions with different importance is calculated through the following formula a′ ij = h T ReLU((v i ⊙ v j ) x i x j + b) Among them, are the parameters of the model, t represents the size of the hidden layer of the attention network, called the attention factor, v i represents the embedding vector corresponding to the feature domain, x i represents the feature value, ⊙ represents the element-wise product of two vectors, ReLU represents the activation function of the network, a ij represents the attention score representing the cross-term weight, and p represents the neural weight of the prediction layer.

8. The Web API recommendation method based on a factorization machine according to claim 1, characterized in that In the step (5), the DNN network part for capturing implicit high-order feature interactions is calculated through the following formula Among them, the vector h represents the neural weights of the prediction layer, L represents the number of hidden layers, W L , b L and σ L represent the weight matrix, bias vector, and activation function of the L-th layer.

9. The Web API recommendation method based on a factorization machine according to claim 1, wherein In the step (5) described above, the CIN network part for capturing explicit high-order feature interactions is calculated through the following formula where, is the input feature vector, m is the number of features, and D is the dimension of the embedding vector; represents the output of the k-th layer of the CIN network, where H k represents the number of features in the k-th layer, which can also be understood as the number of neurons; in addition, ° represents the Hadamard product, that is, the multiplication of corresponding dimension elements between vectors; represents the h-th feature vector of the k-th layer, and adding them from the D dimension gives Let T denote the depth of the network, and a pooling vector of length H k can be obtained for the k-th layer By concatenating all the vectors obtained from different layers, the final output of the CIN is produced, and w0 is the regression parameter.

Citation Information

Patent Citations

  • Mashup Web API personalized recommendation based on collaborative filtering and link prediction

    CN110851719A

  • Developer-oriented network API recommendation method

    CN114676332A