Cold start service recommendation method based on multi-strategy pre-training model

By building hypergraphs and using hypergraph neural networks for comparison learning, combined with contrast learning of description text, service embedding is optimized to solve the recommendation problem of Web API without historical interaction records, and efficient and accurate cold start service recommendations are achieved.

CN120066556APending Publication Date: 2025-05-30CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510085350.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively recommend Web APIs without historical interaction records, and fails to solve the cold start problem.

Method used

Using a method based on a multi-strategy pre-training model, we use hypergraphs to extract cold start service features, use hypergraph neural network for comparison learning and feature optimization, and combine the comparison learning of description text to optimize service embedding to improve recommendation performance.

Benefits of technology

It realizes effective recommendations for the Web API without historical interaction records, improves the accuracy and efficiency of service recommendations, and solves the cold start problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066556A_ABST
    Figure CN120066556A_ABST
Patent Text Reader

Abstract

The invention discloses a cold start service recommendation method based on a multi-strategy pre-training model. The cold start service recommendation method comprises the following steps: step 1, constructing a hypergraph based on labels, extracting cold start service features, and constructing service embedding; step 2, obtaining enhanced service embedding based on hypergraph contrast learning of singular value decomposition; step 3, obtaining service embedding based on content optimization based on comparative learning of the description text; and 4, carrying out fine adjustment on the cold start recommendation model through a downstream recommendation task. The cold start service recommendation method based on the multi-strategy pre-training model is designed by analyzing and optimizing data features of Web APIs without historical interaction records and taking a hypergraph neural network as a technical support, and aims to solve the problem of recommendation of APIs without any historical interaction information and optimize the recommendation performance of services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of service computing, and particularly relates to a cold-start service recommendation method based on a multi-strategy pre-trained model. Background Art

[0002] With the development of service-oriented computing (SOC) and cloud computing, Web services have become an important carrier for IT resource delivery. SOC is committed to using services as basic building blocks to create fast, low-cost, secure, and reliable applications. Compared with traditional software components, services have autonomy, self-description, reusability, and high portability.

[0003] Web services are usually exposed externally in the form of Web APIs. A Web API is an application programming interface that allows Web applications to access and implement storage services, messaging services, computing services, and other functions. For example, technology giants such as Google and Microsoft are opening up their software and data resources as Web services through Web APIs to attract a wider group of Internet users. In addition, with the continuous booming development of the Web API economy, numerous Web API sharing repositories have emerged in the market, such as the ProgrammableWeb, RapidAPI, and ShowAPI websites. Taking the open cloud API platform RapidAPI as an example, it currently supports more than 40,000 APIs, more than 12,000 API publishers, more than 200,000 monthly active subscribers, millions of registered developers, and more than 5 billion API requests per month.

[0004] However, in the face of the rapid development of the information society and the emergence of a large number of complex requirements, the existing Web API functions are insufficient to cope with increasingly complex business scenarios. To fill this gap, Mashup technology has emerged. The core idea of Mashup is to use Web APIs as reusable components to create new products or solutions. It matches user-customized requirements by integrating multiple services, greatly reducing the development threshold, enabling developers without programming skills to quickly build applications using ready-made APIs, and also helping developers shorten the development cycle.

[0005] In recent years, more and more developers have benefited from the reuse of web services. However, in the face of the rapidly expanding web API market with diverse functions, it has become increasingly difficult to quickly screen out the most suitable APIs for user needs from it. Therefore, implementing an efficient and accurate service recommendation mechanism to match user goals and needs has become a hot topic in current research. The main purpose of a web service recommendation system is to use the available relevant information to assist users in making appropriate choices among various web services with different functions. According to the types of information used by the recommendation system, traditional recommendation algorithms are mainly divided into three types: collaborative filtering methods using call history, algorithms based on service content information, and combinations of the two. At the same time, thanks to the outstanding ability of Graph Neural Networks (GNNs) in relational data modeling, GNN-based models have also demonstrated excellent potential in the field of service recommendation systems. Hypergraphs generalize the concept of edges in traditional graphs, allowing edges to connect more than two nodes, providing an intuitive and natural framework for modeling complex high-order relationships between services.

[0006] However, most existing studies only focus on web APIs with interaction records with specific Mashups, while ignoring the problem of effectively recommending new web APIs without interaction records and have not proposed a solution strategy for this cold start problem. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, the present invention specifically analyzes and optimizes the data characteristics of web APIs without historical interaction records, and uses hypergraph neural networks as the technical support to provide a cold start service recommendation method based on a multi-strategy pre-training model, aiming to solve the recommendation problem of APIs without any historical interaction information and optimize the recommendation performance of services.

[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0009] A cold start service recommendation method based on a multi-strategy pre-training model, comprising the following steps:

[0010] The first step is to construct a hypergraph based on tags, extract cold start service features, and construct service embeddings;

[0011] The second step is hypergraph contrast learning based on singular value decomposition to obtain enhanced service embeddings;

[0012] The third step is contrast learning based on descriptive text to obtain content-optimized service embeddings;

[0013] The fourth step is to fine-tune the cold start recommendation model through downstream recommendation tasks.

[0014] Further, the process of the first step is as follows:

[0015] Step (1.1) Crawl API service and Mashup service information from relevant API websites;

[0016] Step (1.2) Analyze the crawled service set, select meta-information that is beneficial to improving the accuracy of service recommendation and is rich in content, and form a service dataset, including a warm-start dataset and a cold-start dataset. The cold-start dataset further deletes the meta-information of the call relationship between Mashup and API on the basis of the warm-start dataset;

[0017] Among them, the retained meta-information includes Mashup, API, the call relationship between Mashup and API, the tags carried by Mashup and API, and the service description text; among them, the tags represent some functions or attributes of Mashup and API, and the tag information also includes the collaboration information between Mashup and API;

[0018] Step (1.3) Since the cold-start service does not contain call information, a subgraph structure in which services have pairwise same-tag relationships and form a triadic closure is defined as a motif for extracting many-to-many relationships between services based on tags;

[0019] Step (1.4) Construct a cold-start service hypergraph G according to the motif proposed in step (1.3); sin ;

[0020] Step (1.5) According to the hypergraph G constructed in step (1.4); sin , initialize a cold-start encoder f using a hypergraph convolutional network; cold , which is used to construct embeddings for cold-start service nodes in the setting of meta-learning and improve the convolutional operation in subsequent steps;

[0021] Step (1.6) Instantiate the meta-aggregator g as a self-attention encoder, the core of which is to use the self-attention mechanism (Self-Attention) to calculate the attention inside neighbor nodes;

[0022] Among them, the meta-aggregator is used to process cold-start neighbor node data. A cold-start neighbor node means that there is a cold-start node in the neighborhood of the target node in the hypergraph. Assume the target node is v; tar , its neighbor node is i. When performing hypergraph convolution, if i is a cold-start node, its embedding will indirectly affect the embedding of v; tar 's embedding;

[0023] Step (1.7) g accepts the initial embeddings of k first-order neighbors of node v; as input, where N(v) represents the neighborhood of v;

[0024] In step (1.8), for each first-order neighbor i of v, calculate the attention scores of all first-order neighbors for i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h of each neighbor i i , that is

[0025] In step (1.9), average the embeddings of all neighbors to obtain the meta-embedding of the node that is

[0026] In step (1.10), use LightGCN as a warm-start encoder, and based on the observed rich interactions on the warm-start dataset constructed in step (1.2), that is, the call relationship between Mashup and API, learn the true embeddings of the nodes for further adjustment Among them, LightGCN is an efficient graph convolutional recommendation model. It obtains the embeddings of Mashup and API through graph convolution and matches API and Mashup according to the distance;

[0027] In step (1.11), use cosine similarity to calculate the difference between and use this as the loss function of the meta-aggregator. The formula is as follows:

[0028]

[0029] Among them, cos(·,·) represents cosine similarity calculation, ||·|| represents finding the modulus length, V represents the set of hypergraph nodes, and the argmax function represents finding the meta-aggregator parameter Q that makes the formula maximum, called Q * , and the meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node;

[0030] In step (1.12), adopt an adaptive sampling strategy to improve the aggregation operation on the hypergraph, and combine it with the meta-aggregator to perform improved hypergraph convolution. The total number of convolution layers is L. Among them, the adaptive sampling strategy is a sampling strategy that dynamically adjusts according to the Mashup-API pairs involved in the matching task. In the matching task of Mashup m and API a, the target node v tar when performing feature aggregation, this strategy not only considers the distance between the neighborhood nodes and v tar but also considers the distance between the neighborhood nodes and m and a;

[0031] Step (1.13) uses the warm-start encoder LightGCN trained in step (1.10) to obtain the true embedding of the nodes

[0032] Step (1.14) defines the loss function of the cold-start encoder using cosine similarity where θ 1 represents the parameters of the cold-start encoder;

[0033] Step (1.15) trains the cold-start encoder to obtain the pre-trained service embedding h p1 .

[0034] Preferably, the process of step (1.4) is as follows:

[0035] Step (1.4.1) defines the matrix which is a 0-1 matrix storing the service-label carrying relationship, where S represents the total number of services, T represents the total number of labels, and for each pair (s, t), r s,t = 1 indicates that service s carries label t, and r s,t = 0 indicates that service s does not carry label t;

[0036] Step (1.4.2) calculates according to the service-label carrying relationship matrix R s to obtain where represents the number of services carrying the same label, represents the transpose of the matrix R s ;

[0037] Step (1.4.3) takes the upper right part of the diagonal of the matrix and sets all numbers greater than 0 to 1 to construct the upper triangular matrix R s * , where the elements in R s * represent whether the corresponding services carry the same label;

[0038] Step (1.4.4) defines S sin as the adjacency matrix of the cold-start service hypergraph G sin ;

[0039] where the adjacency matrix S sin of the hypergraph G sin is calculated as S sin = (R s * ·R s * ) ⊙ R s * , and ⊙ represents the Hadamard product of the left and right matrices.

[0040] More preferably, the process of step (1.12) is as follows:

[0041] Step (1.12.1) Calculate the enhanced embedding of node v Where represents the initial embedding of node v, and || is the concatenation operation;

[0042] Step (1.12.2) Calculate the sampling probability of node v for adjacent node i

[0043] In the matching task of Mashup m and API a, the sampling probability of each node v in the hypergraph for each adjacent node i is defined as follows:

[0044]

[0045] Where exp(·) represents the exponential function of e, φ(·) represents a non-linear transformation, σ(·) is a non-linear activation function, and sim(·,·) uses cosine similarity to calculate the similarity between two vertices; represents the hypergraph G sin Adjacency matrix S sin The element value of the row and column where nodes v and i are located in Similarly; represents the enhanced embeddings of the API service nodes and Mashup service nodes involved in the matching task;

[0046] Step (1.12.3) Aggregate the neighborhood of node v based on the current sampling probability to obtain the neighborhood aggregation embedding

[0047] Where

[0048] Step (1.12.4) Calculate the embedding obtained by convolution Complete the convolution of the first layer;

[0049] Where represents the parameter matrix, l = 1;

[0050] Step (1.12.5) Use the obtained in step (1.12.4) to replace the in step (1.12.1) to update and repeat steps (1.12.2 - 1.12.4) to update and Complete the convolution of the second layer;

[0051] Step (1.12.6) repeats Step (1.12.5) to complete the first L - 1 convolutional layers;

[0052] Step (1.12.7) Since the purpose of hypergraph convolution is to use neighboring nodes to generate the embedding of cold - start nodes, in the last layer, only is used to represent the target embedding to obtain

[0053] Furthermore, the process of the second step is as follows:

[0054] Step (2.1) Perform singular value decomposition (SVD) on the adjacency matrix S sin obtained in Step (1.4) to get S sin = UDV T , where U and V are each an S×S orthogonal matrix, V T represents the transpose of matrix V, and D is a diagonal matrix storing the singular values of S sin ;

[0055] Step (2.2) Truncate the list of singular values, retain the top q singular values, compress the adjacency matrix to an approximation S sin of S sin’ , and obtain the enhanced hypergraph G sin of the original hypergraph G sin‘ ; The reconstructed matrix S sin’ is a low - rank approximation of the adjacency matrix S sin . While highlighting the key interactions in the service hypergraph network to emphasize the core architecture of the hypergraph, it preserves the global cooperation signals therein by considering all connection relationships in the hypergraph. The formula is as follows:

[0056]

[0057] where, and respectively contain the first q columns of U and V, is a diagonal matrix composed of the top q singular values, represents the transpose of matrix V q ;

[0058] Step (2.3) Perform the same convolution as in Step (1.12) on the newly generated adjacency matrix S sin’ to obtain the embedding of node v in the l - th layer convolution on the enhanced hypergraph

[0059] Step (2.4) Conduct contrastive learning on the enhanced hypergraph embedding and the original hypergraph embedding. The loss of contrastive learning is defined as follows:

[0060]

[0061] Among them, sim(·, ·) and τ represent the cosine similarity and the temperature hyperparameter respectively, represents the embedding of node v obtained by the l-th layer convolution on the original hypergraph, represents the embedding of node v obtained by the l-th layer convolution on the enhanced hypergraph, log represents the logarithmic operation, L represents the number of convolution layers, and θ 2 represents the parameter of this contrastive learning task;

[0062] In step (2.5), after the contrastive learning framework is trained, the model discards other parts and only retains the encoder on the trained enhanced hypergraph. When there are new cold start nodes entering, the encoder performs convolution to obtain the enhanced service embedding

[0063] The process of the third step is as follows:

[0064] In step (3.1), use the BERT model as a feature extractor to extract its embedding vector from the text information j, denoted as x j , where BERT is a pre-trained language representation model proposed by Google, which is used to efficiently represent highly unstructured text data as vectors;

[0065] In step (3.2), use the feature encoder to obtain the text embedding of node v where the feature encoder is a multi-layer perceptron (MLP). A multi-layer perceptron is a common deep neural network that includes an input layer, an output layer, and hidden layers, and is fully connected between layers:

[0066]

[0067] Among them, W 2 , W 3 and b 1 , b 2 represent the trainable matrix and bias vector of the encoder respectively, and f(·) represents the activation function;

[0068] In step (3.3), on the warm start dataset, adopt LightGCN and perform joint training with the feature extractor and encoder constructed in steps (3.1 - 3.2) under the same supervision to obtain the collaborative embedding of node v

[0069] In step (3.4), perform contrastive learning between the collaborative embedding and the text embedding, and the loss is defined as follows:

[0070]

[0071] Among them, <·,·> represents the vector inner product, N + (v) represents the positive sample set of node v, N + - (v) represents the negative sample set of node v - represents the collaborative embedding of the positive sample v of node v + represents the collaborative embedding of the negative sample v of node v, ln represents the natural logarithm operation, θ - 3 is the parameter for contrastive learning;

[0072] In step (3.5), after the contrastive learning framework is trained, the model will discard other parts and only retain the parameters of the trained text feature encoder. When a new cold start node enters, the encoder is executed to obtain the content-optimized service embedding

[0073] The process of the fourth step is as follows:

[0074] In step (4.1), first use the pre-trained hypergraph encoder and text feature encoder to generate the corresponding service embeddings

[0075] In step (4.2), connect the generated embeddings and convert them into the final service embedding:

[0076]

[0077] where, W 5 is a trainable parameter matrix, || represents the concatenation operation;

[0078] In step (4.3), calculate the correlation score using the inner product of the final embeddings of Mashup m and API a, that is:

[0079] In step (4.4), define the BPR loss to optimize W 5 , and fine-tune the parameters {θ 1 ,θ 2 ,θ 3} of each module:

[0080] L BPR = ∑-lnσ(y(m,a + ) - y(m,a - ))

[0081] where, a + ​​​​​An API that has an interaction history with m, a - An API that has no interaction history with m.

[0082] The beneficial effects of the present invention are mainly manifested in:

[0083] (1) By leveraging descriptive information and label information, and combining the analysis of service objects, the available information in the cold-start service recommendation process is enriched.

[0084] (2) Optimize the convolution process of the cold-start service hypergraph through a meta-aggregator and an adaptive sampling strategy.

[0085] (3) Capture the internal correlation between Mashup and API based on a singular value decomposition-based enhancement strategy while retaining useful structural information.

[0086] (4) Optimize the descriptive text embedding through joint training and contrastive learning of collaborative information and descriptive text. Description of the Drawings

[0087] Figure 1 It is a schematic diagram of the architecture of the multi-strategy pre-training model in the present invention.

[0088] Figure 2 It is a schematic diagram of the hypergraph contrastive learning framework based on singular value decomposition in the present invention.

[0089] Figure 3 It is a schematic diagram of the contrastive learning framework based on descriptive text in the present invention.

[0090] Figure 4 It is a schematic diagram of the results of the ablation experiment in the embodiment.

[0091] Figure 5 It is a schematic diagram of the results of the hyperparameter experiment in the embodiment, where (a) represents temperature, (b) represents the number of singular values, and (c) represents the number of hypergraph convolution layers. Detailed Embodiment

[0092] The present invention will be further described below with reference to the drawings.

[0093] Referring to Figures 1 to 5 , a cold-start service recommendation method based on a multi-strategy pre-training model (Cold-Start Service Recommendation Based on a Multi-Strategy Pre-Training Model, MSPT) includes the following steps:

[0094] The first step is to construct a hypergraph based on labels, extract cold-start service features, and construct service embeddings, as follows:

[0095] Step (1.1): Crawl API services and Mashup service information from relevant API websites;

[0096] Step (1.2): Analyze the crawled service set, select meta-information that is beneficial to improving the accuracy of service recommendation and is rich in content, and form a service data set, including a warm start data set and a cold start data set. The cold start data set further deletes the meta-information of the call relationship between Mashup and API on the basis of the warm start data set;

[0097] Among them, the retained meta-information includes Mashup, API, the call relationship between Mashup and API, the tags carried by Mashup and API, and the service description text; among them, the tags represent some functions or attributes of Mashup and API, and the tag information also includes the collaborative information between Mashup and API;

[0098] Step (1.3): Since the cold start service does not contain call information, the subgraph structure in which the services have the same tag relationship in pairs and form a triadic closure is defined as a motif, which is used to extract the many-to-many relationship between services based on tags;

[0099] Among them, the motif represents the connection pattern that recurs in the network and the quantity is significantly higher than that of other complex networks. This connection pattern often holds the key information of the network; the triadic closure is a concept in social network theory, referring to a property of a triple composed of three nodes A, B, and C, that is, there are strong connections between A and B, A and C, and B and C;

[0100] Step (1.4): According to the motif proposed in step (1.3), perform steps (1.4.1 - 1.4.4) to construct the cold start service hypergraph G sin , and the process is as follows:

[0101] Step (1.4.1): Define the matrix is a 0-1 matrix storing the relationship between services and tags carried. Among them, S represents the total number of services, T represents the total number of tags. For each pair (s, t), r s,t = 1 indicates that service s carries tag t, and r s,t = 0 indicates that service s does not carry tag t;

[0102] Step (1.4.2): Calculate according to the matrix of the relationship between services and tags carried R s to obtain Among them represents the number of services carrying the same tag, represents the transpose of the matrix R s ;

[0103] Step (1.4.3): Take the matrix Set all numbers greater than 0 to 1 in the upper right of the diagonal to construct the upper triangular matrix R s * , where R s * The elements in represent whether the corresponding position services carry the same label;

[0104] Step (1.4.4) defines S sin as the adjacency matrix of the cold start service hypergraph G sin ;

[0105] where, the adjacency matrix S sin of the hypergraph G sin is calculated as S sin =(R s * ·R s * )⊙R s * , ⊙ represents taking the Hadamard product of the matrices on the left and right sides;

[0106] Step (1.5) According to the hypergraph G constructed in step (1.4) sin , initialize a cold start encoder f cold using a hypergraph convolutional network, which is used to construct embeddings for cold start service nodes in the setting of meta-learning and improve the convolutional operation in subsequent steps. The hypergraph convolutional network can be instantiated using different models, such as HyperGCN, FastHyperGCN;

[0107] Step (1.6) Instantiate the meta-aggregator g as a self-attention encoder, the core of which is to use the self-attention mechanism (Self-Attention) to calculate the attention inside the neighbor nodes;

[0108] where, the meta-aggregator is used to process cold start neighbor node data. Cold start neighbor nodes mean that there are cold start nodes in the neighborhood of the target node in the hypergraph. Assume the target node is v tar , and its neighbor node is i. When performing hypergraph convolution, if i is a cold start node, its embedding will indirectly affect the embedding of v tar ;

[0109] Step (1.7) g takes the initial embeddings of k first-order neighbors of node v as input, where N(v) represents the neighborhood of v;

[0110] Step (1.8) For each first-order neighbor i of v, calculate the attention scores of all first-order neighbors for i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h i of each neighbor i, that is

[0111] Step (1.9) averages the embeddings of all neighbors to obtain the meta-embedding of the node That is

[0112] Step (1.10) uses LightGCN as a warm-start encoder to learn the true embeddings of nodes based on the observed rich interactions, i.e., the call relationships between Mashups and APIs, on the warm-start dataset constructed in step (1.2) For further adjustment Among them, LightGCN is an efficient graph convolutional recommendation model. It obtains the embeddings of Mashups and APIs through graph convolution and matches APIs and Mashups according to the distance

[0113] Step (1.11) calculates using cosine similarity Between And uses this as the loss function of the meta-aggregator. The formula is as follows

[0114]

[0115] Among them, cos(·,·) represents cosine similarity calculation, ||·|| represents finding the modulus length, V represents the set of hypergraph nodes, and the argmax function represents finding the meta-aggregator parameter Q that makes the formula The largest, called Q * , the meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node

[0116] In step (1.12) of the conventional graph convolutional network, the feature aggregation operation cannot reflect the different contributions of different neighbors to the target node, and it also lacks the flexibility to adjust its own convolution strategy for different matching tasks. The adaptive sampling strategy is a sampling strategy that dynamically adjusts according to the Mashup-API involved in the matching task. In the matching task of Mashup m and API a, the target node v tar When performing feature aggregation, this strategy not only considers the distance between the neighborhood nodes and v tar But also considers the distance between the neighborhood nodes and m and a; Therefore, the adaptive sampling strategy is used to improve the aggregation operation on the hypergraph and combine it with the meta-aggregator to perform steps (1.12.1 - 1.12.7) to execute the improved hypergraph convolution, where the total number of convolution layers is L

[0117] Step (1.12.1) calculates the enhanced embedding of node v Among them Represents the initial embedding of node v

[0118] Step (1.12.2) Calculate the sampling probability of node v for adjacent node i

[0119] Define that in the matching task of Mashup m and API a, the sampling probability of each node v in the hypergraph for each adjacent node i is as follows:

[0120]

[0121] where exp(·) represents the exponential function of e, φ(·) represents a non - linear transformation, σ(·) is a non - linear activation function, and sim(·,·) uses cosine similarity to calculate the similarity between two vertices; Denote the hypergraph G sin Adjacency matrix S sin The element value of the row and column where nodes v and i are located in Similarly; Denote the enhanced embeddings of the API service nodes and Mashup service nodes involved in the matching task;

[0122] Step (1.12.3) Aggregate the neighborhood of node v based on the current sampling probability to obtain the neighborhood aggregation embedding

[0123] where,

[0124] Step (1.12.4) Calculate the embedding obtained by convolution Complete the convolution of the first layer;

[0125] where, Denote the parameter matrix, l = 1;

[0126] Step (1.12.5) Use the obtained in step (1.12.4) to replace in step (1.12.1) to update and repeat steps (1.12.2 - 1.12.4) to update successively Calculate and Complete the convolution of the second layer;

[0127] Step (1.12.6) Repeat step (1.12.5) to complete the convolution of the first L - 1 layers;

[0128] Step (1.12.7) Since the purpose of hypergraph convolution is to use neighborhood nodes to generate the embedding of cold - start nodes, so, the last layer only uses to represent the obtained target embedding

[0129] In step (1.13), the pre-trained warm-start encoder LightGCN is used to obtain the true embeddings of the nodes

[0130] In step (1.14), the loss function of the cold-start encoder is defined using cosine similarity where θ 1 represents the parameters of the cold-start encoder;

[0131] In step (1.15), the cold-start encoder is trained to obtain the pre-trained service embeddings This step aims to maximize the loss function and updates the neural network parameters using gradient descent and backpropagation techniques. After training, the encoder performs convolution to obtain the pre-trained service embeddings of the cold-start nodes.

[0132] Second, hypergraph contrastive learning based on singular value decomposition is used to obtain enhanced service embeddings, and the process is as follows:

[0133] In step (2.1), the adjacency matrix S obtained in step (1.4) sin is subjected to singular value decomposition (Singular Value Decomposition, SVD), as Figure 2 shown, to obtain S sin = UDV T , where U and V are respectively an S×S orthogonal matrix, and V T represents the transpose of matrix V, and D is a diagonal matrix storing the singular values of S sin ;

[0134] In step (2.2), the singular value list is truncated, and the first q singular values are retained to compress the adjacency matrix into an approximate S sin of S sin’ , and the original hypergraph G sin is obtained to obtain the enhanced hypergraph G sin‘ of G sin’ . Larger singular values are usually associated with the principal components of the matrix; the reconstructed matrix S sin is a low-rank approximation of the adjacency matrix S

[0135]

[0136] where and respectively contain the first q columns of U and V, is a diagonal matrix composed of the first q singular values, represents the matrix V q transpose;

[0137] Step (2.3) performs the same convolution on the newly generated adjacency matrix S sin’ to obtain the embedding of node v in the l-th layer convolution on the enhanced hypergraph

[0138] Step (2.4) performs contrastive learning on the enhanced hypergraph embedding and the original hypergraph embedding. The overall framework is as Figure 2 shown, and the loss of contrastive learning is defined as follows:

[0139]

[0140] where sim(·, ·) and τ represent the cosine similarity and temperature hyperparameter respectively, represents the embedding of node v obtained by the l-th layer convolution on the original hypergraph, represents the embedding of node v obtained by the l-th layer convolution on the enhanced hypergraph, log represents the logarithmic operation, L represents the number of convolution layers, and θ 2 represents the parameter of this contrastive learning task;

[0141] Step (2.5) After the contrastive learning framework is trained, the model will discard other parts and only retain the encoder on the trained enhanced hypergraph. When there are new cold start nodes entering, the encoder performs convolution to obtain the enhanced service embedding

[0142] Thirdly, based on the contrastive learning of the description text, the service embedding optimized based on content is obtained, and the process is as follows:

[0143] Step (3.1) Use the BERT model as a feature extractor to extract its embedding vector from the text information j, denoted as x j , where BERT is a pre-trained language representation model proposed by Google, which is used to efficiently represent highly unstructured text data as vectors;

[0144] Step (3.2) Use the feature encoder to obtain the text embedding of node v where the feature encoder is a multi-layer perceptron (MLP). The multi-layer perceptron is a common deep neural network, which includes an input layer, an output layer and a hidden layer, and is fully connected between layers:

[0145]

[0146] where, W 2 , W3 and b 1 , b 2 respectively represent the trainable matrix and bias vector of the encoder, and f(·) represents the activation function;

[0147] Step (3.3) On the warm-start dataset, adopt LightGCN, and jointly train it with the feature extractor and encoder constructed in steps (3.1 - 3.2) under the same supervision to obtain the collaborative embedding of node v

[0148] Step (3.4) Conduct contrastive learning between the collaborative embedding and the text embedding. The framework is as Figure 3 shown. The contrastive learning maximizes the mutual information between the two, which essentially captures the triple correlation between the collaborative information of Mashup, the collaborative information of API, and the text description information. Although there is no collaborative information when applying the model, the text module will memorize the triple correlation captured by the contrastive learning during the training phase and correct the ambiguous content embedding according to the memorized correlation during the application phase. The loss is defined as follows:

[0149]

[0150] Among them, <·,·> represents the vector inner product, and N + (v) represents the positive sample set of node v + , and N - (v) represents the negative sample set of node v - , represents the collaborative embedding of the positive sample of node v + , represents the collaborative embedding of the negative sample of node v - , ln represents the natural logarithm operation, and θ 3 is the parameter of the contrastive learning;

[0151] Step (3.5) After the contrastive learning framework is trained, the model will discard other parts and only retain the parameters of the trained text feature encoder. When a new cold-start node enters, the encoder is executed to obtain the content-optimized service embedding

[0152] Fourth step, fine-tune the cold-start recommendation model through the downstream recommendation task. The process is as follows:

[0153] Step (4.1) First, use the pre-trained hypergraph encoder and text feature encoder that have been trained to generate the corresponding service embeddings

[0154] Step (4.2) Concatenate the generated embeddings and convert them into the final service embedding:

[0155]

[0156] Among them, W 5 is a trainable parameter matrix, and || represents the concatenation operation;

[0157] Step (4.3) calculates the correlation score using the inner product of the final embeddings of Mashup m and API a, that is:

[0158] Step (4.4) defines the BPR loss to optimize W 5 , and fine-tunes the parameters {θ 1 , θ 2 , θ 3} of each module:

[0159] L BPR = ∑ -lnσ(y(m, a + )) - y(m, a - ))

[0160] Among them, a + represents the API with an interaction history with m, and a - represents the API without an interaction history with m.

[0161] This embodiment takes the actual effect of the specific service data analysis invention as an example, and the implementation plan is as follows:

[0162] 1) Experimental dataset

[0163] The data crawled from the Programmable Web website from 2019 to 2020 is selected as the experimental dataset, including 6,217 Mashup services, 11,930 API services, and related meta-information.

[0164] 2) Evaluation metrics

[0165] The NDCG metric is a normalized representation of the Discounted Cumulative Gain (DCG). This metric takes into account the recommendation order, and the normalized discounted cumulative gain of the recommended items ranked higher is higher. The specific representation is as follows:

[0166]

[0167]

[0168] Among them, K represents the number of recommended APIs, n represents the ranking of the service in the recommendation list, and IDCG represents the DCG in the ideal case.

[0169] ​HR is a commonly used Top-N evaluation metric, especially in cases where it is independent of the recommendation order. The higher the value of HR, the better the recommendation effect. Its calculation method is as follows:

[0170]

[0171] Among them, hit(i) indicates whether the API is hit in the Mashup's recommendation list, and N represents the actual number of APIs called by the Mashup.

[0172] 3) Experimental results

[0173] The ablation experiment results of the MSPT model in this embodiment are as Figure 4 shown, where MSPT-1 represents the MSPT model with the hypergraph contrast learning module in the second step removed, MSPT-2 represents the MSPT model with the content-optimized service embedding module in the third step removed, MSPT-3 represents the MSPT model with the meta-aggregator proposed in the first step removed, MSPT-4 represents the MSPT model with the adaptive sampling module proposed in the first step removed, and MSPT-5 represents the MSPT model with the collaborative module in the third step removed.

[0174] Analysis Figure 4 Analyzing the data, compared with the MSPT model, removing the hypergraph contrast learning module in the MSPT-1 model causes the model to perform worse in the recommendation task. This is because contrast learning can help the model learn more discriminative feature representations, which is beneficial to the recommendation task. After removing the content-optimized service embedding module in the MSPT-2 model, since the description text is ignored in this variant, the accuracy of the recommendation results drops significantly, indicating the importance of the description text for the cold-start recommendation model. When removing the meta-aggregator, the recommendation effect also decreases slightly. This is because in the MSPT-3 model, the cold-start neighbors of the target node cannot be processed in advance, resulting in inaccurate information being introduced during feature aggregation, which demonstrates the effectiveness of the meta-aggregator. When removing the adaptive sampling module, since the nodes can only treat all neighboring nodes equally when aggregating neighborhood information and cannot reflect the different contributions of different neighbors, the recommendation effect of the MSPT-4 model decreases. Removing the collaborative module in the content-optimized service embedding module in the MSPT-5 model also weakens the model's effect, which demonstrates the advantage of the joint training of the text module and the collaborative module. This is because the two modules are trained under the same supervision, resulting in positive transfer of the collaborative signal, and the collaborative embedding will be adjusted according to the ranking supervision signal, enabling it to better adapt to the recommendation task. Generally speaking, when appropriate parameters are selected, using the MSPT model proposed in the present invention can improve the discrimination between services while ensuring the accuracy of service recommendations, which is an effective solution.

[0175] Figure 5 They are the experimental results of the hyperparameters of the MSPT model, mainly studying the number of hypergraph convolutional layers L, the temperature τ of hypergraph contrastive learning, and the number of singular values q retained.

[0176] From Figure 5 As can be seen from (a) in it, when τ takes different values between 0.1 and 10, the performance of the model is relatively stable, and the effect of the model reaches the best when the value is 0.3. The choice of q determines the number of singular values retained by SVD in the model. Figure 5 The experimental results in (b) in it show that a good result can be obtained without a too large value of q. Specifically, when q = 5, it is already sufficient to retain the important structure in the hypergraph and make the experimental results reach the best. As Figure 5 shown by the experimental results in (c) in it, when L increases from 1 to 4, the performance first increases and then decreases, and the experimental effect is the best when it takes 3.

[0177] The content described in the embodiments of this specification is only a list of the implementation forms of the inventive concept and is only for illustrative purposes. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that those of ordinary skill in the art can think of according to the inventive concept of the present invention.

Claims

1. A cold start service recommendation method based on a multi-strategy pre-training model, characterized in that: The method comprises the following steps: The first step is to build a hypergraph based on labels, extract cold start service features, and build service embeddings; The second step is to obtain enhanced service embeddings through hypergraph contrastive learning based on singular value decomposition; The third step is to obtain content-optimized service embedding based on comparative learning of description texts; The fourth step is to fine-tune the cold start recommendation model through downstream recommendation tasks.

2. The cold start service recommendation method based on a multi-strategy pre-training model according to claim 1, characterized in that: The process of the first step is as follows: Step (1.1) crawl API service and Mashup service information from relevant API websites; Step (1.2) analyzes the crawled service set, selects meta-information that is conducive to improving the accuracy of service recommendations and is rich in content, and forms a service data set, including a hot start data set and a cold start data set. The cold start data set further deletes the meta-information of the call relationship between Mashup and API on the basis of the hot start data set; The retained meta-information includes Mashup, API, calling relationship between Mashup and API, labels and service description text carried by Mashup and API; labels represent some functions or attributes of Mashup and API, and label information also includes collaborative information between Mashup and API; Step (1.3) Since the cold start service does not contain call information, the subgraph structure that has the same label relationship between services and forms a ternary closure is defined as a motif to extract the many-to-many relationship between services based on labels; Step (1.4) constructs the cold start service hypergraph G according to the model proposed in step (1.3) sin ; Step (1.5) The hypergraph G constructed according to step (1.4) sin , initialize a cold start encoder f using a hypergraph convolutional network cold , used to build embeddings for cold-start serving nodes in a meta-learning setting and improve convolution operations in subsequent steps; Step (1.6) instantiates the meta-aggregator g as a self-attention encoder, the core of which is to use the self-attention mechanism to calculate the attention within the neighboring nodes; Among them, the meta-aggregator is used to process the cold start neighbor node data. The cold start neighbor node indicates that there is a cold start node in the neighborhood of the target node in the hypergraph. Assume that the target node is v tar , whose neighbor node is i. When performing hypergraph convolution, if i is a cold start node, its embedding will indirectly affect v tar Embedding Step (1.7) g accepts the initial embeddings of the k first-order neighbors of node v As input, where N(v) represents the neighborhood of v; Step (1.8) For each first-order neighbor i of v, calculate the attention scores of all first-order neighbors to i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h of each neighbor i i ,Right now Step (1.9) averages the embeddings of all neighbors to obtain the meta-embedding of the node Right now Step (1.10) uses LightGCN as the warm-start encoder to learn the true embedding of nodes based on the observed rich interactions, i.e., the call relationship between Mashup and API, on the warm-start dataset constructed in step (1.2). For further adjustment Among them, LightGCN is an efficient graph convolution recommendation model, which obtains the embedding of Mashup and API through graph convolution, and matches API and Mashup based on distance; Step (1.11) uses cosine similarity calculation and The difference between them is used as the loss function of the meta-aggregator, and the formula is as follows: Among them, cos(·,·) represents the cosine similarity calculation, ||·|| represents the modulus length, V represents the hypergraph node set, and the argmax function represents finding the formula The largest meta-aggregator parameter Q, called Q * ,The meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node; Step (1.12) uses an adaptive sampling strategy to improve the aggregation operation on the hypergraph, and combines it with the meta-aggregator to perform an improved hypergraph convolution, with a total number of convolution layers of L. The adaptive sampling strategy is a sampling strategy that is dynamically adjusted according to the Mashup-API pair involved in the matching task. In the matching task of Mashup m and API a, the target node v tar When performing feature aggregation, this strategy not only considers the neighborhood nodes and v tar The distance between the neighboring nodes and m and a will also be considered; Step (1.13) uses the warm-start encoder LightGCN trained in step (1.10) to obtain the true embedding of the node Step (1.14) defines the loss function of the cold start encoder using cosine similarity Where θ1 represents the parameter of the cold start encoder; Step (1.15) Train the cold start encoder to obtain the pre-trained serving embedding 3. The cold start service recommendation method based on a multi-strategy pre-training model as claimed in claim 2, characterized in that: The process of step (1.4) is: Step (1.4.1) defines the matrix is a 0-1 matrix of storage services and label carrying relationships, where S represents the total number of services and T represents the total number of labels. For each pair (s, t), r s,t =1 indicates that service s carries label t, r s,t =0 means that service s does not carry label t; Step (1.4.2) is based on the service and label carrying relationship matrix R s Calculated in Indicates the number of services that carry the same label. Represents the matrix R s The transpose of Step (1.4.3) Take the matrix Go to the upper right corner of the diagonal and set all numbers greater than 0 to 1 to construct the upper triangular matrix R s * , where R s * The middle element indicates whether the corresponding location services carry the same label; Step (1.4.4) defines S sin Supergraph G for cold start service sin The adjacency matrix of Among them, the hypergraph G sin The adjacency matrix S sin The calculation formula is S sin =(R s * ·R s * )⊙R s * , ⊙ represents the Hadamard product of the left and right matrices.

4. The cold start service recommendation method based on a multi-strategy pre-training model as claimed in claim 2, characterized in that: The process of step (1.12) is: Step (1.12.1) computes the enhanced embedding of node v in represents the initial embedding of node v, || is the connection operation; Step (1.12.2) calculates the sampling probability of node v to adjacent node i In the matching task between Mashup m and API a, the sampling probability of each node v in the hypergraph for each adjacent node i is as follows: Among them, exp(·) represents the exponential function of e, φ(·) represents a nonlinear transformation, σ(·) is a nonlinear activation function, and sim(·,·) uses cosine similarity to calculate the similarity between two vertices; Represents the hypergraph G sin Adjacency matrix S sin The element value of the row and column where the middle node v, i is located, Similarly; Indicates the enhanced embedding of API service nodes and Mashup service nodes involved in the matching task; Step (1.12.3) is based on the current sampling probability Aggregate the domain of node v and obtain the domain aggregation embedding in, Step (1.12.4) calculates the embedding obtained by convolution Complete the first layer of convolution; in, represents the parameter matrix, l = 1; Step (1.12.5) uses the data obtained in step (1.12.4) Replace step (1.12.1) Update Repeat steps (1.12.2-1.12.4) to update calculate and Complete the convolution of the second layer; Step (1.12.6) repeats step (1.12.5) to complete the first L-1 convolutions; Step (1.12.7) Since the purpose of the hypergraph convolution is to use the neighborhood nodes to generate the embedding of the cold start node, the last layer only uses To represent the target embedding 5. The cold start service recommendation method based on a multi-strategy pre-training model according to any one of claims 2 to 4, characterized in that: The process of the second step is as follows: Step (2.1) is used to obtain the adjacency matrix S from step (1.4) sin Perform singular value decomposition (SVD) to obtain S sin =UDV T , where U and V are S×S orthogonal matrices, V T represents the transpose of the matrix V, and D is the matrix S sin The diagonal matrix of the singular values ​​of ; Step (2.2) truncates the list of singular values, retains the first q singular values, and compresses the adjacency matrix to be approximately S sin S sin ', and obtain the original hypergraph G sin The enhanced hypergraph G sin '; Reconstruct the matrix S sin ' is the adjacency matrix S sin It emphasizes the core architecture of the hypergraph by highlighting the key interactions in the service hypergraph network, while maintaining the global collaborative signal by considering all the connection relationships in the hypergraph. The formula is as follows: in, and Contains the first q columns of U and V respectively, is a diagonal matrix consisting of the first q singular values, Represents the matrix V q The transpose of Step (2.3) in the newly generated adjacency matrix S sin 'Perform the same convolution as step (1.12) to obtain the embedding of node v in the lth convolution layer on the enhanced hypergraph Step (2.4) performs contrastive learning on the enhanced hypergraph embedding and the original hypergraph embedding. The loss of contrastive learning is defined as follows: Among them, sim(·,·) and τ represent the cosine similarity and temperature hyperparameters, respectively. represents the embedding of node v obtained by the l-th convolution on the original hypergraph, represents the embedding of node v obtained by the l-th convolution on the enhanced hypergraph, log represents the logarithmic operation, L represents the number of convolution layers, and θ2 represents the parameters of the contrastive learning task; Step (2.5) When the contrastive learning framework is trained, the model will discard the other parts and only keep the encoder on the trained enhanced hypergraph. When a new cold start node enters, the encoder performs convolution to obtain the enhanced service embedding 6. The cold start service recommendation method based on a multi-strategy pre-training model as claimed in claim 5, characterized in that: The process of the third step is as follows: Step (3.1) uses the BERT model as a feature extractor to extract the embedding vector from the text information j, denoted as x j , where BERT is a pre-trained language representation model proposed by Google, which is used to efficiently represent highly unstructured text data as vectors; Step (3.2) uses the feature encoder to obtain the text embedding of node v The feature encoder is a multi-layer perceptron, which includes an input layer, an output layer, and a hidden layer, and the layers are fully connected: Where W2, W3 and b1, b2 represent the trainable matrix and bias vector of the encoder respectively, and f(·) represents the activation function; Step (3.3) uses LightGCN on the warm start dataset and jointly trains it with the feature extractor and encoder built in steps (3.1-3.2) under the same supervision to obtain the collaborative embedding of node v Step (3.4) performs contrastive learning between collaborative embedding and text embedding, and the loss is defined as follows: Among them, <·,·> represents the vector inner product, N + (v) represents the positive sample v of node v + Set, N - (v) represents the negative sample v of node v - gather, represents the positive sample v of node v + The collaborative embedding of Represents the negative sample v of node v - The collaborative embedding of , ln represents the natural logarithm operation, and θ3 is the parameter of contrastive learning; Step (3.5) When the contrastive learning framework is trained, the model will discard the other parts and only retain the parameters of the trained text feature encoder. When a new cold start node enters, the encoder is executed to obtain the content-optimized service embedding 7. The cold start service recommendation method based on a multi-strategy pre-training model according to claim 6, characterized in that: The process of the fourth step is as follows: Step (4.1) first uses the hypergraph encoder and text feature encoder trained in the pre-training task to generate the corresponding service embedding Step (4.2) concatenates the generated embeddings to transform them into the final serving embedding: Among them, W5 is a trainable parameter matrix, || represents the connection operation; Step (4.3) ends with the final embedding of Mashup m and API a The inner product of is used to calculate the correlation score, that is: Step (4.4) defines the BPR loss to optimize W5 and fine-tune the parameters {θ1,θ2,θ3} of each module: L BPR =∑-lnσ(y(m,a + )-y(m,a - )) Among them, a + Indicates the API that has a history of interaction with m, a - Indicates that there is no interaction history with the API m.