Cloud native scene cold start data service feature extraction method

By using hypergraph neural network and adaptive sampling strategies to extract cold start service features in cloud-native scenarios, the recommendation problem of Web API without historical interaction records is solved, and high-quality cold start API embedding and more accurate service recommendations are achieved.

CN120066557APending Publication Date: 2025-05-30CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510085356.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively recommend Web APIs without historical interaction records, which makes it difficult to solve the cold start recommendation problem.

Method used

The hypergraph neural network in cloud-native scenarios is adopted to extract cold start service hypergraphs based on label information, use meta-aggregators and adaptive sampling strategies to improve hypergraph convolution operations, optimize cold start encoder parameters, and fine-tune models through downstream recommended tasks to build high-quality cold start service feature embedding.

Benefits of technology

It effectively solves the recommendation problem of pure cold start API without historical interactive information, improves the accuracy and efficiency of service recommendations, and improves the embedding quality of cold start APIs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066557A_ABST
    Figure CN120066557A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud native scene cold start data service feature extraction method, which comprises the following steps of: 1, extracting a cold start service hypergraph based on label information; 2, processing cold start neighbor node data based on a meta aggregator; step 3, improving hypergraph aggregation operation based on an adaptive sampling strategy; 4, optimizing the parameters of the cold start encoder according to the embedding predicted by the hot start encoder; and 5, carrying out fine adjustment on the cold start recommendation model through a downstream recommendation task. According to the method, the recommendation problem of the pure cold start API without any historical interaction information is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of service computing, and particularly to a method for extracting cold start data service features in a cloud native scenario. Background Art

[0002] Cloud native is a software approach for building, deploying, and managing modern applications in a cloud computing environment. Modern enterprises hope to build highly scalable, flexible, and resilient applications that can be quickly updated to meet customer needs. To this end, they use modern tools and technologies that inherently support application development on cloud infrastructure. These cloud native technologies enable rapid and frequent changes to applications without affecting service delivery, thus providing adopters with an innovative competitive advantage.

[0003] With the development of cloud computing, more and more manufacturers are providing on-demand cloud native services online through cloud platforms. Service-oriented computing aims to use services as basic building blocks to create fast, low-cost, secure, and reliable applications. Compared with traditional software components, services have autonomy, self-descriptiveness, reusability, and high portability.

[0004] In recent years, more and more developers have benefited from the reuse of Web services, usually in the form of Web APIs. A Web API is an application programming interface that allows Web applications to access and implement storage services, messaging services, computing services, and other functions. For example, technology giants such as Google and Microsoft are opening up their software and data resources as Web services through Web APIs to attract a wider Internet user group. In addition, with the continuous booming of the Web API economy, numerous Web API sharing repositories have emerged in the market, such as ProgrammableWeb, RapidAPI, and ShowAPI websites. Taking the open cloud API platform RapidAPI as an example, it currently supports more than 40,000 APIs, more than 12,000 API publishers, more than 200,000 monthly active subscribers, millions of registered developers, and more than 5 billion API requests per month.

[0005] Facing the rapid development of the information society and the emergence of a large number of additional demands, the existing Web API functions are insufficient to cope with increasingly complex business scenarios. To fill this gap, Mashup technology has emerged. The core concept of Mashup is to use Web APIs as reusable components to create new products or solutions. It matches user-customized needs by integrating multiple services, greatly reducing the development threshold, enabling developers without programming skills to quickly build applications using ready-made APIs, and also helping developers shorten the development cycle.

[0006] However, in the face of the rapidly expanding Web API market with diverse functions, it has become increasingly difficult to quickly screen out the most suitable APIs for user needs, that is, service recommendation. In recent years, due to the outstanding ability of Graph Neural Networks (GNN) in relational data modeling, GNN-based models have demonstrated excellent potential in the field of service recommendation. Hypergraphs generalize the concept of edges in traditional graphs, allowing edges to connect more than two nodes, providing an intuitive and natural framework for modeling complex high-order relationships between services. However, most existing studies only focus on Web APIs with interaction records with specific Mashups, ignoring the problem of effectively recommending new and non-interacted Web APIs, that is, the cold-start recommendation problem. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, to solve the cold-start problem in the actual service recommendation scenario, to construct high-quality embeddings for newly created or APIs without any historical interaction records, and thus achieve service recommendation, the present invention proposes a method for extracting cold-start data service features in a cloud-native scenario, with hypergraph neural networks as the technical support, aiming to solve the pure cold-start problem existing in the actual service recommendation scenario, and focusing on core issues such as how to efficiently extract cold-start service feature information.

[0008] To solve the above problems, the present invention provides the following technical solutions:

[0009] A method for extracting cold-start data service features in a cloud-native scenario, comprising the following steps:

[0010] First step, extract the cold-start service hypergraph based on label information;

[0011] Second step, process the cold-start neighbor node data based on a meta-aggregator. A cold-start neighbor node means that there is a cold-start node in the neighborhood of the target node in the hypergraph. Assume the target node is v tar , and its neighbor node is i. When performing hypergraph convolution, if i is a cold-start node, its embedding will indirectly affect the embedding of v tar ;

[0012] Third step, improve the hypergraph aggregation operation based on an adaptive sampling strategy. The adaptive sampling strategy is a sampling strategy that dynamically adjusts according to the Mashup-API involved in the matching task. In the matching task of Mashup m and API a, when the target node v tar performs feature aggregation, this strategy not only considers the distance between the neighborhood nodes and v tar , but also considers the distance between the neighborhood nodes and m and a;

[0013] Step 4. Optimize the cold-start encoder parameters according to the embeddings predicted by the warm-start encoder;

[0014] Step 5. Fine-tune the cold-start recommendation model through downstream recommendation tasks.

[0015] Furthermore, the process of the first step is as follows:

[0016] Step (1.1) Crawl API service and Mashup service information from relevant API websites;

[0017] Step (1.2) Analyze the crawled service set, select meta-information that is conducive to improving the accuracy of service recommendation and is rich in content, and form a service dataset, including a warm-start dataset and a cold-start dataset; the cold-start dataset further deletes the meta-information of the call relationship between Mashup and API on the basis of the warm-start dataset;

[0018] Among them, the retained meta-information includes Mashup, API, the call relationship between Mashup and API, the tags carried by Mashup and API, and service description texts; among them, the tags represent some functions or attributes of Mashup and API, and the tag information also includes the collaboration information between Mashup and API;

[0019] Step (1.3) Since the cold-start service does not contain call information, a subgraph structure in which services have the same tag relationship in pairs and form a triadic closure is defined as a motif to extract the many-to-many relationship between services based on tags;

[0020] Among them, a motif represents a connection pattern that recurs in a network and whose quantity is significantly higher than that of other complex networks. This connection pattern often holds the key information of the network; a triadic closure is a concept in social network theory, referring to a property of a triad composed of three nodes A, B, and C, that is, there are strong connections between A and B, A and C, and B and C;

[0021] Step (1.4) Construct a cold-start service hypergraph G according to the motif proposed in step (1.3) sin ;

[0022] Step (1.5) According to the hypergraph G constructed in step (1.4) sin , initialize a cold-start encoder f using a hypergraph convolutional network cold , which is used to construct embeddings for cold-start service nodes in the meta-learning setting and improve the convolutional operation in subsequent steps.

[0023] Preferably, the process of step (1.4) is as follows:

[0024] Step (1.4.1) Define a matrix is a 0-1 matrix storing the relationship between storage services and label carriers, where S represents the total number of services, T represents the total number of labels, and for each pair (s, t), r s,t = 1 indicates that service s carries label t, and vice versa;

[0025] Step (1.4.2) According to the matrix R of the relationship between services and label carriers s Calculate to obtain where represents the number of identical labels carried between services, represents the transpose of matrix R s ;

[0026] Step (1.4.3) Take the upper right part of the diagonal of matrix and set all numbers greater than 0 to 1 to construct an upper triangular matrix R s * , where the elements in R s * indicate whether the corresponding services carry the same label;

[0027] Step (1.4.4) Define S sin as the adjacency matrix of the cold start service hypergraph G sin ;

[0028] where the adjacency matrix S sin of the hypergraph G sin is calculated as S sin =(R s * ·R s * )⊙R s * , and ⊙ represents the Hadamard product of the matrices on the left and right sides.

[0029] Furthermore, the process of the second step is as follows:

[0030] Step (2.1) Instantiate the meta-aggregator g as a self-attention encoder, whose core is to use the self-attention mechanism (Self-Attention) to calculate the attention within neighbor nodes;

[0031] Step (2.2) g takes the initial embeddings of k first-order neighbors of node v as input, where N(v) represents the neighborhood of v;

[0032] Step (2.3) For each first-order neighbor i of v, calculate the attention scores of all first-order neighbors for i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h i of each neighbor i, that is

[0033] Step (2.4) averages the embeddings of all neighbors to obtain the meta-embedding of the node That is

[0034] Step (2.5) uses LightGCN as a warm-start encoder to learn the true embeddings of nodes based on the observed rich interactions, i.e., the call relationships between Mashup and API, on the warm-start dataset constructed in step (1.2). For further adjustment Among them, LightGCN is an efficient graph convolutional recommendation model that obtains the embeddings of Mashup and API through graph convolution and matches API and Mashup according to the distance;

[0035] Step (2.6) calculates using cosine similarity Between And take this as the loss function of the meta-aggregator. The formula is as follows

[0036]

[0037] Among them, cos(·,·) represents cosine similarity calculation, ||·|| represents finding the modulus length, V represents the set of hypergraph nodes, and the argmax function represents finding the meta-aggregator parameter Q that maximizes the formula Called Q * , The meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node.

[0038] Furthermore, in the third step, the feature aggregation operation in the conventional graph convolutional network cannot reflect the different contributions of different neighbors to the target node, and also lacks the flexibility to adjust its own convolution strategy for different matching tasks. The adaptive sampling strategy is a sampling strategy that dynamically adjusts according to the Mashup-API involved in the matching task. In the matching task of Mashupm and API a, the target node v tar When performing feature aggregation, this strategy not only considers the distance between the neighborhood node and v tar , but also considers the distance between the neighborhood node and m and a. The process is as follows

[0039] Step (3.1) combines the meta-embedding With the embedding cold Obtained by the l-th layer convolution in f To obtain the enhanced embedding of node v at the l-th layer Among them, || is the concatenation operation, l = 0 Denote the initial embedding of the node. Let \(l = 1,\cdots,L\) denote the embedding of the node obtained by convolution at the \(l\)-th layer, and \(L\) denote the total number of convolutional layers;

[0040] Step (3.2) defines that in the matching task of Mashup \(m\) and API \(a\), the sampling probability of each node \(v\) in the hypergraph for each adjacent node \(i\) at the \(l\)-th layer is as follows:

[0041]

[0042] where \(\exp(\cdot)\) represents the exponential function of \(e\), \(\varphi(\cdot)\) represents a non-linear transformation, \(\sigma(\cdot)\) is a non-linear activation function, and \(\text{sim}(\cdot,\cdot)\) uses cosine similarity to calculate the similarity between two vertices; Denote the hypergraph \(G\) sin Adjacency matrix \(S\) sin The element value at the row and column where nodes \(v\) and \(i\) are located Similarly; Denote the enhanced embedding of the API service node and Mashup service node involved in the matching task at the \(l\)-th layer; when \(l = 0\), Denote the sampling probability of the initial embedding of the neighbor node \(i\). Let \(l = 1,\cdots,L\) denote the sampling probability of the embedding of the neighbor node \(i\) obtained by convolution at the \(l\)-th layer;

[0043] Step (3.3) combines the meta-aggregator and the adaptive sampling strategy to perform a convolution operation on the hypergraph;

[0044] where the improved convolution formula is defined as

[0045]

[0046] where Denote the aggregated embedding of the neighbor nodes of node \(v\) at the \(l\)-th layer, Denote the embedding obtained by convolution of the neighbor nodes of \(v\) in the previous layer \(l - 1\), Denote the sampling probability of node \(v\) for the neighbor nodes at the \((l - 1)\)-th layer, Denote the embedding of node \(v\) obtained by convolution at the \(l\)-th layer, Is the parameter matrix of the \(l\)-th layer convolution, Denote the embedding of node \(v\) obtained by convolution in the previous layer \(l - 1\); since the purpose of hypergraph convolution is to use neighborhood nodes to generate the embedding of cold-start nodes, therefore, only Is used to represent the target embedding obtained

[0047] The process of the fourth step is as follows:

[0048] Step (4.1) Use the warm-start encoder LightGCN trained in step (2.5) to obtain the true embedding of the nodes

[0049] Step (4.2) Define the loss function of the cold-start encoder using cosine similarity where θ 1 represents the parameters of the cold-start encoder;

[0050] Step (4.3) Train the cold-start encoder to obtain the pre-trained embedding

[0051] The process of the fifth step is as follows:

[0052] Step (5.1) Obtain the final embedding of the service according to the pre-trained embedding obtained in step (4.3) where W is the parameter matrix; where W is the parameter matrix;

[0053] Step (5.2) Calculate the correlation score using the inner product of the final embeddings of Mashup and API, that is the inner product of the final embeddings of Mashup and API, that is

[0054] Step (5.3) Define the BPR loss as the loss function of the fine-tuning task

[0055] where the loss function is L BPR = ∑ -lnσ(y(m,a + ) - y(m,a - ))), a + represents the API with an interaction history with m, a - represents the API without an interaction history with m, ln represents the natural logarithm operation, and the model is fine-tuned using the downstream recommendation task to obtain higher-quality cold-start service feature embeddings.

[0056] The beneficial effects of the present invention are: In the actual service recommendation scenario, there are newly created or APIs without any historical interaction records. This method constructs a hypergraph through tags and uses a hypergraph convolutional network to extract service features to construct service embeddings, thereby effectively solving the recommendation problem of pure cold-start APIs without any historical interaction information. The main advantages of this method are: 1) constructing a service hypergraph using motifs composed of tags; 2) using the meta-aggregator g to enhance the embeddings of the cold-start neighbors of the nodes; 3) adjusting the self-convolution strategy according to the Mashup-API pairs involved in the matching task with a dynamically adjusted adaptive sampling strategy. Description of the Drawings

[0057] Figure 1 is a schematic diagram of the motif structure based on tags in the present invention.

[0058] Figure 2 It is a schematic diagram of the aggregation process of the meta aggregator in the present invention.

[0059] Figure 3 It is a schematic diagram of the comparison result of the cosine similarity of the service embedding constructed by the method proposed in this embodiment and the comparison baseline. Detailed implementation manners

[0060] To make the service feature extraction process of the present invention more obvious and understandable, specific embodiments are given below and detailed descriptions are made in conjunction with the accompanying drawings.

[0061] Referring to Figures 1 to 3 , a method for extracting cold start data service features in a cloud native scenario, includes the following steps:

[0062] The first step is to extract the cold start service hypergraph based on label information, and the process is as follows:

[0063] Step (1.1) Crawl API service and Mashup service information from relevant API websites;

[0064] Step (1.2) Analyze the crawled service set, select meta information that is beneficial to improving the accuracy of service recommendation and is rich in content, and form a service data set, including a warm start data set and a cold start data set. The cold start data set further deletes the meta information of the call relationship between Mashup and API on the basis of the warm start data set;

[0065] Among them, the retained meta information includes Mashup, API, the call relationship between Mashup and API, the labels carried by Mashup and API, and the service description text; among them, the label represents part of the functions or attributes of Mashup and API, and the label information also includes the collaboration information between Mashup and API;

[0066] In step (1.3), since the cold start service does not contain call information, a subgraph structure in which services have the same label relationship pairwise and form a triadic closure is defined as a motif, and its structure is as Figure 1 shown, which is used to extract the many-to-many relationship between services based on labels;

[0067] Among them, the motif represents a connection pattern that repeatedly appears in the network and the quantity is significantly higher than other complex networks. This connection pattern often holds the key information of the network; the triadic closure is a concept in social network theory, referring to a property of a triad composed of three nodes A, B, and C, that is, there are strong connections between A and B, A and C, and B and C;

[0068] Step (1.4) constructs the cold start service hypergraph \(G\) in steps (1.4.1 - 1.4.4) based on the motif proposed in step (1.3). sin , and the process is as follows:

[0069] Step (1.4.1) defines the matrix as a 0 - 1 matrix storing the relationship between services and label carrying, where \(S\) represents the total number of services, \(T\) represents the total number of labels, and for each pair \((s, t)\), \(r\) s,t = 1 indicates that service \(s\) carries label \(t\), and vice versa;

[0070] Step (1.4.2) calculates based on the service - label carrying relationship matrix \(R\) s to obtain where represents the number of services carrying the same label, represents the transpose of the matrix \(R\) s ;

[0071] Step (1.4.3) takes the upper right part of the diagonal of the matrix and sets all numbers greater than 0 to 1 to construct the upper triangular matrix \(R\) s * , where the elements in \(R\) s * indicate whether the corresponding services carry the same label;

[0072] Step (1.4.4) defines \(S\) sin as the adjacency matrix of the cold start service hypergraph \(G\) sin ;

[0073] where the adjacency matrix \(S\) of the hypergraph \(G\) sin is calculated by the formula \(S\) sin =(R sin s * ·R s * )⊙R s * , and ⊙ represents taking the Hadamard product of the matrices on the left and right sides;

[0074] Step (1.5) initializes a cold start encoder \(f\) sin using the hypergraph convolutional network based on the hypergraph \(G\) constructed in step (1.4) cold for constructing embeddings for cold start service nodes in the meta - learning setting and improving the convolutional operation in subsequent steps. The hypergraph convolutional network can be instantiated using different models, such as HyperGCN, FastHyperGCN.

[0075] Step 2: Process the cold-start neighbor node data based on the meta-aggregator. Here, the cold-start neighbor node means that there are cold-start nodes in the neighborhood of the target node in the hypergraph. Assume the target node is v tar , and its neighbor node is i. When performing hypergraph convolution, if i is a cold-start node, its embedding will indirectly affect the embedding of v tar . The process is as follows:

[0076] Step (2.1) Instantiate the meta-aggregator g as a self-attention encoder. Its core is to use the self-attention mechanism (Self-Attention) to calculate the attention within the neighbor nodes; the overall aggregation process is as Figure 2 shown, where the higher or lower the weight is the size of the attention score, and the darker the color represents the higher the weight;

[0077] Step (2.2) g takes the initial embeddings of k first-order neighbors of node v as input, where N(v) represents the neighborhood of v;

[0078] Step (2.3) For each first-order neighbor i of v, calculate the attention scores of all first-order neighbors for i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h of each neighbor i i , that is

[0079] Step (2.4) Average the embeddings of all neighbors to obtain the meta-embedding of the node That is

[0080] Step (2.5) Use LightGCN as a warm-start encoder to learn the true embeddings of nodes based on the observed rich interactions, that is, the call relationship between Mashup and API, on the warm-start dataset constructed in step (1.2) for further adjustment Among them, LightGCN is an efficient graph convolutional recommendation model that obtains the embeddings of Mashup and API through graph convolution and matches API and Mashup according to the distance;

[0081] Step (2.6) Calculate the difference between and using cosine similarity, and use this as the loss function of the meta-aggregator. The specific formula is as follows

[0082]

[0083] Among them, cos(·,·) represents cosine similarity calculation, ||·|| represents finding the modulus length, V represents the set of hypergraph nodes, and the argmax function represents finding the value that makes the formula The maximum meta-aggregator parameter Q, denoted as Q * , the meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node.

[0084] Step 3: Improve the hypergraph aggregation operation based on the adaptive sampling strategy. In a conventional graph convolutional network, the feature aggregation operation does not reflect the different contributions of different neighbors to the target node and lacks the flexibility to adjust its own convolution strategy for different matching tasks; the adaptive sampling strategy is a sampling strategy dynamically adjusted according to the Mashup-API involved in the matching task. In the matching task of Mashup m and API a, for the target node v tar when performing feature aggregation, this strategy not only considers the distance between the neighborhood node and v tar but also considers the distance between the neighborhood node and m and a;

[0085] The process of the third step is as follows:

[0086] Step (3.1) Combine the meta-embedding with the embedding cold obtained from the l-th layer convolution in f to obtain the enhanced embedding of node v at the l-th layer, where || is the concatenation operation, l = 0, represents the initial embedding of the node, l = 1,..., L represents the embedding obtained by the node from the l-th layer convolution, and L represents the total number of convolution layers;

[0087] Step (3.2) Define the sampling probability of each node v in the hypergraph for each adjacent node i at the l-th layer in the matching task of Mashup m and API a as follows,

[0088]

[0089] where exp(·) represents the exponential function of e, φ(·) represents a non-linear transformation, σ(·) is a non-linear activation function, and sim(·,·) uses the cosine similarity to calculate the similarity between two vertices; represents the hypergraph G sin adjacency matrix S sin the element value at the row and column where nodes v and i are located, similarly; represents the enhanced embedding of the API service node and Mashup service node involved in the matching task at the l-th layer; when l = 0, represents the sampling probability of the initial embedding of neighbor node i, and when l = 1,..., L, it represents the sampling probability of the embedding obtained by neighbor node i from the l-th layer convolution;

[0090] Step (3.3) performs a convolution operation on the hypergraph by combining the meta-aggregator and the adaptive sampling strategy;

[0091] Among them, the improved convolution formula is defined as,

[0092]

[0093] Among them, represents the aggregated embedding of the neighbor nodes of node v at the l-th layer, represents the embedding obtained by convolving the neighbor nodes of v at the previous layer l - 1, represents the sampling probability of node v for its neighbor nodes at the l - 1 layer, represents the embedding of node v obtained by convolution at the l-th layer, is the parameter matrix of the l-th layer convolution, represents the embedding of node v obtained by convolution at the previous layer l - 1; since the purpose of hypergraph convolution is to use neighborhood nodes to generate the embedding of cold start nodes, therefore, only is used to represent the target embedding to obtain

[0094] Fourthly, optimize the cold start encoder parameters according to the embedding predicted by the warm start encoder, and the process is as follows:

[0095] Step (4.1) uses the warm start encoder LightGCN trained in step (2.5) to obtain the true embedding of the nodes

[0096] Step (4.2) defines the loss function of the cold start encoder using cosine similarity where θ 1 represents the parameters of the cold start encoder;

[0097] Step (4.3) trains the cold start encoder to obtain the pre-trained embedding

[0098] Fifthly, fine-tune the cold start recommendation model through downstream recommendation tasks, and the process is as follows:

[0099] Step (5.1) obtains the final embedding of the service according to the pre-trained embedding obtained in step (4.3) where W is the parameter matrix;

[0100] Step (5.2) calculates the correlation score by taking the inner product of the final embeddings of Mashup and API, that is

[0101] ​​Step (5.3) defines the BPR loss as the loss function for the fine-tuning task.

[0102] where the loss function is L BPR = ∑ -lnσ(y(m,a + ) - y(m,a - ))), a + represents the API with an interaction history with m, a - represents the API without an interaction history with m, ln represents the natural logarithm operation, and the downstream recommendation task is used to fine-tune the model to obtain higher-quality cold start service feature embeddings.

[0103] This embodiment analyzes the actual effects of the invention with specific service data as follows:

[0104] 1) Experimental dataset: The data crawled from the Programmable Web website from 2019 to 2020 is selected as the experimental dataset, including 6,217 Mashup services, 11,930 API services, and related meta-information.

[0105] 2) Baseline methods:

[0106] NCF: This is a recommendation model that combines neural networks and traditional collaborative filtering techniques. It uses a multi-layer perceptron and matrix factorization to learn the representation of items.

[0107] PinSage: This is a general graph neural network model. This model randomly samples neighbor nodes and uses an average function to achieve information aggregation, so as to effectively process large-scale graph data.

[0108] GAT: This is a general graph attention model. It aggregates information of neighbor nodes without sampling through the attention mechanism, and can adaptively assign different weights to different neighbors, so as to more accurately capture the semantic information in the graph structure.

[0109] NGCF: It is a collaborative filtering algorithm based on graph neural networks, but second-order interactions are added during the message passing process to enhance the expression ability of the model and can simultaneously capture the complex non-linear relationships between users and items.

[0110] LightCGN: It is an efficient graph convolutional recommendation model. It discards the feature transformation and non-linear activation functions in NGCF, obtains the embeddings of Mashup and API through graph convolution, and matches API and Mashup according to the distance.

[0111] 3) Evaluation metrics:

[0112] The NDCG metric is a normalized representation of Discounted Cumulative Gain (DCG). This metric takes into account the recommendation order, where a higher normalized discounted cumulative gain for recommendations ranked higher indicates better performance, as shown below.

[0113]

[0114] Where K represents the number of recommended APIs, n represents the ranking of the service in the recommendation list, and IDCG represents the DCG in the ideal case.

[0115] HR is a commonly used Top-N evaluation metric, especially in cases where it is independent of the recommendation order. A higher HR value indicates better recommendation performance, and its calculation method is as follows.

[0116]

[0117] Where hit(i) indicates whether the API is hit in the Mashup's recommendation list, and N represents the actual number of APIs called by the Mashup.

[0118] 4) Experimental results: The comparison results of the cosine similarity of the service embeddings constructed by the method proposed in this invention and the comparison baseline are as Figure 3 shown. Where My. represents the method proposed in this invention; the y-axis label COS represents the cosine similarity. The larger the COS, the higher the embedding quality.

[0119] Table 1 shows the experimental comparison results of the method proposed in this invention and the baseline method in terms of recommendation performance, with the selected evaluation metrics being NDCG@5 and HR@5. The best result for each metric is shown in bold, and the second-best result is underlined.

[0120] Method NDCG@5 HR@5 NCF 0.066 0.081 PinSage 0.087 0.096 GAT <![CDATA 0.142 > <![CDATA 0.164 > NGCF 0.139 0.161 LightGCN 0.141 0.163 My. 0.231 0.253

[0121] Table 1

[0122] Analysis Figure 3 From the data in Table 1, it can be seen that compared with the baseline, the cold start data service feature extraction method proposed in this invention has achieved better results in all aspects. This is mainly due to the construction of the motif-based hypergraph, which effectively extracts the originally loose and difficult-to-use label information, and the hypergraph helps the transmission of high-order information during the convolution process. In terms of the constructed embedding quality, the cosine similarity of My. has increased by 15.82% compared to the best baseline. In terms of the NDCG@5 metric and the HR@5 metric, My. has increased by 62.68% and 54.27% respectively compared to the best baseline.

[0123] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is for illustrative purposes only. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment, and the protection scope of the present invention also covers equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.

Claims

1. A method for extracting cold start data service features in a cloud native scenario, characterized in that: The method comprises the following steps: The first step is to extract the cold start service hypergraph based on label information; The second step is to process the cold start neighbor node data based on the meta-aggregator. The cold start neighbor node indicates that there is a cold start node in the neighborhood of the target node in the hypergraph. Assume that the target node is v tar , whose neighbor node is i. When performing hypergraph convolution, if i is a cold start node, its embedding will indirectly affect v tar Embedding The third step is to improve the hypergraph aggregation operation based on the adaptive sampling strategy. The adaptive sampling strategy is a sampling strategy that is dynamically adjusted according to the Mashup-API pairs involved in the matching task. In the matching task of Mashup m and API a, the target node v tar When performing feature aggregation, this strategy not only considers the neighborhood nodes and v tar The distance between the neighboring nodes and m and a will also be considered; Step 4: Optimize the cold start encoder parameters based on the embedding predicted by the warm start encoder. Step 5: Fine-tune the cold start recommendation model through downstream recommendation tasks.

2. The method for extracting cold start data service features in a cloud native scenario according to claim 1, characterized in that: The process of the first step is as follows: Step (1.1) crawl API service and Mashup service information from relevant API websites; Step (1.2) analyzes the crawled service set, selects meta-information that is conducive to improving the accuracy of service recommendations and is rich in content, and forms a service data set, including a hot start data set and a cold start data set; the cold start data set further deletes the meta-information of the call relationship between Mashup and API on the basis of the hot start data set; The retained meta-information includes Mashup, API, calling relationship between Mashup and API, labels carried by Mashup and API, and service description text; labels represent some functions or attributes of Mashup and API, and label information also includes collaborative information between Mashup and API; Step (1.3) Since the cold start service does not contain call information, the subgraph structure that has the same label relationship between services and forms a ternary closure is defined as a motif to extract the many-to-many relationship between services based on labels; Among them, motifs represent connection patterns that appear repeatedly in the network and are significantly higher in number than other complex networks. This connection pattern often holds key information about the network; ternary closure is a concept in social network theory, which refers to a property of a triple consisting of three nodes A, B, and C, that is, there are strong connections between A and B, A and C, and B and C; Step (1.4) constructs the cold start service hypergraph G according to the model proposed in step (1.3) sin ; Step (1.5) The hypergraph G constructed according to step (1.4) sin , initialize a cold start encoder f using a hypergraph convolutional network cold , which is used to build embeddings for cold-start serving nodes in a meta-learning setting and improve the convolution operation in subsequent steps.

3. The method for extracting cold start data service features in a cloud native scenario according to claim 2, characterized in that: The process of step (1.4) is: Step (1.4.1) defines the matrix is a 0-1 matrix of storage services and label carrying relationships, where S represents the total number of services and T represents the total number of labels. For each pair (s, t), r s,t =1 means service s carries label t, and vice versa; Step (1.4.2) is based on the service and label carrying relationship matrix R s Calculated in Indicates the number of services that carry the same label. Represents the matrix R s The transpose of Step (1.4.3) Take the matrix The upper right corner of the diagonal line sets all numbers greater than 0 to 1 to construct the upper triangular matrix R s * , where R s * The middle element indicates whether the corresponding location services carry the same label; Step (1.4.4) defines S sin Supergraph G for cold start service sin The adjacency matrix of Among them, the hypergraph G sin The adjacency matrix S sin The calculation formula is S sin =(R s * ·R s * )⊙R s * , ⊙ represents the Hadamard product of the left and right matrices.

4. The method for extracting cold start data service features in a cloud native scenario according to any one of claims 1 to 3, characterized in that: The process of the second step is as follows: Step (2.1) instantiates the meta-aggregator g as a self-attention encoder, the core of which is to use the self-attention mechanism to calculate the attention within the neighboring nodes; Step (2.2) g accepts the initial embeddings of the k first-order neighbors of node v As input, where N(v) represents the neighborhood of v; Step (2.3) For each first-order neighbor i of v, calculate the attention scores of all first-order neighbors to i, and aggregate the embeddings of all neighbors according to the attention scores to generate the embedding h of each neighbor i i ,Right now Step (2.4) averages the embeddings of all neighbors to obtain the meta-embedding of the node Right now Step (2.5) uses LightGCN as the warm-start encoder to learn the true embedding of nodes based on the observed rich interactions, i.e., the call relationship between Mashup and API, on the warm-start dataset constructed in step (1.2). For further adjustment Among them, LightGCN is an efficient graph convolution recommendation model, which embeds Mashup and API through graph convolution and matches API and Mashup based on distance; Step (2.6) uses cosine similarity calculation and The difference between them is used as the loss function of the meta-aggregator. The formula is as follows: Among them, cos(·,·) represents the cosine similarity calculation, ||·|| represents the modulus length, V represents the hypergraph node set, and the argmax function represents finding the formula The largest meta-aggregator parameter Q, called Q * ,The meta-aggregator extends the original hypergraph convolution by emphasizing the representation of cold-start neighbors in each convolution step to improve the final embedding of the target node.

5. The method for extracting cold start data service features in a cloud native scenario according to claim 4, characterized in that: The process of the third step is as follows: Step (3.1) embeds the element With f cold The embedding obtained by the l-th convolution layer Combined, we get the enhanced embedding of node v at layer l Among them, || is the connection operation, l = 0, represents the initial embedding of the node, l=1,...,L, represents the embedding of the node obtained by the lth convolution layer, and L represents the total number of convolution layers; Step (3.2) defines that in the matching task between Mashup m and API a, the sampling probability of each node v in the hypergraph for each adjacent node i at layer l is as follows: Among them, exp(·) represents the exponential function of e, φ(·) represents a nonlinear transformation, σ(·) is a nonlinear activation function, and sim(·,·) uses cosine similarity to calculate the similarity between two vertices; Represents the hypergraph G sin Adjacency matrix S sin At node v, the value of the element in the row and column where i is located, Similarly; Indicates the enhanced embedding of the API service nodes and Mashup service nodes involved in the matching task at the lth layer; when l = 0, represents the sampling probability of the initial embedding of neighbor node i, l = 1, ..., L, represents the sampling probability of the embedding obtained by convolution of neighbor node i at the lth layer; Step (3.3) combines the meta-aggregator and the adaptive sampling strategy to perform convolution operations on the hypergraph; Among them, the improved convolution formula is defined as, in, represents the aggregate embedding of the neighbor nodes of node v at layer l, represents the embedding of v’s neighbor nodes obtained by convolution in the previous layer l-1, represents the sampling probability of node v to its neighbor node at the l-1 layer, represents the embedding of node v obtained by convolution at layer l, is the parameter matrix of the l-th layer convolution, represents the embedding of node v obtained by the previous layer l-1 convolution; because the purpose of hypergraph convolution is to use neighborhood nodes to generate the embedding of cold start nodes, the last layer only uses To represent the target embedding 6. The method for extracting cold start data service features in a cloud native scenario according to claim 5, characterized in that: The process of the fourth step is as follows: Step (4.1) uses the warm-start encoder LightGCN trained in step (2.5) to obtain the true embedding of the node Step (4.2) defines the loss function of the cold start encoder using cosine similarity Where θ1 represents the parameter of the cold start encoder; Step (4.3) trains the cold start encoder to obtain the pre-trained embedding 7. The method for extracting cold start data service features in a cloud native scenario according to claim 6, characterized in that: The process of the fifth step is as follows: Step (5.1) is based on the pre-trained embedding obtained in step (4.3) Get the final embed of the service Where W is the parameter matrix; Step (5.2) Final Embedding of Mashup and API The inner product of is used to calculate the correlation score, that is Step (5.3) defines BPR loss as the loss function for the fine-tuning task, Among them, the loss function is L BPR =∑-lnσ(y(m,a + )-y(m,a - )), a + Indicates the API that has a history of interaction with m, a - represents an API that has no interaction history with m, ln represents the natural logarithm operation, and the model is fine-tuned using downstream recommendation tasks to obtain higher quality cold start service feature embeddings.