Service data loss function positive and negative example sampling for web 3.0 and re-recommendation method
The service data loss function positive and negative example sampling and re-recommendation method for Web 3.0 addresses the challenge of balancing diversity and correlation in service recommendations by employing a novel sampling and optimization approach, resulting in improved recommendation accuracy and diversity.
Patent Information
- Application Number
- JP2023216478
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2023-12-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-12-22
AI Technical Summary
Existing service recommendation technologies in the field of service computing struggle to balance the diversity and correlation of recommendation results, leading to suboptimal service discovery for developers amidst vast amounts of data.
A service data loss function positive and negative example sampling and re-recommendation method for Web 3.0, which uses a designed loss function positive and negative example sampling rule to sample service data, construct diversity sample pairs, and optimize the convergence process of the loss function, ultimately generating re-recommendation results using a determinantal point process.
This method effectively optimizes service recommendation models by enhancing the diversity and correlation of recommendation results, improving recommendation accuracy and diversity without relying on additional data, and efficiently rearranging re-recommendation results.
Smart Images

Figure 2025080204000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a service data loss function positive and negative example sampling and re-recommendation method for web 3.0. [Background technology]
[0002] The combination of service computing, big data, blockchain, mobility Internet and other technologies in different fields has promoted the entry of web technology into the "3.0 era". The update and replacement of web technology has brought about a large amount of data, further promoting the development of information technology, but it has also increased the amount of invalid resources and noise data on the Internet to a certain extent. Therefore, in the field of service computing, the challenge is how to quickly and accurately find API services that meet the needs of developers from a large amount of data.
[0003] In this context, service recommendation technology has attracted widespread attention. Conventional service recommendation technology uses collaborative filtering methods to calculate user similarity and service similarity to recommend services. For example, a service information recommendation device and method based on collaborative filtering algorithm, with patent number 202011122955.6, proposes a service recommendation method using a GRU neural network to screen potential people and combine Pearson correlation coefficients to calculate the similarity of potential people. A collaborative filtering method for web service recommendation, with patent number 201710211954.0, can estimate user preferences using service quality information and use a top-k algorithm to find API services that match the user's preferences. Guo Yanhong et al. of the Institute of System Engineering, Dalian University of Technology, propose a personalized improvement policy based on sparse matrices to limit the relationships between users with low similarity density, thereby mitigating the problem of data sparseness in the service recommendation process.
[0004] Most of the above recommendation methods realize service recommendation by modifying the recommendation algorithm or recommendation model, but do not make any innovative improvements to the service recommendation process. From the perspective of the recommendation process, the above recommendation methods all belong to single recommendation methods, and the diversity and correlation of the recommendation results need to be improved. Re-recommendation technology can optimize single recommendation results by using prior information such as service combination, service function, and service structure, and the generated re-recommendation results have high service recommendation quality. Re-recommendation technology has been widely applied in fields such as social networks, e-commerce sales, and streaming platforms, but in the field of service computing, there is still little research and application on re-recommendation at present. Summary of the Invention [Problem to be solved by the invention]
[0005] In order to overcome the shortcomings of the prior art, balance the diversity and correlation of service recommendation results, and improve the quality of service recommendation, the present invention provides a service data loss function positive and negative example sampling and re-recommendation method for web3.0, which is highly efficient and versatile, and which firstly uses a loss function positive and negative example sampling rule to sample service data and obtain loss function positive and negative examples, then uses the loss function positive and negative examples to construct diversity sample pairs and optimize the convergence process of the loss function, and finally generates re-recommendation results and sorts them using a method based on determinant point process. [Means for solving the problem]
[0006] The technical solutions used in the present invention are as follows:
[0007] A service data loss function positive and negative example sampling and re-recommendation method for web 3.0, In step 1, a loss function positive and negative example sampling rule is designed based on the functional domain and description document information of the API service in the service data. The process is as follows: 1.1 Service data: refers to the data used by the service recommendation model in the training process, including API services, functional areas, service combinations, call sequences, and description documents; 1.2 API service: Application Programming Interface (API), represented by the symbol a. 1.3 Functional field: The functional type and application field to which the API service belongs, represented by the symbol c, is the prior information of the API service; 1.4 Service Combination: Combining two or more API services to form a new API service, represented by the symbol m, is the prior information of the API service. 1.5 Call sequence: A sequence of API services accessed by a user, ordered by time, denoted by the symbol s. 1.6 Description document: This is text data that describes the functional area and interface information of the API service, represented by the symbol d, and is the advance information of the API service. 1.7 Positive and negative examples of loss function: In the convergence process of the loss function, the sampled service data is called the positive example of the loss function, and the unsampled service data is called the negative example of the loss function. + and - Each loss function positive example has multiple corresponding loss function negative examples, and the combination of the two is called a loss function positive / negative example. 1.8 Loss function positive and negative example sampling rule: Select the APIs directly called by users and indirectly called by service combinations from the service data as loss function positive examples. For other service data, first obtain the similarity using the loss function positive example similarity calculation method based on prior information, and then select the service data with high similarity as loss function negative examples; In step 2, according to the loss function positive and negative example sampling rule created in step 1.8, the service data is sampled for positive and negative loss function examples, and the number of loss function negative examples is limited using the similarity threshold ζ; In step 3, a diversity sample pair is constructed using the positive and negative examples of the loss function, and the convergence process of the loss function in the service recommendation model is optimized to obtain a diversity service recommendation result set; In step 4, the initial service recommendation result set and the diversified service recommendation result set are integrated to generate re-recommendation results, and finally, the re-recommendation results are re-sorted using the determinantal point process method.
[0008] Preferably, in the above 1.8, the calculation process of the loss function positive and negative example similarity based on prior information is specifically as follows: 1.8.1 Loss function Arbitrarily select one API service that is a positive example and + The functional area and description document to which each belongs are denoted by the symbol c + and d + It is expressed as 1.8.2 Arbitrarily select one API service from the other service data, denoted as a, and denote the functional area and description document to which it belongs by the symbols c and d, respectively; 1.8.3d + and d are input to the pre-trained language model, and the result is a + and a are the embedding vectors of a+ and the symbol e a Among them, the pre-trained language model is a commonly used deep learning method that can convert text data into word vectors, 1.8.4 e a+ and e a Calculate the cosine distance between: e a+ The transpose vector of and e a Multiply by and divide by the magnitude of both, and write the result as
number
number
number
number
number
number
number
number
[0009] Furthermore, the process of step 2 is as follows: 2.1 Loss function positive and negative example sampling: In the training process of the service recommendation model, the service data used in the loss function can be extracted and classified according to the loss function positive and negative example sampling rules, and can be divided into two processes: loss function positive example sampling and loss function negative example sampling. The processes are as described in 2.2 to 2.4. 2.2 Loss function positive example sampling: According to the loss function positive and negative example sampling rules, select loss function positive examples from the service data. Preferably, in 2.2, the process of loss function positive example sampling is as follows: 2.2.1 Loss function We define a set of positive examples and denote it by the symbol set + It is expressed as 2.2.2 Traverse the API services in the service data and find the API service taken for the i-th time as a i year, 2.2.3 Traverse the call sequence in the service data and define the jth call sequence as s j year, 2.2.4 s j ni a i If it contains, a i is called directly, and a i The loss function positive examples
number
number
[0010] 2.3 Loss function negative example sampling: According to the loss function positive and negative example sampling rules, select loss function negative examples from the service data. Preferably, in the above 2.3, the process of loss function negative example sampling is as follows: 2.3.1 Take all API services in the service data and construct set A; 2.3.2 Set A and loss function positive example set + The difference set is defined as the loss function negative example candidate set, denoted by preSet. 2.3.3 Loss function positive example set + , and the i-th loss function positive example is denoted by the symbol
number
number
number
number
number
number
number
[0011] 2.4 Limiting the number of loss function negative examples: after obtaining the loss function negative examples in 2.3, further limit the number of loss function negative examples by the similarity threshold ζ. Preferably, in 2.4, the limiting process of the number of loss function negative examples is as follows: 2.4.1 Define a similarity threshold ζ to be used to control the number of negative examples in the loss function, 2.4.2 Loss function negative example constraint set
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0012] Furthermore, the process of step 3 is as follows: 3.1 Diversity sample pair construction: A sample pair obtained by combining a loss function positive example and a loss function negative example after limiting the number can increase the diversity of service recommendation results compared with a general sample pair. Preferably, in 3.1, the process of constructing a diversity sample pair is as follows: 3.1.1 Define the diversity sample pair set multiSet, 3.1.2 Loss function positive example set + traverses the i-th loss function positive example
number
number
number
number
number
number
number
number
number
number
number
[0013] 3.2 Defining the initial service recommendation result set: Based on the user query, the service recommendation model gives a set of initial API service recommendations, denoted by R d The service recommendation model is a conventional model that can learn and mine service data and provide service recommendation results based on user queries. The model is constructed using methods based on collaborative filtering and neural networks. 3.3 Optimization of the loss function convergence process: First, obtain a diversity sample pair based on 3.1, then replace the original data used in the loss function in 3.1 with the diversity sample pair, and retrain the model obtained in 3.2; 3.4 Diversified service recommendation result set: After 3.3, we define the set of recommended API services given by the service recommendation model as the diversified service recommendation result set, denoted by R m It is expressed as:
[0014] Furthermore, the process of step 4 is as follows: 4.1 Re-recommendation result: In order to improve the recommendation quality, we optimize the service recommendation model, loss function, or data sample, and then perform two or more recommendations. The result is called the re-recommendation result, which is denoted by R. c It is expressed as 4.2 Integration of the initial service recommendation result set and the diversified recommendation result set: The initial service recommendation result set and the diversified recommendation result set are integrated according to the functional field and the service combination. In the above 4.2, the integration process is as follows: 4.2.1 Re-recommendation result set R c Define 4.2.2 Define the correlation hyperparameter θ, 4.2.3 Define the diversity hyperparameter η, 4.2.4 Input the user's query into the service recommendation model to generate the initial service recommendation result set R d Given 4.2.5 Using step 3, the diversified service recommendation result set R m Obtained, 4.2.6 R d traverses the i-th API service and returns a i year, 4.2.7 R m traverses the jth API service and returns the jth API service j year, 4.2.8a i and a j If a is the same, i R c and jump to 4.2.7, 4.2.9a i and a j When the functional areas of the two are different, i and c jLet M be the symbol i In a i represents the set of service combinations corresponding to M j In a j represents a set of service combinations corresponding to 4.2.10 c i and c j Calculate the Jaccard similarity coefficient between and and compare it with η. If the Jaccard similarity coefficient is smaller than η, jump to 4.2.7; 4.2.11 M i and M. j Calculate the Jaccard similarity coefficient between and θ, and compare it with θ. If the Jaccard similarity coefficient is smaller than θ, jump to 4.2.7; 4.2.12a i and a j R c Add to 4.2.13a j R m If this is the last API service in 4.2.14a i R d If this is the last API service in 4.2.15 R c and 4.3 Re-sorting of recommendation results: Using the score matrix and similarity matrix, we calculate the determinant point process to re-sort the recommendation results to balance the diversity and correlation of the recommendation results. Preferably, in 4.3, the process of rearranging the re-recommendation results using a determinantal point process is as follows: 4.3.1 Define the similarity matrix, whose row and column elements are the similarities of the recommendation services, and are denoted by the symbol Q. 4.3.2 Define a score matrix, whose row and column elements are the recommendation scores given by the service recommendation model, denoted by the symbol P. 4.3.3 Define the kernel matrix, denoted by the symbol L. The kernel matrix is obtained by performing matrix operations on the score matrix and the similarity matrix. The operation formula is L=P·Q·P, where the symbol “·” represents matrix multiplication and the symbol “=” represents matrix substitution. 4.3.4 Define the balance coefficient, which represents the degree of balance between diversity and correlation of re-recommendation results, and is denoted by the symbol σ. 4.3.5 Using 4.2, re-recommended result set R c We obtain the symbol
number
number
number
number
number
number
number
number
[0015] The beneficial effects of the present invention are as follows: (1) It is possible to optimize a service recommendation model by performing loss function positive and negative example sampling based on existing service data without relying on additional data; (2) It is highly efficient and the amount of calculation required to rearrange the re-recommendation results using a determinantal point process is small; (3) It has good versatility and can be used in hybrid with conventional service recommendation models to generate re-recommendation results that are both diverse and correlated. [Brief description of the drawings]
[0016]
Figure 1
Figure 2
[0017] The present invention will now be further described in conjunction with the following specification, drawings and examples.
[0018] Referring to FIG. 1 and FIG. 2, a service data loss function positive and negative example sampling and re-recommendation method for web 3.0, In step 1, a loss function positive and negative example sampling rule is designed based on the functional domain and description document information of the API service in the service data. The process is as follows: 1.1 Service data: refers to the data used by the service recommendation model in the training process, including API services, functional areas, service combinations, call sequences, and description documents; 1.2 API service: Application Programming Interface (API), represented by the symbol a. 1.3 Functional field: The functional type and application field to which the API service belongs, represented by the symbol c, is the prior information of the API service; 1.4 Service Combination: Combining two or more API services to form a new API service, represented by the symbol m, is the prior information of the API service. 1.5 Call sequence: A sequence of API services accessed by a user, ordered by time, denoted by the symbol s. 1.6 Description document: This is text data that describes the functional area and interface information of the API service, represented by the symbol d, and is the advance information of the API service. 1.7 Positive and negative examples of loss function: In the convergence process of the loss function, the sampled service data is called the positive example of the loss function, and the unsampled service data is called the negative example of the loss function. + and - Each loss function positive example has multiple corresponding loss function negative examples, and the combination of the two is called a loss function positive / negative example. 1.8 Loss function positive and negative example sampling rule: The APIs directly called by the user and indirectly called by the service combination are selected as loss function positive examples from the service data, and for other service data, the similarity is first obtained using the loss function positive example similarity calculation method based on prior information, and then the service data with high similarity is selected as loss function negative examples. In the above 1.8, the calculation process of the loss function positive and negative example similarity based on prior information is as follows: 1.8.1 Loss function Arbitrarily select one API service that is a positive example and + The functional area and description document to which each belongs are denoted by the symbol c + and d + It is expressed as 1.8.2 Arbitrarily select one API service from the other service data, denoted as a, and denote the functional area and description document to which it belongs by the symbols c and d, respectively; 1.8.3d + and d are input to the pre-trained language model, and the result is a + and a are the embedding vectors of a+ and the symbol e a The pre-trained language model is a commonly used deep learning method that can convert text data into word vectors, for example, the BERT model. 1.8.4 e a+ and e a Calculate the cosine distance between: e a+ The transpose vector of and e a Multiply by and divide by the magnitude of both, and write the result as
number
number
number
number
number
number
number
number
[0019] In step 2, according to the loss function positive and negative example sampling rule created in step 1.8, the service data is sampled for loss function positive and negative examples, and the number of loss function negative examples is limited using the similarity threshold ζ. The process is as follows: 2.1 Loss function positive and negative example sampling: In the training process of the service recommendation model, the service data used in the loss function can be extracted and classified according to the loss function positive and negative example sampling rules, and can be divided into two processes: loss function positive example sampling and loss function negative example sampling. The specific steps are as described in 2.2 to 2.4. 2.2 Loss function positive example sampling: According to the loss function positive and negative example sampling rules, select loss function positive examples from the service data. In the above 2.2, the process of loss function positive example sampling is as follows: 2.2.1 Loss function We define a set of positive examples and denote it by the symbol set + It is expressed as 2.2.2 Traverse the API services in the service data and find the API service taken for the i-th time as a i year, 2.2.3 Traverse the call sequence in the service data and define the jth call sequence as s j year, 2.2.4 s j ni a i If it contains, a i is called directly, and a i The loss function positive examples
number
number
[0020] 2.3 Loss function negative example sampling: According to the loss function positive and negative example sampling rules, select loss function negative examples from the service data. In the above 2.3, the process of loss function negative example sampling is as follows: 2.3.1 Take all API services in the service data and construct set A; 2.3.2 Set A and loss function positive example set + The difference set is defined as the loss function negative example candidate set, denoted by preSet. 2.3.3 Loss function positive example set + , and the i-th loss function positive example is denoted by the symbol
number
number
number
number
number
number
number
[0021] 2.4 Limiting the number of negative examples of the loss function: After obtaining the negative examples of the loss function in 2.3, the number of negative examples of the loss function is further limited by the similarity threshold ζ. In 2.4, the limiting process of the number of negative examples of the loss function is as follows: 2.4.1 Define a similarity threshold ζ to be used to control the number of negative examples in the loss function, 2.4.2 Loss function negative example constraint set
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0022] In step 3, a diversified sample pair is constructed using the positive and negative examples of the loss function, and the convergence process of the loss function in the service recommendation model is optimized to obtain a diversified service recommendation result set. The process is shown in Figure 1(a) and Figure 1(a). The process is as follows: 3.1 Diversity sample pair construction: The sample pair obtained by combining the loss function positive example and the loss function negative example after limiting the number can increase the diversity of the service recommendation results compared to the general sample pair. In the above 3.1, the process of constructing the diversity sample pair is as follows: 3.1.1 Define the diversity sample pair set multiSet, 3.1.2 Loss function positive example set + traverses the i-th loss function positive example
number
number
number
number
number
number
number
number
number
number
number
[0023] 3.2 Defining the initial service recommendation result set: Based on the user query, the service recommendation model gives a set of initial API service recommendations, denoted by R d In this case, the service recommendation model refers to a model that can learn and mine service data and provide service recommendation results based on user queries. Currently, in this research field, there are multiple recommendation models to choose from, such as a service recommendation model based on clustering, a service recommendation model based on theme modeling, etc.
[0024] 3.3 Optimization of the loss function convergence process: First, obtain a diversity sample pair based on 3.1, then replace the original data used in the loss function in 3.1 with the diversity sample pair, and retrain the model obtained in 3.2. The pair-wise Bayesian Personalized Ranking loss function (BPR) can be used as the object to be optimized. 3.4 Diversified service recommendation result set: After 3.3, we define the set of recommended API services given by the service recommendation model as the diversified service recommendation result set, denoted by R m It is expressed as Step 4: The initial service recommendation result set and the diversified service recommendation result set are integrated to generate re-recommendation results, and finally, the re-recommendation results are re-sorted using the determinantal point process method, as shown in Figure 1(c), and the process is as follows: 4.1 Re-recommendation result: In order to improve the recommendation quality, we optimize the service recommendation model, loss function, or data sample, and then perform two or more recommendations. The result is called the re-recommendation result, which is denoted by R. c It is expressed as 4.2 Integration of the initial service recommendation result set and the diversified recommendation result set: The initial service recommendation result set and the diversified recommendation result set are integrated according to the functional field and the service combination. In the above 4.2, the integration process is as follows: 4.2.1 Re-recommendation result set R c Define 4.2.2 Define the correlation hyperparameter θ, 4.2.3 Define the diversity hyperparameter η, 4.2.4 Input the user's query into the service recommendation model to generate the initial service recommendation result set R d Given 4.2.5 Using step 3, the diversified service recommendation result set R m Obtained, 4.2.6 R d traverses the i-th API service and returns a i year, 4.2.7 R m traverses the jth API service and returns the jth API service j year, 4.2.8a i and a j If a is the same, i R c and jump to 4.2.7, 4.2.9a i and a j When the functional areas of the two are different, i and c j Let M be the symbol i In a i represents the set of service combinations corresponding to M j In a j represents a set of service combinations corresponding to 4.2.10 c i and c j Calculate the Jaccard similarity coefficient between and and compare it with η. If the Jaccard similarity coefficient is smaller than η, jump to 4.2.7; 4.2.11 M i and M. j Calculate the Jaccard similarity coefficient between and θ, and compare it with θ. If the Jaccard similarity coefficient is smaller than θ, jump to 4.2.7; 4.2.12a i and a j R c Add to 4.2.13a j R m If this is the last API service in 4.2.14a i R d If this is the last API service in 4.2.15 R c and 4.3 Re-sorting of re-recommendation results: Calculate the determinant point process using the score matrix and the similarity matrix, and re-sort the re-recommendation results to balance the diversity and correlation of the recommendation results. In the above 4.3, the process of re-sorting the re-recommendation results using the determinant point process is as follows: 4.3.1 Define the similarity matrix, whose row and column elements are the similarities of the recommendation services, and are denoted by the symbol Q. 4.3.2 Define a score matrix, whose row and column elements are the recommendation scores given by the service recommendation model, denoted by the symbol P. 4.3.3 Define the kernel matrix, denoted by the symbol L. The kernel matrix is obtained by performing matrix operations on the score matrix and the similarity matrix. The operation formula is L=P·Q·P, where the symbol “·” represents matrix multiplication and the symbol “=” represents matrix substitution. 4.3.4 Define the balance coefficient, which represents the degree of balance between diversity and correlation of re-recommendation results, and is denoted by the symbol σ. 4.3.5 Using 4.2, re-recommended result set R c We obtain the symbol
number
number
number
number
number
number
number
number
[0025] Below, we will analyze the actual effects of the invention based on specific service data.
[0026] 1. The service data used includes 1423 service combinations, 1032 RESTful services, and 324 pairs of service call sequences, and is obtained from the Internet based on crawler technology.
[0027] 2. The hyperparameters in this method are set as follows: Embedding vector e a+ and e aThe feature dimension of is set to 64, the similarity threshold ζ is selected from [0.2, 0.25, 0.3, 0.35, 0.4], and the correlation hyperparameter θ and diversity hyperparameter η are set to 0.1 and 0.3, respectively.
[0028] 3. Four representative service recommendation methods are selected as comparison methods.
[0029] 3.1 service-KNN method: A typical service recommendation method based on similarity.
[0030] 3.2 MMR Method: Maximum Boundary Correlation Method.
[0031] 3.3 AFM model: a service recommendation model based on factorization machine.
[0032] 3.4 NGCF model: A service recommendation method based on collaborative filtering.
[0033] 4. Hit rate is used to evaluate the accuracy of service recommendation, and the calculation formula is as follows:
number
[0034] 5. The average distance within the list is used to indicate the degree of diversification of the service recommendation results, and the calculation formula is as follows:
number
number
[0035] 6. Randomly select 80% of the service data for training, and the other 20% of the data is used as the test sample. The experimental effect is shown in Figure 2.
[0036] 7. As can be seen from analyzing Fig. 2, the present invention has significantly higher recommendation accuracy and recommendation diversity than the service-KNN method, the AFM model, and the NGCF model. The recommendation diversity of the MMR method is slightly higher than that of the present invention, but the present invention has a greater advantage in recommendation accuracy.
[0037] The contents described in the embodiments of this specification are merely enumerations of the implementation forms of the concept of the invention and are merely for the purpose of describing applications. The scope of protection of the present invention is not limited to the specific forms described in the embodiments, and the scope of protection of the present invention also extends to equivalent technical means that a person skilled in the art can think of based on the concept of the present invention.
[0038] (Additional Note) (Appendix 1) A service data loss function positive and negative example sampling and re-recommendation method for web 3.0, In step 1, a loss function positive and negative example sampling rule is designed based on the functional domain and description document information of the API service in the service data. The process is as follows: 1.1 Service data: refers to the data used by the service recommendation model in the training process, including API services, functional areas, service combinations, call sequences, and description documents; 1.2 API Service: Application Program Interface API, represented by the symbol a; 1.3 Functional field: The functional type and application field to which the API service belongs, represented by the symbol c, is the prior information of the API service; 1.4 Service Combination: Combining two or more API services to form a new API service, represented by the symbol m, is the prior information of the API service; 1.5 Call sequence: A sequence of API services accessed by a user, ordered by time, denoted by the symbol s. 1.6 Description document: This is text data that describes the functional area and interface information of the API service, represented by the symbol d, and is the advance information of the API service. 1.7 Positive and negative examples of loss function: In the convergence process of the loss function, the sampled service data is called the positive example of the loss function, and the unsampled service data is called the negative example of the loss function. + and - Each loss function positive example has multiple corresponding loss function negative examples, and the combination of the two is called a loss function positive / negative example. 1.8 Loss function positive and negative example sampling rule: Select the APIs directly called by users and indirectly called by service combinations from the service data as loss function positive examples; for other service data, first obtain the similarity using the loss function positive example similarity calculation method based on prior information, and then select the service data with high similarity as loss function negative examples; In step 2, according to the loss function positive and negative example sampling rule created in step 1.8, the service data is sampled for positive and negative loss function examples, and the number of loss function negative examples is limited using the similarity threshold ζ; In step 3, a diversity sample pair is constructed using the positive and negative examples of the loss function, and the convergence process of the loss function in the service recommendation model is optimized to obtain a diversity service recommendation result set; In step 4, the initial service recommendation result set and the diversified service recommendation result set are integrated to generate re-recommendation results, and finally, the re-recommendation results are re-sorted by using a determinantal point process method.
[0039] (Appendix 2) In the above 1.8, the calculation process of the loss function positive and negative example similarity based on prior information is as follows: 1.8.1 Loss function Arbitrarily select one API service that is a positive example and + The functional area and description document to which each belongs are denoted by the symbol c + and d + It is expressed as 1.8.2 Arbitrarily select one API service from the other service data, denoted as a, and denote the functional area and description document to which it belongs by the symbols c and d, respectively; 1.8.3d + and d are input to the pre-trained language model, and the result is a + and a are the embedding vectors of a+ and the symbol e a Among them, the pre-trained language model is a commonly used deep learning method that can convert text data into word vectors, 1.8.4 e a+ and e a Calculate the cosine distance between: e a+ The transpose vector of and e a Multiply by and divide by the magnitude of both, and write the result as
number
number
number
number
number
number
number
number
[0040] (Appendix 3) The process of step 2 is as follows: 2.1 Loss function positive and negative example sampling: In the training process of the service recommendation model, the service data used in the loss function can be extracted and classified according to the loss function positive and negative example sampling rules, and can be divided into two processes: loss function positive example sampling and loss function negative example sampling. 2.2 Loss function positive example sampling: According to the loss function positive and negative example sampling rules, select loss function positive examples from the service data, 2.3 Loss function negative example sampling: According to the loss function positive and negative example sampling rules, select loss function negative examples from the service data; 2.4 Limiting the number of loss function negative examples: After performing 2.3 and obtaining the loss function negative examples, the method for Web 3.0 service data loss function positive and negative example sampling and re-recommendation described in Appendix 1 or 2 is characterized in that the number of loss function negative examples is further limited by the similarity threshold ζ.
[0041] (Appendix 4) In 2.2 above, the process of loss function positive example sampling is as follows: 2.2.1 Loss function We define a set of positive examples and denote it by the symbol set + It is expressed as 2.2.2 Traverse the API services in the service data and find the API service taken for the i-th time as a i year, 2.2.3 Traverse the call sequence in the service data and define the jth call sequence as s j year, 2.2.4 s j ni a i If it contains, a i is called directly, and a i The loss function positive examples
number
number
[0042] (Appendix 5) In 2.3 above, the process of loss function negative example sampling is as follows: 2.3.1 Take all API services in the service data and construct set A; 2.3.2 Set A and loss function positive example set + The difference set is defined as the loss function negative example candidate set, denoted by preSet. 2.3.3 Loss function positive example set + , and the i-th loss function positive example is denoted by the symbol
number
number
number
number
number
number
number
[0043] (Appendix 6) In 2.4 above, the process of limiting the number of negative examples of the loss function is as follows: 2.4.1 Define a similarity threshold ζ to be used to control the number of negative examples in the loss function, 2.4.2 Loss function negative example constraint set
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0044] (Appendix 7) The process of step 3 is as follows: 3.1 Construction of Diverse Sample Pairs: The sample pairs obtained by combining the loss function positive examples and the loss function negative examples after limiting the number of them can increase the diversity of service recommendation results compared to general sample pairs. 3.2 Defining the initial service recommendation result set: Based on the user query, the service recommendation model gives a set of initial API service recommendations, denoted by R d The service recommendation model is a conventional model that can learn and mine service data and provide service recommendation results based on user queries. The model is constructed using methods based on collaborative filtering and neural networks. 3.3 Optimization of the loss function convergence process: First, obtain a diversity sample pair based on 3.1, then replace the original data used in the loss function in 3.1 with the diversity sample pair, and retrain the model obtained in 3.2; 3.4 Diversified service recommendation result set: After 3.3, we define the set of recommended API services given by the service recommendation model as the diversified service recommendation result set, denoted by Rm The service data loss function positive and negative example sampling and re-recommendation method for web 3.0 described in appendix 1 or 2, characterized in that:
[0045] (Appendix 8) In the above 3.1, the process of constructing a diversity sample pair is as follows: 3.1.1 Define the diversity sample pair set multiSet, 3.1.2 Loss function positive example set + traverses the i-th loss function positive example
number
number
number
number
number
number
number
number
number
number
number
[0046] (Appendix 9) The process of step 4 is as follows: 4.1 Re-recommendation result: In order to improve the recommendation quality, we optimize the service recommendation model, loss function, or data sample, and then perform two or more recommendations. The result is called the re-recommendation result, which is denoted by R. c It is expressed as 4.2 Integration of the initial service recommendation result set and the diversified recommendation result set: The initial service recommendation result set and the diversified recommendation result set are integrated according to the functional field and the service combination. In the above 4.2, the integration process is as follows: 4.2.1 Re-recommendation result set R c Define 4.2.2 Define the correlation hyperparameter θ, 4.2.3 Define the diversity hyperparameter η, 4.2.4 Input the user's query into the service recommendation model to generate the initial service recommendation result set R d Given 4.2.5 Using step 3, the diversified service recommendation result set R m Obtained, 4.2.6 R d traverses the i-th API service and returns a i year, 4.2.7 R m traverses the jth API service and returns the jth API service j year, 4.2.8a i and a j If a is the same, i R c and jump to 4.2.7, 4.2.9a i and a j When the functional areas of the two are different, i and c j Let M be the symbol i In a i represents the set of service combinations corresponding to M j In a j represents a set of service combinations corresponding to 4.2.10 c i and c j Calculate the Jaccard similarity coefficient between and and compare it with η. If the Jaccard similarity coefficient is smaller than η, jump to 4.2.7; 4.2.11 M i and M. j Calculate the Jaccard similarity coefficient between and θ, and compare it with θ. If the Jaccard similarity coefficient is smaller than θ, jump to 4.2.7; 4.2.12a i and a j R c Add to 4.2.13a j R m If this is the last API service in 4.2.14a i R d If this is the last API service in 4.2.15 R c and 4.3 Rearrangement of the recommendation results: Calculating a determinantal point process using a score matrix and a similarity matrix, and rearranging the recommendation results to balance the diversity and correlation of the recommendation results, which is characterized by the web3.0 service data loss function positive and negative example sampling and recommendation method described in Appendix 1 or 2.
[0047] (Appendix 10) In the above 4.3, the process of rearranging the recommendation results using the determinantal point process is as follows: 4.3.1 Define a similarity matrix, where both the row and column elements are the similarity of the recommendation services, represented by the symbol Q. 4.3.2 Define a score matrix, where both the row and column elements are the recommendation scores given by the service recommendation model, represented by the symbol P. 4.3.3 Define a kernel matrix, represented by the symbol L. The kernel matrix is obtained by performing matrix operations on the score matrix and the similarity matrix, and the formula is L = P·Q·P. The "·" symbol represents matrix multiplication, and the "=" symbol represents matrix substitution. 4.3.4 Define a balance coefficient, which represents the degree of balance between the diversity and correlation of the recommendation results, represented by the symbol σ. 4.3.5 Use 4.2 to obtain the recommendation result set R c and, the symbol
Number
Number
Number
number
number
number
number
number
Claims
1. A service data loss function positive and negative example sampling and re-recommendation method for web 3.0, In step 1, a loss function positive and negative example sampling rule is designed based on the functional domain and description document information of the API service in the service data, and the process is as follows: 1.1 Service data: Data used by the service recommendation model in the training process, including API services, functional areas, service combinations, call sequences, and description documents; 1.2 API service: an application program interface API, represented by the symbol a; 1.3 Functional field: The functional type and application field to which the API service belongs, represented by the symbol c, which is the prior information of the API service; 1.4 Service Combination: Combining two or more API services to form a new API service, represented by the symbol m, is the prior information of the API service; 1.5 Call sequence: A sequence of API services accessed by a user, ordered by time, represented by the symbol s. 1.6 Description document: Text data describing the functional field and interface information of an API service, represented by the symbol d, which is advance information of the API service; 1.7 Positive and negative examples of loss function: In the convergence process of the loss function, the sampled service data is called the positive example of the loss function, and the unsampled service data is called the negative example of the loss function. + And - Each loss function positive example has multiple corresponding loss function negative examples, and the combination of the two is called a loss function positive / negative example. 1.8 Loss function positive and negative example sampling rule: Select the APIs directly called by users and indirectly called by service combinations from the service data as loss function positive examples; for other service data, first obtain the similarity using the loss function positive example similarity calculation method based on prior information, and then select the service data with high similarity as loss function negative examples; In step 2, sampling of the service data is performed according to the loss function positive and negative example sampling rule created in step 1.8, and the number of loss function negative examples is limited using a similarity threshold value ζ; In step 3, a diversity sample pair is constructed using the positive and negative examples of the loss function, and the convergence process of the loss function in the service recommendation model is optimized to obtain a diversity service recommendation result set; In step 4, the initial service recommendation result set and the diversified service recommendation result set are integrated to generate re-recommendation results, and finally, the re-recommendation results are re-sorted by using a determinantal point process method.
2. In the above 1.8, the calculation process of the loss function positive and negative example similarity based on prior information is as follows: 1.8.1 Loss Function Arbitrarily select one API service that is a positive example, and + The functional area and description document to which each belongs are denoted by c + and + It is expressed as 1.8.2 Arbitrarily select one API service from the other service data, denoted as a, and denote the functional area and description document to which it belongs by the symbols c and d, respectively; 1.8.3d + and d are input to the pre-trained language model, and the obtained results are a + and a are the embedding vectors of a+ and the symbol e a Among them, the pre-trained language model is a commonly used deep learning method that can convert text data into word vectors, 1.8.4e a+ and a Calculating the cosine distance between e a+ The transpose vector of e a Multiply by and divide by the magnitude of both, and write the result as [0010] It is expressed as 1.8.5 c + Calculating the Jaccard similarity coefficient between and c: First, + After adding the size of and the size of c, + Subtract the size of the intersection of and c, set the result as temp, and then + Divide the size of the intersection of and c by temp, and [0025] year, 1.8.6 Linearly transform the Jaccard similarity coefficient of the cosine distance, restrict the result to the range of 0 to 1, and use the symbol [0030] and symbol [0045] A linear transformation is a mathematical operation that has the property that it preserves the addition of vectors and the multiplication of numbers. 1.8.7 [0050] and [006] Compare the size of + and a are the similarity coefficients, and the symbol [0070] It is expressed as 1.8.8 [0080] The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 1, further comprising:
3. The process of step 2 is as follows: 2.1 Loss function positive and negative example sampling: In the training process of the service recommendation model, the service data used in the loss function can be extracted and classified according to the loss function positive and negative example sampling rules, and can be divided into two processes: loss function positive example sampling and loss function negative example sampling. 2.2 Loss function positive example sampling: Select loss function positive examples from the service data according to the loss function positive and negative example sampling rules; 2.3 Loss function negative example sampling: Select loss function negative examples from the service data according to the loss function positive and negative example sampling rules; 2.4 Limiting the number of loss function negative examples: after obtaining the loss function negative examples in accordance with 2.3, further limit the number of loss function negative examples according to similarity threshold ζ.
4. In the above 2.2, the process of loss function positive example sampling is as follows: 2.2.1 Loss function We define a positive example set and use the symbol set + It is expressed as 2.2.2 Traverse the API services in the service data, and the API service taken for the i-th time is i year, 2.2.3 Traverse the call sequence in the service data and define the call sequence taken at the jth time as s j year, 2.2.4 s j Nia i When i is called directly, and a i The loss function positive examples [0090] And set it + After adding it, jump to 2.2.2, 2.2.5 s j Terminate the traversal if is the last call sequence in the service data; 2.2.6 Traverse the service combinations in the service data, and define the service combination taken at the kth time as m k year, 2.2.7 m k Nia i When i is called indirectly, and a i The loss function positive examples [0089] And set it + After adding it, jump to 2.2.2, 2.2.8 m k Terminate the traversal if is the last call sequence in the service data; 2.2.9a i If is the last API service in the service data, terminate the traversal; 2.2.10 Loss function positive example set + The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 3, further comprising:
5. In 2.3 above, the process of loss function negative example sampling is as follows: 2.3.1 Take all API services in the service data and construct a set A; 2.3.2 Set A and loss function positive example set + The difference set is defined as the loss function negative example candidate set, and is represented by the symbol preSet. 2.3.3 Loss function positive example set + , and the i-th loss function positive example is denoted by the symbol ##EQU00011## It is expressed as 2.3.4 ##EQU00012## We define the loss function negative example set corresponding to ##EQU00013## It is expressed as 2.3.5 Traverse the preSet and denote the jth loss function negative example candidate by the symbol ps j It is expressed as 2.3.6 According to the loss function positive and negative example similarity calculation method based on prior information defined in step 1.8, p j and ##EQU14## The similarity between ##EQU00015## It is expressed as 2.3.7 ps j of ##EQU00016## Add to 2.3.8 ps j If ,is the last loss function negative example candidate example in the preSet, the traversal is terminated; 2.3.9 ##EQU00017## is set + The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 3, characterized in that, if the last one in is the last one in, the traversal is terminated.
6. In the above 2.4, the process of limiting the number of loss function negative examples is as follows: 2.4.1 Define a similarity threshold ζ to be used to control the number of negative examples in the loss function, 2.4.2 Loss function negative example constraint set [0018] Define 2.4.3 Loss function positive example set + traverses the i-th loss function positive example [0019] Then, the corresponding loss function negative example set is [0020] year, 2.4.4 ##EQU00021## The n loss function negative examples with the highest similarity in - Select as 2.4.5 N - traverses the jth loss function negative example [0022] year, 2.4.6 [0023] and [0024] Similarity to [0025] year, 2.4.7 [0026] Compare the size of and ζ, [0027] When is large, [0028] and hold it [0029] In addition, when ζ is large, [0030] Throw away, 2.4.8 [0031] N - If the last loss function negative example in 2.4.9 [0032] is set + If it is the last loss function positive example in 2.4.10 [Equation 33] The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 3, further comprising:
7. The process of step 3 is as follows: 3.1 Diverse sample pair construction: The sample pairs obtained by combining the loss function positive examples and the loss function negative examples after limiting the number of them can increase the diversity of service recommendation results compared to general sample pairs. 3.2 Definition of initial service recommendation result set: Based on the user's query, the service recommendation model gives a set of initial API service recommendations, denoted by R d The service recommendation model is a conventional model that can learn and mine service data and provide service recommendation results based on user queries. The model is constructed using methods based on collaborative filtering and neural networks. 3.3 Optimization of the loss function convergence process: Firstly, obtain a diversity sample pair based on 3.1, then replace the original data used in the loss function in 3.1 with the diversity sample pair, and retrain the model obtained in 3.2; 3.4 Diversified service recommendation result set: After 3.3, we define the set of recommendation API services given by the service recommendation model as a diversified service recommendation result set, denoted by R m The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 1 or 2, characterized in that:
8. In the above 3.1, the process of constructing a diversity sample pair is as follows: 3.1.1 Define a diversity sample pair set multiSet, 3.1.2 Loss function positive example set + traverses the i-th loss function positive example [0034] Then, the corresponding loss function negative example set after limiting the number is 【Number 35】 year, 3.1.3 [0036] Randomly select any number of loss function negative examples from 3.1.4 Traverse the randomly selected loss function negative examples and determine the loss function negative example taken at the jth iteration. [Equation 37] year, 3.1.5 [Equation 38] and [0039] We construct diversified sample pairs with the symbol [0040] The symbols "<" and ">" are [0041] and [0042] is used to represent the bigram relationship between 3.1.6 [0043] Add to multiSet, 3.1.7 [0044] is set + If it is the last loss function positive example in 3.1.8 The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 7, characterized in that it outputs a multiSet.
9. The process of step 4 is as follows: 4.1 Re-recommendation result: In order to improve the recommendation quality, after optimizing the service recommendation model, loss function, or data sample, we make two or more recommendations. The result is called the re-recommendation result, which is denoted by R. c It is expressed as 4.2 Integration of the initial service recommendation result set and the diversified recommendation result set: The initial service recommendation result set and the diversified recommendation result set are integrated according to the function field and the service combination. In the above 4.2, the integration process is as follows: 4.2.1 Re-recommended result set R c Define 4.2.2 Define the correlation hyperparameter θ, 4.2.3 Define the diversity hyperparameter η, 4.2.4 Input the user's query into the service recommendation model to obtain the initial service recommendation result set R d Given 4.2.5 Using step 3, the diversified service recommendation result set R m Obtained, 4.2.6 R d , and the API service taken for the i-th time is i year, 4.2.7 R m traverses the j-th API service and j year, 4.2.8a i and j is the same, then a i R c Add to and jump to 4.2.7, 4.2.9a i and j When the functional areas of the two are different, i and j Let M i So i represents a set of service combinations corresponding to M j So j represents a set of service combinations corresponding to 4.2.10c i and j Calculate the Jaccard similarity coefficient between and η, and compare it with η. If the Jaccard similarity coefficient is smaller than η, jump to 4.2.7; 4.2.11 M i and M. j Calculate the Jaccard similarity coefficient between and θ, and compare it with θ. If the Jaccard similarity coefficient is smaller than θ, jump to 4.2.7; 4.2.12a i and j R c Add to 4.2.13a j R m If it is the last API service in 4.2.14a i R d If it is the last API service in 4.2.15 R c and 4.3 Re-sorting of re-recommendation results: The method for the service data loss function positive and negative example sampling and re-recommendation for Web 3.0 described in claim 1 or 2, characterized in that, the method uses the score matrix and the similarity matrix to calculate the determinant point process, and re-sorts the re-recommendation results to balance the diversity and correlation of the recommendation results.
10. In 4.3, the process of rearranging the re-recommendation results using a determinantal point process is as follows: 4.3.1 Define a similarity matrix, whose row and column elements are the similarities of the recommended services, and are represented by the symbol Q. 4.3.2 Define a score matrix, whose row and column elements are the recommendation scores given by the service recommendation model, denoted by the symbol P; 4.3.3 Define the kernel matrix, denoted by the symbol L. The kernel matrix is obtained by performing a matrix operation on the score matrix and the similarity matrix. The operation formula is L=P·Q·P, where the symbol "·" represents matrix multiplication and the symbol "=" represents matrix substitution. 4.3.4 Define the balance coefficient, which represents the degree of balance between diversity and correlation of re-recommendation results, and is denoted by the symbol σ; 4.3.5 Using 4.2, re-recommended result set R c We obtain the symbol [0045] In the core matrix, R c represents the sub-matrix indexed by 4.3.6 Kernel Matrix, R c Traverse all submatrices indexed by and take the result taken for the i-th time. [0046] year, 4.3.7 [0047] Perform a determinant calculation on and take the logarithm of the result, [0048] where the symbol det denotes the determinant calculation, which is a basic matrix operation, the product of all elements in a matrix and the sum of the corresponding cofactors, and the symbol log denotes the logarithm operation. 4.3.8 [0049] and add σ, 4.3.9 [Number 50] Terminate the traversal if is the last submatrix; 4.3.10 Perform a determinantal calculation on Q and take the logarithm of the result; [0051] year, 4.3.11 From σ [0052] Subtract 4.3.12 Perform maximum a posteriori estimation for σ, i.e., R c The re-recommendation results are screened from the above, the value of σ is maximized, and the symbol R f represents the re-recommendation result after screening, 4.3.13 R f The method for web 3.0 service data loss function positive and negative example sampling and re-recommendation according to claim 9, further comprising:
Citation Information
Patent Citations
Graph embedding enhanced Web API (Application Program Interface) recommendation method and system
CN114817745A
Application interface determination method and device, medium and equipment
CN116049273A
Web API recommendation method and device based on functional semantics and structure interaction
CN116628328A
Cited By
Efficient lightweight federal recommendation method
CN121598334A