A cross-library API recommendation method during code migration

By crawling multi-source information and using deep learning models to vectorize the similarity calculation of the API, combined with graph model and theme model, the problem that existing API recommendation methods are difficult to adapt to multi-store scenarios in code migration is solved, and more efficient and accurate API recommendation effects are achieved.

CN117235138BActive Publication Date: 2025-06-06ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311039688.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-06-06
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

The existing API recommendation method is difficult to adapt to the scenarios of two or more software libraries involved in code migration, and it is impossible to effectively recommend the mapping of one software library's API to another software library.

Method used

A cross-library API recommendation method is adopted to form metadata by crawling multi-source information, and vectorized similarity calculations are used to calculate the API using ALBERT and Word2Vec models. Combining open source project and question-and-answer community data, a graph model and theme model of API usage patterns are built, multiple similarities are calculated and weighted, and the target API is finally sorted.

Benefits of technology

It improves the accuracy and efficiency of API recommendations during code migration, reduces the time and error risk for developers to learn new APIs and match APIs similar to each other, and ensures that the system runs normally after migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235138B_ABST
    Figure CN117235138B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-library API recommendation method in a code migration process, comprising: (1) using a crawler framework to crawl the official document information, open source projects and question-and-answer community data of a source software library and a migration target software library; (2) obtaining the API document similarity S1 of the source software library and the migration target software library through the official document information; (3) obtaining the API code snippet similarity S2 through the open source project; (4) obtaining the topic similarity S3 between the two APIs through the question-and-answer community data; (5) passing the similarities S1, S2 and S3 through a weight matrix W to obtain the final API similarity S; (6) inputting an API of the source software library, calculating the final similarity S of the API of the source software library and each API in the migration target software library, and sorting all APIs from large to small according to S for recommendation. The present invention can combine multi-source information to learn the deep feature representation of the API, thereby providing more effective information for the recommendation task and improving the effect of API recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent software migration, and in particular relates to a cross-library API recommendation method in a code migration process. Background Art

[0002] With the continuous development of Internet technology and the further improvement of information construction, the application scope of various software is also expanding, and it has been widely used and promoted in the fields of economy, education, medical care, transportation, etc. These software are not just simple tools, but also catalysts. While changing people's work and lifestyle, they also profoundly affect the social operation mode and industrial structure. At the same time, the demand for software has also increased significantly, and new software models and development models have continued to emerge, and their scale and number are expanding and expanding at an alarming rate. The emerging software has become an important engine for promoting the development of the digital economy. In the future, with the continuous advancement and promotion of technology, software will be more deeply integrated into various industries and fields, providing people with more efficient, intelligent, and high-quality services and experiences, and further promoting social development and progress.

[0003] Code migration is the process of transferring an existing software system or application from one platform or environment to another. In modern software development, code migration is one of the important ways to develop software quickly. Code migration can achieve the reuse of software components, reduce the workload of software development and testing, and improve software quality and reliability. With the popularity of mobile devices and cloud computing technology, developers need to migrate software applications to different platforms to meet diverse user needs. Code migration can help developers transfer existing code bases to the target platform, thereby speeding up the development and deployment of applications. Not only that, when an enterprise acquires or merges with other companies, it needs to integrate the new code base with the existing code base. Code migration can help enterprises merge the new code base with the existing code base and ensure the functional integrity and stability of the code. Moreover, with the advancement of technology and changes in market demand, software systems need to be continuously upgraded and updated to meet the growing user needs. Code migration can help developers transfer the existing code base to the new development environment to achieve the upgrade and expansion of software system functions. In short, code migration is an important way to develop software quickly. By transferring the existing code base to the target platform, software development efficiency can be improved, development costs can be reduced, and software quality and user experience can be improved.

[0004] In the process of code migration, cross-library API recommendation is of great significance. The scenario of code migration is often that a developer who is familiar with a certain development programming language needs to migrate to another programming language or programming framework that he is not familiar with. To successfully migrate the code to a new programming language or programming framework, implement the original code logic, and reproduce the original code function, the developer needs to be very familiar with both software libraries. Using the migration-oriented cross-library API recommendation method, developers can quickly and accurately find the target software library API based on the API of the original software library, without having to spend time and effort to learn new APIs or spend a lot of time matching APIs with similar functions. This can greatly improve the efficiency of migration and make the migration project progress faster. Moreover, the migration-oriented cross-library API recommendation method can accurately match the target API, thereby avoiding errors that may occur during manual matching, reducing migration risks, and ensuring that the system can operate normally after migration.

[0005] Currently, most API recommendation methods accept natural language input and recommend the API of a certain software library. They are not suitable for scenarios involving two or more software libraries in code migration, nor can they recommend a mapping from one software library to another. Summary of the invention

[0006] The present invention provides a cross-library API recommendation method in the process of code migration, which can combine multi-source information and learn deep feature representation of API, so as to provide more effective information for recommendation tasks and improve the effect of API recommendation.

[0007] A cross-library API recommendation method during code migration includes the following steps:

[0008] (1) Use the crawler framework to crawl the official document information, open source projects, and Q&A community data of the source software library and the migration target software library, and process the crawled information to form metadata;

[0009] (2) Based on official document information, the API usage pattern is divided into three modules: API name, method information, and input and output information. The API name and method information are vectorized using a three-layer fine-tuned ALBERT model, and the input and output information is vectorized using a Word2Vec model.

[0010] The similarity of the API vectors of the source software library and the migration target software library is calculated to obtain the similarity of the three modules. Finally, the three similarities are weighted to obtain the API document similarity S1;

[0011] (3) Screen out relevant usage segments of APIs through open-source projects, construct a graph model of API usage patterns, and then calculate the graph node similarity between two API graph models; calculate the code semantic similarity based on the annotation information and method names of code segments, and finally obtain the API code segment similarity S2 by weighting the graph node similarity and the code semantic similarity;

[0012] (4) Through the Q&A community data, obtain the relevant Q&A descriptions of APIs, get the <API, description information> data pairs, then obtain their topic distribution vectors through the LDA topic model for these data pairs, and finally calculate the topic similarity S3 between two APIs;

[0013] (5) Obtain the final API similarity S by passing the similarities S1, S2, and S3 through the weight matrix W;

[0014] (6) Input an API of the source software library, calculate the final similarity S between the API of the source software library and each API in the migration target software library, and sort all the APIs in descending order of S for recommendation.

[0015] In step (1), the official document information refers to the official API usage documents of the software library, the open-source community refers to open-source hosting platforms such as Github that contain the source software library and the target software library, and the Q&A community refers to programmer professional field communities such as StackOverflow; the information processing methods include removing duplicate information, cleaning and filtering data, extracting keywords, formatting data, etc.

[0016] Step (1) specifically includes: using the Selenium framework to simulate browser behavior, writing a crawler program to automatically open the official document website, traversing the document pages of the two software libraries, and extracting relevant data such as API interfaces, function descriptions, and example codes. Collect all source software library projects and target software library projects in Github within half a year, and filter the projects according to the star count, removing all projects with 0 stars, and then screening out the core files from these projects. Crawl the relevant question and answer data of the two software libraries on Stack Overflow. Retrieve and crawl the content of relevant questions and answers in the Q&A community by searching for the names and related tags of APIs in the source software library and the migration target software library. Then remove the noise and irrelevant information in the text, such as HTML tags, special characters, punctuation marks, stop words, etc., and convert all words into lowercase letter forms.

[0017] In step (2), the fine-tuning process of the ALBERT model and the BERT model is as follows:

[0018] First, ALBERT is pre-trained using the question-answering community data obtained in step (1), and the dynamic MLM model is used to generate training data. All the data are copied into ten equal parts, and each copy of the data uses a different masking method;

[0019] Then, the Softmax objective function is used to fine-tune the ALBERT model on the SNLI dataset. First, the question-answer pairs in the dataset are input into the model respectively, and the average pooling operation is used to obtain two sentence vectors u and v; then the two sentence vectors are concatenated and fused to obtain a comprehensive vector representation. Finally, the Softmax layer is used to predict the annotation of the input sentence pair. The objective function is as follows

[0020]

[0021] Among them, |TD| represents the number of samples, y j Indicates whether the true annotation is of the jth category, Represents the probability that the model predicts the jth class;

[0022] Finally, cosine similarity loss is used as the training loss function for fine-tuning on the STS dataset; the specific calculation method is as follows

[0023]

[0024] Among them, score i Represents the similarity value of the real annotation of the sentence pair, cosine(u i ,v i ) represents the cosine similarity value between the sentence vectors generated by the model embedding two sentences.

[0025] In step (2), the calculation formula of API document similarity S1 is as follows:

[0026]

[0027] Among them, sim a ,sim b ,sim c They are respectively represented as the similarity of API name, method information and input and output information; a 、w b 、w c are weights corresponding to similarities, which are set to 1.0, 1.2, and 0.8 respectively in the present invention.

[0028] In step (3), the specific process of building the graph model of API usage pattern is as follows:

[0029] First, select projects containing the source software library and the target software library from the open-source projects obtained in step (1), determine the entire set of methods U, the set of API nodes I, and the relationship edges E between methods and APIs in the project. Starting from the starting number of the methods, sequentially extract the API calls involved in the methods. Then, determine whether the API exists. If it does not exist, add a new API node number to I. Otherwise, take out the existing API node number from I. Finally, establish the relationship edges between methods and APIs. For example, if API i is extracted from method u, then there is an edge between u and i, that is, e = (u, i), e ∈ E.

[0030] In step (3), after obtaining the call graph of APIs using the graph model, use the SDNE algorithm for graph embedding to calculate the graph node similarity between two APIs. Then, use the fine-tuned ALBERT model in step (2) to calculate the code semantic similarity between the annotation information and the method name.

[0031] The calculation formula for the API code snippet similarity S2 is as follows:

[0032] S2 = w d × sim d + w e × sim e

[0033] Among them, sim d and sim e respectively represent the graph node similarity and the code semantic similarity, and w d 、w e are the weights of the corresponding information similarity, which are set to 0.6 and 0.4 in the present invention.

[0034] In step (4), the calculation process of the topic similarity S3 is as follows:

[0035] First, screen out the Q&A pairs of the source software library and the target software library from the Q&A community data. First, perform text processing on the Q&A pairs, remove noise data such as special characters, punctuation marks, and HTML tags in the text, unify all texts to lowercase or uppercase, and remove stop words, such as prepositions, articles, and other words that have no actual meaning in the context. Then, use them as the description information of the APIs to construct <API, description information> pairs.

[0036] Determine the number of topics k = max{t, s} according to the number of classes t in the source software library and the number of classes s in the migration target software library; for each data pair <API, description information>, randomly shuffle the terms contained therein to generate n new copies, which are represented as a set. Each data pair in the set is used as an independent sample and is a non-repeating permutation of each other; obtain the probability distribution of the API in different topics through the LDA algorithm, use it as the word embedding vector of the API, and finally calculate the cosine similarity between the source software library API and the migration target software library API to obtain S3.

[0037] In step (5), the weight matrix is W = [w1, w2, w3] T , where the values of w1, w2, and w3 are obtained by training with the training set, and in the present invention, they are set to 0.4, 0.3, and 0.3. Multiply the similarity matrix [S1, S2, S3] by the weight matrix to obtain the final API similarity S.

[0038] In step (6), recommend the top 5 APIs with the highest final similarity S to the developer, and at the same time provide the official document information of the API and the code snippets in the open-source software library to the developer as additional information, so that the developer can understand the usage method of the target software library API and improve the migration efficiency.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The present invention can recommend the corresponding target library API during the migration process according to the API usage patterns of the two software libraries, improving the development efficiency during the code migration process.

[0041] 2. The present invention can combine multi-source information to learn the deep feature representation of the API, thereby providing more effective information for the recommendation task and improving the effect of API recommendation.

[0042] 3. The present invention adopts different algorithm models for different types of API information and then fuses them, which can more accurately capture the usage characteristics of the API, improve the recommendation accuracy of the target software library API, and thus improve the efficiency of code migration. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of a cross-library API recommendation method during the code migration process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] The following further describes the present invention in detail with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention, but do not limit it in any way.

[0045] The migration scenario of the embodiment of the present invention is to migrate WPF to Avalonia. WPF is a framework for creating client applications launched by Microsoft. It is developed on the basis of .NET Framework and provides a set of powerful tools and functions for building modern and highly customizable user interfaces. Avalonia is an open source cross-platform UI framework for building modern rich client applications. It is based on the .NET platform and is similar to WPF (Windows Presentation Foundation), but unlike WPF, Avalonia can run not only on Windows, but also on other operating systems such as MacOS and Linux.

[0046] like Figure 1 As shown, a cross-library API recommendation method in the code migration process includes the following steps:

[0047] Step 1: Use the Selenium framework to simulate browser behavior and obtain the official documentation information of the WPF and Avalonia software libraries through automated operations. By writing a crawler program, the official documentation website is automatically opened, and the document pages of the target software library are searched or traversed to extract relevant content, such as API interfaces, function descriptions, sample codes, and other related data. Collect all WPF projects and Avalonia projects in Github from January 1, 2023 to June 30, 2023, and filter the projects according to the number of stars, removing all 0-star projects. Filter out the core files from these projects, all of which end with ".cs" or ".xaml". Crawl the relevant questions and answers data of the target software library on Stack Overflow. By searching the name and related tags of the API in the target software library, search and crawl the content of related questions and answers in the Q&A community. Remove noise and irrelevant information in the text, such as HTML tags, special characters, punctuation, stop words (common but meaningless words), etc., and convert all words to lowercase.

[0048] Step 2, use the question-and-answer community data to pre-train ALBERT. The present invention uses a masked language model for pre-training. During training, the words of each sentence in the corpus will be converted into corresponding tags (tokens), and the tokens of the sentences will be randomly masked according to a certain masking rate (mask rate). All selected tokens will be replaced with [mask] tags with a probability of 80%, remain unchanged with a probability of 10%, and replaced with a random token with a probability of another 10%. The present invention uses a dynamic MLM model to generate training data, and copies all data into ten equal parts, using a different masking method for each copy of the data. After pre-training the model with unlabeled text corpus, the model is further fine-tuned using labeled paired sentences to obtain a deep sentence-level semantic vector representation.

[0049] Then the SNLI dataset is used for fine-tuning. SNLI is a natural language inference dataset provided by Stanford University, which contains 570k sentence pairs and their labels, and can be used to train natural language inference models. The present invention uses the Softmax objective function to fine-tune the model on the SNLI dataset. First, the question-answer pairs are input into the model respectively, and the sentence vectors u and v are obtained by average pooling operation; then the two sentence vectors are concatenated and fused to obtain a comprehensive vector representation, and finally the Softmax layer is used to predict the label of the input sentence pair. The objective function is as follows:

[0050]

[0051] Finally, the STS dataset is used for fine-tuning. STS is a dataset published by Cer et al. for evaluating the performance of models in semantic similarity calculation tasks. The present invention uses cosine similarity loss as the training loss function, which is a minimum mean square error loss. The specific calculation method is as follows, where score i Represents the similarity value of the real annotation of the sentence pair, cosine(u i ,v i ) represents the cosine similarity value between the sentence vectors generated by the model embedding two sentences.

[0052]

[0053] The WPF and Avalonia official document information processed in step 1 is further processed. For each API, the description information and method name of the API are extracted separately and input into the ALBERT fine-tuned in step 2 to obtain the word embedding vector u of the description information. d and the word embedding vector u of the method name m. Because the text types of input parameters and return values ​​are relatively fixed and simpler than description information, the word2Vec method is used to embed them, and the text corpus preprocessed in step 1 is used to train the word2Vec model. This patent uses the Skip-gram model, which can try to predict context words through the current word. Input the input parameters and return parameters into word2Vec, and then concatenate the two vectors to obtain the word embedding vector u p Then calculate the API word vector u in WPF and Avalonia software library d 、u m 、u p Finally, the document similarity S1 is obtained using the following weighted formula:

[0054]

[0055] Among them, sim a ,sim b ,sim c They are respectively expressed as the similarity of API name, method information and input and output information. a 、w b 、w c are the weights of the corresponding information similarities. According to experiments, the present invention sets these three weights to 1.0, 1.2, and 0.8.

[0056] Step 3, determine all the method sets U1 and U2 from the open source projects of WPF and Avalonia, as well as the API sets I1 and I2 of the WPF software library and the Avalonia software library. Number each method of U1 and U2, extract the API calls designed in the method in turn, and add an edge e=(u,i) if API i is used in method u. After obtaining the API call graph, use the SDNE algorithm for graph embedding. The SDEN algorithm is a graph embedding algorithm that uses an autoencoder to simultaneously optimize the first-order and second-order similarities. The learned vector can retain local and global structural information, and obtain the API graph embedding, which is used to calculate the similarity between two APIs.

[0057] For a given graph network, define the vector The model uses an unsupervised autoencoder to learn the second-order similarity information in the network. Given an input x i , first pass through several layers of neural networks to get the vector representation Afterwards, the output is obtained through a neural network with the opposite structure The goal of the autoencoder is to fit the input x i and Learn the low-dimensional vector representation of the input. The corresponding loss function is:

[0058]

[0059] Some properties of graph networks prevent autoencoders from being directly used for graph embedding. Due to the sparsity of the network, there are a large number of zeros in the matrix S. If a traditional autoencoder is used, these zero elements are easier to reconstruct, while non-zero elements are likely to be ignored. To this end, the weights of different elements in the loss function need to be modified:

[0060]

[0061] Among them, ⊙ represents the dot multiplication operation. The graph embedding algorithm must not only maintain the global structure of the network, but also capture the local network structure, that is, the first-order similarity. For this purpose, the supervised learning method can be used, and its loss function is defined as follows:

[0062]

[0063] Use L2 regularization To prevent overfitting, the final loss function of the model is:

[0064]

[0065] Graph embedding can map the nodes in the graph to a low-dimensional vector space, so that the nodes can be represented as continuous numerical vectors in the vector space. The advantage of this is that it can capture the similarities and relationships between API nodes. By inputting the WPF and Avalonia APIs into the model, the graph embedding vector of the API can be obtained, and then the cosine similarity is calculated. Then, the semantic similarity between the annotation information and the method name is calculated using the ALBERT algorithm fine-tuned in step 2, and finally these two similarities are weighted to obtain S2

[0066] S2=w d ×sim d +w e ×sim e

[0067] Among them, sim d ,sim e They are respectively expressed as graph structure similarity and code semantic similarity. d 、w e are the weights of the corresponding information similarity. According to experiments, the present invention sets these two weights to 0.6 and 0.4.

[0068] Step 4: Construct <API, description information> pairs from the Q&A pairs after text preprocessing. Determine the number of topics k = max{t, s} according to the number of WPF classes t and the number of classes in Avalonia s. For each data pair, shuffle the terms contained therein to generate n new copies. All the copies generated after shuffling are represented as a set. Each data pair in the set is used as an independent sample and is a non-repeating permutation of each other. In the present invention, n is set to 10. Through the LDA algorithm, the probability distribution of the API in different topics can be obtained, which is used as the word embedding vector of the API. Calculate the cosine similarity of the APIs in WPF and Avalonia to obtain S3.

[0069] Step 5: Use the three similarities obtained in Steps 2 - 4 to construct a similarity feature vector v = [S1, S2, S3], and the weight matrix is W = [w1, w2, w3]T, where the values of w1, w2, and w3 are obtained by training with the training set. In the present invention, they are set to 0.4, 0.3, and 0.3. Multiply the similarity matrix [S1, S2, S3] by the weight matrix to obtain the final API similarity S. The calculation formula is as follows:

[0070] S = V · W

[0071] Step 6: The developer inputs the API in WPF. The model calculates the final similarity S between this API and all the APIs in Avalonia, and then sorts the similarities from high to low. In this patent, the top 5 APIs ranked by similarity are recommended to the developer as candidate APIs. At the same time, the relevant API description information and code blocks collected in Step 1 are returned as additional information, so that the developer can better understand the use of the APIs in Avalonia and help with the migration work.

[0072] The above-described embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, and equivalent replacements made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A cross-library API recommendation method during code migration process, characterized in that, it includes the following steps: (1) Use a crawler framework to crawl the official document information, open source projects, and Q&A community data of the source software library and the migration target software library, and process the crawled information; (2) Through the official document information, split the usage pattern of the API into three modules: API name, method information, and input / output information; vectorize the API name and method information through a three-layer fine-tuned ALBERT model, and vectorize the input / output information through a Word2Vec model; Calculate the similarity of the API vectors of the source software library and the migration target software library to obtain the similarity of the three modules, and finally weight these three similarities to obtain the API document similarity S1; (3) Through open source projects, screen out the relevant usage segments of the API, construct a graph model of the API usage pattern, and then calculate the graph node similarity between two API graph models; and calculate the code semantic similarity based on the annotation information and method name of the code segment. Finally, weight the graph node similarity and the code semantic similarity to obtain the API code segment similarity S2; (4) Through the Q&A community data, obtain the relevant Q&A descriptions of the API, obtain <API, description information> data pairs, and then obtain their topic distribution vectors through the LDA topic model for these data pairs. Finally, calculate the topic similarity S3 between two APIs; (5) Pass the similarities S1, S2, and S3 through the weight matrix W to obtain the final API similarity S; (6) Input an API of the source software library, calculate the final similarity S between the API of the source software library and each API in the migration target software library, and sort all the APIs in descending order of S for recommendation.

2. The cross-library API recommendation method during code migration process according to claim 1, characterized in that, in step (1), the official document information refers to the official API usage document of the software library, the open source community refers to the open source hosting platform including the source software library and the target software library, and the Q&A community refers to the professional community of programmers; the information processing methods include removing duplicate information, cleaning and filtering data, extracting keywords, and formatting data.

3. The cross-library API recommendation method during code migration process according to claim 1, characterized in that, in step (2), the fine-tuning process of the ALBERT model and the BERT model is as follows: First, use the Q&A community data obtained in step (1) to pre-train ALBERT, generate training data using a dynamic MLM model, and copy all the data into ten equal parts, and each part uses a different masking method; Then, use the Softmax objective function to fine-tune the ALBERT model on the SNLI dataset. First, input the Q&A pairs of the dataset into the model respectively, and use average pooling operation to obtain two sentence vectors u and v; then splice and fuse the two sentence vectors to obtain a comprehensive vector representation, and finally use the Softmax layer to predict the annotation of the input sentence pair. The objective function is as follows Among them, |TD| represents the number of samples, y j Indicates whether the true annotation is of the jth category, Represents the probability that the model predicts the jth class; Finally, the cosine similarity loss is used as the training loss function for fine-tuning on the STS dataset; the specific calculation method is as follows Among them, score i Represents the similarity value of the real annotation of the sentence pair, cosine(u i ,v i ) represents the cosine similarity value between the sentence vectors generated by the model embedding two sentences.

4. The cross-library API recommendation method during code migration according to claim 1, characterized in that, in step (2), the calculation formula of the API document similarity S1 is as follows: Among them, sim a ,sim b ,sim c They are respectively represented as the similarity of API name, method information and input and output information; a 、w b 、w c is the weight corresponding to the similarity.

5. The cross-library API recommendation method during code migration according to claim 1, characterized in that, in step (3), the specific process of constructing the graph model of the API usage pattern is as follows: First, select the projects containing the source software library and the target software library from the open-source projects obtained in step (1), determine the set U of all methods, the set I of API nodes, and the relationship edges E between methods and APIs in the project. Starting from the starting number of the method, extract the API calls involved in the method in sequence; then judge whether the API exists. If it does not exist, add a new API node number to I. Otherwise, take out the existing API node number from I; finally, establish the relationship edge between the method and the API.

6. The cross-library API recommendation method during code migration according to claim 1, characterized in that, in step (3), use the graph model to obtain the call graph of the API, and then use the SDNE algorithm for graph embedding to calculate the graph node similarity between two APIs; then use the fine-tuned ALBERT model in step (2) to calculate the code semantic similarity between the annotation information and the method name; The calculation formula of the API code snippet similarity S2 is as follows: S2=w d ×sim d +in e ×sim e Among them, sim d and sim e Respectively represent the graph node similarity and code semantic similarity, w d 、w e is the weight of the corresponding information similarity.

7. The cross-library API recommendation method during code migration according to claim 1, characterized in that, in step (4), the calculation process of the topic similarity S3 is as follows: First, screen out the Q&A pairs of the source software library and the target software library from the Q&A community data, perform text processing on the Q&A pairs, and then use them as the description information of the API to construct the <API, description information> pair; According to the number t of classes in the source software library and the number s of classes in the migration target software library, determine the number of topics k = max{t, s}; for each data pair <API, description information>, randomly shuffle the terms contained in it to generate n new copies, which are represented as a set. Each data pair in the set is used as an independent sample and is a non-repeating full permutation of each other; obtain the probability distribution of the API in different topics through the LDA algorithm, use it as the word embedding vector of the API, and finally calculate the cosine similarity between the source software library API and the migration target software library API to obtain S3.

8. The cross-library API recommendation method during code migration according to claim 1, characterized in that, In step (5), the weight matrix is ​​W = [w1, w2, w3] T , where the values ​​of w1, w2, and w3 are obtained from the training set. The similarity matrix [S1, S2, S3] is multiplied by the weight matrix to obtain the final API similarity S.

9. The cross-library API recommendation method during code migration according to claim 1, characterized in that, in step (6), recommend the top 5 APIs with the highest final similarity S to the developers, and at the same time provide the official document information of the API and the code snippets in the open-source software library to the developers as additional information.

Citation Information

Patent Citations

  • Code search recommendation device and method based on open source knowledge

    CN112051986A

  • Apparatus and methods for implementation of network software interfaces

    US20060020950A1