A context-aware API recommendation method with implicit feedback mechanism
Patent Information
- Application Number
- CN202410060704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-01-16
AI Technical Summary
[0021]1)本发明公开了一种基于代码的API推荐方法,该方法不需要用户的显式输入,也不需要额外的知识工件的支持,仅需要开源代码作为背景数据,在用户开发的过程中自动地进行API的推荐,相比于基于查询的技术,本发明可以直接帮助用户提高开发效率。
Smart Images

Figure CN117892014B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an automatic API recommendation method based on context-aware collaborative filtering and with an implicit feedback mechanism, belonging to the fields of software engineering and software automation technology. Background Technology
[0002] In modern software development, developers often use third-party libraries to implement required functionality rather than building new systems from scratch. These libraries expose their functionality through Application Programming Interfaces (APIs) and control the interaction with clients via the API. However, correctly learning and mastering API usage is not easy: official documentation often only provides API descriptions without detailed usage examples, and the content of API documentation may be ambiguous, incomplete, or incorrect; searching informal sources like Stack Overflow can be very time-consuming and error-prone, as API usage examples found on Q&A websites may be of poor quality; furthermore, some complex APIs are difficult to learn and master, even for professional developers working at large software companies like Microsoft; some studies suggest that developers spend up to 19% of their programming time searching the internet for source code, especially API usage examples.
[0003] In recent years, the problem of automatic recommendation of API function calls and usage patterns has attracted considerable effort and attention from the research community. Many API recommendation techniques have been proposed to facilitate the learning process for developers and improve development efficiency. These methods are generally divided into two categories: 1) Query-based API recommendation methods, which accept natural language queries from users as input and retrieve and recommend relevant APIs based on the input query. However, in practice, these methods have significant drawbacks: First, they require developers to explicitly provide precise natural language queries as input, which places an additional burden on developers, especially those with limited development experience; second, to train the model to overcome the semantic gap between natural language and programming language, additional knowledge artifacts will be explicitly required: such as API documentation, corpora consisting of code and comments, posts from question-and-answer websites (e.g., SO), etc. Therefore, the performance of query-based recommendation methods is largely limited by the length of the input query and by ambiguous, incomplete, and incorrect documentation or descriptions. 2) Code-based methods, which recommend the next API to be used in a method declaration by treating the surrounding code of the recommendation point as context and open-source code repositories as background data. Clearly, these methods do not require explicit user input or additional knowledge components; instead, they automatically provide an API recommendation list during the user's programming process, solely based on relevant API usage information from open-source code. However, existing code-based API recommendation methods still suffer from high redundancy and poor performance, while those based on recommendation algorithms (such as collaborative filtering) exhibit significant issues like low recommendation accuracy, coarse feature granularity, limited feature types, and the inability to self-update recommendation results. Summary of the Invention
[0004] Purpose of the Invention: Code-based methods automatically provide API recommendation lists during user programming, relying solely on relevant API usage information from open-source code, without requiring explicit user input or additional knowledge components. However, existing code-based API recommendation algorithms still suffer from significant problems such as poor recommendation accuracy, coarse feature granularity, single features, and inability to self-update recommendation results. To overcome the limitations of existing work and further improve the accuracy of API recommendations, this invention innovatively proposes a fine-grained, personalized, context-aware collaborative filtering-based API recommendation method. This method, based on open-source project code, collects multi-dimensional feature information and contextual information at the method declaration level, aiming to automatically and accurately recommend API usage patterns for the functional implementation of the method declaration being edited by the developer. Furthermore, during interaction with the developer, the method records implicit feedback information based on observable user behavior and updates the recommendation process accordingly, achieving dynamic self-updating of recommendation results.
[0005] Technical Solution: A context-aware API recommendation method with implicit feedback mechanism. This method treats APIs as items, method declarations as users, and enclosed items as context. It uses collaborative filtering for API recommendation and employs an implicit feedback mechanism to achieve self-updating of the recommendation process. The method includes the following steps:
[0006] Step 1: First, perform static analysis on the source code of the target Java project to extract the method information and API call sequence information declared in each project;
[0007] Step 2: Use the TF-IDF algorithm to encode the API call information of each item and declaration into a feature vector, and select the N items most relevant to the active item based on the cosine similarity of the feature vectors.
[0008] Step 3: Use the Jaccard coefficient-based method to calculate the similarity between method declarations, and select the M declarations that are most similar to the activity declaration from the N projects;
[0009] Step 4: Based on the API call sequence of similar claims and activity claims, construct a "user-item" rating matrix, calculate the rating of the missing items in the matrix based on a decentralized collaborative rating algorithm, and output the API recommendation list in descending order of rating;
[0010] Step 5: Update the features of the campaign declaration based on the recommendation list and implicit feedback from developers, and return to Step 3 to re-execute. Depending on user needs, feedback may be performed in one or more rounds.
[0011] The specific implementation process of step 1 is as follows:
[0012] 11) Using Rascal M 3 Construct and traverse the abstract syntax tree of the target Java project;
[0013] 12) Analysis of M 3 The model provides method declarations and API call relationships. 3 The method declaration data provided by the model all contain a<v1,v2> Yes, where v1 and v2 are values representing locations. These locations are uniform resource identifiers, representing artifact identifiers (also called logical locations) or physical pointers on the file system pointing to the corresponding artifact (also called physical locations). Declaration relationships map the logical location of an artifact (e.g., a method) to its physical location. "Method call" relationships map the logical location of the caller to the logical location of the callee. All "method call" relationships are categorized and organized based on the caller.
[0014] 13) For each project, output the signature information and API call sequence of all methods declared within it. A method declaration typically consists of a return value (which may be void), a method name, a parameter list (which may be empty), and a method body, with the first three items constituting the method signature. (Based on Rascal M...) 3 The model provides a method call mapping relationship, outputting the method signature information of each caller (i.e., the declaration) and the callee information (i.e., the API call sequence) in its method body.
[0015] In step 2, the N most relevant items to the activity item are selected based on the cosine similarity of feature vectors. This selection involves calculating the cosine similarity between feature vectors to define the similarity between contexts and the relevance between contexts and claims, and then selecting the N most relevant contexts to the activity item. The specific implementation process of step 2 is as follows: First, API calls are treated as terms (keywords), and items and claims are treated as documents. The TF-IDF algorithm is used to encode the API call information of items and claims into corresponding feature vectors, where each dimension represents information about one type of API call in the document. Then, the similarity between candidate items and the activity item, and the relevance between candidate items and the activity claim are calculated based on the cosine similarity of the feature vectors of items and claims. The relevance between candidate items and the activity item is then calculated based on these two relevance indicators, and the N most relevant items (i.e., contexts) to the activity item are selected accordingly.
[0016] In step 3, the Jaccard coefficient is used to measure the similarity between claims, and the M claims that are most similar to the active claim are selected from the N relevant items.
[0017] First, extract the API call set and parameter set for each declaration in the relevant projects, and calculate the Jaccard coefficient for each candidate declaration and active declaration on the call set and parameter set respectively; then, calculate the similarity between the candidate declaration and the active declaration based on the two calculated Jaccard coefficients, sort the candidate declarations according to the similarity, and output the M declarations that are most similar to the active declaration.
[0018] In step 4, by treating APIs as items, method declarations as users, and enclosed items as context, a collaborative filtering algorithm is used to construct a "user-item" rating matrix based on similar and active declarations. The task of recommending APIs for active declarations is achieved by calculating missing values in the matrix. Each row in the rating matrix represents a user (i.e., a method declaration), and each column represents an item (i.e., an API call). For similar declarations, if a declaration calls an API, the value at the intersection of the corresponding row and column is set to 1; otherwise, it is set to 0. For active declarations to be recommended, if they have already called some APIs, the corresponding position is set to 1; otherwise, it is set to -1, indicating a candidate API to be recommended. A decentralized collaborative filtering rating algorithm is used to calculate the score at the missing values in the matrix (i.e., the positions marked -1). This score indicates the probability that the active declaration has called the corresponding API. Therefore, the candidate APIs can be sorted using this score, and then an API recommendation list is output.
[0019] In step 5, the features of the activity declaration are updated based on the recommendation list and implicit feedback from developers, and step 3 is returned for re-execution. First, based on the API recommendation list obtained in step 4, the developers' selections and usage are recorded. If a developer selects certain APIs from the list to complete the development of the activity declaration, these APIs are labeled 1 and stored in the feedback repository; APIs not selected in the list are labeled 0 and stored in the feedback repository. Then, based on the information in the feedback repository, APIs labeled 1 are returned to step 3 to supplement the activity declaration's call set, the similarity between other declarations and the activity declaration is recalculated, and the M most similar declarations are selected; APIs labeled 0 are returned to step 4, and these APIs will be filtered out when constructing the scoring matrix and will not be encoded into the matrix. Finally, based on the reselected similar declarations and the reconstructed scoring matrix, step 4 is re-executed to generate the API recommendation list. Depending on user needs, feedback can be performed in one or more rounds.
[0020] Compared with the prior art, the present invention has the following advantages:
[0021] 1) This invention discloses a code-based API recommendation method. This method does not require explicit user input or additional knowledge artifacts. It only requires open-source code as background data and automatically recommends APIs during the user's development process. Compared with query-based techniques, this invention can directly help users improve development efficiency.
[0022] 2) This invention uses multi-dimensional features at the declaration level (such as call sets and parameter sets) for collaborative filtering, which enables more granular, accurate and efficient API recommendations to help users complete the development and implementation of the current declaration.
[0023] 3) This invention records and uses implicit user feedback information to achieve self-updating of the recommendation process. In practical recommendation applications, this invention can continuously improve the recommendation effect as the project development progresses. Attached Figure Description
[0024] Figure 1 This is an overview diagram of the context-aware API recommendation method with implicit feedback mechanism according to an embodiment of the present invention. Detailed Implementation
[0025] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0026] This invention is a code-based API recommendation technology. Targeting the large number of open-source projects in an Open Source Software Repository (OSS), it first selects candidate projects most relevant to the currently developing project (hereinafter referred to as active projects). Then, it selects declarations from these relevant projects that are most similar to the method declarations currently being implemented (hereinafter referred to as active declarations). Finally, it constructs a "user-item" rating co-occurrence matrix based on the API call information of these similar declarations, and completes the API recommendation task for active declarations by calculating missing values in the matrix. Simultaneously, this invention records implicit feedback information from developers after receiving the API recommendation list and updates the recommendation process based on this information to achieve more accurate recommendations. Compared to other code-based API recommendation technologies, this invention is based on the idea of collaborative filtering, using contextual information and method declaration-level feature information to filter training set projects and declarations within those projects, and employs a classic decentralized collaborative scoring algorithm for API recommendation, significantly improving recommendation efficiency and accuracy. Furthermore, this invention achieves self-updating of the recommendation process, continuously improving the recommendation effect as project development progresses in practice.
[0027] This embodiment provides a context-aware API recommendation method with an implicit feedback mechanism, which improves the accuracy and efficiency of API recommendation and enables the recommendation process to be self-updated.
[0028] Step 1: Perform static analysis on the Java project source code, extracting the declared method information and API call sequence information for each project. Use Rascal M... 3 Construct and traverse the abstract syntax tree of the target Java project, and analyze M 3 The model provides method declarations and API call relationships. 3The model provides a method data that all contain a<v1,v2> Yes, where v1 and v2 are values representing locations. These locations are uniform resource identifiers, representing artifact identifiers (also called logical locations) or physical pointers on the file system pointing to the corresponding artifact (also called physical locations). Declaration relationships map the logical location of an artifact (e.g., a method) to its physical location. "Method call" relationships map the logical location of the caller to the logical location of the callee. All "method call" relationships are categorized and organized based on the caller. For each item, the signature information and API call sequence of all methods declared within it are output. A method declaration typically consists of a return value (which may be void), a method name, a parameter list (which may be empty), and a method body, with the first three items constituting the method signature. Based on Rascal M... 3 The model provides a method call mapping, outputting the method signature information of each caller (i.e., the declaration) and the callee (i.e., the API sequence) information in its method body.
[0029] Step 2: Use the TF-IDF algorithm to encode project and method declarations containing API calls into feature vectors. Define the similarity between contexts and the relevance between contexts and declarations by calculating the cosine similarity between feature vectors. Then select the N contexts most relevant to the active project. The specific implementation process of Step 2 is as follows:
[0030] 21) Treat API calls as terms (keywords) and items and claims as documents. For an API call i and an item (or claim) d, calculate its TF-IDF value and use that TF-IDF value as the value of the i-th dimension of the feature vector of d.
[0031] 22) Based on candidate project p and activity project p a The feature vectors (denoted as respectively) and ), and the event statement d a eigenvectors Calculate candidate item p and active item p a The similarity between sim(p,p) a ) and p and activity statement d a The correlation between rel(p,d) a )as follows:
[0032]
[0033]
[0034] 23) Based on sim(p,p) a ) and rel(p,d aCalculate the relevance rel(p, p) between the active item and the candidate item. a )as follows:
[0035] rel(p,p a )=I(rel(p,d a )>0)*sim(p,p a )
[0036] in,
[0037] Then, based on the relevance rel(p,p) a The size of the candidate items is used to sort all candidate items and output the N items with the highest relevance to the activity. N is a configurable parameter. Users can set an appropriate value according to the actual application scenario. The default value is N=2.
[0038] Step 3: Use the Jaccard coefficient to measure the similarity between claims, and select the M claims most similar to the active claim from the N retained relevant items. The specific implementation process is as follows:
[0039] 31) Based on the N related items obtained in step 2, extract the API call sequence and parameter type list of each method declaration d, which are denoted as the call set Call(d) and the parameter set Parm(d) respectively.
[0040] 32) Regarding the activity declaration d a And candidate declaration d, calculate their Jaccard coefficients on the call set and parameter set respectively, denoted as J call (d a ,d) and J parm (d a ,d), the calculation is as follows:
[0041]
[0042]
[0043] 33) Calculation activity declaration d a Similarity to candidate statement d: sim(d) a ,d)=w1*J call (d a ,d)+w2*J parm (d a ,d). Among them, w1 and w2 are configurable weight parameters that users can adjust according to the application task. The default settings are w1=0.7 and w2=0.3.
[0044] 34) According to sim(d) ad) Sort all candidate declarations and output the M declarations most similar to the active declaration. Similarly, M is a configurable parameter that users can adjust according to their application scenario; the default setting is M=4.
[0045] Step 4: By treating APIs as items, method declarations as users, and enclosed items as context, a collaborative filtering algorithm is used to construct a "user-item" rating matrix based on similarity and activity declarations. The task of recommending APIs based on activity declarations is then accomplished by calculating missing values in the matrix. The specific implementation process is as follows:
[0046] 41) Let L be the number of all API calls appearing in similar and active declarations. Construct a two-dimensional scoring matrix RM of size (M+1)*L, as shown below:
[0047]
[0048] Each row in the matrix represents a method declaration d, and each column represents an API call i. The first M rows of the matrix represent M similar methods d. sim The last line represents the activity method d. a .
[0049] Setting the values of the first M rows of matrix RM: For a given position RM x,y If the x-th line represents the declaration d x The API call represented by column y was invoked. y Then the position RM x,y Set to 1, otherwise set to 0.
[0050] Setting the value of the last row of matrix RM: For a certain position RM M+1,y If the activity declaration d a The API call represented by column y was invoked. y Then the position RM M+1,y Set to 1 otherwise -1. The positions marked -1 represent the missing values in the matrix to be calculated, and the calculated score represents the activity declaration d. a The probability of calling API call i represented by the corresponding column.
[0051] 42) Based on the rating matrix constructed in step 41), calculate the missing values (marked as -1) in the matrix using the following formula:
[0052]
[0053] Where D is the set of all similar declarations, r d,i To declare the score of API call i by d, To declare the average score of d for all API calls, rel(p,p) a ) and sim(d,d a The results are obtained in steps 23) and 24 respectively.
[0054] 43) Based on the missing value scores calculated in step 42), sort the corresponding API calls and output an API recommendation list.
[0055] Step 5: Update the features of the activity declaration based on the recommendation list and implicit feedback from developers, and return to step 3 to re-execute. The specific implementation process is as follows:
[0056] 51) Based on the API recommendation list obtained in step 4, record the developers' selections and usage. If a developer selects certain APIs from the list to complete the development of the activity declaration, these APIs are labeled as 1 and stored in the feedback repository; APIs not selected from the list are labeled as 0 and stored in the feedback repository.
[0057] 52) Based on the information in the feedback repository, return the API with tag 1 to step 3 and supplement the activity declaration d. a Call set Call(d a ), recalculate the similarity of other claims and activity claims and select the M most similar claims; return the APIs with a label of 0 to step 4. When constructing the rating matrix, these APIs will be filtered out and will not be encoded into the matrix.
[0058] 53) Based on the reselected similarity claims and the reconstructed rating matrix, repeat step 4 to generate the API recommendation list. Feedback can be performed in one or more rounds, depending on user needs.
Claims
1. A context-aware API recommendation method with implicit feedback mechanism, characterized in that, By treating APIs as items, method declarations as users, and enclosed projects as context, collaborative filtering is used for API recommendations, and an implicit feedback mechanism is used to achieve self-updating of the recommendation process; this includes the following steps: Step 1: First, perform static analysis on the source code of the target Java project to extract the method information and API call sequence information declared in each project; Step 2: Use the TF-IDF algorithm to encode the API call information of each item and declaration into a feature vector, and select the N items most relevant to the active item based on the cosine similarity of the feature vectors. Step 3: Use the Jaccard coefficient-based method to calculate the similarity between method declarations, and select the M declarations that are most similar to the active method declaration from the N items; Step 4: Based on the API call sequence of similar method declarations and activity method declarations, construct a "user-item" rating matrix, calculate the rating of the missing items in the matrix based on a decentralized collaborative rating algorithm, and output the API recommendation list in descending order of rating; Step 5: Update the features of the activity method declaration based on the recommendation list and implicit feedback from developers, and return to step 3 to re-execute; the specific implementation process is as follows: 51) Based on the API recommendation list obtained in step 4, record the developers' choices and store them in the feedback repository; 52) Return the selected API call information to step 3, supplement the API call set of the active method declaration, recalculate the similarity between the candidate method declaration and the active method declaration, and select M similar method declarations; 53) Based on the reselected similar method declaration, continue to step 4: When constructing the rating matrix, read the API calls that were not selected in the feedback repository. These API calls will be filtered out and will no longer be encoded into the rating matrix, that is, they will no longer be recommended to users. 54) Based on user needs, the feedback rounds are executed in a loop for one or more rounds.
2. The context-aware API recommendation method with implicit feedback mechanism according to claim 1, characterized in that, The specific implementation process of step 1 is as follows: 11) Use Construct and traverse the abstract syntax tree of the target Java project; 12) Analysis The model provides method declarations and API call relationships; 13) For each project, output the signature information and API call sequence of all methods declared therein.
3. The context-aware API recommendation method with implicit feedback mechanism according to claim 1, characterized in that, In step 2, the TF-IDF algorithm is used to encode project and method declarations containing API calls into feature vectors. The cosine similarity between feature vectors is calculated to define the similarity between contexts and the relevance between contexts and declarations. Then, the N contexts most relevant to the active project are selected. The specific implementation process of step 2 is as follows: 21) Projects currently under development are called active projects, and complete projects from open-source code repositories are called candidate projects; 22) Treat API calls as keywords, and projects and declarations as documents; For an API call and a project or statement Calculate its TF-IDF value and use this value as The eigenvector of the first The values of the dimensional components; 23) Calculate candidate items based on the cosine similarity of feature vectors. With activities and projects similarity between and candidate projects Statement of Activity Methods correlation between ; 24) Based on and Calculate the correlation between active items and candidate items: Then, based on the relevance, all candidate items are sorted, and the N items with the highest relevance are output, where... .
4. The context-aware API recommendation method with implicit feedback mechanism according to claim 1, characterized in that, In step 3, the Jaccard coefficient is used to measure the similarity between claims, and the M claims most similar to the activity method claim are selected from the N retained relevant items; the specific implementation process is as follows: 31) The method declaration currently being edited and implemented in the active project is called the active method declaration, and all method declarations from related projects are called candidate method declarations; 32) Extract the API call sequence and parameter type list of each method declaration, and denote them as the call set and parameter set, respectively; 33) Regarding the activity method statement and candidate method declarations Calculate the Jaccard coefficients for the call sets and parameter sets of both, denoted as . and ; 34) Declaration of Calculation Activity Methods and candidate method declarations Similarity: ;in and These are configurable weight parameters; 35) According to Sort all candidate method declarations and output the M declarations that are most similar to the active method declaration.
5. The context-aware API recommendation method with implicit feedback mechanism according to claim 1, characterized in that, In step 4, a user-item rating matrix is constructed based on similarity method declarations and activity method declarations. The API recommendation task for activity method declarations is then implemented by calculating missing values in the matrix. The specific implementation process is as follows: 41) Let L be the number of all API calls appearing in similar method declarations and active method declarations, then the construction... Two-dimensional rating matrix of size Each line represents a method declaration. Each column represents an API call. The first M rows of the matrix represent M similarity method declarations. The last line represents the activity method statement. ; 42) Matrix Setting the values of the first M rows: for a given position If the first Statement by the representative The first one was called API calls represented by the column Then the position Set as Otherwise set to ; 43) Matrix The value setting of the last line: for a certain position If the activity method statement The first one was called API calls represented by the column Then the position Set as Otherwise set to ; marked as The position represents the missing value in the matrix that needs to be calculated; 44) Use a decentralized collaborative scoring algorithm to calculate the score at the missing values in the matrix. The calculated score represents the activity method declaration. Call the API represented by the corresponding column Based on the probability of the API calls being ranked, the corresponding API calls are sorted from highest to lowest score, and an API recommendation list is generated and output.
6. The context-aware API recommendation method with implicit feedback mechanism according to claim 2, characterized in that, In step 12), the analysis The model provides method declarations and API call relationships; the data for the declarations or methods provided by the model all contain a...<v1, v2> Yes, where v1 and v2 are values representing positions; the position is a unified resource identifier that represents the logical or physical location of the workpiece; the declaration relationship maps the logical location of the workpiece to its physical location; the "method call" relationship maps the logical location of the caller to the logical location of the callee.
7. The context-aware API recommendation method with implicit feedback mechanism according to claim 2, characterized in that, In step 13), the signature information and API call sequence of all methods declared for each project are output, specifically: 131) A method declaration consists of a return value, a method name, a parameter list, and a method body, with the first three items constituting the method signature; 132) Based on The provided information outputs the method signature information for each method, as well as the sequence of API calls within its method body.
Citation Information
Patent Citations
API recommendation method and device for cooperatively filtering and balancing data information
CN112269946A
Service recommendation system and training method thereof
CN116975461A