API sequence recommendation method based on user intention perception

By employing a user intent-aware API sequence recommendation method, and leveraging CodeBERT and query expansion techniques, we have addressed the shortcomings of traditional methods in understanding user intent and code context. This approach enables efficient and accurate API sequence recommendation, thereby improving development efficiency and program quality.

CN121052255APending Publication Date: 2025-12-02HANGZHOU DIANZI UNIVERSITY BINJIANG INSTITUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511026253.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately and quickly recommend suitable API sequences, resulting in low development efficiency and poor program quality. Traditional methods lack a comprehensive understanding of user intent and code context.

Method used

By constructing multi-source information training samples, using the pre-trained language model CodeBERT for joint modeling, and combining program structure representation and natural language description, API call sequences that conform to user intent are generated, and semantic enhancement is performed through a query expansion mechanism.

Benefits of technology

It significantly improves the accuracy and usability of API recommendations, and the generated API sequences are logically ordered, comprehensively covering user needs and improving development efficiency and program quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052255A_ABST
    Figure CN121052255A_ABST
Patent Text Reader

Abstract

The invention provides an API sequence recommendation method based on user intention perception. The method comprises the steps of firstly constructing a training sample based on multi-source information; then carrying out joint modeling on the natural language description and the program structure representation by utilizing a pre-training language model to obtain unified representation fusing user intention and code semantics; secondly, predicting a candidate API sequence in an autoregression mode through a sequence generation model; secondly, in the reasoning stage, semantic enhancement is carried out on the input natural language query by adopting a query expansion technology so as to generate an expansion query; and finally, jointly inputting the extended query and the current code context into the sequence generation model, and outputting an API calling sequence conforming to the user intention. Natural language query and code context information are comprehensively utilized, the development intention is obtained from the natural language description input by the user, modeling is conducted on the programming environment in combination with the current code context, and the API recommendation result better meeting the actual development requirement is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of API recommendation technology, and more specifically to an API sequence recommendation method based on user intent awareness. Background Technology

[0002] As software development becomes increasingly complex, developers often need to call application programming interfaces (APIs) provided by multiple software libraries or frameworks when implementing a certain function. However, faced with a large number of APIs with complex functions, developers often find it difficult to accurately and quickly find and correctly combine the required APIs, thus affecting development efficiency and program quality.

[0003] Traditional API lookup methods include consulting official documentation or searching for examples in search engines, code completion techniques, and recommended methods using only natural language or only programming languages. When developers are unfamiliar with the API of a library or framework, they typically need to manually consult official documentation, technical blogs, or search for relevant code examples in search engines. Consulting official documentation or searching for examples in search engines relies on the developer's experience and search skills, is inefficient, time-consuming, and the information obtained often lacks context or is difficult to reuse directly.

[0004] Modern integrated development environments (IDEs) generally include code completion functionality, which can predict the next API call based on information such as variable types and the current context. However, most of these methods are based on static type inference or syntax rules, lacking an understanding of the user's intent, and the completion suggestions are usually local and focused on a single API, making it difficult to cover the complete functional logic.

[0005] Some recommendation systems rely on natural language queries to analyze developers' query intent and recommend relevant APIs, but they neglect code context. Other methods predict the next API based on the current code environment, but fail to understand the semantic intent of the user's current task. This "unimodal input" approach suffers from incomplete information and insufficient understanding, resulting in limited accuracy and usability of the recommendation results.

[0006] Therefore, there is an urgent need for a recommendation method that can comprehensively understand user intent, analyze contextual information, and consider API call relationships to improve the accuracy and practicality of API sequence recommendations. Summary of the Invention

[0007] To fully understand the functional requirements expressed by developers in natural language, obtain code context information, and recommend API call sequences that are correct in order and meet user needs, this invention proposes an API sequence recommendation method based on user intent awareness, comprising the following steps:

[0008] Training samples are constructed based on multi-source information, which includes at least: natural language description, code context, program structure representation, and corresponding API call sequence;

[0009] Furthermore, the natural language description is obtained by cleaning, segmenting, removing stop words, and stemming the source code comments.

[0010] Furthermore, the code context includes the method signature of the current method and the first N lines of code located at the beginning of the method, where N is a preset positive integer.

[0011] Furthermore, the program structure is represented by a feature sequence generated by a depth-first traversal of the program dependency graph of the code fragments. The traversal priority order is as follows: backtracking edges are superior to forward edges, control dependency edges are superior to data dependency edges, and node priority is determined by the order in which statements appear.

[0012] By using a pre-trained language model to jointly model the natural language description and program structure representation, a unified representation that integrates user intent and code semantics is obtained;

[0013] Furthermore, the pre-trained language model is CodeBERT, and the joint modeling includes:

[0014] Fine-tune CodeBERT by simultaneously inputting (comment, code) pairs and (comment, feature sequence) pairs;

[0015] The binary cross-entropy loss is used to measure the semantic similarity between annotations and code or feature sequences.

[0016] Based on the unified representation, candidate API sequences are predicted in an autoregressive manner using a sequence generation model;

[0017] Furthermore, the sequence generation model consists of a fine-tuned CodeBERT as the encoder and a Transformer decoder as the decoder, and the word-by-word generation of API sequences is supervised by cross-entropy loss during training.

[0018] The training phase employs the AdamW optimizer, along with a learning rate warm-up and linear decay scheduling strategy.

[0019] During the inference phase, query expansion is used to semantically enhance the input natural language query in order to generate expanded queries;

[0020] Furthermore, the query expansion includes:

[0021] Collect question-and-answer pairs from question-and-answer communities and build an index;

[0022] Stop words are removed and word segments are split from the original query;

[0023] Retrieve Top-K question-answer pairs as pseudo-relevance feedback;

[0024] Candidate API classes are extracted based on TF-IDF and PageRank weights;

[0025] Calculate the semantic similarity between the query and candidate API classes using a word embedding model;

[0026] The Borda count and semantic similarity score are combined to generate an expanded query.

[0027] Furthermore, the word embedding model is obtained by training fastText on the question-answering corpus in a Skip-gram manner.

[0028] The extended query and the current code context are input into the sequence generation model, which outputs an API call sequence that matches the user's intent.

[0029] Furthermore, the API sequence retains only API calls belonging to the standard library and excludes APIs from third-party libraries.

[0030] The beneficial effects of this invention are:

[0031] This invention's method comprehensively utilizes natural language queries and code context information. It not only extracts development intent from user-input natural language descriptions but also models the programming environment in conjunction with the current code context, thereby generating API recommendations that better match actual development needs, significantly improving the relevance and usability of the recommendations. By transforming the program dependency graph into a feature sequence for modeling, it effectively captures structural information such as the call order and dependencies between APIs, ensuring that the generated API sequences possess a reasonable logical order and execution semantics, thus improving the completeness and executability of the recommendation results. The introduction of a query expansion mechanism semantically enhances and expands the user-input natural language query, improving the model's depth of understanding of user intent. This allows for a more comprehensive coverage of functionalities not explicitly expressed by the user but actually required, enhancing the accuracy and context adaptability of the recommendations. Attached Figure Description

[0032] Figure 1 This is an overall flowchart of the API sequence recommendation method based on user intent awareness of the present invention;

[0033] Figure 2 This is an example diagram illustrating the transformation of the program dependency graph into a feature sequence according to the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0035] like Figure 1 As shown, this invention provides a user intent-aware API sequence recommendation method, comprising the following steps:

[0036] S1: Retrieve Java files with the ".java" extension from the open-source code repository.

[0037] S2: Extract the code comment, code context, code snippet, and API sequence of each Java method in the Java file, and construct a dataset in which code comments, code context, code snippets, and API sequences correspond one-to-one;

[0038] Furthermore, the specific implementation of step S2, which extracts code comments, code context, code snippets, and API sequences, is as follows:

[0039] S21: Extract code comments and preprocess them, specifically including filtering out empty comments or comments containing only one word; excluding comments for specific methods such as setters, getters, constructors, test methods, and overridden methods; selecting only the first line of the comment; and removing HTML tags from code comments, such as... Tags, etc.; remove various noise elements, including numbers, punctuation, and links, from the comments; convert all English text to lowercase; remove stop words from the comments based on a recognized list of English stop words; perform stemming operations to restore each word to its basic form.

[0040] S22: Extract code context and code snippets, that is, divide the method-level code snippets into two parts. The first part is the method declaration and the first three lines of code, which simulates the developer's current code context. The second part is the remaining code after removing the method declaration and the first three lines of code, which contains the API sequence that the developer needs to call to complete the software function. In addition, data with less than 4 lines of code are discarded.

[0041] S23: Obtain the API sequence. That is, for the second part obtained in S22, use the code parsing tool JDT to convert the second part of the code into an abstract syntax tree, and traverse all nodes in the method body of the abstract syntax tree to obtain the corresponding API sequence.

[0042] S24: Preprocess the obtained API sequence, i.e. retain the two types of APIs: method calls and instance creation; remove third-party APIs and retain APIs from the Java standard library. These retained APIs are usually prefixed with "java", "javax" or "org".

[0043] S25: Combine the preprocessed code comments, code context, code snippets, and API sequences into a one-to-one correspondence.

[0044] S3: Generate a program dependency graph of code snippets and transform it into a feature sequence.

[0045] Furthermore, the specific implementation of step S3, which transforms the program dependency graph of the code snippet into a feature sequence, is as follows:

[0046] Specifically, for a given program dependency graph, the root node is selected first. As the starting node of the entire program dependency graph sequence Then from node To begin, recursively select an edge to expand according to the following rules:

[0047] (1) Backtracking edges (edges whose endpoints have been traversed) are selected before forward edges (edges whose endpoints have not been traversed);

[0048] (2) When there are multiple backtracking edges or only multiple forward edges, you can choose to traverse the control dependency edge or the data dependency edge after the first step. If both types of edges exist at the same time, the control dependency edge should be selected first.

[0049] (3) If the edge cannot be determined after the second step, the edge with the highest endpoint priority is selected as the edge to be expanded first. The priority of the node is determined by the order of its corresponding statements in the code snippet.

[0050] (4) In particular, if there are no connected edges, the traversal will continue to the nearest ancestor node, which has at least one edge that has not been traversed.

[0051] Simultaneously, a sequence form of the program dependency graph is generated according to the following rules:

[0052] In this embodiment of the invention, if the traversed edge 𝑒 is a control dependency edge, then only the attributes of the node pointed to by 𝑒 are added to the sequence P.

[0053] In this embodiment of the invention, if the traversed edge 𝑒 is a data-dependent edge, then the attributes of 𝑒 and the attributes of the node pointed to by 𝑒 are added to the sequence P in sequence.

[0054] like Figure 2 As shown in the figure, this embodiment of the invention illustrates a specific example of converting a program dependency graph into a feature sequence, where "a", "b", "c", "d", and "f" all represent the corresponding lines of code. ( (where is a natural number) represents a data dependency edge.

[0055] S4: Use both comment-code pairs and comment-feature sequence pairs as input for fine-tuning CodeBERT.

[0056] Furthermore, step 4 fine-tunes CodeBERT. Specifically, each code snippet is broken down into two pairs, namely ( , )and( , ). , and The structure is as follows:

[0057]

[0058]

[0059]

[0060] in, Representing natural language annotations, each It is a word or phrase in a natural language description. Represents programming language code, each It is a code markup in the code. Represents the feature sequence generated by the program dependency graph, each It is an element in the feature sequence.

[0061] In this embodiment of the invention, all ( , ) and ( , For each input into the CodeBERT model, a binary cross-entropy loss function is used to fine-tune the model, where a softmax layer is used to map the output class label vector to scores between [0,1].

[0062] Specifically, the fine-tuning task consists of two parts. One is natural language description. With characteristic sequence The similarity between them, another is With code sequence To assess their similarity, both used a similar loss function, calculated as follows:

[0063]

[0064]

[0065] in, Indicates model parameters; Represents the training set; It is a commonly used slack variable in machine learning, usually set to 0.05; Representing natural language annotations; It can be a code sequence or characteristic sequence ; express Word embedding; express Word embedding. All All are positive samples, and and These are negative samples, which have an equal number of instances and are replaced by random substitutions. In or for or To create.

[0066] S5: Train the model using a finely tuned CodeBERT encoder and a Transformer decoder.

[0067] Further, in step 5, the model is trained using the finely tuned CodeBERT as the encoder and the Transformer's decoder as the decoder. The specific training process is as follows:

[0068] S51: Input natural language annotations and code context into the model's CodeBERT encoder;

[0069] S52: The Transformer decoder takes word vectors obtained by linear transformation of word vectors generated by the encoder's code annotations and code context as input, and generates the target API sequence step by step through an autoregressive approach;

[0070] S53: Cross-entropy loss is used to supervise the training of the predicted API sequence and the target API sequence, and AdamW is used as the optimizer to control gradient updates and improve the generalization ability of the model.

[0071] S54: Employs a learning rate scheduling strategy of learning rate warmup and linear decay. This involves gradually increasing the learning rate in the early stages of training to stabilize the model and gradually decreasing it in the later stages to prevent overfitting.

[0072] S6: Extend existing queries through query expansion, using the expanded queries as input, and let the model output the API sequence required by the user.

[0073] Furthermore, step S6 expands the query, specifically as follows:

[0074] S61: Collect question-and-answer pairs related to programming tasks from Stack Overflow and build a corpus;

[0075] S62: Perform natural language preprocessing on the corpus constructed in step S61, including removing stop words, punctuation marks and programming keywords, as well as operations such as word segmentation;

[0076] S63: Use the Lucene search engine to index the corpus that has undergone natural language preprocessing;

[0077] S64: For a given natural language query, after performing standardized preprocessing such as stop word removal and lexical segmentation, retrieve the Top-K question-answer pairs from the Stack Overflow corpus as pseudo-relevance feedback, that is, the search results are temporarily assumed to be relevant regardless of whether they are actually relevant;

[0078] S65: Extract code snippets from Top-K question-answer pairs and use TF-IDF and PageRank weighting methods to evaluate the importance of API classes and generate a candidate API class list;

[0079] S66: Use the fastText method to train a word embedding model on the corpus data that has undergone natural language preprocessing in step S62, and use the Skip-gram model to learn the word embedding model, mapping each word to a vector in a high-dimensional semantic space;

[0080] S67: Obtain the query-API semantic similarity score by calculating the semantic similarity between the vector of the query keyword and the vector of the candidate API class in the candidate API class list;

[0081] S68: Combining Borda score and semantic similarity score, the candidate API classes are sorted, and the most relevant API classes are selected as part of the query expansion. These API classes are then appended to the original query to form a rewritten query for API sequence recommendation.

[0082] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

[0083] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for recommending API sequences based on user intent awareness, characterized in that... The method includes the following steps: Training samples are constructed based on multi-source information, which includes at least: natural language description, code context, program structure representation, and corresponding API call sequence; By using a pre-trained language model to jointly model the natural language description and program structure representation, a unified representation that integrates user intent and code semantics is obtained; Based on the unified representation, candidate API sequences are predicted in an autoregressive manner using a sequence generation model; During the inference phase, query expansion is used to semantically enhance the input natural language query in order to generate expanded queries; The extended query and the current code context are input into the sequence generation model, which outputs an API call sequence that matches the user's intent.

2. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The natural language description is obtained by cleaning, segmenting, removing stop words, and stemming source code comments.

3. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The code context includes the method signature of the current method and the first N lines of code at the beginning of the method, where N is a preset positive integer.

4. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The program structure is represented by a feature sequence generated by a depth-first traversal of the program dependency graph of code fragments. The traversal priority order is: backtracking edges are superior to forward edges, control dependency edges are superior to data dependency edges, and node priority is determined by the order in which statements appear.

5. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The pre-trained language model is CodeBERT, and the joint modeling includes: Fine-tune CodeBERT by simultaneously inputting (comment, code) pairs and (comment, feature sequence) pairs; The binary cross-entropy loss is used to measure the semantic similarity between annotations and code or feature sequences.

6. The API sequence recommendation method based on user intent awareness according to claim 5, characterized in that: The sequence generation model consists of a fine-tuned CodeBERT as the encoder and a Transformer as the decoder. During training, cross-entropy loss is used to supervise the word-by-word generation of API sequences.

7. The API sequence recommendation method based on user intent awareness according to claim 6, characterized in that: The AdamW optimizer is used during the training phase, along with a scheduling strategy that combines learning rate warm-up and linear decay.

8. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The query expansion includes: Collect question-and-answer pairs from question-and-answer communities and build an index; Stop words are removed and word segments are split from the original query; Retrieve Top-K question-answer pairs as pseudo-relevance feedback; Candidate API classes are extracted based on TF-IDF and PageRank weights; Calculate the semantic similarity between the query and candidate API classes using a word embedding model; The Borda count and semantic similarity score are combined to generate an expanded query.

9. The API sequence recommendation method based on user intent awareness according to claim 8, characterized in that: The word embedding model was trained using fastText on a question-and-answer corpus in a Skip-gram manner.

10. The API sequence recommendation method based on user intent awareness according to claim 1, characterized in that: The API sequence retains only API calls belonging to the standard library and excludes APIs from third-party libraries.