Repository level augmentation for prompts for code completion

By incorporating repository-level contextual data, including few-shot examples and focus context, into a large language model, the problem of inaccurate predictions on a private repository by large language models is solved, achieving more efficient code completion.

CN120936981APending Publication Date: 2025-11-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480022160.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-28
Filing Date
2024-04-06
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Large language models perform poorly on source code in private repositories and are unable to effectively predict completion candidates because the training data lacks relevant methods, categories, and types.

Method used

By incorporating repository-level contextual data into the prompts, including few-sample examples and focus context, the large language model is guided to generate candidates. The prompts are augmented to improve prediction accuracy by leveraging method signatures and namespace information from the repository database.

Benefits of technology

It improves the accuracy of code completion candidates and reduces the consumption and cost of computing resources without the need to train or fine-tune a large language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120936981A_ABST
    Figure CN120936981A_ABST
Patent Text Reader

Abstract

A code completion system uses a large language model to complete source code fragments formed by part of a source code program given prompts including a repository-level context, an extended context, and a local context. The context of the repository-level extension includes a few-sample instance and a focus context. A few sample example is a code segment from a repository that has a close similarity to a partially formed source code segment. The focus context includes method signatures and namespace information for methods of custom categories defined in the repository. Expansion of cues with various contextual data enables the model to predict more relevant code completion candidates for the custom data without training the model on the custom data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Software development environments (SDAs) are typically used to assist software developers (i.e., users, programmers, etc.) in developing program code. An SDA can include a source code editor and other tools that developers use to write and test their programs. Some SDAs include code completion features, which provide assistance while the developer is editing code by automatically presenting a list of possible candidates based on one or more characters (e.g., letters, symbols, etc.) that the developer has already typed into the source code editor. A pop-up menu may appear along with several suggested code elements that the developer can utilize. This assistance is beneficial because it speeds up development time and reduces common errors, such as typos.

[0002] Given a partial source code snippet, code completion features can leverage a large language model to predict candidates to complete the snippet. Typically, a large language model is trained on a large training dataset of source code to learn the source code needed to predict the complete source code snippet. Large training datasets often consist of source code from publicly available code repositories. However, large language models perform poorly when used with source code from private repositories containing methods, categories, and types not visible in the training dataset. Summary of the Invention

[0003] This summary is provided to introduce, in a simplified form, some concepts that will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0004] The large language model is augmented with repository-level context, comprised of few-shot examples and focus context, to generate candidates for completing partially formed source code snippets. The large language model is pre-trained on publicly available source code. Few-shot examples are code snippets from source code files in a private repository that have close similarity to the partially formed source code snippets and are not part of the large language model's training dataset. Few-shot examples include data associated with the code snippets, such as suffix codes following the snippets and method signatures and namespace information associated with the methods containing the snippets.

[0005] Focus context includes the method signature and namespace information of methods of custom categories defined in the repository. In the absence of training the model on the task, focus context from few-shot examples and hints is used to guide how a large language model performs code completion tasks.

[0006] In addition, the hints include both the local context and the extended context. The local context includes the context of the current scope from the completion point, and the extended context includes method signatures and namespace information that are defined in the file but are not included in the local context.

[0007] These and other features and advantages will become apparent from reading the following detailed description and reviewing the associated drawings. It should be understood that the foregoing general description and the following detailed description are illustrative only and are not intended to limit the claimed aspects. Attached Figure Description

[0008] Figure 1 This is a schematic diagram illustrating an exemplary system for repository-level context extensions for hints used for code completion.

[0009] Figure 2 This is a schematic diagram illustrating an exemplary system for generating a repository (repo) database of few sample examples.

[0010] Figure 3 This is a schematic diagram illustrating a search for a small number of samples.

[0011] Figure 4 This is a schematic diagram illustrating an exemplary configuration of a large language model configured as an attention-based decoder neural converter model.

[0012] Figure 5 This is a flowchart illustrating an exemplary method for a system that provides repository-level contextual extensions for code completion suggestions.

[0013] Figure 6 This is a flowchart illustrating an exemplary method for creating a repository database.

[0014] Figure 7 This is a flowchart illustrating an exemplary method for generating code completion candidates.

[0015] Figure 8 This is a block diagram illustrating an exemplary operating environment. Detailed Implementation

[0016] summary

[0017] Various aspects of this disclosure relate to augmenting hints for a large language model using repository-level context data for completing partially formed source code snippets, the repository-level context data being derived from a private repository associated with a target source code program containing the partially formed source code snippets.

[0018] Large language models trained for code completion typically utilize the context immediately preceding the current cursor position (i.e., the completion point) or the partially formed source code snippet. However, the context required to predict accurate completion candidates often comes from outside the target source code program. To enable the model to predict relevant candidates, it is crucial to incorporate custom data (e.g., method signatures, methods, categories, namespaces) from private repositories, directories, and / or projects that are not visible to the large language model.

[0019] Repository-level contextual data is incorporated into the prompts given to the large language model, guiding the model toward generating candidates aligned with the customized contextual data. Repository-level context includes few-shot examples and focus context. Few-shot examples are code snippets from source code files in a private repository that have a close similarity to partially formed source code fragments. Without training the model, the few-shot examples and focus context from the prompts are used to explicitly guide the large language model on how it should perform the code completion task.

[0020] The focus context includes method signatures and namespace information from the repository. A namespace is a declarative area that provides scope for identifiers (names of types, functions, variables, etc.) within it. Namespaces are used to organize code into logical groups and to prevent potential name conflicts, especially when a codebase includes multiple libraries. Namespace information includes the module header / definition of the namespace, the namespace definition, and custom category definitions.

[0021] The focus context includes the method signatures and namespace information of custom class methods defined in the repository. The method signatures and namespace information of custom class methods defined in the repository and invoked in the program are ordered based on the distance from the call point to the completion point. The method signatures and namespace information of custom class methods defined in the repository but not invoked in the program are added to the beginning of the prompt in random order.

[0022] The hint also includes an extended context and a local context. The extended context includes the method signatures and namespace information of the custom classes defined in the current file, which is prioritized in terms of distance from the completion point. The local context includes the method signature of the method at the current cursor position and the method body of the method from the current cursor position to the current cursor position.

[0023] Few-sample examples are stored in the repository database, and custom method signatures and namespace information are extracted from files in a private repository associated with the target source code program. Code segments from the repository files, along with their associated method signatures and namespace information, are extracted and stored in the repository database. Code segment embeddings are used to index each entry in the repository database. The repository database is searched to obtain the top-most closely related source code snippets. k Each code segment and related data.

[0024] It should be noted that the descriptions in this article use terminology associated with the Python programming language. However, it should be noted that the techniques disclosed herein are not limited to the Python programming language and can be applied to any programming language.

[0025] Now let’s turn our attention to a more detailed description of the extended systems, methods, and components for code completion suggestions.

[0026] system

[0027] Figure 1 A block diagram of an exemplary code completion system 100 with repository-level extensions for code completion suggestions is shown. System 100 includes a source code editor 102 and a code completion system 104.

[0028] Source code editor 102 may be part of an integrated development environment (“IDE”), application, or tool used to develop, test, or maintain software. In one aspect, source code editor 102 may include user interface 106 and parser 108. User interface 106 includes a set of features or functions for developing (e.g., writing, editing, testing) source code programs. User interface 106 may utilize pop-up windows to present a list of possible candidates 110 for completion, allowing the developer to browse candidates and select one from the list. Alternatively, as the user is typing characters into the source code program, the candidate may appear in sync with the current line of source code.

[0029] Parser 108 reads characters input into the source code program via source code editor 102 and generates the corresponding concrete syntax tree 112. As the developer creates and edits source code in source code editor 102, parser 108 also updates concrete syntax tree 112.

[0030] At certain points during the editing process, the user interface 106 will request candidates to complete the source code at the current cursor position. The user interface can detect that the user has entered a specific character or string and automatically initiate a request for a candidate to complete the partially formed source code segment. This character is called a marker character. Subsequently, the user interface 102 will send a request for candidates to present to the developer's query 114. Alternatively, the user can request candidates by entering a specific keystroke or keystroke sequence (e.g., a combination of the control (CTRL) key and the space bar).

[0031] In another aspect, the system can automatically display a single top candidate at the end of the current source code line in a dim color, regardless of the marker character. The system builds and continuously updates the candidate tree in the background, regardless of whether the user decides to trigger the candidate. Candidates are automatically displayed in the user interface when the developer has been idle for some time. If the developer wants to accept a candidate, they can type a specific keystroke or combination of keystrokes (e.g., CTRL and I) to accept it. In this case, the cursor position will move to the end of the suggested code sequence, and the dim color of the candidate code will change to the normal color of the code. If the developer does not want to use the candidate, it disappears as the user continues typing. In this case, the system refines the code sequence based on a prefix filter of the candidate tree based on the newly typed code.

[0032] Code completion system 104 tracks characters input into the source code editor and serves queries or requests 114 for candidates to complete the code at the completion location. Code completion system 104 includes a code completion engine 116, a suggestion generator 118, a decoding engine 120, and a large language model 122. Code completion engine 116 receives queries 114 for candidates to complete the source code fragments and concrete syntax trees 112 formed by the portions of the source code currently residing in source code editor 102. Suggestion generator 118 constructs suggestions 124 for the large language model 122 to autoregressively generate candidates to complete the portion of the source code fragments. Candidates are sorted according to their corresponding probabilities, with the highest probability candidates at the top. Subsequently, a selected number of candidates 110 are returned to source code editor 102 and displayed in user interface 106.

[0033] The prompt generator 118 generates a prompt 124 for the large language model 122, which includes a focus context, a few-shot example, an extended context, and a local context. The prompt generator 118 utilizes an encoder 134 to generate an encoding or embedding of a token for a source code fragment of a query used to search the repository database 136 to find few-shot examples.

[0034] Decoding engine 120 performs a search for candidates to complete the partially formed code snippet. Searching for all possible candidate output sequences based on probability is an NP-hard completion search. Instead, the decoding engine uses a heuristic search algorithm that approximates the best candidates. Decoding engine 120 can generate candidates using beam search, core sampling, random sampling, temperature-controlled random sampling, and / or top-k sampling.

[0035] In one aspect, the large language model 122 is a neural converter model configured with attention to decoder blocks. The attention-based decoder neural converter model is pre-trained on source code programs and source code comments (i.e., natural language text). The attention-based decoder neural converter model is an autoregressive model that produces output one token at a time based on the output of previous time steps. Code completion is best suited for decoder neural converter models because it is an autoregressive task of predicting ordered sequences of tokens, where the order depends on previous tokens in the sequence. Examples of attention-based decoder neural converter models include GitHub's Copilot model, OpenAI's GPT model, and others.

[0036] In one aspect, the large language model is a publicly accessible model located on an external server. The decoding engine 120 may reside on the same external server as the large language model, or within the same computing device as the code completion engine. The code completion engine 116 communicates with the decoding engine 120 over a network via an application programming interface (API).

[0037] Figure 2 A system 200 used to generate a repository (“repo”) database is shown. (Reference) Figure 1 and Figure 2 The repository database 136 contains source code segments from associated file sets 204, which are used to create software applications or services that are not publicly accessible. The associated file sets can be source code repositories or portions of projects associated with source code programs in source code editors. The source code repository 202 is a privately archived file and web hosting facility for storing large amounts of source code. The source code repository 202 can be configured as a version control system, such as Git, Mercurial, etc. An IDE project is a collection of associated files (such as portions of an application or service). The source code repository 202 and projects can include source code files, documentation files, scripts, tests, etc.

[0038] The repository database generator 206 extracts modules, categories, and methods from various files in the private repository 202. Code segments are extracted from each file 204 in the repository 202. The files include modules, categories, and methods used in various source code programs within the private repository. Each file in the repository database 202 containing source code is parsed into a concrete syntax tree. Byte-pair encoding tokenization is used to extract code segments consisting of pre-configured sizes (such as 256 tokens) from the concrete syntax tree. An encoder 134 is used to generate embeddings (encodings (Ci)) of the code segments, which are used as indexes to the repository database for the code segments. Each entry in the repository includes the code segment (Ci), its method signature and namespace information (H), and the line of code (S) following the code segment.

[0039] Figure 3 A system 300 is shown for retrieving small sample examples from a storage database. (Reference) Figure 1 and Figure 3 System 300 utilizes encoder 134, hint generator 118, and repository database 136. Encoder 134 encodes a query 302 containing source code into an embedded encoding (query), and hint generator 118 uses this embedding to search for closely matching embeddings in repository database 136. In one aspect, cosine similarity is used to determine the similarity between the embedding of the query and the embedding of each code segment in the repo database. The closest matching embedding is used to extract top-... k code snippets ( C k ), and its associated method signature and namespace information (H k ), and the continuation of code segment suffixes (S k ).

[0040] As shown in box 312, top-ranking is done by the closest similarity. k The code snippets and related data are sorted. Box 312 contains the top-ranked snippets sorted by descending similarity. k A few-sample instance. Each few-sample instance contains namespace information and a method signature (H). i ), associated code segment (C i ), and the suffix code that follows the code segment until the end of the code block (S i As shown in box 312, the closest few-sample example contains a module header and signature ( module_name_of_ example_k ) and custom category definitions ( class_name_of_example_k ), and the associated method signature ( defmethod_name_of_example_k(args) Namespace information composed of )

[0041] Figure 4An exemplary configuration of a large language model as an attention-based decoder neural converter is shown. A large language model is a deep machine learning model containing billions or more parameters. The parameters are portions of the model learned from a training dataset that defines the model's skills used to generate predictions for the target task.

[0042] Deep machine learning models differ from traditional machine learning models that do not use neural networks. Machine learning involves the use and development of computer systems that analyze patterns in data and derive inferences from those patterns using algorithms and statistical models. These computer systems are capable of learning and adapting without following explicit instructions. Machine learning uses different types of statistical methods to learn from data and predict future decisions. Traditional machine learning includes statistical techniques, data mining, Bayesian networks, Markov models, clustering, support vector machines, and data visualization mapping.

[0043] Deep machine learning differs from traditional machine learning because it uses multiple stages of data processing through many hidden layers of neural networks to learn and interpret relationships between features. Deep machine learning embodies neural networks in a way that differs from traditional machine learning techniques that do not utilize neural networks. Various types of deep machine learning models exist that generate source code, such as recurrent neural network (RNN) models, convolutional neural network (CNN) models, long short-term memory (LSTM) models, and neural transducers with attention.

[0044] The neural decoder-transformer model 400 includes multiple stacked decoder blocks 402A to 402N (“402”). Decoder 400 uses all previously generated target tokens. t 1 … t i-1 As a condition, predict each token in the target language one by one at each time step. t i Each decoder block 402 consists of two layers. The first layer includes a mask multi-head self-attention unit 404 followed by a layer normalization unit 406. The output of the layer normalization unit 406 is fed into a second layer, which includes a feedforward neural network 408 with residual connections to the layer normalization unit 410.

[0045] The masked multi-head self-attention unit 404 receives the output embeddings from previous time steps. The masked multi-head self-attention unit 404 masks the output embeddings from future time steps. The feedforward neural network 408 processes each output encoding separately. Layer normalization units 406 and 410 are used between layers to normalize the input across features.

[0046] Output layer 412 comprises a linear layer 414 and a softmax layer 416. Linear layer 414 projects the vectors generated by the decoder stack onto a logical value vector. Softmax layer 412 then converts the scores of the logical value vectors into a vocabulary-based algorithm. V The output probability of each token in the array is positive and normalized to 418.

[0047] The input layer 420 of the first decoder block 402A includes an input embedding layer 422 containing embeddings of the input sequence, a positional embedding layer 424, and a context tensor 426. The positional embedding 424 is used to maintain the order of tokens in the input sequence. The context tensor 426 contains positional embeddings added to the input embedding 422.

[0048] During inference, the initial input to the first decoder block 404A contains a <start> token and a cue 428, which includes a focus context 430, a few-sample example 432, an extended context 434, and a local context 436. At each subsequent time step, the input is a shifted sequence of the output embeddings from the previous time step, and positional embeddings are added to this shifted sequence to form the context tensor 426.

[0049] method

[0050] Attention now turns to a more detailed description of the methods used in the system. It will be understood that, unless otherwise instructed, the representative methods need not necessarily be performed in the presented order or in any particular order. Furthermore, the various activities described regarding the methods can be performed serially or in parallel, or any combination of serial and parallel operations. In one or more aspects, the methods illustrate operation with respect to the systems and devices disclosed herein.

[0051] Figure 5 This is an example method for generating a System 500 error. (See reference.) Figure 1 and Figure 5 A large language model 122 and an encoder 134 are obtained (box 502). In one aspect, the large language model 122 is a decoder-neural converter model with attention, and the encoder 134 is an encoder-neural converter model with attention (box 502). A storage library is identified (box 504), and a storage library database 136 is generated for the storage library (box 506). When the storage library database is generated, the encoder 134, the storage library database 136, and the large language model 122 are deployed in a code completion system 104 (box 508) for use in prediction candidates to complete partially formed source code fragments (box 510).

[0052] Figure 6An exemplary method 600 for generating a storage database for a specific storage database is shown. (Reference) Figure 2 and Figure 6 The method processes each file 204 (box 602) in the source code repository 202. The source code in each file is parsed into a concrete syntax tree (box 604). Extraction is performed from the concrete syntax tree of a predetermined length. T A sequence of tokens and / or sub-tokens is used, and this is treated as a code segment (box 606). Subsequently, encoder 134 is used to decode the code segment... T An ordered sequence of tokens of length is mapped into a numeric vector, and then into an embedding (box 608).

[0053] Obtain the method signature of the method containing the code segment and the namespace information associated with the code segment (box 610). Subsequently, store the code segment, the method signature of the method containing the code segment, and the namespace information associated with the code segment in a store database 136 indexed by the embedding 212 of the code segment (box 612).

[0054] Figure 7 This is an exemplary method 700 for generating prompts for large language models. (See reference...) Figure 1 and Figure 7 The code completion system 104 receives queries for candidates to complete partially formed code snippets (box 702). Query 114 consists of tokens of a predetermined length preceding the current cursor position or completion point. The partially formed code snippet can be a partially formed method signature, a partially formed expression, a partially formed method body, etc.

[0055] Hint generator 118 obtains a local context 132, an extended context 130, a few-sample example 128, and a focus context 126. Local context 132 includes the method signature of the method at the completion point and the method body up to the completion point (box 704). Hint generator 118 obtains an extended context, which includes the method signature and namespace information defined in the current file outside the local context (box 706). Extended context 130 is prioritized by the absolute line spacing up to the completion point, arranged in descending order (box 706).

[0056] The prompt generator 118 obtains few-sample examples from the repository database (box 708). An encoder is used to generate embeddings for the query (encoder(Q), where Q is the query). The prompt generator 118 calculates the cosine similarity between the embeddings for the query and the embeddings for each code segment in the repository database 136. The cosine similarity measure, or L2-normalized Euclidean distance, calculates the distance between two embeddings as the difference of squared vector values, which is represented as:

[0057] ,in Q It is a query, and c i It is a code segment in the repo database.

[0058] The prompt generator 118 obtains a focus context, which includes the method signatures and namespace information of the custom class methods defined in the repository (box 710). Based on the distance from the call point to the completion point, it prioritizes the method signatures and namespace information of the custom class methods defined in the repository and invoked in the program. It then randomly places the method signatures and namespace information of the custom class methods defined in the repository but not invoked in the program at the beginning of the prompt.

[0059] The prompt generator 118 then assembles the prompt in the following order: focus context 126, few-shot example 128, extended context 130, and local context 132 (box 712). The prompt is sent to the decoding engine 120, which applies the prompt 124 to the large language model 122 (box 714). The large language model 122 interacts with the decoding engine 120 to generate completion candidates 110, which are then returned to the user interface 106 (box 716).

[0060] Technical effects / technical improvements

[0061] The aspects of the subject matter disclosed in this paper relate to the technical problem of generating prompts for large language models to generate candidates for completing parts of source code snippets with custom source code. Technical features associated with addressing this problem include incorporating repository-level context into the prompts. The achieved technical effect is increased accuracy in predicting code completion candidates without the computational burden of training or fine-tuning large language models on custom source code.

[0062] For feasibility, the code completion system must execute within strict time constraints. In scenarios where the large language model resides on an external server accessed via a network, the operations used to generate prompts must be performed on the computing device. Therefore, the operations performed are inherently digital. Human thought cannot directly interface with the CPU, network interface card, other processors, or RAM or digital storage devices to read and write the necessary data and perform the necessary operations and processing steps described herein.

[0063] It is also assumed that the embodiments are capable of operating "at scale" in a production environment or in a test laboratory for a production environment, i.e., capable of handling larger volumes, rather than being considered merely as experiments.

[0064] The technique described in this paper is an improvement on existing solutions that utilize local context as a cue for a large language model or to fine-tune the large language model using custom data. Local context alone is insufficient for a large language model to make predictions on private data. Fine-tuning a large language model using custom data is not always possible due to the significant resources required to construct fine-tuning data and the cost of fine-tuning the large language model. In some scenarios, it is not possible to fine-tune a publicly accessible large language model with restrictions on its use. Enhancing the cue in the manner described in this paper avoids the costly fine-tuning steps and improves predictions by utilizing custom data to enhance the cue.

[0065] Exemplary operating environment

[0066] Now let’s turn our attention to the discussion of the exemplary operating environment 800. Figure 8 An exemplary operating environment 800 is shown in which one or more client computing devices 802 communicate with one or more computing devices. However, it should be noted that the aspects disclosed herein are not limited to any particular configuration of the computing devices. In an alternative embodiment, the large language model is hosted on an external server, and the code completion system is hosted on a separate server. The code completion system communicates with the large language model over a network using API(s), etc.

[0067] The computing device 802 can be any type of electronic device, such as, but not limited to, mobile devices, personal digital assistants, mobile computing devices, smartphones, cellular phones, handheld computers, servers, server arrays or server farms, web servers, network servers, blade servers, internet servers, workstations, microcomputers, mainframe computers, supercomputers, network devices, web devices, distributed computing systems, multiprocessor systems, or combinations thereof. The operating environment 800 can be configured in a network environment, a distributed environment, a multiprocessor environment, or a standalone computing device with access to remote or local storage devices.

[0068] Computing device 802 may include one or more processors 806, one or more communication interfaces 808, one or more storage devices 810, one or more memory devices or memories 814, and one or more input / output devices 812. Processor 806 may be any commercially available or custom processor and may include dual-microprocessor and multiprocessor architectures. Communication interface 808 facilitates wired or wireless communication between computing devices and with other devices. Storage device 810 may be a computer-readable medium that does not contain propagating signals (such as modulated data signals transmitted via a carrier wave). Examples of storage devices 810 include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, magnetic tape cassettes, magnetic tape, and disk storage devices, all of which do not contain propagating signals (such as modulated data signals transmitted via a carrier wave). Multiple storage devices 810 may be present in computing device 802. Input / output devices 812 may include a keyboard, mouse, pen, voice input device, touch input device, display, speaker, printer, etc., and any combination thereof.

[0069] The memory device or memory 814 can be any non-transitory computer-readable storage medium capable of storing executable programs, applications, and data. The computer-readable storage medium does not involve propagating signals (such as modulated data signals transmitted via a carrier wave). It can be any type of non-transitory memory device (e.g., random access memory, read-only memory, etc.), magnetic storage device, volatile storage device, non-volatile storage device, optical storage device, DVD, CD, floppy disk drive, etc., which does not involve propagating signals (such as modulated data signals transmitted via a carrier wave). The memory device 814 may also include one or more external storage devices or remotely located storage devices that do not involve propagating signals (such as modulated data signals transmitted via a carrier wave).

[0070] Memory device 814 may contain instructions, components, and data. A component is a software program that performs a specific function and is otherwise referred to as a module, program, and / or application. Memory device 814 includes an operating system 816, a source code editor 818, a code completion engine 820, a prompt generator 822, an encoder 824, a storage database 826, a decoding engine 828, a large language model 830, and other applications and data 832.

[0071] The computing device 802 can be communicatively coupled via network 804. Network 804 can be configured as an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the public switched telephone network (PSTN), a common old-style telephone service (POTS) network, a wireless network, a WiFi® network, or any other type of network or combination of networks.

[0072] Network 804 can employ a wide variety of wired and / or wireless communication protocols and / or technologies. The different communication protocols and / or technologies that can be adopted by the network at different generations may include, but are not limited to, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access 2000 (CDMA-2000), High-Speed ​​Downlink Packet Access (HSDPA), Long Term Evolution (LTE), Universal Mobile Telecommunications System (UMTS), Evolved Data Optimized (Ev-DO), Global Microwave Interconnection Access (WiMAX), Time Division Multiple Access (TDMA), Orthogonal Frequency Division Multiplexing (OFDM), Ultra Wideband (UWB), Wireless Application Protocol (WAP), User Datagram Protocol (UDP), Transmission Control Protocol / Internet Protocol (TCP / IP), any part of the Open Systems Interconnection (OSI) model protocol, Session Initiation Protocol / Real-Time Transport Protocol (SIP / RTP), Short Message Service (SMS), Multimedia Messaging Service (MMS), or any other communication protocols and / or technologies.

[0073] in conclusion

[0074] A system is disclosed, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for performing the following actions: obtaining a partially formed source code fragment of a source code program, wherein the source code program is associated with a repository having multiple files; extracting a local context of the partially formed source code fragment, wherein the local context includes the partially formed source code fragment; extracting a repository-level context from the repository, wherein the repository-level context includes multiple few-sample instances and a focus context, wherein the few-sample instances include code segments from the repository having a close similarity to the partially formed source code fragment, wherein the focus context includes method signatures of methods of a custom class defined in the repository, wherein the code segment is outside the source code program; creating a hint for a large language model to complete the partially formed source code fragment, wherein the hint includes the repository-level context and the local context; and generating at least one code completion candidate from the large language model, given the hint.

[0075] In one aspect, the focus context includes namespace information for methods of a custom category defined in the repository. In another aspect, one or more programs include additional instructions for performing the following actions: ranking multiple few-sample instances based on the closest similarity to a partially formed source code fragment; and selecting some of the few-sample instances from the ranked few-sample instances with the closest similarity.

[0076] In one aspect, one or more programs include additional instructions for performing the following actions: extracting the method signature, namespace information, and suffix code of each code segment of each of a few selected few examples from a sorted few examples that have the closest similarity to the source code fragment formed in part; and augmenting the hint with the extracted method signature, namespace information, and suffix code for each of the few selected few examples from the sorted few examples.

[0077] In one aspect, one or more programs include additional instructions to perform the following actions: extract an extended context from the source code program, wherein the extended context includes method signatures and namespace information of methods of custom classes defined in the source code program and outside the scope of the local context.

[0078] In one aspect, one or more procedures include additional instructions to perform the following actions: select some extended contexts in the extended context based on the nearest distance to the completion point; and expand the prompt using some of the selected extended contexts in the extended context.

[0079] In one aspect, the local context includes the method signature of the method containing the partially formed source code fragment and the method body of the method containing the partially formed source code fragment.

[0080] In one aspect, the large language model is a neural converter model with attention.

[0081] A computer-implemented method is disclosed, comprising: obtaining a partially formed source code fragment from a source code program, wherein the source code program is associated with a repository having multiple files; extracting a local context of the partially formed source code fragment, wherein the local context includes the context of the partially formed source code fragment; extracting a repository-level context from the repository, wherein the repository-level context includes at least one few-sample instance extracted from the repository and a focus context, wherein the at least one few-sample instance has the closest similarity to the partially formed source code fragment, wherein the focus context includes the method signature of a method of a custom category defined in the repository; and generating at least one code completion candidate from a large language model, given a hint having the repository-level context and the local context, to complete the partially formed source code context.

[0082] In one aspect, the computer-implemented method further includes: expanding the focus context using namespace information of methods of custom categories defined in the repository. In this aspect, the namespace information includes module definitions, namespace definitions, and custom category definitions of methods of custom categories defined in the repository.

[0083] In one aspect, the context of the partially formed source code fragment includes the method signature of the method containing the partially formed source code fragment and the method body of the method containing the partially formed source code fragment.

[0084] In one aspect, at least one few-sample example includes a code segment similar to a partially formed source code fragment, a method signature of a method containing the code segment, namespace information of the code segment, and a suffix code following the code segment.

[0085] In one aspect, the focus context includes the method signatures of custom-class methods defined in the repository and not invoked in the source code program.

[0086] In one aspect, the computer-implemented method also includes prioritizing the focus context based on the distance from the call point to the completion point, wherein the method signature of a custom class method defined in the repository and not called in the source code program is higher than the method signature of the custom class method defined in the repository.

[0087] A computer-implemented method is disclosed, comprising: accessing a large language model over a network to predict code completion candidates for a partially formed source code snippet, wherein the partially formed source code snippet is associated with a repository having multiple files; accessing a database of few examples, wherein the few examples include code segments from the repository, suffix codes following the code segments, and method signatures and namespace information associated with the code segments; selecting a subset of few examples that have few examples of code segments closely similar to the partially formed source code snippet; extracting a focus context for the partially formed source code snippet containing method signatures and namespace information of methods of a custom category defined in the repository; extracting a local context for the partially formed source code snippet; constructing a hint that includes the focus context of the partially formed source code snippet, the selected subset of few examples of the few examples, and the local context; sending the hint to the large language model for code completion candidates to complete the partially formed source code snippet; and receiving the code completion candidates from the large language model.

[0088] In one aspect, the large language model is an attention-based neural converter model. In another aspect, the computer-implemented method further includes: augmenting the prompt with an extended context of a partially formed source code fragment, wherein the extended context includes method signatures of methods of a custom class defined in the source code program and outside the scope of the local context. In another aspect, the focus context is prioritized based on the distance from the call point to the completion point, and the extended context is prioritized based on the distance between the call point and the completion point of the custom class method defined in the source code program.

[0089] Although the subject matter has been described in language specific to structural features and / or methodological actions, it is to be understood that the subject matter as defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as exemplary forms of implementing the claims.

[0090] It is understood that, unless otherwise instructed, the representative methods described herein must not necessarily be performed in the presented order or in any particular order. Furthermore, the various activities described regarding the methods can be performed serially or in parallel, or any combination of serial and parallel operations.

Claims

1. A system comprising: One or more processors; as well as A memory storing one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing the following actions: Obtain a portion of the source code program formed by the source code program, wherein the source code program is associated with a repository having multiple files; Extract the local context of the source code fragment formed by the portion, wherein the local context includes the source code fragment formed by the portion; Extract repository-level context from the repository, wherein the repository-level context includes multiple few-sample instances and focus contexts, wherein the few-sample instances include code segments from the repository that have a close similarity to source code fragments formed by the portion, and wherein the focus context includes method signatures of methods of custom categories defined in the repository, wherein the code segments are outside the source code program; Create hints for a large language model to complete the source code snippets formed by the aforementioned parts, wherein the hints include the repository-level context and the local context; as well as Given the aforementioned prompt, at least one code completion candidate is generated from the large language model.

2. The system of claim 1, wherein the focus context includes namespace information of methods of a custom category defined in the repository.

3. The system of claim 1, wherein the one or more programs include additional instructions for performing the following actions: The plurality of few-sample examples are ranked based on the closest similarity to the source code fragments formed by the aforementioned portions; and Select some of the sorted few sample examples that have the closest similarity.

4. The system of claim 3, wherein the one or more programs include additional instructions for performing the following actions: Extract the method signature, namespace information, and suffix code for each code segment of each of the selected few examples from the sorted few examples that have the closest similarity to the source code fragment formed with the part; and The hint is augmented using the extracted method signature, namespace information, and suffix code for each of the few examples selected from the sorted few examples.

5. The system of claim 1, wherein the one or more programs include additional instructions for performing the following actions: Extract the extended context from the source code program, wherein the extended context includes the method signatures and namespace information of methods of custom categories defined in the source code program and outside the scope of the local context.

6. The system of claim 5, wherein the one or more programs include additional instructions for performing the following actions: Select some extended contexts from the extended contexts based on the nearest distance to the completion point; and The prompt is expanded using selected extended contexts from the extended contexts.

7. The system of claim 1, wherein the local context includes a method signature of a method comprising a source code fragment formed by the portion and a method body of the method comprising the source code fragment formed by the portion.

8. The system of claim 1, wherein the large language model is a neural converter model with attention.

9. A computer-implemented method, comprising: A partial source code fragment is obtained from a source code program, wherein the source code program is associated with a repository having multiple files; Extract the local context of the source code fragment formed by the portion, wherein the local context includes the context of the source code fragment formed by the portion; Extract repository-level context from the repository, wherein the repository-level context includes at least one few-sample instance and a focus context extracted from the repository, wherein the at least one few-sample instance has the closest similarity to the source code fragment formed by the part, and wherein the focus context includes the method signature of the method of the custom category defined in the repository; as well as Given a hint with the repository-level context and the local context, at least one code completion candidate is generated from the large language model to complete the source code context formed by the portion.

10. The computer-implemented method according to claim 9, further comprising: The focus context is expanded using the namespace information of the methods of the custom categories defined in the repository.

11. The computer-implemented method of claim 10, wherein the namespace information includes the module definition, namespace definition, and custom category definition of the method for the custom category defined in the repository.

12. The computer-implemented method of claim 9, wherein the local context of the partially formed source code fragment includes a method signature of the method containing the partially formed source code fragment and a method body of the method containing the partially formed source code fragment.

13. The computer-implemented method of claim 9, wherein the at least one few-sample example includes a code segment similar to the source code fragment formed by the portion, a method signature of the method containing the code segment, namespace information of the code segment, and a suffix code following the code segment.

14. The computer-implemented method of claim 9, wherein the focus context includes the method signature of a method of a custom class defined in the repository and not invoked in the source code program.

15. The computer-implemented method of claim 14, wherein the focus context includes the method signature of a custom class of methods defined in the repository and invoked in the source code program.