Verification and provision of proactively generated code proposals

By employing a higher latency programming language proposal model with proactive verification and constrained prefix matching, along with subword regularization and evaluation dataset conversion, the solution addresses inefficiencies in code proposal generation, ensuring accurate and timely code suggestions.

JP2025524444AActive Publication Date: 2025-07-30AMAZON TECH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024575175
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-22
Filing Date
2023-06-21
Publication Date
2025-07-30
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

Existing code development tools face challenges in providing high-quality and timely code suggestions, especially when dealing with unfamiliar programming languages or contexts, leading to inefficiencies and inaccuracies in code proposal generation.

Method used

Implementing a higher latency programming language proposal model with proactive verification and constrained prefix matching techniques, along with subword regularization and evaluation dataset conversion, to enhance the accuracy and speed of code proposals.

Benefits of technology

The solution ensures that code proposals are valid and accurate, reducing perceived latency and improving the user experience by providing real-time, high-quality suggestions without sacrificing model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524444000001_ABST
    Figure 2025524444000001_ABST
Patent Text Reader

Abstract

Code completion proposals can be proactively obtained and verified. An event can be detected that triggers obtaining code completion proposals for inclusion in a code file being edited using an integrated development environment. Code completion proposals can be obtained. The characters of the code completion proposals can be compared with the characters added to the code file after the detection of the event that triggered the obtaining of the code completion proposals to determine whether the code completion proposals are valid. Then, valid code completion proposals can be displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Programming languages provide developers, designers, and other users with the ability to precisely specify the operation of various hardware or software designs for many different applications. Given a wide variety of programming languages, these developers, designers, and other users may encounter code written in a programming language that the developer may not be very familiar with or may use it differently. Code development tools provide developers, designers, and other users with different capabilities for improving code performance and identifying errors, which can help overcome the situation where the developer is not proficient in the programming language (or the environment in which the programming language is deployed) so that high-performance code can still be written in the exemplary scenarios above.

Brief Description of the Drawings

[0002]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17

[0003] Embodiments are described herein as examples of multiple embodiments and exemplary drawings. Those skilled in the art will recognize that the embodiments are not limited to the embodiments or drawings described. The drawings and their detailed descriptions are not intended to limit the embodiments to the specific forms disclosed. On the contrary, the intention is to cover all modifications, equivalents, and alternatives within the spirit and scope as defined by the appended claims. The headings used herein are for merely organizational purposes and are not meant to limit the scope of the description or the claims. As used throughout this application, the word "may" is used in a permissive sense (i.e., meaning having the possibility) rather than an obligatory sense (i.e., meaning must). Similarly, the words "include", "including", and "includes" mean including but not limited to.

Best Mode for Carrying Out the Invention

[0004] Various techniques for code generation for code development and verification are described herein. Sophisticated code development tools may rely on features powered by machine learning to assist in the design and development of new applications, systems, or services. To truly improve the user experience using these features, the speed and quality of the machine learning-enabled features can depend on various aspects of their development and implementation forms. One such machine learning-enabled feature for code development is code suggestion, which can generate code and recommend it to developers. Techniques for improving the speed and quality of machine learning-enabled code suggestions can improve the user experience as well as the quality of the applications, systems, or services generated using the features.

[0005] FIG. 1 is a logical block diagram showing code generation for code development according to some embodiments. An integrated development environment 110, such as a development application implemented locally on a user device (e.g., a computer or laptop) or hosted as part of a provider network service, may utilize code proposal generation 120 when a code file input is received. As will be discussed in detail below, code proposal handling 112 may proactively obtain and verify (114) code proposals 116 before providing them to display 104, as will be discussed in detail below with respect to FIGS. 3, 4, 8, and 9. In this way, a higher latency programming language proposal model 122, implemented as part of code proposal generation 120 to provide better and usable code proposals 116, can make the proactive request bring the apparent latency of the code proposal to 0 or near 0, even if their latency is longer than that of a model with lower latency but lower accuracy, while still guaranteeing through verification that the proposal is still valid in front of display 104 (in light of potentially changing contexts such as other code file inputs 102). As will be discussed below with respect to FIG. 3, other techniques, such as storing and providing paginated results, can also reduce the latency for waiting for code proposals to display 104.

[0006] Other performance improvements to the use and implementation of code proposal generation 120 can increase the quality and accuracy of proposals without affecting model performance. For example, as will be discussed in detail below with respect to FIGS. 4, 11, and 12, constrained prefix matching techniques can improve proposals in partial word scenarios without increasing the time to generate proposals in a meaningful way.

[0007] Code proposal development 130, such as training 134 and deployment 136 of the programming language proposal model 122, can improve the impact on code proposal performance. For example, as will be discussed in detail below with respect to FIGS. 6, 13, and 14, techniques for modifying the training dataset to train considering subwords can improve the code proposals made in those scenarios. Further, the evaluation of code proposal generation 120 and model 122 can be improved by increasing the number of high-quality evaluation datasets 132 that can be programmatically generated from other evaluation datasets, as will be discussed in detail below with respect to FIGS. 7 and 15 - 16B.

[0008] Note that the foregoing description is not intended to be limiting and is merely provided as an example of an integrated development environment, a code proposal generation system, and tools for code proposal development. As will be discussed in detail below, various other embodiments may implement these techniques.

[0009] This specification then includes a general description of a provider network that may implement code development services. Next, various examples of code development services are considered, including different components / modules, or configurations of components / modules, that may be employed as part of implementing code development services in a provider network. Next, several different methods and techniques are considered, some of which are shown in the accompanying flowcharts. Finally, a description of an exemplary computing system in which various components, modules, systems, devices, and / or nodes may be implemented is provided. Various examples are provided throughout this specification.

[0010] FIG. 2 is a logical block diagram showing a provider network implementing different services including code development services according to some embodiments. Provider network 200 (which may be referred to as a "cloud provider network" or simply "cloud" in some implementations) refers to a pool of network-accessible computing resources (such as computing, storage, and networking resources, applications, and services) that can be virtualized or bare metal. Provider network 200 can provide convenient on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and released in response to customer commands. These resources can be dynamically provisioned and reconfigured to adapt to variable loads.

[0011] The provider network 200 can be formed as several regions, where a region is a distinct geographical area where a cloud provider clusters data centers. Each region can include two or more availability zones connected to each other via a private high-speed network, such as a fiber-optic communication connection. An availability zone (also known as an availability domain or simply a "zone") refers to a separate failure domain that includes one or more data center facilities with separate power, separate networking, and separate cooling from those within another availability zone. Preferably, the availability zones within a region are positioned far enough apart from each other such that the same natural disaster should not take offline two or more availability zones simultaneously. A customer can connect to an availability zone of the provider network 200 via a publicly accessible network (e.g., the Internet, a cellular communication network). A region is connected to a global network that includes private networking infrastructure (e.g., a fiber connection controlled by a cloud provider) that connects each region to at least one other region. The provider network 200 can deliver content from points of presence that are external to these regions but networked with these regions, via edge locations and regional edge cache servers. This compartmentalization and geographical dispersion of computing hardware enables the provider network 200 to provide customers with low-latency resource access at a global scale with a high degree of fault tolerance and stability.

[0012] As described above, the provider network 210 can implement various other services 230 that can be any other type of network-based service, including various computing resources or services such as code development services 210, and various other types of storage (e.g., database services or object storage services), computing, data processing, machine learning, analytics, communication, event handling, visualization, and security services (not shown).

[0013] In various embodiments, the components shown in FIG. 2 may be implemented directly in computer hardware (e.g., as instructions executable directly or indirectly by computer hardware (e.g., a microprocessor or computer system), or using a combination of these techniques. For example, the components of FIG. 2 may be implemented by a system including several computing nodes (or simply nodes), each of which may be similar to the computer system embodiment shown in FIG. 17 and described below. In various embodiments, the functions of a given system or service component (e.g., a component of code development service 210) may be implemented by a particular node or may be distributed across multiple nodes. In some embodiments, a given node may implement the functions of more than one service system component (e.g., more than one data store component).

[0014] Code development service 210, in some embodiments, may be implemented by provider network 200. Code development service 210 may implement various features for writing code for different systems, applications, or devices, and may provide features for recommending, identifying, reviewing, building, and deploying code. For example, code development service 210 may implement development environment 211. Code development environment 211 may provide various code entry tools (e.g., text, diagram / graphics-based application development) for specifying, invoking, or otherwise writing (or having written) code for different hardware or software applications.

[0015] The code development service 210 may implement a code proposal delivery 214 that implements various computing resources to host and implement the code proposal 213 in a scalable manner to deliver on-demand code proposals across a number of clients using a high-performance machine learning model for high-quality code proposal results. For example, the code proposal delivery 214 may implement workload distribution and demand management features to timely handle and return code proposals to provide real-time code proposals with little or no apparent latency to the code proposal handling 220 (inside or outside the provider network 200).

[0016] To avoid having the development environment wait for multiple code proposals to be sent in one communication, in some embodiments, the code proposal delivery 214 may implement a paging feature for code proposals to enable multiple code proposals to be delivered over time via multiple communications from the host implementing the code proposal 213 or other computing resources to the recipient development environments 219 and 211. In this way, a valid code proposal can be created, presented, and then updated as more are received. Such techniques provide a simulated streaming experience without actually requiring the bidirectional streaming supported in the development environment. In this way, the benefits of fast delivery and updating of code proposals can be provided without introducing additional requirements to the development environment, which may not necessarily be maintained by the operator of the provider network 200.

[0017] To implement paging, the code proposals, once generated, may be stored in the service 210 and then returned via multiple exchanges by utilizing a paging token associated with the request for the code proposal to enable additional code proposals to be retrieved from storage and sent back to the development environment 219 or 211.

[0018] In various embodiments, code suggestion 213 may generate code suggestions based on text input in development environment 211 or 219, as will be discussed in detail below with respect to FIG. 3 (e.g., when code is input into development environment 211 or 219, a plugin or other connection that can provide real-time analysis and code suggestions may be utilized). Code suggestion 213 may use a generative model trained to generate code suggestions, a machine learning model such as a generative pre-trained transformer (GPT). The generative model is often trained on a large corpus of data for a particular task. When generating code recommendations, this corpus (e.g., from code suggestion code repository 215 or other code repositories used to train the generative model) may be composed of code repositories or snippets from various sources. Depending on the source or owner, the code may be subject to a particular license that requires attribution for any use or reproduction. Since the generative model can sometimes reproduce a word-for-word or near word-for-word match to the training data, metadata for attributing the original source may also need to be provided as part of the suggestion. Code suggestion metadata (not shown) may provide the ability to provide metadata for the code suggestions that may be provided.

[0019] Code development service 210 may implement (or have access to) code repository 215. Code repository 215 may store various code files, objects, or other code that may be interacted with by various other features of code development service 210 (e.g., development environment 211 for writing, building, compiling, and / or testing code). Code repository 215 may implement various versions and / or other access controls in some embodiments to track and / or maintain a consistent version of a collection of code for various development projects. In some embodiments, the code repository may be stored or implemented outside of provider network 200 (e.g., hosted in a private network or other location).

[0020] The code development service 210 may implement an interface for accessing and / or utilizing various features of the code development service 210. Such an interface may include various types of interfaces such as a command-line interface, a graphical user interface, and / or a program interface (e.g., an application programming interface (API)) to perform the required operations including the operation of the development environment 211. The API refers to an interface and / or communication protocol between a client and a server. When the client makes a request in a predefined format, the client should receive a response in a specific format or initiate a defined action. In the context of a cloud provider network, the API provides a gateway for a customer to access the cloud infrastructure by enabling the customer to obtain data from or cause an action within the cloud provider network, and enables the development of applications that interact with resources and services hosted within the cloud provider network. The API can also enable different services of the cloud provider network to exchange data with each other.

[0021] Generally speaking, client 250 can include any type of client that can be configured to submit network-based requests, including requests for services (such as requests for code search or suggestions), to provider network 200 via network 260. For example, a given client 250 can include a suitable version of a web browser, or can include an extension to the execution environment provided by a web browser, or a plugin module or other type of code module that can execute within the execution environment. Alternatively, client 250 can include an application (or its user interface), a media application, an office application, or any other application that can utilize resources within provider network 200 to implement various applications. In some embodiments, such applications may not necessarily implement full browser support for all types of network-based data, but can include sufficient protocol support (such as for a suitable version of the Hypertext Transfer Protocol (HTTP)) for generating and processing network-based service requests. That is, client 250 can be an application that can interact directly with provider network 200. In some embodiments, client 250 can generate network-based service requests according to a Representational State Transfer (REST)-style network-based service architecture, a document or message-based network-based service architecture, or another suitable network-based service architecture.

[0022] In some embodiments, client 250 may provide access to provider network 200 to other applications in a manner that is transparent to those applications. For example, client 250 may integrate with code development service 210. However, the operating system or file system may present different storage interfaces to the application, such as a conventional file system hierarchy of files, directories, and / or folders. In such embodiments, the application may not need to be modified to utilize the storage system service model. Instead, the details of the interface to the data storage service may be coordinated by client 250 and the operating system or file system, rather than by the application running within the operating system environment.

[0023] Client 250 can communicate network-based service requests to provider network 200 via network 260 and receive responses therefrom. In various embodiments, network 260 can include any suitable combination of networking hardware and protocols necessary to establish network-based communication between client 250 and provider network 200. For example, network 260 can generally include various telecommunications networks and service providers that collectively implement the Internet. Network 260 can also include private networks such as local area networks (LANs) or wide area networks (WANs), as well as public or private wireless networks. For example, both a given client 250 and provider network 200 can be provisioned within an enterprise that has its own internal network. In such an embodiment, network 260 can include the hardware (e.g., modems, routers, switches, load balancers, proxy servers, etc.) and software (e.g., protocol stacks, accounting software, firewall / security software, etc.) necessary to establish networking links between a given client 250 and the Internet, as well as between the Internet and provider network 200. Note that in some embodiments, client 250 can communicate with provider network 200 using a private network rather than the public Internet.

[0024] In some embodiments, the provider network 200 may include the hardware (e.g., modems, routers, switches, load balancers, proxy servers, etc.) and software (e.g., protocol stacks, accounting software, firewall / security software, etc.) necessary to establish networking links between different components of the provider network 200, such as virtual hosts, control plane components, and an external network 260 (e.g., the Internet). In some embodiments, the provider network 200 may employ Internet Protocol (IP) tunneling techniques to provide an overlay network through which encapsulated packets may pass through an internal network using tunnels. The IP tunneling technique may provide a mapping and encapsulation system for creating an overlay network and may provide separate name spaces for the overlay layer and the internal network layer. Packets in the overlay layer may be checked against a mapping directory to determine what their tunnel targets should be. The IP tunneling technique provides a virtual network topology, and the interface presented to the client 250 can communicate with a mapping service that knows where the IP overlay address is when the client 250 provides the IP address to which it wants to send a packet, so that the IP address can be attached to the overlay network to be executed within the virtual space.

[0025] The perceived latency of code suggestions can reduce the use of code suggestions as a feature. For example, if the user has to wait a detectable period after requesting a code suggestion, the user workflow may be interrupted. To eliminate the perceived latency, code suggestions can be prefetched. However, since the input context may have changed since the code suggestion was requested, techniques for verifying proactively fetched code completion suggestions that ensure that a given recommendation is no longer consistent with the current state of the code can be implemented. FIG. 3 is a logical block diagram showing code suggestion handling according to some embodiments.

[0026] Code suggestion handling 220 may implement automatic suggestion event detection 310 that can evaluate keystrokes 342, elapsed time, special keys, or user-specific data to detect an event. This information may be maintained as part of a user-specific state that can be updated or reset when a code suggestion request is submitted in various embodiments. For example, keystrokes, elapsed time, or other metrics may be reset. Special keys may also be trigger events (and may be evaluated in combination with other criteria such as elapsed time). For example, events that trigger obtaining a code completion suggestion may include the input of "{", "[", "(", ":", "ENTER KEY" or "TAB KEY" and an elapsed time threshold. In some embodiments, automatic suggestion event detection 310 may use client-specific events such as entering a specific key or character in a client-specific pattern (or configured / described by the client in a request that configures suggestion handling 220).

[0027] The code proposal request execution 320 can handle the formation, assembly, transmission, and processing of the response from the code proposal 213, including obtaining the code completion proposal and sending a request 322 to process the returned code proposal 324. For example, as shown at 344, the code proposal request execution 320 can obtain a context window of tokens (e.g., the N previous tokens before the cursor) from the file state 340. In some embodiments, file and other context information can be sent as provided by the file and other context extraction 350.

[0028] The file and other context extraction 350 can utilize different techniques to obtain file and other context information outside the context window (e.g., outside the N previous tokens). For example, the file context can be obtained from the same file for which the code proposal is being generated. Information that can be obtained for the file context can include boundaries within the current scope (e.g., code and comments restricted by the current function to provide a local context), class - level information including class declarations, class constructors (e.g., __init__ function), and function - level information for all other public or protected methods defined on the class, function - level information including all functions declared in the current file on both sides of the cursor. In some embodiments, signatures, docstrings, and return statements can be extracted, and / or variable - level information including all previous variable declarations visible in the current generation focus can be extracted.

[0029] Other contexts that can be extracted at 350 can be the in-project context. In modern code development, classes and functions are usually defined in hierarchical files. A simple backward context does not include information outside the current file, which can cause certain scenarios where the machine learning model is less likely to generate correct code. Since code files can use imported classes / functions / variables, adding this context can significantly improve code generation performance. Thus, in some embodiments, the in-project context can be added, in which case all imported classes, functions, and variables from the same project are used to obtain code completion suggestions.

[0030] Other contexts that can be extracted at 350 can be the out-of-project context. The out-of-project context can refer to classes / functions / variables imported from other packages into the current file. This can have an impact on the proposal quality when the imported package is under zero-shot settings (e.g., when the pre-trained model does not have prior knowledge about the package). Thus, other contexts can be obtained by scanning the out-of-project context for packages not included in the pre-training data and including the corresponding classes / functions / variables as context in the request.

[0031] File and other context extraction 350 can perform regular expression-based searches (e.g., for keywords such as "import") and extractions to obtain the various types of contexts discussed above. In some embodiments, parsing-based extraction can be used (e.g., by generating a symbol tree or other parse graph of the code to obtain other context information).

[0032] The code proposal request execution 320 can interact with code proposals provided in a paginated format. For example, the response 324 to a request for a code proposal can include a pagination token indicating that further proposals can be retrieved. The code proposal request execution 320 can further proceed to verify and provide the proposal via the code proposal verification 330, and also submit a subsequent request 322 with the pagination token to obtain further code proposal results, which can then be returned, verified, and provided. In this way, multiple proposals can be made, enabling different execution times for generating code proposals, including potentially better code proposals that can be provided while the user is reviewing the initial proposal.

[0033] The file state 340 can provide information at various stages and, in some embodiments, can include both the code file and its associated metadata. The file state 340 can also provide information for the code proposal verification 330, such as the current character before the cursor.

[0034] The code proposal verification 330 can verify the received code proposal before providing it for display. For example, the code proposal verification can use one or more verification criteria to determine whether the additional characters of the code proposal match or approximately match the characters before the cursor (and added after the time the code proposal request was made). As shown at 352, a valid code proposal can be provided for display. In some embodiments, as shown at 354, approval (or rejection) of these proposals is received and, as shown at 356, can be passed to or included in the file state 340.

[0035] The code proposal verification 330 can identify, display (348), and handle the approval or rejection of valid coded proposals.

[0036] Code proposal 213 can be implemented to provide various code proposals in different scenarios. FIG. 4 is a logical block diagram showing code proposals according to some embodiments. The code proposal request 401 can be received by the tokenizer 410. For example, various different tokenizers or tokenization techniques can be used. The token can be an individual word in the text that includes the preceding space character "[space]word" before the word as the token. Punctuation marks, whitespace, carriage returns, or various other characters can also be individually grouped or considered as tokens.

[0037] The programming language token prediction model 420 can use the provided code 422 and other context 424 such as a token window or the file context outside of other files obtained using techniques such as regular expressions or syntactic analysis as discussed above with respect to FIG. 3. A token prediction model 420 specified in a programming language can be used (e.g., model A for language A, model B for language B, etc.). Each of them can train the code in their respective programming languages and other contexts to generate recommendations. The programming language token prediction model 420 can also be trained with other context information (e.g., the same file, the same project, or other contexts) as discussed above with respect to FIG. 3.

[0038] In some embodiments, the code development service 210 can support a custom programming language model. For example, training data or code data from a user's specific code repository can be provided to train the custom programming language model and, as a result, be used for code proposals.

[0039] The prediction can be provided to the selection 430 and one can be selected based on a confidence score to be provided as a code proposal. In some embodiments, multiple predictions can be provided in a paginated format or other multiple result formats as discussed above with respect to FIG. 3.

[0040] One scenario in which a machine learning model that generates text recommendations can occur is when the input has partial words such as "Syst". In these scenarios, the machine learning model tends to provide insufficient predictions and thus tends to provide insufficient suggestions (e.g., generating gibberish or incoherent generations). This occurs because the model only sees word tokens as input units. To overcome this scenario, backtrack to the last complete token and constrain the generation to match the prompt suffix, which in this case is "Syst". As discussed below, constraining the generation helps improve the accuracy for subword data metrics without sacrificing gains for a general evaluation set.

[0041] When a string prompt is given, non - consistencies from normal decoding can be caused by the suffix of that prompt which can potentially occur with sub - words that are not complete tokens. Matching of the input string suffix can be performed starting with that suffix or using all available tokens that start with the suffix. In some embodiments, the matching is performed efficiently using a character trie data structure (e.g., using native Pytorch arrays as a list of nodes for fast concatenation). Based on the list of matching tokens, other tokens can be masked during the next token prediction, thus ensuring that the generation matches the suffix character - by - character. Further latency optimizations can be achieved by caching very frequent suffixes such as a single space by holding a boolean mask. For each step that matches the suffix, after each token generation step, the matching tokens are removed (character - by - character) from the left of the suffix and constrained generation is performed until the suffix becomes an empty string. In some embodiments, sub - tokens (e.g., suffixes) are determined by using the same pre - tokenization as the pre - tokenization strategy of the tokenizer 410 that performs splitting using word boundaries, which allows for efficient backtracking for character matching since it can be deterministically known that no sub - token can cross a pre - token boundary.

[0042] Constrained prediction generation 440 can implement these techniques to provide code proposals, as will be discussed in more detail below with respect to FIGS. 11 and 12. For example, constrained prediction generation 440 can be a feature that is enabled (or disabled) for a code proposal. In some embodiments, constrained prediction 440 can be configured (e.g., via a proposal request) to utilize a specified maximum number of backtrack tokens (which can be determined dynamically up to a maximum for a code proposal). Constrained prediction generation 440 can receive tokenized input data from the tokenizer 410 and then determine backtrack tokens and sub - tokens.

[0043] The constrained prediction generation 440 can then, in various embodiments, identify one or more possible tokens that match a partial token that can be identified from the possible tokens. For example, the possible tokens can be the vocabulary of a language (e.g., programming or human) that can have different words that can be predicted. A match can be identified when the partial token matches either at the beginning or the end of the possible token (e.g., the possible token matches when the partial token [SYS] matches the beginning [SYS* of the possible token, or the end *SYS] of the possible token). The matching can be performed using a trie structure as considered above and below. A cache for common matches can also be utilized.

[0044] The constrained prediction generation 440 then performs one or more iterations of the next token prediction, filters against the identified possible matches, evaluates the remaining next token predictions using the partial token, and can subtract matching characters from the partial token until all characters from the partial token match. The code proposal can then be provided (e.g., to a prediction selection 430 that can simply transfer or transmit the next token prediction as the code proposal 401).

[0045] FIG. 5 is a logical block diagram showing an exemplary interface of a development environment according to some embodiments. The integrated development environment interface 500 can be implemented on a client of the code development service 210 as depicted in FIG. 2, or can be hosted as part of the code development service 210 as depicted in FIG. 2. The integrated development environment interface 500 can implement a code editor 510 (e.g., a text editor) that can enable a user to enter code in a programming language. The code proposal feature 213 of the code development service 210 can analyze the entered characters to determine a code proposal 520, which can be displayed and added as shown at 522. Although not shown, various other information about the proposal, such as the source of the code, the licensing information of the code, and / or various other code metadata (e.g., style guidelines), can be displayed.

[0046] Code suggestion and other next-word or token prediction techniques may encounter scenarios where partial tokens are the closest contextual input for code suggestion (or other next-word / token prediction). In the following example of Java code, code suggestion is performed after the cursor, and the cursor is at " <t>Consider the following exemplary scenario, represented as "」". GetRecordsResult result = streamClient.getRecords(streamName); while(result.getNextRecordMarker()!=null){ / / (2) System.out.println(result.getNextRecordMarker()); result = streamClient.getRecords(streamName); S <t><- current cursor The most likely prediction after the cursor, and the most likely here, should be the very common module System, which is referenced in Java and also appears in the above context. However, since "S" is a partial token, the current cursor code proposal provided immediately before may not handle partial tokens well. Consider the following exemplary input with proposed code underlined after the cursor. / / Input -> S S <t> leeper sleeper = new Sleeper(1000); / / Input -> Sy Sy <t> strace.endSection(); / / Input -> Sys Sys <t> out.println(result.getNextRecordMarker()); / / Input -> Syst Syst <t> en.sleep(1000); / / Input -> Syste Syste <t> .sleep(1000); It should be noted that in the original text, "en.sleep(1000);" and ".sleep(1000);" seem to have some incorrect spellings. It might be "Thread.sleep(1000);" in a proper Java context.

[0047] In the above example of creating longer partial tokens, possible proposals may still not provide the expected "System" result. This can occur because the machine learning model used to generate the proposals was not trained on partial tokens such as "S" or "Sy".

[0048] This lack of training can occur because of how the tokenizer breaks down the training dataset for the machine learning model used to generate proposals. Consider the following exemplary tokens generated from the input text (left side of "->"). System -> [' System'] Sleeper -> [' S', 'le', 'eper'] / / Prefix S<> Systrace -> [' Sy','str', 'ace'] / / Prefix Sy<> Sysout -> [' Sys', 'out'] / / Prefix Sys<> Systen -> [' S', 'yst', 'en'] / / Prefix Syst Syste -> [' S', 'yst', 'e'] / / Prefix Syste

[0049] To address this issue, a randomly segmented version of some tokens can be applied to the training dataset for training or fine - tuning the machine learning model for code (or other text) proposals. For example, one example of code text could be as follows. System.out.prinln("test"); Tokenizing this as is results in the following. [' System', '.', 'out', '.', 'pr', 'in', 'ln', '(\"', 'test', '\");'] Instead, a random split can be inserted into the sentence, splitting the sentence into two parts. ‘System.out.prinln("test”);’ -> [’S’, ’ystem.out.prinln("test”);’] Each segment is then tokenized individually and then concatenated, which provides the following tokenization result. [‘ S’, ‘ y’, ‘st’, ‘em’, ‘.’, ‘out’, ‘.’, ‘pr’, ‘in’, ‘ln’, ‘(”’, ‘test’, ‘");’] This can be used as part of a training dataset for pre-training or continuous fine-tuning.

[0050] Implementing a randomized segmentation for tokens can, in various embodiments, train a machine learning model on how to compose various configurations of sub-tokens (e.g., from “Sy” to “System”) with respect to the original tokens. Such techniques used to train a machine learning model can significantly improve the accuracy of partial token completion without hurting the full token prediction scenario. Further, these techniques can be implemented without reducing the speed of the machine learning model at the inference stage.

[0051] FIG. 6 is a logical block diagram showing an example of code proposal development for sub-word regularization when training a machine learning model according to some embodiments. The random sub-word tokenizer 610 and model training may implement the tools, systems, or features of code proposal development 217. In other embodiments, a separate training system, application, or service, such as a machine learning service implemented as part of the provider network 200 of FIG. 2, may implement these techniques, as well as the techniques discussed below with respect to FIGS. 13 and 14.

[0052] The random subword tokenizer 610 can obtain training data 602 and randomly segment it, as will be discussed below with respect to FIGS. 13 and 14 for model training 620. For example, a request to generate partial token optimization training data may be received that specifies a storage location or other information describing a source training data set that includes text data. According to some embodiments, the random subword tokenizer can determine multiple tokens from the text data. For example, various different tokenizers can be applied to generate tokens from the input text. For example, a token can be an individual word in a sentence that includes the preceding space character "[space]word" before the word as a token. Punctuation marks, whitespace, carriage returns, or various other characters can also be individually grouped or considered as tokens.

[0053] In some embodiments, the random subword tokenizer 610 can randomly select different tokens among the multiple tokens and segment them into respective subtokens. For example, subword regularization techniques can be performed to sample or identify different tokens for non-optimal segmentation. Such techniques can include randomly selecting tokens from the tokens determined for the text data (e.g., according to configurable variables or parameters that can be represented as percentage values), where the percentage value indicates the likelihood that any one token will be selected for random segmentation. When a token is selected, the token can be randomly segmented into subtoken components that are treated as tokens instead of the source token from which the token was generated.

[0054] Model training 620 can implement various machine learning training frameworks that can execute machine learning jobs, applications, or programs on the initial model 604 using the training dataset generated by the random subword tokenizer 610. Note that the training can be executed to train the model from scratch or can be trained given an unregularized checkpoint. For example, the initial model can already be pre-trained and thus can include various neural network-based machine learning models provided for fine-tuning, or can be a new model that has not been pre-trained. Various different hyperparameters or other configurations of the model training can be specified as part of the training job or request and can be used to execute training against the initial model. Upon completion, model training 620 can provide a trained model with subword regularization, as shown at 606.

[0055] Another tool, system, or feature of Code Proposal Development 217 can be Programming Language Conversion 710, which can convert a source evaluation dataset of one programming language to another programming language. High-quality evaluation datasets are time-consuming to create and typically require the time and effort of a large number of annotators. This is also true for execution-based function completion evaluation sets. In various embodiments, a programmatic test conversion tool from a source programming language such as Python to another target programming language is applicable to tests that perform accuracy evaluation based on the return value of a function with a ground truth value (thus, value-oriented conversion). These embodiments can be used to convert many test cases, which helps reduce annotation time and increase the number of evaluation datasets for building and testing additional code generation systems for many different languages. These techniques are widely applicable and can be used to convert existing datasets such as MBPP (Most Basic Python Programming) to Javascript, Java, Typescript, Ruby, Go, C#, or any other programming language for which a conversion rule set is generated.

[0056] In some embodiments, the conversion process begins by inferring the types of the function arguments, which can be done by examining the argument values within each test case. Type mappings from Python to each language, such as from "list" to "ArrayList" in Java or from "dictionary" to "HashMap". The values of different test cases can have different types, and thus, the common superclass of all observed types of each argument can be inferred according to the type hierarchy. Since there can also be many levels of types (due to containers such as lists or sets), types can be recursively inferred to be consistent at each level. For example, "list of list" and "list of object" have a common type of "list of object". The type of the expected return value can also be inferred by examining the expected return value within the test case that matches the value of the function executed with the given input of that test case.

[0057] In addition to types, the conversion of arguments and return values from a source programming language to a target programming language by generating strings representing objects of the target language that can be parsed by each interpreter / compiler. For example, Python's [1,2] is converted to 'Arrays.aslist(1, 2)', or Python's {1 : 2, 3: ["foo", "bar”]} is converted to 'new HashMap(){{put(1, 2);put(3, Arrays.aslist("foo", "bar”))' with recursive support for any nested structure.

[0058] For test case conversion, in some embodiments, all the information regarding the return type and arguments / expected return values can be combined to construct code representing the input / output objects in the appropriate target programming language using an appropriate comparator for equivalence.

[0059] In addition, in some embodiments, the conversion of source programming language prompt strings, including function signatures and docstrings that include input / output examples, can be converted to prompt strings in other target programming languages. The style of the function signature can be emulated in each language, along with the appropriate return / argument types where applicable, and the function / argument / class names can be converted to a style-appropriate form (camel case or Pascal case). The docstring can be formatted so that the input and output appear as close as possible to the input / output format for that particular language.

[0060] FIG. 7 is a logical block diagram showing an example of code proposal development for evaluation dataset conversion according to some embodiments. As discussed above, and as will be discussed in detail below with respect to FIGS. 15-16B, the programming language conversion 710 can convert a given evaluation dataset from a source programming language to a target programming language. For example, the programming language conversion 710 can receive a conversion request 702. The conversion request 702 can specify the source programming language and target programming language of the evaluation dataset, as well as the storage location or other access information of the source evaluation dataset, such as the source evaluation dataset of programming language A772, and the target storage location, format, and / or other access information for generating a new evaluation dataset, such as the storage location of the new evaluation dataset of programming language B774.

[0061] The programming language conversion 710 can utilize different conversion techniques for different parts of the items in the evaluation dataset, such as techniques for inferring or mapping (e.g., recursively) the types of function signatures in the test statement conversion 730 and the natural language conversion 750. Each of these features can utilize specific conversion rules, mappings, and / or type hierarchies for the specified source programming language and target programming language.

[0062] The function signature conversion 720 can identify the function signature in the source 772 by parsing the items in the evaluation dataset according to a parser or rule set for the first (source) programming language and locating the function signature. In the Python programming language, for example, to locate the function signature, a search for "def" (e.g., a regular expression search) can be performed, and the function signature can be delimited by various other symbols (e.g., it can include arguments within parentheses). Once the function signature is located, different techniques can be executed to determine what the type of each argument or parameter of the function is. For example, the test cases of the function can identify the values of the arguments. To complete the conversion, one or the mapping rules specific to the conversion of the function signature of the source programming language to the target programming language can be applied.

[0063] The test statement conversion 730 can convert the test statements using the knowledge determined as part of converting the function signature 720. For example, using the argument format of the function signature in the source, various test values can be extracted from the source test statements and inserted into the target programming language version of the test, which can be obtained as a template test statement that approves the arguments and triggers an error or other indication if the test statement fails.

[0064] Natural language conversion 750 can be implemented as part of converting a prompt from source 772 into a target programming language evaluation dataset 774. For example, the conversion of the prompt can include changing features such as symbols used to denote code comments (e.g., non-executable statements within the code), for example, changing from “””” to / * *. The conversion of the prompt can also include changing natural language statements to replace source programming language terms with target programming language terms, such as from “Write a function in Python” to “Write a function in Java”, or changing between terms such as from “none” to “null”. In some scenarios, the conversion rules can remove unnecessary or uncovered table source programming language specific statements.

[0065] The function body of the target programming language can be generated by sending a request from 740 to code suggestion 213, and code suggestion 213 can receive the request 704 and return the generated code 706. The request can include, in some embodiments, the converted prompt and the converted function signature of the test item.

[0066] Examples of validating and proactively providing the code suggestions discussed above with respect to FIGS. 2-7 are given with respect to an example of a code development service. Various other types of code development tools, systems, or applications can implement these techniques. FIG. 8 is a high-level flowchart showing techniques and methods for implementing the validation and proactive provision of code suggestions according to some embodiments. These techniques, as well as the techniques discussed below with respect to FIGS. 9-16B, can be implemented using various components of a provider network as discussed above with respect to FIGS. 2-7, or other types or systems implementing code development tools or other applications.

[0067] As shown at 810, in some embodiments, an event can be detected that triggers obtaining code completion suggestions for inclusion in a code file being edited using an integrated development environment. In some embodiments, the event that triggers obtaining code completion suggestions can be based on one or more criteria. For example, a keystroke count (since the last code completion suggestion request was made) can be maintained. Only this keystroke count can be the event that triggers when the number of keystrokes exceeds a threshold. In some embodiments, other criteria can be considered. For example, the time elapsed since the last trigger can also be used, which can obtain code suggestions after a period of time has elapsed since the last recommendation was made. In some embodiments, a combination of criteria (e.g., keystrokes and elapsed time) can be used. In some embodiments, the event trigger can be user- or client-specific based on heuristics such as entering or using a particular character or key (e.g., after selecting the TAB key for indenting).

[0068] As shown at 820, in some embodiments, generation of code completion suggestions can be caused. The code completion suggestions, in some embodiments, can be based on the character immediately preceding the cursor at a first time when an event that triggers a request for the code completion suggestions is detected, and the code completion suggestions include the proposed characters to be entered into the code file immediately after the cursor at the first time. In some embodiments, the code suggestions can be implemented and executed locally (e.g., by a local subsystem). In some embodiments, the code suggestions can be generated remotely (e.g., as a feature of the code development service 210 of FIG. 2).

[0069] As shown at 830, a determination can be made as to whether a comparison of the number of proposed characters and the number of actual characters entered into the code file after the first time meets one or more verification criteria. For example, the verification criteria can be an exact match, as considered below with respect to FIG. 9. In other embodiments, the verification criteria can allow for a fuzzy match or a near match (e.g., a 3-out-of-4 character matching). If no match is found, as shown at 850, the code proposal can be discarded. Otherwise, as shown at 840, the code proposal can be displayed.

[0070] In some embodiments, the code completion proposal can be approved or rejected by the user and can itself trigger a recommendation for a further code completion proposal.

[0071] FIG. 9 is an exemplary timeline for detecting, verifying, and displaying a code completion proposal according to some embodiments. At time T1, the input code 910 within the editor is shown along with the location of the cursor. An event is detected that triggers obtaining a code proposal. At time T2, the input code 920 has changed with additional characters added as indicated by the moved cursor. The code proposal 930 can be verified at T2 by comparing a number (e.g., 4) of characters with the additional characters "temp".

[0072] Code proposals for programming languages can be provided in various embodiments by the code proposal features of services such as the code development service 210 or as a stand-alone code generation system. Many of the techniques considered above and below for improving the performance of various stages of code proposals can be integrated with code proposals generated using techniques such as the technique of FIG. 10. FIG. 10 is a high-level flowchart showing techniques and methods for implementing generating a code proposal for an input programming code according to some embodiments.

[0073] As shown at 1010, in some embodiments, a request to generate a code proposal for input programming code may be received. For example, the code proposal request may be generated, in some embodiments, as part of the eager or anticipated code proposal request techniques discussed below with respect to FIGS. 3, 8, and 9, or as a manual request for a recommendation. In some embodiments, the request may be used to perform various techniques, such as generating a code proposal to provide a transformed function body for evaluation dataset conversion, as discussed below with respect to FIGS. 15 - 16B.

[0074] As shown at 1020, in some embodiments, tokens may be determined from the input programming code. Various tokenization techniques discussed above and below may be performed. For example, various different tokenizers or tokenization techniques may be used. A token may be an individual word in the text that includes the preceding space character "[space]word" before the word as the token. Punctuation, whitespace, carriage returns, or various other characters may also be grouped or considered individually as tokens.

[0075] As shown at 1030, in some embodiments, a machine learning model trained to generate a next token prediction for a programming language corresponding to the programming language of the input programming code is applied to the tokens of the input programming code to generate a next token prediction for the input programming code. This machine learning model may be trained using randomized token segmentation, as discussed above with respect to FIG. 6 and below with respect to FIGS. 13 and 14. Various different types of machine learning models may be used, such as a Generative Pretrained Transformer (GPT), a sequence-to-sequence model, or other neural network-based models such as Long Short-Term Memory (LSTM).

[0076] As shown in 1040, in some embodiments, one of the next token predictions can be selected according to each confidence score for returning as a code proposal for the input programming code. In some embodiments, multiple recommendations can be generated, and some recommendations including multiple recommendations (e.g., top three according to the trust store) can be provided.

[0077] As discussed above with respect to FIGS. 1, 5, and 6, the partial token scenario can cause the machine learning model to generate incorrect text proposals. As discussed above with respect to FIG. 6 and below with respect to FIGS. 13 and 14, several techniques can be applied to improve the training side problems, but inference time techniques can also be implemented. In various embodiments, constrained prefix matching for generating the next token prediction can be used in such situations to correct partial token proposal errors. FIG. 11 is a high-level flowchart showing techniques and methods for implementing constrained prefix matching for generating the next token prediction according to some embodiments.

[0078] As shown in 1110, in some embodiments, an input text can be received for performing the next token prediction on the input text. For example, the input text can be received as part of a request for a code proposal, as discussed above with respect to FIGS. 4, 8, and 10, and can be the amount of code written in the programming language before the cursor when the code proposal request is triggered. The input text can also be received for other text prediction or completion scenarios (e.g., text proposals for drafting various documents, text auto-completion for various use cases such as providing input for other requests or forms in natural language).

[0079] As shown at 1120, word boundaries for a tokenizer of input text can be determined. The rightmost boundary potentially includes a partial token. This partial token can be used as a prompt suffix for constraining the generation of the next token. Note: These word boundaries are units larger than tokens and are often referred to as pretokens. The last token can be a partial token. For example, various different tokenizers or tokenization techniques can be used. A token can be an individual word in a text that includes the preceding space character "[space]word" before the word as the token. Punctuation, whitespace, carriage returns, or various other characters can also be grouped or considered individually as tokens.

[0080] In some embodiments, "pretokens" that occur immediately before a partial token can be identified. From these pretokens, backtrack tokens can be determined. For example, starting from the pretoken immediately before the prompt suffix, one or more of the pretokens can be added to the backtrack token as it operates in reverse order of the tokens until the maximum number of backtrack tokens is reached or a special character (e.g., carriage return) is reached. In some embodiments, the maximum number of backtrack tokens can be a configurable parameter for the prediction of the next token (e.g., as part of the requirements for the next token prediction or as a separate configuration requirement).

[0081] As shown at 1130, in various embodiments, one or more possible tokens that match the prompt suffix can be identified from the possible tokens. For example, the possible tokens can be the vocabulary of a language (e.g., programming or human) that can have different words that can be predicted. A match can be identified when the prompt suffix matches either at the beginning or the end of a possible token (e.g., the possible token matches when the prompt suffix [SYS] matches the beginning [SYS* of the possible token or the end *SYS] of the possible token).

[0082] In some embodiments, one or more different data structures may be used to identify matching possible tokens. For example, a trie data structure may be used to store different possible tokens. A trie may be a search tree for prefixes (or suffixes), the trie may be string-indexed against a vocabulary of words, and individual nodes include links to suffix child nodes that add additional characters to the suffix at each child node. Another example of a data structure that may be used to efficiently identify matching tokens may be a cache of possible matching tokens (e.g., as a boolean mask).

[0083] As shown in 1140, in some embodiments, the next token prediction may be filtered according to one or more identified possible tokens, and the next token prediction is generated by applying a machine learning model to the remaining portion of the input text that does not include the number of backtrack tokens corresponding to the pre-token. For example, a given input to the machine learning model may have some input tokens (including sub-tokens) from the text before the cursor, such as 15 tokens if token 15 is a sub-token. If the number of backtrack tokens is 3, tokens 14, 13, and 12 (adjacent to token 15) may not be used as input for the next token prediction machine learning model, and as a result, the input may instead be tokens 11 to 1 leading backward (and may include three more leading tokens to supplement the backtrack tokens) and sub-token 15.

[0084] The result of the next token prediction given the input tokens may include several different token predictions with varying confidence values. Token predictions that are not one of the identified possible tokens may be excluded from consideration. The highest confidence score remaining among the predictions may be identified.

[0085] For the next token prediction of this remainder, as shown at 1150, the number of characters matching the partial token can be subtracted from the left of the partial token. As shown at 1160, if no further characters remain, as shown at 1170, the next token prediction can be provided as the next token prediction. If none remain, as indicated by the positive exit from 1160, another iteration of the next token prediction can be performed using the remaining characters until no characters remain.

[0086] FIG. 12 is a logical block diagram showing different iterations of constrained prefix matching according to some embodiments, as discussed above with respect to FIG. 11. Iteration 1201 shows an input that includes the remainder but does not include backtrack token 1202 and partial token (e.g., "[SPACE]SYSTE"). The most highly trusted filtered prediction can be "[SPACE]SY". This can be used to subtract matching characters from the partial token such that "STE" remains, as shown in iteration 1203. For this iteration, the prediction can be "[SPACE]SYST". After subtraction, the partial token is "E". For iteration 1205, the prediction can be "[SPACE]SYSTEM", which matches the remaining character "E", ending the iteration and using the prediction "[SPACE]SYSTEM" as the next token prediction.

[0087] As discussed above with respect to FIGS. 1 and 6, training data for code suggestion and other text generation systems can encounter difficulties when partial tokens are included as part of the input for text generation inference, which can occur in various autocomplete scenarios where text (e.g., code) is generated to provide the input context of previously input text up to some partial tokens (e.g., partial or incomplete words) to automatically complete the next part of the text. To improve the ability of a machine learning model to consider these scenarios, a training data set may need to include data items that train these partial token scenarios. FIG. 13 is a high-level flowchart showing techniques and methods for implementing random token segmentation for training a next token prediction model according to some embodiments.

[0088] As shown at 1310, in some embodiments, text data may be received for training a machine learning model to predict a next text token given an input text token. For example, a request to generate partial token optimized training data may be received that specifies a storage location or other information describing a source training data set that includes the text data. In some embodiments, this request may be received or specified as part of a training job submitted to a machine learning system or service, such as a machine learning service implemented as part of a provider network like provider network 200 of FIG. 2.

[0089] As shown at 1320, according to some embodiments, a plurality of tokens may be determined from the text data. For example, various different tokenizers may be applied to generate tokens from the input text. For example, a token may be an individual word in the text that includes the preceding space character "[space]word" before the word as the token. Punctuation, whitespace, carriage returns, or various other characters may also be individually grouped or considered as tokens.

[0090] As shown at 1330, in some embodiments, different tokens among a plurality of tokens may be randomly segmented into respective subtokens. For example, subword regularization techniques may be performed to sample or identify different tokens for non-optimal segmentation. FIG. 14, discussed below, provides an example of random token segmentation. Such techniques may include randomly selecting tokens from tokens determined for text data (e.g., according to configurable variables or parameters that may be shown as percentage values), where the percentage value indicates the likelihood that any one token will be selected for random segmentation. If a token is selected, the token may be randomly segmented into subtoken components that are treated as tokens instead of the source token from which the token was generated.

[0091] As shown at 1340, in some embodiments, a machine learning model may be trained to predict a next token given an input text token using a plurality of tokens each including a respective subtoken as a training data set. In some embodiments, the trained machine learning model may be stored at a location specified in a training request. In some embodiments, the machine learning model may be deployed for different applications, including an auto-complete application for code suggestions or text as discussed above. Various different training techniques and machine learning model types for next token prediction may be used, such as sequence-to-sequence models or other neural network-based models like long short-term memory (LTSM).

[0092] FIG. 14 is a logical block diagram showing possible random tokenization of text according to some embodiments. As shown at 1410, text data such as "New York" can be an example of a longer text string that can be tokenized for a training data set, as discussed above with respect to FIG. 13. Each word "New" and "York" can be a separate token. Since tokenization can be performed randomly, for example by using subword regularization techniques, different examples can be generated as shown at 1420, 1430, and 1440 (each block being a subtoken). In these examples, a random selection is made as to whether a token is selected for further segmentation (e.g., both at 1420, "York" at 1430, and "New" at 1440). Random segmentation can result in two or more segments and, in some embodiments, can include treating individual characters as subtokens.

[0093] As discussed above with respect to FIGS. 1 and 7, it can be difficult to obtain high-quality data sets for training and evaluating different systems, services, or applications, such as those implemented to provide code or other text suggestions as discussed above. Some data sets, such as evaluation data sets, can be highly specialized. For example, an evaluation data set for a code suggestion system can depend on various code prompts or problems to be solved. A code prompt can describe a programming problem that can be taken in by a code suggestion system and then have a corresponding portion of the code generated to solve or satisfy that problem. For example, the prompt may desirably be a function that checks whether a given number is odd or even. Various exemplary assertions or unit tests can be included that can be used to determine whether the generated code returns the correct value to pass the test. To increase the number of available high-quality evaluation data sets, techniques for generating new high-quality evaluation data sets can be highly desirable.

[0094] FIG. 15 is a high-level flowchart showing techniques and methods for programmatically generating an evaluation data set for a code generation system according to some embodiments. This technology can be implemented by various types of systems for testing, developing, or implementing code proposals or other code generation systems. In some embodiments, a programming language conversion system may include a suite of tools, including tools for programmatically generating an evaluation data set for a code generation system.

[0095] As shown at 1510, in some embodiments, an evaluation data set specified in a first programming language may be received, and different items of the evaluation data set may correspond to different respective evaluation tests of the code generation system. For example, one or more files, objects, locations, or other information for accessing and obtaining the evaluation data set may be provided as part of a request to perform a conversion of the evaluation data set from a first (e.g., source) programming language to a second (e.g., target) programming language. In some embodiments, multiple target programming languages may be specified as part of the request, and thus, multiple executions of the technology may be specified, as discussed below.

[0096] As shown at 1520, in some embodiments, individual items of a dataset can be converted to a second programming language. For example, the conversion of a prompt can include changing features such as symbols used to indicate code comments (e.g., non-executable statements within the code), e.g., changing from “””” to / * *. The conversion of a prompt can also include changing natural language statements to replace source programming language terms with target programming language terms, e.g., from “Write a function in Python” to “Write a function in Java”, or changing between terms such as from “none” to “null”. In some scenarios, the conversion rules can remove unnecessary or uncovered table source programming language-specific statements.

[0097] As shown at 1530, the function signature of an item in a first programming language can be converted to a second programming language. For example, the function signature can be identified by parsing the items of the evaluation dataset and locating the function signature according to a parser or set of rules for the first (source) programming language. In the Python programming language, for example, a search for “def” (e.g., a regular expression search) can be performed to locate the function signature, and the function signature can be delimited by various other symbols (e.g., can include arguments within parentheses).

[0098] Once the function signature is located, different techniques can be performed to determine what the type of each argument or parameter of the function is. For example, the test cases of the function can identify the values of the arguments. In FIG. 16A, for example, the function signature 1612 can provide the data set source item 1610 being transformed. The function signature 1612 may be located (as discussed above) and then evaluated to determine the values of the arguments "cost, m, n". One of the test cases indicates that the candidate function has an input such as "[[1,2,3],[4,8,2],[1,5,3]],2,2", which can be decomposed into the cost as "[[1,2,3],[4,8,2],[1,5,3]]", m as "2", and n as "2". Thus, cost can be a list and m and n can be integers. The return value can also be inferred from the test, for example, "[[1,2,3],[4,8,2],[1,5,3]],2,2 == 8", where "8" is the desired return value for that test case. In some embodiments, determining the return value can include executing both the source test item function using a test to identify the returned value.

[0099] To complete the conversion, one or mapping rules specific to the conversion of the function signature of the source programming language to the target programming language may be applied. In the illustrated examples of FIGS. 16A and 16B, at 1624, mapping rules for converting the text of the Python function signature 1612 to Java are applied, and "def" is replaced with "class MinCost{public static in MinCost(List<List <integer>>cost, int m, int n{」 is converted as. As discussed above, the determination of the arguments can enable the mapping of the determined source value type to the target value types explicitly declared in Java example 1624 (e.g., "List", "integer", and "integer"). In some scenarios, the arguments are a list of heterogeneous data (which can be a list, set, etc. as elements). To perform the conversion, the inference technique can examine the data type by recursively looking at the elements. In such a technique, the most general type given all the values can be selected. For example, if one argument value is List <integer>is of type, and the other argument value is List <double>If it is of type, the type is List <double>should be. In another example, one argument type is List<List <integer>> and when the other argument type is List<HashMap<Integer, String>>, the type is List <object>It becomes

[0100] As shown in 1540, in some embodiments, test statements of items in a first programming language can be converted into a second programming language. Some knowledge determined as part of converting the function signature can be used to convert the test statements. For example, using the argument format "(cost, m, n)", convert from "assert candidate" to "class Main { public static void main(String[] args) throws Exception { if (!(MinCost.MinCost(Arrays.asList(Arrays.asList(1, 2, 3), Arrays.asList(4,8,2),Arrays.asList(1,5,3)),2,2)==8) throw new java.lang.Exception(”Exception -- test case 0 did not pass”);}", extract various test values from 1618, and as shown in test statement 1628, they can be inserted into the target programming language version of the test. This can be repeated for each test.

[0101] As shown in 1550, in some embodiments, the body of the transformed function signature can be generated in a second programming language according to a prompt within an item used as input to a machine learning model trained to generate code in the second programming language. For example, as will be discussed in detail below with respect to FIGS. 1, 2, 4, and 10, code generation techniques can take as input a given portion of the code to be generated or a natural language statement describing the code and generate the proposed code. Thus, a code conversion system (or other system, service, or application that executes the techniques enumerated in FIG. 15) can call, via an interface, a code generation system or a locally implemented machine learning model to generate function body code in a target programming language (e.g., by specifying the target programming language to cause the use of a machine learning model specific to that target programming language).

[0102] In various embodiments, the assembly of different transformed item parts can be completed according to one or more conversion rules to the target programming language of the items in the evaluation dataset in source programming. For example, the ordering of parts can vary from one programming language. In FIG. 16A, for example, a test item is the ordered function signature 1612, then the prompt 1614, then the function body 1616, and finally the test statement 1618, and as in the transformed version 1620, the ordering is the prompt 1622, then the function signature 1624, then the function body 1626, and finally the test statement 1628. Different programming languages can have different required orderings that can be considered when assembling the transformed versions.

[0103] As shown in 1560, in some embodiments, as part of a new evaluation dataset, the transformed individual items of different evaluation datasets of the evaluation dataset can be stored. For example, each item in the evaluation dataset can be a different file, document, or other object. When each new transformed item is created, the corresponding different file, document, or object can be added to a target storage location for the new evaluation dataset. In some embodiments, various errors can trigger storing the source item in a separate storage location for notification and / or manual conversion (e.g., sending a notification that the source item should be reviewed).

[0104] The methods described herein can be implemented, in various embodiments, by any combination of hardware and software. For example, in one embodiment, the method can be implemented by a computer system (e.g., a computer system such as that shown in FIG. 17) including one or more processors that execute program instructions stored on a computer-readable storage medium coupled to the processor. The program instructions can be configured to implement the functions described herein (e.g., the functions of the various servers and other components that implement the provider network described herein). The various methods shown in and described herein represent exemplary embodiments of the methods. The order of any method can be changed and various elements can be added, rearranged, combined, omitted, modified, etc.

[0105] The techniques discussed above may be implemented on one or more computer systems that can interact with various other devices. FIG. 17 is a block diagram showing an exemplary computer system according to various embodiments. For example, computer system 2000 may be configured to implement a compute cluster, a distributed key-value data store, and / or a client node in different embodiments. Computer system 2000 may be any of a variety of types of devices including, but not limited to, a personal computer system, a desktop computer, a laptop or notebook computer, a mainframe computer system, a handheld computer, a workstation, a network computer, a consumer device, an application server, a storage device, a telephone, a cellular phone, or generally any type of computing device.

[0106] The computer system 2000 includes one or more processors 2010 (any of which may include multiple cores that may be single or multi-threaded) coupled to a system memory 2020 via an input / output (I / O) interface 2030. The computer system 2000 further includes a network interface 2040 coupled to the I / O interface 2030. In various embodiments, the computer system 2000 can be a uniprocessor system including one processor 2010, or a multiprocessor system including multiple processors 2010 (e.g., two, four, eight, or another suitable number). The processor 2010 can be any suitable processor capable of executing instructions. For example, in various embodiments, the processor 2010 can be a general-purpose or embedded processor implementing any of various instruction set architectures (ISA), such as x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multiprocessor system, each of the processors 2010 can, but need not, commonly implement the same ISA. The computer system 2000 also includes one or more network communication devices (e.g., network interface 2040) for communicating with other systems and / or components via a communication network (e.g., the Internet, a LAN, etc.). For example, a client application running on the system 2000 can use the network interface 2040 to communicate with a server application running on a single server or a cluster of servers implementing one or more of the components of the provider network described herein. In another example, an instance of a server application running on the computer system 2000 can use the network interface 2040 to communicate with another instance of a server application (or another server application) that can be implemented on another computer system (e.g., computer system 2090).

[0107] In the illustrated embodiment, computer system 2000 also includes one or more persistent storage devices 2060 and / or one or more I / O devices 2080. In various embodiments, persistent storage device 2060 may correspond to a disk drive, tape drive, solid state memory, other mass storage device, or any other persistent storage device. Computer system 2000 (or a distributed application or operating system running thereon) may store instructions and / or data on persistent storage device 2060 as needed, and may retrieve the stored instructions and / or data as needed. For example, in some embodiments, computer system 2000 may host a storage system server node, and persistent storage 2060 may include an SSD attached to that server node.

[0108] The computer system 2000 includes one or more system memories 2020 configured to store instructions and data accessible by the processor 2010. In various embodiments, the system memory 2020 may be implemented using any suitable memory technology (e.g., cache, static random access memory (SRAM), DRAM, RDRAM, EDO RAM, DDR 20 RAM, synchronous dynamic RAM (SDRAM), Rambus RAM, EEPROM, non-volatile / flash type memory, or one or more of any other type of memory). The system memory 2020 may include program instructions 2025 executable by the processor 2010 to implement the methods and techniques described herein. In various embodiments, the program instructions 2025 may be encoded in platform native binary, any interpreted type language such as Java (trademark) bytecode, or any other language such as C / C++, Java (trademark), or any combination thereof. For example, in the illustrated embodiment, the program instructions 2025 include program instructions executable to implement functions of a provider network in different embodiments. In some embodiments, the program instructions 2025 may implement multiple distinct clients, server nodes, and / or other components.

[0109] In some embodiments, the program instructions 2025 may include instructions executable to implement an operating system (not shown), which may be any of various operating systems such as UNIX, LINUX®, Solaris™, MacOS™, Windows™. Any or all of the program instructions 2025 may be provided as a computer program product or software that includes non-transitory computer-readable storage media storing instructions for programming a computer system (or other electronic device) to perform processes according to various embodiments, such as various techniques for discovering matching code sources according to an index and comparison similarity. The non-transitory computer-readable storage media may include any mechanism for storing information in a form readable by a machine (e.g., a computer), such as software, a processing application. Generally speaking, the non-transitory computer-accessible media may include magnetic or optical media, such as computer-readable storage media or memory media such as disks or DVD / CD-ROMs, coupled to the computer system 2000 via the I / O interface 2030. The non-transitory computer-readable storage media may also include any volatile or non-volatile media such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM, which may be included in some embodiments of the computer system 2000 as system memory 2020 or another type of memory. In other embodiments, the program instructions may be communicated using optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) transmitted via a communication medium such as a network and / or a wireless link, such that the program instructions may be implemented via the network interface 2040.

[0110] In some embodiments, system memory 2020 may include a data store 2045 that may be configured as described herein. Generally, system memory 2020 (e.g., data store 2045 within system memory 2020), persistent storage 2060, and / or remote storage 2070 may store data blocks, replicas of data blocks, metadata associated with data blocks and / or their states, configuration information, and / or any other information that may be used in implementing the methods and techniques described herein.

[0111] In one embodiment, I / O interface 2030 may be configured to coordinate I / O traffic between processor 2010, system memory 2020, and any peripheral devices within the system, including via network interface 2040 or other peripheral interfaces. In some embodiments, I / O interface 2030 may perform any necessary protocol, timing, or other data conversions to transform data signals from one component (e.g., system memory 2020) into a format suitable for use by another component (e.g., processor 2010). In some embodiments, I / O interface 2030 may include support for devices attached through various types of peripheral buses, such as variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, the functionality of I / O interface 2030 may be split among two or more separate components, such as a north bridge and a south bridge. Also, in some embodiments, some or all of the functions of I / O interface 2030, such as the interface to system memory 2020, may be incorporated directly into processor 2010.

[0112] The network interface 2040 may be configured to enable data to be exchanged between the computer system 2000 and other devices attached to a network, such as, for example, another computer system 2090 (which may implement one or more storage system server nodes, database engine head nodes, and / or clients of the database systems described herein). Additionally, the network interface 2040 may be configured to enable communication between the computer system 2000 and various I / O devices 2050 and / or remote storage 2070. The input / output device 2050 may, in some embodiments, include one or more display terminals, keyboards, keypads, touch pads, scanning devices, voice or optical recognition devices, or any other device suitable for inputting or reading data by the one or more computer systems 2000. The plurality of input / output devices 2050 may be present within the computer system 2000 or may be distributed across various nodes of a distributed system including the computer system 2000. In some embodiments, similar input / output devices may be separate from the computer system 2000 and may interact with one or more nodes of a distributed system including the computer system 2000 via a wired or wireless connection, such as via the network interface 2040. The network interface 2040 may generally support one or more wireless networking protocols (e.g., Wi-Fi / IEEE 802.11, or another wireless networking standard). However, in various embodiments, the network interface 2040 may support communication via any suitable wired or wireless general data network, such as, for example, other types of Ethernet networks. Additionally, the network interface 2040 may support communication via a telecommunications / telephony network, such as an analog voice network or a digital fiber communication network, communication via a storage area network, such as a fiber channel SAN, or communication via any other suitable type of network and / or protocol.In various embodiments, computer system 2000 may include more, fewer, or different components than those shown in FIG. 17 (e.g., other network interfaces such as a display, video card, audio card, peripheral devices, ATM interface, Ethernet interface, frame relay interface, etc.).

[0113] Note that any of the embodiments of the distributed systems described herein, or any of their components, may be implemented as one or more network-based services. For example, a computing cluster within a computing service may present a computing service and / or other types of services that employ the distributed computing systems described herein to a client as a network-based service. In some embodiments, the network-based service may be implemented by a software and / or hardware system designed to support interoperable machine-to-machine interactions over a network. The network-based service may have an interface described in a machine-processable format such as the Web Services Description Language (WSDL). Other systems may interact with the network-based service in a manner defined by the description of the interface of the network-based service. For example, the network-based service may define various operations that other systems may call, and may define a specific application programming interface (API) that other systems are expected to conform to when requesting various operations.

[0114] In various embodiments, a network-based service may be requested or invoked through the use of messages that include parameters and / or data associated with the network-based service request. Such messages may be formatted according to a particular markup language such as Extensible Markup Language (XML) and / or encapsulated using a protocol such as Simple Object Access Protocol (SOAP). To execute a network-based service request, a network-based service client assembles a message that includes the request and transmits the message to an addressable endpoint (e.g., a Uniform Resource Locator (URL)) corresponding to the network-based service using an Internet-based application layer transfer protocol such as the Hypertext Transfer Protocol (HTTP).

[0115] In some embodiments, a network-based service may be implemented using Representational State Transfer (“RESTful”) technology rather than message-based technology. For example, a network-based service implemented according to RESTful technology may be invoked through parameters included within HTTP methods such as PUT, GET, or DELETE rather than being encapsulated within a SOAP message.

[0116] While the above embodiments have been described in considerable detail, numerous modifications and variations can be made as will be apparent to those skilled in the art upon a complete understanding of the above disclosure. The following claims are to be construed to embrace all such modifications and changes, and accordingly, the foregoing description is to be regarded in an illustrative rather than a limiting sense.

[0117] Embodiments of the present disclosure can be described in light of the following clauses. Clause 1. A system comprising: at least one processor; A memory for storing program instructions, and the program instructions, when executed by at least one processor, cause the at least one processor to implement an integrated development environment, and the integrated development environment detecting an event that triggers a request for a code completion suggestion to be included in a code file edited using a code development system sending a request for a code completion suggestion to a code generation system, the request including one or more characters immediately preceding the cursor at a first time when an event that triggers the request for a code completion suggestion is detected receiving a code completion suggestion from the code generation system, the code completion suggestion including one or more proposed characters for inputting into the code file immediately after the cursor at the first time comparing the number of one or more proposed characters with the corresponding number of one or more actual characters input into the code file after the first time, and determining that the comparison of the number of one or more proposed characters with the corresponding number of one or more actual characters meets one or more verification criteria displaying the code completion suggestion for input into the code file in response to determining that one or more verification criteria are met. A system configured to perform the above Clause 2. The event that triggers the request is detected based at least in part on the number of keystrokes, or the amount of time elapsed since the previous code suggestion, or a combination of the number of keystrokes and the amount of time elapsed since the previous code suggestion. The system according to Clause 1 Clause 3. The integrated development environment is further configured to obtain one or more portions of the code file outside of some previous tokens as additional context, and the code completion suggestion is further based on the additional context of the one or more obtained portions of the code file outside of some previous tokens. The system according to Clause 1 or 2 Clause 4. The code generation system is implemented as part of a code development service provided by a provider network. The system according to any one of Clauses 1 to 3 Clause 5. A method comprising: detecting, by an integrated development environment, an event that triggers obtaining a code completion proposal for inclusion in a code file being edited using the integrated development environment; causing, by the integrated development environment, at a first time when an event that triggers a request for a code completion proposal is detected, generation of a code completion proposal based at least in part on one or more characters immediately preceding a cursor, the code completion proposal including at least one proposed character for input into the code file immediately after the cursor at the first time; comparing, by the integrated development environment, a number of one or more proposed characters with a corresponding number of one or more actual characters input into the code file after the first time, and determining that the comparison of the number of one or more proposed characters with the corresponding number of one or more actual characters meets one or more verification criteria; and in response to determining that one or more verification criteria are met, causing, by the integrated development environment, display of the code completion proposal for input into the code file. Clause 6. The method according to clause 5, wherein the event that triggers the request is detected based at least in part on a number of keystrokes, or an amount of time elapsed since a previous code proposal, or a combination of the number of keystrokes and the amount of time elapsed since a previous code proposal. Clause 7. The method according to clause 5 or 6, wherein the event that triggers the request is detected based at least in part on input of a particular character. Clause 8. The method according to any one of clauses 5 to 7, further comprising discarding a different code proposal generated for inclusion in the code file according to a determination that one or more verification criteria are not met according to a comparison between further characters input into the code file and the different code proposal. Clause 9. Further includes obtaining one or more portions of a code file outside of some previous tokens as additional context, and the code completion proposal is further based on the additional context of the one or more obtained portions of the code file outside of some previous tokens, and is the method according to any one of Clauses 5 to 8. Clause 10. The method according to Clause 9, wherein obtaining one or more portions of a code file outside of some previous tokens includes performing one or more regular expression searches. Clause 11. Further includes obtaining one or more portions of different code files as additional context, and the code completion proposal is further based on the additional context of the one or more obtained portions of the different code files, and is the method according to Clause 5. Clause 12. Causing the generation of a code completion proposal includes sending a request for the code completion proposal to a code development service provided by a provider network, and is the method according to any one of Clauses 5 to 11. Clause 13. One or more non-transitory computer-readable storage media storing program instructions, wherein the program instructions, when executed on or across one or more computing devices, cause the one or more computing devices to detect an event that triggers obtaining a code completion proposal for inclusion in a code file being edited using a code development system; obtain, at a first time when an event that triggers a request for a code completion proposal is detected, a code completion proposal based at least in part on one or more characters immediately preceding a cursor, the code completion proposal including at least one proposed character for inputting into the code file immediately after the cursor at the first time; compare the number of one or more proposed characters with the corresponding number of one or more actual characters input into the code file after the first time, and determine that the comparison of the number of one or more proposed characters with the corresponding number of one or more actual characters meets one or more verification criteria; One or more non - transitory computer - readable storage media that cause, in response to determining that one or more verification criteria are met, a code completion suggestion to be displayed for input into a code file. Clause 14. The one or more non - transitory computer - readable storage media according to clause 13, wherein an event that triggers a request is detected based at least in part on the number of keystrokes, or the amount of time elapsed since the previous code suggestion, or a combination of the number of keystrokes and the amount of time elapsed since the previous code suggestion. Clause 15. The one or more non - transitory computer - readable storage media according to clause 13 or 14, wherein an event that triggers a request is detected based at least in part on the input of a specific character. Clause 16. One or more non - transitory computer - readable storage media according to any one of clauses 13 - 15, which store further program instructions that cause, when executed on or across one or more computing devices, the one or more computing devices to discard different code suggestions generated for inclusion in a code file according to a determination that a comparison between a further character input into the code file and the different code suggestions does not meet one or more verification criteria. Clause 17. One or more non - transitory computer - readable storage media according to any one of clauses 13 - 16, which store further program instructions that cause, when executed on or across one or more computing devices, the one or more computing devices to obtain one or more portions of a code file outside of some previous tokens as additional context, and wherein the code completion suggestion is further based on the additional context of the one or more obtained portions of the code file outside of some previous tokens. Clause 18. The one or more non - transitory computer - readable storage media according to clause 17, wherein the program instructions cause the one or more computing devices to perform one or more regular expression searches when obtaining one or more portions of a code file outside of some previous tokens. Clause 19. When executed on or across one or more computing devices, further store program instructions that cause the one or more computing devices to obtain one or more portions of different code files as additional context, and the code completion proposals are further based on the additional context of the one or more obtained portions of different code files, one or more non-transitory computer-readable storage media as described in Clause 13. Clause 20. When obtaining code completion proposals, the program instructions cause the one or more computing devices to send a request to a code development service provided by a provider network for the code completion proposals, one or more non-transitory computer-readable storage media as described in any one of Clauses 13 to 19. Clause 21. A system comprising: at least one processor; a memory storing program instructions that cause the at least one processor to implement a programming language conversion system when executed by the at least one processor, the programming language conversion system receives a request to convert an evaluation data set specified in a first programming language via an interface of the programming language conversion system, wherein different items of the evaluation data set correspond to different respective evaluation tests of a code generation system; converts individual items of the different items of the data set into a second programming language; converts the function signature of an item in the first programming language into the second programming language; converts one or more test statements of an item in the first programming language into the second programming language; sends a request to the code generation system to generate a body of the function signature converted in the second programming language according to a prompt within the item; Receiving, from a code generation system, a body of a function signature in a second programming language, Storing, as part of a new evaluation dataset, the converted individual items of different items of an evaluation dataset, Clause 22. The second programming language is the system described in Clause 21, specified in the requirements for converting an evaluation dataset. Clause 23. The programming language conversion system is further configured to identify, for each type in a first programming language, one or more parameters in a function signature mapped to a corresponding type in a second programming language, in order to convert a function signature of an item in the first programming language to the second programming language, the system described in Clause 21 or 22. Clause 24. The programming language conversion system is further configured to identify, for each type in a first programming language, one or more parameters in one or more test statements mapped to a corresponding type in a second programming language, in order to convert one or more test statements of an item in the first programming language to the second programming language, the system described in any one of Clauses 21 to 23. Clause 25. A method comprising: Receiving, via an interface of a programming language conversion system, an evaluation dataset specified in a first programming language, wherein different items of the evaluation dataset correspond to different respective evaluation tests of a code generation system; Converting individual items of different items of the dataset to a second programming language; Converting, by the programming language conversion system, a function signature of an item in the first programming language to the second programming language; The programming language conversion system converts one or more test statements of an item in a first programming language into a second programming language, causes the body of the converted function signature to be generated in the second programming language according to a prompt within the item used as input to a machine learning model trained to generate code in the second programming language by the programming language conversion system, and stores the converted individual items of different items in the evaluation dataset as part of a new evaluation dataset. The method according to clause 25, further comprising receiving a request to convert an evaluation dataset specifying a second programming language. Clause 27. converting individual items of different items in the dataset into a third programming language, converting the function signature of an item in the first programming language into the third programming language by the programming language conversion system, converting one or more test statements of an item in the first programming language into the third programming language by the programming language conversion system, causes the body of the function signature converted in the third programming language to be generated in the third programming language according to a prompt within the item used as input to a second machine learning model trained to generate code in the third programming language by the programming language conversion system, and stores the converted individual items among the different items in the evaluation dataset in the third programming language as part of a second new evaluation dataset. The method according to clause 25 or 26. The method according to any one of clauses 25 to 27, wherein converting the function signature of an item in a first programming language into a second programming language includes identifying the respective types in the first programming language of one or more parameters in the function signature that are mapped to corresponding types in the second programming language. Clause 29. Converting one or more test statements of an item in a first programming language into a second programming language includes identifying the respective types in the first programming language of one or more parameters within the test statement that are mapped to corresponding types in the second programming language, the method according to any one of clauses 25 to 28. Clause 30. The method according to any one of clauses 25 to 29, wherein the programming language conversion system is implemented as part of a code development service provided by a provider network, and a request to perform the conversion is received from a client of the provider network. Clause 31. The method according to any one of clauses 25 to 30, further including performing natural language conversion on a portion of a prompt according to a second programming language. Clause 32. Causing the body of the converted function signature to be generated in a second programming language includes sending a request to a code generation system implemented as part of a code development service provided by a provider network, the method according to any one of clauses 25 to 31. Clause 33. One or more non-transitory computer-readable storage media storing program instructions, which, when executed on or across one or more computing devices, cause the one or more computing devices to receive, via an interface of the programming language conversion system, an evaluation data set specified in a first programming language, wherein different items of the evaluation data set correspond to different evaluation tests of the code generation system, and receive Converting individual items of different items in a dataset to a second programming language, and converting, by a programming language conversion system, the function signature of an item in a first programming language to a second programming language, converting, by a programming language conversion system, one or more test statements of an item in a first programming language to a second programming language, causing, by a programming language conversion system, the body of the converted function signature to be generated in a second programming language according to a prompt within the item used as input to a machine learning model trained to generate code in the second programming language, and storing the converted individual items of different items in an evaluation dataset as part of a new evaluation dataset, on one or more non-transitory computer-readable storage media. Clause 34. Further programming instructions stored on one or more non-transitory computer-readable storage media according to clause 33, the further programming instructions causing one or more computing devices to further implement receiving, when executed on or across one or more computing devices, a request to convert an evaluation dataset specifying a second programming language. Clause 35. Further program instructions stored, the program instructions causing one or more computing devices to, when executed on or across one or more computing devices, convert individual items of different items in a dataset to a third programming language, and convert, by a programming language conversion system, the function signature of an item in a first programming language to a third programming language, convert, by a programming language conversion system, one or more test statements of an item in a first programming language to a third programming language, The body of a function signature converted into a third programming language is generated in the third programming language according to a prompt within an item used as input to a second machine learning model trained to generate code in the third programming language by a programming language conversion system, One or more non-transitory computer-readable storage media according to clause 33 or 34, further implementing storing, as part of a second new evaluation dataset, each converted item among different items of an evaluation dataset in a third programming language. Clause 36. When converting the function signature of an item in a first programming language into a second programming language, program instructions cause one or more computing devices to identify, in the first programming language, the respective types of one or more parameters in the function signature mapped to corresponding types in the second programming language. One or more non-transitory computer-readable storage media according to any one of clauses 33 to 36. Clause 37. When converting the function signature of an item in a first programming language into a second programming language, program instructions cause one or more computing devices to identify, in the first programming language, the respective types of one or more parameters in a test statement mapped to corresponding types in the second programming language. One or more non-transitory computer-readable storage media according to any one of clauses 33 to 36. Clause 38. A programming language conversion system is implemented as part of a code development service provided by a provider network, and a request to execute the conversion is received from a client of the provider network. One or more non-transitory computer-readable storage media according to any one of clauses 33 to 37. One or more non-transitory computer-readable storage media according to any one of clauses 33 to 38, storing further programming instructions that cause one or more computing devices, when executed on or across one or more computing devices, to further implement natural language conversion on a portion of a prompt according to a second programming language. Clause 40. When generating the body of the transformed function signature in a second programming language, the program instructions cause one or more computing devices to send a request to a code generation system implemented as part of a code development service provided by a provider network. One or more non-transitory computer-readable storage media according to any one of clauses 33 to 39. Clause 41. A system comprising: at least one processor; a memory storing program instructions that cause the at least one processor to implement a code generation system when executed by the at least one processor, the code generation system receives input programming code and performs next token prediction on the input programming code; determines word boundaries with respect to a tokenizer of the input text, where the rightmost boundary includes a partial token, and the partial token is used as a prompt suffix; identifies one or more possible tokens that match the prompt suffix by starting with or ending with the prompt suffix from a plurality of possible tokens; Filtering the next token prediction according to one or more identified possible tokens, where the next token prediction is generated by applying a machine learning model trained to predict the next token of programming code to the remaining part of the input programming code that does not include some backtrack tokens corresponding to the pre-token, and the filtering is performed for one or more iterations that remove one or more characters from the left of the partial token that matches one of the next token predictions after each iteration until there are no remaining characters within the partial token, and filtering; A system configured to provide the last token prediction among the next token predictions as the next token prediction for the input programming code. Clause 42. The system according to clause 41, wherein to identify one or more possible tokens that match a partial token from a plurality of possible tokens, the code generation system is configured to access a trie data structure that stores the plurality of possible tokens. Clause 43. The system according to clause 41 or 42, wherein to identify one or more possible tokens that match a partial token from a plurality of possible tokens, the code generation system is configured to access a mask cache that stores possible partial tokens. Clause 44. The system according to any one of clauses 41 to 43, wherein the code generation system is implemented as part of a code development service provided by a provider network, and the input programming code is received as part of a request to generate a code proposal for a code file received from a client of the provider network. Clause 45. A method, Receiving input text by a text generation system and performing a next token prediction for the input text; Determining word boundaries for a tokenizer for the input text by the text generation system for the input text, where the rightmost boundary includes a partial token, and the partial token is used as a prompt suffix. By a text generation system, identifying one or more possible tokens that match a prompt suffix from a plurality of possible tokens by starting with or ending with the prompt suffix; By a text generation system, filtering the next token prediction according to the identified one or more possible tokens, wherein the next token prediction is generated by applying a machine learning model to the remaining part of the input text that does not include some backtrack tokens corresponding to pre-tokens, and the filtering is performed for one or more iterations that remove one or more characters from the left of the partial token that matches one of the next token predictions after each iteration until there are no remaining characters in the partial token; By a text generation system, providing the last one of the next token predictions as the next token prediction, a method comprising. Clause 46. The method according to clause 45, wherein identifying one or more possible tokens that match a partial token from a plurality of possible tokens includes accessing a trie data structure that stores the plurality of possible tokens. Clause 47. The method according to clause 46 or 45, wherein the trie data structure is used to generate the next token prediction using different machine learning models. Clause 48. The method according to any one of clauses 45 to 47, wherein identifying one or more possible tokens that match a partial token from a plurality of possible tokens includes accessing a mask cache that stores possible partial tokens. Clause 49. The method according to any one of clauses 45 to 48, further comprising determining, by the text generation system, the number of backtrack tokens up to the maximum number of backtrack tokens. Clause 50. The method according to any one of clauses 45 to 49, wherein the input text is received as part of a request for providing the next token prediction, and the next token prediction is provided as a response to the request. Clause 51. The text generation system is implemented as part of an auto-complete application, and is the method described in any one of Clauses 45 to 50. Clause 52. The text generation system is implemented as part of a code development service provided by a provider network, and the input text is received as part of a request to generate a code proposal for a code file received from a client of the provider network, and is the method described in any one of Clauses 45 to 51. One or more non-transitory computer-readable storage media storing program instructions, which, when executed on one or more computing devices or across one or more computing devices, cause the one or more computing devices to receive input text and perform next token prediction on the input text; determine word boundaries with respect to a tokenizer for the input text, where the rightmost boundary includes a partial token, and the partial token is used as a prompt suffix; identify from a plurality of possible tokens one or more possible tokens that match the prompt suffix by starting with or ending with the prompt suffix; filter the next token prediction according to the identified one or more possible tokens, where the next token prediction is generated by applying a machine learning model to the remaining portion of the input text that does not include some backtrack tokens corresponding to a pre-token, and the filtering is performed for one or more iterations that remove one or more characters from the left of a partial token that matches one of the next token predictions after each iteration until there are no remaining characters within the partial token; provide the last token prediction among the next token predictions as the next token prediction for the input text. One or more non-transitory computer-readable storage media. Clause 54. When identifying one or more possible tokens that match a partial token from a plurality of possible tokens, the program instructions cause one or more computing devices to access one or more non-transitory computer-readable storage media described in Clause 53 that store a trie data structure of the plurality of possible tokens. Clause 55. The trie data structure is one or more non-transitory computer-readable storage media described in Clause 54 that are used to generate next token predictions using different machine learning models. Clause 56. When identifying one or more possible tokens that match a partial token from a plurality of possible tokens, the program instructions cause one or more computing devices to access a mask cache that stores possible partial tokens, the one or more non-transitory computer-readable storage media described in any one of Clauses 53 to 55. Clause 57. When executed on or across one or more computing devices, further program instructions are stored that cause one or more computing devices to further determine, by a text generation system, the number of backtrack tokens up to a maximum number of backtrack tokens, the one or more non-transitory computer-readable storage media described in any one of Clauses 53 to 56. Clause 58. The input text is received as part of a request for providing next token predictions, and the next token predictions are provided as a response to the request, the one or more non-transitory computer-readable storage media described in any one of Clauses 53 to 57. Clause 59. One or more computing devices are implemented as part of an autocomplete application, the one or more non-transitory computer-readable storage media described in any one of Clauses 53 to 58. Clause 60. One or more computing devices are implemented as part of a code development service provided by a provider network, and the input text is received as part of a request to generate a code proposal for a code file received from a client of the provider network, one or more non-transitory computer-readable storage media according to any one of Clauses 53 to 59. Clause 61. A system comprising: at least one processor; a memory storing program instructions that, when executed by the at least one processor, cause the at least one processor to implement a machine learning system, the machine learning system comprising: receiving programming code for training a machine learning model to predict a next programming code token when an input programming code token is provided; analyzing the programming code to determine a plurality of tokens from the programming code; randomly segmenting different tokens of the plurality of tokens into respective plurality of subtokens; using the plurality of tokens each including a respective plurality of subtokens as a training data set to train a machine learning model to predict a next programming code token when a programming code token input is provided. Clause 62. The system according to Clause 61, wherein the random segmentation of different tokens of the tokens is performed according to subword regularization techniques. Clause 63. The system according to Clause 61 or 62, wherein the machine learning model is a pre-trained machine learning model. Clause 64. The system according to any one of Clauses 61 to 63, wherein the machine learning system is implemented as part of a code development service provided by a provider network to train a machine learning model for generating code proposals. Clause 65. A method comprising: In a machine learning system, receiving text data for training a machine learning model to predict the next text token given an input text token; determining, by the machine learning system, a plurality of tokens from the text data; randomly segmenting, by the machine learning system, different tokens among the plurality of tokens into respective pluralities of subtokens; training, by the machine learning system, the machine learning model to predict the next text token when an input text token is given, using the plurality of tokens each including respective pluralities of subtokens as a training data set. A method comprising: Article 66. The method according to Article 65, wherein the random segmentation of different tokens among the tokens is performed according to subword regularization techniques. Article 67. The method according to Article 64 or 65, wherein the machine learning model is a pre-trained machine learning model. Article 68. The method according to any one of Articles 65 to 67, wherein the text data is code written in a programming language, and the next text token and the input text token are respective programming code tokens. Article 69. Determining a plurality of tokens from the text data and randomly segmenting different tokens among the plurality of tokens into respective pluralities of subtokens are performed by a tokenizer, and the tokenizer is applicable for training a second machine learning model to predict the next code token in a second programming language. The method according to Article 68. Article 70. The method according to any one of Articles 65 to 69, further comprising storing the trained machine learning model at a location specified in a requirement for training the machine learning model. Article 71. The method according to any one of Articles 65 to 70, further comprising deploying the trained machine learning model as part of an auto-complete application. Clause 72. The machine learning system is implemented as part of a code development service provided by a provider network to train a machine learning model for generating code proposals, in accordance with any one of Clauses 65 to 71. Clause 73. One or more non-transitory computer-readable storage media storing program instructions, the program instructions, when executed on one or more computing devices or across one or more computing devices, cause the one or more computing devices to receive text data for training a machine learning model to predict the next text token when an input text token is provided; determine a plurality of tokens from the text data; randomly segment different tokens among the plurality of tokens into respective sub-tokens; use the plurality of tokens each including respective sub-tokens as a training data set to train a machine learning model to predict the next text token when an input text token is provided. One or more non-transitory computer-readable storage media implementing the above. Clause 74. The random segmentation of different tokens among the tokens is performed according to sub-word regularization techniques. One or more non-transitory computer-readable storage media according to Clause 73. Clause 75. The machine learning model is a pre-trained machine learning model. One or more non-transitory computer-readable storage media according to Clause 73 or 74. Clause 76. The text data is code written in a programming language, and the next text token and the input text token are respective programming code tokens. One or more non-transitory computer-readable storage media according to any one of Clauses 73 to 75. Clause 77. Determining a plurality of tokens from text data and randomly segmenting different tokens among the plurality of tokens into respective plural sub-tokens is performed by a tokenizer, and the tokenizer is applicable for training a second machine learning model for predicting the following code tokens in a second programming language, one or more non-transitory computer-readable storage media according to Clause 76. Clause 78. One or more non-transitory computer-readable storage media according to any one of Clauses 73 to 77, storing further program instructions that, when executed on or across one or more computing devices, further cause the one or more computing devices to store a machine learning model trained at a location specified in requirements for training the machine learning model. Clause 79. One or more non-transitory computer-readable storage media according to any one of Clauses 73 to 78, storing further program instructions that, when executed on or across one or more computing devices, further cause the one or more computing devices to deploy a trained machine learning model as part of an autocomplete application. Clause 80. One or more non-transitory computer-readable storage media according to any one of Clauses 73 to 79, wherein a machine learning system is implemented as part of a code development service provided by a provider network for training a machine learning model for generating code proposals.< / object> < / integer> < / double> < / double> < / integer> < / integer> < / t> < / t> < / t> < / t> < / t> < / t> < / t>

Claims

1. A system comprising: at least one processor; and a memory storing program instructions that, when executed by the at least one processor, cause the at least one processor to implement an integrated development environment, the integrated development environment comprising: detecting an event that triggers obtaining a code completion proposal for inclusion in a code file being edited using the integrated development environment; causing generation of the code completion proposal, at a first time when the event triggering the request for the code completion proposal is detected, based at least in part on one or more characters immediately preceding a cursor, the code completion proposal including, at the first time, one or more proposed characters for inputting into the code file immediately after the cursor; comparing the number of the one or more proposed characters to a corresponding number of one or more actual characters input into the code file after the first time, and determining that the comparison of the number of the one or more proposed characters and the corresponding number of the one or more actual characters meets one or more verification criteria; and in response to determining that the one or more verification criteria are met, displaying the code completion proposal for input into the code file.

2. The system of claim 1, wherein the event triggering the request is detected based at least in part on a number of keystrokes, or an amount of time elapsed since a previous code proposal, or a combination of the number of keystrokes and the amount of time elapsed since the previous code proposal.

3. The system of claim 1 or 2, wherein the integrated development environment is further configured to obtain one or more portions of the code file outside of some previous tokens as additional context, and the code completion proposal is further based on the additional context of the one or more obtained portions of the code file outside of some previous tokens.

4. The system of any one of claims 1 to 3, wherein the code generation system is implemented as part of a code development service provided by a provider network.

5. A method comprising: Detecting, by the integrated development environment, an event that triggers obtaining a code completion suggestion for inclusion in a code file being edited using the integrated development environment; causing, by the integrated development environment, generation of the code completion suggestion at a first time when the event that triggers the request for the code completion suggestion is detected, based at least in part on one or more characters immediately preceding a cursor, wherein the code completion suggestion includes, at the first time, one or more proposed characters for input into the code file immediately after the cursor; comparing, by the integrated development environment, a number of the one or more proposed characters with a corresponding number of one or more actual characters input into the code file after the first time, and determining that the comparison of the number of the one or more proposed characters and the corresponding number of the one or more actual characters meets one or more verification criteria; displaying, by the integrated development environment, the code completion suggestion for input into the code file in response to determining that the one or more verification criteria are met. A method comprising.

6. The method of claim 5, wherein the event that triggers the request is detected based at least in part on a number of keystrokes, or an amount of time elapsed since a previous code suggestion, or a combination of the number of the keystrokes and the amount of time elapsed since the previous code suggestion.

7. The method according to claim 5 or 6, wherein the event that triggers the request is detected based at least in part on input of a specific character.

8. Further comprising discarding the different code suggestion generated for inclusion in the code file according to a determination that the one or more verification criteria are not met according to a comparison between the additional characters input into the code file and the different code suggestion. The method according to any one of claims 5 to 7.

10. Further comprising obtaining, as additional context, one or more portions of the code file outside of some previous tokens, wherein the code completion suggestion is further based on the additional context of the one or more obtained portions of the code file outside of some previous tokens. The method according to any one of claims 5 to 8.

10. The method of claim 9, wherein obtaining the one or more portions of the code file outside of the previous several tokens includes performing one or more regular expression searches.

11. The method of claim 5, further comprising obtaining one or more portions of different code files as additional context, wherein the code completion proposal is further based on the additional context of the one or more obtained portions of the different code files.

12. The method according to any one of claims 5 to 11, wherein causing the generation of the code completion proposal includes sending a request for the code completion proposal to a code development service provided by a provider network.

13. One or more non-transitory computer-readable storage media storing program instructions, wherein the program instructions, when executed on or across one or more computing devices, cause the one or more computing devices to detect an event that triggers obtaining a code completion proposal for inclusion in a code file being edited using the code development system; obtaining the code completion proposal, at a first time when the event that triggers the request for the code completion proposal is detected, at least partially based on one or more characters immediately preceding a cursor, wherein the code completion proposal includes, at the first time, one or more proposed characters for input into the code file immediately after the cursor; comparing the number of the one or more proposed characters with a corresponding number of one or more actual characters input into the code file after the first time, and determining that the comparison of the number of the one or more proposed characters and the corresponding number of the one or more actual characters meets one or more verification criteria; displaying the code completion proposal for input into the code file in response to determining that the one or more verification criteria are met.

14. The one or more non-transitory computer-readable storage media of claim 13, wherein the event that triggers the request is detected based at least in part on the number of keystrokes, or the amount of time elapsed since the previous code suggestion, or a combination of the number of keystrokes and the amount of time elapsed since the previous code suggestion.

15. The one or more non-transitory computer-readable storage media of claim 13 or 14, wherein, when obtaining the code completion suggestion, the program instructions cause the one or more computing devices to send a request for the code completion suggestion to a code development service provided by a provider network.

Citation Information

Patent Citations

  • Japanese predication input method / System and recording medium programming and recording method

    JP2000010969A

  • Editor program, editing method, editor device, and recording medium

    JP2004341939A

  • Computer program, and device and method for receiving input of source program

    JP2010097426A

  • Character recommendation method, character recommendation device, computer device, and program

    JP2022540736A

  • Character recommending method and apparatus, and computer device and storage medium

    US20210294432A1