Methods, apparatus, and articles of manufacture to generate usage dependent code embedding
By considering the usage and invocation context of code snippets in code embedding technology, higher-quality embeddings are generated, solving the problem of failure to utilize context in existing technologies and improving the performance of code intelligence tasks.
Patent Information
- Application Number
- CN202511996625.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-17
- Filing Date
- 2022-12-02
- Publication Date
- 2026-03-20
AI Technical Summary
Existing code embedding techniques fail to effectively utilize the contextual information of code snippets, resulting in low-quality embedded code that negatively impacts the performance of code intelligence tasks.
By taking into account the usage and calling context of code snippets, usage-dependent code embeddings are generated, including selecting lines of code surrounding the code snippet as context and using transformer models for embedding processing.
It improves the quality of code embedding, thereby enhancing the performance of code intelligence tools, particularly in tasks such as code summarization, clone detection, and repair.
Smart Images

Figure CN121704898A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to language processing, and more specifically to methods, apparatus, and artifacts for generating usage-dependent code embeddings. Background Technology
[0002] Artificial intelligence (AI), including machine learning (ML), deep learning (DL), and / or other artificial machine-driven logic, enables machines (e.g., computers, logic circuits, etc.) to use models to process input data to generate outputs based on patterns and / or associations previously learned by the model through a training process. For example, the model may be trained with data to recognize patterns and / or associations and follow these patterns and / or associations when processing input data, such that (one or more) other inputs result in (one or more) outputs consistent with the recognized patterns and / or associations.
[0003] With the development of AI, developers have applied it to many different fields. One application area of AI is code intelligence tasks. For example, AI can be used for automatic code captioning, code clone detection, code completion, and so on. To aid in code intelligence tasks, AI models utilize code embeddings. Code embeddings are one or more vectors that capture information within a piece of code. DL models may require one or more input vectors to be real-valued vectors (e.g., vectors of real numbers). Therefore, many code embeddings consist of one or more fixed-size real-valued vectors that capture information from a piece of code. Summary of the Invention
[0004] One aspect of this disclosure provides an apparatus for generating usage-dependent code embeddings. The apparatus includes: at least one memory; instructions; and processor circuitry for executing the instructions to at least: obtain code comprising a code segment to be processed by an artificial intelligence (AI) model; select a usage context for the code segment, the usage context including at least one LOC preceding the code segment or preceding a line of code (LOC) that calls the code segment, the code segment itself, and at least one LOC following the code segment or following the LOC that calls the code segment; generate a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code segment, and generate a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a turn-off token; and concatenate a transformed token embedding vector of the turn-off token and a fifth list of one or more transformed token embedding vectors for the first list of one or more token embedding vectors.
[0005] Another aspect of this disclosure provides a method for generating usage-dependent code embeddings. The method includes: obtaining code comprising a code snippet to be processed by an artificial intelligence (AI) model; selecting a usage context for the code snippet, the usage context including at least one LOC preceding or before a line of code (LOC) invoking the code snippet, the code snippet itself, and at least one LOC following or after the LOC invoking the code snippet; generating a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code snippet, and generating a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a closing token; and concatenating a transformed token embedding vector of the closing token and a fifth list of one or more transformed token embedding vectors from the first list of one or more token embedding vectors.
[0006] Another aspect of this disclosure provides an apparatus for generating usage-dependent code embeddings. The apparatus includes: means for parsing code, configured to: obtain code comprising code segments to be processed by an artificial intelligence (AI) model; and select a usage context for the code segment, the usage context including at least one LOC preceding the code segment or preceding a line of code (LOC) that calls the code segment, the code segment, and at least one LOC following the code segment or following the LOC that calls the code segment; means for embedding code, configured to generate a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code segment, and a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a turn-off token; and means for concatenating vectors, configured to concatenate a transformed token embedding vector of a turn-off token and a fifth list of one or more transformed token embedding vectors for the first list of one or more token embedding vectors.
[0007] Another aspect of this disclosure provides a machine-readable medium. This machine-readable medium includes code that, when executed, causes a machine to perform the aforementioned method of generating code that depends on the embedded code used. Attached Figure Description
[0008] Figure 1 The network graph includes example service providers with example code embedded circuits and example semantic search engines.
[0009] Figure 2 The block diagram illustrates Figure 1 This is an example implementation of a code embedding circuit used to generate code embeddings that depend on the code used.
[0010] Figure 3 The flowchart illustrates the trainingFigure 1 and / or Figure 2 The code embeds an example of the circuit process.
[0011] Figure 4 The flowchart illustrates an example process for generating code embeddings that depend on the code used.
[0012] Figure 5 The flowchart illustrates the process of generating example code embeddings that depend on the example function that is used twice in the example code.
[0013] Figure 6 The flowchart represents what can be implemented by example processor circuitry and / or instantiated. Figure 1 and / or Figure 2 The code is embedded in the circuit to execute the trained example machine-readable instructions and / or example operations.
[0014] Figure 7 The flowchart represents what can be implemented by example processor circuitry and / or instantiated. Figure 1 and / or Figure 2 The code embedding circuitry is used to generate example machine-readable instructions and / or example operations that depend on the code embedding used.
[0015] Figure 8 This is a block diagram of an example processing platform including processor circuitry configured to perform and / or instantiate... Figure 6 and / or Figure 7 Example machine-readable instructions and / or example operations to implement Figure 1 and / or Figure 2 The code is embedded in the circuit.
[0016] Figure 9 yes Figure 8 A block diagram illustrating an example implementation of the processor circuit.
[0017] Figure 10 yes Figure 8 A block diagram of another example implementation of the processor circuit.
[0018] Figure 11 This is a block diagram of an example software distribution platform (e.g., one or more servers) used to distribute software (e.g., with...) Figure 6 and / or Figure 7 The software corresponding to the example machine-readable instructions is distributed to client devices associated with end users and / or consumers (e.g., for licensing, selling and / or using), retailers (e.g., for selling, reselling, licensing and / or sublicensing) and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, for example, retailers and / or other end users such as direct purchase customers).
[0019] Generally, the same reference numerals will be used throughout all (one or more) the accompanying drawings and the accompanying written description to refer to the same or similar parts. The drawings are not to scale. As used herein, references to connections (e.g., attachment, coupling, joining, engagement) may include intermediate members between the elements mentioned by the connection and / or relative movement between these elements, unless otherwise indicated. Therefore, references to connections do not necessarily imply that two elements are directly connected and / or have a fixed relationship with each other.
[0020] Unless otherwise specifically stated, this document uses descriptive terms such as “first,” “second,” “third,” etc., without indicating or otherwise suggesting any priority, physical order, arrangement in a list, and / or any sorting, but merely as labels and / or arbitrary names to distinguish elements for ease of understanding of the disclosed examples. In some examples, the descriptive term “first” may be used to refer to an element in a detailed description, while the same element may be referred to in the claims using different descriptive terms, such as “second” or “third.” In such cases, it should be understood that such descriptive terms are only used to explicitly identify those elements that may, for example, share the same name in other cases.
[0021] As used herein, the phrase “communicate with”—including its variations—covers direct communication and / or indirect communication via one or more intermediate components, without requiring direct physical (e.g., wired) communication and / or continuous communication, but also including selective communication at periodic intervals, scheduled intervals, non-periodic intervals, and / or one-off events. As used herein, “processor circuitry” is defined as including (i) one or more dedicated electrical circuits configured to perform one or more specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based electrical circuits programmed with instructions to perform specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuitry include programmable microprocessors, field-programmable gate arrays (FPGAs) with instantiable instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers and integrated circuits, such as application-specific integrated circuits (ASICs). For example, an XPU can be implemented by a heterogeneous computing system that includes multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or combinations thereof) and one or more application programming interfaces (APIs) that can assign computational tasks to any one or more of these multiple types of processing circuitry best suited to perform the computational task. Detailed Implementation
[0022] Code intelligence tasks (e.g., automatic code captioning, code clone detection, code completion, etc.) rely on high-quality code embeddings. For example, when training AI-based models (e.g., neural networks (NNs)) to perform code intelligence tasks, these AI-based models may expect real-valued code embeddings as input. To generate one or more code embeddings, code embedding circuitry processes the input code, as further described herein. After generating one or more code embeddings, the code embedding circuitry provides one or more code embeddings to one or more downstream AI / ML models that implement the code intelligence task.
[0023] As mentioned above, code intelligence tasks include clone detection, code summarization, and code remediation. Clone detection refers to detecting whether two code snippets implement the same functionality. For example, a user can input a code snippet into a code intelligence system (e.g., one or more AI / ML models) to determine if that snippet implements the same functionality as any code indexed by the system. Clone detection is useful in situations where the user has requested the code intelligence system to recommend alternative implementations of the same functionality as the input code snippet. Thus, clone detection can help developers identify vulnerabilities and / or find more efficient implementations of code snippets.
[0024] Code summarization refers to generating a natural language description of the functionality of an input code snippet. Generating such a description reduces code documentation work and makes understanding legacy uncommented code easier by adding comments. Code fixing refers to modifying an input code snippet to fix functional and / or syntax flaws. One or more downstream AI / ML models fix the input code snippet and output (e.g., emit) a fixed version of the input code snippet (e.g., to a user).
[0025] This is not an exhaustive list of code intelligence tasks. Many code intelligence tasks rely on high-quality code embeddings as input. The quality of the code embeddings strongly influences the performance of downstream AI / ML models that implement code intelligence tasks. Therefore, the examples disclosed in this paper include code embedding circuits that capture high-level semantic information associated with input code snippets.
[0026] Existing code embedding techniques generally fall into two categories: embeddings built from structured code representations and embeddings built from code text. Structured techniques construct a graph-based intermediate representation of the input code, which is then used to generate code embeddings. The structured representation can be purely syntactic (e.g., reflecting the arrangement of text tokens in the code). For example, the AROMA system extracts hand-designed features from syntactic code representations to construct code embeddings, which are then used to measure the similarity between two pieces of code, while code2seq models use a standard tree-like code representation called an abstract syntax tree and use random paths within the tree to construct embeddings. Structured representations can also include semantic information, such as indicating that bracket tokens "(" and ")" in one line of code (LOC) (e.g., in LOC: Z = find max(arr)) signify a function call, while bracket tokens in another LOC (e.g., in LOC: C = (A+B)*2) indicate mathematical grouping and / or priority levels. MISIM technology uses structured code representations with rich and customizable semantic structures to construct code embeddings, while contextual flow graphs (XFG) enrich the structured representations with data and control flow information. The semantic-based structured representations result in higher-quality embeddings, as evidenced by performance improvements in downstream code intelligence tasks built upon these embeddings.
[0027] Embedding techniques that construct embeddings from code text include techniques like CodeBERT, which construct code embeddings directly from code text without using any intermediate structural representations. Existing techniques for extracting code embeddings are trainable (e.g., existing techniques include free parameters that are tuned using several optimization and / or training iterations to improve the quality of the generated embeddings). Embedding optimization objectives can be self-supervised (e.g., the embedding optimization objective does not depend on any downstream task) or supervised objectives that depend on a specific downstream task.
[0028] Existing embedding techniques do not consider the context in which code snippets are used. Because they fail to take context into account, existing techniques miss potentially relevant information included in the usage context(s) of the code snippet(s). While theoretically the function body (e.g., the function body itself) might be sufficient for an expert programmer to infer the types of possible function arguments, the usage context provides a more readily available source of information about function arguments. This source of information is not being utilized by existing techniques. Furthermore, when embedding code snippets, existing techniques do not improve embedding over time as programmers use the code snippets.
[0029] Conversely, the examples disclosed herein leverage an increasing amount of usage context to improve code embedding. The examples disclosed herein include methods, devices, and artifacts for generating usage-dependent code embeddings. While existing code embedding methods only consider the code snippet being embedded, the examples disclosed herein consider the context in which the code snippet is used and / or invoked. By considering context, the examples disclosed herein produce higher-quality code embeddings, thus improving the performance of a wide range of code intelligence tools. The context in which a code snippet is used and / or invoked is referred to as the usage context of the code snippet. As used herein, the usage context of a code snippet refers to one or more LOCs surrounding one or more lines of code (LOCs) that use and / or invoke the code snippet (e.g., before and / or after them) and one or more LOCs that use and / or invoke the code snippet. When the code snippet is a function, the usage context includes the LOCs surrounding the LOC that invokes the function and the LOC that invokes the function. If the code snippet is used multiple times in the input code, then the examples disclosed herein select multiple usage contexts for the code snippet.
[0030] Figure 1 Network graph 100 includes an example service provider 102 with an example code embedding circuit 104 and an example semantic search engine 106. Network graph 100 includes the example service provider 102, an example network 108, an example database 110, an example version control system (VCS) 112, and an example user device 114. Figure 1 In the example, example service provider 102, example database 110, example VCS 112, example user equipment 114 and / or one or more additional devices are communicatively coupled via example network 108.
[0031] exist Figure 1 In the illustrated example, service provider 102 is implemented by processor circuitry. For example, service provider 102 is implemented by one or more servers executing one or more trained AI-based models (e.g., ML models, NN models, DL models, etc.) and / or instantiating and / or executing instructions to implement one or more AI-based models as peripheral components. As described above, service provider 102 includes code embedding circuitry 104 and a semantic search engine 106. In some examples, service provider 102 includes interface circuitry to obtain code including code snippets to be processed by the AI / ML model (e.g., semantic search engine 106) and / or return results generated by the AI / ML model (e.g., semantic search engine 106).
[0032] exist Figure 1In the illustrated example, service provider 102 provides one or more services and / or products to the end user. For example, service provider 102 provides one or more trained models for download, hosts a web interface through which the user can access the one or more models, and so on. The one or more models provided by service provider 102 may include one or more models implementing code embedding circuit 104 and / or one or more models implementing semantic search engine 106. In some examples, service provider 102 provides the end user with plug-ins implementing code embedding circuit 104 and / or semantic search engine 106. In this way, the end user can implement code embedding circuit 104 and / or semantic search engine 106 locally (e.g., at user device 114).
[0033] In some examples, end users may implement the code embedding circuit 104 and / or the semantic search engine 106 as plugins to an integrated development environment (IDE) installed on user device 114. In some examples, the instructions for implementing the code embedding circuit 104 and / or the semantic search engine 106 may be included in the IDE. In such examples, when an end user purchases, leases, or otherwise obtains the IDE from its developer, the end user also receives the instructions for implementing the code embedding circuit 104 and / or the semantic search engine 106. Additionally or alternatively, VCS 112 may implement the code embedding circuit 104 and / or the semantic search engine 106. In such examples, end users accessing VCS 112 (e.g., via user device 114) may utilize the functionality of the code embedding circuit 104 and / or the semantic search engine 106 when editing code maintained by VCS 112.
[0034] exist Figure 1 In the illustrated example, code embedding circuitry 104 is implemented by processor circuitry that preprocesses code snippets for subsequent processing by semantic search engine 106. Code embedding circuitry 104 is coupled to semantic search engine 106 and / or network 108 (e.g., via wired and / or wireless connections). Given a code snippet to be embedded for further processing by an AI / ML model (e.g., implemented by semantic search engine 106), code embedding circuitry 104 samples the usage context of the code snippet within a larger code body. This code snippet can be a function (e.g., the body of a function) or code written within the body of a program. Code embedding circuitry 104 then performs and / or instantiates the embedding process (described further below) within the usage context of the code snippet to generate usage-dependent code embeddings. Code embedding circuitry 104 forwards one or more usage-dependent code embeddings to semantic search engine 106, which performs one or more code intelligence tasks.
[0035] exist Figure 1 In the illustrated example, code embedding circuit 104 samples the use context of a code snippet (e.g., code written in the body of a program or function) and other instances of the code snippet, or the use context of the LOC that calls the function when the code snippet is a function. Code embedding circuit 104 performs and / or instantiates the embedding process on the code snippet, one or more use contexts of the code snippet, and / or one or more use contexts of the LOC that calls the code snippet (when the code snippet is a function) to generate use-dependent code embeddings. Code embedding circuit 104 forwards one or more use-dependent code embeddings to semantic search engine 106. The example use contexts disclosed herein advantageously provide additional information about the code snippet to be processed, including information about function parameters, information about how the function's output is used, and / or general information about the programming context of using and / or calling the function.
[0036] In the examples disclosed in this article, Figure 1 The code embedding circuit 104 executes and / or instantiates one or more machine learning models and / or related circuits to identify the usage context of the code of interest and generate usage-dependent code embeddings. In the examples disclosed herein, one or more machine learning models and / or related circuits are language-dependent. For example, if the input code is written in C++, the code embedding circuit 104 implements one or more machine learning models and / or related circuits trained on the C++ code. In another example, if the input code is written in Python, the code embedding circuit 104 implements one or more machine learning models and / or related circuits trained on the Python code. In some examples, one or more machine learning models are stored in a model data repository prior to execution and / or instantiation (e.g., during deployment and / or training).
[0037] There are many different types of machine learning models and / or machine learning architectures. In the examples disclosed herein, a transformer model is used. Using a transformer model enables the transformation of one vector or matrix into another. However, other AI-based and / or machine learning models / architectures are suitable for use in the example methods disclosed herein to transform an input sequence into an output sequence with the same number of elements as the input sequence. Other types of machine learning models that can be used additionally or alternatively include sequence-to-sequence (seq2seq) models, recurrent neural networks (RNNs), long short-term memory (LSTM) models, gate recurrent unit (GRU) models, and so on.
[0038] Generally, implementing an ML / AI system involves two phases: a learning / training phase and an inference phase. In the learning / training phase, training algorithms are used to train a model to operate based on patterns and / or associations, for example, in the training data. Training data is input data that has been classified or labeled to predict the output data from the machine learning model. Typically, the model includes internal parameters that guide how the input data is transformed into output data, for example, through a series of nodes and connections within the model. Furthermore, hyperparameters are used as part of the training process to control how learning is performed (e.g., learning rate, the number of layers to be used in the machine learning model, etc.). Hyperparameters are defined as training parameters determined before initiating the training process.
[0039] Different types of training can be performed based on the type of ML / AI model and / or its expected output. For example, supervised training uses inputs and corresponding expected (e.g., labeled) outputs to select parameters for an ML / AI model (e.g., by iterating over combinations of selected parameters that reduce model error). As used in this paper, labelling refers to the expected output of a machine learning model (e.g., classification, expected output value, etc.). Alternatively, unsupervised training (e.g., in deep learning, subsets of machine learning, etc.) involves inferring patterns from inputs to select parameters for an ML / AI model (e.g., without the benefit of expected (e.g., labeled) outputs). Another form of training is self-supervised training. In self-supervised training, training data is autonomously labeled (e.g., by leveraging relationships between different input signals). In self-supervised training, the model learns to predict a portion of the input from other parts of the input.
[0040] In the examples disclosed herein, the ML / AI model is trained using self-supervised training. However, any other training algorithms may be used additionally or alternatively. In the examples disclosed herein, training is performed until an acceptable amount of error is achieved. For example, a portion of the training data is used to train the model, and another portion of the training data is used to test the model's error. If the error does not meet (e.g., is above) a threshold, the examples disclosed herein use additional training data to further train and / or adjust the model until the error meets (e.g., is below) the threshold. In the examples disclosed herein, training is performed at service provider 102. However, in additional or alternative examples, training may be performed at user device 114. Training is performed using hyperparameters that control how learning is performed (e.g., learning rate, number of layers to be used in the machine learning model, etc.). In the examples disclosed herein, hyperparameters control the learning rate and regularization. Example hyperparameters controlling regularization control decay, weight fitting, etc. Such hyperparameters are selected, for example, by a grid search method that samples many hyperparameter choices and trains a different model for each hyperparameter choice, then finally uses the best-performing model. In some examples, retraining can be performed. This retraining can be performed in response to the availability of new training data (e.g., taking training data for a language for which one or more machine learning models have not yet been trained).
[0041] Training is performed using training data. In the examples disclosed herein, the training data originates from any source. For example, training data can be obtained from public code repositories (e.g., GitHub, GitLab, other open-source code repositories, etc.), such as code repositories maintained by database 110 and / or VCS112. In some examples, training data can be obtained from code repositories within an organization. Once training is complete, the model is deployed as an executable and / or instantiable construct that processes inputs and provides outputs based on a network of nodes and connections defined in the model. In the examples disclosed herein, the model is stored at service provider 102 for execution and / or instantiation by code embedding circuit 104. In some examples, the model may be stored at service provider 102, where it can be licensed and / or sold to end users. For example, after a sale or license, the model may be transferred to user device 114 and executed and / or instantiated by it to implement code embedding circuit 104.
[0042] Once trained, the deployed model can be manipulated to process data during the inference phase. In the inference phase, the data to be analyzed (e.g., real-world data) is input into the model, and the model is executed and / or instantiated to create output. This inference phase can be thought of as the AI “thinking” to generate output based on what it has learned from training (e.g., applying learned patterns and / or associations to real-world data by executing and / or instantiating the model). In some examples, the input data undergoes preprocessing before being used as input to the machine learning model. Furthermore, in some examples, the output data may undergo post-processing after it has been generated by the AI model to transform the output into a useful result (e.g., a display of data, instructions to be executed and / or instantiated by the machine, etc.). For example, the output of the AI model can be processed by a semantic search engine 106 to perform code intelligence tasks.
[0043] In some examples, the output of the deployed model can be captured and provided as feedback. By analyzing the feedback, the accuracy of the deployed model can be determined. If the feedback indicates that the accuracy of the deployed model is below a threshold or other criterion, the feedback, along with the updated training dataset, hyperparameters, etc., can be used to trigger the training of an updated model (e.g., by generating a new model or adjusting a previously deployed model) to generate the updated deployed model.
[0044] In the examples disclosed herein, code and / or code snippets that include code segments to be processed by an AI / ML model (e.g., a semantic search engine 106) can be commented code, self-documented code, uncommented code, and / or non-self-documented code. Commented code refers to code that includes many comments relative to the number of LOCs. Self-documented code refers to code that includes (a) many functions and / or variables with labels relating to the use and / or meaning of the functions and / or variables, relative to (b) the number of functions and / or variables in the code. Uncommented code refers to code that: (1) does not include comments, (2) includes very few comments relative to the number of LOCs, or (3) includes comments in a manner specific to the code developer and not clearly understood by others (e.g., comments marked as inaccurate and / or misleading). Non-documented code refers to (1) functions and / or variables that do not include labels relating to the use and / or meaning of functions and / or variables, or (2) functions and / or variables that include (a) very few labels relating to the use and / or meaning of functions and / or variables compared to the number of functions and / or variables in (b) code, and is referred to herein as non-documented code.
[0045] exist Figure 1In the illustrated example, the semantic search engine 106 is implemented by processor circuitry that trains other components of the semantic search engine 106, such as one or more Bayesian neural networks (BNNs), to generate searchable representations of the VCS 112, determine the intent of natural language (NL) queries, and / or interpret the code embeddings used in the code snippets included in queries against the semantic search engine 106. In additional or alternative examples, the semantic search engine 106 may implement any other ML / AI model. After performing analysis on the query (e.g., one or more code intelligence tasks), the semantic search engine 106 returns results (e.g., to the end user who submitted the query). Figure 1 In the example, the semantic search engine 106 is coupled to the code embedding circuit 104.
[0046] exist Figure 1 In the example, the code embedding circuit 104 uses information about how and / or how the code snippet is actually used when generating the code embedding. Therefore, as one or more programmers change how the code snippet is used in a larger program, the output generated by the semantic search engine 106 changes. For example, if several usage contexts of the code snippet to be embedded use it in a preferred manner (e.g., correctly), the performance of the semantic search engine 106 (e.g., for code summarization and / or code repair) will improve. Furthermore, in Figure 1 In the example, as the number of usage contexts of code snippets in larger programs and / or codebases increases, the embedding quality of code snippets will improve, resulting in improved performance of downstream code intelligence tasks performed by the semantic search engine 106.
[0047] In some examples, Figure 1 Service provider 102 requests sample use cases for the code snippet before semantic search engine 106 performs code intelligence tasks (e.g., code completion, code clone detection, vulnerability patching, etc.) on the code snippet. For example, if code embedding circuit 104 and / or semantic search engine 106 are not integrated with a code editor (e.g., IDE, VCS 112, etc.) and instead receive input as an explicit code snippet provided by the user, service provider 102 requests sample use cases for the code snippet. In some examples, service provider 102 may request the user to provide access to an entire file or an entire codebase, which includes code snippets (e.g., code within programs and / or functions) to be processed by code embedding circuit 104 and / or semantic search engine 106.
[0048] exist Figure 1In the illustrated example, network 108 is the Internet. However, example network 108 can be implemented using any suitable wired and / or wireless network(s), including, for example, one or more data buses, one or more local area networks (LANs), one or more wireless LANs, one or more cellular networks, one or more private networks, one or more public networks, and so on. In additional or alternative examples, network 108 is an enterprise network (e.g., within an enterprise, company, etc.), a home network, and so on. Example network 108 enables service provider 102 (including code embedding circuit 104 and / or semantic search engine 106), database 110, VCS 112, and user equipment 114 to communicate.
[0049] exist Figure 1In the illustrated example, database 110 stores data related to VCS 112. Database 110 may be implemented using volatile memory (e.g., Synchronous Dynamic Random-Access Memory (SDRAM), Dynamic Random-Access Memory (DRAM), RAMBUS Dynamic Random-Access Memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory). Database 110 may additionally or alternatively be implemented using one or more double data rate (DDR) memories, such as DDR, DDR2, DDR3, DDR4, DDR5, mobile DDR (mDDR), DDR SDRAM, etc. Database 110 may be implemented additionally or alternatively by one or more mass storage devices, such as one or more hard disk drives (HDDs), one or more compact disk (CD) drives, one or more digital versatile disk (DVD) drives, one or more solid-state disk (SSD) drives, one or more Secure Digital (SD) cards, one or more CompactFlash (CF) cards, and so on. Although database 110 is illustrated as a single database in the example shown, database 110 may be implemented by any number and / or type of databases. Furthermore, the data stored in database 110 may take any data format, such as binary data, comma-separated data, tab-separated data, structured query language (SQL) structures, and so on.
[0050] In some examples, database 110 is implemented using a graph database (GDB). When implemented using GDB, database 110 associates the data stored in database 110 with various nodes and edges, where edges represent relationships between nodes. These relationships allow data stored in database 110 to be linked together, so that related data can be retrieved in a single query. In examples where database 110 is implemented using GDB, database 110 may be implemented using one or more Neo4J graph databases. In additional or alternative examples where database 110 is implemented using GDB, database 110 may be implemented using one or more ArangoDB graph databases, one or more OrientDB graph databases, one or more AmazonNeptune graph databases, and so on. In examples where database 110 is implemented using GDB, an appropriate implementation of database 110 will be able to implicitly or explicitly store the probability distribution of source code intent through text (e.g., string) similarity metrics.
[0051] exist Figure 1 In the illustrated examples, VCS112 is implemented by one or more computers and / or one or more memories associated with the VCS platform. In some examples, the components of VCS112 may be distributed (e.g., geographically dispersed). Figure 1 In the example, VCS112 manages changes to computer programs, websites, and / or other sets of information. Users of VCS112 (e.g., developers accessing VCS112 via user device 114) can edit programs and / or other code managed by VCS112. To edit the code, the developer operates on a working copy of the latest version of the code managed by VCS112.
[0052] exist Figure 1 In the illustrated example, when a developer wants to merge their edits with the latest version of the code in VCS112, the developer submits their changes to VCS112. VCS112 then updates the code to the latest version to reflect the working copy of the code in all instances of VCS112. In some examples, VCS112 can roll back commits (e.g., when a developer wants to review a previous version of the program). VCS112 users (e.g., reviewers, other users who did not draft the code, etc.) can apply comments to the code in a commit and / or send messages to the code's drafter to review and / or otherwise improve the code in the commit.
[0053] exist Figure 1In the illustrated example, VCS112 is implemented by one or more computers and / or one or more storage devices associated with a Git platform. In additional or alternative examples, the one or more computers and / or one or more storage devices implementing VCS112 may be associated with another VCS platform, such as AWS CodeCommit, Microsoft TeamFoundation Server, Gerrit Code Review, Subversion, etc.
[0054] As mentioned above, in some examples, Figure 1 The VCS112 can implement the code embedding circuit 104 and / or the semantic search engine 106. In such an example, while the user is editing code using the VCS112, the code embedding circuit 104 and / or the semantic search engine 106 can process the code to provide suggestions to the user. Additionally or alternatively, in such an example, the user can select code snippets (e.g., by highlighting the code snippet, by copying and pasting the code snippet into a search field of the graphical user interface (GUI), by selecting files and / or code libraries that include the code snippet, etc.) for processing by the code embedding circuit 104 and / or the semantic search engine 106.
[0055] exist Figure 1 In the illustrated example, user equipment 114 is implemented by a laptop computer. In additional or alternative examples, user equipment 114 may be implemented by a mobile phone, tablet computer, desktop computer, server, etc., including one or more analog or digital circuits, logic circuits, one or more programmable processors, one or more programmable controllers, one or more GPUs, one or more DSPs, one or more ASICs, one or more programmable logic devices (PLDs), and / or one or more field-programmable logic devices (FPLDs). User equipment 114 may additionally or alternatively be implemented by a CPU, GPU, accelerator, heterogeneous system, etc.
[0056] exist Figure 1In the illustrated example, user device 114 subscribes to, purchases, and / or leases products and / or services from service provider 102 to access one or more machine learning models trained to model VCS, identify intents of NL queries, return code snippets retrieved from the database based on the intents of NL queries, process queries including unannotated and / or non-self-documented code snippets, and return code snippets and / or the intents of related VCS submissions.
[0057] For example, Figure 1 User device 114 accesses one or more trained models through technologies such as downloading one or more models from service provider 102, accessing a web interface hosted by service provider 102 and / or another device. In some examples, user device 114 installs plugins to implement machine learning applications. In such examples, the plugin implements code embedding circuit 104 and / or semantic search engine 106. In some examples, user device 114 can access code embedding circuit 104 and / or semantic search engine 106 via VCS 112 as described above. User device 114 can access VCS 112 via a web interface, via an application, or similar means.
[0058] Figure 2 The block diagram illustrates Figure 1 The example implementation of code embedding circuit 104 is used to generate code embeddings that depend on the usage. Figure 2 In the illustrated example, the code embedding circuit 104 includes an example parsing circuit 202, an example embedding circuit 204, an example transformation circuit 206, an example concatenation circuit 208, an example mask generation circuit 210, and an example classifier circuit 212. Figure 2 The code-embedded circuit 104 can be instantiated (e.g., create an instance of itself, make it exist for any length of time, materialize it, implement it, etc.) by executing instructions from a processor circuit such as a central processing unit. Additionally or alternatively, Figure 2 The code-embedded circuit 104 can be instantiated (e.g., instantiated, materialized, implemented, etc.) by an ASIC or FPGA configured to perform operations corresponding to instructions. It should be understood that... Figure 2 Some or all of the circuits can be instantiated at the same or different times. Some or all of the circuits can be instantiated, for example, in one or more threads that execute concurrently on the hardware and / or serially on the hardware. Furthermore, in some examples, Figure 2 Some or all of the circuitry can be implemented by one or more virtual machines and / or containers that execute on the microprocessor.
[0059] exist Figure 2In the illustrated example, the analytical circuit 202 is coupled to the network 108 and the embedded circuit 204. Figure 2 In some examples, the parsing circuit 202 obtains training data and / or code that will be processed by an AI / ML model (e.g., semantic search engine 106). In some examples, the interface circuitry of the service provider 102 obtains training data and / or code that will be processed by an AI / ML model (e.g., semantic search engine 106).
[0060] exist Figure 2 In the illustrated example, parsing circuit 202 parses the code to select the usage context of the code snippet, including a first number of LOCs before the code snippet or the LOC of the calling code snippet, the code snippet itself, and a second number of LOCs after the code snippet or the LOC of the calling code snippet. Parsing circuit 202 also determines whether the code includes additional instances of the code snippet or additional LOCs of the calling code snippet (e.g., when the code snippet is a function) to be used as the usage context. Figure 2 In the example, in response to determining that the code includes additional instances of the code snippet or additional LOCs of calling the code snippet (e.g., a function), the parsing circuit 202 selects a usage context for the next instance of the code snippet or the next LOC of calling the code snippet (e.g., a function), wherein the usage context includes a first number of LOCs before the next instance of the code snippet or the next LOC of calling the code snippet, the next LOC of the next instance of the code snippet or the next LOC of calling the code snippet, and a second number of LOCs after the next instance of the code snippet or the next LOC of calling the code snippet. After the code embedding circuit 104 completes the embedding process, the example parsing circuit 202 determines whether there is any additional code to be processed by the AI / ML model (e.g., implemented by the semantic search engine 106).
[0061] exist Figure 2 In the illustrated example, when processing code is obtained via network 108, parsing circuit 202 identifies an example first usage context 214 of the code segment to be processed by semantic search engine 106 and an example second usage context 216 of the code segment to be processed by semantic search engine 106. Furthermore, in Figure 2 In the example, the code snippet is a function represented by the example body of function 218. The parsing circuit 202 selects a first usage context 214, a second usage context 216, and a function body 218, and transmits the first usage context 214, the second usage context 216, and the function body 218 to the embedding circuit 204.
[0062] exist Figure 2In the example, the first LOC before the LOC of the code snippet and / or the code snippet that calls it as a function is one (e.g., the first LOC = 1). Figure 2 In the example, the second number of LOCs following the LOC of the code snippet and / or the call to the code snippet as a function is one (e.g., the second number of LOCs = 1). Therefore, an individual usage context includes at least three LOCs: the LOC describing the code snippet to be embedded (which may be one or more LOCs), an LOC preceding the code snippet, and an LOC following the code snippet. However, in alternative examples, the first number of LOCs preceding the LOC of the code snippet and / or the call to the code snippet as a function and / or the second number of LOCs following the LOC of the code snippet and / or the call to the code snippet as a function can be different. For example, the usage context includes at least one LOC preceding the LOC of the code snippet or the call to the code snippet, the code snippet itself, and at least one LOC following the LOC of the code snippet or the call to the code snippet.
[0063] In some examples, the designers of code embedding circuit 104 and / or semantic search engine 106 can choose to use the size of the context (e.g., the number of LOCs before the LOC of the code snippet to be embedded and / or the number of LOCs after ... the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOCs after the LOC of the code snippet to be embedded and the number of LOC In some examples, the first number of LOCs before the LOC of a code snippet and / or the LOC after the LOC of a code snippet and / or the LOC after the LOC of a code snippet and / or the LOC after the LOC of a code snippet as a function can be any integer based on user and / or manufacturer preferences.
[0064] In the examples disclosed herein, the first number of LOCs preceding the LOC of the code snippet and / or the call to the code snippet as a function corresponds to a first threshold number of LOCs. Furthermore, in the examples disclosed herein, the second number of LOCs following the LOC of the code snippet and / or the call to the code snippet as a function corresponds to a second threshold number of LOCs. In some examples, the second threshold number of LOCs differs from the first threshold number of LOCs.
[0065] The following pseudocode 1, pseudocode 2, and pseudocode 3 illustrate example implementations of the first usage context 214, the second usage context 216, and the function body 218, respectively.
[0066] error_arr=(yx)**2
[0067] max_error=find_max(error_arr)
[0068] print('Maximum error is',max_error)
[0069] Pseudocode 1
[0070] result_arr = np.array(result)
[0071] max_elem=find_max(result_arr)
[0072] normalized_result_arr=result_arr / max_elem
[0073] Pseudocode 2
[0074] def find_max(arr):
[0075] l = arr[1]
[0076] for z in arr[1:]:
[0077] if z>l:
[0078] l=z
[0079] return l
[0080] Pseudocode 3
[0081] exist Figure 2 In the illustrated example, embedded circuit 204 is coupled to parsing circuit 202, transformation circuit 206, mask generation circuit 210, and classifier circuit 212. Figure 2In the example, the embedded circuit receives one or more usage contexts and / or code snippets (e.g., function bodies) from the parsing circuit 202. For example, the embedded circuit 204 receives a first usage context 214, a second usage context 216, and a function body 218 from the parsing circuit 202.
[0082] exist Figure 2 In the illustrated example, embedded circuit 204 generates a list of one or more tokens for the selected use context, for the code snippet, and / or, when the code snippet is a function, for the body of the function (e.g., the function body). For example, for a use context, embedded circuit 204 separates the text string selected by parsing circuit 202 into one or more groups of one or more characters. An example list of one or more tokens is further described below. For a use context, embedded circuit 204 also appends a close token (cls) to the list of one or more tokens for that use context. The close token indicates the termination of the use context.
[0083] After generating a list of one or more tokens for the selected use context, code snippet, and / or the body of a function (when the code snippet is a function), and appending a close token (cls) to the list of one or more tokens for the use context, the embedding circuit 204 generates a list of one or more token embedding vectors for the list of one or more tokens. The list of one or more token embedding vectors comprises the token embedding vectors of the tokens in the list of one or more tokens. For example, the embedding circuit 204 initializes an embedding matrix and / or table that maps possible tokens to real-valued vectors, referred to as token embedding vectors. The embedding circuit 204 transforms the list of one or more tokens into a list of one or more token embedding vectors by mapping the tokens to their corresponding token embedding vectors. The token embedding vectors are learnable.
[0084] exist Figure 2 In the illustrated example, the transformation circuit 206 is coupled to the embedded circuit 204, the cascaded circuit 208, the mask generation circuit 210, the classifier circuit 212, and the semantic search engine 106. Figure 2 In the example, the transform circuit 206 implements one or more transformer models to transform the token embedding vector, as further described below. However, in additional or alternative examples, the transform circuit 206 may implement one or more other AI / ML models, such as one or more RNNs, one or more LSTM models, and / or one or more GRU models. Figure 2 In the example, given a list of one or more token embedding vectors, the transformation circuit 206 generates a list of one or more transformed token embedding vectors. Figure 2In the example, the transformed token embedding vector is a real-valued vector. The real-valued output vector of a list of one or more token embedding vectors depends on the other token embedding vectors. After generating one or more lists of transformed token embedding vectors for a list of one or more token embedding vectors, the transformation circuit 206 forwards the list of one or more transformed token embedding vectors to the cascade circuit 208.
[0085] exist Figure 2 In the example, series circuit 208 is coupled to conversion circuit 206. Figure 2 In the example, cascade circuit 208 concatenates a transformed token embedding vector for the use context and a list of one or more transformed token embedding vectors for the code snippet or the body of a function (when the code snippet is a function). Therefore, the transformed token embedding vector for the close token represents a fixed-size embedding for that use context. Cascade circuit 208 also concatenates the close vector (z... sp This is appended to a concatenated list of one or more transformed token embedding vectors. In the example disclosed herein, the closing vector indicates the termination of the concatenated list of one or more transformed token embedding vectors. The concatenating circuit 208 then forwards the concatenated list of one or more transformed token embedding vectors to the transforming circuit 206.
[0086] exist Figure 2 In the example, transformation circuit 206 generates a transformed concatenated list of one or more transformed token embedding vectors. The length of the transformed concatenated list depends on the number of tokens in the function body (e.g., code snippet) and the number of usage contexts. Transformation circuit 206 then transmits at least one of the transformed concatenated list or the transformed closed vector to semantic search engine 106. For example, some downstream tasks, such as code summarization or code repair, anticipate a variable-size embedding list. In this way, semantic search engine 106 can use the transformed concatenated list as usage-dependent embeddings. In some examples, processor circuitry performing other tasks (e.g., clone detection) may anticipate a fixed-size embedding vector. In such an example, the task can use the transformed closed vector (e.g., the last real-valued vector in the transformed concatenated list, e.g., corresponding to z). sp The vector is a fixed-size vector that depends on the embedding used.
[0087] exist Figure 2In the example, semantic search engine 106 performs one or more code intelligence tasks based on usage-dependent embeddings generated by code embedding circuitry 104. For example, when performing a code repair task, semantic search engine 106 determines that the body 218 of a function (e.g., the find_max function) has a functional defect where the first element in the array (e.g., element 0) is not considered. Furthermore, as described above, by incorporating usage context into the embedding process, code embedding circuitry 104 provides additional information to semantic search engine 106. For example, a second usage context 216 of the code snippet indicates that the function (e.g., the body 218 of the function) accepts NumPy arrays (e.g., arrays compatible with Python's NumPy library, which adds support for large multidimensional arrays and matrices, as well as support for high-level mathematical functions that operate on such arrays and matrices) as input. This is valuable information that is typically missing in the function bodies of untyped languages and would otherwise be missed by existing embedding techniques.
[0088] Figure 8 The example code embedding circuit 104 also includes a mask generation circuit 210 and a classifier circuit 212. Figure 9 In the example, the mask generation circuit 210 is coupled to the embedding circuit 204 and the transformation circuit 206. Figure 6 In the example, classifier circuit 212 is coupled to embedding circuit 204 and transformation circuit 206. Code embedding circuit 104 utilizes mask generation circuit 210 and classifier circuit 212 during training.
[0089] During training, Figure 7The mask generation circuit 210 generates a bit mask (e.g., a vector of some combination of zeros and ones, including at least one zero) for a code segment (e.g., the body of a function), and multiplies a list of one or more token embedding vectors of the code segment (e.g., the body of a function) by the bit mask. Therefore, at least one of the values in the list of one or more token embedding vectors of the code segment (e.g., the body of a function) is masked (e.g., set to zero). Furthermore, during training, the classifier circuit 212 classifies the masked tokens after the code embedding process is complete. The embedding circuit 204 compares the classification values of one or more masked tokens with the actual values of one or more masked tokens. For each masked token, the classifier circuit 212 generates a probability distribution (e.g., a vector with positive entries whose sum equals one, and whose size is the number of possible tokens) that covers all possible tokens. The loss function for each masked token is the distance between the predicted probability vector of that masked token and the ideal probability vector, which assigns a probability of -(1) to the correct token and a probability of zero(0) to all other tokens. The total loss function is the sum of the loss functions for all masked tokens. The goal of training is to minimize the total loss to encourage the model to assign the highest probability (e.g., a probability of -(1)) to the correct token and the lowest probability (e.g., a probability of zero(0)) to the incorrect token. This can be achieved by using backpropagation to adjust the parameters of the embedding circuit 204 and / or the transformation circuit 206 to minimize the total loss. Training continues until the total loss becomes below a threshold and / or until the total loss stops decreasing. After training is complete, the AI / ML model implementing the code embedding circuit 104 is deployed.
[0090] Although it has already been Figure 10 The example embedding process is discussed, but in additional or alternative examples, other embedding processes can be implemented to embed function bodies, code snippets, and / or one or more usage contexts.
[0091] In some examples, the code embedding circuit 104 includes means for parsing the code. For example, the means for parsing the code may be implemented by parsing circuit 202. In some examples, parsing circuit 202 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, parsing circuit 202 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6 Blocks 602, 604, 606, and 608 and at least Figure 7Blocks 702, 704, 706, 708, and 726 implement machine-executable instructions. In some examples, the parsing circuit 202 can be instantiated by hardware logic circuitry configured to perform operations corresponding to the machine-readable instructions. Figure 10 The parsing circuit 202 can be implemented using an ASIC or FPGA circuit 1000. Alternatively or additionally, the parsing circuit 202 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the parsing circuit 202 can be implemented by at least one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGA, application-specific integrated circuit (ASIC), comparator, operational amplifier, logic circuitry, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable. In some examples, ASIC refers to an application-specific integrated circuit.
[0092] In some examples, the code embedding circuit 104 includes means for embedding code. For example, the means for embedding code may be implemented by the embedding circuit 204. In some examples, the embedding circuit 204 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, embedded circuit 204 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6 Blocks 610, 612, 614, 630, and 632 and at least Figure 7 Blocks 710, 712, and 714 implement machine-executable instructions. In some examples, embedded circuitry 204 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 10 The embedded circuit 204 can be implemented by an ASIC or FPGA circuit 1000. Alternatively or additionally, the embedded circuit 204 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the embedded circuit 204 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, application-specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally appropriate.
[0093] In some examples, the code embedding circuit 104 includes means for transforming vectors. For example, the means for transforming vectors may be implemented by transformation circuit 206. In some examples, transformation circuit 206 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, conversion circuit 206 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6 Blocks 620, 626, and 634 and at least Figure 7 Blocks 716, 722, and 724 implement machine-executable instructions. In some examples, the transformation circuit 206 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 10 The transformation circuit 206 can be implemented by an ASIC or FPGA circuit 1000. Alternatively or additionally, the transformation circuit 206 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the transformation circuit 206 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, application-specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0094] In some examples, the code embedding circuit 104 includes means for concatenating vectors. For example, the means for concatenating vectors may be implemented by concatenating circuit 208. In some examples, concatenating circuit 208 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, series circuit 208 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6 Blocks 622 and 624 and / or at least Figure 7 Blocks 718 and 720 implement machine-executable instructions. In some examples, cascaded circuitry 208 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 10The cascade circuit 208 can be implemented using an ASIC or FPGA circuit 1000. Alternatively or additionally, the cascade circuit 208 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the cascade circuit 208 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, application-specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0095] In some examples, the code embedding circuit 104 includes means for generating a mask. For example, the means for generating a mask may be implemented by a mask generation circuit 210. In some examples, the mask generation circuit 210 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, mask generation circuit 210 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6 Blocks 616 and 618 implement machine-executable instructions. In some examples, the mask generation circuit 210 can be instantiated by hardware logic circuitry configured to perform operations corresponding to machine-readable instructions. Figure 10 The mask generation circuit 210 is implemented by an ASIC or FPGA circuit 1000. Alternatively or additionally, the mask generation circuit 210 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the mask generation circuit 210 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, application-specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0096] In some examples, the code embedding circuit 104 includes means for classification. For example, the means for classification may be implemented by classifier circuit 212. In some examples, classifier circuit 212 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812. For example, classifier circuit 212 can be used... Figure 9 Example general-purpose microprocessor circuit 900 is instantiated by executing machine-executable instructions, such as by at least Figure 6Block 628 implements the machine-executable instructions. In some examples, the classifier circuit 212 can be instantiated by hardware logic circuitry configured to perform operations corresponding to the machine-readable instructions. Figure 10 The classifier circuit 212 can be implemented using an ASIC or FPGA circuit 1000. Alternatively or additionally, the classifier circuit 212 can be instantiated by any other combination of hardware, software, and / or firmware. For example, the classifier circuit 212 can be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, application-specific integrated circuits (ASICs), comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0097] In some examples, service provider 102 includes means for processing searches. For example, the means for processing searches may be implemented by semantic search engine 106. In some examples, semantic search engine 106 may be instantiated by processor circuitry, for example... Figure 8 Example processor circuit 812 and / or other processor circuits. For example, semantic search engine 212 can be used through, for example... Figure 9 Example general-purpose microprocessor circuits such as the general-purpose microprocessor circuit 900 are instantiated by executing machine-executable instructions. In some examples, the semantic search engine 106 can be instantiated by hardware logic circuitry, which can be implemented by ASIC or FPGA circuitry configured to perform operations corresponding to machine-readable instructions, for example... Figure 10 The FPGA circuit 1000. Additionally or alternatively, the semantic search engine 106 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the semantic search engine 106 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGA, application-specific integrated circuit (ASIC), comparator, operational amplifier, logic circuit, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0098] Figure 3 The flowchart illustrates the training Figure 1 and / or Figure 2 The code is embedded in the example process 300 of circuit 104. Figure 3In the illustrated example, the code embedding circuit 104 implements a K-layer transformer model to generate the top-level embedding of the code snippet. Figure 3 In the example, the code snippet is the function body. However, the code snippet can be a LOC (Level of Context) within the body of the program that does define the function body. The example code embedding circuit 104 implements self-supervised training. Figure 3 The code embedding circuit 104 performs and / or instantiates a masked token prediction task, wherein some (e.g., at least one) tokens in the function body are masked (e.g., set to zero), and the code embedding circuit 104 is trained to predict the masked (e.g., hidden) token(s). Figure 3 In the example, the second token in the function body is masked (e.g., the embedding vector of the second token in the function body is set to zero (0)), and the task of the code embedding circuit 104 is to use the top-level embedding to predict the masked token.
[0099] exist Figure 3 In the illustrated example, parsing circuit 202 obtains code including code snippets to be processed by semantic search engine 106. Therefore, parsing circuit 202 selects a usage context for the code snippet and the body of the function representing the code snippet. For the usage context, parsing circuit 202 selects one or more LOCs before the LOC of the function call, the LOC of the function call, and one or more LOCs after the LOC of the function call. Parsing circuit 202 forwards the selected code to embedding circuit 204.
[0100] exist Figure 3 In the illustrated example, the parsing circuit 202 selects M use contexts in which functions are used and / or called. Use contexts may have variable sizes, but in the example disclosed herein, all use contexts include 2L LOCs and the LOC of the use and / or call of the function. For example, a use context includes L LOCs before the LOC of the use and / or call of the function body, L LOCs after the line containing the use and / or call of the function body, and the LOC of the use and / or call of the function body.
[0101] exist Figure 3 In the illustrated example, embedded circuit 204 generates a first list 302a for one or more tokens for a first usage context, a second list 302b for one or more tokens for a second usage context, a third list 302c for one or more tokens for an Mth usage context, and a list 304 for one or more tokens for the body of the function. In the examples disclosed herein, the list of one or more tokens for the usage context is represented as... Where N iLet be the number of tokens in the i-th usage context, and let i range from 1 to M. In the examples disclosed herein, a list of one or more tokens used for the body of a function is represented as Where N body This is the number of tokens in the function body. Lists 1, 2, 3, and 4 below illustrate examples of a first list 302a for one or more tokens used in a first usage context, a second list 302b for one or more tokens used in a second usage context, a third list 302c for one or more tokens used in an Mth usage context, and a list 304 for one or more tokens used in the function body.
[0102]
[0103] List 1
[0104]
[0105] List 2
[0106]
[0107] List 3
[0108]
[0109] List 4
[0110] exist Figure 3 In the illustrated example, for the selected use context, the embedded circuit 204 appends a shutdown token to a list of one or more tokens for that use context. Figure 3 In the example, the embedding circuit 204 initializes a token embedding matrix 306(R) with a fixed-size, learnable embedding vector of size E for the possible tokens in the source code, where e t is the embedding vector of token t. The embedding circuit 204 arranges these embedding vectors into rows of a token embedding matrix 306(R) of size V×E, where V is the size of the source code token vocabulary.
[0111] exist Figure 3 In the illustrated example, example embedding circuit 204 generates a list of one or more token embedding vectors for a list of one or more tokens. Figure 3 In the example, a list of one or more token embedding vectors may include the token embedding vectors of the tokens in that list. Figure 3 In the example, the embedding circuit 204 converts the input token into a corresponding token embedding vector by looking up the token embedding vector of a given token in the token embedding matrix 306(R). A list of one or more token embedding vectors is represented as...
[0112] exist Figure 3 In the illustrated example, embedding circuit 204 generates a first list 308a of one or more token embedding vectors for a first list 302a of one or more tokens, a second list 308b of one or more token embedding vectors for a second list 302b of one or more tokens, a third list 308c of one or more token embedding vectors for a third list 302c of one or more tokens, and a list 310 of one or more token embedding vectors for a list 304 of one or more tokens for the body of a function. Embedding circuit 204 constructs token embedding vectors from code text (e.g., from the function body and the token sequence in the usage context) without using any intermediate structural representation. For example, the first list 308a of one or more token embedding vectors for the first list 302a of one or more tokens, the second list 308b of one or more token embedding vectors for the second list 302b of one or more tokens, the third list 308c of one or more token embedding vectors for the third list 302c of one or more tokens, and the list 310 of one or more token embedding vectors for the list 304 of one or more tokens for the body of a function are illustrated in Lists 5, 6, 7, and 8 below, respectively.
[0113]
[0114] List 5
[0115]
[0116] List 6
[0117]
[0118] List 7
[0119]
[0120] List 8
[0121] exist Figure 3 In the illustrated example, the embedded circuit 204 is initialized to an N of size E. max N positional embedding vectors, where N max It is the maximum length of the token sequence fed into the code embedding circuit 104. The position embedding vector is represented as... Then, the embedding circuit 204 adds the position embedding vector to the token embedding vector using the context to obtain the sequence. Embedded circuit 204 also adds position embedding vectors to list 310 of one or more token embedding vectors of list 304 of one or more tokens used for the body of the function to obtain a sequence.
[0122] exist Figure 3 In the illustrated example, to train the transform circuit 206, the code embedding circuit 104 uses masked token prediction training in multiple training iterations. In each training iteration, the mask generation circuit 210 generates a mask of size N. body A zero and one bit mask 312 (e.g., a random bit mask). The bit mask 312 is multiplied by a position-adjusted list 310 of one or more token embedding vectors of a list 304 of one or more tokens used in the function body (e.g., ...). Furthermore, the token embedding corresponding to the zero position in the bit mask 312 is set as a zero vector, as shown in the masked list 314 of one or more token embedding vectors in the list 304 of one or more tokens for the function body.
[0123] exist Figure 3 In the illustrated example, the transformation circuit 206 implements one or more transformer layers of the K-layer transformer model. For i from 1 to K, the transformation circuit 206 performs and / or instantiates the following equation 1.
[0124]
[0125] In equation 1, and These are the learnable query matrix, learnable key matrix, and learnable value projection matrix of the attention head h in transformer layer i, respectively. Figure 3 In the example, h ranges from 1 to H, where H is the number of heads of interest. In Equation 1, q i,h ,、k i,h , and v i, The size of h is N×d model , where d model It is the width of the K-layer transformer model executed and / or instantiated by the code embedding circuit 104. In Equation 1, z i-1 This is the output of the previous transformer layer (e.g., the (i-1)th layer). Then, the transformer circuit 206 executes and / or instantiates the softmax function row by row, as shown in Equations 2 and 3 below.
[0126]
[0127] Equation 2
[0128] u i =concat(ui,0 ,...,u i,H )
[0129] Equation 3
[0130] exist Figure 3 In the example, the softmax function is the function that normalizes the input values to a probability distribution. The concatenation function is the function that concatenates the input values. Figure 3 In the example, the output of Equation 3 has a size NH×d model Then, the transformation circuit 206 generates one or more transformed token embedding vectors z from a list of one or more token embedding vectors according to Equation 4 below. i A list.
[0131]
[0132] Equation 4
[0133] In Equation 4, the list of one or more transformed token embedding vectors has a size of N×d. model .exist Figure 3 In the example, for the context, the list of one or more transformed token embedding vectors is composed of sequences It indicates. In Figure 3 In the example, the list of one or more transformed token embedding vectors used for the function body is composed of a sequence express.
[0134] exist Figure 3 In the illustrated example, the cascade circuit 208 is cascaded with a transformed token embedding vector (e.g., for using a context-sensitive close token) for the close token. ) and a list of one or more transformed token embedding vectors for the function body (e.g., For example, the serial circuit 208 will embed the first transformed token embedding vector 316a of the close token used in the first usage context. The second transformed token embedding vector 316b for the close token used in the second usage context And the third transformed token embedding vector 316c used for the close token of the Mth usage context Prepended to a list of one or more transformed token embedding vectors used for the function body (e.g., ), to generate a concatenated list 318 of one or more transformed token embedding vectors. In this way, the transformed token embedding vector used for the context-based closing token (e.g., This indicates the usage context. In some examples, the cascaded circuit 208 embeds the transformed token of the turn-off token into a vector (e.g., ...) in any order. ) is placed before the list of one or more transformed token embedding vectors used for the function body.
[0135] exist Figure 3 In the illustrated example, the series circuit 208 will also turn off the vector (e sp ) appended to one or more concatenated lists of transformed token embedding vectors 318 (e.g., The transformation circuit 206 then processes one or more concatenated lists 318 of transformed token embedding vectors to generate one or more transformed concatenated lists 322 of transformed token embedding vectors (e.g., For example, converter circuit 206 implements the second converter layer of a K-layer converter model. In the examples disclosed herein, the first and second converter layers may include the same or different parameters. Figure 3 In the example, classifier circuit 212 (e.g., performs and / or instantiates multilayer perception) processes the top-level embedding of the masked token (e.g., q2 324) and predicts the identity of the masked token.
[0136] exist Figure 3 In the illustrated example, code embedding circuit 104 implements a classification loss (e.g., cross-entropy loss) to train classifier circuit 212, and more generally, code embedding circuit 104. In response to a classification loss greater than a threshold, code embedding circuit 104 backpropagates the error through a K-layer transformer model to update one or more weights of the model. In response to a classification loss less than or equal to the threshold at the end of training, a K-layer transformer model is deployed to be executed and / or instantiated by code embedding circuit 104. After training, code embedding circuit 104 may not utilize mask generation circuit 210 or classifier circuit 212. However, in additional or alternative examples, code embedding circuit 104 may utilize mask generation circuit 210 and / or classifier circuit 212 after training.
[0137] Figure 4 The flowchart illustrates an example process 400 that generates code that depends on the embedded code used. Figure 4 In the illustrated example, the code embedding circuit 104 implements a K-layer transformer model to generate the top-level embedding of the code snippet. Figure 4 In the example, parsing circuit 202 obtains code including code snippets to be processed by semantic search engine 106. Figure 4 In the example, the code snippet is the function body. Therefore, the parsing circuit 202 selects a usage context for the code snippet, including the LOC surrounding the calling function, the calling function's LOC, and the function body. The parsing circuit 202 forwards the selected code to the embedding circuit 204.
[0138] existFigure 4 In the illustrated example, the analytical circuit 202 selects M usage contexts in which functions are used and / or called. Figure 4 In the example, the embedded circuit 204 generates a first list 402a for one or more tokens for a first use context, a second list 402b for one or more tokens for a second use context, a third list 402c for one or more tokens for an Mth use context, and a list 404 for one or more tokens for the body of the function.
[0139] exist Figure 4 In the illustrated example, for the selected use context, the embedded circuit 204 appends a shutdown token to a list of one or more tokens for that use context. Figure 4 In the example, the embedding circuit 204 initializes a token embedding matrix 406(R) with a fixed-size, learnable embedding vector of size E for the possible tokens in the source code, where e t These are the embedding vectors of token t. The embedding circuit 204 arranges these embedding vectors into rows of a token embedding matrix 406(R) of size V×E, where V is the size of the source code token vocabulary.
[0140] exist Figure 4 In the illustrated example, example embedding circuit 204 generates a list of one or more token embedding vectors for a list of one or more tokens. Figure 4 In the example, the list of one or more token embedding vectors includes the token embedding vectors of the tokens in that list. Figure 4 In the example, the embedding circuit 204 converts the input token into the corresponding token embedding vector by looking up the token embedding vector of a token in the token embedding matrix 406(R).
[0141] exist Figure 4 In the illustrated example, the embedding circuit 204 generates a first list 408a of one or more token embedding vectors for a first list 402a of one or more tokens, a second list 408b of one or more token embedding vectors for a second list 402b of one or more tokens, a third list 408c of one or more token embedding vectors for a third list 402c of one or more tokens, and generates a list 410 of one or more token embedding vectors for a list 404 of one or more tokens for the body of the function. Figure 4 In the example, the embedding circuit 204 constructs a token embedding vector from the code text (e.g., from the function body and the token sequence in the usage context) without using any intermediate structural representation.
[0142] exist Figure 4 In the illustrated example, the embedded circuit 204 initializes the size E of N. maxN positional embedding vectors, where N max It is the maximum length of the token sequence fed into the code embedding circuit 104. The position embedding vector is represented as... Then, the embedding circuit 204 adds the position embedding vector to the token embedding vector using the context to obtain the sequence. Embedded circuit 204 also adds position embedding vectors to list 410 of one or more token embedding vectors of list 404 of one or more tokens used for the body of the function to obtain a sequence.
[0143] exist Figure 4 In the illustrated example, the converter circuit 206 implements one or more converter layers of the K-layer converter model. Figure 4 In the example, the transformation circuit 206 executes and / or instantiates Equations 1, 2, 3, and 4 described above and below. Figure 4 In the example, for the context, the list of one or more transformed token embedding vectors is composed of sequences It indicates. In Figure 4 In the example, the list of one or more transformed token embedding vectors used for the function body is composed of a sequence Therefore, the transformation circuit 206 processes a sequence of N token embedding vectors and generates a sequence of N transformed token embedding vectors, where the size of the transformed token embedding vectors is d. model ,(For example,
[0144] exist Figure 4 In the illustrated example, a converter function refers to one or more converter layers executed by converter circuit 206. Figure 4 In the example, the serial circuit 208 is serialized with the transformed token embedding vector (e.g., for individual use context) of the close token. ) and a list of one or more transformed token embedding vectors for the function body (e.g., For example, the serial circuit 208 will embed the first transformed token embedding vector 412a of the close token used in the first usage context. The second transformed token embedding vector 412b for the close token used in the second usage context And the third transformed token embedding vector 412c for the close token of the Mth use context A list of one or more transformed token embedding vectors placed in the body of the function (e.g., (414) preceding this, to generate a concatenated list of one or more transformed token embedding vectors. In this way, the transformed token embedding vector used for the individual's use context of the closing token (e.g., This indicates the usage context. In some examples, the cascaded circuit 208 embeds the transformed token of the turn-off token into a vector (e.g., ...) in any order. ) is placed before the list of one or more transformed token embedding vectors used for the function body.
[0145] exist Figure 5 In the illustrated example, the cascaded circuit 208 also embeds one or more transformed tokens into a cascaded list 414 of a vector list (e.g., )Additional closing vector (e sp The transformation circuit 206 then processes one or more concatenated lists 414 of transformed token embedding vectors to generate one or more transformed concatenated lists 418 of transformed token embedding vectors (e.g., For example, the converter circuit 206 implements the second converter layer of the K-layer converter model.
[0146] exist Figure 5 In the illustrated example, one or more transformed token embedding vectors are transformed concatenated list 418 (e.g., This includes the transformed closed vector 420(q) sp In the example disclosed in this paper, the transformed closed vector 420(q) sp The transformed concatenation list 418 of one or more transformed token embedding vectors represents the final dependency of the function body and / or code snippet on the used embedding in a code intelligence task (e.g., clone detection) that utilizes a fixed-size embedding vector. In additional or alternative examples, for code intelligence tasks (e.g., code summarization and / or code repair) that utilize top-level embeddings that include information about tokens in a code snippet or function body, the transformed concatenation list 418 represents the final dependency of the function body and / or code snippet on the used embedding.
[0147] Figure 5 The flowchart illustrates the example process 500 for generating code embeddings that depend on the example function that is used twice in the example code. Figure 5 The example code embedding circuit 104 implements a K-layer transformer model to generate top-level embeddings of code snippets. Figure 5 In the example, the parsing circuit 202 obtains code including code snippets to be processed by the semantic search engine 106. Figure 5 The example parsing circuit 202 selects the usage context of the code snippet and the body of the function representing the code snippet. The parsing circuit 202 forwards the selected code to the embedding circuit 204.
[0148] exist Figure 5In the illustrated example, the analytical circuit 202 selects two (2) usage contexts for using and / or calling functions. Figure 5 In the example, embedded circuit 204 generates a first list 502a of one or more tokens for a first usage context, a second list 502b of one or more tokens for a second usage context, and a list 504 of one or more tokens for the body of the function. Figure 5 In the example, for the first use context, the embedded circuit 204 appends the first closing token 506a(cls) to a first list 502a of one or more tokens. For the second use context, Figure 5 The example embedded circuit 204 appends a second closing token 506b (cls) to a second list 502b of one or more tokens.
[0149] For example, for the first usage context, the second usage context, and the body of the function, the embedded circuit 204 generates the following lists 9, 10, and 11, respectively.
[0150] [error_arr,=,(,y,...,max_error,),cls]
[0151] List 9
[0152] [result_arr,=,np,...,max_elem,cls]
[0153] List 10
[0154] [def,find_max,(,arr,),:,l,=,...return,l]
[0155] List 11
[0156] In the examples disclosed in this article, there are different ways to tokenize the same text. For example, in Figure 5 In the example, embedded circuit 204 treats result_arr as a single token. In another example, different tokenization techniques can cause embedded circuit 204 to generate three tokens: "result", "_", and "arr". The designer of embedded circuit 104 can specify the appropriate tokenization technique.
[0157] exist Figure 5 In the illustrated example, the embedding circuit 204 generates a first list 508a of one or more token embedding vectors for a first list 502a of one or more tokens. Furthermore, Figure 5The example embedding circuit 204 generates a second list 508b of one or more token embedding vectors for a second list 502b of one or more tokens. The example embedding circuit 204 also generates a list 510 of one or more token embedding vectors for a list 504 of one or more tokens used in the body of the function.
[0158] exist Figure 5 In the illustrated example, the converter circuit 206 implements one or more converter layers of the K-layer converter model. Figure 5 In the example, the transformation circuit 206 executes and / or instantiates Equations 1, 2, 3, and 4 described above and below. Figure 2 In the example, cascade circuit 208 concatenates a transformed token embedding vector for a close token used in a context and a list of one or more transformed token embedding vectors for the function body. For example, cascade circuit 208 concatenates a first transformed token embedding vector 512a(x) for a close token used in a first context. cls ) and the second transformed token embedding vector 512b(y) for the second use context of the close token cls The transformed token embedding vectors are placed before a list 510 of one or more tokens in a list 504 for the body of the function to produce a concatenated list 514 of one or more transformed token embedding vectors. In this way, the transformed token embedding vector of the closing token used for the usage context represents that usage context. In some examples, the concatenated circuit 208 places the transformed token embedding vector of the closing token before the list of one or more transformed token embedding vectors used for the function body in any order.
[0159] exist Figure 1 In the illustrated example, the cascade circuit 208 also appends a shut-off vector 516 (z) to a cascade list 514 of one or more transformed token embedding vectors. sp The transform circuit 206 then processes one or more concatenated lists 514 of transformed token embedding vectors to generate one or more transformed concatenated lists 518 of transformed token embedding vectors. For example, the transform circuit 206 implements the second transformer layer of the K-layer transformer model.
[0160] exist Figure 2 In the illustrated example, the transformed concatenation list 518 of one or more transformed token embedding vectors includes a transformed closing vector 520 (e.g., In the example disclosed in this paper, the transformed closed vector is 520. This indicates that in code intelligence tasks (e.g., clone detection) utilizing fixed-size embedding vectors, the final dependency of function bodies and / or code snippets on the used embeddings is represented. In additional or alternative examples, for code intelligence tasks (e.g., code summarization and / or code repair) utilizing top-level embeddings that include information about tokens in code snippets or function bodies, a transformed concatenated list 518 of one or more transformed token embedding vectors represents the final dependency of function bodies and / or code snippets on the used embeddings.
[0161] Although Figure 1 The diagram illustrates the implementation. Figure 2 The code is embedded in the example of circuit 104, but Figure 1 One or more of the elements, processes, and / or devices shown may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, example parsing circuit 202, example embedding circuit 204, example transformation circuit 206, example cascading circuit 208, example mask generation circuit 210, example classifier circuit 212, and / or more generally, Figure 2 and / or Figure 1 The example code embedding circuit 104 and / or semantic search engine 106 can be implemented in hardware alone, or in combination with hardware and software and / or firmware. Thus, for example, the example parsing circuit 202, the example embedding circuit 204, the example transformation circuit 206, the example cascading circuit 208, the example mask generation circuit 210, the example classifier circuit 212, and / or more generally... Figure 2 and / or Figure 2 The example code embedded in either circuit 104 or semantic search engine 106 can be implemented by processor circuitry, analog circuitry, digital circuitry, logic circuitry, programmable processors, programmable microcontrollers, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field-programmable logic devices (FPLDs) (e.g., field-programmable gate arrays (FPGAs)). Furthermore, Figure 6 The example code embedded circuit 104 may include, in addition to Figure 7 Those other than or replacing those shown Figure 1 One or more of the elements, processes and / or devices shown, and / or may include any or all of more than one of the elements, processes and devices shown.
[0162] exist Figure 2 and Figure 8 The diagram shows a representative used for implementation.Figure 9 and / or Figure 10 The code embedding circuit 104 is a flowchart of example hardware logic circuitry, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof. Machine-readable instructions may be one or more executable and / or instantiable programs, or portions thereof, for execution and / or instantiation by processor circuitry, such as those described below. Figure 6 The processor circuit 812 shown in the example processor platform 800 discussed below and / or the contact information below. Figure 7 and / or Figure 6 The example processor circuitry is described. The program may be embodied in software stored on one or more non-transitory computer-readable storage media associated with the processor circuitry located in one or more hardware devices, such as compact disks (CDs), floppy disks, hard disk drives (HDDs), solid-state drives (SSDs), digital versatile disks (DVDs), Blu-ray discs, volatile memory (e.g., any type of random-access memory (RAM), etc.) or non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, HDDs, SSDs, etc.), but the entire program and / or a portion thereof may alternatively be executed and / or instantiated and / or embodied in firmware or dedicated hardware by one or more hardware devices other than the processor circuitry. Machine-readable instructions may be distributed across multiple hardware devices and / or executed and / or instantiated by two or more hardware devices (e.g., server and client hardware devices). For example, client hardware devices can be implemented by endpoint client hardware devices (e.g., hardware devices associated with a user) or intermediate client hardware devices (e.g., radio access network (RAN) gateways that facilitate communication between servers and endpoint client hardware devices). Similarly, non-transitory computer-readable storage media can include one or more media located in one or more hardware devices. Additionally, although referenced... Figure 7 and / or Figure 6The flowchart shown illustrates the example program, but many other methods of implementing the example code embedded circuit 104 may be used alternatively. For example, the execution order of the blocks may be changed, and / or some of the blocks described may be altered, eliminated, or combined. Additionally or alternatively, any or all blocks may be implemented by one or more hardware circuits (e.g., processor circuitry, discrete and / or integrated analog and / or digital circuitry, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuitry, etc.) configured to perform the corresponding operations without executing software or firmware. The processor circuitry may be distributed across different network locations and / or local to: one or more hardware devices in a single machine (e.g., a single-core processor (e.g., a single-core central processing unit (CPU)), a multi-core processor (e.g., a multi-core CPU), etc.), multiple processors distributed across multiple servers in a server rack, multiple processors distributed across one or more server racks, CPUs and / or FPGAs located in the same package (e.g., the same integrated circuit (IC) package or two or more separate housings, etc.).
[0163] Machine-readable instructions described herein may be stored in one or more of the following formats: compressed format, encrypted format, segmented format, compiled format, executable and / or instantiable format, packaged format, etc. Machine-readable instructions as described herein may be stored as data or data structures (e.g., as parts of instructions, code, representations of code, etc.) that can be used to create, manufacture, and / or produce machine-executable and / or instantiable instructions. For example, machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located in the same or different locations within a network or set of networks (e.g., in the cloud, at an edge device, etc.). Machine-readable instructions may require installation, modification, adaptation, updating, combination, supplementation, configuration, decryption, decompression, unpacking, distribution, reassignment, compilation, etc., to make them directly readable, interpretable, and / or executable and / or instantiable by computing devices and / or other machines. For example, machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and / or stored on separate computing devices, wherein when these parts are decrypted, decompressed, and / or combined, they form a set of machine-executable instructions that together form one or more operations of a program such as that described herein.
[0164] In another example, machine-readable instructions may be stored in a state in which they can be read by processor circuitry, but require the addition of libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc., to execute and / or instantiate these machine-readable instructions on a specific computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., storage settings, input data, recording network addresses, etc.) before they can be fully or partially executed and / or instantiated. Thus, the machine-readable medium used herein may include machine-readable instructions and / or / one or more programs, regardless of the specific format or state of these machine-readable instructions and / or / one or more programs at the time of storage or otherwise at rest or in transit.
[0165] The machine-readable instructions described in this article can be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, machine-readable instructions can be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0166] As described above, executable and / or instantiable instructions (e.g., computer and / or machine-readable instructions) stored on one or more non-transitory computer and / or machine-readable media can be used to implement... Figure 1 and / or Figure 2 Example operations, where the medium is, for example, an optical storage device, a magnetic storage device, an HDD, flash memory, read-only memory (ROM), a CD, a DVD, a cache, any type of RAM, a register, and / or any other storage device or disk in which information may be stored for any duration (e.g., long time period storage, permanent storage, for short instances, for temporary buffering, and / or for caching information). As used herein, the terms non-transitory computer-readable medium and non-transitory computer-readable storage medium are explicitly defined to include any type of computer-readable storage device and / or disk, excluding propagated signals and transmission media.
[0167] "Comprising" and "including" (and all their forms and tenses) are used herein as introductory terms. Thus, whenever a claim uses any form of "comprising" or "including" (e.g., including, comprising, having, etc.) as a preamble or in any kind of claim recitation, it is understood that additional elements, terms, etc., may be present and not fall outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, the preamble of a claim, it is introductory in the same way that the terms "comprising" and "including" are introductory. The term "and / or" when used, for example, in the form of, say, A, B, and / or C, refers to any combination or subset of A, B, and C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. As used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A or B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein in the context of describing the execution or operation of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A and B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, for the purposes of this document in the context of describing the execution or operation of a process, instruction, action, activity and / or step, the phrase “at least one of A or B” is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.
[0168] For the purposes of this document, singular references (e.g., “a,” “an,” “first,” “second,” etc.) do not exclude pluralism. As used herein, the term “a” or “an” refers to one or more of that object. The terms “a” (or “an”), “one or more,” and “at least one” are used interchangeably herein. Furthermore, although listed separately, multiple means, elements, or methodological actions may be implemented by, for example, the same entity or object. Moreover, while individual features may be included in different examples or claims, they may be combined, and inclusion in different examples or claims does not imply that the combination of features is infeasible and / or not advantageous.
[0169] Figure 6 The flowchart represents what can be implemented by example processor circuitry and / or instantiated. Figure 6 and / or Figure 6 The code is embedded in the circuit 104 to execute the trained example machine-readable instructions and / or example operations 600. Figure 6 The machine-readable instructions and / or operations 600 begin at block 602, where parsing circuitry 202 obtains training data comprising one or more code snippets. For example, in block 602, parsing circuitry 202 obtains training data comprising one or more code snippets from VCS 112 via network 108. In some examples, interface circuitry of service provider 102 obtains training data comprising one or more code snippets.
[0170] exist Figure 6 In the illustrated example, at block 604, parsing circuitry 202 selects a usage context for the code snippet, wherein the usage context includes a first number of LOCs preceding the code snippet or the LOC that calls the code snippet, the code snippet itself, and a second number of LOCs following the LOC that calls the code snippet. At block 606, parsing circuitry 202 determines whether the code (e.g., training data) includes an additional instance of the code snippet or an additional LOC that calls the code snippet (e.g., if the code snippet is a function body). In response to parsing circuitry 202 determining that the code includes an additional instance of the code snippet or an additional LOC that calls the code snippet (block 606: "Yes"), machine-readable instructions and / or operations 600 proceed to block 608. At block 608, parsing circuitry 202 selects a usage context for the next instance of the code snippet or the next LOC that calls the code snippet (e.g., a function). For example, the use context selected by the parsing circuit 202 at block 608 includes a first number of LOCs before the next instance of the code segment or the next LOC of the calling code segment (e.g., a function), the next LOC of the next instance of the code segment or the next LOC of the calling code segment (e.g., a function), and a second number of LOCs after the next instance of the code segment or the next LOC of the calling code segment (e.g., a function).
[0171] Returning to block 606, in response to parsing circuitry 202 determining that the code does not include additional instances of the code snippet or additional LOCs calling the code snippet (Block 606: "No"), machine-readable instructions and / or operations 600 proceed to block 610. In block 610, embedding circuitry 204 generates a list of one or more tokens for the selected use context and code snippet. In block 612, for the use context, embedding circuitry 204 appends a close token (cls) indicating the termination of that use context to the list of one or more tokens for that use context.
[0172] exist Figure 6 In the illustrated example, at block 614, embedding circuit 204 generates a list of one or more token embedding vectors for a list of one or more tokens, wherein the list of one or more token embedding vectors includes the token embedding vectors of the tokens in the list of one or more tokens. At block 616, mask generation circuit 210 generates a bitmask for the code segment. For example, the bitmask includes a random combination of one and zero, including at least one zero. At block 618, mask generation circuit 210 multiplies the list of one or more token embedding vectors of the code segment by the bitmask.
[0173] exist Figure 6 In the illustrated example, at block 620, transformation circuit 206 generates one or more transformed token embedding vectors for a list of one or more token embedding vectors. At block 622, concatenation circuit 208 concatenates the transformed token embedding vector for a context-based closing token and the list of one or more transformed token embedding vectors for a code snippet. At block 624, concatenation circuit 208 appends a closing vector to the concatenated list of one or more transformed token embedding vectors, the closing vector indicating the termination of the concatenated list of one or more transformed token embedding vectors.
[0174] exist Figure 7 In the illustrated example, at block 626, transform circuit 206 generates one or more transformed concatenated lists of transformed token embedding vectors and closing vectors. At block 628, classifier circuit 212 classifies one or more masked token embedding vectors of the code snippet. Figure 1 In the example, at block 630, embedding circuit 204 determines whether the classification meets the training termination criterion. Classifier circuit 212 generates a probability distribution over all possible tokens (e.g., a vector with positive entries whose sum equals one, and whose size is the number of possible tokens). The loss function for each masked token is the distance between the predicted probability vector of that masked token and the ideal probability vector, which assigns a probability of one (1) to the correct token and a probability of zero (0) to all other tokens. The total loss function is the sum of the loss functions for all masked tokens. The goal of training is to minimize this total loss to encourage the model to assign the highest probability (e.g., a probability of one (1)) to the correct token and the lowest probability (e.g., a probability of zero (0)) to the incorrect token. Training can be achieved by using backpropagation techniques to adjust the parameters in embedding circuit 204 and / or transformation circuit 206 to minimize the total loss. Training continues until the total loss becomes below a threshold and / or until the total loss stops decreasing. In response to the embedded circuit 204 determining that the classification does not meet the training termination criteria (block 630: "No"), machine-readable instructions and / or operations 600 proceed to block 632.
[0175] exist Figure 2 In the illustrated example, at block 632, the embedded circuit 204 adjusts its parameters to meet the training termination criterion. At block 634, the transformation circuit 206 adjusts its parameters to meet the training termination criterion. In some examples, the central training control circuit adjusts the parameters of the embedded circuit 204 and / or the transformation circuit 206 to meet the training termination criterion. Returning to block 630, in response to the embedded circuit 204 determining that the classification meets the training termination criterion (block 630: "Yes"), machine-readable instructions and / or operations 600 terminate, and the code embedded circuit 104 is ready for deployment.
[0176] Figure 7 The flowchart represents what can be implemented by example processor circuitry and / or instantiated. Figure 7 and / or Figure 7 The code embedding circuit 104 generates example machine-readable instructions and / or example operations that depend on the code embedding used. Figure 7 The machine-readable instructions and / or operations 700 begin at block 702, where parsing circuitry 202 obtains code including code snippets to be processed by the artificial intelligence model. For example, in block 702, parsing circuitry 202 obtains code including code snippets to be processed by the artificial intelligence model from VCS 112 via network 108. In some examples, interface circuitry of service provider 102 obtains code including code snippets to be processed by the artificial intelligence model.
[0177] exist Figure 7 In the illustrated example, at block 704, the parsing circuit 202 selects the usage context of the code snippet, wherein the usage context includes a first number of LOCs before the code snippet or the LOC that calls the code snippet, the code snippet itself, and a second number of LOCs after the code snippet or the LOC that calls the code snippet. Figure 7 In the example, at block 706, parsing circuit 202 determines whether the code includes an additional instance of the code snippet or an additional LOC that calls the code snippet (e.g., in the case that the code snippet is a function body). Figure 7In the example, in response to parsing circuitry 202 determining that the code includes an additional instance of the code segment or an additional LOC of the calling code segment (block 706: "Yes"), machine-readable instructions and / or operations 700 proceed to block 708. At block 708, parsing circuitry 202 selects a use context for the next instance of the code segment or the next LOC of the calling code segment (e.g., a function). For example, the use context selected by parsing circuitry 202 at block 708 includes a first number of LOCs preceding the next instance of the code segment or the next LOC of the calling code segment (e.g., a function), the next LOC of the next instance of the code segment or the next LOC of the calling code segment (e.g., a function), and a second number of LOCs following the next instance of the code segment or the next LOC of the calling code segment (e.g., a function).
[0178] Returning to block 706, in response to parsing circuitry 202 determining that the code does not include additional instances of the code snippet or additional LOCs calling the code snippet (Block 706: "No"), machine-readable instructions and / or operations 700 proceed to block 710. In block 710, embedding circuitry 204 generates a list of one or more tokens for the selected use context and code snippet. In block 712, for an individual use context, embedding circuitry 204 appends a close token (cls) indicating the termination of that use context to the list of one or more tokens for that use context.
[0179] exist Figure 8 In the illustrated example, at block 714, embedding circuit 204 generates a list of one or more token embedding vectors for an individual list of one or more tokens, wherein the list of one or more token embedding vectors includes the token embedding vectors of the individual tokens in the list of one or more tokens. At block 716, transformation circuit 206 generates a list of one or more transformed token embedding vectors for the individual list of one or more token embedding vectors. At block 718, concatenation circuit 208 concatenates the transformed token embedding vectors for a close token used in an individual usage context and the list of one or more transformed token embedding vectors for a code snippet. At block 720, concatenation circuit 208 appends a close vector to the concatenated list of one or more transformed token embedding vectors, the close vector indicating the termination of the concatenated list of one or more transformed token embedding vectors.
[0180] exist Figure 6 In the illustrated example, at block 722, the transformation circuit 206 generates a transformed concatenated list of one or more transformed token embedding vectors and closing vectors. At block 724, the transformation circuit 206 transmits at least one of the transformed concatenated list of one or more transformed token embedding vectors or transformed closing vectors to the semantic search engine 106. Figure 7In the example, at block 726, parsing circuit 202 determines whether there is additional code to be processed by the AI model. In response to parsing circuit 202 determining that there is additional code to be processed by the AI model (block 726: "Yes"), machine-readable instructions and / or operations 700 return to block 702. In response to parsing circuit 202 determining that there is no additional code to be processed by the AI model (block 726: "No"), machine-readable instructions and / or operations 700 terminate.
[0181] Figure 1 This is a block diagram of an example processor platform 800, which is constructed to perform and / or instantiate... Figure 2 and / or Figure 1 Machine-readable instructions and / or operations to achieve Figure 6 and / or Figure 7 The code is embedded in circuit 104. Processor platform 800 can be, for example, a server, personal computer, workstation, self-learning machine (e.g., neural network), mobile device (e.g., cellular phone, smartphone, such as iPad). TM Tablet devices, personal digital assistants (PDAs), internet-connected appliances, DVD players, CD players, digital video recorders, Blu-ray players, game consoles, set-top boxes, headphones (e.g., augmented reality (AR) headphones, virtual reality (VR) headphones, etc.) or other wearable devices, or any other type of computing device.
[0182] The illustrated processor platform 800 includes processor circuitry 812. The illustrated processor circuitry 812 is hardware. For example, processor circuitry 812 may be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. Processor circuitry 812 may be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, processor circuitry 812 implements... Figure 9 The illustrated examples include the parsing circuit 202, the embedding circuit 204, the transformation circuit 206, the cascading circuit 208, the mask generation circuit 210, the classifier circuit 212, and / or more generally, the code embedding circuit 104.
[0183] The illustrated processor circuitry 812 includes local memory 813 (e.g., cache, registers, etc.). The illustrated processor circuitry 812 communicates via bus 818 with main memory, which includes volatile memory 814 and non-volatile memory 816. The volatile memory 814 may be Synchronous Dynamic Random-Access Memory (SDRAM), Dynamic Random-Access Memory (DRAM), etc. Dynamic Random Access Memory (DRAM) Dynamic Random-AccessMemory, The main memory 814, 816 can be implemented using flash memory and / or any other type of RAM device. Access to the illustrated main memory 814, 816 is controlled by the memory controller 817.
[0184] The processor platform 800 illustrated also includes interface circuitry 820. Interface circuitry 820 can be implemented in hardware according to any type of interface standard, such as an Ethernet interface, a universal serial bus (USB) interface, etc. Interfaces include near field communication (NFC) interfaces, peripheral component interconnect (PCI) interfaces, and / or peripheral component interconnect express (PCIe) interfaces.
[0185] In the illustrated example, one or more input devices 822 are connected to interface circuitry 820. The input devices 822 allow a user to input data and / or commands into processor circuitry 812. The input devices 822 may be implemented as, for example, audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, touchpads, trackballs, isopoint devices, and / or voice recognition systems.
[0186] One or more output devices 824 are also connected to the interface circuitry 820 illustrated in the figure. The output devices 824 may be implemented, for example, by display devices (e.g., light-emitting diodes (LEDs), organic light-emitting diodes (OLEDs), liquid crystal displays (LCDs), cathode ray tube (CRT) displays, in-place switching (IPS) displays, touchscreens, etc.), haptic output devices, printers, and / or speakers. The interface circuitry 820 illustrated in the figure thus typically includes a graphics driver card, a graphics driver chip, and / or graphics processor circuitry, such as a GPU.
[0187] The interface circuit 820 illustrated also includes communication devices, such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces, to facilitate data exchange with external machines (e.g., any kind of computing device) via network 826. Communication can be made via, for example, Ethernet connections, digital subscriber line (DSL) connections, telephone line connections, coaxial cable systems, satellite systems, line-to-line wireless systems, cellular telephone systems, optical connections, and so on.
[0188] The processor platform 800 illustrated also includes one or more mass storage devices 828 for storing software and / or data. Examples of such mass storage devices 828 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disc drives, redundant array of independent disks (RAID) systems, solid-state storage devices (such as flash memory devices and / or SSDs), and DVD drives.
[0189] can be Figure 8 and / or Figure 8 The machine-executable instructions 832 implemented by the machine-readable instructions can be stored in a mass storage device 828, a volatile memory 814, a non-volatile memory 816, and / or on a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0190] Figure 6 yes Figure 7 A block diagram of an example implementation of the processor circuit 812. In this example, Figure 2 The processor circuit 812 is implemented by the general-purpose microprocessor circuit 900. The general-purpose microprocessor circuit 900 executes...Figure 2 and / or Figure 6 The flowchart contains part or all of the machine-readable instructions to effectively translate Figure 7 The circuit is instantiated as a logic circuit to perform operations corresponding to these machine-readable instructions (e.g., corresponding to instructions). In some such examples, Figure 8 The circuit is instantiated by the hardware circuitry of the microprocessor circuitry 900 in conjunction with instructions. For example, the microprocessor circuitry 900 can implement multi-core hardware circuitry, such as a CPU, DSP, GPU, XPU, etc. While it may include any number of example cores 902 (e.g., one core), this example of the microprocessor circuitry 900 is a multi-core semiconductor device including N cores. The cores 902 of the microprocessor circuitry 900 can operate independently or collaboratively to execute machine-readable instructions. For example, machine code corresponding to firmware, embedded software, or a software program can be executed by one of the cores 902, or by multiple cores 902 at the same or different times. In some examples, the machine code corresponding to the firmware, embedded software, or software program is divided into threads and executed in parallel by two or more cores 902. The software program may correspond to... Figure 9 and / or Figure 10 The flowchart represents part or all of the machine-readable instructions and / or operations.
[0191] Core 902 can communicate via a first example bus 904. In some examples, the first bus 904 can implement a communication bus to enable communication associated with one or more of cores 902. For example, the first bus 904 can implement at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 904 can implement any other type of computing or electrical bus. Core 902 can obtain data, instructions, and / or signals from one or more external devices via example interface circuitry 906. Core 902 can output data, instructions, and / or signals to one or more external devices via interface circuitry 906. While the core 902 of this example includes example local memory 920 (e.g., a Level 1 (L1) cache that can be partitioned into an L1 data cache and an L1 instruction cache), the microprocessor circuitry 900 also includes example shared memory 910 (e.g., a Level 2 (L2) cache) that can be shared by the cores for high-speed access to data and / or instructions. Data and / or instructions can be transferred (e.g., shared) by writing to and / or reading from shared memory 910. The local memory 920 and shared memory 910 of each core 902 can be a multi-level cache memory and main memory (e.g., Figure 8 The cache hierarchy is part of the main memory (814, 816). Typically, higher-level memories in this hierarchy exhibit lower access times and smaller storage capacities compared to lower-level memories. Variations in the various levels of the cache hierarchy are managed by cache coherency strategies (e.g., reconciliation).
[0192] Each core 902 may be referred to as a CPU, DSP, GPU, etc., or any other type of hardware circuitry. Each core 902 includes a control unit circuitry 914, an arithmetic and logic (AL) circuitry (sometimes called ALU circuitry) 916, multiple registers 918, an L1 cache 920, and a second example bus 922. Other structures may also be present. For example, each core 902 may include vector unit circuitry, single instruction multiple data (SIMD) unit circuitry, load / store unit (LSU) circuitry, branch / jump unit circuitry, floating-point unit (FPU) circuitry, etc. The control unit circuitry 914 includes semiconductor-based circuitry configured to control (e.g., coordinate) data movement within the corresponding core 902. In some examples, the control unit circuitry 914 is referred to as control circuitry. The AL circuitry 916 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on the data within the corresponding core 902. In some examples, the AL circuit 916 performs integer-based operations. In other examples, the AL circuit 916 also performs floating-point operations. In still other examples, the AL circuit 916 may include a first AL circuit performing integer-based operations and a second AL circuit performing floating-point operations. In some examples, the AL circuit 916 may be referred to as an Arithmetic Logic Unit (ALU). Register 918 is a semiconductor-based structure used to store data and / or instructions, such as the results of one or more operations performed by the AL circuit 916 of the corresponding core 902. For example, register 918 may include one or more vector registers, one or more SIMD registers, one or more general-purpose registers, one or more flag registers, one or more segment registers, one or more machine-specific registers, one or more instruction pointer registers, one or more control registers, one or more debug registers, one or more memory management registers, one or more machine check registers, and so on. Register 918 may be as follows: Figure 9 The diagram shows the arrangement as a bank. Alternatively, register 918 can be organized in any other arrangement, format, or structure, including distribution throughout core 902 to reduce access time. The second bus 922 can implement at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.
[0193] Each core 902 and / or more generally, the microprocessor circuitry 900 may include additional and / or alternative structures as shown and described above. For example, one or more clock circuits, one or more power sources, one or more power gates, one or more cache home agents (CHAs), one or more converged / common mesh stops (CMSs), one or more shifters (e.g., one or more barrel shifters), and / or other circuitry may be present. The microprocessor circuitry 900 is a semiconductor device fabricated to include a plurality of interconnected transistors to implement the above-described structures in one or more integrated circuits (ICs) contained in one or more packages. The processor circuitry may include one or more accelerators and / or work in conjunction with one or more accelerators. In some examples, accelerators are implemented by logic circuitry to perform certain tasks faster and / or more efficiently than a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. GPUs or other programmable devices may also be accelerators. Accelerators may be on the same board as the processor circuitry, in the same chip package as the processor circuitry, and / or in one or more packages separate from the processor circuitry.
[0194] Figure 9 yes Figure 6 A block diagram of another example implementation of the processor circuit 812. In this example, the processor circuit 812 is implemented by the FPGA circuit 1000. For example, the FPGA circuit 1000 can be used to, for example, perform actions that would otherwise be possible via... Figure 7 The example microprocessor circuit 900 executes operations by executing corresponding machine-readable instructions. However, once configured, the FPGA circuit 1000 instantiates the machine-readable instructions in hardware, thus often executing operations faster than a general-purpose microprocessor executing the corresponding software.
[0195] More specifically, as described above Figure 10 The microprocessor circuit 900 (it is a general-purpose device that can be programmed to execute...) Figure 6 and / or Figure 7 The flowchart represents part or all of the machine-readable instructions, but its interconnections and logic circuitry are fixed once manufactured. In contrast, Figure 6 The example FPGA circuit 1000 includes interconnects and logic circuits, which can be configured and / or interconnected in different ways after manufacturing to instantiate, for example, by... Figure 7 and / or Figure 6The flowchart represents part or all of the machine-readable instructions. Specifically, the FPGA circuit 1000 can be considered as an array of logic gates, interconnects, and switches. Switches can be programmed to change the way logic gates are interconnected, effectively forming one or more dedicated logic circuits (unless and until the FPGA circuit 1000 is reprogrammed). The configured logic circuits enable the logic gates to cooperate in different ways to perform different operations on the data received by the input circuits. These operations can correspond to... Figure 7 and / or Figure 6 The flowchart represents part or all of the software. Therefore, the FPGA circuit 1000 can be constructed to effectively utilize... Figure 7 and / or Figure 10 The flowchart's machine-readable instructions, in whole or in part, are instantiated into dedicated logic circuits to perform operations corresponding to these software instructions in a dedicated manner similar to that of an ASIC. Therefore, the FPGA circuit 1000 performs operations related to... Figure 10 and / or Figure 9 The speed at which some or all of the machine-readable instructions correspond to operations can be faster than the speed at which a general-purpose microprocessor executes those instructions.
[0196] exist Figure 6 In the example, the FPGA circuit 1000 is configured to be programmed (and / or reprogrammed once or multiple times) by the end user using a hardware description language (HDL) (such as Verilog). Figure 7 The FPGA circuit 1000 includes example input / output (I / O) circuitry 1002 to obtain and / or output data to example configuration circuitry 1004 and / or external hardware (e.g., external hardware circuitry) 1006. For example, configuration circuitry 1004 may implement interface circuitry that can obtain machine-readable instructions to configure FPGA circuitry 1000, or parts thereof. In some such examples, configuration circuitry 1004 may obtain machine-readable instructions from a user, a machine (e.g., hardware circuitry (e.g., programmable or dedicated circuitry)) that can implement artificial intelligence / machine learning (AI / ML) models to generate instructions, etc. In some examples, external hardware 1006 may implement… Figure 10 The microprocessor circuit 900. The FPGA circuit 1000 also includes an array of example logic gates 1008, multiple example configurable interconnects 1010, and example memory circuitry 1012. The logic gates 1008 and interconnects 1010 can be configured to instantiate and... Figure 10 and / or Figure 9At least some of the machine-readable instructions correspond to one or more operations, and / or other expected operations. Figure 10 The logic gate circuits 1008 shown are manufactured in groups or blocks. Each block includes semiconductor-based electrical structures that can be configured into logic circuits. In some examples, the electrical structures include logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that provide basic building blocks for the logic circuits. Each logic gate circuit 1008 contains electrically controllable switches (e.g., transistors) to enable the configuration of the electrical structures and / or logic gates to form a circuit that performs the desired operation. The logic gate circuits 1008 may include other electrical structures such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.
[0197] The interconnects 1010 in the illustrated example are conductive paths, traces, vias, etc., which may include electrically controllable switches (e.g., transistors). The states of these switches can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more logic gates 1008 to program the desired logic circuit.
[0198] The illustrated storage circuit 1012 is configured to store the results of one or more operations performed by the corresponding logic gates. The storage circuit 1012 can be implemented using a register or similar device. In the illustrated example, the storage circuit 1012 is distributed among the logic gates 1008 to facilitate access and improve execution speed.
[0199] Figure 8 The example FPGA circuit 1000 also includes example dedicated operating circuitry 1014. In this example, dedicated operating circuitry 1014 includes dedicated circuitry 1016, which can be invoked to implement common functions to avoid the need for field programming of these functions. Examples of such dedicated circuitry 1016 include memory (e.g., DRAM) controller circuitry, PCIe controller circuitry, clock circuitry, transceiver circuitry, memory, and multiplier-accumulator circuitry. Other types of dedicated circuitry may also be present. In some examples, FPGA circuitry 1000 may also include example general-purpose programmable circuitry 1018, such as example CPU 1020 and / or example DSP 1022. Other general-purpose programmable circuitry 1018 may additionally or alternatively exist, such as GPUs, XPUs, etc., which can be programmed to perform other operations.
[0200] Although Figure 10 and Figure 8 The diagram shows Figure 9These are two example implementations of the processor circuit 812, but many other approaches are envisioned. For example, as mentioned above, modern FPGA circuits may include an onboard CPU, such as... Figure 10 One or more example CPUs 1020. Therefore, Figure 6 The processor circuit 812 can be additionally combined Figure 7 Example microprocessor circuit 900 and Figure 9 The example FPGA circuit 1000 is used for implementation. In some such hybrid examples, it is implemented by... Figure 6 and / or Figure 7 The first part of the machine-readable instruction represented by the flowchart can be derived from... Figure 10 One or more core 902s execute, by Figure 6 and / or Figure 7 The second part of the machine-readable instruction represented by the flowchart can be derived from... Figure 2 The FPGA circuit 1000 executes, and / or is performed by... Figure 2 and / or Figure 8 The third part of the machine-readable instructions shown in the flowchart can be executed by the ASIC. It should be understood that... Figure 9 Some or all of the circuits can thus be instantiated at the same or different times. Some or all of the circuits can be instantiated, for example, in one or more threads that execute simultaneously and / or sequentially. Furthermore, in some examples, Figure 10 Some or all of the circuitry can be implemented within one or more virtual machines and / or containers that execute on the microprocessor.
[0201] In some examples, Figure 8 The processor circuitry 812 can be housed in one or more packages. For example, Figure 11 Microprocessor circuit 900 and / or Figure 8 The FPGA circuitry 1000 can be housed in one or more packages. In some examples, the XPU can be... Figure 8 The processor circuitry 812 is implemented and can be in one or more packages. For example, the XPU may include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in yet another package.
[0202] Figure 6 The diagram illustrates a sample software distribution platform 1105, used to distribute software such as... Figure 7Example machine-readable instructions 832, such as software, are distributed to hardware devices owned and / or operated by a third party. Example software distribution platform 1105 can be implemented by any computer server, data facility, cloud service, etc., capable of storing software and transferring it to other computing devices. The third party can be a customer of the entity that owns and / or operates the software distribution platform 1105. For example, the entity owning and / or operating the software distribution platform 1105 can be the software (e.g.,...) Figure 6 The developer, seller, and / or licensor of the example machine-readable instructions 832. A third party may be a consumer, user, retailer, OEM, etc., who purchases and / or licenses the software for use and / or resells and / or sublicenses it. In the illustrated example, the software distribution platform 1105 includes one or more servers and one or more storage devices. The storage devices store the machine-readable instructions 832, which may correspond to the instructions described above. Figure 7 Example machine-readable instructions and / or operations 600 and / or Figure 8 Example machine-readable instructions and / or operations 700. One or more servers of the example software distribution platform 1105 communicate with network 1110, which may correspond to the Internet and / or any one or more of the example networks 108 described above. In some examples, as part of a business transaction, one or more servers respond to a request to transfer software to a requesting party. Payment for the delivery, sale, and / or licensing of the software may be processed by one or more servers of the software distribution platform and / or by a third-party payment entity. These servers enable purchasers and / or licensors to download machine-readable instructions 832 from the software distribution platform 1105. For example, it may be compatible with... Example machine-readable instructions and / or operations 600 and / or Software corresponding to the example machine-readable instructions and / or operations 700 can be downloaded to the example processor platform 800, which will execute the machine-readable instructions 832 to implement the code-embedded circuit 104. In some examples, one or more servers of the software distribution platform 1105 periodically provide, transmit, and / or force software updates (e.g., Example machine-readable instructions (832) are provided to ensure that improvements, patches, updates, etc., are distributed and applied to the software at the end-user device.
[0203] As will be clear from the foregoing, example systems, methods, apparatuses, and artifacts for generating code embeddings that depend on usage have been disclosed. AI models for achieving code intelligence tasks are beginning to play a more significant role in improving programmer productivity and code quality. AI-driven code assistance tools are gaining widespread adoption, and a large number of code intelligence tools rely on high-quality code embeddings. Therefore, the examples disclosed in this paper improve the performance of a wide range of code intelligence tools by improving the quality of code embeddings.
[0204] Capturing the semantic intent of an input code snippet is relevant to downstream tasks such as code fixing, as this intent informs the processor circuitry performing the downstream task of information such as the programmer's intent. Therefore, since the examples disclosed herein improve code embedding quality by capturing information about the programmer's intent, they provide downstream tasks with a strong indicator of the programmer's intent and what he / she believes the code snippet should do. By leveraging this information, the examples disclosed herein increase the chances of generating embeddings that accurately capture the programmer's intent.
[0205] As mentioned above, code embedding quality plays a significant role in the success of many code intelligence tasks. The high-quality embeddings disclosed herein go beyond capturing the syntactic properties of code and encode high-level semantic information about the code's functionality. The disclosed methods, apparatuses, and artifacts include an embedding generation technique that provides an additional source of code-related information not captured by existing techniques. For example, the examples disclosed herein capture the code's usage context. The example methods, apparatuses, and artifacts disclosed herein consider not only the code body but also the code's usage scenario to generate more informative embeddings. Therefore, the disclosed systems, methods, apparatuses, and artifacts improve the efficiency of using computing devices by improving the performance of downstream code intelligence networks. The disclosed systems, methods, apparatuses, and artifacts thus represent one or more improvements to the operation of machines such as computers or other electronic and / or mechanical devices.
[0206] This article discloses example methods, apparatuses, systems, and artifacts for generating code embeddings that depend on usage. Further examples and combinations thereof include the following:
[0207] Example 1 includes an apparatus for generating usage-dependent code embeddings, the apparatus comprising interface circuitry for obtaining code including code fragments to be processed by an artificial intelligence (AI) model, and processor circuitry including one or more of the following: at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP), wherein the CPU, GPU, or DSP has: control circuitry for controlling the movement of data within the processor circuitry; arithmetic and first logic circuitry for performing one or more first operations corresponding to instructions; and one or more registers for storing a first result of the one or more first operations, the instructions being in the apparatus; a field-programmable gate array (FPGA) circuitry including second logic gates, a plurality of configurable interconnects, and storage circuitry, the second logic gates and interconnects performing one or more second operations, the storage circuitry storing a second result of the one or more second operations, or a dedicated set of... An integrated circuit (ASIC) includes a third logic gate circuit for performing one or more third operations, wherein the processor circuit performs at least one of the first operation, the second operation, or the third operation to instantiate a parsing circuit for selecting a use context for the code segment, the use context including at least one LOC before the code segment or before the line of code (LOC) that calls the code segment, the code segment, and at least one LOC after the code segment or after the LOC that calls the code segment; an embedding circuit for generating a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code segment and a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the use context, the fourth list including a turn-off token; and a concatenation circuit for concatenating the transformed token embedding vector of the turn-off token and a fifth list of one or more transformed token embedding vectors for the first list of one or more token embedding vectors.
[0208] Example 2 includes the apparatus as described in Example 1, wherein the processor circuitry performs at least one of the first operation, the second operation, or the third operation to instantiate the cascaded circuitry to place the transformed token embedding vector of the turn-off token in front of a fifth list of the one or more transformed token embedding vectors.
[0209] Example 3 includes an apparatus as described in any one of Example 1 or 2, wherein the processor circuitry performs at least one of the first operation, the second operation, or the third operation to instantiate a transformation circuitry to generate a fifth list of the one or more transformed token embedding vectors for a first list of the one or more token embedding vectors, and to generate a sixth list of the one or more transformed token embedding vectors for a third list of the one or more token embedding vectors.
[0210] Example 4 includes an apparatus as described in any one of Examples 1, 2, or 3, wherein the processor circuit is a first processor circuit, and the first processor circuit performs at least one of the first operation, the second operation, or the third operation to instantiate the concatenation circuit for appending a turn-off vector to a concatenation list of a transformed token embedding vector of the turn-off token and a fifth list of the one or more transformed token embedding vectors, and a transformation circuit for generating a transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and a transformed concatenation list of the turn-off vectors, and sending at least one of the transformed concatenation list or the transformed turn-off vectors to a second processor circuit implementing the AI model.
[0211] Example 5 includes an apparatus as described in any one of Examples 1, 2, 3, or 4, wherein at least one of at least one LOC before or after the LOC of the code segment is invoked, or at least one LOC after the code segment is invoked, provides the AI model with information about the code segment to be processed, including information about parameters of the code segment, information about how the output of the code segment is used, or information about the programming context of using the code segment.
[0212] Example 6 includes an apparatus as described in any one of Examples 1, 2, 3, 4 or 5, wherein at least one LOC before the code segment or before the LOC of calling the code segment and at least one LOC after the code segment or after the LOC of calling the code segment correspond to a threshold number of LOCs.
[0213] Example 7 includes an apparatus as described in any one of Examples 1, 2, 3, 4 or 5, wherein at least one LOC preceding the code segment or before the LOC that calls the code segment corresponds to a first LOC threshold number, and at least one LOC following the code segment or after the LOC that calls the code segment corresponds to a second LOC threshold number different from the first LOC threshold number.
[0214] Example 8 includes a non-transitory computer-readable medium comprising machine-readable instructions that, when executed, cause processor circuitry to obtain code including a code segment to be processed by an artificial intelligence (AI) model, select a use context for the code segment, the use context including at least one LOC before or before a line of code (LOC) that calls the code segment, the code segment, and at least one LOC after or after the LOC that calls the code segment, generate a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code segment, and generate a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the use context, the fourth list including a turn-off token, and concatenating a transformed token embedding vector of the turn-off token with a fifth list of one or more transformed token embedding vectors for the first list of one or more token embedding vectors.
[0215] Example 9 includes a non-transitory computer-readable medium as described in Example 8, wherein the instructions cause the processor circuitry to place the transformed token embedding vector of the turn-off token before a fifth list of the one or more transformed token embedding vectors.
[0216] Example 10 includes a non-transitory computer-readable medium as described in any of Examples 8 or 9, wherein the instructions cause the processor circuitry to generate a fifth list of the one or more transformed token embedding vectors for a first list of the one or more token embedding vectors, and to generate a sixth list of the one or more transformed token embedding vectors for a third list of the one or more token embedding vectors.
[0217] Example 11 includes a non-transitory computer-readable medium as described in any of Examples 8, 9, or 10, wherein the processor circuitry is a first processor circuitry, and the instructions cause the first processor circuitry to append a turn-off vector to a concatenated list of a transformed token embedding vector of the turn-off token and a fifth list of the one or more transformed token embedding vectors, generate a transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and a transformed concatenated list of the turn-off vectors, and send at least one of the transformed concatenated list or the transformed turn-off vectors to a second processor circuitry implementing the AI model.
[0218] Example 12 includes a non-transitory computer-readable medium as described in any of Examples 8, 9, 10, or 11, wherein at least one of at least one LOC before or after the code segment or the LOC that calls the code segment provides the AI model with information about the code segment to be processed, including information about parameters of the code segment, information about how the output of the code segment is used, or information about the programming context of using the code segment.
[0219] Example 13 includes a non-transitory computer-readable medium as described in any one of Examples 8, 9, 10, 11 or 12, wherein at least one LOC preceding the code segment or before the LOC that calls the code segment and at least one LOC following the code segment or after the LOC that calls the code segment correspond to a threshold number of LOCs.
[0220] Example 14 includes a non-transitory computer-readable medium as described in any one of Examples 8, 9, 10, 11 or 12, wherein at least one LOC preceding the code segment or before the LOC that calls the code segment corresponds to a first LOC threshold number, and at least one LOC following the code segment or after the LOC that calls the code segment corresponds to a second LOC threshold number different from the first LOC threshold number.
[0221] Example 15 includes an apparatus for generating usage-dependent code embeddings, the apparatus including at least one memory, instructions, and processor circuitry for executing instructions to at least obtain code comprising a code segment to be processed by an artificial intelligence (AI) model, selecting a usage context for the code segment, the usage context including at least one LOC before or before a line of code (LOC) that calls the code segment, the code segment, and at least one LOC after the code segment or the LOC that calls the code segment, generating a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code segment, and generating a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a turn-off token, and concatenating a transformed token embedding vector of the turn-off token with a fifth list of one or more transformed token embedding vectors for the first list of one or more token embedding vectors.
[0222] Example 16 includes the apparatus as described in Example 15, wherein the processor circuitry executes the instructions to place the transformed token embedding vector of the turn-off token before a fifth list of the one or more transformed token embedding vectors.
[0223] Example 17 includes an apparatus as described in any of Examples 15 or 16, wherein the processor circuitry executes the instructions to generate a fifth list of the one or more transformed token embedding vectors for a first list of the one or more token embedding vectors, and to generate a sixth list of the one or more transformed token embedding vectors for a third list of the one or more token embedding vectors.
[0224] Example 18 includes an apparatus as described in any one of Examples 15, 16, or 17, wherein the processor circuitry is a first processor circuitry, and the first processor circuitry executes the instructions to append a turn-off vector to a concatenated list of a transformed token embedding vector of the turn-off token and a fifth list of the one or more transformed token embedding vectors, generate a transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and a transformed concatenated list of the turn-off vectors, and send at least one of the transformed concatenated list or the transformed turn-off vectors to a second processor circuitry implementing the AI model.
[0225] Example 19 includes an apparatus as described in any one of Examples 15, 16, 17 or 18, wherein at least one of at least one LOC before or after the code segment or after the LOC of the code segment is invoked provides the AI model with information about the code segment to be processed, including information about parameters of the code segment, information about how the output of the code segment is used, or information about the programming context of using the code segment.
[0226] Example 20 includes an apparatus as described in any one of Examples 15, 16, 17, 18 or 19, wherein at least one LOC before the code segment or before the LOC of calling the code segment and at least one LOC after the code segment or after the LOC of calling the code segment correspond to the number of LOC thresholds.
[0227] Example 21 includes an apparatus as described in any one of Examples 15, 16, 17, 18 or 19, wherein at least one LOC preceding the code segment or before the LOC of calling the code segment corresponds to a first LOC threshold number, and at least one LOC following the code segment or after the LOC of calling the code segment corresponds to a second LOC threshold number different from the first LOC threshold number.
[0228] Example 22 includes a method for generating usage-dependent code embeddings, the method comprising obtaining code including a code snippet to be processed by an artificial intelligence (AI) model, selecting a usage context for the code snippet, the usage context including at least one LOC before or before a line of code (LOC) that calls the code snippet, the code snippet, and at least one LOC after or after a LOC that calls the code snippet, generating a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code snippet, and generating a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a turn-off token, and concatenating a transformed token embedding vector of the turn-off token with a fifth list of one or more transformed token embedding vectors from the first list of one or more token embedding vectors.
[0229] Example 23 includes the method as described in Example 22, further comprising placing the transformed token embedding vector of the turn-off token before a fifth list of the one or more transformed token embedding vectors.
[0230] Example 24 includes the method as described in any of Examples 22 or 23, further comprising generating a fifth list of the one or more transformed token embedding vectors from a first list of the one or more token embedding vectors, and generating a sixth list of the one or more transformed token embedding vectors from a third list of the one or more token embedding vectors.
[0231] Example 25 includes the method as described in any of Examples 22, 23, or 24, further comprising appending a turn-off vector to a concatenated list of a transformed token embedding vector of the turn-off token and a fifth list of the one or more transformed token embedding vectors, generating a transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and a transformed concatenated list of the turn-off vectors, and sending at least one of the transformed concatenated list or the transformed turn-off vectors to processor circuitry implementing the AI model.
[0232] Example 26 includes a method as described in any of Examples 22, 23, 24, or 25, wherein at least one of at least one LOC before or after the code snippet or the LOC after the call to the code snippet provides the AI model with information about the code snippet to be processed, including information about parameters of the code snippet, information about how the output of the code snippet is used, or information about the programming context of using the code snippet.
[0233] Example 27 includes the method described in any one of Examples 22, 23, 24, 25 or 26, wherein at least one LOC before the code segment or before the LOC of calling the code segment and at least one LOC after the code segment or after the LOC of calling the code segment correspond to the number of LOC thresholds.
[0234] Example 28 includes the method described in any one of Examples 22, 23, 24, 25 or 26, wherein at least one LOC preceding the code segment or before the LOC of calling the code segment corresponds to a first LOC threshold number, and at least one LOC following the code segment or after the LOC of calling the code segment corresponds to a second LOC threshold number different from the first LOC threshold number.
[0235] Example 29 includes an apparatus for generating usage-dependent code embeddings, the apparatus including means for parsing code to obtain code including code snippets to be processed by an artificial intelligence (AI) model, and selecting a usage context for the code snippets, the usage context including at least one LOC before or before a line of code (LOC) that calls the code snippets, the code snippets, and at least one LOC after or after a LOC that calls the code snippets; means for embedding code to generate a first list of one or more token embedding vectors for a first token of a second list of one or more tokens of the code snippets, and a third list of one or more token embedding vectors for a second token of a fourth list of one or more tokens of the usage context, the fourth list including a turn-off token; and means for concatenating vectors to concatenate transformed token embedding vectors of the turn-off tokens with a fifth list of one or more transformed token embedding vectors of the first list of one or more token embedding vectors.
[0236] Example 30 includes the device as described in Example 29, wherein the means for concatenating vectors is used to place the transformed token embedding vector of the closing token before a fifth list of the one or more transformed token embedding vectors.
[0237] Example 31 includes the device as described in any of Examples 29 or 30, and further includes means for transforming vectors to generate a fifth list of the one or more transformed token embedding vectors for a first list of the one or more token embedding vectors, and to generate a sixth list of the one or more transformed token embedding vectors for a third list of the one or more token embedding vectors.
[0238] Example 32 includes an apparatus as described in any of Examples 29, 30, or 31, wherein the means for concatenating vectors is used to append a turn-off vector to a concatenated list of a transformed token embedding vector of the turn-off token and a fifth list of the one or more transformed token embedding vectors, and the apparatus further includes means for transforming vectors to generate a transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and a transformed concatenated list of the turn-off vectors, and to send at least one of the transformed concatenated list or the transformed turn-off vectors to processor circuitry implementing the AI model.
[0239] Example 33 includes a device as described in any one of Examples 29, 30, 31 or 32, wherein at least one of at least one LOC before or after the code snippet or the LOC after the call to the code snippet provides the AI model with information about the code snippet to be processed, including information about the parameters of the code snippet, information about how to use the output of the code snippet, or information about the programming context of using the code snippet.
[0240] Example 34 includes a device as described in any one of Examples 29, 30, 31, 32 or 33, wherein at least one LOC before the code segment or before the LOC of calling the code segment and at least one LOC after the code segment or after the LOC of calling the code segment correspond to the number of LOC thresholds.
[0241] Example 35 includes a device as described in any one of Examples 29, 30, 31, 32 or 33, wherein at least one LOC preceding the code segment or before the LOC that calls the code segment corresponds to a first LOC threshold number, and at least one LOC following the code segment or after the LOC that calls the code segment corresponds to a second LOC threshold number different from the first LOC threshold number.
[0242] The appended claims are hereby incorporated into this Detailed Description section by reference. While certain example systems, methods, apparatuses, and articles of manufacture are disclosed herein, the scope of this patent is not limited thereto. Rather, this patent covers all systems, methods, apparatuses, and articles of manufacture that fairly fall within the scope of the claims of this patent.
Claims
1. One or more computer-readable media, the computer-readable media containing instructions that, in response to being executed by one or more processors, cause the one or more processors to: Obtain code that includes code snippets that will be processed by an artificial intelligence (AI) model; Select the usage context of the code snippet, in which the code snippet is used and / or invoked; A first list of one or more token embedding vectors is generated for the first token of a second list of one or more tokens of the code snippet, and a third list of one or more token embedding vectors is generated for the second token of a fourth list of one or more tokens of the context, the fourth list including a turn-off token; as well as A fifth list of one or more transformed token embedding vectors concatenated with the transformed token embedding vector of the closed token and a first list of the one or more token embedding vectors.
2. The computer-readable medium of claim 1, further comprising instructions that, in response to being executed by one or more processors, cause the one or more processors to: The transformed token embedding vector of the closed token is placed before the fifth list of the one or more transformed token embedding vectors.
3. The computer-readable medium of claim 1, further comprising instructions that, in response to being executed by one or more processors, cause the one or more processors to: A fifth list of transformed token embedding vectors is generated for the first list of the one or more token embedding vectors, and a sixth list of transformed token embedding vectors is generated for the third list of the one or more token embedding vectors.
4. The computer-readable medium of claim 1, further comprising instructions that, in response to being executed by one or more processors, cause the one or more processors to: The closing vector is appended to the concatenated list of the following: the transformed token embedding vector of the closing token, and a fifth list of the one or more transformed token embedding vectors; Generate a transformed concatenated list of the following: the transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and the turn-off vector; and At least one of the following is transmitted to the processor circuitry implementing the AI model: the transformed concatenated list or the transformed closed vector.
5. The computer-readable medium of claim 1, wherein the usage context of the code snippet provides the AI model with at least one of the following information: information about the parameters of the code snippet, information about how the output of the code snippet is used, and information about the programming context of using the code snippet.
6. A calculation method, comprising: Obtain code that includes code snippets that will be processed by an artificial intelligence (AI) model; Select the usage context of the code snippet, in which the code snippet is used and / or invoked; A first list of one or more token embedding vectors is generated for the first token of a second list of one or more tokens of the code snippet, and a third list of one or more token embedding vectors is generated for the second token of a fourth list of one or more tokens of the context, the fourth list including a turn-off token; as well as A fifth list of one or more transformed token embedding vectors concatenated with the transformed token embedding vector of the closed token and a first list of the one or more token embedding vectors.
7. The method of claim 6, further comprising placing the transformed token embedding vector of the closing token before a fifth list of the one or more transformed token embedding vectors.
8. The method of claim 6, further comprising generating a fifth list of the one or more transformed token embedding vectors for a first list of the one or more token embedding vectors, and generating a sixth list of the one or more transformed token embedding vectors for a third list of the one or more token embedding vectors.
9. The method according to claim 6, further comprising: The closing vector is appended to a concatenated list of the following: the transformed token embedding vector of the closing token and a fifth list of the one or more transformed token embedding vectors; Generate a transformed concatenated list of the following: the transformed token embedding vector of the turn-off token, a fifth list of the one or more transformed token embedding vectors, and the turn-off vector; and At least one of the following is transmitted to the processor circuitry implementing the AI model: the transformed concatenated list or the transformed closed vector.
10. The method of claim 6, wherein the usage context of the code snippet provides the AI model with at least one of the following information: information about the parameters of the code snippet, information about how the output of the code snippet is used, and information about the programming context of using the code snippet.
11. A computing system comprising: At least one memory for storing instructions; A processor circuit for executing the instructions to perform the method described in any one of claims 6 to 10.
12. A computing device comprising means for performing the method of any one of claims 6 to 10.
13. A computer program product comprising instructions that, in response to being executed on a computing device, cause the computing device to perform the method of any one of claims 6 to 10.