System for assisting in executing software engineering tasks

By integrating machine learning models and software tools, the system achieves automated conversion from textual descriptions of software engineering tasks to source code, solving the problems of insufficient automation and natural language input support in existing tools, and improving the accuracy and efficiency of source code generation.

CN121039680APending Publication Date: 2025-11-28LARIDO LABORATORIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093259.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-30
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing software tools struggle to automate and efficiently convert software engineering task requirements and specifications into source code, and lack direct support for natural language input from software developers.

Method used

By integrating machine learning models with software tools, and by specializing the textual descriptions of software engineering tasks, machine learning models are used to predict source code changes and interact with developers, enabling pipelined source code generation and repair.

Benefits of technology

It improves the automation of software engineering tasks, enhances support for natural language input from developers, and improves the accuracy and efficiency of source code generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121039680A_ABST
    Figure CN121039680A_ABST
Patent Text Reader

Abstract

A system is described that trains a machine learning model to assist in performing software engineering tasks. The system retrieves data from a data source associated with a software engineering task. The system links data by linking each problem report describing any one of the software engineering tasks with source code, the source code being associated with the any one of the software engineering tasks. The system transforms data to be compatible with a data format for training a machine learning model to assist in performing software engineering tasks. The system trains the machine learning model with the transformed data to assist in performing a software engineering task by making a prediction of source code changes associated with the software engineering task.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Software engineering is the process of constructing functional software from requirements and / or specifications, which are human-centric, typically in the form of natural language documents and supporting data. The constructed software is machine-centric, in the form of textual source code and supporting data. Software developers typically use software tools to assist in the complex tasks of creating requirements and specification documents and constructing source code.

[0002] Interpretation and refinement of requirements and specification documents is required, a process that involves understanding ideas at various levels of abstraction (vague and precise) and organizing them into a coherent software design. Software engineers need to understand technical natural language documents and reconcile their implicit requirements with the capabilities of the underlying computing platform of the source code.

[0003] Software tools for requirements and specification management and design / architecture are a broad category of products that include process-focused management systems that facilitate communication and cataloging of requirements and specifications, modeling tools that allow visualization of potential software designs at various levels of fidelity, and project management systems that are typically used to store requirements and specifications and track their implementation progress. These early software tools are similar in that they facilitate a relatively narrow range of tasks and do not even attempt to fully automate these tasks.

[0004] Software tools for software construction form an even broader category of products. Most of these tools are implemented in code editors that provide assistance for many tasks. Code editors make source code more readable by organizing and highlighting it and facilitate source code navigation by hyperlinks. Source code can include readme files and other types of text files such as boilerplate license headers that are typically associated with source code files. These software construction tools edit code to make it syntactically consistent and add dependency source code constructs such as import statements. Code editors generate boilerplate source code from fixed templates and reduce typing by completing partially typed words or source code lines. Software construction tools streamline the process of writing source code but do not automate its writing or assist software developers in writing source code that is more appropriately linked to requirements or specifications. BRIEF DESCRIPTION OF DRAWINGS

[0005] Figure 1A and Figure 1B illustrates an example of a source code line that can be used to assist in performing software engineering tasks in one embodiment;

[0006] Figure 2FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0007] Figure 3 FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0008] Figure 4A and Figure 4B FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0009] Figure 5A and Figure 5B FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0010] Figure 6A and Figure 6B FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0011] Figure 7A and Figure 7 FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0012] Figure 8A and Figure 8 FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented.

[0013] FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented. DETAILED DESCRIPTION

[0014] Figure 1A and Figure 1B FIG. 9 is a block diagram illustrating an example hardware device in which the subject matter can be implemented. Figure 1A As depicted, source code line 102 illustrates an example source code for performing standard artificial intelligence (AI) driven code completion that interprets the comment to generate a complete source code line. As depicted, Figure 1A As depicted, source code line 104 illustrates an example source code for performing more sophisticated AI driven code completion that generates an entire function body from a function name and documentation lines contained in the comment, which is similar to existing state-of-the-art code completion style tools.

[0015] Embodiments of the present disclosure provide a system that includes a higher-order concept tool unlike any other tool on the market. For example, the system integrates more directly with software issue trackers and enables developers to experiment more with pure natural language input. The system is significantly different from software issue trackers that generate bug reports in that the system is able to specialize its behavior through the text description of the requirements and specifications of a software engineering task. Recognizing that changes are central to modern software development, the system extends the standard code completion convention with machine learning models that can propose changes to bug reports and source code and tool chains that can present these changes in a pipelined fashion.

[0016] As Figure 1A depicted, the bug report 106 describes a software engineering task that the system can automatically read from issue tracking software that software developers use to edit bug reports, requirements, and specifications. The text of the bug report has a strong influence on what source code is generated because this capability is derived entirely from the system’s innovative use of the project history. The source code suggestions or predictions focus on the specific software engineering task. As Figure 1A depicted, the source code line 108 illustrates a specific prediction for source code that is based on the bug report 106 for the software engineering task and appears in the context of existing source code.

[0017] Sometimes, the machine learning model is confident about the scope of the source code change — not just what needs to change, but where the change ends. As Figure 1B depicted, the source code line 110 illustrates a predicted source code change that is fully isolated into a single line and presented as a boxed-in pop-up.

[0018] When confident, the machine learning model can suggest source code changes at locations other than where the software developer is currently working. As Figure 1B depicted, the source code line 112 illustrates a prediction of a complete source code change 8 lines below the current cursor. And in some cases, the predicted source code change can be hundreds of lines away — or entirely in another source code file. Being able to locate predictions of source code changes — rather than just predicting them — is key to realizing the significant challenge of making artificial intelligence self-sufficient in solving problems.

[0019] In addition to predicting new source code lines, the system can also help software developers adapt and fix their existing source code. As Figure 1BAs depicted, source code line 114 illustrates a machine learning model that is highly confident in a prediction of a portion of source code that does not match a corresponding portion in existing source code, and therefore renders the corresponding portion in existing source code in boldface and underlined. A software developer will notice this obvious signal on the corresponding portion in existing source code, select the highlighted portion in source code to see a comparison between the existing portion in source code and the predicted portion in source code, realize any errors, and correct any errors.

[0020] Embodiments of the present disclosure provide a system and an agent that automates general software engineering tasks that broadly encompass work involving interpreting problem reports, requirements, and specifications, and implementing the interpreted problem reports, requirements, and specifications into source code. The system learns, is trained from various software engineering data sources. Past performed and recorded work is standardized, abstracted, and built into one or more machine learning models. The agent then acts as an interface between a software developer and the one or more machine learning models, thereby leveraging knowledge learned from past work to streamline and automate current and future work.

[0021] Embodiments herein provide a system that trains a machine learning model to assist in performing a software engineering task. The system retrieves data from data sources associated with the software engineering task. The system links the data by linking each problem report describing any one of the software engineering tasks with source code (associated with the any one of the software engineering tasks). The system transforms the data to make it compatible with a data format for training the machine learning model to assist in performing the software engineering task. The system trains the machine learning model with the transformed data to assist in performing the software engineering task by making predictions of source code changes associated with the software engineering task.

[0022] For example, the system retrieves data from a large corpus of open source software engineering data through a dynamic data pipeline, such that the scale of learned knowledge can grow to the knowledge of collective software engineering intelligence, thereby capturing the state-of-the-art in the field of software engineering as a whole. The system links the retrieved data by linking commits of source code through a version control system to issue reports in an issue tracker that the commits resolve, and linking issue reports to more general documents that the issue reports cite. The system transforms the linked data into a general uniform format, such as representing commits as a difference between a before state and an after state in a textual format that is encoded. The system trains a machine learning model with the transformed data that enables making various predictions about source code changes for a software engineering task. The more links between historical issue reports and historical source code changes used to train the machine learning model, the more current source code changes the machine learning model can predict for a current issue report, and the more accurate the predictions will be.

[0023] Embodiments herein provide a system that uses a machine learning model to assist in performing a software engineering task. The system receives a request from an issue tracker or a code editor of a software developer to start working on a software engineering task, and outputs an issue report that describes the software engineering task and / or source code associated with the software engineering task to the issue tracker and / or the code editor. The system stores updates to the issue report or source code changes associated with the software engineering task received from the issue tracker and / or the code editor. The system receives a request from the issue tracker or the code editor for a predicted completion of the software engineering task, retrieves context data that establishes a context of the software engineering task, and transforms the context data to be compatible with a data format used to train the machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed context data to predict any number of completions of the software engineering task. The system enables the software developer to complete the software engineering task by outputting the predicted completion of the software engineering task to the issue tracker and / or the code editor.

[0024] For example, the system outputs the requested problem report for the software engineering task (to fix the subject check error response) and / or the source code to a code editor or issue tracker of a software developer named Sofia, and then the system stores updates of Sofia, including her source code changes. In response to a request by Sofia for any predicted completion of the software engineering task, the system retrieves contextual data that establishes a context of the software engineering task, such as the source code changes of Sofia. Next, the system transforms the contextual data (which includes the source code changes of Sofia) to make it compatible with a data format used to train a machine learning model to assist in performing software engineering tasks. The trained machine learning model uses the transformed contextual data to predict a set of source code changes with a 95% predicted confidence level for the software engineering task and another set of source code changes with a 90% predicted confidence level for the software engineering task. The system outputs a succinct representation of these predicted sets of source code changes to Sofia, who can use her issue tracker to expand any selected succinct representation to provide her with predicted source code changes within the context of the existing source code, such as Sofia can accept the predicted source code changes that she submits, Figure 1A the predicted source code change at line 651 in the depicted source code line 108. When enabling Sofia to complete the software engineering task, the machine learning model can provide multiple alternative sets of predicted source code changes (that have sufficiently high predicted confidence levels to be reasonable alternatives) as completion of the software engineering task, from which Sofia can select any one of the sets of predicted source code changes to submit.

[0025] Embodiments herein provide a system that assists in authoring a problem report that describes software engineering tasks and problem information. The system receives a request from a stakeholder associated with a problem tracker of a software engineering task to start working on an incomplete problem report that describes software engineering tasks and problem information, and that is intended for a software developer, and assigns the software engineering task to the software developer. The system receives a request from the stakeholder for a predicted completion of the incomplete problem report, retrieves contextual data that establishes a context of the software engineering task, and transforms the contextual data to be compatible with a data format used to train a machine learning model to assist with the software engineering task. The trained machine learning model uses the transformed contextual data to predict a completion of the incomplete problem report that describes software engineering tasks and problem information. The system enables the software developer to complete the software engineering task by outputting the accepted completion of the incomplete problem report to the software developer's problem tracker based on the predicted completion of the incomplete problem report that describes software engineering tasks and problem information.

[0026] For example, the system initiates a requested problem report for a software engineering task (to fix the subject check error response) to a problem tracker of Stacy Holder, who is a stakeholder of the software engineering task, and then the system responds to Stacy's next request by assigning the software engineering task to a software developer named Sofia. In response to Stacy's subsequent request to predict a completion of her incomplete problem report, the system retrieves contextual data that establishes a context of the software engineering task, and transforms the contextual data (which includes the incomplete problem report) to be compatible with a data format used to train a machine learning model to assist with performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict a completion of the incomplete problem report for the software engineering task (to fix the subject check error response) that corrects the description of the software engineering task from requiring removal of variable names when they are omitted to requiring removal of variable names when they are embedded. The system enables Sofia to complete her assigned software engineering task by outputting the corrected description. In addition to the machine learning model being trained to predict source code changes for the software engineering task, the machine learning model has been trained on a sufficiently diverse set of problem reports to be able to predict a completion of the incomplete problem report that can require deletion and / or replacement of a portion of the incomplete problem report, such as deletion / replacement of the word "omitted" that was mistakenly included in the software engineering task.

[0027] Embodiments herein provide a system that assists in automating software engineering tasks. The system receives a request for a review issue report from an issue tracker of a software developer and outputs an issue report to the issue tracker, the issue report describing a software engineering task. The system receives a request of a predicted source code change for the software engineering task from the issue tracker, retrieves contextual data that establishes a context of the software engineering task, and transforms the contextual data to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict a source code change for the software engineering task. The system outputs the predicted source code change for the software engineering task to the issue tracker. The system commits the source code change into source code associated with the software engineering task based on the predicted source code change accepted by the issue tracker.

[0028] For example, the system outputs a requested issue report describing a software engineering task (to fix a subject check error response) to an issue tracker of a software developer named Sofia. In response to a request of Sofia for a predicted source code change for the software engineering task (to fix a subject check error response), the system retrieves contextual data that establishes a context of the software engineering task, such as an update of an issue report of the software engineering task by Sofia. Next, the system transforms the contextual data (which includes the updated issue report of Sofia) to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict a source code change, such as Figure 1A a new source code at line 651 in the depicted source code line 108. The system outputs the predicted source code change for the software engineering task to Sofia, who commits any source code changes she accepts. Even if Sofia only works on the issue report describing her software engineering task without generating any source code changes for the software engineering task and then requests a predicted source code change for the software engineering task, the system can use the updated issue report to predict each source code change needed to complete the software engineering task, thereby automating the generation of source code changes.

[0029] Embodiments herein provide a system that assists in authoring source code for a software engineering task. The system stores source code changes received from a code editor of a software developer at a location in source code associated with the software engineering task. The system receives a request from the code editor for a predicted source code change at the location in the source code, retrieves contextual data that establishes a context for the software engineering task, and transforms the contextual data to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict the source code change at the location in the source code. The system outputs the predicted source code change at the location in the source code to the code editor of the software developer. The system commits the source code change to the source code associated with the software engineering task based on the predicted source code change being accepted by the software developer at the location in the source code.

[0030] For example, the system stores some source code changes required for a software engineering task to fix a subject check error response, which are generated by a code editor of a software developer named Sofia. In response to a request by Sofia for a predicted source code change at line 701 in the source code where she is working on the software engineering task to fix a subject check error response, the system retrieves contextual data that establishes a context for the software engineering task, such as source code changes by Sofia. Next, the system transforms the contextual data, which includes source code changes by Sofia, to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. As Figure 1B As illustrated by line 704 in the depicted source code lines 110, the trained machine learning model uses the transformed contextual data to predict a source code change at line 701 in the source code to complete the localized software engineering task to fix a subject check error response. The system outputs the predicted source code change at line 704 in the source code to the code editor of Sofia, which enables her to complete her localized software engineering task, commit the predicted source code change that she accepts, and move on to any other location in the source code that she wishes to focus on to complete her software engineering task to fix a subject check error response.

[0031] Embodiments herein provide a system that assists in identifying unpredicted portions of source code files of software engineering tasks. The system receives source code changes from a code editor associated with a software developer and stores the source code changes in a source code file associated with a software engineering task. The system receives a request from the code editor to predict source code of the source code file, retrieves contextual data that establishes a context of the software engineering task, and transforms the contextual data to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict source code of the source code file, where portions in the source code file correspond to portions in the predicted source code. The system identifies, via the code editor, each portion in the source code file that is determinative of a corresponding portion in the predicted source code. The system submits any disparate portions in the predicted source code requested and accepted by the code editor into the source code file.

[0032] For example, the system stores source code changes in a source code file of a software engineering task to fix subject check error responses, the source code changes generated by a code editor of a software developer named Sofia. In response to a request by Sofia to predict source code of the source code file of the software engineering task (to fix subject check error responses), the system retrieves contextual data that establishes a context of the software engineering task, such as source code changes in the source code file by Sofia. Next, the system transforms the contextual data (that includes source code changes by Sofia) to be compatible with a data format used to train a machine learning model to assist in performing the software engineering task. The trained machine learning model uses the transformed contextual data to predict source code of the source code file, where each portion in the source code file corresponds to a portion in the predicted source code. Sofia uses her code editor to review the predicted source code and select a portion of the predicted source code that is determinative of a corresponding portion in the source code file. Sofia accepts the portion of the predicted source code to replace the corresponding portion in the source code file. If the source code changes by Sofia begin to fix subject check error responses, then many highlighted portions in the source code file can identify predicted source code needed to complete her software engineering task, but if the source code changes by Sofia complete her task, then few highlighted portions in the source code file can identify predicted source code that corrects her possible errors and / or makes her source code more efficient. Figure 1B The bolded and underlined portion of line 701 in the depicted source code line 114 Sofia selects the bolded portion of line 701 (that shows a portion in the predicted source code that is different from the bolded portion of line 701 ), and accepts the portion in the predicted source code to replace the bolded portion of line 701. If the source code changes by Sofia begin to fix subject check error responses, then many highlighted portions in the source code file can identify predicted source code needed to complete her software engineering task, but if the source code changes by Sofia complete her task, then few highlighted portions in the source code file can identify predicted source code that corrects her possible errors and / or makes her source code more efficient.

[0033] Figure 2A block diagram of a system 200 that assists in performing software engineering tasks in one embodiment is illustrated. As shown Figure 2 The system 200 can illustrate a cloud computing environment in which data, applications, services and other resources are stored and delivered through shared data centers and appear as a single access point for end users, as shown. The system 200 can also represent any other type of distributed computer network environment in which servers control storage and distribution of resources and services for different client users.

[0034] In embodiments, the system 200 represents a cloud computing system that includes a first client 202, a second client 204, and a first server 206 and a second server 208 that can be provided by a hosting company. The clients 202-204 and servers 206-208 communicate via a network 210.

[0035] The client 202, which can be referred to as a client 202 of a software developer, includes word processor-like software for editing source code, such as a version control system, which can be referred to as a code editor 212, that contains a code assistant 214 that can be implemented as a plug-in that extends the native functionality of the code editor 212 by focusing on helping the user write the most appropriate source code to accomplish a given task. Similarly, the client 202 includes word processor-like software, which can be referred to as an issue tracker 216, for writing issue reports and natural language descriptions of project requirements and specifications for software engineering tasks that contains a writing assistant 218 that can be implemented as a plug-in that extends the native functionality of the issue tracker 216 by focusing on helping the user write the most appropriate natural language text to accomplish a given task. Similarly, the client 204, which can be referred to as a client 204 of a stakeholder of a software engineering task, can include a code editor 220 that can contain a code assistant 222 and an issue tracker 224 that can contain a writing assistant 226.

[0036] The server 206, which can be referred to as a training server 206, can include a dynamic data pipeline 228, one or more machine learning models 230, a retrieval component 232, a linking component 234, a transformation component 236, and a training component 238. The server 208, which can be referred to as a production server 208, can include a dynamic data pipeline 240, one or more machine learning models 242, and an agent 244. The machine learning models 230 and 242 can be based on any of a variety of models, such as a gradient boosting classifier, a k-nearest neighbor classifier, a neural network, a random forest, a support vector machine, a naive Bayes classifier, and a logistic regression model. The assistants 214, 218, 222, and 226 can be provided by the training server 206 and / or the production server 208 to assist the code editor 212, the issue tracker 216, the code editor 220, and the issue tracker 224 in interacting with the components 228-238 that reside on the training server 206 and / or the system components 240-244 that reside on the production server 208. The system 200 can include any number of clients 202-204, any number of servers 206-208, any number of networks 210, and any number of components 212-244 depicted in Figure 2 as residing on any of the clients 202-204 or servers 206-208.

[0037] The clients 202-204 and servers 206-208 can each be substantially similar to the system 900 depicted in FIG. 9 and described below. Figure 2 The system components 228-238 are depicted as residing entirely on the first server 206, and the system components 240-244 are depicted as residing entirely on the second server 208, but the system components 228-244 can reside entirely on the first server 206, entirely on the second server 208, entirely on the clients 202-204, entirely on Figure 2 another server not depicted in FIG. 9, or partially on the servers 206-208, partially on the clients 202-204, and partially on other servers, in any combination.

[0038] Software engineers, software developers, and software project stakeholders typically maintain digital records of their work in various information repositories, such as issue trackers, project documentation, textual communications, and version control systems. Issue trackers, which were originally designed to track defects in software, are now used to archive and track progress of all types of software engineering work. Issue trackers contain numbered natural language issue reports, which can be assigned as work items to personnel such as software developers. These issue reports have meta-tags such as priority or subsystem, and they have annotation and editing capabilities that can enable updates as work progresses.

[0039] Project documentation consists of documents authored by software developers themselves or other software project stakeholders. These documents can contain information about project requirements, specifications, schedules, work progress, and / or descriptions of the software being built itself. Textual communications occur between team members, including both software developers and other software project stakeholders, and can be synchronous or asynchronous and direct (one-to-one) or group. Major examples include email and chat platforms.

[0040] Version control systems are temporal databases that hold authoritative copies of source code for software engineering projects. Version control systems are notable in that they store a complete history of all data they contain. The history is partitioned into atomic units, commonly referred to as commits, which contain a set of file-level changes coupled with a textual description. This history allows software developers to track the provenance of source code changes and to retrieve any recorded version of the source code. In addition to software source code, version control systems often contain configuration, support data, and even some forms of documentation.

[0041] Embodiments of the system 200 learn by processing heterogeneous software engineering data from different sources into a common stream and using the data stream to train the machine learning model(s) 230 to assist in performing software engineering tasks. Figure 3 FIG. 1 is a block diagram of an embodiment of a system 200 that trains machine learning model(s) 230 to assist in performing software engineering tasks. The system 200 includes a client 202-204 and / or a server 206-208. The client 202-204 and / or the server 206-208 are involved in and / or there are certain steps between them. Figure 2 FIG. 2 is a flow diagram of an embodiment of a method that trains machine learning model(s) 230 to assist in performing software engineering tasks. The flow diagram 300 illustrates method actions, which are illustrated as flow diagram blocks, that are used to

[0042] The system retrieves data from a plurality of data sources associated with software engineering tasks (block 302). The system collects various types and quantities of data. For example and without limitation, this can include retrieving all data by the component 232 through the dynamic data pipeline 228 to a central database.

[0043] Data can be information that has been translated into a form efficient for movement or processing. A data source can be the location where the information originates. Software engineering tasks can be a set of work performed on the design, development, testing, and / or maintenance of computer programs.

[0044] Each data source connects to a central location. Local data sources are accessed by reading files, while network sources are accessed via network file transfer or a more standardized application programming interface (API). Retrieval component 232 can either completely replicate the data to the central database or connect to the data, making it accessible on demand.

[0045] Retrieval component 232 can retrieve data from data sources, including those associated with multiple software engineering projects (associated with a single enterprise). For example, training server 206 can extend data acquisition and subsequent learning processes from information from a single project to the acquisition and training of data from more than one project. Within an organization, this process can facilitate institutional memory, enabling the transfer of work and knowledge learned from past and cross-organizational projects to new and current projects. Software engineering projects can be carefully planned efforts to achieve specific goals in the design, development, testing, and / or maintenance of computer programs. A single enterprise can be a business or company.

[0046] Retrieval component 232 can retrieve data from data sources including open-source software projects. For example, training server 206 can further extend the same data acquisition and learning process by using a wide range of open software engineering data, such as information sources publicly available from open-source software projects. As training server 206 is enhanced and expanded, the learned knowledge grows from institutional memory to a collective software engineering intelligence, thereby capturing the current state-of-the-art level of technology in the field of software engineering as a whole. Open-source software projects can be carefully planned efforts to achieve specific goals in the design, development, testing, and / or maintenance of computer programs whose code is set to be freely available for possible modification and redistribution.

[0047] Training server 206 encodes the item history as text and allows it to be learned within a large language model. Training server 206 can scale this process indefinitely and achieve any level of accuracy (limited by computational cost and data availability). The item history model follows the same scaling laws as other large language modeling tasks, meaning that scaling it with more data and more computational power will make it more accurate, as more resources invested lead to higher accuracy. As training server 206 diversifies its data according to even more diverse items, accuracy can reach breakthrough levels.

[0048] After retrieving the data, the system links the data by linking each problem report describing any one software engineering task with the source code (associated with any one software engineering task) (block 304). The system links the data to enable prediction of source code data based on problem report data. By way of example and not limitation, this can include linking component 234 linking the data together to establish locality. A problem report can be a textual description of a difficulty. Source code can be a textual list of commands that will be compiled or assembled into an executable computer program. Further, source code can include readme files and other types of text files such as boilerplate license headers that are typically associated with source code files.

[0049] Many data entries will be more strongly related to each other than they are to other data entries, such as a document about subsystem A is more strongly related to the source code of subsystem A than it is to any other source code. Linking component 234 can link commits of source code made through a version control system (such as code editor 212) to problem reports in problem tracker 216 that the commits addressed, and link problem reports to more general documents that the problem reports reference. Training server 206 can expand such links and store the expanded links in a central database.

[0050] After linking the data, the system transforms the data to make it compatible with the data format used to train any machine learning model(s) (block 306). The system unifies the various data formats of the various types of data. In an embodiment, this can include transforming component 236 transforming the data into a common unified format that is compatible with the learning method in use.

[0051] Compatibility can be the ability to use with a specified software without special adaptation or modification. Data format can be a structured organization of information. Machine learning model can be the application of artificial intelligence to dynamic data that provides a system with the ability to automatically learn and improve from experience without explicit programming.

[0052] While each data source is primarily based on text, the structure and content of that text varies. Transformations vary with the input data and the learning method, but example transformations can involve: the transformation component 236 taking a source code commit from a version control system (which can generally be viewed as a before state and an after state), and representing that commit as a text formatted, encoded difference of those two states. Another example of a transformation can involve: the training server 206 filtering out irrelevant and unhelpful parts of individual data items (such as boilerplate license headers from source code files). The transformation component 236 can transform data independently, thereby producing a processed dataset, or the transformation component 236 can transform data on demand as part of a pipeline.

[0053] After transforming the data, the system trains the machine learning model(s) with the transformed data to assist in performing software engineering tasks by making predictions of source code changes for the software engineering tasks (block 308). The system trains the machine learning model(s) to predict source code for a current issue based on history of source code created for issues. For example and without limitation, this can include the training component 238 applying a learning method to train the machine learning model(s) 230. The training component 238 can feed a complete pipeline to a learning method that trains the machine learning model(s) 230 with data that is capable of making various predictions about that data and similar data, such as source code changes for software engineering tasks. Examples of learning methods include rule-based engines and machine learning methods such as deep neural networks.

[0054] Transformed data can be information that has been converted from one format to another. Predictions can be forecasts. Source code changes can be modifications to a text list of commands that will be compiled or assembled into an executable computer program.

[0055] The training component 238 trains each instance of the machine learning model(s) 230 on data that open source developers contributed strictly before a cutoff date, and then performs an evaluation / experiment on all data strictly after that date. In simple terms, the training component 238 evaluates the ability of the machine learning model(s) 230 to perform “new” work when trained on “old” work.

[0056] Software engineering is a continually updating discipline. Across a broad range, best practices evolve, new programming languages emerge, and the popularity of libraries and other components rises and falls. More urgent updates occur when previously unknown security vulnerabilities are discovered and must be addressed in current software and prevented in future source code. The training server 206 can adapt the retrieval, linking, and transformation of data steps of the general learning process by adding new source data, thereby refreshing the database with new entries. Those same steps of retrieval, linking, and transformation of data can also serve to remove existing data, thereby discarding bad or outdated software engineering knowledge.

[0057] Accordingly, after the machine learning model(s) are initially trained, the system iteratively retrieves additional data from the data source, links the additional data, transforms the additional data, and trains the machine learning model(s) with the transformed additional data (block 310). The system can update the training of the machine learning model(s). As an example and not by way of limitation, this can include the training server 206 applying the learning process using the updated data to create updated machine learning model(s) 230. Iteratively training the machine learning model(s) 230 with transformed additional data can include initializing the machine learning model(s) 230 and then training the initialized machine learning model(s) 230 using only the transformed additional data, training the machine learning model(s) 230 using both the transformed additional data and previously transformed data, or incrementally training the machine learning model(s) 230 using the transformed additional data. For example, the training server 206 can train the machine learning model(s) 230 from scratch using the additional data, thereby creating a brand new machine learning model(s) 230 from the additional data, or incrementally train the machine learning model(s) 230 using the additional data.

[0058] The transformed additional data can be supplemental information that has been converted from one format to another. Incrementally can be a regular increase, addition, or staging. The initialized machine learning model can be an artificial intelligence applied to dynamic data that provides the system with the ability to automatically learn and improve from experience without explicit programming and that has been placed in a starting condition.

[0059] The training server 206 adds, removes, and updates data in the database as annotation data, so that the data is recognizable by the learning process. If supported, the training of the machine learning model(s) 230 can then restart its process using the current machine learning model(s) 230 as input, focusing only on the updates to the data. This focus can be facilitated by the training server 206 making the learning process only able to access the updated data, or the training server 206 can continue to access all data but adjust the weight of its updated data accordingly.

[0060] A standard learning process can enable the machine learning model(s) 230 to learn the knowledge of a single project or team, while an extended learning process can enable the machine learning model(s) 230 to learn general software engineering knowledge. In many cases, both types of machine learning model 230 will be beneficial. Thus, the machine learning model(s) 230 can include a set of machine learning model(s) 230 trained with data associated with multiple software engineering projects (associated with a single enterprise) and machine learning model(s) 230 trained with data associated with general software engineering knowledge, or machine learning model(s) 230 trained with both data associated with multiple software engineering projects (associated with a single enterprise) and data associated with general software engineering knowledge. For example, downstream agents 244 can utilize these machine learning models 230 that are adapted to use more than one of the machine learning model(s) 242, such as by using one focused machine learning model 242 for team-specific tasks and one broad machine learning model 242 for general knowledge. However, this process can also be facilitated by the learning process itself, such as by precisely producing one hybrid machine learning model 242 that is trained with both team-specific tasks and broad general knowledge.

[0061] Assuming the training server 206 has performed the above wide-ranging adaptation and generated the general machine learning model(s) 242, the training server 206 can then perform the steps of retrieving, linking, and transforming data on potentially private and confidential data for one project, team, or individual. Thus, training the machine learning model(s) 230 can include learning with data associated with general software engineering knowledge and then learning with data associated with a single enterprise, with the data associated with the single enterprise weighted less than the data associated with the general software engineering knowledge. For example, the training server 206 can adapt by using the general machine learning model 230 as an initialization point and then applying the learning method to the updated data based on training the specialized machine learning model 230 with specialized data. To facilitate the blending of the knowledge captured in the machine learning model(s) 230 without overwriting it, the training server 206 can adapt the learning process to reduce the weight of its new specialized data. The general software engineering knowledge can be a broad understanding of the design, development, testing, and / or maintenance of computer programs. The more links between historical problem reports and historical source code changes used to train the machine learning model(s) 230, the more current source code changes the machine learning model(s) can predict for a current problem report, and the more accurate the prediction will be.

[0062] The system 200 is built around the project history machine learning model(s) 242, which will assist in smoothly integrating with software developers’ existing workflows. The system 200 exposes the machine learning model(s) 242 via tools (such as the code editors 212 and 220) that are similar to, but functionally encompass, code completion style tools. The advantage of these code completion style tools is that they enable software developers to very smoothly switch between assisted and non-assisted development. When such a tool makes an unexpected action, there is no jarring moment that breaks their flow: the software developer simply ignores the assistance and continues typing.

[0063] While Figure 3 Blocks 302-310 are depicted as occurring in a particular order, but blocks 302-310 can occur in other orders. In other embodiments, each of blocks 302-310 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0064] To facilitate and automate software engineering tasks, the system 200 connects the machine learning model(s) 242 to software developers and software developers' data through the agent 244. The role of the agent 244 is to: collect real-time data about the software engineering task to be processed, pass the real-time data to the machine learning model(s) 242 to query the way the software engineering task is to be completed or assisted, and present that automation or assistance to a human software developer.

[0065] The general form of the agent 244 involves a dynamic data pipeline 240 that mirrors the dynamic data pipeline 228 used to build the machine learning model(s) 242. The agent 244 uses the dynamic data pipeline 240 to establish a context for the current software engineering task. The agent 244 applies that context to the machine learning model(s) 242, resulting in some form of fully or partially automated work. That work is then presented by the agent 244 through the instantiation of various specific tools.

[0066] The machine learning model(s) 242 that result from the learning process capture a wide variety of software engineering knowledge. The machine learning model(s) 242 learn about problem reports, requirements, specifications, and designs from natural language sources, and about software constructs from structured source code sources via programming languages. Linking the data sources enables the machine learning model(s) 230 to connect these different types of knowledge.

[0067] The steps of retrieval, linking, and transforming data of the general learning process describe the connection, centralization, linking, and normalization of software engineering data as it is consumed by the machine learning model(s) 230 during the learning process. The use of such machine learning model(s) 242 by the agent 244 requires the extraction or querying of knowledge from the machine learning model(s) 242, which in turn requires a similar but localized data extraction process. The software engineering work performed or facilitated by the agent 244 is done in the context of a specific project and software engineering task. A subset of the project's linked data will be relevant to that software engineering task and must be retrieved when assisting that software engineering task. For example, a software engineering task assigned through a problem tracking software can include the text of a problem report as dynamic data, a software engineering task discussed in a specific thread of a chat or email system can include the text of that thread, and a task related to a specific subsystem of a software project can include the programming language and source code of that subsystem.

[0068] Collection of such data is performed in the dynamic data pipeline 240. The collection or retrieval mechanism can be similar to the one used in the instantiation of the general learning process, but it must be adapted to additional constraints. The data that will be processed is typically much less than the data consumed during training, and the current data must be updated and collected in real time. For example, if a participating project stakeholder creates a comment that elucidates a report of the original problem when interacting with the machine learning model(s) 242 in the future, the dynamic data pipeline 240 can include that elucidation.

[0069] If the software developer seeks additional elucidation about the software engineering task from a colleague through a chat system, future model queries by the agent 244 can include the text of that interaction. If the software developer performs a portion of the software engineering task by writing and saving an amount of source code, further interactions between the agent 244 and the machine learning model(s) 242 will include that portion of the work done. If the software developer deletes an amount of source code to partially complete the software engineering task, continuing interactions will include the fact that the source code has been removed.

[0070] The dynamic data pipeline 240 is responsible for maintaining a constant data connection to each of these sources. Example forms of these connections include application programming interfaces (APIs), database connections, and local or network file system read operations.

[0071] The agent 244 uses the dynamic data pipeline 240 to establish a context for a given software engineering task that is presented as a query to the machine learning model(s) 242. The training server 206 can have (and will likely or ideally) exposed the machine learning model(s) 230 to similar or similarly adaptable contexts. This enables the machine learning model(s) 242 to use its embedded knowledge to make predictions about how the software engineering task should be performed.

[0072] FIG. 4 is a flowchart that illustrates a method of using machine learning model(s) to assist in performing a software engineering task, in one embodiment. The flowchart 400 illustrates method actions that are illustrated as flowchart blocks that are used to Figure 2 certain steps involved with and / or between the clients 202-204 and / or servers 206-208.

[0073] After a system, such as an issue tracking system, assigns a software engineering task to a software developer, the software developer logs into an issue tracker 216 on a client 202 and begins reading the task description. The system receives a request from a code editor or issue tracker associated with the software developer to begin working on the software engineering task (block 402). The system stores the source code and issue report for the software engineering task. For example and without limitation, this can include the agent 244 receiving a request from a code editor 212 or issue tracker 216 on a client 202 of a software developer named Sofia to begin working on the software engineering task to fix the subject check error response and remove variable names when not embedded by writing source code or editing an issue report.

[0074] The request can be a message soliciting information or resources. The software developer can be a person who designs, creates, and / or maintains an application (which allows a user to perform a specific task on a computer). The code editor can be a tool used to write software. The issue tracker can be a tool that helps manage and resolve difficulties.

[0075] After receiving the request from the software developer to begin working on the software engineering task, the system outputs the issue report describing the software engineering task and / or the source code for the software engineering task to an issue tracker and / or code editor associated with the software developer (block 404). The system provides the source code and issue report to the software developer. For example and without limitation, this can include the agent 244 outputting the issue report and / or source code for the software engineering task to fix the subject check error response and remove variable names when not embedded to Sofia.

[0076] After optionally outputting the issue report to the software developer, the system optionally enables the software developer to clarify the issue report describing the software engineering task and / or update the issue report to describe a strategy for completing the software engineering task via an issue tracker associated with a stakeholder of the software engineering task (block 406). The system enables updates to the issue report describing the software engineering task. In an embodiment, this can include the agent 244 enabling Sofia to communicate with the issue tracker 224 on a client 204 of Stacy Holder, a stakeholder of the software engineering task to fix the subject check error response and remove variable names when not embedded. Sofia and Stacy agree on a high priority (but not the highest priority) for the software engineering task to fix the subject check error response and remove variable names when not embedded.

[0077] A stakeholder can be a person who has an interest or concern in something, especially a business. A strategy can be a plan of action or policy designed to achieve a major or overall aim.

[0078] After optionally elaborating on the problem report, the system stores any updates to the problem report describing the software engineering task and / or any source code for the software engineering task received from the issue tracker and / or code editor associated with the software developer (block 408). The system stores any updates to the source code and problem report for the software engineering task. For example and without limitation, this can include the agent 244 storing an update by Sofia and a limited number of source code changes by Sofia to the dynamic data pipeline 232, the update including a high priority (but not the highest priority) software engineering task to fix the subject check error response and remove variable names when not embedded. The update can be replacing an older version with a newer version.

[0079] After outputting the problem report and / or source code for the software engineering task to the software developer, the system receives an implicit or explicit request from the software developer's issue tracker or code editor for a predicted completion scheme for the software engineering task (block 410). The system is requested to predict the source code for the software engineering task. As an example and without limitation, this can include the agent 244 receiving an explicit request from Sofia's issue tracker 216 or code editor 212 for a predicted completion scheme for the software engineering task to fix the subject check error response and remove variable names when not embedded.

[0080] An explicit request can be a message that overtly seeks information or resources. An implicit request can be a message that implicitly seeks information or resources. A completion can be an act of finishing an action, activity, or process.

[0081] In addition to receiving an explicit request for a predicted completion scheme for the software engineering task, the agent 244 can also identify an implicit request for a predicted completion scheme for the software engineering task. For example, Sofia starts her issue tracker 216 and begins editing the problem report for the software engineering task (to fix the subject check error response and remove variable names when not embedded), and / or starts her code editor 212 and begins writing the source code for the software engineering task (to fix the subject check error response and remove variable names when not embedded), and then stops and waits - at which point the agent 244 notifies her that a lack of editing activity during a displayed timer counts down uninterrupted will result in the machine learning model(s) 242 predicting a completion scheme for the software engineering task (to fix the subject check error response and remove variable names when not embedded) based on the problem report and source code for the software engineering task.

[0082] Upon receiving the request to predict a completion solution for the software engineering task, the system retrieves context data that establishes a context for the software engineering task (block 412). The system retrieves the context data for the software engineering task to accurately predict the software engineering task. In an embodiment, this can include the agent 244 retrieving, from the dynamic data pipeline 232, context data that establishes a context for the software engineering task, such as the description for fixing the subject check error response and removing variable names when not embedded and the update to the problem report that is targeted for a high priority (but not the highest priority) for the software engineering task (for fixing the subject check error response and removing variable names when not embedded), and retrieving the source code changes for Sofia from the dynamic data pipeline 240. The context can be the environment in which the action setting is formed. The context data can be information about the environment in which the action setting is formed, and the context data has been translated into a form that is efficient for movement or processing.

[0083] Upon retrieving the context data that establishes a context for the software engineering task, the system transforms the context data to be compatible with a data format used to train the machine learning model(s) to assist in performing the software engineering task (block 414). The system transforms the context data for the software engineering task to be in a format used to predict the software engineering task. For example and without limitation, this can include the agent 244 transforming the retrieved problem report and source code changes to be compatible with a data format used to train the machine learning model(s) 230 to learn to predict a completion solution for the software engineering task.

[0084] Upon the system transforming the format of the various types of context data for the software engineering task, the machine learning model(s) use the transformed context data to predict a completion solution for the software engineering task (block 416). The system uses the transformed data to predict the source code for the software engineering task. As an example and without limitation, this can include the machine learning model(s) 242 using the transformed context data to predict a first set of source code changes and a second set of source code changes, the first set of source code changes having a 95% prediction confidence level for the software engineering task (for fixing the subject check error response and removing variable names when not embedded), the second set of source code changes having a 90% prediction confidence level for the software engineering task (for fixing the subject check error response and removing variable names when not embedded). The transformed context data can be information about the environment in which the action setting is formed, and the transformed context data has been translated from one format to another format that is efficient for movement or processing.

[0085] A prediction of zero completion solutions for a software engineering task can indicate that the machine learning model(s) 242 are unable to serve the software engineering task or lack any predicted completion solutions for the software engineering task with a prediction confidence level above a threshold. For example, if none of the initially determined completion solutions have a prediction confidence level above a threshold of 50%, the machine learning model(s) 242 can predict that the software engineering task has no completion solutions for fixing the subject check error response and removing variable names when variable names are not embedded.

[0086] The lack can be an absence. The predicted completion solution can be a forecast of a behavior that completes an action, activity, or process. The prediction confidence level can be a probability that the forecast is accurate. The threshold can be an intensity or value of a signal that will produce a response or specified effect.

[0087] Predicting multiple completion solutions for a software engineering task can indicate that the multiple predicted completion solutions for the software engineering task have corresponding prediction confidence levels above a threshold. Each of the multiple completion solutions for the software engineering task with a prediction confidence level above a threshold can be interpreted as a reasonable alternative.

[0088] After predicting a completion solution for a software engineering task, the system enables a software developer to complete the software engineering task by outputting the predicted completion solution for the software engineering task to a code editor or issue tracker associated with the software developer (block 418). The system outputs the predicted completion solution for the software engineering task. In embodiments, this can include the agent 244 outputting two different sets of predicted completion solutions (for fixing the subject check error response and removing variable names when variable names are not embedded) for the software engineering task within or alongside the issue tracker 216, where each predicted completion solution is depicted in a differential form that is a succinct representation of many possible source code changes that have been analyzed by the machine learning model(s) 242 as being able to complete the software engineering task and having a prediction confidence level above a threshold.

[0089] In addition to enabling Sofia to complete software engineering tasks, the machine-learned model(s) 242 can also provide multiple alternative predicted source code change sets (that have a sufficiently high predicted confidence level of being a reasonable alternative) to use as a completion of the software engineering task from which Sofia can select any one of the predicted source code change sets to commit. Since the predicted source code change sets are low-level implementations of a completion of the software engineering task, and the agent 244 outputs the predicted source code changes to Sofia's high-level problem tracker 216, the predicted source code changes can be depicted as a succinct representation of all the predicted source code change sets for each predicted completion of the software engineering task. When Sofia uses her high-level problem tracker 216 to review the succinct representation of all the predicted source code change sets for any predicted completion of the software engineering task, the selection of the single representation causes the agent 244 to expand the selected representation to provide Sofia with the full section of the selected source code changes within the context of the existing source code, such as Figure 1A the predicted source code change depicted in line 651 of the source code lines 108.

[0090] After predicting the source code changes for multiple completions of the software engineering task, the machine-learned model(s) optionally predict other source code changes for another completion of the software engineering task based on modifications to the problem report describing the software engineering task and / or the source code changes associated with the software engineering task received from the problem tracker and / or code editor associated with the software developer (block 420). For example and without limitation, this can include the machine-learned model(s) 242 responding to Sofia modifying the problem report for the software engineering task (Sofia's modification to specify that by selecting a predicted completion of her software engineering task, she completed the source code changes to remove the variable name when the variable name was not embedded) and using this modification as contextual data to predict other source code changes for another completion of the software engineering task (but only to fix the subject check error response without needing to remove the embedded variable name). The modifications can be changes.

[0091] After optionally predicting another completion of the software engineering task on the other source code change, the system optionally outputs the predicted other source code change of the other completion of the software engineering task to the source code editor and / or issue tracker associated with the software developer (block 422). The system responds to the issue report on the software engineering task and / or the modification of the source code by outputting a new prediction for the software engineering task. For example and without limitation, this can include the agent 244 outputting the predicted other source code change of the other completion of the software engineering task (but only for the fix subject check error response without the need to remove the embedded variable name) to Sofia.

[0092] After outputting all of the predicted source code changes, the system optionally commits the predicted source code change associated with one of the completions of the software engineering task based on the acceptance of the predicted source code change by the code editor and / or issue tracker associated with the software developer into the source code associated with the software engineering task (block 424). The system commits the predicted source code accepted by the software developer. By way of example and without limitation, this can include the agent 244 committing the predicted source code change of the completion accepted by Sofia to the source code control repository of the software engineering task and closing the software engineering task on behalf of Sofia. The predicted source code change can be a forecast of modifications to a text list of commands that will be compiled or assembled into an executable computer program.

[0093] The source code change can include source code to be added to the source code associated with the software engineering task, source code to be removed from the source code associated with the software engineering task, and / or source code to be replaced at the source code associated with the software engineering task. In addition to enabling Sofia to complete the software engineering task, the machine-learned model(s) 242 can also provide multiple alternative sets of predicted source code changes (that have a sufficiently high predicted confidence level to be a reasonable alternative) to be used as a completion of the software engineering task from which Sofia can select any one of the sets of predicted source code changes to commit.

[0094] While FIG. 4 depicts blocks 402-424 occurring in a particular order, blocks 402-424 can occur in other orders. In other embodiments, each of blocks 402-424 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0095] The agent 244 presents the prediction or recommendation of the machine learning model(s) 242 through one or more software tools. The user experience of the software developer need not be fully specified within this general architecture, and its ideal form can depend on many factors, including but not limited to the personal preferences of the software developer, the particular type of task being performed within this general framework, or the effectiveness and / or accuracy of the particular instantiation of the machine learning model(s) 242.

[0096] In its general form, the machine learning model(s) 242 are able to read many types of software engineering data and make predictions of its ideal form. While most instantiations focus on making predictions of what source code should be written, thereby automating the work of the software developer, some instantiations can work directly on natural language problem reports and / or project requirements and / or specifications. Clear and accurate natural language problem reports and / or project requirements and / or specifications greatly facilitate the work of the software developer. Natural language focused instantiations of the tool, such as the authoring assistants 218 and / or 226 that can be able to assist project stakeholders in authoring the process, will be presented directly within the issue tracking system.

[0097] FIG. 5 is a flowchart illustrating a method of assisting in authoring a problem report describing software engineering task and problem information, in one embodiment. The flowchart 500 illustrates method actions, which are illustrated as flowchart blocks, for Figure 2 certain steps involved with and / or between the clients 202-204 and / or servers 206-208.

[0098] The system receives a request from an issue tracker associated with a stakeholder of a software engineering task to begin processing an incomplete problem report describing software engineering task and problem information, and the incomplete problem report is intended for a software developer (block 502). The system stores source code and problem reports for software engineering tasks. For example and without limitation, this can include the agent 244 receiving a request from a stakeholder named Stacy Holder, who is using her issue tracker 224 to create a new problem report for a software engineering task (to fix a subject verification error response), and beginning to fill out a textual description of the work she expects a software developer named Sofia to perform. The incomplete problem report can be an unfinished textual specification describing a difficult problem. The problem information can be specified or learned facts about the difficult problem.

[0099] After starting processing of the incomplete problem report describing the software engineering task, the system assigns the software engineering task to a software developer (block 504). The system provides the software developer with the source code and the problem report. By way of example and without limitation, this can include the agent 244 assigning the software engineering task (to fix the subject check error response) to Sofia.

[0100] After assigning the software engineering task, the system receives a request for a completion solution to the predicted incomplete problem report from the issue tracker associated with the stakeholder, the incomplete problem report describing the software engineering task and the problem information (block 506). The system receives the request for a completion solution to the predicted incomplete problem report. In an embodiment, this can include the agent 244 responding to an intentional button press on the issue tracker 224 by receiving a request for a completion solution to the predicted software engineering task to fix the subject check error response from Stacy.

[0101] In addition to receiving an explicit request for a completion solution to the predicted incomplete problem report, the agent 244 can also identify an implicit request for a completion solution to the predicted incomplete problem report. For example, Stacy launches her issue tracker 224 and starts processing of the incomplete problem report describing the software engineering task, and then stops and waits - at which point the agent 244 notifies her that, during the displayed countdown of the timer without interruption, the continued lack of editing activity will cause the machine learning model(s) 242 to predict a completion solution to the incomplete problem report based on any edits made by Stacy to the incomplete problem report.

[0102] After receiving the request for a completion solution to the predicted incomplete problem report, the system retrieves context data that establishes a context for the software engineering task (block 508). The system retrieves the context data for the software engineering task to make a prediction about the software engineering task. By way of example and without limitation, this can include the agent 244 retrieving the incomplete problem report for the software engineering task (to fix the subject check error response) and the already existing source code (for the subject check error response) from the dynamic data pipeline 240 as context data for the software engineering task (to fix the subject check error response).

[0103] After retrieving the context data establishing the context of the software engineering task, the system transforms the context data to be compatible with a data format used to train the machine learning model(s) to assist the software engineering task (block 510). The system transforms the context data to enable a prediction about the software engineering task. By way of example and without limitation, this can include the agent 244 transforming the context data of the software engineering task (to fix the subject check error response) and the source code of the software engineering task (to fix the subject check error response) to be compatible with a data format used to train the machine learning model(s) 230 to predict a completion solution for an incomplete problem report describing the software engineering task.

[0104] After transforming the context data of the software engineering task to be compatible with a format used to train the machine learning model(s), the machine learning model(s) use the transformed context data to predict a completion solution for an incomplete problem report describing the software engineering task and the problem information (block 512). The system makes a prediction about the software engineering task. In an embodiment, this can include the machine learning model(s) 242 using the transformed context data to predict a completion solution for an incomplete problem report of the software engineering task (to fix the subject check error response) that corrects the description of the software engineering task from requiring the variable name to be removed when the variable name is not omitted to requiring the variable name to be removed when the variable name is not embedded, instead of an incorrect description of the software engineering task that requires the variable name to be removed when the variable name is not omitted.

[0105] After using the transformed context data to predict a completion solution for an incomplete problem report describing the software engineering task, the system enables a software developer to complete the software engineering task based on the predicted completion solution of the incomplete problem report describing the software engineering task and the problem information by outputting the accepted completion solution of the incomplete problem report to a problem tracker associated with the software developer (block 514). The system enables a software developer to complete the software engineering task. For example and without limitation, this can include the agent 244 enabling Sofia to more easily and quickly complete her assigned software engineering task by outputting the enhanced description provided by the predicted completion solution of the incomplete problem report of the software engineering task (to fix the subject check error response). The predicted completion solution of the incomplete problem report corrects the description of the software engineering task from requiring the variable name to be removed when the variable name is not omitted to requiring the variable name to be removed when the variable name is not embedded, instead of an incorrect description of the software engineering task that requires the variable name to be removed when the variable name is not omitted. The accepted completion solution can be a selected action, activity, or process to complete.

[0106] An accepted completion of an incomplete problem report can include at least one of a deletion or a replacement of a portion of the incomplete problem report. For example, when the intelligent agent 244 outputs a predicted completion of an incomplete problem report of a software engineering task (for fixing the subject check error response), an incorrect description of the software engineering task (which incorrectly describes: require removing variable name when variable name is omitted) is revised by a strikeout that identifies a proposed deletion of the word "omitted" and a proposed replacement with the word "embedded." A deletion can be the removal of data from a computer. A replacement can be the substitution of one entity for another. A portion can be a segment of something (such as an object) that, in combination with other segments, makes up a whole. In addition to the machine learning model(s) 242 being trained to predict source code changes of software engineering tasks, the machine learning model(s) have also been trained on a sufficiently diverse set of problem reports to be able to predict completion of incomplete problem reports that can require a deletion and / or a replacement of a portion of the incomplete problem report, such as a deletion / replacement of the word "omitted" that was erroneously included in the software engineering task.

[0107] An accepted completion of an incomplete problem report can be based on at least one edit to a predicted completion of the incomplete problem report that describes a software engineering task and problem information, the at least one edit being from a problem tracker associated with a stakeholder, and the problem information can include whether the problem is reproducible, steps to reproduce the problem, and / or a current priority associated with the problem. For example, Stacy reviews a predicted completion of an incomplete problem report and notices that several items that would be important for Sofia to know are still not mentioned in the incomplete problem report, such as whether the problem is reliably reproducible, steps to reproduce the problem, and how high the current priority of the problem is. Accordingly, Stacy accepts the predicted completion of the incomplete problem report and edits the generated text to specify that the problem is reproducible, steps to reproduce the problem, and a high priority (but not the highest priority) of the software engineering task (for fixing the subject check error response).

[0108] An edit can be a change to text. A problem can be a difficult question. Reproducible can be the ability to show, do, or make again. Current priority can be a fact or condition that is currently considered or regarded as more important than others.

[0109] In addition to enabling the software developer to complete the software engineering task, the system optionally enables the software developer to clarify an incomplete problem report describing the software engineering task via the stakeholder-associated problem tracker, and / or update the incomplete problem report to describe a strategy for completing the software engineering task (block 516). The system enables clarification and updating of the incomplete problem report. By way of example and not limitation, this can include the agent 244 enabling Sofia to contact Stacy through her respective problem trackers 216 and 224 to precisely clarify how to take the second step needed to reproduce the (occasionally occurring) difficulty with the subject check error response, and record that clarification detail as part of the incomplete problem report.

[0110] After predicting a completion of the incomplete problem report, the machine learning model(s) optionally predict another completion of the incomplete problem report describing the software engineering task and problem information based on modifications to the predicted completion of the incomplete problem report describing the software engineering task and problem information received from the software developer-associated problem tracker (block 518). The system can revise the prediction of the completion of the incomplete problem report. In embodiments, this can include the machine learning model(s) 242 revising the predicted completion of the incomplete problem report to require the software engineering task to only fix the subject check error response without requiring removal of the embedded variable name because Sofia modified the incomplete problem report to indicate that she had already completed removal of the variable name when the variable name was not embedded in the subject check error response.

[0111] After optionally predicting another completion of the incomplete problem report, the system can output the other completion of the incomplete problem report describing the software engineering task and problem information to the software developer-associated problem tracker (block 520). The system outputs the revised prediction of the completion of the incomplete problem report. By way of example and not limitation, this can include the agent 244 outputting the new prediction of the completion of the incomplete problem report to Sofia's problem tracker 216 that only requires the software engineering task to fix the subject check error response without requiring removal of the embedded variable name.

[0112] While FIG. 5 depicts blocks 502-520 occurring in a particular order, blocks 502-520 can occur in other orders. In other embodiments, each of blocks 502-520 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0113] The most general tool instantiation of this framework is to enable software developers to automatically perform general software engineering tasks given their natural language description and peripheral software-related data. This software developer interface can be presented alongside or within the system through which the software developer is assigned work. For example, many software developers receive their tasks through issue tracking systems, which also often serve project management duties. In this instantiation, the software developer's user experience can be manifested in the method described in the flowchart of Figure 6.

[0114] Figure 6 is a flowchart illustrating a method that assists in automating software engineering tasks, in one embodiment. The flowchart 600 illustrates method actions, which are illustrated as flowchart blocks, that are involved with and / or between the clients 202-204 and / or servers 206-208. Figure 2

[0115] After a system, such as an issue tracking system, assigns a software engineering task to a software developer, the software developer logs into the issue tracker 216 on the client 202 and begins reading the issue report that describes the software engineering task. The system receives a request to review the issue report from the issue tracker associated with the software developer, the issue report describing the software engineering task (block 602). The system stores the source code and the issue report for the software engineering task. For example and without limitation, this can include the agent 244 receiving a request to review the issue report from the issue tracker 216 on the client 202 of the software developer named Sofia, the issue report describing the software engineering task to fix the subject check error response and remove variable names when not embedded.

[0116] After receiving the request to review the issue report from the software developer's issue tracker, the system outputs the issue report describing the software engineering task to the issue tracker associated with the software developer (block 604). The system provides the source code and the issue report to the software developer. For example and without limitation, this can include the agent 244 outputting the issue report to the issue tracker 216 of Sofia, the issue report describing the software engineering task to fix the subject check error response and remove variable names when not embedded.

[0117] ​After outputting the problem report describing the software engineering task to the software developer, the system optionally enables the software developer to clarify the problem report describing the software engineering task via the problem tracker associated with the stakeholder of the software engineering task, and / or update the problem report to describe a strategy for completing the software engineering task (block 606). The system enables updates to the problem report describing the software engineering task. In an embodiment, this can include the agent 244 enabling Sofia to communicate with the problem tracker 216 on the client 204 of Stacy Holder, the stakeholder of the software engineering task, to clarify the problem report for the software engineering task to fix the subject check error response and remove variable names when not embedded. Sofia and Stacy agree on a high priority (but not highest priority) software engineering task to fix the subject check error response and remove variable names when not embedded.

[0118] After optionally clarifying the problem report, the system optionally stores any updates to the problem report describing the software engineering task received from the problem tracker associated with the software developer (block 608). Any updates to the source code and the problem report for the software engineering task are stored. For example and without limitation, this can include the agent 244 storing the updates by Sofia for the high priority (but not highest priority) software engineering task to fix the subject check error response and remove variable names when not embedded to the dynamic data pipeline 240.

[0119] After outputting the problem report describing the software engineering task to the software developer's problem tracker, the system receives an implicit or explicit request from the problem tracker associated with the software developer to predict a source code change for the software engineering task (block 610). The system is requested to predict a source code change for the software engineering task. By way of example and without limitation, this can include the agent 244 receiving an explicit request for a predicted source code change for the software engineering task to fix the subject check error response and remove variable names when not embedded from a button (pressed on the problem tracker 216 of Sofia).

[0120] In addition to receiving an explicit request for a source code change for a predicted software engineering task, the agent 244 can also identify an implicit request for a source code change for a predicted software engineering task. For example, Sofia starts her issue tracker 216 and begins updating the problem report describing (one or more) software engineering tasks to fix subject check error responses and remove variable names when not embedded, and then stops and waits - at which point the agent 244 notifies her that, during the displayed countdown timer without interruption, continued lack of editing activity will result in the machine learning model(s) 242 predicting a source code change based on any modifications Sofia makes to the problem report describing the software engineering task to fix subject check error responses and remove variable names when not embedded.

[0121] After receiving a request for a source code change for a predicted software engineering task, the system retrieves context data establishing a context for the software engineering task (block 612). The system retrieves the context data for the software engineering task to make accurate predictions for the software engineering task. In an embodiment, this can include the agent 244 retrieving, from the dynamic data pipeline 240, context data establishing a context for the software engineering task, such as the description to fix subject check error responses and remove variable names when not embedded and updates to the problem report that are pertinent to the high-priority (but not highest-priority) software engineering task to fix subject check error responses and remove variable names when not embedded.

[0122] After retrieving the context data establishing a context for the software engineering task, the system transforms the context data to be compatible with a data format used to train the machine learning model(s) to assist in performing the software engineering task (block 614). The system transforms the context data for the software engineering task to be in a format for predicting the software engineering task. For example and without limitation, this can include the agent 244 transforming the retrieved problem report for the software engineering task to be compatible with a data format used to train the machine learning model(s) 230 to learn to predict a source code change for the software engineering task.

[0123] After transforming the various types of context data for the software engineering task, the machine learning model(s) use the transformed context data to predict a source code change for the software engineering task (block 616). The system uses the transformed problem report to predict a source code change for the software engineering task. As an example and without limitation, this can include the machine learning model(s) 242 using the transformed context data to predict a source code change, such as Figure 1Anew source code at line 651 in depicted source code line 108. The source code changes can include: a) new source code that will be added to the source code associated with the software engineering task, b) some existing source code that will be removed from the source code associated with the software engineering task, and / or c) new source code for replacing some source code associated with the software engineering task.

[0124] After predicting the source code changes for the software engineering task, the system outputs the predicted source code changes for the software engineering task to the issue tracker associated with the software developer (block 618). The system outputs the predicted source code changes for the software engineering task. In embodiments, this can include the agent 244 outputting the predicted source code changes for the software engineering task to the issue tracker 216 of Sofia, the predicted source code changes including Figure 1A new source code at line 651 in depicted source code line 108.

[0125] These predicted or suggested source code changes can be presented in the issue tracker 216 or alongside in a diff form, which is a succinct representation of the many possible source code changes that have been analyzed by the machine learning model(s) 242 as being able to complete the software engineering task. Since the predicted source code changes are low-level implementations of the completion solution for the software engineering task, and the agent 244 outputs the predicted source code changes to the high-level issue tracker 216 of Sofia, the predicted source code is depicted as a succinct representation of the predicted source code changes for the software engineering task. When Sofia uses her high-level issue tracker 216 to review the succinct representation of the predicted source code changes for the software engineering task, the selection of an individual representation causes the agent 244 to expand the selected representation to provide Sofia with the full section of predicted source code within the context of the existing source code, such as Figure 1A predicted source code changes at line 651 in depicted source code line 108.

[0126] After predicting the source code changes for the software engineering task, the machine learning model(s) optionally predict other source code changes for the software engineering task based on a problem report received from a problem tracker associated with the software developer that describes the software engineering task and / or a modification to the source code of the software engineering task (block 620). The system responds to the problem report for the software engineering task and / or the modification to the source code by making new predictions for the software engineering task. In embodiments, this can include the machine learning model(s) 242 responding to a problem report for the software engineering task by Sofia modifying (to fix the subject check error response and to remove the variable name when the variable name is not embedded) the software engineering task (such modification by Sofia to specify that she made the source code changes to remove the variable name when the variable name is not embedded) and using such modification as contextual data to predict other source code changes for another completion of the software engineering task (but only to fix the subject check error response without removing the embedded variable name).

[0127] After optionally predicting other source code changes for the software engineering task, the system optionally outputs the predicted other source code changes for the software engineering task to a problem tracker associated with the software developer (block 622). The system responds to the problem report for the software engineering task and / or the modification to the source code by outputting new predictions for the software engineering task. As an example and not by way of limitation, this can include the agent 244 outputting other source code changes for another completion of the software engineering task but only to fix the subject check error response without removing the embedded variable name.

[0128] After outputting all predicted source code changes for the software engineering task, the system commits the source code changes to the source code associated with the software engineering task based on the predicted source code changes accepted by the problem tracker (block 624). The system commits the predicted source code accepted by the software developer. In embodiments, this can include the agent 244 committing the predicted source code changes for removing the embedded variable name accepted by Sofia to the source code control repository for the software engineering task and closing the software engineering task on behalf of Sofia. Even if Sofia only worked on the problem report describing her software engineering task without generating any source code changes for the software engineering task and then requested predicted source code changes for the software engineering task, the machine learning model(s) 242 can use the updated problem report to predict each source code change needed to complete the software engineering task, thereby automating the generation of source code changes.

[0129] Although FIGURE 6 depicts the blocks 602-624 occurring in a particular order, the blocks 602-624 can occur in other orders. In other embodiments, each of the blocks 602-624 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0130] Software developers can tend to monitor their automation tools more closely. Additional instantiations can reside within a software developer's code editor 212, which is a word processor-like software that the software developer uses to edit source code. Such a tool can be implemented within the code editor 212 as a plug-in that extends the native functionality of the code editor 212. Such a code editor-focused tool will focus on helping the software developer write the most appropriate source code to accomplish a given task.

[0131] Figure 7 FIGURE 7 is a flowchart that illustrates a method that assists in writing code for a software engineering task, in an embodiment. The flowchart 700 illustrates method actions that are illustrated as flowchart blocks that are used to Figure 2 certain steps involved with and / or between the clients 202-204 and / or servers 206-208.

[0132] After a system, such as a bug tracking system, assigns a software engineering task to a software developer, the software developer logs into a bug tracker 216 on the client 202 and begins reading the task description. The system optionally receives a request from a code editor associated with the software developer to begin working on a location in source code associated with the software engineering task (block 702). The system stores the source code and bug report for the software engineering task. As an example and not by way of limitation, this can include the agent 244 receiving a request from the code editor 212 on the client 202 of the software developer named Sofia who decides that she wants to begin working on the software engineering task to fix the subject verification error response and remove the variable name when it is not embedded, at a location in the source code of the software. The location can be a region.

[0133] After optionally receiving a request from a code editor of a software developer to begin working on a location in source code associated with a software engineering task, the system optionally outputs the source code at the location in the source code associated with the software engineering task to the code editor associated with the software developer (block 704). The system provides the source code and bug report to the software developer. As an example and not by way of limitation, this can include the agent 244 outputting a section of the source code for a location in the source code that begins at line 701 where Sofia has positioned a cursor in her code editor 212.

[0134] After optionally outputting the source code of the software engineering task to the code editor of the software developer, the system optionally enables the software developer to articulate a problem report describing the software engineering task via the issue tracker associated with the stakeholder of the software engineering task, and / or update the problem report to describe a strategy for completing the software engineering task (block 706). The system implements the update to the problem report describing the software engineering task. In an embodiment, this can include the agent 244 enabling Sofia to communicate with the issue tracker 216 on Stacy Holder's client 204, Stacy Holder being a stakeholder of the software engineering task to fix the subject check error response and remove the variable name when it is not embedded. Sofia and Stacy agree on a high-priority (but not the highest-priority) software engineering task to fix the subject check error response and remove the variable name when it is not embedded.

[0135] After optionally articulating the problem report describing the software engineering task, the system stores the source code changes received from the code editor associated with the software developer at the location in the source code associated with the software engineering task (block 708). The system stores the source code of the software engineering task and any updates to the problem report. For example and without limitation, this can include the agent 244 storing Sofia's source code changes to line 651 in the source code, where Sofia has been working on the same software engineering task, the source code changes being a small part of the source code changes required to fix the subject check error response and remove the variable name when it is not embedded.

[0136] After storing the source code changes at the location in the source code, the system receives an implicit or explicit request from the code editor for a predicted source code change at the location in the source code associated with the software engineering task (block 710). The system is requested to predict the source code of the software engineering task. As an example and without limitation, this can include the agent 244 receiving an explicit request from the button (pressed on Sofia's code editor 212) for a predicted source code change of the software engineering task at line 701 in the source code where her cursor is currently located. The location in the source code can be the same or different from the source code location, such as after Sofia stores the source code changes at one location, she can stay at the same location or move to a new location (at which she requests the predicted source code change).

[0137] In addition to receiving explicit requests for source code changes at predicted source code locations, the agent 244 can also identify implicit requests for source code changes at predicted source code locations. For example, Sofia uses her code editor 212 and writes source code changes at various source code locations including line 651, positions her cursor at line 701 in the same source code, and then stops and waits - at which point the agent 244 notifies her that the continued lack of coding activity during the displayed timer countdown without interruption will cause the machine learning model(s) 242 to predict a source code change for the software engineering task at line 701 where her cursor is located. The source code location can be a region of a text listing of commands that will be compiled or assembled into an executable computer program.

[0138] After receiving the request to predict a source code change at a source code location associated with the software engineering task, the system retrieves context data that establishes a context for the software engineering task (block 712). The system retrieves the context data for the software engineering task to make accurate predictions for the software engineering task. In an embodiment, this can include the agent 244 retrieving context data that establishes a context for the software engineering task from the dynamic data pipeline 232, such as source code changes at various locations in the source code by Sofia from the dynamic data pipeline 240 that are part of a partial solution needed for the software engineering task to: fix the subject verification error response and remove the variable name when it is not embedded.

[0139] After retrieving the context data that establishes a context for the software engineering task, the system transforms the context data to be compatible with a data format used to train the machine learning model(s) to assist in performing the software engineering task (block 714). The system transforms the context data for the software engineering task to be in a format used to predict the software engineering task. For example and without limitation, this can include the agent 244 transforming the source code changes at various locations in the source code to be compatible with a data format used to train the machine learning model(s) 230 to predict the source code changes for the software engineering task.

[0140] After the system transforms the various types of context data for the software engineering task, the machine learning model(s) use the transformed context data to predict a source code change at a source code location associated with the software engineering task (block 716). The system uses the transformed data to predict the source code for the software engineering task. As an example and without limitation, this can include the machine learning model(s) 242 using the transformed context data to predict a source code change at line 701 in the source code for the software engineering task. Figure 1AThe depicted source code change at line 701 in source code line 108 to complete the localized software engineering task to fix the subject check error response and remove the variable name when it is not embedded. The source code change can include: a) source code to be added to the source code location, b) source code to be removed from the source code location, c) source code to be replaced in the source code location, and / or d) any type of source code change to another source code location. The type can be a category or classification result.

[0141] After predicting the source code change at the source code location, the system outputs the predicted source code change at the source code location to the code editor (block 718). The system outputs the predicted source code change at the source code location. In embodiments, this can include the agent 244 outputting the predicted source code change at the source code location to the code editor 212 of the software developer associated with the software engineering task. Figure 1B The depicted predicted source code change at line 704 in source code line 110 is output to Sofia’s code editor 212, which enables her to complete her localized software engineering task.

[0142] After predicting and outputting the source code change at the source code location associated with the software engineering task, the machine learning model(s) optionally predict other source code changes at another source code location associated with the software engineering task based on modifications to the predicted source code change and / or issue report at the source code location received from the code editor and / or issue tracker associated with the software developer (block 720). The system responds to the issue report and / or modifications to the source code for the software engineering task by making new predictions for the software engineering task. For example, and without limitation, this can include the machine learning model(s) 242 responding to Sofia’s modification of the issue report for the software engineering task (to fix the subject check error response and remove the variable name when it is not embedded) and using this modification as contextual data to predict new source code changes at various source code locations associated with the software engineering task (but only to fix the subject check error response without removing the embedded variable name).

[0143] After optionally predicting other source code changes at another source code location associated with the software engineering task, the system optionally outputs the predicted other source code changes at the other source code location of the software engineering task to the code editor (block 722). The system responds to the problem report for the software engineering task and / or the modification of the source code by outputting new predictions for the software engineering task. As an example and not by way of limitation, this can include the agent 244 outputting the slightly different source code changes for the localized software engineering task to Sofia, but only for the fix subject check error response without requiring removal of the embedded variable name.

[0144] After outputting the predicted source code changes, the system commits the source code changes based on any predicted source code changes at any source code location accepted by the code editor (block 724). The system commits the predicted source code accepted by the software developer. In an embodiment, this can include the agent 244 committing the predicted source code changes for removal of the embedded variable name accepted by Sofia to the source code control repository for the software engineering task and closing the software engineering task on behalf of Sofia.

[0145] Predicting source code changes, outputting the predicted source code changes, and committing the source code changes based on any predicted source code changes at the source code location can include predicting additional source code changes, outputting the additional source code changes, and committing the source code changes based on any additional source code changes at another source code location associated with the software engineering task. For example, in addition to predicting the source code changes at line 704 in the depicted source code line 110, Figure 1B In addition to predicting the source code changes at line 704 in the depicted source code line 110, the machine learning model(s) 242 also predict Figure 1A the source code changes at line 676 in the depicted source code line 112 to complete the localized software engineering task for the fix subject check error response.

[0146] After the source code change at the source code location is committed, the system has the following options: iteratively retrieve context data for source code changes at other source code locations, transform the context data for source code changes at other source code locations, predict source code changes at other source code locations, output predicted source code changes at other source code locations, and commit source code changes at other source code locations until the software developer completes the software engineering task (block 726). The system enables prediction of source code changes at other locations until the software developer completes the software engineering task. For example and without limitation, this can include the agent 244 repeatedly continuing the above process, which enables Sofia to turn to any other location in the source code that she wishes to focus on to complete each of her localized software engineering tasks. Then, the agent 244 commits the source code changes to the source code control repository for the software engineering task and closes the software engineering task on behalf of the software developer.

[0147] Although Figure 7 While blocks 702-726 are depicted as occurring in a particular order, blocks 702-726 can occur in other orders. In other embodiments, each of blocks 702-726 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0148] The agent 244 can help the software developer find anomalies in the source code immediately after the software developer writes the source code. These anomalies can be bugs or other difficulties (such as inefficiently written code), and finding these bugs and / or difficulties early on helps the software developer complete the software engineering task more efficiently. The wide-ranging version(s) of the machine learning model(s) 242 contain a broad range of knowledge such that, with enough data, almost no software engineering task or source code base will be truly unique. Thus, the small portion of the software developer’s source code that is different from what the machine learning model(s) 242 predict can be considered unique, surprising, or unanticipated, and thus worthy of further review by the software developer, who can find such anomalies to be errors or inefficiently written source code.

[0149] Figure 8 FIG. 8 is a flowchart illustrating one method that assists in identifying unanticipated portions of source code files for a software engineering task, in an embodiment. The flowchart 800 illustrates method actions, which are illustrated as flowchart blocks, that are used to Figure 2 certain steps involved with and / or between the clients 202-204 and / or servers 206-208 of FIG. 1.

[0150] After a system, such as an issue tracking system, assigns a software engineering task to a software developer, the software developer logs into an issue tracker 216 on a client 202 and begins reading the task description. The system then optionally receives a request from a code editor associated with the software developer to begin working on a source code file associated with the software engineering task (block 802). The system stores the source code and issue report for the software engineering task. For example and without limitation, this can include the agent 244 receiving the following request from the code editor 212 on the client 202 of the software developer named Sofia to begin writing source code for a source code file for the software engineering task to fix the subject check error response and remove variable names when not embedded. The source code file can be an object in a computer system for storing a text list of commands that will be compiled or assembled into an executable computer program.

[0151] After optionally receiving the request for the source code file, the system optionally outputs the source code file associated with the software engineering task to the code editor associated with the software developer (block 804). The system provides the source code and issue report to the software developer. As an example and without limitation, this can include the agent 244 sending the source code file to Sofia, who wants to work on her software engineering task to fix the subject check error response and remove variable names when not embedded.

[0152] After optionally outputting the source code file for the software engineering task to the code editor of the software developer, the system optionally enables the software developer to clarify the issue report describing the software engineering task via the issue tracker associated with a stakeholder of the software engineering task and / or update the issue report to describe a strategy for completing the software engineering task (block 806). The system enables updates to the issue report describing the software engineering task. In an embodiment, this can include the agent 244 enabling Sofia to communicate with the issue tracker 216 on the client 204 of Stacy Holder, who is a stakeholder of the software engineering task to fix the subject check error response and remove variable names when not embedded. Sofia and Stacy agree on a high priority (but not the highest priority) for the software engineering task to fix the subject check error response and remove variable names when not embedded.

[0153] After optionally clarifying the problem report describing the software engineering task, the system stores source code changes received from a code editor associated with the software developer in a source code file associated with the software engineering task (block 808). Any updates to the source code and problem report for the software engineering task are stored. For example and without limitation, this can include agent 244 storing Sofia's source code changes to the source code file for the software engineering task to fix the subject check error response and remove variable names when they are not embedded.

[0154] After storing the source code changes to the source code file, the system receives an implicit or explicit request from the code editor to predict the source code of the source code file (block 810). The system is requested to predict the source code of the software engineering task. By way of example and without limitation, this can include agent 244 receiving an explicit request from a button (pressed on Sofia's code editor 212) to predict the source code of the source code file for the software engineering task to fix the subject check error response and remove variable names when they are not embedded.

[0155] In addition to receiving an explicit request to predict the source code of the source code file, agent 244 can also identify an implicit request to predict the source code of the source code file. For example, Sofia is using her code editor 212 and writing source code changes to the source code file and then stops and waits - at this point agent 244 notifies her that the continued lack of editing activity during the displayed timer's uninterrupted countdown will result in the machine learning model(s) 242 predicting the source code of the source code file.

[0156] After receiving a request for a prediction of the source code of the source code file for the software engineering task, the system retrieves context data that establishes a context for the software engineering task (block 812). The system retrieves the context data for the software engineering task to accurately make predictions for the software engineering task. In an embodiment, this can include agent 244 retrieving from dynamic data pipeline 240 context data that establishes a context for the software engineering task, such as Sofia's source code changes in the source code file.

[0157] After retrieving the context data establishing the context of the software engineering task, the system transforms the context data to be compatible with a data format used to train the machine learning model(s) to assist in performing the software engineering task (block 814). The system transforms the context data of the software engineering task to be in a format used to predict the software engineering task. For example and without limitation, this can include the agent 244 transforming the retrieved source code changes of the source code file to be compatible with a data format used to train the machine learning model(s) 230 to learn to predict the source code of the source code file associated with the software engineering task.

[0158] After the system transforms the various types of context data of the software engineering task, the machine learning model(s) then use the transformed context data to predict the source code of the source code file associated with the software engineering task, where portions in the source code file correspond to portions in the predicted source code (block 816). The system uses the transformed data to predict the source code of the software engineering task. As an example and without limitation, this can include the machine learning model(s) 242 using the transformed context data to predict the source code of the source code file of the software engineering task to fix the subject check error response and remove variable names when they are not embedded. Portions in the predicted source code, such as the predicted “else: #pragma: no cover”, match portions of the source code file, such as the 702nd line of the source code file “else: #pragma: no cover”. A portion can be a part of a whole. The predicted source code can be a forecast of a text list of commands that will be compiled or assembled into an executable computer program.

[0159] The predicted source code can include: a) source code to be added to the source code file, b) source code to be removed from the source code file, c) source code to be replaced in the source code file, and / or d) any type of source code of any other source code file. A portion of the source code file can be a source code line, a source code word, and / or a single text character of source code. A line can be a horizontal line of text. A word can be a single unique meaningful text element. A single text character can be a printed or written symbol or letter.

[0160] After predicting the source code of the source code file, the system identifies each portion in the source code file that is determined to be different from a corresponding portion in the predicted source code via a code editor (block 818). The system identifies where the predicted source code is different from the source code file. In embodiments, this can include the agent 244 rendering the source code file with the predicted source code using boldface highlighting and underlining to indicate where the predicted source code is different from the source code file. Figure 1B a portion of the 701st line in the depicted source code line 114 because the portion of line 701 differs from its counterpart in the predicted source code (loc) that has a predicted confidence score of 95% above the 75% probability threshold.

[0161] Based on the determination of whether the counterpart in the predicted source code has a predicted confidence level that meets the threshold, each portion in the source code file that differs from the counterpart in the predicted source code is conditionally identified. For example, for each portion in the predicted source code that differs from its counterpart in the source code file and has a predicted confidence score below the 75% probability threshold, the agent 244 highlights the counterpart in the source code file with only underlining instead of with bold on the user interface in the code editor 212 of Sofia. For portions in the predicted source code that have no difference from their counterparts in the source code file, the agent 244 does not render anything different from what was previously rendered on the user interface of the code editor 212 of Sofia.

[0162] After identifying the portions in the source code file that differ from the portions in the predicted source code, the machine learning model(s) optionally predict other source code of the source code file based on any accepted modifications to the predicted source code and / or issue reports received from the code editor and / or issue tracker associated with the software developer, where the portions in the source code file correspond to the portions in the predicted other source code (block 820). The system responds to issue reports and / or modifications to the source code of the software engineering task by making new predictions for the software engineering task. For example and without limitation, this can include the machine learning model(s) 242 responding to an issue report of the software engineering task by Sofia modifying (to fix the subject check error response and remove the variable name when it is not embedded) and using this modification as contextual data to predict other source code of the source code file (but only to fix the subject check error response without removing the embedded variable name). An accepted portion can be a part of a whole that is approved for a certain purpose.

[0163] After optionally predicting other source code of the source code file, the system optionally outputs the predicted other source code of the source code file to the code editor (block 822). The system responds to the problem report for the software engineering task and / or the modification of the source code by outputting new predictions for the software engineering task. As an example and not by way of limitation, this can include the agent 244 outputting the slightly revised source code of the source code file of the software engineering task to Sofia, but only for the fix subject check error response without requiring removal of the embedded variable name.

[0164] After outputting the predicted source code of the source code file to the code editor, the system commits any differing portions of the predicted source code that the code editor requested and accepted into the source code file (block 824). The system commits the predicted source code that the software developer accepted into the source code file. In an embodiment, this can include the agent 244 responding to Sofia using her code editor 212 to review the bold highlighted portion of line 701 , selecting the bold highlighted portion of line 701 (which causes display of the portion of the predicted source code that differs from the bold highlighted portion of line 701), and selecting to accept the portion of the predicted source code to replace Figure 1B the bold highlighted portion of line 701 in the depicted source code line 114. The agent 244 commits the portion of the predicted source code that Sofia accepted to the source code control repository of the software engineering task and closes the software engineering task (for the fix subject check error response) on behalf of Sofia. If the source code changes of Sofia begin to fix the subject check error response, then many of the highlighted portions in the source code file can identify the predicted source code needed to complete her software engineering task, but if the source code changes of Sofia complete her task, then few of the highlighted portions in the source code file can identify the predicted source code (that corrects her possible errors and / or makes her source code more efficient). The differing portions can be part of the whole that can be distinguished from the corresponding portion.

[0165] Although Figure 8 blocks 802-824 are depicted as occurring in a particular order, blocks 802-824 can occur in other orders. In other embodiments, each of blocks 802-824 can also be performed in combination with other blocks, and / or some blocks can be divided into different groups of blocks.

[0166] Software engineering is a complex process that can be facilitated by tools. A typical system includes focused vertical tools that assist in partial aspects of the process to varying degrees. The system 200 subsumes existing technology by learning the entire software development process end-to-end in a general-purpose model. The relevant agent 244 and one or more relevant tool instantiations use the machine learning model(s) 242 and software developer-owned engineering data to assist and automate in the representation of highly general-purpose software development tasks. Embodiments of the system 200 assist and automate to varying degrees, with the most general-purpose being able to perform an entire software engineering task assigned to a software developer by reading a problem report (describing a software engineering task using natural language).

[0167] An example hardware device in which the subject matter can be implemented will be described. Those skilled in the art will appreciate that the elements illustrated in FIG. 9 can be varied from one implementation to another. Referring to FIG. 9, an example system for implementing the subject matter disclosed herein includes a hardware device 900 that includes a processing unit 902, a memory 904, a storage 906, a data input module 908, a display adapter 910, a communication interface 912, and a bus 914 coupling the elements 904-912 to the processing unit 902.

[0168] The bus 914 can include any type of bus architecture. Examples include a memory bus, a peripheral bus, a local bus, etc. The processing unit 902 refers to a machine, device, or apparatus that executes instructions, and can include a microprocessor, a digital signal processor (DSP), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc. The processing unit 902 can be configured to execute program instructions stored in the memory 904 and / or the storage 906 and / or received via the data input module 908.

[0169] The memory 904 can include read-only memory (ROM) 916 and random access memory (RAM) 918. The memory 904 can be configured to store program instructions and data during the operation of the device 900. In various embodiments, the memory 904 can include any of a variety of memory technologies, such as static random access memory (SRAM) or dynamic RAM (DRAM), including variants such as double data rate synchronous DRAM (DDR SDRAM), error-correcting code synchronous DRAM (ECC SDRAM), or RAMBUS DRAM (RDRAM), for example.

[0170] The memory 904 can also include nonvolatile memory technology such as flash memory, or NVRAM, among others. In some embodiments, it is contemplated that memories 904 can include a combination of technologies, such as the aforementioned technologies, as well as others not specifically mentioned here. When the subject matter is implemented in software, as is shown in the exemplary embodiment, the basic input / output system (BIOS) 920 being stored in ROM 916 contains the basic routines that help to transfer information between elements within the computer system, such as during start-up.

[0171] The storage device 906 can include a flash data storage device for reading from and writing to a flash memory, a hard disk drive for reading from and writing to a hard disk, a magnetic disk drive for reading from or writing to a removable magnetic disk, and / or an optical disk drive for reading from or writing to a removable optical disk (such as a CD ROM, DVD, or other optical media). The drives and their associated computer-readable media provide nonvolatile storage of computer-readable instructions, data structures, program modules and other data for the hardware environment 900.

[0172] It should be noted that the methods described herein can be embodied in executable instructions stored in computer-readable media for use by or in connection with an instruction execution machine, apparatus, or device, such as a computer-based or processor-containing machine, apparatus, or device. A skilled artisan will appreciate that for some embodiments, other types of computer- readable media that can store data that is accessible to a computer, such as magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, RAM, ROM, and the like, can also be used in the exemplary operating environment. As used herein, "computer-readable media" can include one or more of any suitable media for storing the executable instructions of a computer program in one or more of electronic, magnetic, optical, and electromagnetic formats, such that an instruction execution machine, system, apparatus or device can read (or access) the instructions from the computer-readable media and execute the instructions for performing the described methods. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; RAM; ROM; erasable programmable read-only memory (EPROM or flash memory); optical storage devices, including portable compact discs (CDs), portable digital video discs (DVDs), High Definition DVDs (HD-DVDs), Blu-ray discs; and the like. TM

[0173] ​A number of program modules can be stored on the storage device 906, ROM 916, or RAM 918, including an operating system 922, one or more application programs 924, program data 926, and other program modules 928. A user can enter commands and information into the hardware device 900 through a data input module 908. The data input module 908 can include mechanisms such as a keyboard, touchscreen, pointing device, etc.

[0174] Other external input devices (not shown) can be connected to the hardware device 900 via external data input interface 930. As examples and not by way of limitation, external input devices can include a microphone, joystick, game pad, satellite dish, scanner, or the like. In some embodiments, external input devices can include video or audio input devices such as a camera, still camera, etc. The data input module 908 can be configured to receive input from one or more users of the device 900 and deliver such input to the processing unit 902 and / or memory 904 via the bus 914.

[0175] A display 932 is also connected to the bus 914 via a display adapter 910. The display 932 can be configured to display output of the device 900 to one or more users. In some embodiments, a given device such as a touchscreen can function as both the data input module 908 and the display 932. An external display device can also be connected to the bus 914 via an external display interface 934. Other peripheral output devices (such as speakers and printers) can be connected to the hardware device 900, not shown.

[0176] The hardware device 900 can operate in a networked environment using logical connections to one or more remote nodes (not shown) via the communication interface 912. The remote nodes can be another computer, a server, a router, a peer device or other common network node, and typically include many or all of the elements described above relative to the hardware device 900. The communication interface 912 can interface to a wireless network and / or a wired network.

[0177] Examples of wireless networks include, for example, a BLUETOOTH network, a wireless personal area network, a wireless 802.11 local area network (LAN), and / or a wireless telephonic network (e.g., cellular, PCS or GSM network). Examples of wired networks include, for example, a LAN, a fiber optic network, a wired personal area network, a telephonic network, and / or a wide area network (WAN). Such networking environments are commonplace in intranets, the Internet, office networks, enterprise-wide computer networks, etc. In some embodiments, the communication interface 912 can include logic configured to support direct memory access (DMA) transfers between the memory 904 and other devices.

[0178] In a networking environment, program modules depicted relative to the hardware device 900, or portions thereof, can be stored in a remote storage device, such as for example on a server. It should be understood that other hardware and / or software can be used to establish a communication link between the hardware device 900 and other devices.

[0179] It should be understood that the arrangement of the hardware device 900 illustrated in FIG. 9 is but one possible implementation, and that other arrangements are possible. It should also be understood that the various system components (and means) described below and illustrated in the various diagrams, as defined by the claims, represent logical components configured to perform the functions described herein. For example, one or more of these system components (and means) can be implemented in whole or part by at least some of the components illustrated in the arrangement of the hardware device 900.

[0180] In addition, while at least one of these components is implemented at least partially as an electronic hardware component, and thus constitutes a machine, other components can be implemented in software, hardware, or a combination of software and hardware. More particularly, at least one component defined by a claim is implemented at least partially as an electronic hardware component, such as an instruction execution machine (e.g., a processor-based or processor-containing machine) and / or as a special purpose circuit or circuitry (e.g., discrete logic gates interconnected to perform a specialized function), such as those illustrated in FIG. 9.

[0181] Other components can be implemented in software, hardware, or a combination of software and hardware. Moreover, some or all of these other components can be combined, some can be omitted altogether, and additional components can be added, while still achieving the functionality described herein. Thus, the subject matter described herein can be embodied in a multitude of different variations and all such are contemplated within the claimed scope.

[0182] In the description above, subject matter is described in reference to symbolic representations of actions and operations of one or more devices that are performed by the device(s) executing instructions. Thus, it will be understood that such actions and operations, while sometimes referred to as being performed by the computer, include the manipulation by the processing unit of the computer of electrical signals representing data in a structured form. This manipulation transforms the data or maintains it at locations in the memory system of the computer, which reconfigures or otherwise alters the operation of the device in a manner well understood by those skilled in the art. The data structures where data is maintained are physical locations of the memory that have particular properties defined by the format of the data. However, while the subject matter is being described in the foregoing context, it is not meant to be limiting as those of skill will understand the claims and descriptions are intended to include both hardware and software based mechanisms.

[0183] To facilitate an understanding of the subject matter described herein, a number of aspects are described in the sequence of acts. At least one of these aspects is performed by an electronic hardware component. For example, it will be recognized that various actions can be performed by specific circuits or circuit systems, by program instructions being executed by one or more processors, or by a combination of both. The description herein of any sequence of acts is not intended to imply that the specific order described for performing that sequence is the only order in which the sequence can be performed. Unless otherwise indicated herein, or unless clearly contradicted by context, all methods described herein can be performed in any suitable order. The claims should not be construed to cover only those embodiments which can be practically carried out.

[0184] While one or more embodiments have been described by way of example and in terms of specific embodiments, it is to be understood that one or more embodiments are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the embodiments. The scope of the appended claims should therefore be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures.

Claims

1. A system for training machine learning models to assist in performing software engineering tasks, the system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: Retrieve data from multiple data sources associated with software engineering tasks; The data is linked by linking each problem report describing any one of the software engineering tasks to the source code associated with any one of the software engineering tasks; Transform the data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; as well as The machine learning model is trained using the transformed data to assist in the execution of software engineering tasks by making predictions about source code changes associated with those tasks.

2. The system of claim 1, wherein the plurality of instructions further cause the one or more processors to iteratively perform the following operations: retrieving additional data from the plurality of data sources, linking the retrieved additional data, transforming the additional data, and training the machine learning model using the transformed additional data.

3. The system of claim 2, wherein iteratively training the machine learning model using the transformed additional data comprises one of the following: initializing the machine learning model and then training the initialized machine learning model using only the transformed additional data; training the machine learning model using both the transformed additional data and the previously transformed data; or incrementally training the machine learning model using the transformed additional data.

4. The system of claim 1, wherein the data is retrieved from a plurality of data sources including open-source software projects.

5. The system of claim 1, wherein the data is retrieved from a plurality of data sources, the plurality of data sources including data sources associated with a plurality of software engineering projects, the plurality of software engineering projects being associated with a single enterprise.

6. The system of claim 1, wherein the machine learning model comprises one of the group consisting of: The machine learning model trained using data associated with multiple software engineering projects related to a single enterprise, and the machine learning model trained using data associated with general software engineering knowledge, or the machine learning model trained using both data associated with multiple software engineering projects related to a single enterprise and data associated with general software engineering knowledge.

7. The system of claim 6, wherein training the machine learning model comprises: The training is performed using data associated with general software engineering knowledge, and then using data associated with the individual enterprise, with the data associated with the individual enterprise having a lower weight than the data associated with general software engineering knowledge.

8. A computer-based method for training a machine learning model to assist in performing software engineering tasks, the method comprising: Retrieve data from multiple data sources associated with software engineering tasks; The data is linked by linking each problem report describing any one of the software engineering tasks to the source code associated with any one of the software engineering tasks; Transform the data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; as well as The machine learning model is trained using the transformed data to assist in performing software engineering tasks by making predictions about source code changes associated with those tasks.

9. The method of claim 8, wherein the computer-implemented method further comprises iteratively performing the following operations: retrieving additional data from the plurality of data sources, linking the additional data, transforming the additional data, and training the machine learning model using the transformed additional data.

10. The method of claim 9, wherein iteratively training the machine learning model using the transformed supplementary data comprises one of: initializing the machine learning model and then training the initialized machine learning model using only the transformed supplementary data; training the machine learning model using both the transformed supplementary data and the previously transformed data; or incrementally training the machine learning model using the transformed supplementary data.

11. The method of claim 8, wherein the data is retrieved from a plurality of data sources including open-source software projects.

12. The method of claim 8, wherein the data is retrieved from a plurality of data sources, the plurality of data sources including data sources associated with a plurality of software engineering projects, the plurality of software engineering projects being associated with a single enterprise.

13. The method of claim 8, wherein the machine learning model comprises one of the group consisting of: The machine learning model trained using data associated with multiple software engineering projects related to a single enterprise, and the machine learning model trained using data associated with general software engineering knowledge, or the machine learning model trained using both data associated with multiple software engineering projects related to a single enterprise and data associated with general software engineering knowledge.

14. The method of claim 13, wherein training the machine learning model comprises: The training is performed using data associated with general software engineering knowledge, and then using data associated with the individual enterprise, which has a lower weight than the data associated with general software engineering knowledge.

15. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for performing the following operations: Retrieve data from multiple data sources associated with software engineering tasks; The data is linked by linking each problem report describing any one of the software engineering tasks to the source code associated with any one of the software engineering tasks; Transform the data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; as well as The machine learning model is trained using the transformed data to assist in performing software engineering tasks by making predictions about source code changes associated with those tasks.

16. The computer program product of claim 15, wherein the program code further includes instructions for iteratively performing the following operations: retrieving additional data from the plurality of data sources, linking the additional data, transforming the additional data, and training the machine learning model using the transformed additional data.

17. The computer program product of claim 16, wherein iteratively training the machine learning model using the transformed additional data comprises one of the following: initializing the machine learning model and then training the initialized machine learning model using only the transformed additional data; training the machine learning model using both the transformed additional data and the previously transformed data; or incrementally training the machine learning model using the transformed additional data.

18. The computer program product of claim 15, wherein the data is retrieved from a plurality of data sources including open-source software projects.

19. The computer program product of claim 15, wherein the data is retrieved from a plurality of data sources, the plurality of data sources including data sources associated with a plurality of software engineering projects associated with a single enterprise.

20. The computer program product of claim 15, wherein the machine learning model comprises one of the group consisting of: The machine learning model trained using data associated with multiple software engineering projects related to a single enterprise and the machine learning model trained using data associated with general software engineering knowledge, or the machine learning model trained using both data associated with multiple software engineering projects related to a single enterprise and data associated with general software engineering knowledge, wherein training the machine learning model includes: training with data associated with general software engineering knowledge and then training with data associated with the single enterprise, wherein the weight of the data associated with the single enterprise is lower than that of the data associated with general software engineering knowledge.

21. A system using a machine learning model, said machine learning model assisting in performing software engineering tasks, said system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: In response to receiving a request from either an issue tracker or a code editor to begin processing work related to a software engineering task, at least one of an issue report describing the software engineering task or source code associated with the software engineering task is output to at least one of the issue tracker and the code editor associated with the software developer. The system stores at least one of the following: an update to a problem report describing the software engineering task received from at least one of the problem tracker or the code editor; or at least one of the source code changes associated with the software engineering task. In response to receiving either an implicit or explicit request from either the issue tracker or the code editor for a predicted completion scheme for the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict any number of completion schemes for the software engineering task. as well as By outputting a predicted completion scheme for the software engineering task to at least one of the problem's tracker and the code editor, the software developer is able to complete the software engineering task.

22. A computer-implemented method using a machine learning model, said machine learning model assisting in performing software engineering tasks, the method comprising: In response to receiving a request from either the issue tracker or the code editor to begin processing work related to the software engineering task, at least one of an issue report describing the software engineering task or source code associated with the software engineering task is output to at least one of the issue tracker and the code editor associated with the software developer. Store at least one of the following: an update to a problem report describing the software engineering task received from at least one of the problem tracker or the code editor; or at least one of the source code changes associated with the software engineering task. In response to receiving either an implicit or explicit request from either the issue tracker or the code editor for a predicted completion scheme for the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict any number of completion schemes for the software engineering task. as well as The software developer is able to complete the software engineering task by outputting the predicted completion scheme of the software engineering task to at least one of the issue tracker and the code editor.

23. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for: In response to receiving a request from either an issue tracker or a code editor to begin processing work related to the software engineering task, at least one of an issue report describing the software engineering task or source code associated with the software engineering task is output to at least one of the issue tracker and the code editor associated with the software developer. Store at least one of the following: an update to a problem report describing the software engineering task received from at least one of the problem tracker or the code editor; or at least one of the source code changes associated with the software engineering task. In response to receiving either an implicit or explicit request from either the issue tracker or the code editor for a predicted completion scheme for the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict any number of completion schemes for the software engineering task. as well as The software developer is able to complete the software engineering task by outputting the predicted completion scheme of the software engineering task to at least one of the issue tracker and the code editor.

24. A system for assisting in writing problem reports, the problem reports describing software engineering tasks and problem information, the system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: In response to receiving a request from an issue tracker associated with a stakeholder of a software engineering task to begin processing an incomplete issue report, the software engineering task is assigned to a software developer, the incomplete issue report describing the software engineering task and issue information, and the incomplete issue report is intended for use by the software developer; In response to receiving either an implicit or explicit request from the issue tracker for a predicted completion scheme for an incomplete issue report describing the software engineering task and issue information, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses transformed contextual data to predict a completion scheme for the incomplete problem report, which describes software engineering tasks and problem information; and Based on the predicted completion scheme of the incomplete issue report describing software engineering tasks and issue information, the software developer is able to complete the software engineering task by outputting the accepted completion scheme of the incomplete issue report to the issue tracker associated with the software developer.

25. A computer-implemented method for assisting in the writing of problem reports, the problem reports describing software engineering tasks and problem information, the method comprising: In response to receiving a request from an issue tracker associated with a stakeholder of a software engineering task to begin processing an incomplete issue report, the software engineering task is assigned to a software developer, the incomplete issue report describing the software engineering task and issue information, and the incomplete issue report is intended for use by the software developer; In response to receiving either an implicit or explicit request from the issue tracker for a predicted completion scheme for an incomplete issue report describing the software engineering task and issue information, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses transformed contextual data to predict a completion scheme for the incomplete problem report, which describes software engineering tasks and problem information; and Based on the predicted completion scheme of the incomplete issue report describing the software engineering task and issue information, the software developer is able to complete the software engineering task by outputting the accepted completion scheme of the incomplete issue report to the issue tracker associated with the software developer.

26. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for: In response to receiving a request from an issue tracker associated with a stakeholder of the software engineering task to begin processing an incomplete issue report, the software engineering task is assigned to a software developer, the incomplete issue report describing the software engineering task and issue information, and the incomplete issue report is intended for use by the software developer; In response to receiving either an implicit or explicit request from the issue tracker for a predicted completion scheme for an incomplete issue report describing the software engineering task and issue information, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses transformed contextual data to predict a completion scheme for the incomplete problem report, which describes the software engineering task and problem information; and Based on the predicted completion scheme of the incomplete issue report describing the software engineering task and issue information, the software developer is able to complete the software engineering task by outputting the accepted completion scheme of the incomplete issue report to the issue tracker associated with the software developer.

27. A system for assisting in the automation of software engineering tasks, the system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: In response to receiving a request for reviewing an issue report from the issue tracker, the issue report describing the software engineering task is output to the issue tracker associated with the software developer; In response to receiving either an implicit or explicit request from the issue tracker for predicting source code changes to the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict source code changes for software engineering tasks. The predicted source code changes for software engineering tasks are output to the issue tracker; as well as Based on any predicted source code changes accepted by the issue tracker, the source code changes are submitted to the source code associated with the software engineering task.

28. A method for assisting in the computer-based implementation of automated software engineering tasks, the method comprising: In response to receiving a request for reviewing an issue report from the issue tracker, the issue report describing the software engineering task is output to the issue tracker associated with the software developer; In response to receiving either an implicit or explicit request from the issue tracker for predicting source code changes to the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict source code changes for the software engineering task. The predicted source code changes for the software engineering task are output to the issue tracker; as well as Based on any predicted source code changes accepted by the issue tracker, the source code changes are submitted to the source code associated with the software engineering task.

29. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for: In response to receiving a request for reviewing an issue report from the issue tracker, the issue report describing the software engineering task is output to the issue tracker associated with the software developer; In response to receiving either an implicit or explicit request from the issue tracker for predicting source code changes to the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed contextual data to predict source code changes for the software engineering task. The predicted source code changes for the software engineering task are output to the issue tracker; as well as Based on any predicted source code changes accepted by the issue tracker, the source code changes are submitted to the source code associated with the software engineering task.

30. A system for assisting in writing source code for software engineering tasks, the system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: Source code changes received from the code editor associated with the software developer are stored at the location in the source code associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor for a predicted source code change at a source code location associated with the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed context data to predict source code changes at the source code location; Output the predicted source code changes at the specified source code location to the code editor; as well as Submit source code changes based on any predicted source code changes at any source code location accepted by the code editor.

31. A computer-implemented method for assisting in writing source code for software engineering tasks, the method comprising: Source code changes received from the code editor associated with the software developer are stored at the location in the source code associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor for a predicted source code change at a source code location associated with the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed context data to predict source code changes at the source code location; Output the predicted source code changes at the specified source code location to the code editor; as well as Submit source code changes based on any predicted source code changes at any source code location accepted by the code editor.

32. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for: Source code changes received from the code editor associated with the software developer are stored in the location within the source code associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor for a predicted source code change at a source code location associated with the software engineering task, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing the software engineering tasks; The machine learning model uses the transformed context data to predict source code changes at the source code location; Output the predicted source code changes at the specified source code location to the code editor; as well as Submit source code changes based on any predicted source code changes at any source code location accepted by the code editor.

33. A system for assisting in identifying unforeseen portions of source code files for software engineering tasks, the system comprising: One or more processors; as well as A non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause the one or more processors to: Source code changes received from the code editor associated with the software developer are stored in the source code file associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor to predict the source code of the source code file, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing the software engineering tasks; The machine learning model uses the transformed context data to predict the source code of the source code file, wherein portions of the source code file correspond to portions of the predicted source code. Using the code editor, each different part is identified by comparing the corresponding part in the source code file with the predicted source code. as well as Submit any discrepancies in the predicted source code requested and accepted by the code editor to the source code file.

34. A computer-implemented method for assisting in identifying unforeseen portions of source code files for software engineering tasks, the method comprising: Source code changes received from the code editor associated with the software developer are stored in source code files associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor to predict the source code of the source code file, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed context data to predict the source code of the source code file, wherein portions of the source code file correspond to portions of the predicted source code. Using the code editor, each different part is identified by comparing the corresponding part in the source code file with the predicted source code. as well as Submit any discrepancies in the predicted source code requested and accepted by the code editor to the source code file.

35. A computer program product comprising a non-transitory computer-readable medium having computer-readable program code embodied therein, executable by one or more processors, the program code including instructions for: Source code changes received from the code editor associated with the software developer are stored in the source code file associated with the software engineering task; In response to receiving either an implicit or explicit request from the code editor to predict the source code of the source code file, context data for establishing a context for the software engineering task is retrieved. Transform the context data to make it compatible with the data format used to train machine learning models to assist in performing software engineering tasks; The machine learning model uses the transformed context data to predict the source code of the source code file, wherein portions of the source code file correspond to portions of the predicted source code. Using the code editor, each different part is identified by comparing the corresponding part in the source code file with the predicted source code. as well as Submit any discrepancies in the predicted source code requested and accepted by the code editor to the source code file.